Image processing method and device

By receiving image processing tasks in the image processing method, inputting multiple target images into the image processing model, and determining the detection results based on local and global image features, the problem of difficult to realize automatic detection of lymph node metastasis in CT images in the prior art is solved, and the accuracy and efficiency of detection are improved.

CN116797554BActive Publication Date: 2025-05-09ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310628573.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2025-05-09
Estimated Expiration
2043-05-31

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate automatic detection of whether a single lymph node in CT images is metastasized, resulting in the inability to obtain less costly and more accurate training data.

Method used

By providing an image processing method, receiving an image processing task, inputting a plurality of target images into the image processing model, determining the detection result based on local and global image features, and realizing automatic detection of whether there are abnormal objects to be detected in the target to be detected partition.

Benefits of technology

The accuracy and detection efficiency of detection results of the subjects to be detected are improved, and accurate automatic detection of lymph node metastasis in CT images can be achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116797554B_ABST
    Figure CN116797554B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide an image processing method and device, wherein the method includes: receiving an image processing task, wherein the image processing task carries multiple target images corresponding to a target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected; inputting multiple target images into an image processing model to obtain a detection result corresponding to the target partition to be detected, wherein the image processing model determines the detection result based on the local image features and global image features of each target image. By inputting multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, and automatic detection of whether there is an abnormal object to be detected in the target partition to be detected can be achieved. The image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy and efficiency of the detection result for the object to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present specification relate to the field of computer technology, and in particular to an image processing method. Background Art

[0002] Lymph node metastasis is a common type of cancer, and usually requires professional doctors to determine the test results from the pathology report based on their experience.

[0003] However, at present, it takes a lot of effort for professionals to determine whether a single lymph node in a CT image has metastasized based on the pathology report, which is very difficult to achieve clinically. As a result, it is impossible to accurately label single lymph nodes in a large number of CT images based on the pathology report, and it is impossible to obtain low-cost and more accurate training data. Therefore, it is difficult to achieve automatic detection of whether a single lymph node has metastasized. Summary of the invention

[0004] In view of this, an embodiment of the present specification provides an image processing method. One or more embodiments of the present specification also relate to an image processing apparatus, a computing device, a computer-readable storage medium and a computer program to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, there is provided an image processing method, including:

[0006] receiving an image processing task, wherein the image processing task carries a plurality of target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected;

[0007] Multiple target images are input into an image processing model to obtain detection results corresponding to the target partitions to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image.

[0008] According to a second aspect of the embodiments of this specification, a CT image processing method is provided, comprising:

[0009] receiving a CT image processing task, wherein the CT image processing task carries a plurality of CT images corresponding to the target partition to be detected, and the CT image processing task is used to detect whether there are abnormal cells to be detected in the target partition to be detected;

[0010] Multiple CT images are input into a CT image processing model to obtain detection results corresponding to the target partition to be detected, wherein the CT image processing model determines the detection results based on local image features and global image features of each CT image.

[0011] According to a third aspect of an embodiment of this specification, a training method for an image processing model is provided, which is applied to a cloud-side device, including:

[0012] Acquire training sample pairs and guidance information of the training sample pairs, wherein the training sample pairs include training samples and sample labels, the training samples include multiple sample images corresponding to the target partition to be detected, and the sample labels are used to identify sample results of the target partition to be detected;

[0013] Input the training sample pairs and the guidance information into the initial image processing model to obtain the sample prediction results output by the initial image processing model and the target attention information of the target partition to be detected;

[0014] According to the sample prediction results and the sample results, a first loss value of the initial image processing model is calculated;

[0015] According to the target attention information and the guidance information, a second loss value of the initial image processing model is calculated;

[0016] Adjusting the model parameters of the initial image processing model according to the first loss value and the second loss value until a preset stop condition is reached, thereby obtaining the model parameters of the image processing model;

[0017] The model parameters of the image processing model are sent to the end-side device.

[0018] According to a fourth aspect of the embodiments of this specification, there is provided an image processing method, including:

[0019] Receive an image processing request sent by a user, wherein the image processing request includes an image processing task, the image processing task carries multiple target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected;

[0020] Inputting multiple target images into an image processing model to obtain detection results corresponding to the target partitions to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image;

[0021] Send the detection result corresponding to the target partition to be detected to the user.

[0022] According to a fifth aspect of the embodiments of this specification, there is provided an image processing device, including:

[0023] A receiving module is configured to receive an image processing task, wherein the image processing task carries a plurality of target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected;

[0024] The detection module is configured to input multiple target images into an image processing model to obtain detection results corresponding to the target partitions to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image.

[0025] According to a sixth aspect of an embodiment of this specification, a computing device is provided, including:

[0026] Memory and processor;

[0027] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.

[0028] According to a seventh aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and the computer-executable instructions can implement the steps of the above method when executed by a processor.

[0029] According to an eighth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the steps of the above method are implemented when the computer is executed.

[0030] An image processing method provided by an embodiment of the present specification receives an image processing task, wherein the image processing task carries multiple target images corresponding to a target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected; the multiple target images are input into an image processing model to obtain detection results corresponding to the target partition to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image.

[0031] In this way, by carrying multiple target images corresponding to the target partition to be detected in the image processing task and inputting the multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, thereby realizing automatic detection of whether there are abnormal objects to be detected in the target partition to be detected, and the image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy and efficiency of the detection results for the objects to be detected. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is an architecture diagram of an image processing system provided by an embodiment of this specification;

[0033] Figure 2 is a flow chart of an image processing method provided by an embodiment of this specification;

[0034] Figure 3 is a structural diagram of an image processing model provided by an embodiment of this specification;

[0035] Figure 4 is a structural diagram of an image processing model provided by an embodiment of this specification;

[0036] Figure 5 is a structural diagram of an image processing model provided by an embodiment of this specification;

[0037] Figure 6 is a structural diagram of an image processing model provided by an embodiment of this specification;

[0038] Figure 7 is a structural diagram of an image processing model provided by an embodiment of this specification;

[0039] Figure 8 is a structural diagram of an image processing model provided by an embodiment of this specification;

[0040] Fig. 9 is a structural diagram of an image processing model provided by an embodiment of this specification;

[0041] Fig.10 is a flow chart of a CT image processing method provided by an embodiment of this specification;

[0042] Fig.11 is a flowchart of a training method for an image processing model provided by an embodiment of this specification;

[0043] Fig.12 is a flow chart of an image processing method provided by an embodiment of this specification;

[0044] Fig.13 is a processing flow chart of an image processing method provided by an embodiment of this specification;

[0045] Fig.14 is a training diagram of an image processing model provided by an embodiment of this specification;

[0046] Fig.15 is a flow chart of an esophageal cancer detection method provided by an embodiment of this specification;

[0047] Fig.16 is a structural schematic diagram of an image processing device provided by an embodiment of this specification;

[0048] Fig.17 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0049] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0050] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0051] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0052] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0053] First, the terms involved in one or more embodiments of this specification are explained.

[0054] CT (Computed Tomography): Computer tomography uses a precisely collimated X-ray beam and a highly sensitive detector to perform cross-sectional scans around a certain part of the human body one by one. It has the characteristics of fast scanning time and clear images, and can be used to detect a variety of diseases.

[0055] EC (esophageal cancer): Esophageal cancer is a highly lethal cancer. According to incomplete statistics, the 5-year survival rate is low. However, if resectable / curable esophageal cancer is discovered early, the mortality rate will be greatly reduced. Among them, lymph node metastasis is a relatively common and typical disease.

[0056] LN (lymphnode): Lymph node. The main functions of lymph nodes are: filtering lymph, removing bacteria and foreign matter, producing lymphocytes and antibodies, etc. The lymphatic vessels of various parts or organs of the human body generally converge to nearby local lymph nodes. When a part of the body or an organ is diseased or inflamed, foreign matter such as bacteria and toxins can spread to the corresponding nearby lymph nodes through the lymphatic vessels. The local lymph nodes have the function of intercepting and removing these foreign matter such as bacteria or toxins, and become a defense barrier to prevent the spread and diffusion of lesions. If the local lymph nodes cannot intercept and remove these bacteria or toxins, the lesions can also spread and diffuse to distant places along the output tubes of the local lymph nodes.

[0057] LNS (lymphnode station): Lymph node station. In a human organ or tissue structure with a lymphatic system, the organ or tissue structure is divided into regions according to actual needs, and each partition containing lymph nodes in the organ or tissue structure is called a lymph node station. For example, by dividing the mediastinal tissue structure into regions, 9 mediastinal lymph nodes can be obtained. Different lymph nodes can have different partition numbers. For example, the 9 mediastinal lymph nodes can be numbered as: S1 (left+right), S2 (left+right), S3anterior, S3posterior, S4 (left+right), S5, S6, S7, S8.

[0058] RECIST (response evaluation criteria in solid tumor) is a standard for evaluating the clinical efficacy of solid tumors. It determines the size and number of measurable lesions at the baseline level, standardizes the measurement method, and evaluates the efficacy by the changes in target lesions during treatment and follow-up.

[0059] CAD (computer aided diagnosis): Computer-aided diagnosis refers to the use of imaging, medical image processing technology and other possible physiological and biochemical means, combined with computer analysis and calculation, to assist in the discovery of lesions and improve the accuracy of diagnosis.

[0060] Deep-Stationing model: a model that can accurately segment lymph nodes.

[0061] 2.5D Mask R-CNN framework (2.5-dimensional deep learning detection framework): is a two-stage framework. The first stage scans the image and generates a region that is likely to contain an object. The second stage classifies the region and generates a bounding box and mask. Mask R-CNN is a popular instance segmentation framework.

[0062] AdamW optimizer: A model training optimizer in PyTorch that is used to gradient update the parameters of the neural network to minimize the loss function.

[0063] PyTorch: is an environment that provides a running environment for building and training models.

[0064] Predicting lymph node (LN) metastasis in computed tomography (CT) is very important for the formulation of esophageal cancer staging treatment plans and prognosis. However, determining whether a single lymph node has metastasized based on CT images requires a lot of effort and cost for professional medical staff, and relying solely on experience to make judgments is prone to inaccurate positioning, resulting in immeasurable consequences. Currently, RECIST and morphological and texture features can be used to predict lymph node metastasis, but research reports show that the accuracy of the prediction results obtained by the above methods is still low and cannot be well applied to the prediction of whether lymph nodes have metastasized.

[0065] With the continuous development of computer technology, various learning models can gradually be applied to the prediction of results. As deep learning has achieved remarkable success in various medical imaging computer-aided diagnosis (CAD) tasks, it is currently proposed to apply deep learning models to the prediction of lymph node metastasis. However, deep learning models need to be trained with large-scale accurately labeled data to achieve accurate prediction results. For the prediction of lymph node metastasis, professional medical staff need to determine whether a single lymph node has metastasized from various medical images (such as CT images, magnetic resonance imaging, etc.) based on pathology reports, and label whether a single lymph node in various medical images is abnormal.

[0066] However, currently, from the perspective of surgical operation practice of lymph node dissection, the pathology report only shows the number of lymph nodes removed in each lymph node station and the number of metastatic lymph nodes found in the corresponding lymph node station (LNS). Therefore, even for experienced medical experts, it is extremely difficult to establish a one-to-one pairing between the lymph nodes observed in CT images and the individual metastatic lymph nodes indicated in the pathology report. As the data set expands, the classification of individual lymph node metastases using pathologically confirmed labels may become unscalable. Therefore, it is difficult for deep learning models to obtain large-scale, accurately labeled training data, resulting in the inability to output more accurate prediction results, and it is difficult to actually apply them to the prediction of lymph node metastasis.

[0067] Based on this, an embodiment of the present specification provides an image processing method, which receives an image processing task, wherein the image processing task carries multiple target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected; the multiple target images are input into an image processing model to obtain detection results corresponding to the target partition to be detected, wherein the image processing model determines the detection results based on the local image features and global image features of each target image.

[0068] In this way, by carrying multiple target images corresponding to the target partition to be detected in the image processing task and inputting the multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, thereby realizing automatic detection of whether there are abnormal objects to be detected in the target partition to be detected, and the image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy and efficiency of the detection results for the objects to be detected.

[0069] See also Figure 1 , Figure 1 An architecture diagram of an image processing system provided by an embodiment of the present specification is shown, and the image processing system may include a client 100 and a server 200;

[0070] The client 100 is used to send an image processing task to the server 200, wherein the image processing task carries multiple target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected.

[0071] The server 200 is used to receive an image processing task, input multiple target images into an image processing model, and obtain detection results corresponding to the target partitions to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image.

[0072] The client 100 is also used to receive the detection result of the image processing task sent by the server 200.

[0073] Apply the solution of the embodiments of this specification to receive an image processing task, wherein the image processing task carries multiple target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected; input the multiple target images into the image processing model to obtain the detection results corresponding to the target partition to be detected, wherein the image processing model determines the detection results based on the local image features and global image features of each target image.

[0074] In this way, by carrying multiple target images corresponding to the target partition to be detected in the image processing task and inputting the multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, thereby realizing automatic detection of whether there are abnormal objects to be detected in the target partition to be detected, and the image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy and efficiency of the detection results for the objects to be detected.

[0075] In practical applications, the image processing system may include multiple clients 100 and a server 200. Multiple clients 100 may establish communication connections through the server 200. In the image processing scenario, the server 200 is used to provide information extraction services between multiple clients 100. Multiple clients 100 may serve as senders or receivers, respectively, and realize communication through the server 200.

[0076] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the image processing scenario, the user can publish a data stream to the server 200 through the client 100, and the server 200 generates a detection result based on the data stream and pushes the detection result to other clients that have established communication.

[0077] The client 100 and the server 200 are connected via a network. The network provides a medium for a communication link between the client 100 and the server 200. The network may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, etc. before being released to the server 200.

[0078] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language Version 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The client 100 can be based on the software development kit (SDK, Software Development Kit) of the corresponding service provided by the server 200, such as based on the real-time communication (RTC, Real Time Communication) SDK development and acquisition. The client 100 can be deployed in an electronic device and needs to rely on the device to run or some APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer and other end-side devices. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0079] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers for background training that provide support for models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server (cloud-side device) for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0080] It is worth noting that the image processing method provided in the embodiments of this specification is generally executed by the server, but in other embodiments of this specification, the client may also have similar functions to the server, thereby executing the image processing method provided in the embodiments of this specification. In other embodiments, the image processing method provided in the embodiments of this specification may also be jointly executed by the client and the server.

[0081] In this specification, an image processing method is provided. This specification also relates to an image processing apparatus, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0082] See also Figure 2 , Figure 2 A flowchart of an image processing method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0083] Step 202: receiving an image processing task, wherein the image processing task carries a plurality of target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected.

[0084] In actual applications, the image processing tasks sent by the user can be received through the server or the client.

[0085] Specifically, the image processing task is a task for detecting whether there are abnormal objects to be detected in the target partition to be detected, and the image processing task carries multiple target images corresponding to the target partition to be detected. Furthermore, the target partition to be detected can be understood as a partition used to predict whether there are abnormal objects to be detected. The objects to be detected can be in normal or abnormal states, that is, when the human body is healthy, the objects to be detected should be in a normal state. When there are symptoms in the target partition to be detected, the objects to be detected may be in an abnormal state. By predicting whether there are abnormal objects to be detected in the target partition to be detected, the state of the objects to be detected can be further judged based on the prediction results, thereby providing assistance for the positioning and precise treatment of abnormal objects to be detected.

[0086] It should be noted that in one or more embodiments of the present specification, the image processing task can be applied to the recognition of various types of medical images, and judge whether there are abnormal objects to be detected in the target partition to be detected in the medical image based on the image features. For example, in the application scenario of treating fractures, it is possible to predict whether a fracture has occurred in the target partition to be detected based on the medical image of the target partition to be detected, thereby locating the fracture site. In the application scenario of lymph node metastasis, it is possible to predict whether lymph node metastasis has occurred in the target partition to be detected based on the medical image of the target partition to be detected, thereby helping doctors to accurately locate and treat abnormal lymph nodes.

[0087] In one or more embodiments of this specification, in order to improve the accuracy of the prediction results of the target partition to be detected, and to make the prediction and positioning of abnormal objects to be detected more accurate, before receiving the image processing task, a specific area to be detected can also be partitioned. Specifically, the area to be detected can be understood as an organ, tissue, or structure of the human body, such as the esophagus, aortic arch, ascending aorta, heart, spine, mediastinum, etc. The area to be detected can be partitioned according to actual conditions, and the area to be detected can be divided into multiple partitions to be detected, and one of the multiple partitions to be detected can be determined as the target partition to be detected. It should be noted that determining a target partition to be detected from multiple partitions to be detected can be randomly determined or specified according to the needs of actual applications, and this specification does not impose any restrictions on this.

[0088] For example, in the application scenario of predicting lymph node metastasis, assuming that the area to be detected is the mediastinal tissue of the human body, a CT image of the patient's entire mediastinal tissue can be collected, and one of the 9 lymph nodes divided into the mediastinal tissue is selected as the target partition to be detected. Multiple target images corresponding to the target partition to be detected are carried into the image processing task for input, thereby realizing the prediction of whether there are abnormal objects to be detected in the target partition to be detected.

[0089] In practical applications, the object to be detected may be a certain type of cell, a certain type of tissue structure, etc. For example, depending on the actual application scenario, it may be a fracture area, a single lymph node, a single cancer cell, etc.

[0090] By receiving the image processing task, multiple target images corresponding to the target to-be-detected partition carried in the image processing task can be used as input to detect whether there is an abnormal to-be-detected object in the target to-be-detected partition.

[0091] Step 204: Input multiple target images into the image processing model to obtain detection results corresponding to the target partitions to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image.

[0092] In practical applications, when an image processing task is received, multiple target images carried by the image processing task can be obtained, and the multiple target images can be input into the image processing model to obtain the detection results corresponding to the target partitions to be detected output by the image processing model.

[0093] Specifically, the image processing model can extract local image features and global image features of each target image based on multiple input target images, and determine the detection result of the target partition to be detected based on the local image features and global image features of each target image. Further, the detection result of the target partition to be detected can be understood as a prediction result of whether there is an abnormal object to be detected in the target partition to be detected. The detection result can be a probability value output by the image processing model. In practical applications, it can be specified as a value from 0 to 1. The closer the value output by the image processing model is to 0, the lower the probability of the existence of an abnormal object to be detected in the partition to be detected. The closer the value output by the image processing model is to 1, the higher the probability of the existence of an abnormal object to be detected in the partition to be detected.

[0094] In order to improve the task execution effect of the image processing model in performing the image processing task and improve the accuracy of the output detection result, in one or more embodiments of the present specification, the image processing model may include a feature extraction layer, a feature fusion layer and an output layer, wherein the feature fusion layer may be used to fuse the local image features and the global image features of each target image;

[0095] See also Figure 3 , Figure 3 A structural diagram of an image processing model provided by an embodiment of the present specification is shown. The model structure of the image processing model may include a feature extraction layer 302, a feature fusion layer 304 and an output layer 306. The feature extraction layer 302, the feature fusion layer 304 and the output layer 306 may be connected in sequence. A plurality of target images may be first input into the feature extraction layer 302 to obtain a result processed by the feature extraction layer 302. The feature extraction layer 302 then inputs the processed result into the feature fusion layer 304 to obtain a result processed by the feature fusion layer 304. The feature fusion layer 304 then inputs the processed result into the output layer 306 to obtain a detection result of the target partition to be detected that is finally output by the output layer.

[0096] Accordingly, inputting multiple target images into the image processing model to obtain the detection results corresponding to the target partitions to be detected may include the following steps S3002 to S3006:

[0097] S3002: Input multiple target images into the feature extraction layer to obtain image feature information corresponding to each target image.

[0098] In practical applications, multiple target images can be input into the feature extraction layer to obtain image feature information corresponding to each target image.

[0099] Specifically, the feature extraction layer can be understood as a structure for extracting image feature information in the target image. Further, the feature extraction layer can include one or more full convolution blocks for mapping the image information in each target image through convolution to obtain image feature information corresponding to each target image.

[0100] Optionally, the feature extraction layer may include a full convolution block, and multiple target images are input into the feature extraction layer to obtain image feature information corresponding to each target image, which may include the following steps:

[0101] A target image to be processed is determined from among multiple target images;

[0102] The target image to be processed is input into the full convolution block to obtain image feature information.

[0103] Specifically, the target image to be processed may be any one of the multiple target images. The multiple target images are input into the full convolution block of the feature extraction layer, and the image information can be mapped through the convolution structure in the full convolution block to extract the image feature information corresponding to the multiple target images.

[0104] Optionally, in order to further reduce the number of network parameters, enhance feature expressiveness, and obtain better model performance, the feature extraction layer may include n fully convolutional blocks connected in sequence, where n is a positive integer greater than or equal to 2.

[0105] Accordingly, inputting a plurality of target images into the feature extraction layer to obtain image feature information corresponding to each target image may further include the following steps:

[0106] determining a target image to be processed among a plurality of target images;

[0107] Input the target image to be processed into the first full convolution block to obtain the first image feature information;

[0108] Input the i-1th image feature information into the i-th full convolution block to obtain the i-th image feature information, where 2≤i≤n;

[0109] Determine whether i is equal to n;

[0110] If not, i is incremented by 1, and the operation of inputting the i-1th image feature information into the i-th full convolution block to obtain the i-th image feature information is continued;

[0111] If so, the i-th image feature information is determined as the image feature information corresponding to the target image to be processed.

[0112] See also Figure 4 , Figure 4 A structural diagram of an image processing model provided by an embodiment of this specification is shown.

[0113] Specifically, the feature extraction layer may include a plurality of full convolution blocks 400 connected in sequence, determine a target image to be processed from a plurality of target images, input the target image to be processed into the first full convolution block 400, and obtain the first image feature information output by the first full convolution block 400; input the first image feature information into the second full convolution block 400, and obtain the second image feature information output by the second full convolution block 400... and so on, until the n-1th image feature information is input into the nth full convolution block 400, and the nth image feature information output by the nth full convolution block 400 is obtained, and the nth image feature information is determined as the image feature information corresponding to the target image to be processed.

[0114] In practical applications, image feature information corresponding to multiple target images can be obtained according to the above method.

[0115] By inputting multiple target images into one or more full convolution blocks in the feature extraction layer, image feature information corresponding to the multiple target images can be obtained for subsequent feature processing to obtain a prediction of the detection results of the target partition to be detected.

[0116] In practical applications, image feature information corresponding to multiple target images is obtained, and each image feature information can be input into a feature fusion layer to obtain image fusion information corresponding to each target image.

[0117] S3004: Input feature information of each image into a feature fusion layer to obtain image fusion information corresponding to each target image.

[0118] Specifically, the feature fusion layer is used to fuse the local image features and global image features of each target image. By inputting each image feature information into the feature fusion layer, the local features and global features in each image feature information can be extracted respectively, so as to obtain local-global fusion information according to the extracted local features and global features.

[0119] Since the local feature information of the image extracted through the convolution layer is likely to lose information at a fine scale during the process of encoding the global feature information by the codec, in order to encode the feature information of the image without losing the local feature information, in one or more embodiments of the present specification, the feature information of each image is input into a feature fusion layer, the local feature information and the global feature information in each image feature information are extracted respectively, and the local feature information and the global feature information are fused to obtain the image fusion information corresponding to each target image, thereby improving the accuracy of the feature information and further improving the accuracy of the model output results.

[0120] Furthermore, in one or more embodiments of the present specification, the feature fusion layer may include one or more local-global fusion blocks to respectively extract and fuse local feature information and global feature information in the image feature information.

[0121] Optionally, the feature fusion layer may include a local global fusion block, and inputting the feature information of each image into the feature fusion layer to obtain the image fusion information corresponding to each target image may include the following steps:

[0122] Determine target image feature information from each image feature information;

[0123] The target image feature information is input into the local-global fusion block to obtain image fusion information.

[0124] Specifically, the target image feature information is any one of the image feature information. The target image feature information is input into the local-global fusion block, and the obtained image fusion information is more accurate.

[0125] Optionally, the feature fusion layer may include m local-global fusion blocks connected sequentially, where m is a positive integer greater than or equal to 2.

[0126] Accordingly, inputting each image feature information into the feature fusion layer to obtain the image fusion information corresponding to each target image may include the following steps:

[0127] Determine target image feature information from each image feature information;

[0128] Input the target image feature information into the first local-global fusion block to obtain the first image fusion information;

[0129] Input the j-1th image fusion information into the jth local-global fusion block to obtain the jth image fusion information, where 2≤j≤m;

[0130] Determine whether j is equal to m;

[0131] If not, j is incremented by 1, and the operation of inputting the j-1th image fusion information into the jth local-global fusion block to obtain the jth image fusion information is continued;

[0132] If so, the j-th image fusion information is determined as the image fusion information corresponding to the target image feature information.

[0133] See also Figure 5 , Figure 5 A structural diagram of an image processing model provided by an embodiment of this specification is shown;

[0134] Specifically, the feature fusion layer may include a plurality of local-global fusion blocks 500 connected in sequence, determine the target image feature information from each image feature information, input the target image feature information into the first local-global fusion block 500, and obtain the first image fusion information output by the first local-global fusion block 500; input the first image fusion information into the second local-global fusion block 500, and obtain the second image fusion information output by the second local-global fusion block 500... and so on, until the m-1th image fusion information is input into the mth local-global fusion block 500, and the mth image fusion information output by the mth local-global fusion block 500 is obtained, and the mth image fusion information is determined as the image fusion information corresponding to the target image feature information.

[0135] In practical applications, the image fusion information corresponding to each image feature information can be obtained according to the above method.

[0136] By inputting the feature information of each image into one or more local-global fusion blocks in the feature fusion layer, image fusion information corresponding to multiple target images can be obtained. The global features in the image feature information can be encoded while ensuring the accuracy of the local feature information. This can be used for subsequent feature processing to obtain a prediction of the detection results of the target partition to be detected, thereby improving the accuracy of the output results of the image processing model.

[0137] In one or more embodiments of the present specification, the local-global fusion block may include at least one local-global feature fusion layer, and the local-global feature fusion layer may include a local feature extraction sublayer and a global feature extraction sublayer and a fusion layer.

[0138] Specifically, the local-global feature fusion layer is used to extract local feature information and global feature information from each image feature information, and fuse the local feature information and the global feature information to obtain image fusion information corresponding to each target image. Further, the local feature extraction sublayer is used to extract local feature information corresponding to each target image, the global feature extraction sublayer is used to extract global feature information corresponding to each target image, and the fusion layer is used to fuse the extracted local feature information and global feature information to obtain image fusion information corresponding to each target image.

[0139] Optionally, a local-global fusion block may include a local-global feature fusion layer. In order to improve the accuracy of feature information and train a more accurate image processing model, a local-global fusion block may also include multiple local-global feature fusion layers connected in sequence.

[0140] Accordingly, for the t-th local-global feature fusion layer, the following steps may be included:

[0141] Acquire feature information of the image to be processed, wherein the feature information of the image to be processed includes feature information of the target image or fusion information of the t-1th image;

[0142] Inputting the feature information of the image to be processed into the local feature extraction sublayer to obtain the local feature information corresponding to the feature information of the image to be processed;

[0143] Inputting the feature information of the image to be processed into the global feature extraction sublayer to obtain the global feature information corresponding to the feature information of the image to be processed;

[0144] The local feature information and the global feature information are input into the fusion layer to obtain the image fusion information corresponding to the t-th local-global feature fusion layer.

[0145] Specifically, the image feature information to be processed includes target image feature information or the t-1th image fusion information, wherein the target image feature information is image feature information output by the feature extraction layer and input to the feature fusion layer; the t-1th image fusion information is image fusion information output by the t-1th local-global feature fusion layer.

[0146] See also Figure 6 , Figure 6 The structure diagram of an image processing model provided by an embodiment of the present specification is shown. The local-global feature fusion layer 600 may include a local feature extraction sublayer 6002 , a global feature extraction sublayer 6004 , and a fusion layer 6006 .

[0147] The local-global feature fusion layer 600 can receive feature information X of the image to be processed, input the feature information X of the image to be processed into the local feature extraction sublayer 6002 and the global feature extraction sublayer 6004 respectively, extract the local feature information of the feature information X of the image to be processed through the local feature extraction sublayer 6002; extract the global feature information of the feature information X of the image to be processed through the global feature extraction sublayer 6004, and then input the local feature information and the global feature information into the fusion layer 6006, and obtain the local-global fusion information of the feature information X of the image to be processed through the fusion layer 6006.

[0148] In practical applications, by inputting the feature information of the image to be processed into the local-global feature fusion layer in the local-global fusion block, the local feature information and the global feature information of the image to be processed can be extracted respectively, and the local feature information and the global feature information are fused to obtain the image fusion information output by the local-global feature fusion layer. By extracting and fusing the local feature information and the global feature information respectively, the global feature information encoding without losing the fine-scale information can be obtained, thereby improving the accuracy of the image processing model.

[0149] In one or more embodiments of the present specification, the local feature extraction sublayer may include a convolution processing unit.

[0150] Inputting the feature information of the image to be processed into the local feature extraction sublayer to obtain the local feature information corresponding to the feature information of the image to be processed may include the following steps:

[0151] Inputting the feature information of the image to be processed into the convolution processing unit, and extracting the initial local feature information corresponding to the feature information of the image to be processed;

[0152] Perform feature mapping on the initial local feature information to obtain local feature information corresponding to the feature information of the image to be processed.

[0153] See also Figure 7 , Figure 7 A structural diagram of an image processing model provided by an embodiment of this specification is shown.

[0154] Specifically, the local feature extraction sublayer 700 may include a convolution processing unit 7002 and a first convolution kernel 7004. The convolution processing unit 7002 may be understood as a convolution kernel for extracting local feature information. The feature information of the image to be processed is input into the convolution processing unit 7002, and the initial local feature information corresponding to the feature information of the image to be processed may be extracted and output through the convolution processing unit 7002; the initial local feature information may be understood as feature information with the same dimension as the feature information of the image to be processed.

[0155] In order to facilitate data classification of local feature information and obtain feature information with a finer scale, in one or more embodiments of the present specification, a first convolution kernel 7004 may be connected to the convolution processing unit 7002, and the first convolution kernel 7004 is used to perform feature mapping on the initial local feature information to obtain local feature information corresponding to the feature information of the image to be processed.

[0156] Optionally, the local feature information may be feature information of a higher dimension than the initial local feature information. For example, the dimension of the feature information X to be processed is c-dimensional, the dimension of the initial local feature information obtained by the convolution processing unit 7002 is also c-dimensional, and after feature mapping by the first convolution kernel 7004, the dimension of the obtained local feature information may be d-dimensional.

[0157] Optionally, the convolution processing unit 7002 may be a Conv3*3*3 convolution unit, and the first convolution kernel 7004 may be a Conv1*1*1 convolution kernel.

[0158] By inputting the feature information of the image to be processed into the local feature extraction sublayer, and obtaining the local feature information corresponding to the feature information of the image to be processed through the convolution processing unit and the first convolution kernel, the fineness of the image feature information extraction can be improved, and a more accurate image processing result can be obtained, thereby improving the accuracy of the image processing model.

[0159] In one or more embodiments of the present specification, the global feature extraction sublayer may include a codec image processing unit.

[0160] Inputting the feature information of the image to be processed into the global feature extraction sublayer to obtain the global feature information corresponding to the feature information of the image to be processed may include the following steps:

[0161] Perform feature mapping on the feature information of the image to be processed to obtain the feature information of the target image to be processed corresponding to the feature information of the image to be processed;

[0162] The feature information of the target image to be processed is input into the encoding and decoding image processing unit to obtain the global feature information corresponding to the feature information of the target image to be processed.

[0163] In practical applications, in order to facilitate the splicing and fusion of local feature information and global feature information, before extracting the global feature information through the codec image processing unit, the feature information of the image to be processed can be mapped to feature information with the same dimension as the local feature information.

[0164] Specifically, the target image feature information to be processed is feature information of higher dimension than the image feature information to be processed obtained after feature mapping. The codec image processing unit has a coding-decoding structure and can extract global feature information from the target image feature information to be processed.

[0165] See also Figure 8 , Figure 8 A structural diagram of an image processing model provided by an embodiment of this specification is shown.

[0166] Specifically, the global feature extraction sublayer 800 may include a second convolution kernel 8002 and a codec image processing unit 8004. A second convolution kernel 8002 may be connected before the codec image processing unit 8004. After the feature information of the image to be processed is input into the global feature extraction sublayer, it may be first mapped into target feature information of the image to be processed with a higher dimension through the second convolution kernel 8002, and then the target feature information of the image to be processed is input into the codec image processing unit 8004. The global feature information in the target feature information of the image to be processed is extracted through the encoder and decoder in the codec image processing unit 8004 to obtain the global feature information of the feature information of the image to be processed.

[0167] Optionally, the second convolution kernel 8002 may be a convolution kernel of Conv1*1*1, and the encoding and decoding image processing unit 8004 may be a transformer processing unit, which may extract global feature information from the feature information of the target image to be processed by Unfold->transformer->Fold.

[0168] By inputting the feature information of the image to be processed into the global feature extraction sublayer, and obtaining the global feature information corresponding to the feature information of the image to be processed through the second convolution kernel and the codec image processing unit, the integrity of the image feature information extraction can be improved, and more comprehensive image feature information can be obtained, thereby improving the reliability and accuracy of the image processing model.

[0169] In practical applications, after obtaining local feature information and global feature information, the local feature information and the global feature information can be input into a fusion layer to obtain image fusion information.

[0170] In one or more embodiments of the present specification, inputting local feature information and global feature information into a fusion layer to obtain image fusion information corresponding to the t-th local-global feature fusion layer may include the following steps:

[0171] Splice local feature information and global feature information to obtain initial image fusion information;

[0172] Perform feature mapping on the initial image fusion information to obtain the image fusion information corresponding to the tth local-global feature fusion layer.

[0173] Specifically, the initial image fusion information is feature information in the same dimension as the local feature information and the global feature information obtained by splicing the local feature information and the global feature information. The image fusion information can be understood as feature mapping the initial image fusion information to obtain feature information with the same dimension as the feature information of the image to be processed.

[0174] See also Fig. 9 , Fig. 9 A structural diagram of an image processing model provided by an embodiment of this specification is shown;

[0175] Specifically, the fusion layer 900 includes a splicing unit 9002 and a third convolution kernel 9004. The splicing unit 9002 is used to receive the local feature information and the global feature information output by the local feature extraction sublayer and the global feature extraction sublayer, and splice the local feature information and the global feature information to obtain the initial image fusion information. The third convolution kernel 9004 receives the initial image fusion information output by the splicing unit 9002, performs feature mapping on the initial image fusion information, and obtains the image fusion information corresponding to the t-th local global feature fusion layer. Optionally, the third convolution kernel 9004 can be a convolution kernel of Conv1*1*1.

[0176] By inputting local feature information and global feature information into the fusion layer, splicing them through the fusion layer to obtain initial image fusion information, and performing feature mapping on the initial image fusion information to obtain image fusion information, the local information of the target image at a finer scale and the more complete global information can be fused to obtain more accurate image feature information, which can make the prediction results output by the image processing model more accurate.

[0177] In practical applications, the image fusion information corresponding to each target image is obtained, and each image fusion information can be input into the output layer to obtain the detection result corresponding to the target partition to be detected output by the output layer.

[0178] S3006: Input the fusion information of each image into the output layer to obtain the detection result corresponding to the target partition to be detected.

[0179] Specifically, the output layer can calculate and output the detection result corresponding to the target partition to be detected based on the image fusion information corresponding to each target image. Furthermore, the output layer can classify the image fusion information through a softmax (normalization) operation to obtain the predicted probability of the existence of an abnormal object to be detected in the target partition to be detected, and output it as the detection result corresponding to the target partition to be detected.

[0180] Multiple target images are input into the image processing model, and the input multiple target images are processed into image feature information through the feature extraction layer of the image processing model. The image feature information is processed into image fusion information through the feature fusion layer, and then the image fusion information is processed through the output layer to output the detection result corresponding to the target partition to be detected. This can improve the accuracy of prediction of whether there are abnormal objects to be detected in the target partition to be detected, avoid inaccurate judgment results caused by human factors, and improve the detection efficiency of the target partition to be detected and the objects to be detected, so as to provide assistance for further medical diagnosis and treatment of the objects to be detected.

[0181] With the continuous development of computer technology, deep learning can gradually be applied to various medical imaging computer-aided diagnosis (CAD) tasks. However, deep learning models rely on large-scale accurately labeled training samples and sample labels as training data for model training. For image processing application scenarios where it is difficult to match the abnormal states of the objects to be detected one by one from medical images, it costs a lot to obtain large amounts of accurate training data, so it is more difficult to train the model.

[0182] Based on this, in one or more embodiments of this specification, the image processing model is obtained by the following S20402-S20410 training:

[0183] S20402: Acquire training sample pairs and guidance information of the training sample pairs, wherein the training sample pairs include training samples and sample labels, the training samples include multiple sample images corresponding to the target partition to be detected, and the sample labels are used to identify sample results of the target partition to be detected.

[0184] Specifically, the training sample pair includes a training sample and a sample label. The training sample can be understood as a plurality of sample images corresponding to the target partition to be detected, and the sample label is used to identify the sample result of the target partition to be detected.

[0185] In practical applications, in order to increase the amount of training data for training the model, reduce the annotation cost, and improve the annotation accuracy, the model training method of the image processing model provided in one or more embodiments of this specification can be used to mark whether there is an abnormal object to be detected in the target partition to be detected in the medical image according to the pathology report. That is, the sample result can be understood as whether there is an abnormal object to be detected in the target partition to be detected, and the sample label of the sample image with the abnormal object to be detected can be set to 1, and the sample label of the sample image without the abnormal object to be detected can be set to 0.

[0186] Since in the training sample pair, the training sample is the target partition to be detected for anomaly detection, the sample label is the probability of whether there is an abnormal object to be detected in the target partition to be detected, and the target partition to be detected may also contain other human cells or tissues and structures other than the object to be detected, which may interfere with the model training. Therefore, the model training can be implicitly or explicitly guided by the guidance information. The guidance information of the training sample pair can be understood as the information for a priori guidance of the attention of the image processing model during the model training process based on the training sample pair. Furthermore, the guidance information can be understood as the position information of the object to be detected in the target partition to be detected. Through the guidance information, the model's attention can be focused on the object to be detected in the target partition to be detected, preventing the model from being interfered by other objects in the target partition to be detected other than the object to be detected, resulting in inaccurate training results.

[0187] In order to obtain more accurate guidance information, thereby improving the accuracy of model training, in one or more real-time examples of this specification, obtaining guidance information of training sample pairs may include the following steps:

[0188] Determine the position information of the object to be detected in the target to be detected partition of each sample image;

[0189] The position information of the object to be detected is used as the guidance information corresponding to each sample image.

[0190] It should be noted that in order to focus the model's attention on the objects to be detected in the target to-be-detected partitions of each sample image and improve the accuracy of model training, the position information of the objects to be detected can be determined in the target to-be-detected partitions according to the medical image, and the position information of the objects to be detected can be used as the guidance information corresponding to each sample image. It is also possible to pre-train a depth detection model corresponding to the objects to be detected, use the depth detection model of the objects to be detected, automatically detect and segment all visible objects to be detected in the target to-be-detected partitions of the medical image, obtain mask information corresponding to the objects to be detected, and use the mask information obtained by the above method as the guidance information corresponding to each sample image.

[0191] S20404: Input the training sample pairs and guidance information into the initial image processing model to obtain the sample prediction results output by the initial image processing model and the target attention information of the target partition to be detected.

[0192] In practical applications, after obtaining the guidance information of the training sample pairs, the training sample pairs and the guidance information can be input into the initial image processing model. After being processed by the initial image processing model, the sample prediction results output by the initial image processing model and the target attention information of the target partition to be detected can be obtained.

[0193] Specifically, the initial image processing model can be understood as an image processing model that has not yet completed training, the sample prediction result can be understood as the predicted probability that there are abnormal objects to be detected in the training samples output by the initial image processing model based on the input training sample pairs and the guidance information, and the target attention information can be understood as the location information of the object to be detected in the target partition to be detected.

[0194] In one or more embodiments of the present specification, the image processing model may include a feature extraction layer, a feature fusion layer, and an output layer.

[0195] Accordingly, the training sample pairs and the guidance information are input into the initial image processing model to obtain the sample prediction results output by the initial image processing model and the target attention information of the target partition to be detected, which may include the following steps:

[0196] Splicing training samples and guidance information to obtain target splicing information;

[0197] Input the target splicing information into the feature extraction layer to obtain the feature extraction information output by the feature extraction layer;

[0198] Inputting the feature extraction information into the feature fusion layer to obtain feature fusion information output by the feature fusion layer and at least one attention feature map;

[0199] Input the feature fusion information into the output layer to obtain the sample prediction results corresponding to the target partition to be detected;

[0200] Aggregate at least one attention feature map to obtain target attention information corresponding to the target partition to be detected.

[0201] Specifically, the image processing model can splice the training samples and the guidance information in the training sample pair to obtain the target splicing information, and the target splicing information can be understood as the training sample that integrates the position information of the object to be detected. The feature extraction information can be understood as the feature information obtained by the feature extraction layer performing feature extraction based on the input target splicing information. The feature fusion information can be understood as the feature information obtained by the feature fusion layer performing local feature information extraction, global feature information extraction, and fusing local feature information and global feature information based on the input feature extraction information. The attention feature map can be understood as the intermediate product output by each local-global fusion block in the feature fusion layer in the process of processing the feature information. The attention feature map can be used to represent the position information of the object to be detected in the target partition to be detected.

[0202] It should be noted that each local-global fusion block in the feature fusion layer may include at least one local-global feature fusion layer. In order to improve the accuracy of training and obtain a more accurate attention feature map, in one or more embodiments of the present specification, the attention feature map output by the last local-global feature fusion layer in each local-global fusion block may be obtained.

[0203] In practical applications, after obtaining the feature fusion information and at least one attention feature map output by the feature fusion layer, the feature fusion information can be input into the output layer to obtain the sample prediction result corresponding to the target partition to be detected; and at least one attention feature map can be aggregated to obtain the target attention information corresponding to the target partition to be detected.

[0204] Specifically, the sample prediction result is the prediction result output by the initial image processing model on whether there is an abnormal object to be detected in the target partition to be detected. The prediction result can be set to: the higher the probability of detecting that there is an abnormal object to be detected in the target partition to be detected, the closer the prediction result is to 1; the lower the probability of detecting that there is an abnormal object to be detected in the target partition to be detected, the closer the prediction result is to 0.

[0205] In practical applications, aggregating at least one attention feature map to obtain the target attention information corresponding to the target partition to be detected can be achieved in the following way: interpolating each attention feature map to the same size as the attention feature map output by the first local-global fusion block, thereby achieving attention aggregation of at least one attention feature map and obtaining the target attention information of the target partition to be detected.

[0206] By splicing training samples and guidance information, the target splicing information is input into the initial image processing model, and the guidance information is used to explicitly guide the model training process, thereby improving the accuracy of model training, thereby improving the accuracy of the detection results output by the model, and improving the task processing effect when the model performs image processing tasks. By obtaining the attention feature map output by the feature fusion layer during the training process, and aggregating the attention feature map to obtain the target attention information corresponding to the target partition to be detected, it can be used for subsequent comparison of the target attention information and the guidance information, calculate the model loss value, adjust the model parameters, and realize implicit guidance of the model training process.

[0207] S20406: Calculate the first loss value of the initial image processing model based on the sample prediction results and the sample results.

[0208] Specifically, the first loss value can be understood as comparing the prediction result output by the initial image processing model with the sample label to obtain the error between the model task execution result and the actual result.

[0209] In practical applications, the sample prediction results and sample results can be compared, and the first loss value of the initial image processing model can be calculated through methods such as cross entropy calculation.

[0210] S20408: Calculate the second loss value of the initial image processing model based on the target attention information and the guidance information.

[0211] Specifically, the second loss value can be understood as the model error obtained by comparing the attention position information of the object to be detected in the target partition to be detected with the position information that the model should actually pay attention to during the model training process according to the initial image processing model.

[0212] In practical applications, the target attention information and the guidance information can be compared, and the second loss value of the initial image processing model can be calculated by methods such as cross entropy calculation.

[0213] S20410: Adjust the model parameters of the initial image processing model according to the first loss value and the second loss value until a preset stop condition is reached to obtain the image processing model.

[0214] Specifically, the preset stopping condition can be the preset thresholds corresponding to the first loss value and the second loss value respectively, or it can be the preset number of model cycle training. It can be determined based on actual conditions, and this specification does not impose any limitations on this.

[0215] In actual applications, the first loss value and the second loss value of the initial image processing model are calculated, and the parameters of the initial image processing model can be adjusted according to the first loss value and the second loss value. The parameter adjustment method can be exemplarily reverse transfer parameter adjustment, which can be determined according to the actual application situation, and this specification does not impose any limitation on this.

[0216] By splicing the guidance information with the training samples and calculating the second loss value through the guidance information and the target attention information, it is possible to achieve explicit and implicit guidance of the model training through the guidance information, thereby avoiding the model from being interfered by other objects other than the objects to be detected in the target partition to be detected, and improving the accuracy of model training, thereby improving the accuracy of the output prediction results when the image processing model performs image processing tasks.

[0217] An image processing method provided by an embodiment of the present specification receives an image processing task, wherein the image processing task carries multiple target images corresponding to a target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected; the multiple target images are input into an image processing model to obtain detection results corresponding to the target partition to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image.

[0218] In this way, by carrying multiple target images corresponding to the target partition to be detected in the image processing task and inputting the multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, thereby realizing automatic detection of whether there are abnormal objects to be detected in the target partition to be detected, and the image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy and efficiency of the detection results for the objects to be detected.

[0219] See also Fig.10 , Fig.10 A flowchart of a CT image processing method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0220] Step 1002: receiving a CT image processing task, wherein the CT image processing task carries a plurality of CT images corresponding to the target partition to be detected, and the CT image processing task is used to detect whether there are abnormal cells to be detected in the target partition to be detected.

[0221] Step 1004: Input multiple CT images into a CT image processing model to obtain detection results corresponding to the target partition to be detected, wherein the CT image processing model determines the detection results based on local image features and global image features of each CT image.

[0222] It should be noted that the implementation of step 1002 to step 1004 is the same as the implementation of step 202 to step 204 described above, and will not be described in detail in the embodiments of this specification.

[0223] Exemplarily, assuming that the CT image processing task is to determine whether there is lymph node metastasis in the esophageal cancer (EC) lymph node station, the above CT image processing method can be used to determine the probability of lymph node metastasis in the target lymph node station to be detected.

[0224] By applying the solution of the embodiments of the present specification, by carrying multiple target images corresponding to the target partition to be detected in the image processing task and inputting the multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, thereby enabling automatic detection of whether there are abnormal objects to be detected in the target partition to be detected, and the image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy of the detection result and the detection efficiency for the objects to be detected.

[0225] See also Fig.11 , Fig.11 A flowchart of a training method for an image processing model provided by an embodiment of this specification is shown, which is applied to a cloud-side device and specifically includes the following steps:

[0226] Step 1102: Obtain training sample pairs and guidance information of the training sample pairs, wherein the training sample pairs include training samples and sample labels, the training samples include multiple sample images corresponding to the target partition to be detected, and the sample labels are used to identify sample results of the target partition to be detected.

[0227] Step 1104: Input the training sample pairs and guidance information into the initial image processing model to obtain the sample prediction results output by the initial image processing model and the target attention information of the target partition to be detected.

[0228] Step 1106: Calculate the first loss value of the initial image processing model based on the sample prediction result and the sample result.

[0229] Step 1108: Calculate the second loss value of the initial image processing model based on the target attention information and the guidance information.

[0230] Step 1110: Adjust the model parameters of the initial image processing model according to the first loss value and the second loss value until a preset stop condition is reached, thereby obtaining the model parameters of the image processing model.

[0231] Step 1112: Send model parameters of the image processing model to the end-side device.

[0232] It should be noted that the implementation method of step 1102 to step 1110 is the same as that of the above-mentioned S20402 to S20410, and will not be repeated in this embodiment of the specification.

[0233] In actual applications, after the cloud-side device sends the model parameters of the image processing model to the terminal-side device, the terminal-side device can build the image processing model locally according to the model parameters of the image processing model, and further use the image processing model to perform image processing.

[0234] The scheme of the embodiments of the present specification is applied, by obtaining training sample pairs and guidance information of the training sample pairs, inputting the training samples and the guidance information into the initial image processing model, performing feature extraction and fusion of local and global feature information after splicing, outputting the sample prediction results, and calculating the first loss value of the initial image processing model by comparing the sample prediction results with the sample results. The second loss value of the initial image processing model is calculated by comparing the intermediate product target attention information output by the initial image processing model with the guidance information. The model parameters are adjusted by the first loss value and the second loss value, which can achieve implicit and explicit guidance of the model through the guidance information, and can make the final image processing model more accurate.

[0235] See also Fig.12 , Fig.12 A flowchart of an image processing method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0236] Step 1202: receiving an image processing request sent by a user, wherein the image processing request includes an image processing task, the image processing task carries multiple target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected.

[0237] Step 1204: Input multiple target images into the image processing model to obtain detection results corresponding to the target partitions to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image.

[0238] Step 1206: Send the detection result corresponding to the target partition to be detected to the user.

[0239] It should be noted that the specific implementation method of step 1202 to step 1204 is the same as the implementation method of the above-mentioned step 202 to step 204, and will not be repeated in the embodiment of this specification.

[0240] By applying the solution of the embodiments of the present specification, by carrying multiple target images corresponding to the target partition to be detected in the image processing task and inputting the multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, thereby enabling automatic detection of whether there are abnormal objects to be detected in the target partition to be detected, and the image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy of the detection result and the detection efficiency for the objects to be detected.

[0241] The following combination Fig.13 Taking the application of the image processing method provided in this specification in the detection of abnormal lymph node stations in esophageal cancer as an example, the image processing method is further described. Fig.13 A flowchart of a processing process of an image processing method provided by an embodiment of the present specification is shown, which specifically includes the following steps:

[0242] Step 1302: Receive a mediastinal lymph node CT image processing task, wherein the mediastinal lymph node CT image processing task carries multiple CT images corresponding to the target mediastinal lymph node to be detected, and the mediastinal lymph node CT image processing task is used to detect whether there is lymph node metastasis in the target mediastinal lymph node to be detected.

[0243] In practical applications, a mediastinal lymph node CT image processing task is received, and multiple CT images corresponding to the target mediastinal lymph node to be detected are input into the image processing model, so as to obtain the detection result of whether lymph node metastasis occurs in the target mediastinal lymph node to be detected based on the output of the image processing model based on the multiple CT images.

[0244] Step 1304: Input multiple CT images into a feature extraction layer included in the image processing model to obtain image feature information corresponding to each CT image, wherein the feature extraction layer includes two fully convolutional blocks connected in sequence.

[0245] For example, Fig.14 As shown, the image processing model includes a feature extraction layer, a feature fusion layer and an output layer, wherein the feature extraction layer, the feature fusion layer and the output layer are connected in sequence. Inputting multiple CT images into the feature extraction layer may include the following steps: inputting multiple CT images into full convolution block 1, obtaining image feature information 1 output by full convolution block 1, inputting image feature information 1 into full convolution block 2, obtaining image feature information 2 output by full convolution block 2, and determining image feature information 2 as image feature information output by the feature extraction layer.

[0246] Step 1306: Input feature information of each image into the feature fusion layer included in the image processing model to obtain image fusion information corresponding to each CT image, wherein the feature fusion layer includes 3 local-global fusion blocks connected in sequence, and each local-global fusion block includes at least one local-global feature fusion layer.

[0247] For example, Fig.14 As shown, the feature fusion layer includes three sequentially connected local-global fusion blocks 1, local-global fusion block 2 and local-global fusion block 3, and each local-global fusion block includes at least one local-global feature fusion layer. Each local-global feature fusion layer includes a local feature extraction sublayer, a global feature extraction sublayer and a fusion layer. Furthermore, the local feature extraction sublayer includes a convolution subunit, and the global feature extraction sublayer includes a codec image processing unit.

[0248] Exemplarily, it is assumed that the local-global fusion block 1 includes two local-global feature fusion layers, each image feature information is input into the local-global feature fusion layer 1 in the local-global fusion block 1, each image feature information is respectively input into the local feature extraction sublayer and the global feature extraction sublayer in the local-global feature fusion layer 1, and the local feature information is obtained according to the convolution subunit in the local feature extraction sublayer; the global feature information is obtained according to the codec image processing unit in the global feature extraction sublayer, the local feature information and the global feature information are fused through the fusion layer to obtain image fusion information, the local-global fusion information is input into the local-global feature fusion layer 2, and the image fusion information 1 output by the local-global fusion block 1 is obtained in the same manner as above. After obtaining the image fusion information 1 output by the local-global fusion block 1, the image fusion information 1 can be input into the local-global fusion block 2 to obtain the image fusion information 2 output by the local-global fusion block 2, and then the image fusion information 2 is input into the local-global fusion block 3 to obtain the image fusion information 3 output by the local-global fusion block 3, and the image fusion information 3 is determined as the image fusion information output by the feature fusion layer.

[0249] Step 1308: Input the fusion information of each image into the output layer included in the image processing model to obtain the detection result corresponding to the target mediastinal lymph node station to be detected.

[0250] In practical applications, the image fusion information is input into the output layer of the image processing model, and the image fusion information can be softmax (normalized) processed through the output layer to obtain the detection result corresponding to the target mediastinal lymph node station to be detected output by the image processing model.

[0251] Specifically, the closer the result to be detected is to 0, the lower the possibility that lymph node metastasis exists in the target mediastinal lymph node station to be detected, and the closer the result to be detected is to 1, the higher the possibility that lymph node metastasis exists in the target mediastinal lymph node station to be detected.

[0252] By applying the solution of the embodiments of the present specification, by carrying multiple target images corresponding to the target partition to be detected in the image processing task and inputting the multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, thereby enabling automatic detection of whether there are abnormal objects to be detected in the target partition to be detected, and the image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy of the detection result and the detection efficiency for the objects to be detected.

[0253] The following combination Fig.14 Taking the application of the training method of the image processing model provided in this specification in the abnormal detection of esophageal cancer lymph node stations as an example, the training method of the image processing model is further explained. Fig.14 A training diagram of an image processing model provided by an embodiment of the present specification is shown.

[0254] CT images and pathology reports of the parts to be detected of multiple esophageal cancer patients are obtained, the mediastinal lymph node stations to be detected are determined from the CT images of the parts to be detected, and the multiple CT images corresponding to the target mediastinal lymph node stations to be detected are used as training samples of the image processing model. According to the pathology report, sample labels are annotated for each CT image, wherein the sample label corresponding to the training sample with lymph node metastasis in the target mediastinal lymph node station to be detected is 1, and the sample label corresponding to the training sample without lymph node metastasis in the target mediastinal lymph node station to be detected is 0.

[0255] Through the pre-trained deep lymph node detection model, all visible lymph nodes (short diameter ≥ 5mm) in the above CT images are automatically detected and segmented, and the lymph node masks obtained by these automatic segmentations are used as the guidance information of the attention prior to guide the model learning.

[0256] The constructed training sample pairs and the guidance information corresponding to the training sample pairs are used as the input of the image processing model and input into the initial image processing model to be trained. The initial image processing model receives the training sample pairs and the guidance information, first splices the training samples and the guidance information to obtain the sample splicing information, and then inputs the sample splicing information into two sequentially connected full convolution blocks (full convolution block 1, full convolution block 2) to obtain the image feature information output by the second full convolution block; then the image feature information is input into three sequentially connected local-global fusion blocks (local-global fusion block 1, local-global fusion block 2, local-global fusion block 3) to obtain the image fusion information output by the third local-global fusion block, wherein the local-global fusion block 1 includes 2 local-global feature fusion layers, the local-global fusion block 2 includes 4 local-global feature fusion layers, and the local-global fusion block 3 includes 3 local-global feature fusion layers.

[0257] In the process of obtaining the initial image processing model to process the image feature information through three local-global fusion blocks, the attention feature map output by the last local-global feature fusion layer of each local-global fusion block is interpolated to the same size, and then the target attention information corresponding to the three attention feature maps is aggregated.

[0258] The image fusion information output by the third local-global fusion block is output to the output layer to obtain the probability of lymph node metastasis in the mediastinal lymph node station to be detected output by the initial image processing model, and this probability value is determined as the sample prediction result.

[0259] According to the sample prediction results and sample labels, the first loss value of the initial image processing model is calculated; according to the target attention information and the guidance information, the second loss value of the initial image processing model is calculated. And according to the first loss value and the second loss value, the model parameters of the initial image model are adjusted until the preset stop condition is reached to obtain the image processing model.

[0260] The scheme of the embodiments of the present specification is applied, by obtaining training sample pairs and guidance information of the training sample pairs, inputting the training samples and the guidance information into the initial image processing model, performing feature extraction and fusion of local and global feature information after splicing, outputting the sample prediction results, and calculating the first loss value of the initial image processing model by comparing the sample prediction results with the sample results. The second loss value of the initial image processing model is calculated by comparing the intermediate product target attention information output by the initial image processing model with the guidance information. The model parameters are adjusted by the first loss value and the second loss value, which can achieve implicit and explicit guidance of the model through the guidance information, and can make the final image processing model more accurate.

[0261] See also Fig.15 , Fig.15A flowchart of an esophageal cancer detection method provided by an embodiment of this specification is shown, which specifically includes the following steps:

[0262] Step 1502: Train the image detection model.

[0263] Specifically, multiple esophageal cancer CT images can be obtained, and the lymph nodes in the CT images can be segmented by the Deep-Stationing model to obtain multiple lymph nodes in different positions. Through the pathology report of the esophageal cancer patient, whether there is esophageal cancer lymph node metastasis in each lymph node is marked to obtain a training sample pair, wherein the lymph node with lymph node metastasis is marked as 1, and the lymph node without lymph node metastasis is marked as 0. Further, the training sample pair includes a training sample and a sample label, and the training sample includes a plurality of sample images corresponding to the target lymph node to be detected, and the sample label is used to identify the sample result of the target lymph node to be detected. At the same time, since the pathology report can only determine whether lymph node metastasis has occurred in the lymph node station, and it is impossible to accurately mark whether lymph node metastasis has occurred in a single cell in the lymph node station, therefore, in order to improve the detection accuracy and help doctors locate the position of abnormal cells for subsequent diagnosis and treatment, one or more embodiments of this specification also use an LN detection algorithm to segment the LN instance to obtain an LN mask. By using the LN mask as the guidance information of the training sample pair and inputting it into the initial image detection model together with the training sample pair, the sample prediction results output by the initial image detection model and the target attention information of the target lymph node to be detected can be obtained. By explicitly guiding the detection results of the image detection model through the sample results and implicitly guiding the detection results of the image detection model through the guidance information, the image detection model can aggregate its attention to the lymph node to be detected, improve the position accuracy, and thus improve the accuracy of the output results of the image detection model.

[0264] It should be noted that the LN detection algorithm is based on the 2.5D Mask R-CNN framework, which takes 9 consecutive axial CT slices as input, fuses 2D feature images to obtain 3D context information, predicts the position and mask of LN instances in CT slices, and can obtain accurate LN location information.

[0265] According to the sample prediction results output by the initial image detection model and the sample results corresponding to the sample labels, the first loss value of the initial image detection model can be calculated; according to the target attention information output by the initial image detection model and the guidance information corresponding to the training sample pairs, the second loss value of the initial image detection model can be calculated; according to the first loss value and the second loss value, the model parameters of the initial image detection model are adjusted together, and then the adjusted initial image detection model is used to continue to predict whether there are cancer cells in multiple sample images corresponding to the target lymph node to be detected, until the preset stop condition is reached, and a trained image detection model is obtained.

[0266] Step 1504: receiving an esophageal cancer CT image detection task, wherein the esophageal cancer CT image detection task carries multiple target lymph node station CT images corresponding to the target lymph node to be detected, and the esophageal cancer CT image detection task is used to detect whether lymph node metastasis occurs in the target lymph node to be detected.

[0267] In practical applications, when receiving an esophageal cancer CT image detection task, the lymph nodes in the esophageal cancer CT image can be segmented in advance through the Deep-Stationing model to determine the target lymph nodes to be detected, and multiple target lymph node CT images corresponding to the target lymph nodes to be detected are input into the trained image detection model to detect whether there are cancer cells in the lymph nodes.

[0268] Step 1506: Input multiple target lymph node CT images into the image detection model to obtain detection results corresponding to the target lymph node to be detected, wherein the image detection model determines the detection results based on local image features and global image features of each target lymph node CT image.

[0269] In practical applications, through the image detection model, multiple lymph node CT images can be first input into the feature extraction layer included in the image detection model to obtain the image feature information corresponding to each lymph node CT image, wherein the feature extraction layer includes 2 sequentially connected full convolution blocks. Specifically, multiple lymph node CT images are input into full convolution block 1 to obtain image feature information 1 output by full convolution block 1, image feature information 1 is input into full convolution block 2 to obtain image feature information 2 output by full convolution block 2, and image feature information 2 is determined as the image feature information output by the feature extraction layer.

[0270] Then, the feature information of each image can be input into the feature fusion layer included in the image detection model to obtain the image fusion information corresponding to each lymph node CT image, wherein the feature fusion layer includes 3 local-global fusion blocks connected in sequence, each local-global fusion block includes at least one local-global feature fusion layer, each local-global feature fusion layer includes a local feature extraction sublayer, a global feature extraction sublayer and a fusion layer, the local feature extraction sublayer includes a convolution subunit, and the global feature extraction sublayer includes a codec image processing unit.

[0271] Finally, the image fusion information can be input into the output layer included in the image detection model to obtain the detection result corresponding to the target lymph node to be detected. In practical applications, the image fusion information is input into the output layer included in the image detection model, and the image fusion information can be softmax (normalized) processed by the output layer to obtain the detection result corresponding to the target lymph node to be detected output by the image detection model. Specifically, the closer the result to be detected is to 0, the lower the possibility of lymph node metastasis in the target lymph node to be detected, and the closer the result to be detected is to 1, the higher the possibility of lymph node metastasis in the target lymph node to be detected.

[0272] In the embodiments provided in this specification, training sample pairs and guiding information of sample pairs are constructed based on a data set of 1205 patients with esophageal cancer collected. Based on this, a larger data set for the classification of LN metastasis of esophageal cancer can be constructed, and the label of the LN station is determined by the LN anatomical results in the data set. If all anatomical lymph nodes of the lymph node in the report are shown to be benign, the lymph node is benign, and if there are any metastatic lymph nodes in the report, it is marked as metastasis. The median size of the CT image is 512×512×91mm, and the median resolution is 0.795×0.795×5.0mm. All CT images are resampled to a consistent spatial resolution of 0.787×0.787×5.0mm. The 3D training patches are generated by cropping the 96×96×32ROI of the CT image and the LN mask, respectively, centered on each LN station, and each LN-station binary map is further used to "mask" the CT image by setting the voxels outside the LN-station to a constant value of -1024. Configuring local-global transformer blocks with L=2, 4, and 3 respectively and setting the threshold τ in the mask-guided margin loss to 0.8 can achieve better verification performance. The image detection model can be developed based on PyTorch v1.12.1. For the training of the image detection model, the AdamW optimizer with a learning rate of 1.6e-3 and a weight decay of 5e-4 can be used. It is stipulated that the image detection model converges after 250 trainings.

[0273] See Table 1 below, which shows a quantitative comparison of five-fold cross validation of segmented data at the patient level for different image detection models and input settings. Wherein, LMViT is the model structure of the image detection model provided in the embodiment of this specification, and LMViT+Lmask is the image detection model trained according to the training method of the image detection model provided in the embodiment of this specification. AUROC is the area under the characteristic curve, R@S80 is the recall rate with a specificity rate of 80%, S@R80 is the specificity with a recall rate of 80%, and Mask LN Whether to use LN instance as mask for implicit guidance.

[0274] Table 1

[0275]

[0276] It can be obtained that when the CT image is used alone as input (without the LN instance mask), the LMViT provided in the embodiment of this specification improves the AUROC to 0.8214, which is the highest among all evaluation models, proving the effectiveness of the proposed local-global feature fusion architecture. When the LN instance mask is further used as an auxiliary input, all models show performance gain, and the AUROC exceeds 0.83. By applying the solution of the embodiment of this specification, by carrying multiple target images corresponding to the target partition to be detected in the image processing task, and inputting the multiple target images into the image processing model, the detection result corresponding to the target partition to be detected output by the image processing model can be obtained, thereby realizing automatic detection of whether there is an abnormal object to be detected in the target partition to be detected, and the image processing model determines the detection result based on the local image features and global image features of each target image, which can improve the accuracy and efficiency of the detection result for the object to be detected.

[0277] Corresponding to the above method embodiment, this specification also provides an image processing device embodiment, Fig.16 FIG. 2 shows a schematic diagram of the structure of an image processing device provided by an embodiment of the present specification. Fig.16 As shown, the device comprises:

[0278] The receiving module 1602 is configured to receive an image processing task, wherein the image processing task carries a plurality of target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected;

[0279] The detection module 1604 is configured to input multiple target images into an image processing model to obtain detection results corresponding to the target partitions to be detected, wherein the image processing model determines the detection results based on local image features and global image features of each target image.

[0280] Optionally, the image processing model includes a feature extraction layer, a feature fusion layer and an output layer, wherein the feature fusion layer is used to fuse local image features and global image features of each target image;

[0281] Accordingly, the detection module 1604 is further configured to:

[0282] Input multiple target images into the feature extraction layer to obtain image feature information corresponding to each target image;

[0283] Input the feature information of each image into the feature fusion layer to obtain the image fusion information corresponding to each target image;

[0284] The fusion information of each image is input into the output layer to obtain the detection result corresponding to the target partition to be detected.

[0285] Optionally, the feature extraction layer includes n fully convolutional blocks connected in sequence, where n is a positive integer greater than or equal to 2;

[0286] Accordingly, the detection module 1604 is further configured to:

[0287] determining a target image to be processed among a plurality of target images;

[0288] Input the target image to be processed into the first full convolution block to obtain the first image feature information;

[0289] Input the i-1th image feature information into the i-th full convolution block to obtain the i-th image feature information, where 2≤i≤n;

[0290] Determine whether i is equal to n;

[0291] If not, i is incremented by 1, and the operation of inputting the i-1th image feature information into the i-th full convolution block to obtain the i-th image feature information is continued;

[0292] If so, the i-th image feature information is determined as the image feature information corresponding to the target image to be processed.

[0293] Optionally, the feature fusion layer includes m local-global fusion blocks connected sequentially, where m is a positive integer greater than or equal to 2;

[0294] Accordingly, the detection module 1604 is further configured to:

[0295] Determine target image feature information from each image feature information;

[0296] Input the target image feature information into the first local-global fusion block to obtain the first image fusion information;

[0297] Input the j-1th image fusion information into the jth local-global fusion block to obtain the jth image fusion information, where 2≤j≤m;

[0298] Determine whether j is equal to m;

[0299] If not, j is incremented by 1, and the operation of inputting the j-1th image fusion information into the jth local-global fusion block to obtain the jth image fusion information is continued;

[0300] If so, the j-th image fusion information is determined as the image fusion information corresponding to the target image feature information.

[0301] Optionally, the local-global fusion block includes at least one local-global feature fusion layer, the local-global feature fusion layer includes a local feature extraction sublayer, a global feature extraction sublayer, and a fusion layer; accordingly, the detection module 1604 includes a t-th local-global feature fusion layer, which is configured as follows:

[0302] Acquire feature information of the image to be processed, wherein the feature information of the image to be processed includes feature information of the target image or fusion information of the t-1th image;

[0303] Inputting the feature information of the image to be processed into the local feature extraction sublayer to obtain the local feature information corresponding to the feature information of the image to be processed;

[0304] Inputting the feature information of the image to be processed into the global feature extraction sublayer to obtain the global feature information corresponding to the feature information of the image to be processed;

[0305] The local feature information and the global feature information are input into the fusion layer to obtain the image fusion information corresponding to the t-th local-global feature fusion layer.

[0306] Optionally, the local feature extraction sublayer includes a convolution processing unit; accordingly, the t-th local-global feature fusion layer is further configured as follows:

[0307] Input the feature information of the image to be processed into the local feature extraction sublayer to obtain the local feature information corresponding to the feature information of the image to be processed, including:

[0308] Inputting the feature information of the image to be processed into the convolution processing unit, and extracting the initial local feature information corresponding to the feature information of the image to be processed;

[0309] Perform feature mapping on the initial local feature information to obtain local feature information corresponding to the feature information of the image to be processed.

[0310] Optionally, the global feature extraction sublayer includes a codec image processing unit;

[0311] Accordingly, the tth local-global feature fusion layer is further configured as:

[0312] Perform feature mapping on the feature information of the image to be processed to obtain the feature information of the target image to be processed corresponding to the feature information of the image to be processed;

[0313] The feature information of the target image to be processed is input into the encoding and decoding image processing unit to obtain the global feature information corresponding to the feature information of the target image to be processed.

[0314] Optionally, the tth local-global feature fusion layer is further configured as follows:

[0315] Splice local feature information and global feature information to obtain initial image fusion information;

[0316] Perform feature mapping on the initial image fusion information to obtain the image fusion information corresponding to the tth local-global feature fusion layer.

[0317] Optionally, the image processing apparatus further includes a training module configured to:

[0318] Acquire training sample pairs and guidance information of the training sample pairs, wherein the training sample pairs include training samples and sample labels, the training samples include multiple sample images corresponding to the target partition to be detected, and the sample labels are used to identify sample results of the target partition to be detected;

[0319] Input the training sample pairs and the guidance information into the initial image processing model to obtain the sample prediction results output by the initial image processing model and the target attention information of the target partition to be detected;

[0320] According to the sample prediction results and the sample results, a first loss value of the initial image processing model is calculated;

[0321] According to the target attention information and the guidance information, a second loss value of the initial image processing model is calculated;

[0322] The model parameters of the initial image processing model are adjusted according to the first loss value and the second loss value until a preset stop condition is reached to obtain an image processing model.

[0323] Optionally, the training module is further configured to:

[0324] Determine the position information of the object to be detected in the target to be detected partition of each sample image;

[0325] The position information of the object to be detected is used as the guidance information corresponding to each sample image.

[0326] Optionally, the image processing model includes a feature extraction layer, a feature fusion layer and an output layer;

[0327] Accordingly, the training module is further configured as follows:

[0328] Splicing training samples and guidance information to obtain target splicing information;

[0329] Input the target splicing information into the feature extraction layer to obtain the feature extraction information output by the feature extraction layer;

[0330] Inputting the feature extraction information into the feature fusion layer to obtain feature fusion information output by the feature fusion layer and at least one attention feature map;

[0331] Input the feature fusion information into the output layer to obtain the sample prediction results corresponding to the target partition to be detected;

[0332] Aggregate at least one attention feature map to obtain target attention information corresponding to the target partition to be detected.

[0333] An image processing device provided by an embodiment of the present specification can obtain the detection results corresponding to the target partition to be detected output by the image processing model by carrying multiple target images corresponding to the target partition to be detected in the image processing task and inputting the multiple target images into the image processing model, thereby realizing automatic detection of whether there are abnormal objects to be detected in the target partition to be detected, and the image processing model determines the detection results based on the local image features and global image features of each target image, which can improve the accuracy and efficiency of the detection results for the objects to be detected.

[0334] The above is a schematic scheme of an image processing device of this embodiment. It should be noted that the technical scheme of the image processing device and the technical scheme of the above-mentioned image processing method belong to the same concept, and the details not described in detail in the technical scheme of the image processing device can be referred to the description of the technical scheme of the above-mentioned image processing method.

[0335] Fig.17 The structure block diagram of a computing device provided by an embodiment of the present specification is shown. The components of the computing device 1700 include but are not limited to a memory 1710 and a processor 1720. The processor 1720 is connected to the memory 1710 via a bus 1730, and the database 1750 is used to store data.

[0336] The computing device 1700 also includes an access device 1740, which enables the computing device 1700 to communicate via one or more networks 1760. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1740 may include one or more of any type of network interface (e.g., a network interface card (NIC)) of wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world microwave interconnection access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0337] In one embodiment of the present specification, the above components of the computing device 1700 and Fig.17 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Fig.17 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0338] The computing device 1700 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1700 may also be a mobile or stationary server.

[0339] The processor 1720 is used to execute the following computer executable instructions, which implement the steps of the above method when executed by the processor.

[0340] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above method.

[0341] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which can implement the steps of the above method when executed by a processor.

[0342] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above method.

[0343] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.

[0344] The above is an illustrative solution of a computer program of this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above method belong to the same concept, and the details not described in detail in the technical solution of the computer program can be referred to the description of the technical solution of the above method.

[0345] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0346] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0347] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0348] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0349] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described herein within the scope permitted by applicable laws and regulations, subject to the requirements of applicable laws and regulations of the country where the user is located (for example, with the user's explicit consent, effective notification to the user, etc.).

[0350] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. An image processing method, comprising: Receive an image processing task, wherein the image processing task carries a plurality of target images corresponding to the target to-be-detected partition, and the image processing task is used to detect whether there is an abnormal to-be-detected object in the target to-be-detected partition; Input the multiple target images into the feature extraction layer of the image processing model to obtain image feature information corresponding to each target image, input each image feature information into the feature fusion layer of the image processing model to obtain image fusion information corresponding to each target image, input each image fusion information into the output layer of the image processing model to obtain the detection result corresponding to the target partition to be detected, the feature fusion layer includes at least two local-global fusion blocks connected in sequence, the local-global fusion block is used to receive the target image feature information in each image feature information or the image fusion information output by the previous local-global fusion block, and obtain the image fusion information corresponding to the local-global fusion block, the local-global fusion block includes at least one local-global feature fusion layer, the local-global feature fusion layer includes a local feature extraction sublayer, a global feature extraction sublayer, and a fusion layer; for the t-th local-global feature fusion layer, the method includes: Acquire feature information of an image to be processed, wherein the feature information of the image to be processed includes feature information of a target image or fusion information of the t-1th image; input the feature information of the image to be processed into the local feature extraction sublayer to obtain local feature information corresponding to the feature information of the image to be processed; input the feature information of the image to be processed into the global feature extraction sublayer to obtain global feature information corresponding to the feature information of the image to be processed; input the local feature information and the global feature information into the fusion layer to obtain image fusion information corresponding to the tth local-global feature fusion layer.

2. According to the method of claim 1, the feature extraction layer comprises n sequentially connected full convolution blocks, wherein: n is a positive integer greater than or equal to 2; The step of inputting the plurality of target images into the feature extraction layer to obtain image feature information corresponding to each target image includes: Determining a target image to be processed among the multiple target images; Inputting the target image to be processed into a first full convolution block to obtain first image feature information; Input the i-1th image feature information into the i-th full convolution block to obtain the i-th image feature information, where 2≤i≤n; Determine whether i is equal to n; If not, i is incremented by 1, and the operation of inputting the i-1th image feature information into the i-th full convolution block to obtain the i-th image feature information is continued; If so, the i-th image feature information is determined as the image feature information corresponding to the target image to be processed.

3. According to the method of claim 1, the feature fusion layer comprises m sequentially connected local-global fusion blocks, wherein: m is a positive integer greater than or equal to 2; The step of inputting the feature information of each image into the feature fusion layer to obtain the image fusion information corresponding to each target image includes: Determine target image feature information from each image feature information; Inputting the target image feature information into a first local-global fusion block to obtain first image fusion information; Input the j-1th image fusion information into the jth local-global fusion block to obtain the jth image fusion information, where 2≤j≤m; Determine whether j is equal to m; If not, j is incremented by 1, and the operation of inputting the j-1th image fusion information into the jth local-global fusion block to obtain the jth image fusion information is continued; If so, the j-th image fusion information is determined as the image fusion information corresponding to the target image feature information.

4. The method according to any one of claims 1 to 3, wherein the local feature extraction sublayer comprises a convolution processing unit; The step of inputting the image feature information to be processed into the local feature extraction sublayer to obtain local feature information corresponding to the image feature information to be processed includes: Inputting the feature information of the image to be processed into the convolution processing unit, and extracting initial local feature information corresponding to the feature information of the image to be processed; Feature mapping is performed on the initial local feature information to obtain local feature information corresponding to the feature information of the image to be processed.

5. The method according to any one of claims 1 to 3, wherein the global feature extraction sublayer comprises a codec image processing unit; The step of inputting the feature information of the image to be processed into the global feature extraction sublayer to obtain the global feature information corresponding to the feature information of the image to be processed includes: Performing feature mapping on the feature information of the image to be processed to obtain feature information of the target image to be processed corresponding to the feature information of the image to be processed; The target image feature information to be processed is input into the encoding and decoding image processing unit to obtain global feature information corresponding to the target image feature information to be processed.

6. According to the method of claim 1, the image processing model is trained by the following steps: Obtain training sample pairs and guidance information of the training sample pairs, wherein: The training sample pair includes a training sample and a sample label, the training sample includes a plurality of sample images corresponding to the target partition to be detected, and the sample label is used to identify the sample result of the target partition to be detected; Inputting the training sample pair and the guidance information into an initial image processing model to obtain a sample prediction result output by the initial image processing model and target attention information of the target partition to be detected; Calculating a first loss value of the initial image processing model according to the sample prediction result and the sample result; Calculating a second loss value of the initial image processing model according to the target attention information and the guidance information; The model parameters of the initial image processing model are adjusted according to the first loss value and the second loss value until a preset stop condition is reached to obtain an image processing model.

7. The method according to claim 6, wherein the image processing model comprises a feature extraction layer, a feature fusion layer and an output layer; Inputting the training sample pair and the guide information into an initial image processing model to obtain a sample prediction result output by the initial image processing model and target attention information of the target partition to be detected, including: Splicing the training sample and the guide information to obtain target splicing information; Inputting the target splicing information into the feature extraction layer to obtain feature extraction information output by the feature extraction layer; Inputting the feature extraction information into the feature fusion layer to obtain feature fusion information output by the feature fusion layer and at least one attention feature map; Inputting the feature fusion information into the output layer to obtain the sample prediction result corresponding to the target partition to be detected; Aggregate the at least one attention feature map to obtain target attention information corresponding to the target partition to be detected.

8. A CT image processing method, comprising: receiving a CT image processing task, wherein the CT image processing task carries a plurality of CT images corresponding to the target partition to be detected, and the CT image processing task is used to detect whether there are abnormal cells to be detected in the target partition to be detected; Input the multiple CT images into the feature extraction layer of the CT image processing model to obtain image feature information corresponding to each CT image, input each image feature information into the feature fusion layer of the CT image processing model to obtain image fusion information corresponding to each CT image, input each image fusion information into the output layer of the CT image processing model to obtain the detection result corresponding to the target partition to be detected, the feature fusion layer includes at least two local-global fusion blocks connected in sequence, the local-global fusion block is used to receive the target image feature information in each image feature information or the image fusion information output by the previous local-global fusion block, and obtain the image fusion information corresponding to the local-global fusion block, the local-global fusion block includes at least one local-global feature fusion layer, the local-global feature fusion layer includes a local feature extraction sublayer, a global feature extraction sublayer, and a fusion layer; for the t-th local-global feature fusion layer, the method includes: Acquire feature information of an image to be processed, wherein the feature information of the image to be processed includes feature information of a target image or fusion information of the t-1th image; input the feature information of the image to be processed into the local feature extraction sublayer to obtain local feature information corresponding to the feature information of the image to be processed; input the feature information of the image to be processed into the global feature extraction sublayer to obtain global feature information corresponding to the feature information of the image to be processed; input the local feature information and the global feature information into the fusion layer to obtain image fusion information corresponding to the tth local-global feature fusion layer.

9. A training method for an image processing model, applied to a cloud-side device, comprising: Acquire a training sample pair and guidance information of the training sample pair, wherein the training sample pair includes a training sample and a sample label, the training sample includes a plurality of sample images corresponding to the target partition to be detected, and the sample label is used to identify the sample result of the target partition to be detected; Splicing the training sample and the guide information to obtain target splicing information; Input the target splicing information into the feature extraction layer of the initial image processing model to obtain the feature extraction information output by the feature extraction layer, input the feature extraction information into the feature fusion layer of the initial image processing model to obtain the feature fusion information and at least one attention feature map output by the feature fusion layer, input the feature fusion information into the output layer of the initial image processing model to obtain the sample prediction result corresponding to the target partition to be detected, aggregate the at least one attention feature map, and obtain the target attention information of the target partition to be detected, the feature fusion layer includes at least two local global fusion blocks connected in sequence, the local global fusion block is used to receive the feature extraction information or the feature fusion information output by the previous local global fusion block, and obtain the feature fusion information corresponding to the local global fusion block, the local global fusion block includes at least one local global feature fusion layer, and the local global feature fusion layer includes a local feature extraction sublayer, a global feature extraction sublayer, and a fusion layer; Calculating a first loss value of the initial image processing model according to the sample prediction result and the sample result; Calculating a second loss value of the initial image processing model according to the target attention information and the guidance information; Adjusting the model parameters of the initial image processing model according to the first loss value and the second loss value until a preset stop condition is reached, thereby obtaining the model parameters of the image processing model; The model parameters of the image processing model are sent to the end-side device.

10. An image processing method, comprising: Receive an image processing request sent by a user, wherein the image processing request includes an image processing task, the image processing task carries multiple target images corresponding to the target partition to be detected, and the image processing task is used to detect whether there is an abnormal object to be detected in the target partition to be detected; The multiple target images are input into the feature extraction layer of the image processing model to obtain image feature information corresponding to each target image, each image feature information is input into the feature fusion layer of the image processing model to obtain image fusion information corresponding to each target image, each image fusion information is input into the output layer of the image processing model to obtain the detection result corresponding to the target partition to be detected, the feature fusion layer includes at least two local global fusion blocks connected in sequence, the local global fusion block is used to receive the target image feature information in each image feature information or the image fusion information output by the previous local global fusion block, and obtain the image fusion information corresponding to the local global fusion block, the local global fusion block includes at least one local global feature fusion layer , the local-global feature fusion layer includes a local feature extraction sublayer, a global feature extraction sublayer, and a fusion layer; for the t-th local-global feature fusion layer, the method includes: obtaining feature information of the image to be processed, the feature information of the image to be processed includes feature information of the target image or fusion information of the t-1th image; inputting the feature information of the image to be processed into the local feature extraction sublayer to obtain local feature information corresponding to the feature information of the image to be processed; inputting the feature information of the image to be processed into the global feature extraction sublayer to obtain global feature information corresponding to the feature information of the image to be processed; inputting the local feature information and the global feature information into the fusion layer to obtain image fusion information corresponding to the t-th local-global feature fusion layer; The detection result corresponding to the target partition to be detected is sent to the user.

11. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.

12. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and readable storage medium

    CN115170896A

  • Unified cascade panoramic narrative detection and segmentation method

    CN116050409A