Image processing method, method for training an image processing model

By extracting the texture information of the detection area in the medical image in the image processing model, the problem of low image recognition accuracy in traditional Chinese medicine in the prior art is solved, and more accurate recognition of mild and moderate fatty liver is achieved.

CN116993680BActive Publication Date: 2025-06-17ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310814888.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-04
Publication Date
2025-06-17
Estimated Expiration
2043-07-04

AI Technical Summary

Technical Problem

The prior art recognizes medical images with low recognition accuracy, especially the subtle differences in liver attenuation caused by mild or moderate fatty liver cannot be accurately identified.

Method used

By receiving the image processing task, multiple target images are input into the image processing model, detection area feature information and texture information are generated, and the final detection result is generated. In the process of processing the target image by the image processing model, in addition to considering the detection area feature information, the detection area texture information is also extracted from the detection area feature information to improve the recognition accuracy.

Benefits of technology

By extracting the texture information of the detection area, the characteristics of the target detection area are enriched, thereby improving the recognition accuracy of the image processing model and more accurately identifying subtle differences in medical images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116993680B_ABST
    Figure CN116993680B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide an image processing method and a training method for an image processing model. The image processing method includes: receiving an image processing task, where the image processing task carries multiple target images corresponding to a target detection area, and the target image processing task is used to detect whether the target detection area is abnormal; inputting the multiple target images into the image processing model to obtain a detection result corresponding to the target detection area, where the image processing model generates detection area feature information and detection area texture information based on each target image, and generates the detection result based on the detection area feature information and the detection area texture information. The method provided in this specification combines the detection area feature information and the detection area texture information, improves the dimension of image processing, and further improves the accuracy of the detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and particularly to an image processing method. Background Art

[0002] With the improvement of people's living standards, the incidence rate of various organs of the human body has been increasing year by year, gradually becoming a major hidden danger threatening health, such as fatty liver, pulmonary fibrosis, and so on. Currently, usually, imaging tools such as ultrasonic waves and magnetic resonance imaging are used to assist operators in examinations, but their accuracy depends on the skills and experience of the operators. Based on this, computed tomography (CT) provides an operator-independent, standardized, and general method for quantitatively determining fat content. By discriminating medical images, it has become an important means for screening.

[0003] In the current recognition of medical images, the recognition accuracy is relatively low. For example, the subtle differences in liver attenuation caused by mild fatty liver or moderate fatty liver cannot be accurately recognized and evaluated. Therefore, how to improve the recognition accuracy of medical images has become an urgent problem for technicians to solve. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide an image processing method. One or more embodiments of this specification also relate to an image processing device, a computing device, a computer-readable storage medium, and a computer program, so as to solve the technical defects existing in the prior art.

[0005] According to the first aspect of the embodiments of this specification, an image processing method is provided, including:

[0006] Receiving an image processing task, where the image processing task carries a plurality of target images corresponding to a target detection region, and the target image processing task is used to detect whether the target detection region is abnormal;

[0007] Inputting the plurality of target images into an image processing model to obtain a detection result corresponding to the target detection region, where the image processing model generates detection region feature information and detection region texture information based on each target image, and generates the detection result based on the detection region feature information and the detection region texture information.

[0008] According to the second aspect of the embodiments of this specification, a CT image processing method is provided, including:

[0009] Receiving a CT image processing task, where the CT image processing task carries a plurality of CT images corresponding to a target detection region, and the CT image processing task is used to detect whether there is an abnormality in the target detection region;

[0010] Input the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection region, where the CT image processing model generates detection region feature information and detection region texture information based on each CT image, and generates the detection result based on the detection region feature information and the detection region texture information.

[0011] According to a third aspect of the embodiments of the present specification, there is provided a method for training an image processing model, which is applied to a cloud-side device and includes:

[0012] Obtain a training sample pair, where the training sample pair includes a plurality of training sample images corresponding to a target detection region and a sample detection result of the target detection region, and the sample detection result includes a standard sample detection result or a reference sample detection result;

[0013] Input the multiple training sample images into an image processing model to obtain a predicted detection result corresponding to the target detection region;

[0014] Calculate a model loss value according to the predicted detection result and the sample detection result;

[0015] Adjust model parameters of the image processing model according to the model loss value until a model training stop condition is reached, and obtain model parameters of the image processing model;

[0016] Send the model parameters of the image processing model to a terminal-side device.

[0017] According to a fourth aspect of the embodiments of the present specification, there is provided an image processing method, which includes:

[0018] Receive an image processing request sent by a user, where the image processing request includes an image processing task, the image processing task carries a plurality of target images corresponding to a target detection region, and the target image processing task is used to detect whether the target detection region is abnormal;

[0019] Input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection region, where the image processing model generates detection region feature information and detection region texture information based on each target image, and generates the detection result based on the detection region feature information and the detection region texture information;

[0020] Send the detection result corresponding to the target detection region to the user.

[0021] According to a fifth aspect of the embodiments of the present specification, there is provided an image processing device, which includes:

[0022] A receiving module, configured to receive an image processing task, where the image processing task carries a plurality of target images corresponding to a target detection area, and the target image processing task is used to detect whether the target detection area is abnormal;

[0023] A detection module, configured to input the plurality of target images into an image processing model to obtain a detection result corresponding to the target detection area, where the image processing model generates detection area feature information and detection area texture information based on each target image, and generates the detection result based on the detection area feature information and the detection area texture information.

[0024] According to a sixth aspect of the embodiments of the present specification, a computing device is provided, including:

[0025] A memory and a processor;

[0026] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.

[0027] According to a seventh aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above method are implemented.

[0028] According to an eighth aspect of the embodiments of the present specification, a computer program is provided. When the computer program is executed on a computer, the computer is made to execute the steps of the above method.

[0029] The image processing method provided by an embodiment of the present specification receives an image processing task, where the image processing task carries a plurality of target images corresponding to a target detection area, and the target image processing task is used to detect whether the target detection area is abnormal; inputs the plurality of target images into an image processing model to obtain a detection result corresponding to the target detection area, where the image processing model generates detection area feature information and detection area texture information based on each target image, and generates the detection result based on the detection area feature information and the detection area texture information.

[0030] Through the method provided by the embodiments of the present specification, in the process of recognizing and processing a plurality of target images, detection area feature information is generated based on each target image, then detection area texture information is extracted from the detection area feature information, and the final image coding feature is generated based on the detection area feature information and the detection area feature information. Before image decoding, distilled feature information and classification feature information are added, so that during image decoding, the distilled feature and the classification feature are referenced, making the final detection result more accurate. Description of the Drawings

[0031] Figure 1 is an architecture diagram of an image processing system provided by an embodiment of this specification;

[0032] Figure 2 is a flowchart of an image processing method provided by an embodiment of this specification;

[0033] Figure 3 is a schematic diagram of the model structure of an image processing model provided by an embodiment of this specification;

[0034] Figure 4 is a flowchart of a CT image processing method provided by an embodiment of this specification;

[0035] Figure 5 is a flowchart of a training method for an image processing model provided by an embodiment of this specification;

[0036] Figure 6 is a flowchart of another image processing method provided by an embodiment of this specification;

[0037] Figure 7 is a flowchart of the processing procedure of an image processing method applied to the scenario of detecting fatty liver provided by an embodiment of this specification;

[0038] Figure 8 is a schematic diagram of the structure of an image processing apparatus provided by an embodiment of this specification;

[0039] Figure 9 is a block diagram of the structure of a computing device provided by an embodiment of this specification. Detailed Description of the Embodiments

[0040] In the following description, numerous specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.

[0041] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.

[0042] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0043] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0044] First, the noun terms involved in one or more embodiments of this specification are explained.

[0045] CT (Computed Tomography): Computed Tomography, which uses a precisely collimated X-ray beam and a highly sensitive detector to make successive cross-sectional scans around a certain part of the human body. It has the characteristics of fast scanning time and clear images and can be used for the examination of various diseases.

[0046] CNN (convolutional neural network): Convolutional Neural Network, which is a commonly used deep learning model in the field of computer vision.

[0047] NC CT (non-contrast CT): Plain CT, which is a common scan in CT and is generally used for routine examinations of certain parts.

[0048] CAD (computer aided diagnosis): Computer Aided Diagnosis, which refers to the use of imaging, medical image processing technology, and other possible physiological and biochemical means, combined with the analysis and calculation of a computer, to assist in detecting lesions and improving the accuracy of diagnosis.

[0049] With the continuous development of computer technology, various learning models have gradually been applied to the prediction of various application scenarios. Correspondingly, deep learning models have also achieved remarkable success in the task of computer-aided diagnosis (CAD) of medical images. Pathological analysis and classification from medical images are important topics in computer-aided diagnosis.

[0050] Currently, when detecting organs, computed tomography (CT) is usually used. CT provides an operator-independent, standardized, and general method for quantifying fat content, making it an ideal choice for screening diseased organs in various clinical situations. Taking the detection of fatty liver as an example, radiologists usually measure the CT attenuation in Hounsfield Units (HU) throughout the liver and draw one or more target regions on the representative parenchymal part of the liver. Doctors also use measurements such as the liver-spleen attenuation difference and the liver-spleen attenuation ratio to compare the liver attenuation. However, these processes are subjective. Moreover, during the process of image recognition, some subtle differences and changes in organs cannot be accurately recognized.

[0051] Based on this, in this specification, an image processing method is provided. This specification also relates to an image processing device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.

[0052] See Figure 1 , Figure 1 shows an architecture diagram of an image processing system provided by an embodiment of this specification. The image processing system may include a client 100 and a server 200;

[0053] The client 100 is used to send an image processing task to the server 200. Among them, the image processing task carries multiple target images corresponding to the target detection region, and the target image processing task is used to detect whether the target detection region is abnormal;

[0054] The server 200 is used to input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection region. Among them, the image processing model generates detection region feature information and detection region texture information based on each target image, and generates the detection result based on the detection region feature information and the detection region texture information; send the detection result of the image processing task to the client 100;

[0055] The client 100 is also used to receive the detection result of the image processing task sent by the server 200.

[0056] Applying the solution of the embodiment of this specification, receive an image processing task, where the image processing task carries multiple target images corresponding to the target detection region, and the target image processing task is used to detect whether the target detection region is abnormal; input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection region, where the image processing model generates detection region feature information and detection region texture information based on each target image, and generates the detection result based on the detection region feature information and the detection region texture information.

[0057] Through the solution provided by the embodiments of this specification, the detection area feature information is generated based on the target image, and at the same time, the detection area texture information in the target detection area is extracted. Finally, the final detection result is generated based on the detection area feature information and the detection area texture information. In the process of the image processing model processing the target image, not only the detection area feature information is considered, but also the detection area texture information is further extracted from the detection area feature information. Since the texture information can more accurately represent the state of the target detection area, by extracting and analyzing the features of the texture information of the target detection area, the features of the target detection area can be enriched, thereby improving the accuracy of the image processing model.

[0058] In practical applications, the image processing system may include multiple clients 100 and a server 200. Among them, the client 100 can be called the edge device, and the server 200 can be called the cloud device. Communication connections can be established between multiple clients 100 through the server 200. In the image processing scenario, the server 200 is used to provide image processing services between multiple clients 100. Multiple clients 100 can be used as the sender or receiver respectively, and communicate through the server 200.

[0059] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In the image processing scenario, it can be that the user publishes a data stream to the server 200 through the client 100, and the server 200 generates a detection result based on the data stream and pushes the detection result to other clients that have established communication.

[0060] Among them, a connection is established between the client 100 and the server 200 through the network. The network provides the medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The data transmitted by the client 100 may need to be processed such as encoded, transcoded, compressed, etc. before being published to the server 200.

[0061] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application), or a cloud application, etc. The client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0062] The server 200 can include servers that provide various services. For example, a server that provides communication services for multiple clients, or a server for background training that provides support for models used on the client, or a server that processes data sent by the client, etc. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.

[0063] It is worth noting that the image processing method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client can also have a similar function to the server, so as to execute the image processing method provided in the embodiments of this specification. In other embodiments, the image processing method provided in the embodiments of this specification can also be jointly executed by the client and the server.

[0064] See Figure 2 , Figure 2 which shows the flowchart of an image processing method provided in an embodiment of this specification, specifically including the following steps:

[0065] Step 202: Receive an image processing task, where the image processing task carries multiple target images corresponding to a target detection area, and the target image processing task is used to detect whether the target detection area is abnormal.

[0066] In practical applications, the image processing task sent by the user can be received through the server or the client.

[0067] Specifically, the image processing task is a task for detecting whether the target detection area is abnormal. The image processing task carries multiple target images corresponding to the target detection area. Further, the target detection area can be understood as a partition for predicting whether an abnormality occurs. For example, the target detection area can be any organ in the human body, such as the liver, spleen, lung, stomach, etc. By predicting whether the target detection area is abnormal, the state of the object to be detected can be further assisted in judgment based on the prediction result, thereby helping to determine the state of the object to be detected.

[0068] It should be noted that in one or more embodiments of this specification, the image processing task can be applied to the recognition of various medical images, and whether an abnormality occurs in the target detection area in the medical image can be judged according to the image features. Exemplarily, in the application scenario of pulmonary fibrosis detection, whether pulmonary fibrosis has occurred in the lungs can be predicted based on the medical images of the lungs; in the application scenario of fatty liver detection, whether fatty liver has occurred in the liver can be predicted based on the medical images of the liver; thus helping doctors to assist in judging whether an abnormality occurs in the target detection area, and further facilitating subsequent treatment.

[0069] In a specific embodiment provided in this specification, in the application scenario of detecting pulmonary fibrosis, the obtained target images are the images of the lungs. Specifically, the obtained multiple target images are the CT images of the lungs. The multiple target images can form a 3D image of the lungs. By obtaining multiple CT images of the lungs and performing image detection processing on the multiple CT images, it is detected whether pulmonary fibrosis occurs in the lungs.

[0070] By receiving the image processing task, the terminal can use the multiple target images corresponding to the target detection area carried in the image processing task as input to detect whether there is an abnormality in the target detection area.

[0071] In practical applications, there may be multiple regions in the target image, and it is necessary to pre-determine the target detection region in multiple regions of the target image. For example, in medical images, if the CT image taken is a chest CT, the CT image may include the heart, spleen, liver, etc.; if the CT image taken is an abdominal CT, the CT image may include the liver, spleen, etc. In different CT images, the positions of the target detection regions are different. Therefore, in a specific embodiment provided in this specification, an image segmentation recognition model is used, and through supervised or semi-supervised model training, the image segmentation recognition model is trained to enable it to have the ability to determine the target detection region from the target image. In the process of actual application, multiple target images and target detection region identifiers are input into the image segmentation recognition model, and the image segmentation recognition model can determine the target detection region corresponding to the target detection region identifier from multiple target images, thereby improving the accuracy of segmenting the target detection region from multiple target images.

[0072] Step 204: Input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection region, where the image processing model generates detection region feature information and detection region texture information based on each target image, and generates the detection result based on the detection region feature information and the detection region texture information.

[0073] In practical applications, after receiving an image processing task, multiple target images carried by the image processing task are obtained from the task, and the multiple target images are input into an image processing model for processing, and a detection result corresponding to the target detection region output by the image processing model can be obtained. The detection result specifically includes whether the target detection region has an abnormality, the degree information of the abnormality, and so on.

[0074] Specifically, the image processing model can extract detection region feature information corresponding to the target detection region and detection region texture information corresponding to the target detection region based on the input multiple target images. Among them, the detection region feature information corresponding to the target detection region specifically refers to the local image feature information corresponding to the target detection region, and the detection region feature information is obtained after image feature extraction based on multiple target images; the detection region texture information specifically refers to the texture information corresponding to the target detection region. After obtaining the detection region feature information and the detection region texture information corresponding to the target detection region, the final detection result is generated based on these two pieces of information.

[0075] In the embodiment provided in this specification, in addition to collecting the detection region feature information corresponding to the target detection region, the detection region texture information corresponding to the target detection region is also collected. By extracting the texture information in the target detection region, the features of the target detection region are enriched, thereby improving the detection accuracy of the image processing model.

[0076] Specifically, in a specific embodiment provided in this specification, the image processing model includes an encoder, a decoder, and a classifier;

[0077] Inputting the multiple target images into the image processing model to obtain the detection result corresponding to the target detection area, including S2042 - S2046:

[0078] S2042. Input the multiple target images into the encoder to obtain image encoding features, where the image encoding features are determined based on the detection area feature information and the detection area texture information corresponding to each target image.

[0079] Specifically, the encoder is used to extract the feature information of the target detection area. The encoder encodes the input multiple target images, extracts the image feature information in each target image, and obtains the image encoding features corresponding to the target detection area.

[0080] It should be noted that in one or more specific embodiments of this specification, the image encoding features corresponding to the target detection area are specifically determined according to the detection area feature information and the detection area texture information corresponding to the target detection area. The detection area feature information is obtained by feature extraction from each target image, and the detection area texture information is extracted from the detection area feature information.

[0081] Specifically, the encoder includes an image encoding unit, a texture encoding unit, and a feature fusion unit;

[0082] Inputting the multiple target images into the encoder to obtain image encoding features includes:

[0083] Input the multiple target images into the image encoding unit to obtain the detection area feature information corresponding to the multiple target images;

[0084] Input the detection area feature information into the texture encoding unit to obtain the detection area texture information corresponding to the detection area feature information;

[0085] Input the detection area feature information and the detection area texture information into the feature fusion unit to obtain image encoding features.

[0086] Specifically, the encoder includes an image encoding unit, a texture encoding unit, and a feature fusion unit. The image encoding unit is used to extract image features of the target image. The texture encoding unit is used to extract texture features from the extracted image feature information. The feature fusion unit is used to fuse the image features and the texture features to finally generate the image encoding features output by the encoder.

[0087] Further, the image encoding unit can specifically be based on the feature extraction unit of the 3D convolutional network. After obtaining multiple target images, it identifies the 3D image region of the target detection region in the multiple target images. Where H represents the height information of the target detection region, W represents the width information of the target detection region, and D represents the depth information of the target detection region.

[0088] Input the 3D image region of the target detection region into the image encoding unit (Patch Encoder 3D CNN). In the image encoding unit, based on the preset patch size, divide the 3D image region into several image patches, and then extract the feature information corresponding to each image patch through the 3D convolutional network respectively, so as to obtain the detection region feature information output by the image encoding unit.

[0089] In a specific embodiment provided in this specification, taking the dimension of the feature information corresponding to each image patch as 512 dimensions as an example, if the 3D image region of the target detection region is divided into P image patches, after each image patch is processed by the 3D convolutional network, the dimension of the feature information corresponding to each image patch is 1*512, then the dimension of the detection region feature information composed of P image patches is P*512 dimensions.

[0090] After obtaining the detection region feature information, input the detection region feature information into the texture encoding unit. Extract the texture information of the detection region feature information in the texture encoding unit to obtain the detection region texture information. There is a texture template stored in the texture encoding unit. The essence of the texture encoding unit is to calculate the similarity between the detection region feature information and the texture template, and assemble the detection region feature information into the texture template according to the similarity to obtain the final detection region texture information.

[0091] Finally, process the obtained detection region feature information and detection region texture information through the feature fusion unit to obtain the image encoding feature output by the feature fusion unit. In practical applications, the detection region feature information and the detection region texture information can be fused by means of weighted summation, or by means of direct summation, or by means of feature splicing. In the embodiment provided in this specification, it is preferably fused by means of direct summation. For example, if the dimension of the detection region feature information is P*512 dimensions and the dimension of the detection region texture information is also P*512 dimensions, then after the two are fused, the dimension of the obtained image encoding feature is still P*512 dimensions.

[0092] In an encoder, by extracting the detection region feature information and the detection region texture information corresponding to the target detection region, in subsequent processing, in addition to decoding the image feature information from the detection region feature information, the texture information of the target detection region can also be extracted from the detection region texture information, and the information of the target detection region is characterized from the perspective of texture features, thereby enriching the depth representation of the target detection region and improving the accuracy of analysis.

[0093] In practical applications, in addition to being able to refer to the texture features corresponding to the target detection region, the biomarker information determined from the target image can also be referred to. Specifically, in another specific embodiment provided in this specification, the method further includes:

[0094] Obtaining the biomarker information corresponding to multiple target images;

[0095] Inputting the multiple target images and the biomarker information into an image processing model to obtain the detection result corresponding to the target detection region.

[0096] The biomarker information specifically refers to the information for marking the target detection region and the reference detection region in the target image. In subsequent image processing, the multiple target images and the biomarker information are input into the image processing model together, so that the image processing model outputs the detection result corresponding to the target detection region.

[0097] In a specific embodiment provided in this specification, taking the target detection region as the liver as an example, the reference detection region can be the spleen beside the liver. By performing a series of biometric evaluations on the liver and the spleen, the biomarker information is obtained. Specifically, by calculating the histograms of the CT values of the liver and the spleen, the average attenuation values of the CT values of the liver and the spleen are obtained. At the same time, the attenuation ratios and attenuation differences of the CT values of the liver and the spleen are also calculated; in addition, the regional attenuation of the liver and the spleen, as well as the meta-information of the target to be detected (such as gender, age, etc.) are respectively evaluated. Based on the above information, the biomarker information I can be obtained BIO 。

[0098] S2044. Input the image encoding feature into the decoder to obtain the image decoding feature corresponding to the image encoding feature.

[0099] Among them, the decoder is used to decode the image encoding feature, so as to obtain the image decoding feature corresponding to the image encoding feature from the image encoding feature. Further, in the decoder, multiple decoding layers are used, and the multi-head self-attention mechanism is used in each decoding layer. The image encoding feature is decoded through the multi-head self-attention mechanism to obtain the corresponding image decoding feature. Preferably, there are 4 decoding layers in the decoder. Each decoding layer uses the multi-head self-attention mechanism.

[0100] In one or more specific embodiments provided in this specification, specifically, inputting the image coding feature into the decoder to obtain the image decoding feature corresponding to the image coding feature includes:

[0101] Adding distilled feature information and classification feature information to the image coding feature to obtain a to-be-processed image coding feature;

[0102] Inputting the to-be-processed image coding feature into the decoder to obtain an image decoding feature.

[0103] In practical applications, the image processing model is a pre-trained machine learning model. Limited by the lack of training data, before decoding an image, the image processing model provided in the embodiments of this specification adds distilled feature information and classification feature information to the image coding feature. Among them, the distilled feature information is used to compare with non-standard labels, and the classification feature information is used to compare with standard labels, thereby improving the efficiency in the model training stage.

[0104] Based on this, after obtaining the image coding feature, distilled feature information and classification feature information are added to the image coding feature, which is used to predict the final detection result according to the distilled feature information and classification feature information in the subsequent processing process. Further, the distilled feature information can be added to the beginning of the image coding feature, and the classification feature information can be added to the end of the image coding feature.

[0105] For example, taking the image coding feature as [E0, E1... E P as an example, adding distilled feature information E dis and classification feature information E cls to this image coding feature to obtain the to-be-processed image coding feature [E dis , E0, E1... E P , E cls , and then inputting the to-be-processed image coding feature into the decoder for decoding, finally obtaining the image decoding feature. Taking the to-be-processed image coding feature [E dis , E0, E1... E P , E cls as an example, after being processed by the decoder, the obtained image decoding feature is [D dis , D0, D1... D P , D cls .

[0106] In another specific embodiment provided in this specification, when biomarker information is input into the image processing model, inputting the to-be-processed image coding feature into the decoder to obtain the image decoding feature includes:

[0107] Splice the biomarker information to the image encoding feature to be processed to obtain a spliced image encoding feature;

[0108] Input the spliced image encoding feature into the decoder to obtain an image decoding feature.

[0109] In the above steps, the situation of inputting biomarker information into the image processing model is also mentioned. In this case, the biomarker information is added to the image encoding feature to be processed to obtain a spliced image encoding feature. Further, for example, taking the image encoding feature to be processed [E dis 、E0、E1……E P 、E cls and the biomarker information I BIO as an example, the biomarker information I BIO is converted into the corresponding biomarker feature information E BIO , and then the biomarker feature information and the image encoding feature to be processed are spliced to obtain a spliced image encoding feature [E dis 、E0、E1……E P 、E BIO 、E cls . Input the spliced image encoding feature into the decoder to obtain an image decoding feature [D dis 、D0、D1……D P 、D BIO 、D cls .

[0110] S2046. Input the image decoding feature into the classifier to obtain the detection result corresponding to the target detection area.

[0111] After obtaining the image decoding feature, input the image decoding feature into the classifier, and the classifier classifies the image decoding feature to obtain the final detection result corresponding to the target detection area.

[0112] In a specific embodiment provided in this specification, when the target image is input into the image processing model, the image decoding feature is [D dis 、D0、D1……D P 、D cls . Input the image decoding feature into the classifier for processing, and the classifier classifies according to the distilled decoding feature and the classification decoding feature to obtain the final detection result.

[0113] In another specific embodiment provided in this specification, when the target image and biomarker information are input into the image processing model, the image decoding feature is [D dis 、D0、D1……D P 、D BIO 、Dcls , the image decoding features are input into a classifier for processing, and the classifier classifies according to the distilled decoding features and the classification decoding features, so as to obtain the final detection result.

[0114] Specifically, inputting the image decoding features into the classifier to obtain the detection result corresponding to the target detection area includes:

[0115] Inputting the image decoding features into the classifier to obtain a first classification result corresponding to the distilled feature information and a second classification result corresponding to the classification feature information;

[0116] Generating the detection result according to the first classification result and the second classification result.

[0117] In the method provided in this specification, after the image decoding features are input into the classifier, the classifier classifies according to the distilled feature information and the classification feature information, respectively obtains a first classification result corresponding to the distilled feature information and a second classification result corresponding to the classification feature information, and then performs weighted averaging on the first classification result and the second classification result to obtain the final detection result.

[0118] See Figure 3 , Figure 3 shows a schematic diagram of the model structure of an image processing model provided in an embodiment of this specification. As Figure 3 shown, the image processing model includes an encoder, a decoder, and a classifier. The encoder includes an image encoding unit, a texture encoding unit, and a feature fusion unit. The decoder includes four decoding layers based on the multi-headed self-attention mechanism (MSA).

[0119] Multiple target images are input into the image processing model. After the image encoding unit extracts image features, detection area feature information is obtained. The detection area feature information is input into the texture encoding unit for texture feature extraction to obtain detection area texture information. The detection area feature information and the detection area texture information are input into the feature fusion unit to obtain image encoding features as [E0, E1... E P .

[0120] Biological markers are performed on multiple target images to obtain biological marker information I BIO , which is input into the image processing model and, after embedding processing, biological marker feature information E BIO is obtained. The image encoding features as [E0, E1... E P , the biological marker feature information E BIO are concatenated, and the distilled feature information D dis and the classification feature information Dcls , generate the spliced image encoding features [E dis , E0, E1... E P , E BIO , E cls , and input the spliced image encoding features [E dis , E0, E1... E P , E BIO , E cls into the decoder for decoding to obtain the image decoding features as [D dis , D0, D1... D P , D BIO , D cls , and input the D dis , D0, D1... D P , D BIO , D cls in the image decoding features as [D dis , D cls vectors into the classifier for classification to obtain the first classification result corresponding to D dis and the second classification result corresponding to D cls . Finally, by means of weighted summation, the first classification result and the second classification result are fused to obtain the final detection result.

[0121] In a specific embodiment provided in this specification, in the process of recognizing and processing multiple target images, detection region feature information is generated based on each target image, and then detection region texture information is extracted from the detection region feature information. Based on the detection region feature information and the detection region feature information, the final image encoding features are generated. Before image decoding, distilled feature information and classification feature information are added, so that the distilled features and classification features are referred to during the image decoding process, making the final detection result more accurate.

[0122] In the method provided in an embodiment of this specification, the image processing model is a trained machine learning model. Further, the image processing model is a supervised trained model. Specifically, the image processing model is obtained through the following S2062 - S2068 training:

[0123] S2062. Obtain training sample pairs, where the training sample pairs include multiple training sample images corresponding to the target detection region and the sample detection results of the target detection region, and the sample detection results include standard sample detection results or reference sample detection results.

[0124] Specifically, the training method of the image processing model provided in this specification uses supervised training, which includes training sample pairs. Each training sample pair includes multiple training sample images for the target detection area and sample detection results. It should be noted that in the training method of the image processing model provided in this specification, the sample detection results include two categories, one is the standard sample detection result, and the other is the reference sample detection result. Among them, the standard sample detection result is a verified sample detection result, and the reference sample detection result is a sample detection result predicted by relevant technicians based on experience.

[0125] For example, taking the image processing model for predicting whether there is fatty liver in the liver as an example for explanation, in the model training stage of the image processing model, multiple training sample pairs form a training sample set. The training sample set includes two training sample subsets. The first training sample subset has 680 sample pairs, which are the CT images of the livers of 680 users and the pathological results verified by pathology; the second training sample subset has 1103 sample pairs, including the CT images of the livers of 1103 users and the prediction results predicted by doctors. Among them, the pathological results verified by pathology are the standard sample detection results, and the prediction results predicted by doctors are the reference sample detection results. Both the standard sample detection result and the reference sample detection result are sample detection results.

[0126] In the method provided in this specification, multiple training sample pairs are input into the image processing model according to a preset batch. The training sample image in the training sample pair and the corresponding sample detection result are positive sample pairs, and the sample detection results in other training sample pairs are negative sample pairs. For example, still taking the above example of predicting whether there is fatty liver in the liver, one training batch has 64 training sample pairs. For training sample pair 1, training sample image 1 and sample detection result 1 are positive sample pairs, and training sample image 1 and other sample detection results are negative sample pairs; for training sample pair 2, training sample image 2 and sample detection result 2 are positive sample pairs, and training sample image 2 and other sample detection results are negative sample pairs...

[0127] S2064. Input the multiple training sample images into the image processing model to obtain the predicted detection result corresponding to the target detection area.

[0128] After obtaining the training sample pairs, according to the preset training batches, input the training sample images of multiple training sample pairs into the image processing model. At this time, the image processing model is an untrained image processing model. In the image processing model, first generate the detection region feature information corresponding to the target detection region according to each training sample image, then obtain the detection region texture information from the detection region feature information, then generate the image encoding feature according to the detection region feature information and the detection region texture information, and then add the distillation feature information and the classification feature information to the image encoding feature to obtain the image encoding feature to be processed.

[0129] In a specific embodiment provided in this specification, the biomarker information of the target detection region and the reference detection region corresponding to the target detection region will also be extracted, and the biomarker information will be input into the image processing model and spliced with the image encoding feature to be processed to generate the spliced image encoding feature. Then it is input into the decoder for decoding processing, and finally input into the classifier to obtain the predicted detection result corresponding to the target detection region of the final model output.

[0130] In the method provided in this specification, the model structure of the image processing model is the same as that of the above-mentioned image processing model. Regarding the data processing process of the training sample image in the untrained image processing model, refer to the data processing process of the target image in the image processing model above, and details will not be repeated here.

[0131] S2066. Calculate the model loss value according to the predicted detection result and the sample detection result.

[0132] After obtaining the predicted detection result of the image processing model, the model loss value can be calculated according to the predicted detection result and the sample detection result. In the method provided in this specification, there are many methods for calculating the model loss value, such as the cross-entropy loss function, the maximum loss function, the average loss function, etc. In this specification, the specific manner of the loss function is not limited and shall be subject to actual applications.

[0133] In another specific embodiment provided in this specification, the predicted detection result includes the first predicted classification result corresponding to the distillation feature information and the second predicted classification result corresponding to the classification feature information;

[0134] Calculating the model loss value according to the predicted detection result and the sample detection result includes:

[0135] Calculating the first loss value according to the reference sample detection result and the first predicted classification result; or

[0136] Calculating the second loss value according to the standard sample detection result and the first predicted classification result;

[0137] Specifically, in the method provided in this specification, distilled feature information and classification feature information are added to the encoded feature information in the image processing model, where the distilled feature information corresponds to the detection result of the reference sample, and the classification feature information corresponds to the detection result of the standard sample.

[0138] In the model application stage, the detection result output by the model is determined according to the first predicted classification result corresponding to the distilled feature information and the second predicted classification result corresponding to the classification feature information. In the model training stage, by adjusting the weights, when the training sample pair is the detection result of the standard sample, the second loss value is calculated according to the detection result of the standard sample and the second predicted classification result; when the training sample pair is the detection result of the reference sample, the first loss value is calculated according to the detection result of the reference sample and the first predicted classification result. That is, in the model training stage, if the training sample pair is the detection result of the standard sample, the loss value is calculated with the predicted result of the classification feature information; if the training sample pair is the detection result of the reference sample, the loss value is calculated with the predicted result of the distilled feature information.

[0139] S2068. Adjust the model parameters of the image processing model according to the model loss value until the model training stop condition is reached.

[0140] After obtaining the model loss value, the model parameters of the image processing model can be adjusted according to the model loss value. Specifically, the model loss value can be backpropagated to update the model parameters of the image processing model in turn.

[0141] Correspondingly, adjusting the model parameters of the image processing model according to the model loss value includes:

[0142] Adjust the model parameters of the image processing model according to the first loss value and the second loss value.

[0143] In a specific embodiment provided in this specification, there may be both the first loss value and the second loss value in the model training of the same batch. After the training of the training data of the same batch is completed, the model parameters of the image processing model are adjusted according to the multiple first loss values and / or multiple second loss values of this batch.

[0144] After adjusting the model parameters, the above steps can be continued to repeat the training of the image processing model until the training stop condition is reached. In practical applications, the training stop conditions of the image processing model include:

[0145] The model loss value is less than the preset threshold; and / or

[0146] The number of training rounds reaches the preset number of training rounds.

[0147] Specifically, during the training of the image processing model, the training stop condition of the model can be set to that the model loss value is less than a preset threshold, or the training stop condition can be set to that the number of training rounds is a preset number of training rounds, for example, 10 rounds of training. In this specification, the preset threshold of the loss value and / or the preset number of training rounds are not specifically limited and shall be subject to actual applications.

[0148] In the method provided in the embodiments of this specification, during the training of the image processing model, the detection results of verified standard samples are used, and the detection results of reference samples predicted by relevant technicians based on experience are also used. This enriches the number of the training sample set and provides a data basis for the training of the image processing model. During the training process, the first predicted classification result corresponding to the distilled feature information and the second predicted classification result corresponding to the classification feature information are used to train with the training sample set, which enriches the processing ability of the image processing model, enables it to have richer generalization ability, and thus further improves the prediction accuracy of the image processing model.

[0149] See Figure 4 , Figure 4 shows a flowchart of a CT image processing method provided by an embodiment of this specification, which specifically includes the following steps:

[0150] Step 402: Receive a CT image processing task, where the CT image processing task carries multiple CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormality in the target detection area.

[0151] Step 404: Input the multiple CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, where the CT image processing model generates detection area feature information and detection area texture information based on each CT image, and generates the detection result based on the detection area feature information and the detection area texture information.

[0152] It should be noted that the implementation manners of steps 402 and 404 are the same as those of steps 202 - 204 above, and the embodiments of this specification will not elaborate further.

[0153] Exemplarily, taking the liver as the target detection area, to detect whether there is fatty liver in the liver, a CT image processing task is received. The CT image processing task includes multiple CT images corresponding to the liver of the target user. The multiple CT images can form a 3D map of the liver of the target user, and this CT image processing task is used to detect whether the target user has fatty liver.

[0154] Applying the method of the embodiments of this specification, the CT image processing model is the image processing model in the above embodiments. The model structure of the CT image processing model is the same as that of the image processing model in the above embodiments, which will not be elaborated here. By inputting multiple CT images corresponding to the liver region into the CT image processing model, the detection result corresponding to the liver region output by the image processing model can be obtained, thereby realizing the automatic detection of whether there is fatty liver in the liver and the degree of fatty liver.

[0155] In the CT image processing method provided by the embodiments of this specification, in the CT image processing model, detection region feature information is generated based on the CT image, and then detection region texture information is extracted from the detection region feature information. Based on the detection region feature information and the detection region feature information, the final image coding feature is generated. Before image decoding, distilled feature information and classification feature information are added, so that during the image decoding process, the distilled feature and the classification feature are referred to, making the final detection result more accurate.

[0156] See Figure 5 , Figure 5 shows a flowchart of a training method for an image processing model provided by an embodiment of this specification, which is applied to a cloud-side device and specifically includes the following steps:

[0157] Step 502: Obtain a training sample pair, where the training sample pair includes multiple training sample images corresponding to the target detection region and the sample detection result of the target detection region, and the sample detection result includes a standard sample detection result or a reference sample detection result.

[0158] Step 504: Input the multiple training sample images into the image processing model to obtain the predicted detection result corresponding to the target detection region.

[0159] Step 506: Calculate the model loss value according to the predicted detection result and the sample detection result.

[0160] Step 508: Adjust the model parameters of the image processing model according to the model loss value until the model training stop condition is reached, and obtain the model parameters of the image processing model.

[0161] Step 510: Send the model parameters of the image processing model to the end-side device.

[0162] It should be noted that the implementation manners of steps 502 to 508 are the same as those of S2062 - S2068 above, and the embodiments of this specification will not elaborate further.

[0163] In practical applications, since training a model requires a large amount of data and good computing resources, edge devices may not have the corresponding processing capabilities. Therefore, the process of model training can be implemented on cloud devices. After obtaining the model parameters of the image processing model, the cloud devices can also send the model parameters to the edge devices. The edge devices can construct an image processing model locally based on the model parameters of the image processing model and further use the image processing model to perform image processing.

[0164] In the method provided in the embodiments of this specification, during the process of training the image processing model, the verified standard sample detection results are used, and the reference sample detection results predicted by relevant technicians based on experience are also used. This enriches the quantity of the training sample set and provides a data basis for the training of the image processing model. During the training process, the first predicted classification result corresponding to the distilled feature information and the second predicted classification result corresponding to the classification feature information are used to train with the training sample set, which enriches the processing capabilities of the image processing model and enables it to have more abundant generalization capabilities, thereby further improving the prediction accuracy of the image processing model.

[0165] See Figure 6 , Figure 6 which shows a flowchart of an image processing method provided by an embodiment of this specification, specifically including the following steps:

[0166] Step 602: Receive an image processing request sent by a user, where the image processing request includes an image processing task, the image processing task carries multiple target images corresponding to a target detection area, and the target image processing task is used to detect whether the target detection area is abnormal.

[0167] Step 604: Input the multiple target images into the image processing model to obtain a detection result corresponding to the target detection area, where the image processing model generates detection area feature information and detection area texture information based on each target image, and generates the detection result based on the detection area feature information and the detection area texture information.

[0168] Step 606: Send the detection result corresponding to the target detection area to the user.

[0169] It should be noted that the specific implementation manners of steps 602 - 604 are the same as those of the above steps 202 - 204, and will not be elaborated in the embodiments of this specification.

[0170] In this embodiment, an image processing request sent by a user is received. The image processing request includes an image processing task. After the detection result is obtained through the image processing method of the above embodiment, the detection result needs to be returned to the user so that the user can perform corresponding subsequent processing based on the detection result.

[0171] In the method provided by the embodiments of this specification, in the process of recognizing and processing multiple target images, detection region feature information is generated based on each target image, and then detection region texture information is extracted from the detection region feature information. Based on the detection region feature information and the detection region feature information, the final image coding feature is generated. Before image decoding, distilled feature information and classification feature information are added, so that during the image decoding process, the distilled feature and the classification feature are referred to, making the final detection result more accurate.

[0172] The following combines the attached Figure 7 , taking the application of the image processing method provided in this specification in detecting the presence of fatty liver as an example, to further illustrate the image processing method. Among them, Figure 7 FIG. shows the processing flow chart of an image processing method provided by an embodiment of this specification, which specifically includes the following steps:

[0173] Step 702: Receive a CT image processing task, where the CT image processing task carries multiple CT images corresponding to the liver.

[0174] Step 704: Obtain biomarker information corresponding to the multiple CT images, where the biomarker information includes liver CT information, spleen CT information, and user attribute information.

[0175] Step 706: Input the multiple CT images and the biomarker information into the CT image processing model to obtain the detection result corresponding to the liver.

[0176] Specifically, when the CT image processing task is to detect whether the target user has fatty liver and the degree of fatty liver.

[0177] The CT image processing model is pre-trained in advance, and the source of the training data comes from 680 research subjects diagnosed by pathology and 1103 research subjects predicted by doctors. Among the 680 research subjects diagnosed by pathology, there are 203 healthy research subjects, 150 mild fatty liver research subjects, 138 moderate fatty liver research subjects, and 89 severe fatty liver research subjects. Among the 1103 research subjects predicted by doctors, there are 438 healthy research subjects, 307 mild fatty liver research subjects, 112 moderate fatty liver research subjects, and 246 severe fatty liver research subjects.

[0178] Using multiple CT scanners, under the same contrast conditions, non-enhanced CT scan images of the chest and abdomen of the above research subjects were collected respectively. The collected images were used as training sample images, and the corresponding diagnostic results of each research subject were used as sample detection results. Thus, the CT image processing model was trained. For the specific training process, refer to the description in the above embodiments and will not be elaborated here.

[0179] After the model was trained, a validation dataset was selected. The data source of the validation dataset came from 226 validation research subjects as the validation set UNIFESP-tr. Through statistical analysis of the data of the validation set at three different levels: Mild, Moderate, and Severe, AUC represents the aggregated statistical data, and ACC represents the accuracy value. Refer to Table 1 below.

[0180] Table 1

[0181]

[0182] As shown in Table 1, in a specific embodiment provided in this specification, three ablation verification methods were used, namely the biometric-based method, the deep learning branch method, and the hybrid configuration method.

[0183] In the biometric-based method, three configurations were tested, namely "MAL", "MALRO", "MALRO + ". Among them, "MAL" is a configuration that only uses the average HU of the liver and user attribute information for logical ordinal regression training; "MALRO" adds sampling of the liver target area on the basis of "MAL"; "MALRO + " adds spleen-related biomarker information on the basis of "MALRO".

[0184] In the deep learning branch method, three configurations were also tested, namely "3DN", "3DNT", "3DNT + ". Among them, "3DN" is a basic configuration that only uses 3D-ResNet34; "3DNT" adds a multi-head self-attention mechanism on the basis of "3DN"; "3DNT + " adds texture encoding on the basis of "3DNT".

[0185] In the hybrid configuration method, four configurations were tested, namely "Bio-3DNT", "Bio-3DNT T ", "Bio-3DNT R ", "Bio-3DNT TR", where "Bio-3DNT" is a configuration that combines biometrics and deep learning. In this case, a teacher model "Bio-3DNT" is also set up. T ", a radiologist knowledge model "Bio-3DNT" R ", and a teacher-radiologist knowledge combination model "Bio-3DNT" TR ". "Bio-3DNT" is selected as the teacher model. The sample labels of the teacher model use only the standard sample detection results, the sample labels of the radiologist instruction model use only the reference sample detection results, and the sample labels of the teacher-radiologist knowledge combination model "Bio-3DNT" TR " use both the standard sample detection results and the reference sample detection results.

[0186] Referring to Table 1, the method of introducing deep learning has greatly improved the overall accuracy compared to the biometric-based method. Adding the multi-head self-attention mechanism and texture encoding to the deep learning method further improves the performance. When deep learning and biometrics are combined, the overall performance is further improved. In the teacher-radiologist knowledge combination model "Bio-3DNT" TR " that combines the standard sample detection results and the reference sample detection results, the model performance is jointly improved by refining the radiologist's instructions and the teacher model.

[0187] Through the method provided in the embodiments of this specification, the trained image processing model also refers to the biometric information corresponding to the target image in actual applications. Generate detection region feature information based on each target image, then extract detection region texture information from the detection region feature information, and generate the final image coding feature based on the detection region feature information, detection region feature information, and biometric information. Before image decoding, distilled feature information and classification feature information are added, so that the distilled feature and classification feature are referred to during the image decoding process, making the final detection result more accurate.

[0188] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing device. Figure 8 Shows a schematic structural diagram of an image processing device provided by an embodiment of this specification. As Figure 8 shown, the device includes:

[0189] A receiving module 802, configured to receive an image processing task, where the image processing task carries a plurality of target images corresponding to a target detection region, and the target image processing task is used to detect whether the target detection region is abnormal;

[0190] The detection module 804 is configured to input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection region, where the image processing model generates detection region feature information and detection region texture information based on each target image, and generates the detection result based on the detection region feature information and the detection region texture information.

[0191] Optionally, the image processing model includes an encoder, a decoder, and a classifier;

[0192] The detection module 804 is further configured to:

[0193] Input the multiple target images into the encoder to obtain image encoding features, where the image encoding features are determined based on the detection region feature information and the detection region texture information corresponding to each target image;

[0194] Input the image encoding features into the decoder to obtain image decoding features corresponding to the image encoding features;

[0195] Input the image decoding features into the classifier to obtain a detection result corresponding to the target detection region.

[0196] Optionally, the encoder includes an image encoding unit, a texture encoding unit, and a feature fusion unit;

[0197] The detection module 804 is further configured to:

[0198] Input the multiple target images into the image encoding unit to obtain detection region feature information corresponding to the multiple target images;

[0199] Input the detection region feature information into the texture encoding unit to obtain detection region texture information corresponding to the detection region feature information;

[0200] Input the detection region feature information and the detection region texture information into the feature fusion unit to obtain image encoding features.

[0201] Optionally, the detection module 804 is further configured to:

[0202] Add distilled feature information and classification feature information to the image encoding features to obtain image encoding features to be processed;

[0203] Input the image encoding features to be processed into the decoder to obtain image decoding features.

[0204] Optionally, the detection module 804 is further configured to:

[0205] Input the image decoding feature into the classifier to obtain a first classification result corresponding to the distilled feature information and a second classification result corresponding to the classification feature information;

[0206] Generate the detection result according to the first classification result and the second classification result.

[0207] Optionally, the device further includes:

[0208] An acquisition module, configured to acquire biomarker information corresponding to a plurality of target images;

[0209] The detection module 804 is further configured to input the plurality of target images and the biomarker information into an image processing model to obtain a detection result corresponding to the target detection area.

[0210] Optionally, the detection module 804 is further configured:

[0211] Splice the biomarker information to the image encoding feature to be processed to obtain a spliced image encoding feature;

[0212] Input the spliced image encoding feature into the decoder to obtain an image decoding feature.

[0213] Optionally, the device further includes a training module, configured to:

[0214] Obtain a training sample pair, where the training sample pair includes a plurality of training sample images corresponding to a target detection area and a sample detection result of the target detection area, and the sample detection result includes a standard sample detection result or a reference sample detection result;

[0215] Input the plurality of training sample images into an image processing model to obtain a predicted detection result corresponding to the target detection area;

[0216] Calculate a model loss value according to the predicted detection result and the sample detection result;

[0217] Adjust the model parameters of the image processing model according to the model loss value until a model training stop condition is reached.

[0218] Optionally, the predicted detection result includes a first predicted classification result corresponding to the distilled feature information and a second predicted classification result corresponding to the classification feature information;

[0219] The training module is further configured to:

[0220] Calculate a first loss value according to the reference sample detection result and the first predicted classification result; or

[0221] Calculate a second loss value according to the standard sample detection result and the first prediction classification result;

[0222] Adjust the model parameters of the image processing model according to the first loss value and the second loss value.

[0223] Through the device provided in the embodiments of this specification, in the process of recognizing and processing multiple target images, generate detection region feature information based on each target image, then extract detection region texture information from the detection region feature information, generate the final image coding feature based on the detection region feature information and the detection region feature information, and add distilled feature information and classification feature information before image decoding, so that during the image decoding process, the distilled feature and the classification feature are referenced, making the final detection result more accurate.

[0224] The above is a schematic solution of an image processing device according to this embodiment. It should be noted that the technical solution of this image processing device and the technical solution of the above image processing method belong to the same concept. For the details not described in the technical solution of the image processing device, reference can be made to the description of the technical solution of the above image processing method.

[0225] Figure 9 FIG. shows a structural block diagram of a computing device 900 according to an embodiment of this specification. The components of the computing device 900 include but are not limited to a memory 910 and a processor 920. The processor 920 is connected to the memory 910 through a bus 930, and a database 950 is used to store data.

[0226] The computing device 900 further includes an access device 940, which enables the computing device 900 to communicate via one or more networks 960. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 940 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC).

[0227] In one embodiment of the present specification, the above components of the computing device 900 and Figure 9 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 9 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.

[0228] The computing device 900 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 900 can also be a mobile or stationary server.

[0229] Among them, the processor 920 is used to execute the following computer-executable instructions, which when executed by the processor implement the steps of the above image processing method.

[0230] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above image processing method belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above image processing method.

[0231] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above image processing method are implemented.

[0232] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above image processing method belong to the same concept. For the details not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above image processing method.

[0233] An embodiment of this specification also provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above image processing method.

[0234] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above image processing method belong to the same concept. For the details not described in the technical solution of the computer program, reference can be made to the description of the technical solution of the above image processing method.

[0235] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0236] The computer instructions include computer program code, which can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0237] It should be noted that, for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0238] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0239] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. An image processing method, comprising: Receive an image processing task, where the image processing task carries multiple target images corresponding to a target detection area, and the target image processing task is used to detect whether the target detection area is abnormal; Input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection area. The image processing model generates detection area feature information and detection area texture information based on each target image, determines an image coding feature based on the detection area feature information and the detection area texture information, adds distilled feature information and classification feature information to the image coding feature to obtain a to-be-processed image coding feature, obtains an image decoding feature according to the to-be-processed image coding feature, and generates the detection result according to the image decoding feature.

2. The method according to claim 1, wherein the image processing model comprises an encoder, a decoder and a classifier; Input the multiple target images into an image processing model to obtain a detection result corresponding to the target detection area, including: Input the multiple target images into the encoder to obtain an image coding feature, where the image coding feature is determined based on the detection area feature information and the detection area texture information corresponding to each target image; Input the image coding feature into the decoder to obtain an image decoding feature corresponding to the image coding feature; Input the image decoding feature into the classifier to obtain a detection result corresponding to the target detection area.

3. The method according to claim 2, wherein the encoder comprises an image encoding unit, a texture encoding unit and a feature fusion unit; Input the multiple target images into the encoder to obtain an image coding feature, including: Input the multiple target images into the image coding unit to obtain detection area feature information corresponding to the multiple target images; Input the detection area feature information into the texture coding unit to obtain detection area texture information corresponding to the detection area feature information; Input the detection area feature information and the detection area texture information into the feature fusion unit to obtain an image coding feature.

4. The method according to claim 2, inputting the image encoding feature into the decoder to obtain an image decoding feature corresponding to the image encoding feature, comprising: Add distilled feature information and classification feature information to the image coding feature to obtain a to-be-processed image coding feature; Input the to-be-processed image coding feature into the decoder to obtain an image decoding feature.

5. The method according to claim 4, inputting the image decoding feature into the classifier to obtain a detection result corresponding to the target detection region, comprising: Input the image decoding feature into the classifier to obtain a first classification result corresponding to the distilled feature information and a second classification result corresponding to the classification feature information; Generate the detection result according to the first classification result and the second classification result.

6. The method according to claim 4, further comprising: Obtain biomarker information corresponding to multiple target images; Input the multiple target images and the biomarker information into an image processing model to obtain a detection result corresponding to the target detection area.

7. The method according to claim 6, inputting the to-be-processed image encoding feature into the decoder to obtain an image decoding feature, comprising: Concatenate the biomarker information to the to-be-processed image coding feature to obtain a concatenated image coding feature; Input the concatenated image coding feature into the decoder to obtain an image decoding feature.

8. The method according to claim 1, wherein the image processing model is obtained by training through the following steps: Obtaining a training sample pair, wherein, The training sample pair includes multiple training sample images corresponding to a target detection area and a sample detection result of the target detection area, and the sample detection result includes a standard sample detection result or a reference sample detection result; Input the multiple training sample images into an image processing model to obtain a predicted detection result corresponding to the target detection area; Calculate a model loss value according to the predicted detection result and the sample detection result; Adjust the model parameters of the image processing model according to the model loss value until a model training stop condition is reached.

9. The method according to claim 8, wherein the predicted detection result includes a first predicted classification result corresponding to the distilled feature information and a second predicted classification result corresponding to the classification feature information; Calculating a model loss value according to the predicted detection result and the sample detection result includes: Calculating a first loss value according to the reference sample detection result and the first predicted classification result; or Calculating a second loss value according to the standard sample detection result and the first predicted classification result; Correspondingly, adjusting the model parameters of the image processing model according to the model loss value includes: Adjusting the model parameters of the image processing model according to the first loss value and the second loss value.

10. A CT image processing method, comprising: Receive a CT image processing task, where the CT image processing task carries a plurality of CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormality in the target detection area; Input the plurality of CT images into a CT image processing model to obtain a detection result corresponding to the target detection area, where the CT image processing model generates detection area feature information and detection area texture information based on each CT image, and determines an image coding feature based on the detection area feature information and the detection area texture information, adds distilled feature information and classification feature information to the image coding feature to obtain a to-be-processed image coding feature, obtains an image decoding feature according to the to-be-processed image coding feature, and generates the detection result according to the image decoding feature.

11. A training method of an image processing model, applied to a cloud-side device, comprising: Obtain a training sample pair, where the training sample pair includes a plurality of training sample images corresponding to a target detection area and a sample detection result of the target detection area, and the sample detection result includes a standard sample detection result or a reference sample detection result; Input the plurality of training sample images into an image processing model to obtain a predicted detection result corresponding to the target detection area, where the image processing model generates detection area feature information and detection area texture information based on each training sample image, and determines an image coding feature based on the detection area feature information and the detection area texture information, adds distilled feature information and classification feature information to the image coding feature to obtain a to-be-processed image coding feature, obtains an image decoding feature according to the to-be-processed image coding feature, and generates the predicted detection result according to the image decoding feature; Calculate a model loss value according to the predicted detection result and the sample detection result; Adjust the model parameters of the image processing model according to the model loss value until a model training stop condition is reached, and obtain the model parameters of the image processing model; Send the model parameters of the image processing model to the edge device.

12. An image processing method, comprising receiving an image processing request sent by a user, wherein, The image processing request includes an image processing task, the image processing task carries a plurality of target images corresponding to a target detection area, and the target image processing task is used to detect whether there is an abnormality in the target detection area; Input the multiple target images into an image processing model to obtain the detection results corresponding to the target detection regions. Among them, the image processing model generates detection region feature information and detection region texture information based on each target image, determines an image coding feature based on the detection region feature information and the detection region texture information, adds distilled feature information and classification feature information to the image coding feature to obtain a to-be-processed image coding feature, obtains an image decoding feature according to the to-be-processed image coding feature, and generates the detection results according to the image decoding feature; Send the detection results corresponding to the target detection regions to the user.

13. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.

14. A computer-readable storage medium storing computer-executable instructions, which when executed by a processor implement the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image recognition method based on neural network and related device

    CN113723310A

  • Tranform-based thyroid nodule detection method

    CN114494215A

  • Tumor cell detection equipment based on machine vision and method thereof

    CN115410050A

  • Object detection model training method and object detection method

    CN115965829A