Image detection method, device, computer equipment and storage medium

By filling in the missing sequence of MRI image and using AI technology for detection and segmentation, the misdetection and missed detection of lesions or abnormal areas of MRI image are solved, and more efficient lesion recognition and segmentation effects are achieved.

CN115115575BActive Publication Date: 2025-08-29腾讯医疗健康(深圳)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210456475.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-08-29
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

There are problems of misdetection and missed detection in the existing MRI image lesions or abnormal areas. Especially in the absence of image sequences, it is difficult to effectively identify the lesions or abnormal areas.

Method used

By designing reference image collection and image detection model, missing image sequences are filled and detection segmentation are performed. AI technology is used to assist doctors in identifying lesions or abnormal areas, and image recovery and feature learning are used using mask autoencoder and self-distillation technology.

Benefits of technology

The segmentation effect of MRI image lesions or abnormal areas is improved, the possibility of misjudgment and misjudgment is reduced, and the accuracy and efficiency of image detection are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115575B_ABST
    Figure CN115115575B_ABST
Patent Text Reader

Abstract

This application discloses an image detection method, apparatus, computer equipment, and storage medium that can be applied to the field of artificial intelligence, such as intelligent medicine. The method includes: obtaining a set of images to be detected that includes N sequences, performing detection on the set of images to be detected, and determining missing image regions in the set of images to be detected; then using a reference image set to fill in the missing image regions in the set of images to be detected to obtain a target detection image set; and then calling an image detection model to perform image detection and segmentation on the target detection image set to obtain a segmentation result corresponding to the set of images to be detected. This method can be used to perform intelligent detection and segmentation on multiple sequences of images, which can assist in the identification of lesions or abnormal areas.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and in particular to an image detection method, apparatus, computer equipment, and storage medium. Background Art

[0002] Magnetic resonance imaging (MRI) images are obtained using magnetic resonance imaging technology, which uses static and radiofrequency magnetic fields to image human tissue. During the imaging process, neither electron-ionizing radiation nor contrast agents are used to produce clear, high-contrast images. MRI can reveal organ dysfunction and early-stage lesions from within the human body. MRI images typically include multiple sequences, such as the fluid-attenuated inversion recovery (FLAIR) sequence, T1 sequence, T1c sequence, and T2 sequence. The multiple sequences included in these different sequences can present different tissue images and highlight different lesion areas.

[0003] In actual application, the identification of MRI lesions or abnormal areas is generally performed manually by doctors using their workbench and MRI, which may result in certain false detections or missed detections. Summary of the Invention

[0004] The embodiments of the present application provide an image detection method, apparatus, computer device, and storage medium, which can perform intelligent detection and segmentation on multiple sequence images and assist in the identification of lesions or abnormal areas.

[0005] In one aspect, an embodiment of the present application discloses an image detection method, the method comprising:

[0006] Acquire a set of images to be detected, where the set of images to be detected includes N sequences, each sequence includes sequence images, and N is an integer greater than or equal to 1;

[0007] If it is detected that the image set to be detected is in an image missing state, determining a missing image area in the image set to be detected;

[0008] Using a reference image set to fill in missing image areas in the to-be-detected image set, to obtain a target detection image set, the reference image set comprising N reference sequences, each reference sequence comprising a reference sequence image;

[0009] Abnormal area detection processing is performed in the image according to the target detection image set.

[0010] On the other hand, an embodiment of the present application discloses an image detection device, which includes:

[0011] An acquisition unit, configured to acquire a set of images to be detected, wherein the set of images to be detected includes N sequences, each sequence includes a sequence image, and N is an integer greater than or equal to 1;

[0012] a determining unit, configured to determine a missing image region in the image set to be detected if it is detected that the image set to be detected is in an image missing state;

[0013] a processing unit, configured to perform a filling operation on missing image regions in the to-be-detected image set using a reference image set to obtain a target detection image set, wherein the reference image set includes N reference sequences, each of which includes a reference sequence image;

[0014] The processing unit is further configured to perform abnormal region detection processing in an image based on the target detection image set.

[0015] Correspondingly, an embodiment of the present application also discloses a computer device, including an input interface and an output interface, and the computer device also includes: a processor, suitable for implementing one or more computer programs; and a computer storage medium, wherein the computer storage medium stores one or more computer programs, and the one or more computer programs are suitable for being loaded by the processor and executing the above-mentioned image detection method.

[0016] Correspondingly, an embodiment of the present application further discloses a computer-readable storage medium, wherein the computer storage medium stores one or more computer programs, and the one or more computer programs are suitable for being loaded by a processor and executing the above-mentioned image detection method.

[0017] Accordingly, embodiments of the present application further disclose a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described image detection method.

[0018] In an embodiment of the present application, the disclosed image detection method may mainly include: a computer device first obtains a set of images to be detected including N sequences, and in the case where there may be missing sequence images, the set of images to be detected can be detected, and after determining the missing image area in the set of images to be detected, the reference image set is used to fill in the missing image area in the set of images to be detected, to obtain a target detection image set, and after filling in the images of the missing sequence, the computer device then calls the image detection model to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the set of images to be detected. In this way, the characteristic information of the set of images to be detected can be better obtained, thereby helping to improve the segmentation effect of abnormal or lesion areas of multi-sequence images in the case of missing sequences, and can better assist doctor users in observing lesions or abnormal areas of images such as MRI, and to a certain extent reduce the possibility of missing or misjudging abnormal or lesion areas of images such as MRI. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 This is a schematic diagram of the architecture of an image detection system disclosed in an embodiment of the present application;

[0021] Figure 2 This is a schematic diagram of an application scenario of image detection disclosed in an embodiment of the present application;

[0022] Figure 3a is a schematic diagram of one of the image set relationships disclosed in the embodiments of this application;

[0023] Figure 3b This is a sequence diagram disclosed in the examples of this application;

[0024] Figure 4 This is a flow chart of an image detection method disclosed in an embodiment of the present application;

[0025] Figure 5 This is a schematic diagram of an image segmentation result disclosed in an embodiment of the present application;

[0026] Figure 6 is a schematic diagram of another image segmentation result disclosed in an embodiment of the present application;

[0027] Figure 7 This is a schematic diagram of an interface of an image detection method disclosed in an embodiment of the present application;

[0028] Figure 8 This is a training framework diagram for an image detection method disclosed in an embodiment of the present application;

[0029] Figure 9 This is a schematic diagram of the process of pre-training the image detection model disclosed in the embodiment of the present application;

[0030] Figure 10a It is a synthetic full-sequence synthetic image data disclosed in the embodiment of this application;

[0031] Figure 10b This is an implementation effect diagram disclosed in the embodiment of this application;

[0032] Figure 11 This is a schematic diagram of the process of fine-tuning the image detection model disclosed in the embodiment of this application;

[0033] Figure 12 1 is a schematic structural diagram of an image detection device disclosed in an embodiment of the present application;

[0034] Figure 13 It is a structural diagram of a computer device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0036] The image processing method provided in the embodiment of the present application is used to segment and identify abnormal areas in some images to be detected. Considering the possibility that the images to be detected are missing, a reference image set and an image detection model are designed. On the one hand, the reference image set is used to fill in the missing images. On the other hand, the image detection model can also restore the sequence and perform detection and segmentation to determine the lesion part. In this way, it is possible to better perform segmentation and detection of abnormalities or lesions in multi-sequence images such as MRI, and use AI (Artificial Intelligence) technology to assist doctors and other users in observing images and identifying lesions, reducing the possibility of missing or misjudging abnormalities or lesion areas in images such as MRI.

[0037] AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0038] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0039] The network models involved in AI can be trained and optimized through machine learning. Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0040] At the same time, in order to explain this application more clearly, some professional terms involved in this application are briefly described first, which may include:

[0041] (1) Magnetic resonance imaging (MRI): MRI is a relatively new medical imaging technology that uses static magnetic fields and radiofrequency magnetic fields to image human tissues. During the imaging process, high-contrast, clear images can be obtained without the use of electron ionizing radiation or contrast agents. It can reflect the abnormalities and early lesions of human organs from the inside of human molecules, and is superior to X-ray CT in many aspects. MRI images generally contain multiple sequences, such as FLAIR, T1, T1c, T2, etc. These different sequences can highlight different lesion areas. (2) Missing modality: In clinical applications, MRI often has one or more missing sequences due to image damage, artifacts, acquisition protocols, allergies to contrast agents, or cost. (3) Masked Autoencoder (MAE): As an image self-supervision framework, MAE has achieved great success in the field of self-supervision. Its agent task is to guide the model to restore the original pixel values ​​of an image based on the visible small blocks in an image. (4) Model Inversion (MI): Model inversion has long been used in the field of interpretability of deep learning. The goal of this technology is to synthesize the most representative images predicted by certain networks, such as saliency maps for classification. (5) Self-distillation (SD): Self-distillation is the use of supervised learning for knowledge distillation. Compared with the original knowledge distillation method, its teacher model and student model are one model, that is, one model guides itself to learn and complete knowledge distillation. (6) Multimodal masked autoencoder (M 3 AE): This is the abbreviation of the image detection model proposed in this application. It is an autoencoder that performs masking and restoration on multi-sequence image data, thereby simultaneously learning the associations between different sequences and the structural relationships in the image. (7) GFLOPS (Giga Floating-point Operations Per Second): This is the number of floating-point operations per second (GFLOPS). Floating-point refers to a number with a decimal point. Floating-point operations are the four arithmetic operations on decimals. They are often used to measure computer computing speed or to estimate computer performance, especially in the field of scientific computing that uses a large number of floating-point operations. They are mainly used in the training process of the model of this application.

[0042] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of an image detection system disclosed in an embodiment of the present application. Figure 1As shown, the architecture diagram 100 of the image detection system may include a terminal device 101 and a server 102, wherein the server 102 may be set in the cloud 103. The terminal device 101 is mainly used to receive the image set to be detected and the segmentation results corresponding to the image set to be detected in the embodiment of the present application. The server 102 is mainly used to deploy the image detection model in the embodiment of the present application, so that the image detection model can detect and segment the image set to be detected to obtain the final segmentation results. At the same time, the server 102 can also be responsible for training the image detection model.

[0043] In one possible implementation, the terminal device 101 obtains a set of images to be detected, which includes N sequences, each sequence includes one or more sequence images, and N is an integer greater than or equal to 1; the terminal device 101 then sends the set of images to be detected to the server 102, and the server 102 detects the set of images to be detected. If it is detected that the set of images to be detected is in an image missing state, the missing image area in the set of images to be detected is determined; further, the server 102 uses a reference image set to fill the missing image area in the set of images to be detected to obtain a target detection image set, wherein the reference image set includes N reference sequences, and each reference sequence includes one or more reference sequence images; the server 102 then calls the image detection model to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the set of images to be detected.

[0044] by Figure 3a For example, Figure 3a This is a schematic diagram of one of the image set relationships disclosed in the embodiment of the present application. The image set to be detected is in an image missing state, which may be missing one or more sequences. Figure 3a In the missing image, the sequence in the upper left corner, which is the position of the dotted box, is missing. The filling operation is to fill these missing parts through a trained reference image set. Specifically, the missing sequence in the missing image set is filled with the sequence at the corresponding position in the reference image set. For example, Figure 3a The sequence in the upper left corner of the image set to be detected is filled with the reference image sequence 1 in the upper left corner of the reference image set to obtain the final target detection image set.

[0045] Among them, the image detection model and the reference image set are obtained by server 102 through training optimization. The image detection model is trained based on the full sequence training image set and the missing training image set obtained after performing region masking processing on the full sequence training image set; the reference image set is obtained by optimizing the missing training image set and the initial reference image set. The initial reference image set includes N reference sequences, each initial reference sequence includes one or more initial reference sequence images, and the value of the pixel point on each initial reference sequence image is the value to be optimized. The value to be optimized of the pixel point on each initial reference sequence image is optimized to obtain the reference image set.

[0046] In a possible application scenario, for MRI data or brain tumor data, taking the server 102 as the cloud 103 as an example, an image detection scenario is described, such as Figure 2 When a user uploads a set of images to be segmented (i.e., input), the set of images to be segmented can be multi-sequence image data, in which any zero to multiple sequences may be missing. Based on the present application, a trained image detection model can be used to directly obtain the segmentation results (i.e., output) of abnormal or lesion areas such as segmented brain tumor areas. The segmentation results are specifically distinguished by different colors. The colors of regions 200, 201, and 203 are different. Region 200 (usually purple) is a background color unrelated to the lesion, region 201 (usually blue) represents edema, and region 202 (usually yellow) represents an enhancing tumor. In some cases, regions representing necrosis and non-enhancing tumor cores (usually green) may also appear. Figure 2 Alternatively, in one possible implementation, before the image detection model is finally determined, a pre-trained image detection model is also obtained. Based on this, the segmentation result may include, in addition to the label information, a full sequence of restored images obtained by completing the image set to be detected that is in a state of missing images. The full sequence of restored images is obtained by the pre-trained image detection model.

[0047] In one embodiment, Figure 3b This is a sequence diagram disclosed in an embodiment of the present application. It exemplarily gives 14 combined image sets (i.e., 14 sequence-missing image sets) and a full sequence in 4 sequences, for a total of 15 sequences. The image detection model trained using the embodiment of the present application can be used to process any of the 15 image sets and obtain segmentation results, thereby demonstrating the versatility of the image detection model of the embodiment of the present application in detecting and segmenting sequence-missing images and full-sequence images.

[0048] The terminal device 101 involved in the embodiments of the present application includes but is not limited to user equipment, handheld devices with wireless communication functions, vehicle-mounted devices, wearable devices or computing devices. For example, the terminal device can be a mobile phone, a tablet computer or a computer with wireless transceiver functions. The terminal device can also be a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in unmanned driving, a wireless terminal device in telemedicine, a wireless terminal device in a smart grid, a wireless terminal device in a smart city, a wireless terminal device in a smart home, etc. In the embodiments of the present application, the device for implementing the terminal device can also be a device that can support the terminal device to implement the function, such as a chip system, which can be installed in the terminal device. In the technical solution provided in the embodiments of the present application, the technical solution provided in the embodiments of the present application is described by taking the device for implementing the function of the terminal device as an example.

[0049] The server 102 mentioned in the embodiment of the present application can be a server. The server here can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services, such as Figure 1 The cloud 103 is mentioned as an example, and the present application embodiment does not limit it. In the technical solution provided in the embodiment of the present application, the technical solution provided in the embodiment of the present application is described by taking the server as an example.

[0050] In one embodiment, there may also be a computer device that can simultaneously implement the relevant functions of the terminal device 101 and the server 102 mentioned above, that is, it can interact with users such as doctors, obtain the required image set to be detected and present the segmentation results, and can also implement the image detection method mentioned in the embodiment of this application on the image set to be detected.

[0051] See Figure 4 , Figure 4 This is a flow chart of an image detection method disclosed in an embodiment of the present application, which mainly illustrates the use of an image detection model. The method can be applied to the server mentioned above or to a computer device, and can mainly include the following steps.

[0052] S401: Acquire a set of images to be detected. The acquired set of images to be detected may include N sequences. In this application, each sequence includes sequence images, and N is an integer greater than 1. For example, a nuclear magnetic resonance image may include 4 sequences. For another example, I is a set of images to be detected. Where W is the width of each image in the sequence, H is the height of each image in the sequence, D is the number of slices in the sequence, i.e., the number of images in the sequence, and N is the number of sequences in the image set to be tested. The image set to be tested can be obtained from a database for model prediction; alternatively, the image set to be tested can be obtained by a doctor during a patient examination, primarily used to determine the patient's condition based on the image set to be tested.

[0053] S402: If it is detected that the image set to be detected is in an image missing state, the missing image area in the image set to be detected is determined. After the image set to be detected is obtained, any image set in the image set to be detected is detected to determine whether there is a sequence missing situation. If so, the missing image area in the image set to be detected is determined. The missing situation here refers to the lack of one or several sequences, or the situation where a certain sequence is overlapped, such as an incomplete sequence display. In some possible situations, the image set to be detected may not be missing. In this case, the image set to be detected can be directly regarded as a full sequence image set to be detected. The full sequence image set to be detected can be directly detected and segmented by the image detection model, and the segmentation result can be output without using the reference image set to perform a padding operation on the full sequence of the image set to be detected. In other words, after the reference image set is used to perform a padding operation on the full sequence of the image set to be detected, the output is still a full sequence image set corresponding to the image set to be detected that can be input into the image detection model.

[0054] S403: Fill in the missing image areas in the to-be-detected image set using the reference image set to obtain the target detection image set, wherein the reference image set includes N reference sequences, each of which includes one or more sequence images.

[0055] In one possible implementation, after determining the missing image regions in the set of images to be detected in step S402, the missing image regions in the set of images to be detected are then filled in using the reference image set to obtain a set of target detection images. Filling in missing image regions using the reference image set is a method that saves time and space while producing synthetic data for the missing sequence at a very low cost. This method eliminates the need for additional modules and improves efficiency.

[0056] In one embodiment, the reference image set is obtained through training optimization. The reference image set may be obtained by optimizing the missing training image set and the initial reference image set. The initial reference image set may include N initial reference sequences, each of which includes one or more initial reference sequence images. The value of the pixel point on each initial reference sequence image is a value to be optimized. The value to be optimized of the pixel point on each initial reference sequence image is optimized to obtain the reference image set. For details on obtaining the reference image set, please refer to Figure 9 and Figure 11 The training process described in the corresponding embodiment.

[0057] S404: Perform abnormal area detection processing in the image according to the target detection image set. For the completed target detection image set, the abnormal area of ​​the image can be detected and segmented through some models to determine the abnormal area for display to the user. The abnormal area detection processing described in this application can be considered as the detection processing of the lesion area of ​​some medical images, or it can be the detection processing of image areas that exist in some images and are inconsistent with the image content of other areas or the image content of normal areas, or it can be the detection processing of image areas in some other special cases. The specific purpose of abnormal area detection can be determined according to the training image set and the supervision image set in the corresponding training data used during training.

[0058] In one possible implementation, the S404 may include: calling an image detection model to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the image set to be detected. After obtaining the target detection image set, calling the image detection model to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the image set to be detected. Specifically, the target detection image set may be input into the image detection model, and the image detection model may perform detection segmentation on the target detection image set to obtain a segmentation result. The segmentation result may specifically be the marking information obtained after marking the abnormal area of ​​the image set to be detected. At the same time, in some possible implementations, the segmentation result may also include a full sequence restored image set obtained after completing the image missing state of the image set to be detected, such as the above-mentioned pre-trained image detection model obtained in the model pre-training stage to restore the target detection image set after completing the image missing state of the image set to be detected, and obtain a full sequence restored image set.

[0059] In one embodiment, the image detection model is obtained through training and optimization. Specifically, the image detection model is obtained based on the full sequence training image set, the missing training image set, and the combined training image set. The combined training image set is obtained by combining the missing training image set and the initial reference image set. The missing training image set is obtained by performing region masking processing on the full sequence training image set using a masking technique. The specific acquisition of the image detection model can be found in Figure 9 and Figure 11 The training process described in the corresponding embodiment.

[0060] For example, for MRI data or brain tumor data, the image detection model can be used to obtain the segmentation results of four separate sequence images and one full sequence image, see Figure 5 , is a schematic diagram of an image segmentation result disclosed in an embodiment of the present application, specifically the cancer area segmentation result of the data of BraTS2018 (a data set for a multi-sequence brain tumor segmentation competition) using four separate sequence images and one full sequence image. In addition, the segmentation result can be accurately interpreted based on the provided gold standard (Groundtruth). The FLAIR sequence is a commonly used sequence of nuclear magnetic resonance (MR). The full name is the fluid attenuated inversion recovery sequence, also known as water suppression imaging technology. In layman's terms, it is a water pressure image. In this sequence, the cerebrospinal fluid shows a low signal (darker), and substantial lesions and lesions containing bound water appear as obvious high signals (brighter); T1 and T2 are physical quantities used to measure electromagnetic waves. They can be used as imaging data. Different sequences can highlight different sub-regions of the lesion, so as to assist the doctor user in finally determining the condition of the lesion.

[0061] It is worth noting that the application scenarios of this application are not limited to MRI data or brain tumor data, but can also be other types of multi-sequence medical imaging data combinations (such as PET (Positron Emission Computed Tomography, positron emission tomography), CT (Computed Tomography, electronic computed tomography), MRI, etc.) and other body parts (such as lung tumors), such as Figure 6As shown, (a) is lung tumor segmentation based on multiple PET sequence images, and (b) is lung tumor segmentation based on multiple CT sequence images. That is to say, the image set to be detected may be an MRI sequence, a PET sequence, or a CT sequence, etc., or a combination of two or more of the MRI sequence, PET sequence, and CT sequence. By combining sequences of multiple modalities (such as MRI, PET, CT) together, comprehensive AI recognition is performed to assist users such as doctors in more comprehensive detection of lesions. It should be noted that in the case of two or more combinations of MRI sequences, PET sequences, and CT sequences, the reference image set and the image detection model are based on the corresponding combined images to form training data, for example, the combination of MRI and PET. Then, in the subsequent pre-training stage and fine-tuning stage, the training data obtained by the combination of MRI and PET are used for optimization training to obtain the reference image set and the image detection model.

[0062] In one possible implementation, the image detection method of the present application can be displayed in a visual interface. Specifically, the image set to be detected is displayed in the first display area on the user interface, and at the same time, the segmentation results corresponding to the image set to be detected are displayed in the second display area of ​​the user interface. Similarly, in the present application, the segmentation results may include: marking information obtained after marking abnormal areas of the image set to be detected; or marking information obtained after marking abnormal areas of the image set to be detected, and a full-sequence restored image set obtained after completing the image missing state of the image set to be detected.

[0063] like Figure 7 As shown, it is a schematic diagram of the interface of an image detection method disclosed in an embodiment of the present application, wherein 701 is the first display area, 702 is the second display area, and the user can import the image set to be detected by clicking the import button in the first display area 701, as shown in 703, and then click the start button in the first display area 701, and then the corresponding segmentation results can be seen in the second display area 702, as shown in FIG. Figure 7 As shown, 704 is the restored full sequence image set corresponding to the image set to be detected, and 705 is the segmentation mark information. The segmentation mark information can also be directly superimposed with the full sequence image set 704 and displayed together.

[0064] In an embodiment of the present application, a computer device first obtains a set of images to be detected, which includes N sequences. Specifically, each sequence includes one or more sequence images, and N is an integer greater than 1; then the set of images to be detected is detected. If it is detected that the set of images to be detected is in an image missing state, the missing image area in the set of images to be detected can be determined, and then the missing image area in the set of images to be detected is filled using a reference image set to obtain a target detection image set, wherein the reference image set includes N reference sequences, and each reference sequence includes one or more sequence images. Based on this, the images of the missing sequences are filled; further, the computer device calls the image detection model to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the set of images to be detected.

[0065] In an embodiment of the present application, a computer device first completes an image with a missing sequence based on a reference image set, so as to better obtain feature information of the image set to be detected, thereby helping to improve the multi-sequence image segmentation effect in the case of missing sequences. Furthermore, the target detection image set is detected using a trained image detection model. Since the image detection model has been continuously optimized, the target detection image set can be detected faster to obtain segmentation results, thereby improving the overall image detection efficiency.

[0066] See Figure 8 , is a training framework diagram for an image detection method disclosed in an embodiment of the present application, which is roughly divided into two parts, one of which is pre-training ( Figure 8 The upper part of the straight line shown) and one part for fine adjustment ( Figure 8 lower half of the straight line shown).

[0067] In the pre-training stage, taking a pair of full-sequence training image sets and an initial reference image set as an example, it can mainly include: performing region masking processing on the full-sequence training images to obtain a missing training image set, where the region masking processing can specifically be masking any one or more sequences in multiple sequences, or masking any one or more sequences in multiple sequences and then masking some three-dimensional small blocks, and then combining the missing training image set and the initial reference image set to obtain a combined training image set, inputting the combined training image set into the initial model for training to obtain a predicted image set, and optimizing the initial model and the initial reference image set according to the difference between the predicted image set and the full-sequence training image set as the supervision image, so as to obtain a pre-trained reference image set and a pre-trained image detection model, where the specific difference refers to calculating the loss value between the predicted image set and the initial reference image set, and adjusting the initial model and the initial reference image set according to the loss value. After repeated training, when the loss value reaches the convergence condition, a multi-sequence masked autoencoder (MMA) can be obtained. 3 The AE, or pre-trained image detection model, is designed to learn feature representations for multiple image sequences in the presence of missing sequences. During pre-training, the model is optimized through continuous backpropagation (model derivation), which also generates a set of pre-trained reference images (i.e., the initial reference images are continuously optimized). This set of images can be used to fill in missing sequences during training and inference.

[0068] It should be noted that the initial reference image set can be an initially generated N-sequence image set, or it can be a N-sequence reference image set obtained after training based on the previous full-sequence training image set and the missing training image set, which requires further optimization. In other words, in addition to the final usable reference image set, any image set that requires training and optimization can be referred to as the initial reference image set.

[0069] In the fine-tuning stage, taking a pair of full sequence training image sets and the segmentation supervision information corresponding to the full sequence training image sets as an example, it mainly includes: first inputting the full sequence training image set into the M obtained in the pre-training stage 3Segmentation prediction is performed in AE (i.e., pre-trained image detection model) to obtain first segmentation prediction information, which is stored in the storage space. The full sequence training image set and the pre-trained reference image set are then randomly combined to obtain a combined fine-tuning image set, and the combined fine-tuning image set is input into the pre-trained image detection model for segmentation prediction to obtain second segmentation prediction information. Then, according to the difference between the first segmentation prediction information and the second segmentation prediction information, and the difference between the second segmentation prediction information and the segmentation supervision information configured for the full sequence training image set, the pre-trained reference image set and the pre-trained image detection model are optimized to obtain a reference image set and an image detection model.

[0070] In one embodiment, the difference between the first segmentation prediction information and the second segmentation prediction information, and the difference between the second segmentation prediction information and the segmentation supervision information configured for the full sequence training image set are both reflected in loss values, that is, the loss value between the first segmentation prediction information and the second segmentation prediction information is first calculated, and the loss value between the second segmentation prediction information and the segmentation supervision information configured for the full sequence training image set is calculated, and then the sum of the two loss values ​​is calculated, and the pre-trained image detection model and the pre-trained reference image set are fine-tuned according to the value. After repeated adjustments, when the loss value reaches the convergence condition, the image detection model can achieve a higher-precision segmentation effect of the abnormal or lesion area in the case of missing sequences, thereby obtaining the final image detection model and reference image set. The image detection model and reference image set obtained after the above two stages of training are universal and can be used to process MRI image data in any missing sequence during testing (use).

[0071] The backbone network of the network model used in the training process of this application can be VT-UNet-T, which is a pure Transformer (a self-attention transformation network) architecture. The corresponding parameter number and computational complexity are lower than the commonly used 3DUnet (an image analysis model) or Vnet (an image analysis model). At the same time, this application uses the Adam (Adaptive momentum, an optimization algorithm) algorithm as the optimizer during network training, setting the number of training rounds in the first and second stages to 600 and 400 rounds respectively. The initial learning rate of training is 3e-4, and the cosine annealing learning rate scheduling mechanism is used during training, which has good corresponding convergence. This application trains the model on two 2080Ti NVIDIA graphics cards with a batch size of 2. In order to standardize all data, the pixel values ​​can be clipped to 1% to 99% of the intensity value during training, then minimum or maximum scaling is performed, and finally randomly cropped to a fixed size of 128×128×128 pixels for training. The side length of the random three-dimensional patch can be set to 16 pixels. The images in the corresponding initial reference image set are initialized by Gaussian noise, and λ can be set to 0.1.

[0072] It is worth noting that: 1. When filling in missing sequences, the pre-trained model can be directly used to generate synthetic data of the missing sequences; 2. In addition to using VT-UNet as the backbone network, the network model in this application can also use other commonly used segmentation networks as the backbone network; 3. This application can be expanded to other multi-sequence images with similar application scenarios or other tissue structure segmentation tasks, and is not limited to MIR data or brain tumor data.

[0073] according to Figure 8 The training framework diagram described above, where the flowchart of the pre-training phase can be found in Figure 9 , is a schematic diagram of a process for pre-training an image detection model disclosed in an embodiment of the present application, Figure 9 It can include S901-S905, and the specific steps are as follows:

[0074] S901: Obtain training data for training an image detection model. The training data may include a full sequence of training images (supervisory images) and an initial reference image set. This data may be obtained from a database or a relevant institution. For example, MIR data or brain tumor data may be obtained directly from a hospital.

[0075] S902: Performing a region masking process on the full sequence training image set to obtain a missing training image set. The missing training image set can be obtained by masking one or more sequences in the full sequence training image set using a masking process or other methods. For example, if the full sequence training image set has four sequences, region masking can be performed on one, two, or three of the sequences to obtain the missing training image set. Alternatively, one or more sequences in the full sequence training image set can be masked, and region masking can be performed on the sequence images in the remaining sequences to obtain the missing training image set.

[0076] For example, a full-sequence training image set consists of four sequences. Region masking can be performed by masking one of the sequences and then covering some of the three-dimensional blocks of the remaining three sequences to obtain a missing training image set. The missing training data set may include M missing sequences, where M is an integer greater than or equal to 1 and less than N.

[0077] S903: Obtain a combined training image set based on the missing training image set and the initial reference image set. The combined training image set can be obtained based on the missing training image set and the initial reference image set. Specifically, the image set at the corresponding sequence position in the initial reference image set can be overlaid on the missing sequences in the missing training image set to obtain the combined training image set. During the overlaying process, do not overwrite the remaining portions of the missing training image set. This ensures that the missing sequences can be effectively restored in subsequent training.

[0078] For example, for MRI, if the missing training image set is missing the T1 sequence, the reference sequence set corresponding to T1 in the reference image can be used to cover the position of the T1 sequence of the missing training image set to obtain a completed combined training image set.

[0079] S904: Call the initial model to detect the combined training image set to obtain a set of predicted images. After obtaining the combined training image set, the combined training image set is input into the initial model for processing to obtain a set of predicted images. The initial model is constructed based on a mask autoencoder, and the corresponding backbone network can use VT-UNet. Of course, other commonly used segmentation networks can also be used as the backbone network, and this application is not limited to them.

[0080] S905: Optimizing the initial model and the initial reference image set according to the difference between the predicted image set and the full sequence training image set, so as to obtain a pre-trained reference image set and a pre-trained image detection model.

[0081] In one possible implementation, the initial model and the initial reference image set are optimized based on the difference between the predicted image set and the full sequence training image set, and the initial model and the initial reference image set are optimized based on the loss value between the predicted image set and the full sequence training image set.

[0082] In one embodiment, the loss value between the predicted image set and the full sequence training image set can be calculated. If the loss value between a large number of predicted image sets and the full sequence training image set is less than or equal to a first threshold, the pre-trained reference image set and the pre-trained image detection model are determined. Alternatively, when the initial model is in a certain set of model parameters, the corresponding loss value is the smallest for a large number of missing training image sets or the full sequence training image set, and the corresponding initial model is determined to be the pre-trained image detection model, and the corresponding pre-trained reference image set is obtained.

[0083] In another possible implementation, based on the difference between the predicted image set and the full sequence training image set, the optimization target expression for optimizing the initial model is as shown in formula (1):

[0084]

[0085] Among them, x represents the full sequence training image set, x′ represents the missing training image set, and x sub represents the initial reference image set, S(x′,x sub ) represents the combined training image set, F is the reconstruction function, is the L2 regularization term, γ is the weight, is the mean square error loss function. Through this optimization formula, the model corresponding to the minimum loss value can be determined and determined as the pre-trained image detection model. This target optimization formula can enable the initial model to learn the relationship between sequences in the data and the integrity of the anatomy without any annotation. According to the difference between the predicted image set and the full sequence training image set, the optimization target expression for optimizing the initial reference image set is as shown in formula (2):

[0086]

[0087] Among them, x represents the full sequence training image set, x′ represents the missing training image set, and x sub represents the initial reference image set, represents the set of pre-trained reference images, S(x′,x sub ) represents the combined training image set, F is the reconstruction function, is the L2 regularization term, γ is the weight, is the mean square error loss function. Through this optimization formula, The corresponding content in is used to complete the content of x that was masked during the pre-training process, rather than directly masking it with 0. This can better reconstruct multiple sequences with missing content (sequences or partial blocks). The completed content must capture information that can represent a specific sequence. This will also help improve the performance of multi-sequence segmentation in the case of missing partial sequences. In layman's terms, this is to optimize the initial reference image set through backpropagation, which can be called model derivation. In this way, the model does not need to introduce new modules, and the optimization cost of the initial reference image set is extremely low. See Figure 10a , is a set of full sequence restored images obtained through this optimization method.

[0088] In this application, formula (1) and formula (2) can use a very small regularization weight, that is, γ = 0.005, and at the same time, the mean square error loss function is used. It can make the model better reconstruct the original image, and the regularization term Can get x sub The credibility is higher.

[0089] It is worth noting that for the area masking process in step S902, this application has conducted corresponding experiments and proved that the masking method of this application can achieve better results among various methods. Figure 10b , is the experimental effect of the method of this application under different mask probabilities. In the MAE method, the model can only restore the masked area by referring to the content around the image, while in the training process of this application, the model can also restore the masked area by referring to images of other sequences. Therefore, this application chooses a larger mask probability to make the self-supervised task of this application more difficult, so that the model can learn better features. Figure 10b As shown in the figure, the final model performance obtained by using 0.8125 or 0.875 is better than 0.75 (the mask probability in MAE-related papers). The Dice indicator (a set similarity measurement indicator, such as DSC (Dice Similarity Coefficient)) is used to measure the experimental effect. The higher the Dice, the better. Among them, WT (whole tumor) is the tumor as a whole, including all tumor areas; TC (tumor core) is the tumor core, which consists of enhancing tumor, necrotic area and non-enhancing tumor core; ET (enhancing tumor) is the enhancing tumor.

[0090] This application mainly describes the pre-training process in the model training process, with the aim of obtaining a pre-trained image detection model and a pre-trained reference image set. This application uses a multi-sequence masked autoencoder to learn the rich feature representations in multi-sequence MRI in the case of missing sequences. The model is a single encoder-decoder structure, which reduces the difficulty of training the model. At the same time, the pre-training stage trains the initial model and the initial reference image set based on the training data and the sequence completion rules based on model inversion to obtain the pre-trained reference image set and the pre-trained image detection model. The pre-trained reference image set and the pre-trained image detection model can be used to complete the sequences that may be missing during the training and reasoning process, thereby improving the efficiency of the image detection of this application.

[0091] according to Figure 8 The training framework diagram described in this paper, where the flowchart of the fine-tuning phase can be found in Figure 11 , Figure 11 It may include S1101-S1104, and the specific steps are as follows:

[0092] S1101: Combine the pre-trained reference image set with the full-sequence training image set to obtain a combined fine-tuning image set. In one possible implementation, the pre-trained reference image set and the full-sequence training image set are combined according to rule representation information to obtain a combined fine-tuning image set; or the pre-trained reference image set and the full-sequence training image set are randomly combined to obtain a combined fine-tuning image set. Among the N fine-tuning sequences included in the combined fine-tuning image set, x fine-tuning sequences are from the pre-trained reference image set, and y fine-tuning sequences are from the full-sequence training image set, where x and y are positive integers and x + y = N.

[0093] For example, Figure 8 As shown, rule representation information is displayed. According to the rule representation information, the full sequence training image set and the pre-trained reference image set can be combined to obtain a combined fine-tuning image set. It can be seen that the dark part of the rule representation information indicates that the position corresponding to the full sequence training image set is covered, and the light part indicates that the position corresponding to the full sequence training image set is not covered. According to the rule, the combination is performed to obtain a combined fine-tuning image set.

[0094] S1102: Perform segmentation prediction on the entire training image sequence using a pre-trained image detection model to obtain first segmentation prediction information. In one possible implementation, to obtain better segmentation results, the entire training image sequence is first input into the pre-trained image detection model for segmentation prediction. The first segmentation prediction information is then stored in a storage space. This storage space can also be stored in CPU memory, which is more suitable for hardware that cannot implement joint training due to a lack of GPU memory. Furthermore, the first segmentation prediction information can be updated in real time while the model is fine-tuned.

[0095] S1103: Perform segmentation prediction on the combined fine-tuned image set using the pre-trained image detection model to obtain second segmentation prediction information. In one possible implementation, the combined fine-tuned image set is input into the pre-trained image detection model for segmentation prediction to obtain the second segmentation prediction information.

[0096] S1104: Optimize the pre-trained reference image set and the pre-trained image detection model based on the difference between the first segmentation prediction information and the second segmentation prediction information, and the difference between the second segmentation prediction information and the segmentation supervision information configured for the full sequence training image set to obtain a reference image set and an image detection model.

[0097] In one possible implementation, the pre-trained reference image set and the pre-trained image detection model are optimized based on the difference between the first segmentation prediction information and the second segmentation prediction information, and the difference between the second segmentation prediction information and the segmentation supervision information configured for the full sequence training image set. The loss value between the first segmentation prediction information and the second segmentation prediction information and the loss value between the second segmentation prediction information and the segmentation supervision information are calculated, and the parameter values ​​of the pre-trained reference image set and the pre-trained image detection model are optimized based on these two loss values ​​to obtain the reference image set and the image detection model.

[0098] Specifically, the optimization objective expression for optimizing the pre-trained reference image set and the pre-trained image detection model is formula (3):

[0099]

[0100] in, Represents the first segmentation prediction information (the corresponding segmentation result for the full sequence), Represents the second segmentation prediction information (the corresponding segmentation result when the sequence is missing), s gt represents the segmentation supervision information of the full sequence training image set configuration, f is the backbone network corresponding to the model, and f s It is the segmentation head after the backbone network, λ is the weight, which can be set to 0.1. is the sum of Dice loss and cross loss, is the consistency loss function, As shown in formula (4):

[0101]

[0102] It is the KL distance (Kullback-Leibler Divergence, which measures the difference between two probability distributions in the same event space) between the segmentation results under the full sequence and the segmentation results under the missing sequence. Specifically, W is the width of each image in the image set, H is the height of each image in the image set, D is the number of slices of each image in the image set, and C is the total number of categories of the image set segmentation.

[0103] Steps S1102-S1104 are a computationally efficient self-distillation method that can transfer task-related knowledge from full-sequence data to missing-sequence data in the same network. The model can be fine-tuned into a multi-sequence segmentation model that can simultaneously handle various missing sequences, while also reducing the computational overhead during training and deployment.

[0104] This application mainly describes the fine-tuning process during model training, with the goal of obtaining the final image detection model and reference image set, which can be processed for sequence images in any situation to obtain the corresponding segmentation results. The fine-tuning task of this application is a computationally efficient self-distillation method. During the fine-tuning process of the segmentation task, the information of the full sequence data is distilled to the missing sequence to achieve a higher accuracy segmentation effect in the case of missing sequence.

[0105] go through Figure 9 as well as Figure 11 After two stages of training, the generated image detection model and reference image set are both highly versatile and can be used to process MRI data in any missing sequence when used (i.e., the prediction process). Based on this, the present application conducted specific experiments, specifically experiments completed on the PyTorch neural network framework, and obtained corresponding experimental results. Specifically, the technology corresponding to the image detection method of the present application was experimented on the brain tumor segmentation competitions BraTS 2018 and BraTS 2019 to verify its effectiveness. The BraTS series data set consists of multiple pairs of MRIs including four sequences, namely T1, T1c, T2 and FLAIR. These data have been sorted and organized by the contestants, including stripping the skull, resampling to a uniform resolution (1mm 3 ), and perform preprocessing such as co-registration on the same template.

[0106] In this competition, four types of intratumor structures (edema, enhancing tumor, necrosis and non-enhancing tumor core) are divided into three tumor regions and used as segmentation targets for the competition: 1. Whole tumor (WT), including all tumor regions; 2. Tumor core (TC), consisting of enhancing tumor, necrotic area and non-enhancing tumor core; 3. Enhancing tumor (ET). The BraTS 2018 and BraTS 2019 data sets include 285 and 335 data sets and corresponding tumor region annotations, respectively. In the experiment, the two data sets can be randomly divided into training set and test set in a ratio of 80:20, respectively. In this application, the Dice coefficient and 95% Hausdorff distance (HD95) can be used as evaluation indicators. In addition, an online evaluation system can be used to verify the performance of the technology of this application in the verification set stored in the database under full sequence conditions. The above-mentioned BraTS 2018 and BraTS 2019 are two data sets that already exist in the database and can be used directly.

[0107] Table 1 compares the proposed image detection method with three general methods for brain MRI tumor segmentation in the BraTS 2018 dataset: HVED, LCRL, and FGMF. FGMF is the general method with the highest performance. Since these methods present segmentation results on 20% of the data, their results can be directly extracted from the database. Table 1 shows that the proposed image detection method performs best overall on the test set, achieving the best median in all three tumor regions. Furthermore, the proposed image detection method achieves the best results in the majority of cases (the proposed image detection method achieves the best results in 14, 11, and 10 cases, respectively, for a total of 15 missing sequences). It is worth noting that the proposed image detection method utilizes a basic single encoder-decoder framework, while the three other methods mentioned above all employ multiple encoders or decoders, resulting in a computationally intensive approach compared to the proposed image detection method.

[0108] Table 1

[0109]

[0110] Among them, the existing and missing sequences are represented by · and °, respectively. The p-value is given by the Wilcoxon test of the significance of the corresponding method and the image detection method of the present application. The specific explanations of the above three methods are: HVED (Hetero-ModalVariational Encoder-Decoder for Joint Modality Completion and Segmentation), LCRL (Latent correlation representation learning for brain tumor segmentation with missing MRI modalities), and FGMF (Feature-enhanced generation and multimodality fusion based deep neural network for brain tumor segmentation with missing MR modalities).

[0111] Table 2 compares our image detection method on the BraTS 2019 dataset with LCRL, the only comparative method tested on this dataset. The results show that our image detection method outperforms LCRL in all cases of missing images in all tumor regions, demonstrating its good generalization.

[0112] Table 2

[0113]

[0114] The existing and missing sequences are represented by · and °, respectively, and the p-values ​​are given by the Wilcoxon test of the significance of the corresponding method and the image detection method of the present application.

[0115] In addition, although the image segmentation method of this application proposes a "general" model, in order to reflect the effectiveness of the image detection method of this application, it is also compared with the currently best dedicated model ACN (Adversarial Co-training Network). The training and testing ratio used by this method is different from that of the image detection method of this application, and the results shown in its paper can be directly quoted as a reference.

[0116] The specific results can be seen in Table 3. The image detection method of this application only needs to train one model, and the overall performance is almost the same as that of the ACN that trains a model separately for each missing case (this method requires training 15 models in the experiment of this application).

[0117] Table 3

[0118]

[0119] Among them, the existing and missing sequences are represented by · and ο, respectively, and the p value is given by the Wilcoxon test of the significance of the corresponding method and the image detection method of the present application.

[0120] Furthermore, to objectively demonstrate the performance of our image detection method in full-sequence scenarios, Table 4 compares the online test results of our image detection method with those of several existing methods on two datasets, including LCRL, VT-UNet-T (the backbone network used in our image detection method), and TransBTS (another Transformer model for brain MRI tumor segmentation). The results of the preferred solutions from the corresponding competition (obtained from existing databases) are also included in Table 4 for reference. It is worth noting that these preferred solutions typically undergo extensive engineering, such as fine-tuning parameters. The results show that, compared with other non-competition solutions, our image detection method achieved the best results in nine scenarios (two datasets × two indicators × three tumor regions, for a total of 12 scenarios). In some cases, our image detection method nearly surpassed the preferred solutions from the corresponding competitions, which typically underwent extensive parameter tuning. These results demonstrate that the multi-sequence representations learned by our image detection method are not only robust to missing sequences but also effective for full sequences.

[0121] Table 4

[0122]

[0123] Among them, the comparison results of the image detection method of this application with the existing best methods under the full sequence conditions of BraTS 2018 (left) and BraTS 2019 (right) data. Challenge indicates the preferred solution of the corresponding competition, and NA indicates that it is not available.

[0124] Furthermore, to verify the effectiveness of each module in the technology proposed by the image detection method of this application, an ablation experiment can be completed by removing each module in the overall solution one by one. The results are shown in Table 5, and the following conclusions can be summarized:

[0125] 1. In rows 1 and 2 (a, b), the pre-training stage is removed from the training process, and the pre-training parameters on the ImageNet dataset are added to the latter. The results of both rows are significantly reduced, indicating that the pre-trained image detection model plays an indispensable role in the framework proposed by the image detection method of this application.

[0126] 2. In the third row (c), replacing the full sequence image learned through model inversion in this application with an image containing all zeros resulted in a decrease in the result, while in the fourth row (d), replacing the full sequence image with the average value of all data in the corresponding sequence significantly worsened the result. This shows that the missing sequence filling scheme based on model inversion proposed in the image detection method of this application captures more useful sequence feature information and can effectively serve as a supplement to missing sequences in brain tumor segmentation.

[0127] 3. Finally, compared to row 5 (e), the proposed framework achieves better evaluation metrics across all tumor regions, validating the effectiveness of self-distillation from full to missing sequences. Furthermore, the proposed self-distillation framework saves approximately 52 GFLOPS (floating-point operations) compared to the joint training approach.

[0128] Table 5

[0129]

[0130] Based on the above method embodiment, the present application embodiment also provides a structural diagram of an image detection device. Figure 12 , is a structural diagram of an image detection device provided by an embodiment of the present invention. The device can be applied to the server mentioned above, or can be applied to a computer device, Figure 12 The image detection device 1200 shown can run the following units:

[0131] An acquiring unit 1201 is configured to acquire a set of images to be detected, where the set of images to be detected includes N sequences, each sequence includes a sequence of images, and N is an integer greater than 1.

[0132] The determining unit 1202 is configured to determine a missing image region in the image set to be detected if it is detected that the image set to be detected is in an image missing state;

[0133] The processing unit 1203 is configured to use a reference image set to fill in missing image regions in the to-be-detected image set to obtain a target detection image set, wherein the reference image set includes N reference sequences, each of which includes a reference sequence image.

[0134] The processing unit 1203 is further configured to perform abnormal region detection processing in an image according to the target detection image set.

[0135] In a possible implementation, the processing unit 1203 performs abnormal region detection processing in the image according to the target detection image set, which may specifically be used to:

[0136] The image detection model is called to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the image set to be detected.

[0137] In one possible implementation, the reference image set and the image detection model are obtained through training optimization; wherein, the image detection model is trained based on a full sequence training image set and a missing training image set obtained by performing region masking processing on the full sequence training image set; the reference image set is obtained by optimizing the missing training image set and the initial reference image set, the initial reference image set includes N initial reference sequences, each initial reference sequence includes an initial reference sequence image, and the value of the pixel point on each initial reference sequence image is a value to be optimized, and the reference image set is obtained by optimizing the value to be optimized of the pixel point on each initial reference sequence image.

[0138] In a possible implementation, the image detection device further includes:

[0139] The display unit 1204 is used to display the image set to be detected in the first display area on the user interface; and display the segmentation result corresponding to the image set to be detected in the second display area of ​​the user interface; wherein the segmentation result includes: marking information obtained after marking abnormal areas of the image set to be detected; or, the segmentation result includes: marking information obtained after marking abnormal areas of the image set to be detected, and a full-sequence restored image set obtained after completing the image missing state of the image set to be detected.

[0140] In one possible implementation, the marking information obtained after marking the abnormal areas of the image set to be detected is obtained through the image detection model, and the full-sequence restored image set obtained after filling in the image-missing state of the image set to be detected is obtained through the reference image set.

[0141] In a possible implementation, the acquisition unit 1201 is further configured to acquire training data for training the image detection model, the training data including: a full sequence training image set as supervisory images and an initial reference image set;

[0142] The processing unit 1203 is further configured to:

[0143] Performing a region masking process on the full sequence training image set to obtain a missing training image set, wherein the missing training image set includes M missing sequences, where M is an integer greater than or equal to 1 and less than N;

[0144] Obtaining a combined training image set according to the missing training image set and the initial reference image set;

[0145] Calling the initial model to detect the combined training image set to obtain a predicted image set;

[0146] According to the difference between the predicted image set and the full sequence training image set, the initial model and the initial reference image set are optimized to obtain a pre-trained reference image set and a pre-trained image detection model.

[0147] In a possible implementation, the processing unit 1203 performs region masking processing on the full sequence training image set to obtain a missing training image set, specifically including:

[0148] masking one or more sequences in the full-sequence training image set to obtain a missing training image set;

[0149] Alternatively, one or more sequences in the full sequence training image set are masked, and region masking processing is performed on sequence images in the remaining sequences to obtain a missing training image set.

[0150] In a possible implementation, the processing unit 1203 optimizes the initial model according to the difference between the predicted image set and the full sequence training image set, and the optimization target expression is:

[0151]

[0152] Among them, x represents the full sequence training image set, x′ represents the missing training image set, and x sub represents the initial reference image set, S(x′,x sub ) represents the combined training image set, F is the reconstruction function, is the L2 regularization term, γ is the weight, is the mean square error loss function.

[0153] In a possible implementation, the processing unit 1203 optimizes the initial reference image set according to the difference between the predicted image set and the full sequence training image set. The optimization objective expression is:

[0154]

[0155] Among them, x represents the full sequence training image set, x′ represents the missing training image set, and x sub represents the initial reference image set, represents the set of pre-trained reference images, S(x′,x sub ) represents the combined training image set, F is the reconstruction function, is the L2 regularization term, γ is the weight, is the mean square error loss function.

[0156] In a possible implementation, the processing unit 1203 is further configured to:

[0157] Combining the pre-trained reference image set with the full-sequence training image set to obtain a combined fine-tuning image set, wherein x fine-tuning sequences of N fine-tuning sequences included in the combined fine-tuning image set are from the pre-trained reference image set, and y fine-tuning sequences are from the full-sequence training image set, where x and y are positive integers and x+y=N;

[0158] Performing segmentation prediction on the full sequence training image set using the pre-trained image detection model to obtain first segmentation prediction information;

[0159] Performing segmentation prediction on the combined fine-tuned image set using the pre-trained image detection model to obtain second segmentation prediction information;

[0160] Based on the difference between the first segmentation prediction information and the second segmentation prediction information, and the difference between the second segmentation prediction information and the segmentation supervision information configured for the full sequence training image set, the pre-trained reference image set and the pre-trained image detection model are optimized to obtain a reference image set and an image detection model.

[0161] In a possible implementation, the optimization target expression for optimizing the pre-trained reference image set and the pre-trained image detection model by the processing unit 1203 is:

[0162]

[0163] in, represents the first segmentation prediction information, Represents the second segmentation prediction information, s gt represents the segmentation supervision information configured for the full sequence training image set, λ is the weight, is the sum of Dice loss and cross loss, is the consistency loss function.

[0164] In one possible implementation, the corresponding image detection model and reference image set are obtained by training and optimizing the initial model and the initial reference image set;

[0165] Training and optimizing the initial model and the initial reference image set includes pre-training and fine-tuning; the initial model is constructed based on a mask autoencoder;

[0166] During pre-training, the initial model and the initial reference image set are trained according to the training data and a sequence completion rule based on model inversion to obtain a pre-trained reference image set and a pre-trained image detection model;

[0167] During fine-tuning, the pre-trained reference image set and the pre-trained image detection model are trained according to a self-distillation method from the full-sequence training image set to the missing sequence data set to obtain a reference image set and an image detection model.

[0168] In the embodiment of the present application, the acquisition unit 1201 acquires a set of images to be detected including N sequences, and then detects the set of images to be detected. If it is detected that the set of images to be detected is in an image missing state, the determination unit 1202 determines the missing image area in the set of images to be detected, and the processing unit 1203 then uses the reference image set to fill in the missing image area in the set of images to be detected to obtain a target detection image set, based on which the images of the missing sequences are filled; further, the image detection model is called to perform image detection segmentation on the target detection image set to obtain the segmentation result corresponding to the set of images to be detected. Through this method, the image with missing sequences is first filled in according to the reference image set, so that the feature information of the set of images to be detected can be better obtained, which can help improve the multi-sequence image segmentation effect in the case of missing sequences. Further, the trained image detection model is used to detect the target detection image set. Since the image detection model has been continuously optimized, the target detection image set can be detected faster to obtain the segmentation result, thereby improving the image detection efficiency.

[0169] Based on the above method and device embodiments, the present application provides a computer device. Figure 13 , is a structural diagram of a computer device provided in an embodiment of the present application. Figure 13 The computer device 1300 shown includes at least a processor 1301, an input interface 1302, an output interface 1303, a computer storage medium 1304, and a memory 1305. The processor 1301, the input interface 1302, the output interface 1303, the computer storage medium 1304, and the memory 1305 may be connected via a bus or other means.

[0170] Computer storage medium 1304 may be stored in memory 1305 of computer device 1300. Computer storage medium 1304 is used to store computer programs, including program instructions. Processor 1301 is used to execute the program instructions stored in computer storage medium 1304. Processor 1301 (or CPU (Central Processing Unit)) is the computing and control core of computer device 1300 and is suitable for implementing one or more computer programs, specifically loading and executing one or more computer programs to implement corresponding method flows or corresponding functions.

[0171] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer storage medium provides a storage space, which stores the operating system of the computer device. In addition, one or more computer programs (including program codes) suitable for being loaded and executed by the processor 1301 are also stored in the storage space. It should be noted that the computer storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk memory; optionally, it can also be at least one computer storage medium located away from the aforementioned processor. The computer storage medium can be loaded by the processor 2701 and execute one or more computer programs stored in the computer storage medium to achieve the above-mentioned related Figure 4 、 Figure 9 as well as Figure 11 In a specific implementation, one or more instructions in a computer storage medium are loaded by the processor 1301 and execute the following steps:

[0172] Acquire a set of images to be detected, wherein the set of images to be detected includes N sequences, each sequence includes sequence images, and N is an integer greater than 1;

[0173] If it is detected that the image set to be detected is in an image missing state, determining a missing image area in the image set to be detected;

[0174] Using a reference image set to fill in missing image areas in the to-be-detected image set, to obtain a target detection image set, the reference image set comprising N reference sequences, each reference sequence comprising a reference sequence image;

[0175] Abnormal area detection processing is performed in the image according to the target detection image set.

[0176] In a possible implementation, the processor 1301 performs abnormal region detection processing in an image according to the target detection image set, which may be specifically used to:

[0177] Call the image detection model to perform image detection segmentation on the target detection image set to obtain the segmentation results corresponding to the image set to be detected

[0178] In one possible implementation, the reference image set and the image detection model are obtained through training optimization; wherein, the image detection model is trained based on a full sequence training image set and a missing training image set obtained by performing region masking processing on the full sequence training image set; the reference image set is obtained by optimizing the missing training image set and the initial reference image set, the initial reference image set includes N initial reference sequences, each initial reference sequence includes an initial reference sequence image, and the value of the pixel point on each initial reference sequence image is a value to be optimized, and the reference image set is obtained by optimizing the value to be optimized of the pixel point on each initial reference sequence image.

[0179] In a possible implementation, the processor 1301 is further configured to:

[0180] A set of images to be detected is displayed in a first display area on a user interface; and a segmentation result corresponding to the set of images to be detected is displayed in a second display area of ​​the user interface; wherein the segmentation result includes: marking information obtained after marking abnormal areas of the set of images to be detected; or, the segmentation result includes: marking information obtained after marking abnormal areas of the set of images to be detected, and a full-sequence restored image set obtained after completing a process for restoring the set of images to be detected that is in an image-missing state.

[0181] In one possible implementation, the marking information obtained after marking abnormal areas of the image set to be detected is obtained through the image detection model; the full-sequence restored image set obtained after filling in the image-missing state of the image set to be detected is obtained through the reference image set.

[0182] In a possible implementation, the processor 1301 is further configured to:

[0183] Acquiring training data for training the image detection model, the training data comprising: a full sequence of training images as supervisory images and an initial reference image set;

[0184] Performing a region masking process on the full sequence training image set to obtain a missing training image set; the missing training image set includes M missing sequences, where M is an integer greater than or equal to 1 and less than N;

[0185] Obtaining a combined training image set according to the missing training image set and the initial reference image set;

[0186] Calling the initial model to detect the combined training image set to obtain a predicted image set;

[0187] According to the difference between the predicted image set and the full sequence training image set, the initial model and the initial reference image set are optimized to obtain a pre-trained reference image set and a pre-trained image detection model.

[0188] In a possible implementation, the processor 1301 performs region masking processing on the full sequence training image set to obtain a missing training image set, which can be specifically used to:

[0189] masking one or more sequences in the full-sequence training image set to obtain a missing training image set;

[0190] Alternatively, one or more sequences in the full sequence training image set are masked, and region masking processing is performed on sequence images in the remaining sequences to obtain a missing training image set.

[0191] In a possible implementation, the processor 1301 optimizes the initial model according to the difference between the predicted image set and the full sequence training image set. The optimization objective expression is:

[0192]

[0193] Among them, x represents the full sequence training image set, x′ represents the missing training image set, and x sub represents the initial reference image set, S(x′,x sub ) represents the combined training image set, F is the reconstruction function, is the L2 regularization term, γ is the weight, is the mean squared error loss function.

[0194] In a possible implementation, the processor 1301 optimizes the initial reference image set according to the difference between the predicted image set and the full sequence training image set. The optimization objective expression is:

[0195]

[0196] Among them, x represents the full sequence training image set, x′ represents the missing training image set, and x sub represents the initial reference image set, represents the set of pre-trained reference images, S(x′,x sub ) represents the combined training image set, F is the reconstruction function, is the L2 regularization term, γ is the weight, is the mean squared error loss function.

[0197] In a possible implementation, the processor 1301 is further configured to:

[0198] Combining the pre-trained reference image set with the full-sequence training image set to obtain a combined fine-tuning image set, wherein x fine-tuning sequences of N fine-tuning sequences included in the combined fine-tuning image set are from the pre-trained reference image set, and y fine-tuning sequences are from the full-sequence training image set, where x and y are positive integers and x+y=N;

[0199] Performing segmentation prediction on the full sequence training image set using the pre-trained image detection model to obtain first segmentation prediction information;

[0200] Performing segmentation prediction on the combined fine-tuned image set using the pre-trained image detection model to obtain second segmentation prediction information;

[0201] Based on the difference between the first segmentation prediction information and the second segmentation prediction information, and the difference between the second segmentation prediction information and the segmentation supervision information configured for the full sequence training image set, the pre-trained reference image set and the pre-trained image detection model are optimized to obtain a reference image set and an image detection model.

[0202] In a possible implementation, the optimization target expression for optimizing the pre-trained reference image set and the pre-trained image detection model by the processor 1301 is:

[0203]

[0204] in, represents the first segmentation prediction information, Represents the second segmentation prediction information, s gt represents the segmentation supervision information of the full sequence training image set configuration, λ is the weight, is the sum of Dice loss and cross loss, is the consistency loss function.

[0205] In one possible implementation, the corresponding image detection model and reference image set are obtained by training and optimizing the initial model and the initial reference image set;

[0206] Training and optimizing the initial model and the initial reference image set includes pre-training and fine-tuning; the initial model is constructed based on a mask autoencoder;

[0207] During pre-training, the initial model and the initial reference image set are trained according to the training data and a sequence completion rule based on model inversion to obtain a pre-trained reference image set and a pre-trained image detection model;

[0208] During fine-tuning, the pre-trained reference image set and the pre-trained image detection model are trained according to a self-distillation method from the full-sequence training image set to the missing sequence data set to obtain a reference image set and an image detection model.

[0209] In an embodiment of the present application, the processor 1301 obtains a set of images to be detected including N sequences. If it is detected that the set of images to be detected is in an image missing state, the missing image area in the set of images to be detected can be determined, and then the reference image set is used to perform a filling operation on the missing image area in the set of images to be detected to obtain a target detection image set. Based on this, the image of the missing sequence is filled; further, the image detection model is called to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the set of images to be detected. Through this method, the image with missing sequence can be filled according to the reference image set first, and the feature information of the set of images to be detected can be better obtained, which can help improve the multi-sequence image segmentation effect in the case of missing sequence. Further, the trained image detection model is used to detect the target detection image set. Since the image detection model has been continuously optimized, the target detection image set can be detected faster to obtain the segmentation result, thereby improving the image detection efficiency.

[0210] In an embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the steps performed in all the above embodiments can be executed.

[0211] An embodiment of the present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. When the computer instructions are executed by a processor of a computer device, the methods in all the above embodiments are executed.

[0212] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0213] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that implementing all or part of the processes of the above embodiment and making equivalent changes in accordance with the claims of this application still fall within the scope of the invention.

Claims

1. An image detection method, characterized in that: The method comprises: Acquire a set of images to be detected, wherein the set of images to be detected includes N sequences, each sequence includes sequence images, and N is an integer greater than or equal to 1; If it is detected that the image set to be detected is in an image missing state, determining a missing image area in the image set to be detected; Using a reference image set to fill in missing image areas in the to-be-detected image set, to obtain a target detection image set, the reference image set comprising N reference sequences, each reference sequence comprising a reference sequence image; Calling the image detection model to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the image set to be detected; Acquiring training data for training the image detection model, the training data comprising: a full sequence of training images as supervisory images and an initial reference image set; Performing a region masking process on the full sequence training image set to obtain a missing training image set; the missing training image set includes M missing sequences, where M is an integer greater than or equal to 1 and less than N; Obtaining a combined training image set based on the missing training image set and the initial reference image set; calling the initial model to detect the combined training image set to obtain a predicted image set; According to the difference between the predicted image set and the full sequence training image set, the initial model and the initial reference image set are optimized to obtain a pre-trained reference image set and a pre-trained image detection model.

2. The method according to claim 1, wherein The reference image set and the image detection model are obtained through training and optimization; The image detection model is trained based on a full sequence training image set and a missing training image set obtained by performing region masking processing on the full sequence training image set; The reference image set is obtained by optimizing the missing training image set and the initial reference image set; The initial reference image set includes N initial reference sequences, each initial reference sequence includes an initial reference sequence image, and the values ​​of the pixel points on each initial reference sequence image are values ​​to be optimized. The reference image set is obtained by optimizing the values ​​to be optimized of the pixel points on each initial reference sequence image.

3. The method according to claim 1, wherein The method further comprises: Displaying a set of images to be detected in a first display area on the user interface; Displaying the segmentation results corresponding to the set of images to be detected in the second display area of ​​the user interface; The segmentation result includes: marking information obtained by marking abnormal areas of the image set to be detected; Alternatively, the segmentation result includes: marking information obtained by marking abnormal regions of the image set to be detected, and a full sequence restored image set obtained by completing a process for restoring missing images of the image set to be detected.

4. The method according to claim 3, wherein The marking information obtained after marking abnormal areas of the set of images to be detected is obtained by the image detection model; The full sequence restored image set obtained by performing a filling process on the image set to be detected that is in an image missing state is obtained through the reference image set.

5. The method according to claim 1, wherein The performing region masking processing on the full sequence training image set to obtain a missing training image set includes: masking one or more training image sequences in the full-sequence training image set to obtain a missing training image set; Alternatively, one or more training image sequences in the full-sequence training image set are masked, and region masking processing is performed on sequence images in the remaining training image sequences to obtain a missing training image set.

6. The method according to claim 1, wherein The optimization objective expression for optimizing the initial model according to the difference between the predicted image set and the full sequence training image set is: Among them, x represents the full sequence training image set, x ′ represents the missing training image set, x sub represents the initial reference image set, S(x ′ ,x sub ) represents the combined training image set, F is the reconstruction function, is the L2 regularization term, γ is the weight, is the mean squared error loss function.

7. The method according to claim 1, wherein The optimization objective expression for optimizing the initial reference image set according to the difference between the predicted image set and the full sequence training image set is: Among them, x represents the full sequence training image set, x ′ represents the missing training image set, x sub represents the initial reference image set, represents the pre-training reference image set, S(x ′ ,x sub ) represents the combined training image set, F is the reconstruction function, is the L2 regularization term, γ is the weight, L mse is the mean squared error loss function.

8. The method according to claim 1, wherein The method further comprises: Combining the pre-trained reference image set with the full-sequence training image set to obtain a combined fine-tuning image set, wherein x fine-tuning sequences of N fine-tuning sequences included in the combined fine-tuning image set are from the pre-trained reference image set, and y fine-tuning sequences are from the full-sequence training image set, where x and y are positive integers and x+y=N; Performing segmentation prediction on the full sequence training image set using the pre-trained image detection model to obtain first segmentation prediction information; Performing segmentation prediction on the combined fine-tuned image set using the pre-trained image detection model to obtain second segmentation prediction information; Based on the difference between the first segmentation prediction information and the second segmentation prediction information, and the difference between the second segmentation prediction information and the segmentation supervision information configured for the full sequence training image set, the pre-trained reference image set and the pre-trained image detection model are optimized to obtain a reference image set and an image detection model.

9. The method according to claim 8, wherein The optimization objective expression for optimizing the pre-trained reference image set and the pre-trained image detection model is: in, represents the first segmentation prediction information, Represents the second segmentation prediction information, s gt represents the segmentation supervision information configured for the full sequence training image set, λ is the weight, is the sum of Dice loss and cross loss, is the consistency loss function.

10. The method according to claim 1, wherein The corresponding image detection model and reference image set are obtained by training and optimizing the initial model and the initial reference image set; Training and optimizing the initial model and the initial reference image set includes pre-training and fine-tuning; the initial model is constructed based on a mask autoencoder; During pre-training, the initial model and the initial reference image set are trained according to the training data and a sequence completion rule based on model inversion to obtain a pre-trained reference image set and a pre-trained image detection model; During fine-tuning, the pre-trained reference image set and the pre-trained image detection model are trained according to a self-distillation method from a full-sequence training image set to a missing sequence data set to obtain a reference image set and an image detection model.

11. An image detection device, characterized in that: The device comprises: An acquisition unit is used to acquire a set of images to be detected, wherein the set of images to be detected includes N sequences, each sequence includes a sequence image, and N is an integer greater than 1; a determining unit, configured to determine a missing image region in the image set to be detected if it is detected that the image set to be detected is in an image missing state; a processing unit, configured to perform a filling operation on missing image regions in the to-be-detected image set using a reference image set to obtain a target detection image set, wherein the reference image set includes N reference sequences, each of which includes a reference sequence image; The processing unit is further configured to call an image detection model to perform image detection segmentation on the target detection image set to obtain a segmentation result corresponding to the set of images to be detected; The acquisition unit is further configured to acquire training data for training the image detection model, wherein the training data includes: a full sequence training image set as a supervision image and an initial reference image set; The processing unit is further configured to: Performing a region masking process on the full sequence training image set to obtain a missing training image set, wherein the missing training image set includes M missing sequences, where M is an integer greater than or equal to 1 and less than N; Obtaining a combined training image set according to the missing training image set and the initial reference image set; Calling the initial model to detect the combined training image set to obtain a predicted image set; According to the difference between the predicted image set and the full sequence training image set, the initial model and the initial reference image set are optimized to obtain a pre-trained reference image set and a pre-trained image detection model.

12. A computer device comprising an input interface and an output interface, the computer device further comprising: a processor adapted to execute one or more computer programs; And, a computer storage medium, wherein the computer storage medium stores one or more computer programs, wherein the one or more computer programs are suitable for being loaded by the processor and executing the image detection method according to any one of claims 1 to 10.

13. A computer storage medium, characterized in that The computer storage medium stores one or more computer programs, and the one or more computer programs are suitable for being loaded by a processor and executing the image detection method according to any one of claims 1 to 10.

14. A computer program product comprising computer instructions, characterized in that The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, so that the computer device executes the image detection method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Method for determining to-be-labeled image, and method and device for model training

    CN110517759A

  • Image processing method and device, computer equipment and storage medium

    CN114119578A