Image recognition method and system for urinary calculus
By using migration networks to migrate the feature of different modal images in urinary stone image recognition, the problem of insufficient information complementarity in traditional technology is solved, and the recognition accuracy of urinary stone areas is improved.
Patent Information
- Application Number
- CN202510217744.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional single-modal urinary stone images cannot comprehensively and accurately detect and quantify urinary stone lesions, and the information between the multimodal images cannot be fully integrated, resulting in poor recognition accuracy.
An image recognition method for urinary stones is proposed. By acquiring images of different modalities (such as CT and MRI), using a migration network to perform feature migration, and an intermediate feature map with complementary information is generated, thereby improving the accuracy of the identification of stone areas by the image segmentation model.
By fully realizing information complementarity between different modal images, the identification accuracy of urinary stone areas is improved, and the problems of poor information loss and interpretability in traditional technologies are solved.
Smart Images

Figure CN120163976A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of medical image processing, and in particular, to an image recognition method and system for urinary calculi. Background Art
[0002] Urinary calculi are solid structures formed within the urinary system, which can cause serious complications such as urinary system inflammation and infection. The urinary calculi image segmentation technology based on computer vision plays an important role in assisting doctors in treating urinary calculi. By using different computer vision technologies, doctors can accurately locate urinary calculi, evaluate their size, quantity, and morphology, and then determine the best treatment plan. At the same time, it can assist doctors in monitoring the development of calculi; therefore, the computer analysis and processing of urinary calculi have practical significance.
[0003] Currently, the research on urinary calculi image segmentation technology mainly focuses on specific single modalities, such as computed tomography (CT), positron emission tomography (PET), or magnetic resonance imaging (MRI). However, traditional single-modal images cannot comprehensively and accurately detect and quantify urinary calculi lesions. There is complementary information between different-modal images. If this complementary information can be utilized, it will help improve the recognition accuracy of urinary calculi.
[0004] Currently, there are mainly the following several ways to combine features between multi-modal images: 1) Before inputting multi-modal images into an image segmentation model, first fuse different-modal images into one-modal images, but this method has a large information loss; 2) Input multi-modal images into corresponding image segmentation models respectively, and then fuse them after obtaining their respective segmentation results. However, this method does not fuse the image information of different modalities during the image segmentation process, which will lead to ignoring the information correlation between different-modal images and poor interpretability. Summary of the Invention
[0005] The following is an overview of the topics described in detail in this article. This overview is not intended to limit the protection scope of the claims.
[0006] The main purpose of the embodiments of the present disclosure is to propose an image recognition method and system for urinary calculi, which can more accurately identify the regions of calculi in different-modal images.
[0007] The first aspect of the embodiments of the present application proposes an image recognition method for urinary calculi, and the method includes the following steps: Obtain a first urolith image to be recognized and a second urolith image to be recognized of a target patient; the first urolith image to be recognized is an image of a first modality, and the second urolith image to be recognized is an image of a second modality, and the first modality and the second modality are different; Input the first urolith image to be recognized and the second urolith image to be recognized into a preset image segmentation model to obtain the urolith contour output by the image segmentation model; Generate a corresponding recognition report according to the urolith contour; Wherein, the image segmentation model includes a segmentation network and a migration network, and the segmentation network includes an input layer, an encoder-decoder, and an output layer; The training process of the image segmentation model is as follows: Obtain a first training urolith image of the first modality and a second training urolith image of the second modality; Input the first training urolith image and the second training urolith image into the input layer to obtain a first initial feature map corresponding to the first training urolith image and a second initial feature map corresponding to the second training urolith image; Input the first initial feature map and the second initial feature map into the migration network, and by migrating the first initial feature map to the second initial feature map, obtain a second intermediate feature map; and by migrating the second initial feature map to the first initial feature map, obtain a first intermediate feature map; Input the first intermediate feature map and the second intermediate feature map into the encoder-decoder to obtain a decoded feature map; Input the decoded feature map into the output layer to obtain the output urolith contour.
[0008] The present application proposes an image recognition method for uroliths, which has the following beneficial effects: Different from the existing feature fusion technology in multi-modal images, before encoding the images of the first modality and the second modality, the present application first inputs the first initial feature map extracted from the image of the first modality into the migration network, and by migrating the first initial feature map to the second initial feature map extracted from the image of the second modality, obtains a second initial feature map with the texture feature information in the first initial feature map, and then inputs the second initial feature map into the migration network to migrate the second initial feature map to the first initial feature map, so as to obtain a first initial feature map with the texture feature information in the second initial feature map, fully realizing the complementarity of information between the images of the first modality and the second modality, and further improving the accuracy of the image segmentation model in recognizing the stone area.
[0009] In some embodiments of the present application, the codec includes a first encoder and a first decoder, as well as a second encoder and a second decoder; the first encoder and the second encoder have the same structure, and the first decoder and the second decoder have the same structure; The input data of the first encoder is the first intermediate feature map, the input data of the second encoder is the second intermediate feature map, and the fused feature map between the feature map output by the first decoder and the feature map output by the second decoder is used as the decoded feature map; The first encoder includes 6 convolution modules with a convolution kernel of 3 × 3, and 3 downsampling modules. The first layer and the last layer of the first encoder are both the convolution modules, and there are two connected convolution modules between every two adjacent downsampling modules. The first decoder includes 6 convolution modules with a convolution kernel of 3 × 3, and 3 upsampling modules. The first layer and the last layer of the first decoder are both the convolution modules, and there are two connected convolution modules between every two adjacent upsampling modules; The input feature map of the j-th convolution module of the first encoder is the first fused feature map, and the input feature map of the j-th convolution module of the second encoder is the second fused feature map; where the first fused feature map is the fused feature map between the output feature map of the downsampling module closest to and before the j-th convolution module of the first encoder and the third fused feature map, and the third fused feature map is the fused feature map between the output feature map of the downsampling module closest to and before the j-th convolution module of the first encoder and the output feature map of the downsampling module closest to and before the j-th convolution module of the second encoder; the second fused feature map is the fused feature map between the output feature map of the downsampling module closest to and before the j-th convolution module of the second encoder and the third fused feature map; j is 2, 4, 6; The input feature map of the i-th convolution module of the first encoder is the output feature map of the (i - 1)-th convolution module, and the input feature map of the i-th convolution module of the second encoder is the output feature map of the (i - 1)-th convolution module; i is 3, 5; The input feature map of the j-th convolutional module of the first decoder is the fourth fused feature map, and the input feature map of the j-th convolutional module of the second decoder is the fifth fused feature map; wherein, the fourth fused feature map is the feature map fused between the output feature map of the nearest upsampling module before the j-th convolutional module of the first decoder and the sixth fused feature map, and the sixth fused feature map is the feature map fused between the output feature map of the nearest upsampling module before the j-th convolutional module of the first decoder and the output feature map of the nearest upsampling module before the j-th convolutional module of the second decoder; the fifth fused feature map is the feature map fused between the output feature map of the nearest upsampling module before the j-th convolutional module of the second decoder and the sixth fused feature map; The input feature map of the i-th convolutional module of the first decoder is the output feature map of the (i - 1)-th convolutional module, and the input feature map of the i-th convolutional module of the second decoder is the output feature map of the (i - 1)-th convolutional module.
[0010] In some embodiments of the present application, the output end of at least one convolutional module in the first encoder is connected to the input end of the convolutional module with the same feature map dimension in the first decoder by residual connection, or the output end of at least one convolutional module in the second encoder is connected to the input end of the convolutional module with the same feature map dimension in the second decoder by residual connection.
[0011] In some embodiments of the present application, the migration network includes a region segmentation layer, a migration layer, and a fusion layer; Inputting the first initial feature map and the second initial feature map into the migration network to obtain a second intermediate feature map by migrating the first initial feature map towards the second initial feature map, includes: Inputting the first initial feature map and the second initial feature map into the region segmentation layer to obtain m first segmentation regions from the first initial feature map and n second segmentation regions from the second initial feature map; Inputting the m first segmentation regions and the n second segmentation regions into the migration layer to perform global migration and local migration of the first initial feature map towards the second initial feature map according to the m first segmentation regions and the n second segmentation regions by using a migration algorithm, respectively obtaining a global migration result and a local migration result; Inputting the global migration result and the local migration result into the fusion layer to obtain the second intermediate feature map; Among them, a migration algorithm is used to perform global migration of the first initial feature map to the second initial feature map to obtain a global migration result, including: Performing pixel-by-pixel multiplication on the first initial feature map and the second initial feature map to obtain a first pixel-by-pixel multiplication result, and inputting the first pixel-by-pixel multiplication result into the AdaIN function to obtain a global migration result; Among them, a migration algorithm is used to perform local migration of the first initial feature map to the second initial feature map to obtain a local migration result, including: Finding the first segmentation region most similar to each second segmentation region among the m first segmentation regions; Calculating the pixel-by-pixel multiplication between the mask corresponding to each second segmentation region and the second initial feature map one by one to obtain a second pixel-by-pixel multiplication result, and calculating the pixel-by-pixel multiplication between the mask corresponding to the most similar first segmentation region of each second segmentation region and the first initial feature map one by one to obtain a third pixel-by-pixel multiplication result; Inputting the second pixel-by-pixel multiplication result and the third pixel-by-pixel multiplication result into the AdaIN function to obtain the migration result corresponding to each second segmentation region; Fusing the migration results corresponding to the n second segmentation regions to obtain a local migration result.
[0012] In some embodiments of the present application, the migration network includes a region segmentation layer, a migration layer, and a fusion layer; Inputting the first initial feature map and the second initial feature map into the migration network to perform migration of the second initial feature map to the first initial feature map to obtain a first intermediate feature map, including: Inputting the first initial feature map and the second initial feature map into the region segmentation layer to obtain m first segmentation regions from the first initial feature map and n second segmentation regions from the second initial feature map; Inputting the m first segmentation regions and the n second segmentation regions into the migration layer to perform global migration and local migration of the second initial feature map to the first initial feature map according to the m first segmentation regions and the n second segmentation regions by using a migration algorithm, respectively obtaining a global migration result and a local migration result; Inputting the global migration result and the local migration result into the fusion layer to obtain the first intermediate feature map; Among them, a migration algorithm is used to perform global migration of the second initial feature map to the first initial feature map to obtain a global migration result, including: Perform pixel-by-pixel multiplication on the first initial feature map and the second initial feature map to obtain a first pixel-by-pixel multiplication result, and input the first pixel-by-pixel multiplication result into the AdaIN function to obtain a global migration result; Among them, a migration algorithm is used to perform local migration of the second initial feature map to the first initial feature map to obtain a local migration result, including: Find the second segmentation region that is most similar to each first segmentation region among the n second segmentation regions; Calculate the pixel-by-pixel multiplication between the mask corresponding to each first segmentation region and the first initial feature map one by one to obtain a second pixel-by-pixel multiplication result, and calculate the pixel-by-pixel multiplication between the mask corresponding to the most similar second segmentation region corresponding to each first segmentation region and the second initial feature map one by one to obtain a third pixel-by-pixel multiplication result; Input the second pixel-by-pixel multiplication result and the third pixel-by-pixel multiplication result into the AdaIN function to obtain the migration result corresponding to each first segmentation region; Fuse the migration results corresponding to the m first segmentation regions to obtain a local migration result.
[0013] In some embodiments of the present application, the first modality and the second modality are the MRI and CT modalities respectively.
[0014] To achieve the above object, a second aspect of the embodiments of the present invention provides an image recognition system for urinary calculi, and the system includes the following structures: A data acquisition unit for acquiring a first urinary calculus image to be recognized and a second urinary calculus image to be recognized of a target patient; the first urinary calculus image to be recognized is an image of the first modality, and the second urinary calculus image to be recognized is an image of the second modality, and the first modality and the second modality are different; A model application unit for inputting the first urinary calculus image to be recognized and the second urinary calculus image to be recognized into a preset image segmentation model to obtain the urinary calculus contour output by the image segmentation model; A report generation unit for generating a corresponding recognition report according to the urinary calculus contour; Among them, the image segmentation model includes a segmentation network and a migration network, and the segmentation network includes an input layer, an encoder, a decoder, and an output layer; A model training unit for training the image segmentation model, and the training process of the image segmentation model is: Obtain a first training urinary calculus image of the first modality and a second training urinary calculus image of the second modality; Input the first training urolithiasis image and the second training urolithiasis image into the input layer to obtain a first initial feature map corresponding to the first training urolithiasis image and a second initial feature map corresponding to the second training urolithiasis image; Input the first initial feature map and the second initial feature map into the migration network to obtain a second intermediate feature map by migrating the first initial feature map to the second initial feature map; and obtain a first intermediate feature map by migrating the second initial feature map to the first initial feature map; Input the first intermediate feature map and the second intermediate feature map into the encoder-decoder to obtain a decoded feature map; Input the decoded feature map into the output layer to obtain the output urolithiasis contour.
[0015] To achieve the above object, a third aspect of the embodiments of the present invention provides an electronic device, including: at least one control processor and a memory for communicatively connecting with the at least one control processor; the memory stores instructions executable by the at least one control processor, and the instructions are executed by the at least one control processor to enable the at least one control processor to execute the above image recognition method for urolithiasis.
[0016] To achieve the above object, a fourth aspect of the embodiments of the present invention provides a computer-readable storage medium, which stores computer-executable instructions for causing a computer to execute the above image recognition method for urolithiasis.
[0017] It can be understood that the beneficial effects of the above second aspect to the fourth aspect compared with the related art are the same as those of the above first aspect compared with the related art. For the relevant descriptions, reference can be made to the relevant descriptions in the above first aspect, and details will not be repeated here. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the related art descriptions. Obviously, the following drawings are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is a schematic flowchart of the image recognition method for urolithiasis provided by the embodiments of the present application; Figure 2 is a schematic diagram of the network structure provided by the embodiments of the present application; Figure 3It is a schematic diagram of an embodiment of an image recognition system for urinary calculi provided by this application; Figure 4 It is a schematic diagram of an electronic device provided by an embodiment of this application. Detailed implementation manners
[0020] In order to make the objectives, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0021] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0023] Urinary calculi are solid structures formed in the urinary system, which can cause serious complications such as urinary system inflammation and infection. The computer vision-based urinary calculus image segmentation technology plays an important role in assisting doctors in treating urinary calculi. By using different computer vision technologies, doctors can accurately locate urinary calculi, evaluate their size, quantity and morphology, and then determine the best treatment plan. At the same time, it can assist doctors in monitoring the development of calculi; therefore, the computer analysis and processing of urinary calculi has practical significance.
[0024] Currently, the research on urinary calculus image segmentation technology mainly focuses on specific single modalities, such as computed tomography (CT), positron emission tomography (PET) or magnetic resonance imaging (MRI). However, traditional single-modal images cannot comprehensively and accurately detect and quantify urinary calculus lesions. There is complementary information between different-modal images. If this complementary information can be utilized, it will help improve the recognition accuracy of urinary calculi.
[0025] Currently, there are mainly the following several ways to combine the features between multi-modal images: 1) Before inputting multi-modal images into an image segmentation model, first fuse different-modal images into one-modal images, but this method has a large amount of information loss; 2) Input the multi-modal images into the corresponding image segmentation models respectively. After obtaining their respective segmentation results, they are fused. However, this method does not fuse the image information of different modalities during the image segmentation process, which will lead to ignoring the information correlation between different types of modality images and result in poor interpretability.
[0026] To solve the defects, such as Figure 1 and Figure 2 , the embodiments of the present application provide an image recognition method for urinary calculi. The method includes the following steps: Step S110, obtain the first urinary calculus image to be recognized and the second urinary calculus image to be recognized of the target patient; the first urinary calculus image to be recognized is an image of the first modality, and the second urinary calculus image to be recognized is an image of the second modality, and the first modality and the second modality are different; Step S120, input the first urinary calculus image to be recognized and the second urinary calculus image to be recognized into a preset image segmentation model to obtain the urinary calculus contour output by the image segmentation model; Step S130, generate a corresponding recognition report according to the urinary calculus contour; In step S110 of this embodiment, the target patient refers to a patient diagnosed by a doctor with urinary calculi. It should be noted that the doctor diagnoses that the patient has urinary calculi through the hospital's medical means and the hospital's experience. The recognition of the present application is only used to assist the doctor in locating urinary calculi in medical images.
[0027] The first modality and the second modality are respectively any one of CT, MRI, and PET, or can also be variant images of CT, MRI, or PET. For example, the image of the first modality is a CT image, and the image of the second modality is an MRI.
[0028] There is complementary information in the images of different modalities. Therefore, in step S120 of this embodiment, the first urinary calculus image to be recognized and the second urinary calculus image to be recognized are input into a preset image segmentation model to obtain the urinary calculus contour output by the image segmentation model. It should be noted that the image segmentation model can output the urinary calculus contour corresponding to the first urinary calculus image to be recognized, and can also output the urinary calculus contour corresponding to the second urinary calculus image to be recognized. During the image segmentation process, there will be an interaction process of feature information between the first urinary calculus image to be recognized and the second urinary calculus image to be recognized. For details, see the introduction of the subsequent model training process.
[0029] In step S130, a corresponding recognition report is generated according to the urinary calculus contour. The report includes but is not limited to the relevant content of the size, position, quantity, and shape of the urinary calculus.
[0030] Among them, the image segmentation model includes a segmentation network and a migration network. The segmentation network includes an input layer, an encoder-decoder, and an output layer; The training process of the image segmentation model is as follows: Step S210, obtain the first training urinary calculus image of the first modality and the second training urinary calculus image of the second modality; Step S220, input the first training urinary calculus image and the second training urinary calculus image into the input layer to obtain a first initial feature map corresponding to the first training urinary calculus image and a second initial feature map corresponding to the second training urinary calculus image; Step S230, input the first initial feature map and the second initial feature map into the migration network to obtain a second intermediate feature map by migrating the first initial feature map to the second initial feature map; and obtain a first intermediate feature map by migrating the second initial feature map to the first initial feature map; Step S240, input the first intermediate feature map and the second intermediate feature map into the encoder-decoder to obtain a decoded feature map; Step S250, input the decoded feature map into the output layer to obtain the output urinary calculus contour.
[0031] The output of the encoder is used as the input of the decoder. The main function of the input layer is to extract the feature map; the main function of the output layer is to segment the urinary calculus contour.
[0032] In step S210, the first training urinary calculus image of the first modality and the second training urinary calculus image of the second modality can be sourced from the hospital's database. For example, by extracting the first modality and second modality training urinary calculus images of the same patient in the database, M groups of first training urinary calculus images and second training urinary calculus images of M patients are extracted as the training set, which will not be elaborated here. It should be noted that the process of image preprocessing is not elaborated here. The preprocessing can be carried out based on experience, but since it is not the focus of this application, it will not be elaborated here.
[0033] In step S220, the input layer is a convolutional layer.
[0034] Input the first training urinary calculus image into the input layer to obtain a first initial feature map corresponding to the first training urinary calculus image output by the input layer.
[0035] Input the second training urinary calculus image into the input layer to obtain a second initial feature map corresponding to the first training urinary calculus image output by the input layer.
[0036] In step S230, the first initial feature map is input into the migration network, and then by migrating the first initial feature map towards the second initial feature map, a second initial feature map with the texture feature information in the first initial feature map is obtained; the second initial feature map is input into the migration network to migrate the second initial feature map towards the first initial feature map, that is, the first initial feature map with the texture feature information in the second initial feature map.
[0037] In step S240, the first intermediate feature map and the second intermediate feature map are input into the encoder-decoder to obtain a decoded feature map. Then, the decoded feature map is input into the output layer to obtain the output urinary stone contour.
[0038] It should be noted that when the training process is clear, the feature processing logic of the application process of the model to the image is similar to that of the training process, which will not be elaborated here.
[0039] Different from the existing feature fusion technology in multi-modal images, in this application, before encoding the images of the first modality and the second modality, the first initial feature map extracted from the images of the first modality is first input into the migration network. By migrating the first initial feature map towards the second initial feature map extracted from the images of the second modality, a second initial feature map (i.e., the second intermediate feature map) with the texture feature information in the first initial feature map is obtained. Then, the second initial feature map is input into the migration network to migrate the second initial feature map towards the first initial feature map, so as to obtain a first initial feature map (i.e., the first intermediate feature map) with the texture feature information in the second initial feature map, which fully realizes the complementarity of information between the images of the first modality and the second modality, and further improves the accuracy of the image segmentation model in identifying the stone area.
[0040] In some embodiments of this application, the migration network includes a region segmentation layer, a migration layer, and a fusion layer; In step S230 of this embodiment, inputting the first initial feature map and the second initial feature map into the migration network to migrate the first initial feature map towards the second initial feature map to obtain the second intermediate feature map includes the following steps S310 to S330: Step S310, input the first initial feature map and the second initial feature map into the region segmentation layer to obtain m first segmentation regions from the first initial feature map and n second segmentation regions from the second initial feature map; Step S320, input the m first segmentation regions and the n second segmentation regions into the migration layer, and according to the m first segmentation regions and the n second segmentation regions, use the migration algorithm to perform global migration and local migration of the first initial feature map towards the second initial feature map, respectively obtaining the global migration result and the local migration result; Step S330: Input the global migration result and the local migration result into the fusion layer to obtain the second intermediate feature map; Among them, a migration algorithm is used to perform global migration of the first initial feature map to the second initial feature map to obtain the global migration result, including: Perform pixel-by-pixel multiplication on the first initial feature map and the second initial feature map to obtain the first pixel-by-pixel multiplication result, and input the first pixel-by-pixel multiplication result into the AdaIN function to obtain the global migration result.
[0041] Among them, a migration algorithm is used to perform local migration of the first initial feature map to the second initial feature map to obtain the local migration result, including: Find the first segmentation region in the m first segmentation regions that is most similar to each second segmentation region; Calculate the pixel-by-pixel multiplication between the mask corresponding to each second segmentation region and the second initial feature map one by one to obtain the second pixel-by-pixel multiplication result, and calculate the pixel-by-pixel multiplication between the mask corresponding to the most similar first segmentation region of each second segmentation region and the first initial feature map one by one to obtain the third pixel-by-pixel multiplication result; Input the second pixel-by-pixel multiplication result and the third pixel-by-pixel multiplication result into the AdaIN function to obtain the migration result corresponding to each second segmentation region; Fuse the migration results corresponding to the n second segmentation regions to obtain the local migration result.
[0042] First, introduce the migration network. The output of the input layer of the segmentation network is used as the input of the migration network, that is, the input of the region segmentation layer. The purpose of the region segmentation layer is mainly to extract the first segmentation region from the first initial feature map. It should be noted that this is related to the preset label, because the region division label is a basic operation in this field for training data and will not be elaborated here; the region segmentation layer also extracts the second segmentation region from the second initial feature map.
[0043] The purposes of the migration layer include: 1) Migrate the features in the first initial feature map to the second initial feature map; 2) Migrate the features in the second initial feature map to the first initial feature map; In this embodiment, the processes of 1) and 2) both include global migration and local migration. Adding local migration on the basis of global migration here can enhance the consistency between local parts and improve the effect of image segmentation; The purpose of the fusion layer is to fuse the local migration result and the global migration result.
[0044] In step S310 of this embodiment, according to the preset tags, m first segmentation regions are obtained from the first initial feature map, and n second segmentation regions are obtained from the second initial feature map. The quantities m and n here are determined by the tags.
[0045] It should be noted that the first segmentation region is equal to the per-pixel multiplication between the corresponding mask and the first initial feature map, and the second segmentation region is equal to the per-pixel multiplication between the corresponding mask and the second initial feature map.
[0046] In step S320 of this embodiment, the m first segmentation regions and the n second segmentation regions are input into the migration layer, so as to perform global migration of the first initial feature map to the second initial feature map according to the m first segmentation regions and the n second segmentation regions by using the migration algorithm; Among them, when using the migration algorithm to perform global migration of the first initial feature map to the second initial feature map, the obtained global migration result includes: Performing per-pixel multiplication on the first initial feature map and the second initial feature map to obtain a first per-pixel multiplication result, and inputting the first per-pixel multiplication result into the AdaIN function to obtain the global migration result; AdaIN is a new adaptive instance normalization (AdaIN) layer that aligns the mean and variance of the content features with the mean and variance of the style features. For example, the content features and the style features are the first initial feature map and the second initial feature map respectively, and it specifically depends on the object of migration.
[0047] Global migration means performing per-pixel multiplication on the first initial feature map and the second initial feature map, and then inputting it into the AdaIN function, then a result is obtained.
[0048] Among them, when using the migration algorithm to perform local migration of the first initial feature map to the second initial feature map, the obtained local migration result includes: Finding the first segmentation region most similar to each second segmentation region among the m first segmentation regions; Calculating the per-pixel multiplication between the mask corresponding to each second segmentation region and the second initial feature map one by one to obtain a second per-pixel multiplication result, and calculating the per-pixel multiplication between the mask corresponding to the first segmentation region most similar to each second segmentation region and the first initial feature map one by one to obtain a third per-pixel multiplication result; Inputting the second per-pixel multiplication result and the third per-pixel multiplication result into the AdaIN function to obtain the migration result corresponding to each second segmentation region; Fusing the migration results corresponding to the n second segmentation regions to obtain the local migration result.
[0049] Taking the j-th second segmentation region among the n second segmentation regions as an example, it is necessary to find the first segmentation region that is most similar to the j-th second segmentation region among the m first segmentation regions. The purpose of finding this most similar region is to perform local migration between the j-th second segmentation region and the corresponding first segmentation region.
[0050] To find a first segmentation region corresponding to the j-th second segmentation region, a pre-trained model and a preset label can be used to find the most similar region between the two modalities.
[0051] Then, calculate the pixel-by-pixel multiplication between the mask corresponding to each second segmentation region and the second initial feature map one by one to obtain the second pixel-by-pixel multiplication result. Also, calculate the pixel-by-pixel multiplication between the mask corresponding to the first segmentation region that is most similar to each second segmentation region and the first initial feature map one by one to obtain the third pixel-by-pixel multiplication result. Input the second pixel-by-pixel multiplication result and the third pixel-by-pixel multiplication result into the AdaIN function to obtain the migration result corresponding to each second segmentation region. Fuse the migration results corresponding to the n second segmentation regions to obtain the local migration result.
[0052] In some embodiments of the present application, the migration network includes a region segmentation layer, a migration layer, and a fusion layer; Inputting the first initial feature map and the second initial feature map into the migration network in step S230 to obtain a first intermediate feature map by migrating the second initial feature map towards the first initial feature map includes the following steps S410 to S430: Step S410: Input the first initial feature map and the second initial feature map into the region segmentation layer to obtain m first segmentation regions from the first initial feature map and n second segmentation regions from the second initial feature map; Step S420: Input the m first segmentation regions and the n second segmentation regions into the migration layer to perform global migration and local migration of the second initial feature map towards the first initial feature map according to the m first segmentation regions and the n second segmentation regions, respectively obtaining the global migration result and the local migration result; Step S430: Input the global migration result and the local migration result into the fusion layer to obtain the first intermediate feature map; Among them, performing global migration of the second initial feature map towards the first initial feature map using the migration algorithm to obtain the global migration result includes: Performing pixel-by-pixel multiplication on the first initial feature map and the second initial feature map to obtain the first pixel-by-pixel multiplication result, and inputting the first pixel-by-pixel multiplication result into the AdaIN function to obtain the global migration result; Among them, a migration algorithm is adopted to perform local migration from the second initial feature map to the first initial feature map to obtain a local migration result, including: Find the second segmentation region that is most similar to each first segmentation region among the n second segmentation regions; Calculate the per-pixel dot product between the mask corresponding to each first segmentation region and the first initial feature map one by one to obtain a second per-pixel dot product result, and calculate the per-pixel dot product between the mask corresponding to the second segmentation region most similar to each first segmentation region and the second initial feature map one by one to obtain a third per-pixel dot product result; Input the second per-pixel dot product result and the third per-pixel dot product result into the AdaIN function to obtain the migration result corresponding to each first segmentation region; Fuse the migration results corresponding to the m first segmentation regions to obtain a local migration result.
[0053] It should be noted that this embodiment is similar to the above embodiment. The process of migrating the second initial feature map to the first initial feature map to obtain the first intermediate feature map in this embodiment is similar to the process of migrating the first initial feature map to the second initial feature map above, and will not be described in detail here.
[0054] This embodiment and the above embodiment innovatively use a migration algorithm. It can, while still retaining the two recognition routes for separately recognizing the images of the first modality and the second modality, before encoding the images of the first modality and the second modality, in the migration network, during the process of migrating the first initial feature map to the second initial feature map, global migration and local migration are used. Through the fusion between global information and local information, the second initial feature map can have the global information of the first initial feature map, the second initial feature map can have the local information of the first initial feature map, enhance the consistency between local information, and achieve the supplement of detailed texture features between images of different modalities; in the migration network, during the process of migrating the second initial feature map to the first initial feature map, global migration and local migration are used. Through the fusion between global information and local information, the first initial feature map can have the global information of the second initial feature map, the first initial feature map can have the local information of the second initial feature map, enhance the consistency between local information, and achieve the supplement of detailed texture features between images of different modalities; fully realize the complementarity of information between the images of the first modality and the second modality, and improve the accuracy of the image segmentation model in recognizing the calculus regions of different modality images respectively.
[0055] Such as Figure 2As shown, in some embodiments, the encoder-decoder includes a first encoder and a first decoder, as well as a second encoder and a second decoder; the first encoder and the second encoder have the same structure, and the first decoder and the second decoder have the same structure; The fused feature map between the feature map output by the first decoder and the feature map output by the second decoder serves as the decoded feature map.
[0056] The first encoder includes 6 convolution modules with a convolution kernel of 3 ×3, and 3 downsampling modules (shown as downsampling layers or upsampling layers). The first layer and the last layer of the first encoder are both convolution modules, and there are two connected convolution modules between every two adjacent downsampling modules. The first decoder includes 6 convolution modules with a convolution kernel of 3 ×3, and 3 upsampling modules. The first layer and the last layer of the first decoder are both convolution modules, and there are two connected convolution modules between every two adjacent upsampling modules; The input feature map of the j-th convolution module of the first encoder is the first fused feature map, and the input feature map of the j-th convolution module of the second encoder is the second fused feature map; wherein, the first fused feature map is the fused feature map between the output feature map of the downsampling module that is closest and before the j-th convolution module of the first encoder and the third fused feature map, and the third fused feature map is the fused feature map between the output feature map of the downsampling module that is closest and before the j-th convolution module of the first encoder and the output feature map of the downsampling module that is closest and before the j-th convolution module of the second encoder; the second fused feature map is the fused feature map between the output feature map of the downsampling module that is closest and before the j-th convolution module of the second encoder and the third fused feature map; j is 2, 4, 6; The input feature map of the i-th convolution module of the first encoder is the output feature map of the (i - 1)-th convolution module, and the input feature map of the i-th convolution module of the second encoder is the output feature map of the (i - 1)-th convolution module; i is 3, 5; The input feature map of the j-th convolutional module of the first decoder is the fourth fused feature map, and the input feature map of the j-th convolutional module of the second decoder is the fifth fused feature map. Among them, the fourth fused feature map is the feature map fused between the output feature map of the nearest upsampling module before the j-th convolutional module of the first decoder and the sixth fused feature map, and the sixth fused feature map is the feature map fused between the output feature map of the nearest upsampling module before the j-th convolutional module of the first decoder and the output feature map of the nearest upsampling module before the j-th convolutional module of the second decoder; the fifth fused feature map is the feature map fused between the output feature map of the nearest upsampling module before the j-th convolutional module of the second decoder and the sixth fused feature map. The input feature map of the i-th convolutional module of the first decoder is the output feature map of the (i - 1)-th convolutional module, and the input feature map of the i-th convolutional module of the second decoder is the output feature map of the (i - 1)-th convolutional module.
[0057] The structure of this embodiment includes two sets of identical structures, namely the first encoder - first decoder and the second encoder - second decoder. Among them, the first encoder or the second encoder extracts features through convolution and pooling operations, and the first decoder restores the image resolution through upsampling and convolution operations. In this embodiment, through the two sets of encoding - decoding structures, the separate feature learning of two different modality images of CT and MRI can be achieved simultaneously, and information sharing and interaction (i.e., the first intermediate feature map and the second intermediate feature map after migration) are realized again between the two sets of encoding - decoding structures, which can further learn the information sharing and interaction between the second initial feature map (i.e., the second intermediate feature map) with the texture feature information in the first initial feature map and the first initial feature map (i.e., the first intermediate feature map) with the texture feature information in the second initial feature map, and can improve the accuracy of stone area recognition.
[0058] In some embodiments, the output end of at least one convolutional module in the first encoder is residually connected to the input end of the convolutional module with the same feature map dimension in the first decoder, or the output end of at least one convolutional module in the second encoder is residually connected to the input end of the convolutional module with the same feature map dimension in the second decoder. By splicing the feature maps of the encoder and the decoder through skip connections, not only the detailed information is retained, but also the context information is combined, which helps to restore the details of the image and improve the segmentation accuracy.
[0059] It should be noted that the image segmentation model is completed after backpropagation based on the loss function until convergence, which is common knowledge in this field and will not be elaborated here.
[0060] In this embodiment, in order to improve the accuracy of the model in identifying stones, the loss function of the segmentation network is designed in this embodiment: The loss function of the segmentation network can be:
[0061] Where, represents the loss function; represents the similarity; represents the i-th first intermediate feature map sample; represents the j-th second intermediate feature map sample; represents the set of similar sample pairs; represents the set of dissimilar sample pairs; and are determined during data acquisition, that is, in the process of acquiring the first training urinary calculus image and the second training urinary calculus image.
[0062] The following describes the process of obtaining the similarity: (1) Divide the first intermediate feature map and the second intermediate feature map extracted from the first training urinary calculus image of the first modality and the second training urinary calculus image of the second modality into different groups of sample pairs; Since there are multiple first training urinary calculus images and second training urinary calculus images respectively, multiple first intermediate feature maps and second intermediate feature maps will be obtained, so they can be divided into multiple groups. One group includes one first intermediate feature map and one second intermediate feature map. Assuming there are m first intermediate feature maps and n second intermediate feature maps, there are mn groups.
[0063] (2) Calculate the similarity of each group of sample pairs; The calculation process includes: 1)
[0064] Where, is the feature map formed by splicing the i-th first intermediate feature map sample and the j-th second intermediate feature map sample; is the weight value; is the bias term; is the non-linear activation function.
[0065] 2)
[0066] Where, is the similarity between; is the weight value; is the bias term, is the Sigmoid activation function.
[0067] This embodiment uses to represent a set of
[0068] The advantage of the processing here is that: It can constrain the segmentation network to make the feature maps of similar samples closer and the feature maps of dissimilar samples more separated, thereby improving the fusion effect of the segmentation network on the feature maps.
[0069] This embodiment can be applied to the following scenarios: The target patient is diagnosed with urinary calculi, and the doctor recommends that the patient monitor regularly. The patient collected MRI images of urinary calculi using an MRI instrument and ct images using a CT instrument in different hospitals.
[0070] To assist the treatment of the patient, the doctor uses the image segmentation model provided in this application to identify the contours of the urinary calculi regions in the MRI images and ct images respectively, and identify the contours of the urinary calculi regions in the MRI images and the contours of the urinary calculi regions in the ct images; during the image processing process, information complementarity of the two-modal images is achieved, that is, when processing one modal image, the information in the other modal image will be migrated in, so that when processing one modal image, the information of its own modal and the information of other modalities can be extracted, thereby realizing the interaction of feature information.
[0071] Finally, an MRI or ct urinary calculi report is generated.
[0072] As Figure 3 shown, an embodiment of this application provides an image recognition method system for urinary calculi. The system includes the following structures: The data acquisition unit 1100 is used to acquire the first urinary calculi image to be recognized and the second urinary calculi image to be recognized of the target patient; the first urinary calculi image to be recognized is an image of the first modality, the second urinary calculi image to be recognized is an image of the second modality, and the first modality and the second modality are different; The model application unit 1200 is used to input the first urinary calculi image to be recognized and the second urinary calculi image to be recognized into a preset image segmentation model to obtain the urinary calculi contour output by the image segmentation model; The report generation unit 1300 is used to generate a corresponding recognition report according to the urinary calculi contour; Among them, the image segmentation model includes a segmentation network and a migration network, and the segmentation network includes an input layer, an encoder, a decoder, and an output layer; The model training unit 1400 is used to train the image segmentation model. The training process of the image segmentation model is as follows: Obtain the first training urinary calculus image of the first modality and the second training urinary calculus image of the second modality; Input the first training urinary calculus image and the second training urinary calculus image into the input layer to obtain the first initial feature map corresponding to the first training urinary calculus image and the second initial feature map corresponding to the second training urinary calculus image; Input the first initial feature map and the second initial feature map into the transfer network to obtain a second intermediate feature map by transferring the first initial feature map to the second initial feature map; and obtain a first intermediate feature map by transferring the second initial feature map to the first initial feature map; Input the first intermediate feature map and the second intermediate feature map into the encoder-decoder to obtain a decoded feature map; Input the decoded feature map into the output layer to obtain the output urinary calculus contour.
[0073] It should be noted that the urinary calculus image recognition system provided in this embodiment and the above-mentioned urinary calculus image recognition method are based on the same inventive concept. Therefore, the relevant content of the above-mentioned urinary calculus image recognition method also applies to the content of the urinary calculus image recognition system. Therefore, it will not be elaborated here.
[0074] Such as Figure 4 , this application embodiment also provides an electronic device, which includes: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes at least one program to implement the above-mentioned urinary calculus image recognition method of the present disclosure.
[0075] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0076] The following details the electronic device of the embodiment of the present application.
[0077] The processor 1600 can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present invention; The memory 1700 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1700 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1700 and are called by the processor 1600 to execute the image recognition method for urinary calculi in the embodiments of the present invention.
[0078] The input / output interface 1800 is used to implement information input and output; The communication interface 1900 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); The bus 2000 transmits information between various components of the device (such as the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900); Among them, the processor 1600, the memory 1700, the input / output interface 1800, and the communication interface 1900 are communicatively connected to each other inside the device through the bus 2000.
[0079] The embodiments of the present invention also provide a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the above-mentioned image recognition method for urinary calculi.
[0080] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0081] The embodiments described in the present invention are for more clearly explaining the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided in the embodiments of the present invention. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present invention are equally applicable to similar technical problems.
[0082] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present invention, and may include more or fewer steps than those shown, or combine certain steps, or different steps.
[0083] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0084] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.
[0085] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0086] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0087] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0088] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0089] In addition, each functional unit in various embodiments of this application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0090] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store programs.
[0091] The above is a specific description of the preferred implementation of the embodiments of this application. However, the embodiments of this application are not limited to the above-mentioned implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the embodiments of this application. These equivalent deformations or substitutions are all included in the scope defined by the claims of the embodiments of this application.
Claims
1. A method for image recognition of urinary stones, characterized in that: The method comprises the following steps: Acquire a first urinary stone to-be-identified image and a second urinary stone to-be-identified image of a target patient; the first urinary stone to-be-identified image is an image of a first modality, the second urinary stone to-be-identified image is an image of a second modality, and the first modality and the second modality are different; Inputting the first urinary stone to-be-identified image and the second urinary stone to-be-identified image into a preset image segmentation model to obtain a urinary stone contour output by the image segmentation model; generating a corresponding identification report according to the urinary stone outline; The image segmentation model includes a segmentation network and a migration network, and the segmentation network includes an input layer, an encoder-decoder, and an output layer; The training process of the image segmentation model is: Acquire a first training urinary stone image of the first modality and a second training urinary stone image of the second modality; Inputting the first training urinary calculus image and the second training urinary calculus image into the input layer to obtain a first initial feature map corresponding to the first training urinary calculus image and a second initial feature map corresponding to the second training urinary calculus image; Inputting the first initial feature map and the second initial feature map into the migration network, so as to obtain a second intermediate feature map by migrating the first initial feature map to the second initial feature map; and obtaining a first intermediate feature map by migrating the second initial feature map to the first initial feature map; Inputting the first intermediate feature map and the second intermediate feature map into the encoder-decoder to obtain a decoded feature map; The decoded feature map is input into the output layer to obtain the output urinary stone contour.
2. The method and system for image recognition of urinary stones according to claim 1, characterized in that: The input layer and the output layer are both convolutional layers.
3. The method and system for image recognition of urinary stones according to claim 2, characterized in that: The codec comprises a first encoder and a first decoder, and a second encoder and a second decoder; the first encoder and the second encoder have the same structure, and the first decoder and the second decoder have the same structure; The input data of the first encoder is the first intermediate feature map, the input data of the second encoder is the second intermediate feature map, and a fused feature map between the feature map output by the first decoder and the feature map output by the second decoder is used as the decoded feature map; The first encoder includes 6 convolution kernels with 3 3 convolution modules and 3 downsampling modules, the first layer and the last layer of the first encoder are both the convolution modules, and there are two connected convolution modules between each two adjacent downsampling modules, the first decoder includes 6 convolution kernels with 3 3 convolution modules and 3 up-sampling modules, the first layer and the last layer of the first decoder are both the convolution modules, and there are two connected convolution modules between every two adjacent up-sampling modules; The first encoder The input feature map of the jth convolution module of the second encoder is the first fused feature map, and the input feature map of the jth convolution module of the second encoder is the second fused feature map; wherein the first fused feature map is a feature map fused between the output feature map of the downsampling module that is before the jth convolution module of the first encoder and the closest one and the third fused feature map, and the third fused feature map is a feature map fused between the output feature map of the downsampling module that is before the jth convolution module of the first encoder and the closest one and the output feature map of the downsampling module that is before the jth convolution module of the second encoder and the closest one; the second fused feature map is a feature map fused between the output feature map of the downsampling module that is before the jth convolution module of the second encoder and the closest one and the third fused feature map; j is 2, 4, or 6; The input feature map of the i-th convolution module of the first encoder is the output feature map of the i-1-th convolution module, and the input feature map of the i-th convolution module of the second encoder is the output feature map of the i-1-th convolution module; i is 3 or 5; The input feature map of the j-th convolution module of the first decoder is the fourth fused feature map, and the input feature map of the j-th convolution module of the second decoder is the fifth fused feature map; wherein the fourth fused feature map is a feature map fused between the output feature map of the upsampling module that is before and closest to the j-th convolution module of the first decoder and the sixth fused feature map, and the sixth fused feature map is a feature map fused between the output feature map of the upsampling module that is before and closest to the j-th convolution module of the first decoder and the output feature map of the upsampling module that is before and closest to the j-th convolution module of the second decoder; the fifth fused feature map is a feature map fused between the output feature map of the upsampling module that is before and closest to the j-th convolution module of the second decoder and the sixth fused feature map; The input feature map of the i-th convolution module of the first decoder is the output feature map of the i-1-th convolution module, and the input feature map of the i-th convolution module of the second decoder is the output feature map of the i-1-th convolution module.
4. The method and system for image recognition of urinary stones according to claim 3, characterized in that: The output end of at least one of the convolution modules in the first encoder is residually connected to the input end of the convolution module with the same feature map dimension in the first decoder, or the output end of at least one of the convolution modules in the second encoder is residually connected to the input end of the convolution module with the same feature map dimension in the second decoder.
5. The method and system for image recognition of urinary stones according to claim 1, characterized in that: The migration network includes a region segmentation layer, a migration layer and a fusion layer; Inputting the first initial feature map and the second initial feature map into the migration network to obtain a second intermediate feature map by migrating the first initial feature map to the second initial feature map, including: Inputting the first initial feature map and the second initial feature map into a region segmentation layer to obtain m first segmentation regions from the first initial feature map and to obtain n second segmentation regions from the second initial feature map; Inputting the m first segmented regions and the n second segmented regions into the migration layer, so as to use a migration algorithm to perform global migration and local migration from the first initial feature map to the second initial feature map according to the m first segmented regions and the n second segmented regions, and obtain a global migration result and a local migration result respectively; Inputting the global migration result and the local migration result into the fusion layer to obtain the second intermediate feature map; The migration algorithm is used to perform global migration from the first initial feature map to the second initial feature map to obtain a global migration result, including: Performing pixel-by-pixel multiplication on the first initial feature map and the second initial feature map to obtain a first pixel-by-pixel multiplication result, and inputting the first pixel-by-pixel multiplication result into an AdaIN function to obtain a global migration result; The method of using a migration algorithm to locally migrate the first initial feature map to the second initial feature map to obtain a local migration result includes: Finding the first segmented region that is most similar to each second segmented region among the m first segmented regions; Calculate the pixel-by-pixel dot product between the mask corresponding to each of the second segmented regions and the second initial feature map one by one to obtain a second pixel-by-pixel dot product result, and calculate the pixel-by-pixel dot product between the mask corresponding to the first segmented region that is most similar to each of the second segmented regions and the first initial feature map one by one to obtain a third pixel-by-pixel dot product result; Inputting the second pixel-by-pixel dot product result and the third pixel-by-pixel dot product result into the AdaIN function to obtain a migration result corresponding to each of the second segmented regions; The migration results corresponding to the n second segmented regions are merged to obtain a local migration result.
6. The method and system for image recognition of urinary stones according to claim 1, characterized in that: The migration network includes a region segmentation layer, a migration layer and a fusion layer; Inputting the first initial feature map and the second initial feature map into the migration network to obtain a first intermediate feature map by migrating the second initial feature map to the first initial feature map, including: Inputting the first initial feature map and the second initial feature map into a region segmentation layer to obtain m first segmentation regions from the first initial feature map and to obtain n second segmentation regions from the second initial feature map; Inputting the m first segmented regions and the n second segmented regions into the migration layer, so as to use a migration algorithm to perform global migration and local migration from the second initial feature map to the first initial feature map according to the m first segmented regions and the n second segmented regions, and obtain a global migration result and a local migration result respectively; Inputting the global migration result and the local migration result into the fusion layer to obtain the first intermediate feature map; The migration algorithm is used to perform global migration from the second initial feature map to the first initial feature map to obtain a global migration result, including: Performing pixel-by-pixel multiplication on the first initial feature map and the second initial feature map to obtain a first pixel-by-pixel multiplication result, and inputting the first pixel-by-pixel multiplication result into an AdaIN function to obtain a global migration result; The method of using a migration algorithm to locally migrate the second initial feature map to the first initial feature map to obtain a local migration result includes: Finding the second segmented region that is most similar to each first segmented region among the n second segmented regions; Calculate the pixel-by-pixel dot product between the mask corresponding to each of the first segmented regions and the first initial feature map one by one to obtain a second pixel-by-pixel dot product result, and calculate the pixel-by-pixel dot product between the mask corresponding to the second segmented region that is most similar to each of the first segmented regions and the second initial feature map one by one to obtain a third pixel-by-pixel dot product result; Inputting the second pixel-by-pixel dot product result and the third pixel-by-pixel dot product result into an AdaIN function to obtain a migration result corresponding to each of the first segmented regions; The migration results corresponding to the m first segmented regions are merged to obtain a local migration result.
7. The method for image recognition of urinary stones according to claim 1, characterized in that: The first modality and the second modality are MRI and CT modalities, respectively.
8. A method and system for image recognition of urinary stones, characterized in that: The system comprises the following structure: A data acquisition unit, configured to acquire a first urinary calculus to-be-identified image and a second urinary calculus to-be-identified image of a target patient; the first urinary calculus to-be-identified image is an image of a first modality, the second urinary calculus to-be-identified image is an image of a second modality, and the first modality and the second modality are different; A model application unit, used for inputting the first urinary stone to-be-identified image and the second urinary stone to-be-identified image into a preset image segmentation model to obtain the urinary stone contour output by the image segmentation model; A report generating unit, used for generating a corresponding identification report according to the urinary stone outline; The image segmentation model includes a segmentation network and a migration network, and the segmentation network includes an input layer, an encoder, a decoder, and an output layer; The model training unit is used to perform a training process on the image segmentation model. The training process of the image segmentation model is: Acquire a first training urinary stone image of the first modality and a second training urinary stone image of the second modality; Inputting the first training urinary calculus image and the second training urinary calculus image into the input layer to obtain a first initial feature map corresponding to the first training urinary calculus image and a second initial feature map corresponding to the second training urinary calculus image; Inputting the first initial feature map and the second initial feature map into the migration network, so as to obtain a second intermediate feature map by migrating the first initial feature map to the second initial feature map; and obtaining a first intermediate feature map by migrating the second initial feature map to the first initial feature map; Inputting the first intermediate feature map and the second intermediate feature map into the encoder-decoder to obtain a decoded feature map; The decoded feature map is input into the output layer to obtain the output urinary stone contour.
9. An electronic device, characterized in that: It includes at least one control processor and a memory for communicating with the at least one control processor; the memory stores instructions that can be executed by the at least one control processor, and the instructions are executed by the at least one control processor so that the at least one control processor can execute the image recognition method for urinary stones described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the urinary stone image recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Medical image coloring method based on depth color migration
CN110458906A
Monocular depth estimation method based on multi-mode unsupervised image content decoupling
CN111445476A
Medical image segmentation method, medical image display method, medical image model training method, medical image display system, medical image model training equipment and medium
CN113450359A
Children intussusception diagnosis intelligent analysis system based on multi-modal fusion
CN114188021A
Medical image recognition method and device based on tumor marker, equipment and medium
CN114511569A