A bone imaging generation method based on an improved auxiliary classification generative adversarial network
Through the improved auxiliary classification generation adversarial network and dual-input gating attention structure, combined with dense residual convolution blocks and gradient penalty term loss function, the problems of scarcity of data and imbalance of sample categories in the multi-classification task of bone imaging are solved, and high-quality generation of bone imaging and improvement of multi-classification task are achieved.
Patent Information
- Application Number
- CN202310138371.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-02-20
AI Technical Summary
The prior art has problems of scarce data and unbalanced sample categories in the multi-classification task of bone imaging, resulting in low quality of the image generated by bone imaging and low computational efficiency and specificity.
The improved auxiliary classification generation adversarial network and dual-input gating attention structure are adopted, combining dense residual convolution blocks and gradient penalty term loss function, a U-shaped network generator and dense residual attention convolution block discriminator are constructed to improve the detailed feature generation and image quality of bone imaging.
The image quality of bone imaging is effectively improved, the training process is optimized, and the generated bone imaging is more similar to real bone imaging, improving the accuracy and efficiency of multi-classification tasks.
Smart Images

Figure CN116309362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning, and in particular to a bone imaging generation method based on an improved auxiliary classification generative adversarial network. Background Art
[0002] Bone metastasis has become a major cause of death in cancer patients worldwide. SPECT bone imaging has the advantages of low cost, high sensitivity, and wide imaging field of view, making it the most important examination method for evaluating bone metastasis. At present, doctors mainly rely on film reading experience and other auxiliary information to judge the type of bone imaging lesions, which has certain subjectivity, repetitiveness, and insufficient analysis.
[0003] In order to assist physicians in making rapid and objective diagnoses, many researchers have begun to apply machine learning algorithms and digital image processing technology to bone imaging to achieve automatic recognition in recent years. However, in general, although traditional digital image processing technology can achieve automatic recognition, the accuracy is not high; the classification accuracy of traditional machine learning algorithms has improved, but the computational efficiency and specificity are low.
[0004] In recent years, deep learning has gradually begun to be studied and applied in the field of medical image analysis. Among them, some scholars have begun to apply convolutional neural networks to bone imaging classification tasks. The literature "Zheng ZW, Liu LX, Chen XY, et al. Construction of bisection model of SPECT bone scan image based on VGGNet[C]. IEEE of the International Conference on Artificial Intelligence and Industrial Design(AIID), May 28-30, 2021, Guangzhou, China. New York: IEEE, 2021: 150-154." The small sample bone imaging dataset was enhanced by rotation, mirroring, and translation, and the classification effects of VGG models of different depths were compared. The accuracy, recall rate, and F1 score all reached more than 90%. The literature "Nikolaos P, Elpiniki P, Athanasios A, et al. Bone metastasis classification using whole body images from prostate cancer patients based on convolutional neural networks application [J]. Plos One, 2020, 15 (8): 1-29." fine-tuned the CNN hyperparameters to optimize the model performance, and then used the VGG16, ResNet50, GoogleNet and MobileNet models to judge the benign and malignant nature of bone images in prostate cancer patients.The paper "Charis N, Dimitrios E, Nikolaos P, et al. A lightweight Convolutional Neural Network architecture applied for bone metastasis classification in Nuclear Medicine: A Case Study on Prostate Cancer Patients [J]. Healthcare, 2020, 8 (4): 1-13." uses a lightweight fully convolutional neural network (LB-FCN) as the basic framework to reduce the amount of network computation, and reduces floating-point operations to further reduce the computational burden. Compared with ResNet, VGG16, and MobileNet, this framework has achieved better results in bone imaging classification tasks. The paper "Lin Q, Cao CH G, Li TT, et al. dSPIC: a deep SPECT image classification network for automated multi-disease, multi-lesion diagnosis [J]. BMC Medical Imaging, 2021, 21 (1): 1-17." proposes a custom CNN model to identify the types of bone imaging - healthy, bone metastasis, or arthritis, and uses traditional data enhancement and DCGAN for data enhancement, with a recognition accuracy of more than 70%. It can be seen that the current research mainly focuses on the classification method of a single disease. When multi-classification of bone imaging is performed, bone imaging has the problems of data scarcity and imbalanced sample categories. Summary of the invention
[0005] The present invention discloses a bone imaging generation method based on an improved auxiliary classification generative adversarial network, comprising the following steps: (1) preparing a bone imaging data set; (2) constructing a dense residual convolution block; (3) constructing an improved auxiliary classification generative adversarial network and a dual-input gated attention structure; (4) designing a loss function combined with a gradient penalty term; (5) training a network model and saving the parameters of the model after the training is completed; (6) generating bone imaging using the saved model to generate new bone imaging data.
[0006] The technical solution provided by the present invention is: a bone imaging generation method based on an improved auxiliary classification generative adversarial network, characterized by comprising the following steps:
[0007] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:
[0008] A bone imaging generation method based on an improved auxiliary classification generative adversarial network, characterized in that the calculation method comprises the following steps:
[0009] 1. A bone imaging generation method based on an improved auxiliary classification generative adversarial network, characterized by comprising the following steps:
[0010] Step 1: Preprocess the bone image. The specific processing method is as follows:
[0011] (1) Normalize the bone image and convert the original bone image file into a visible grayscale image;
[0012] (2) cropping the bone image and adjusting the image to N×3N pixel size, where N is a positive integer;
[0013] (3) Doctors classified bone images into three categories: healthy, malignant, and benign based on case information, and used bone images as training samples;
[0014] Where: x i represents the bone image of the input of the i-th layer, [x 0 ,x 1 ,....,x i-1 ] represents the concatenated features of all feature maps before the i-th layer; H i Represents a nonlinear mapping, which is a combination of batch normalization and PReLU activation function operations;
[0015] In order to effectively utilize the features extracted by the densely connected structure, after the third feature concatenation, a coordinated attention mechanism module is used to suppress the extracted useless features and let the model focus on the generation of detail features; finally, the residual structure is used to fuse the module extracted features with the input features;
[0016] Among them, the PReLU activation function is:
[0017]
[0018] Where: a is a constant greater than 0; x represents the bone imaging of the input PReLU activation function;
[0019] Step 3: Construct an improved auxiliary classification generative adversarial network with dual-input gated attention structure:
[0020] The generator uses an L-layer U-type network as the basic framework, where L is a positive integer, and includes three parts: the input end, the encoding part, and the decoding part. At the input end, the bone imaging category information is first embedded in the batch, and then the feature is concatenated with the noise to generate a new feature and reshape it into L features of different scales, which are used as the input of each layer of the U-type network encoding part. In the encoding part, the feature map is extracted and downsampled using the attention-intensive residual convolution block and the maximum pooling layer. At the same time, the downsampled features and the reshaped features at the input end are sent to the dual-input gated attention structure together. In the decoding part, the feature map is extracted and upsampled using the attention-intensive residual convolution and linear interpolation method. In addition, the dual-input gated attention structure is used at the jump connection to guide the high-level global feature x 1 With low-level detail features x 2 Perform efficient fusion; finally, use a 1×1 convolution to output the reconstructed bone image; the calculation process of the dual-input gated attention structure is as follows:
[0021] The high-level global feature x 1 With low-level detail features x 2 Perform feature concatenation, that is, superimpose the features in the channel dimension to obtain a new feature x c ; Then, the connection between the space and channels of the new feature map is learned by coordinating the attention mechanism, and the obtained weight matrix is multiplied by x c Get the updated features and compare them with x c Residual fusion is performed to retain some information of the input features; finally, after 1×1 convolution, batch normalization and PReLU activation function output module, the calculation process is as follows:
[0022] z c =concatetate(x 1 ,x 2 )
[0023] Where: Assume that the size of the two input features is C×H×W; concatetate() represents feature concatenation, that is, concatenating x in the channel dimension C. 1 、x 2 Perform superposition operation, let x 1 c 1 ×h×w、x 2 c 2 ×h×w, then concat(x 1 ,x in The calculation result of (c 1 +c 2 )×h×w;;z c Represents the features after splicing;
[0024]
[0025]
[0026] In the formula: c represents the number of channels; h represents the height; w represents the width; x c (h,i) indicates the i-th pixel when the channel is c high and h high; x c (j,w) indicates the jth pixel when the channel is c and the width is w; Indicates that in channel c, the w-th row of feature information at height h is summed and averaged; The calculation method is the same;
[0027] Indicates that in channel c, the w-th row of feature information at height h is summed and averaged; The calculation method is the same;
[0028]
[0029] Where: F1(·) represents the 1×1 convolution operation after feature concatenation; δ(·) represents the nonlinear transformation, which is the combination of batch normalization and PReLU activation function;
[0030]
[0031]
[0032] Where: f h 、f w They represent the vectors of f after dimension reduction along the h and w directions respectively; σ(·) represents the Sigmoid activation function operation;
[0033]
[0034] Where: x c (i,j) represents the bone image of the input coordinated attention mechanism; represents the attention weight; x' c (i,j) indicates the Represents the bone imaging feature map after the attention weight is updated;
[0035] x out =δ(Conv 1×1 (x c +x' c ))
[0036] Where: x c +x' c Indicates that x c and the updated x c 'Perform feature fusion; Conv1x1 () represents a 1×1 convolution operation; x out Represents the output characteristics of the entire module;
[0037] In the discriminator: first, the bone image is generated by inputting into the model; then, the feature information of the bone image is fully extracted through three dense residual attention convolution blocks; then, the extracted feature matrix is flattened, and its true and false information is output through the fully connected layer 1, and the category information of the bone image is output through the fully connected layer 2;
[0038] Step 4: Design the loss function and introduce the gradient penalty term based on the original loss of the auxiliary generative adversarial network. The improved network model objective function is as follows:
[0039] L S =E[logP(S=real|X real )]+E[logP(S=fake|X fake )]+L gp
[0040] L C =E[logP(C=c|X real )]+E[logP(C=c|X fake )]
[0041] Where: C represents the type of bone imaging; L S Indicates the loss of distinguishing true from false; L C Represents the loss of discriminative categories; X fake Indicates the generated bone image; X real represents the real bone image in the training set; E() represents the expectation; P() represents the probability;
[0042] L gp =λE[(||▽D(X fake )|| 2 -1) 2 ]
[0043] Where: λ is a constant; ▽ represents the gradient, L gp is the gradient penalty term;
[0044] The generator and discriminator losses of the model are as follows:
[0045] L G =L s -L c
[0046] L D =L c +L s
[0047] Where: L Grepresents the generator loss; L D represents the discriminator loss;
[0048] Step 5: Send the test set data preprocessed in step 1 to the network model built in step 4, calculate the loss value using the improved loss function in step 5, and evaluate the generation quality of bone imaging by the structural similarity coefficient, and save the final network model, which is recorded as MU-ACGAN;
[0049] The formula for calculating the structural similarity coefficient is as follows:
[0050]
[0051] Where: represents the grayscale variance between the generated bone image and the real bone image, σ xy represents covariance; μ x , μ y represents the mean value of pixels; c 1 、c 2 is a constant; SSIM represents the similarity between the bone image generated by the image and the real bone image, ranging from [0,1]. The closer the value is to 1, the closer it is to the real bone image;
[0052] Step 6: Input random noise into the saved MU-ACGAN and output the newly generated bone image.
[0053] The present invention provides an improved bone imaging generation method based on an auxiliary classification generative adversarial network. The method uses a dense residual attention convolution block as a basic convolution unit, and proposes an improved bone imaging generation method based on an auxiliary classification generative adversarial network. The method uses U-Net as a generator framework, and combines dense residual connection convolution blocks and dual-input gated attention structures to improve the generation of bone imaging detail features. The discriminator extracts bone imaging features through dense residual connection convolution blocks for discrimination. In addition, a gradient penalty is introduced into the training loss function to improve the stability of model training. The bone imaging generation method based on an auxiliary classification generative adversarial network disclosed by the present invention effectively improves the image quality of the generated bone imaging while generating three types of bone imaging: healthy, malignant, and benign changes, and can provide more bone imaging training data for subsequent bone imaging recognition tasks.
[0054] Beneficial effects:
[0055] Compared with the current mainstream bone imaging generation method, the present invention has the following beneficial effects:
[0056] (1) Improving the bone imaging generation method of auxiliary classification generative adversarial network can generate bone imaging data of different categories, such as healthy, malignant, and benign changes;
[0057] (2) Compared with the generation effect of the original auxiliary classification generative adversarial network, the improved auxiliary classification generative adversarial network improves the image generation quality of bone imaging and optimizes the training process through the dual-input gated attention structure, dense residual convolution block and the loss function with gradient penalty term, and can generate images that are more similar to real bone imaging; BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 The flowchart of the present invention is as follows; the raw data is preprocessed and grayed, and input into the improved auxiliary classification generative adversarial network for training, the final network model is saved, random noise and categories are input into the model again, and finally bone images of different categories are output;
[0059] Figure 2 Generate a schematic diagram of the adversarial network for auxiliary classification; input random noise and categories to the generator, output the generated samples, and send the generated samples to the discriminator for category and true and false discrimination;
[0060] Figure 3 It is a diagram of a dual-input gated attention structure; high-level features and low-level features are input into the dual-input gated attention structure to update feature weights;
[0061] Figure 4 is a dense residual convolution block diagram, CA represents coordinated attention mechanism, and BN represents batch normalization;
[0062] Figure 5 Generate adversarial network graphs for improved auxiliary classification;
[0063] Skip connection: skip connection operation;
[0064] RDA block: dense residual convolution block;
[0065] Max pool: maximum pooling layer;
[0066] Up sampling: Up sampling;
[0067] RCAG: Dual-input gated attention architecture
[0068] 1×1conv: 1×1 convolution operation;
[0069] Figure 6 Visualization of the bone image generation process; input random noise and visualize the generated bone images of the 1st, 100th, 200th and 500th rounds of training;
[0070] Figure 7The generated bone image is compared with the real bone image; from top to bottom, the first line is the bone image generation effect of the auxiliary classification generation adversarial network, the second line is the bone image generation effect of the present invention, and the third line is the real bone image; DETAILED DESCRIPTION
[0071] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 The specific implementation of the present invention is as follows:
[0072] Step 1: Preprocess the bone image. The specific processing method is as follows:
[0073] (1) Normalize the bone image and convert the original bone image file into a visible grayscale image;
[0074] (2) The bone images were cropped and the image size was adjusted to 256 × 768 pixels;
[0075] (3) Use three types of bone images, namely healthy, malignant, and benign changes, as training samples;
[0076] Step 2: Construct a dense residual convolution block, which consists of 3×3 convolution, 1×1 convolution, PReLU activation function, batch normalization and coordinated attention mechanism; first, the module extracts different bone imaging features at a unified scale through three 3×3 convolutions, and batch normalization and PReLU activation function are used after each convolution; secondly, the module aggregates different feature information at a unified scale through dense connections; at the same time, a 1×1 convolution is used to compress the number of channels after each feature aggregation; the attention dense residual convolution block is calculated as follows:
[0077] x i =H i ([x 0 ,x 1 ,…,x i-1 ]),i∈[1,M]
[0078] Where: x i represents the bone image of the input of the i-th layer, [x 0 ,x 1 ,....,x i-1 ] represents the concatenated features of all feature maps before the i-th layer; H i Represents a nonlinear mapping, which is a combination of batch normalization and PReLU activation function operations;
[0079] In order to effectively utilize the features extracted by the densely connected structure, after the third feature concatenation, a coordinated attention mechanism module is used to suppress the extracted useless features and let the model focus on the generation of detail features; finally, the residual structure is used to fuse the module extracted features with the input features;
[0080] Among them, the PReLU activation function is:
[0081]
[0082] Where: a is a constant greater than 0; x represents the bone imaging of the input PReLU activation function;
[0083] Step 3: Construct an improved auxiliary classification generative adversarial network with dual-input gated attention structure:
[0084] The generator uses an L-layer U-type network as the basic framework, where L is a positive integer, and includes three parts: the input end, the encoding part, and the decoding part. At the input end, the bone imaging category information is first embedded in the batch, and then the feature is concatenated with the noise to generate a new feature and reshape it into L features of different scales, which are used as the input of each layer of the U-type network encoding part. In the encoding part, the feature map is extracted and downsampled using the attention-intensive residual convolution block and the maximum pooling layer. At the same time, the downsampled features and the reshaped features at the input end are sent to the dual-input gated attention structure together. In the decoding part, the feature map is extracted and upsampled using the attention-intensive residual convolution and linear interpolation method. In addition, the dual-input gated attention structure is used at the jump connection to guide the high-level global feature x 1 With low-level detail features x 2 Perform efficient fusion; finally, use a 1×1 convolution to output the reconstructed bone image; the calculation process of the dual-input gated attention structure is as follows:
[0085] The high-level global feature x 1 Concatenate the features with the low-level detail features, that is, superimpose the features in the channel dimension to obtain a new feature x c ; Then, the connection between the space and channels of the new feature map is learned by coordinating the attention mechanism, and the obtained weight matrix is multiplied by x c Get the updated features and compare them with x c Residual fusion is performed to retain some information of the input features; finally, after 1×1 convolution, batch normalization and PReLU activation function output module, the calculation process is as follows:
[0086] z c =concatetate(x 1 ,x 2 )
[0087] Where: Assume that the size of the two input features is C×H×W; concatetate() represents feature concatenation, that is, concatenating x in the channel dimension C. 1 、x 2 Perform superposition operation, let x1 c 1 ×h×w、x 2 c 2 ×h×w, then concat(x 1 ,x in The calculation result of (c 1 +c 2 )×h×w;z c Represents the features after splicing;
[0088]
[0089]
[0090] In the formula: c represents the number of channels; h represents the height; w represents the width; x represents the width c (h,i) indicates the i-th pixel when the channel is c high and h high; x c (j,w) indicates the jth pixel when the channel is c and the width is w; Indicates that in channel c, the feature information of the wth row of height h is summed and averaged; The calculation method is the same;
[0091] Indicates that in channel c, the feature information of the wth row of height h is summed and averaged; The calculation method is the same;
[0092]
[0093] Where: F1(·) represents the 1×1 convolution operation after feature concatenation; δ(·) represents the nonlinear transformation, which is the combination of batch normalization and PReLU activation function;
[0094]
[0095]
[0096] Where: f h 、f w They represent the vectors of f after dimension reduction along the h and w directions respectively; σ(·) represents the Sigmoid activation function operation;
[0097]
[0098] Where: x c (i,j) represents the bone image of the input coordinated attention mechanism; represents the attention weight; x' c (i,j) indicates the Represents the bone imaging feature map after the attention weight is updated;
[0099] x out =δ(Conv 1×1 (x c +x' c ))
[0100] Where: x c +x' c Indicates that x c and the updated x c 'Perform feature fusion; Conv 1x1 () represents a 1×1 convolution operation; x out Represents the output characteristics of the entire module;
[0101] In the discriminator: first, the bone image is generated by inputting into the model; then, the feature information of the bone image is fully extracted through three dense residual attention convolution blocks; then, the extracted feature matrix is flattened, and its true and false information is output through the fully connected layer 1, and the category information of the bone image is output through the fully connected layer 2;
[0102] Step 4: Design the loss function and introduce the gradient penalty term based on the original loss of the auxiliary generative adversarial network. The improved network model objective function is as follows:
[0103] L S =E[logP(S=real|X real )]+E[logP(S=fake|X fake )]+L gp
[0104] L C =E[logP(C=c|X real )]+E[logP(C=c|X fake )]
[0105] Where: L S Indicates the loss of distinguishing true from false; L C Represents the loss of discriminative categories; X fake Indicates the generated bone image; X real represents the real bone images in the training set;
[0106]
[0107] Where: λ is a constant, in this example, λ is 10; ▽ represents the gradient, L gp is the gradient penalty term;
[0108] The generator and discriminator losses of the model are as follows:
[0109] L G =L s -Lc
[0110] L D =L c +L s
[0111] Where: L G represents the generator loss; L D represents the discriminator loss;
[0112] Step 5: Send the test set data preprocessed in step 1 to the network model built in step 3, calculate the loss value using the weighted loss function designed in step 4, and evaluate the generation quality of bone imaging using the structural similarity coefficient. Save the final network model, which is recorded as MU-ACGAN.
[0113] The formula for calculating the structural similarity coefficient is as follows:
[0114]
[0115] Where: represents the grayscale variance between the generated bone image and the real bone image, σ xy represents covariance; μ x , μ y represents the pixel mean; c 1 、c 2 is a constant greater than 0; SSIM represents the similarity between the bone image generated by the image and the real bone image, ranging from [0,1]. The closer the value is to 1, the closer it is to the real bone image;
[0116] Step 6: Input random noise into the saved MU-ACGAN and output three types of bone images: healthy, malignant, and benign changes.
[0117] The effect of the implementation method of the present invention is shown in the following table, which shows the bone imaging generation effect of the auxiliary classification generation adversarial network and the method of the present invention, as shown in Table 1:
[0118] Table 1 Average intersection-union ratios of different algorithms
[0119]
[0120] It can be seen from the above table that the improved auxiliary classification generative adversarial network proposed in the present invention has a better effect than the original auxiliary classification generative adversarial network in generating healthy, malignant and benign bone images with structural similarity coefficients increased by 0.0805 (17.6%), 0.0412 (8.1%) and 0.0417 (8.2%) respectively.
Claims
1. A bone imaging generation method based on an improved auxiliary classification generative adversarial network, characterized in that it includes the following steps: Step 1: Preprocess the bone imaging, and the specific processing method is: (1) Normalize the bone imaging to convert the original bone imaging file into a visible grayscale image; (2) Crop the bone imaging and adjust the image to a size of N×3N pixels, where N is a positive integer; (3) The doctor classifies the bone imaging into three categories: healthy, malignant, and benign according to the case information, and uses the bone imaging as a training sample; Step 2: Construct a dense residual convolutional block, which is composed of a 3×3 convolution, a 1×1 convolution, a PReLU activation function, batch normalization, and a coordinated attention mechanism; First, this module extracts different bone imaging features at a unified scale through 3 3×3 convolutions, and batch normalization and the PReLU activation function are used after each convolution; Secondly, the module uses a dense connection method to converge different feature information at a unified scale; At the same time, a 1×1 convolution is used to compress the number of channels after each feature convergence; The calculation of the attention dense residual convolutional block is as follows: x i = H i ([x 0 , x 1 , …, x i-1 ), i ∈ [1, M] where: x i represents the bone scintigraphy of the input of the i-th layer, [x 0 , x 1 ,...., x i-1 represents the feature after concatenating all the feature maps before the i-th layer; H i represents the non-linear mapping, that is, the combination of batch normalization and PReLU activation function operations; To effectively utilize the features extracted by the dense connection structure, after the 3rd feature splicing, a coordinated attention mechanism module is used to suppress the useless features extracted, so that the model focuses on the generation of detailed features; Finally, a residual structure is used to fuse the features extracted by the module with the input features; Among them, the PReLU activation function is: In the formula: a is a constant greater than 0; x represents the bone imaging input to the PReLU activation function; Step 3: Construct an improved auxiliary classification generative adversarial network and a dual-input gated attention structure: The generator uses an L-layer U-shaped network as the basic framework, where L is a positive integer, and it includes three parts: an input end, an encoding part, and a decoding part. At the input end, first, the bone scintigraphy category information is embedded in the batch, and then it is concatenated with noise in terms of features to generate a new feature and reshape it into L features of different scales, which are respectively used as the input of each layer of the encoding part of the U-shaped network. In the encoding part, an attention dense residual convolution block and a max pooling layer are used to extract features and downsample the feature map. At the same time, the downsampled features and the reshaped features at the input end are sent into a dual-input gated attention structure. In the decoding part, an attention dense residual convolution and linear interpolation are used to extract features and upsample the feature map. In addition, a dual-input gated attention structure is adopted at the skip connection to guide the high-level global feature x 1 and the low-level detailed feature x 2 for efficient fusion. Finally, a 1×1 convolution is used to output the reconstructed bone scintigraphy. Among them, the calculation process of the dual-input gated attention structure is as follows: The high-level global feature x 1 and the low-level detailed feature x 2 are subjected to feature concatenation, that is, the features are stacked in the channel dimension to obtain a new feature x c ; then, the relationship between the space and channels of the new feature map is learned through the coordinated attention mechanism, and the obtained weight matrix is multiplied by x c to obtain the updated feature, and it is subjected to residual fusion with x c to retain partial information of the input feature; finally, through the 1×1 convolution, batch normalization, and PReLU activation function output module, the calculation process is as follows: z c = concatenate(x 1 , x 2 ) Where: Assume that the sizes of two input features are both C×H×W; concatetate(·) represents feature concatenation, that is, superimposing operations on x 1 , x 2 . Assume that x 1 is c 1 ×h×w, and x 2 is c 2 ×h×w. Then the calculation result of concat(x 1 , x in ) is (c 1 + c 2 )×h×w; z c represents the concatenated feature; Where: c represents the number of channels; h represents the height; w represents the width; x c (h, i) represents the i-th pixel point when the channel is c and the height is h; x c (j, w) represents the j-th pixel point when the channel is c and the width is w; represents summing and averaging the feature information of the w-th row of height h in channel c; The calculation method is the same; It means to sum and average the feature information of the w-th row with height h in channel c; The calculation method is the same; In the formula: F1(·) represents performing a 1×1 convolution operation after feature splicing; δ(·) represents a non-linear transformation, that is, a combination of batch normalization and the PReLU activation function; where: f h and f w respectively represent the vectors obtained by reducing the dimension of f along the h and w directions; σ(·) represents the Sigmoid activation function operation; Where: x c (i, j) represents the bone scintigraphy input to the coordinated attention mechanism; represents the attention weight; x' c (i, j) represents after represents the bone scintigraphy feature map after the attention weight update; x out = δ(Conv 1×1 (x c + x' c )) Where: x c + x' c represents the feature fusion of x c and the updated x c '; Conv 1x1 (·) represents a 1×1 convolution operation; x out represents the output feature of the entire module; In the discriminator: First, input the generated bone imaging into the model; Then, fully extract the feature information of the bone imaging through 3 dense residual attention convolutional blocks; Next, flatten the extracted feature matrix, and output its true / false information through the fully connected layer 1 and output the category information of the bone imaging through the fully connected layer 2; Step 4: Design a loss function combined with a gradient penalty term, introduce a gradient penalty term on the basis of the original loss of the auxiliary generative adversarial network, and the objective function of the improved network model is as follows: L S = E[log P(S = real|X real )] + E[log P(S = fake|X fake )] + L gp L C = E[log P(C = c|X real )] + E[log P(C = c|X fake )] Where: C represents the type of bone imaging; L S represents the loss for discriminating true or false; L C represents the loss for discriminating categories; X fake represents the generated bone imaging; X real represents the true bone imaging in the training set; E() represents expectation; P(·) represents probability; where: λ is a constant; denotes the gradient, and L gp is the gradient penalty term; The losses of the generator and discriminator of the model are as follows: L G = L s - L c L D = L c + L s Where: L G represents the generator loss; L D represents the discriminator loss; Step 5: Send the test set data preprocessed in Step 1 into the network model built in Step 4, calculate the loss value using the improved loss function in Step 5, and evaluate the generation quality of the bone imaging through the structural similarity coefficient, and save the final network model, denoted as MU-ACGAN; The calculation formula of the structural similarity coefficient is as follows: In the formula: represents the gray variance between the generated bone scintigraphy and the true bone scintigraphy, σ xy represents the covariance; μ x and μ y represent the pixel point mean value; c 1 and c 2 are constants; SSIM represents the similarity between the generated bone scintigraphy and the true bone scintigraphy of the picture, ranging from [0, 1]. The closer its value is to 1, the closer it is to the true bone scintigraphy; Step 6: Input random noise into the saved MU-ACGAN and output a newly generated bone imaging.