Non-tracking freestyle three-dimensional ultrasonic image reconstruction method based on deep learning
Through deep learning, the three-dimensional ultrasonic image reconstruction method of ultrasonic sequence coding, wavelet transform convolution and hybrid attention modules is constructed, which solves the problems of heavy equipment and insufficient reconstruction accuracy in the traditional method, and realizes efficient and accurate ultrasonic image reconstruction to adapt to image processing under different conditions.
Patent Information
- Application Number
- CN202510240753.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-08
AI Technical Summary
The existing 3D ultrasound reconstruction methods rely on motion sensors, resulting in bulky and inconvenient equipment, and insufficient reconstruction accuracy when processing complex scanning paths, manual reconstruction methods are time-consuming and inconsistent, and the existing sensorless methods cannot meet clinical needs when complex and diverse ultrasound images.
Using the trackless freestyle three-dimensional ultrasonic image reconstruction method based on deep learning, the ultrasonic sequence coding module, wavelet transform convolution module and hybrid attention module are constructed, combining the root mean square distance and Pearson correlation loss function, spatial relationship modeling and multi-scale feature extraction between image frames are realized, improving reconstruction accuracy and efficiency.
It improves the efficiency and accuracy of three-dimensional reconstruction of ultrasound images, enhances the adaptability and robustness of the model under different conditions, reduces the dependence on high-end equipment, and provides more accurate diagnostic information.
Smart Images

Figure CN120279167A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a tracking-free free-form three-dimensional ultrasound image reconstruction method based on deep learning. Background Art
[0002] In recent years, 3D ultrasound imaging technology has made remarkable progress in the field of medical imaging and has been widely used in multiple fields such as cardiac, obstetric, abdominal, and vascular examinations. The freehand 3D ultrasound reconstruction technology overcomes the limitations of traditional array probes in terms of field of view and acquisition rate by stitching multiple 2D ultrasound cross-sections, providing a larger and more flexible field of view. However, accurately estimating the scanning trajectory of the probe to maintain high reconstruction quality remains a major technical challenge.
[0003] Traditional 3D ultrasound reconstruction methods usually rely on motion sensors to track the movement of the probe, which is cumbersome and inconvenient in clinical applications. Although deep neural networks (DNNs) and 3D convolutional networks (CNNs) have been introduced to estimate the position of the imaging plane, these methods are not ideal when dealing with out-of-plane (OOP) motion, especially in complex scanning paths. Although existing sensorless methods can provide certain reconstruction effects in some cases, they often cannot meet the clinical requirements when facing the complexity and diversity of ultrasound images.
[0004] With the rapid growth of medical imaging data, traditional manual reconstruction methods have become increasingly impractical. Manual operation is not only time-consuming and labor-intensive but also easily affected by the skills and experience of the operator, resulting in inconsistent and subjective results. Automated reconstruction methods can significantly improve the operation efficiency, reduce human errors, and provide more consistent and objective results. Neural networks play a key role in 3D ultrasound reconstruction and can achieve more efficient reconstruction by learning image features and patterns. Especially when dealing with ultrasound images of different qualities and features, deep learning methods demonstrate strong adaptability and robustness. However, existing methods still have certain limitations in dealing with complex non-linear scanning trajectories, resulting in insufficient reconstruction accuracy. Therefore, there is an urgent need to provide a new sensorless 3D ultrasound reconstruction method to effectively solve the deficiencies of the existing technology. Summary of the Invention
[0005] Aiming at the problems existing in the above-mentioned prior art, the present invention provides a tracking-free free-form three-dimensional ultrasound image reconstruction method based on deep learning. This method can significantly improve the efficiency and accuracy of three-dimensional reconstruction of ultrasound images. At the same time, it can improve the adaptability and robustness in dealing with ultrasound images under different conditions and provide more accurate diagnostic information for clinical use.
[0006] To achieve the above object, the present invention provides a method for reconstructing a free-style three-dimensional ultrasound image based on deep learning, comprising the following steps:
[0007] Step 1: Obtain an image data set;
[0008] Use an image acquisition system to scan similar diagnostic regions of several patients with a non-linear scanning trajectory to obtain two-dimensional ultrasound image sequence data. Store the two-dimensional ultrasound image sequence data to form an image data set, and divide it into a training set, a test set, and a validation set according to a set ratio;
[0009] Step 2: Construct a three-dimensional ultrasound image reconstruction model;
[0010] S21: Construct an ultrasound sequence encoding module;
[0011] Use the frame index (i, j) and the sequence length m as hyperparameters, so that the recurrent neural network f sequentially obtains image frames from the sequence based on the context relationship of the ultrasound sequence, and models the spatial relationship between ultrasound image frames by predicting the spatial transformation matrix between frames, obtaining an ultrasound sequence encoding module for subsequent frame sequence encoding;
[0012] S22: Construct a wavelet transform convolution module;
[0013] Use the Haar wavelet transform combined with small kernel depth convolution to construct a wavelet transform convolution module; in the wavelet transform convolution module, use the Haar wavelet transform to filter and downsample the input low-frequency and high-frequency content, then perform small kernel depth convolution on different frequency maps, and finally use the inverse wavelet transform to construct the output;
[0014] S23: Construct a hybrid attention module;
[0015] Modify the MBConv block network based on the EfficientNet architecture. In the first MBConv block of each stage, replace the original SE module with a CoordAtt module, and at the same time, retain the original SE module in the second MBConv block, and then construct a hybrid attention module by combining the CoordAtt module and the SE module; in the hybrid attention module, the CoordAtt module is used to capture global information through adaptive pooling operations to generate a spatial attention map, and the SE module is used to enhance the expression ability of the feature map by learning the weights between channels. By combining spatial and channel-level information, the focusing ability and recognition ability for key image regions are enhanced;
[0016] S24: Construct a loss function;
[0017] Construct a loss function by combining the weighted sum of the root mean square distance loss function and the Pearson correlation loss function;
[0018] S25: Obtain a three-dimensional ultrasound image reconstruction model through training;
[0019] Construct a deep learning model by combining an ultrasound sequence encoding module, a wavelet transform convolutional module, and a hybrid attention module, and use a training set to train the deep learning module. During the training process, continuously update the model parameters through the backpropagation algorithm to reduce the value of the loss function, and verify the results through a validation set to select the optimal parameters, thereby obtaining a trained deep learning model; use a test set to test the trained deep learning model to obtain a three-dimensional ultrasound image reconstruction model;
[0020] Step 3: Reconstruct the three-dimensional ultrasound image;
[0021] Use an image acquisition system to scan the patient's diagnosis and treatment area with a non-linear scanning trajectory to obtain two-dimensional ultrasound image sequence data;
[0022] Use the two-dimensional ultrasound image sequence data as input data to input into the three-dimensional ultrasound image reconstruction model for image reconstruction. During the image reconstruction process, use the ultrasound sequence encoding module to encode the context relationship between image frames to capture the context relationship in the ultrasound image sequence, use the wavelet transform convolutional module to extract features from the image sequence data. In this process, extract multi-scale features by cascading wavelet transform and small kernel depth convolution to pay attention to the features in different frequency bands, use the hybrid attention module to enhance the recognition ability of the focusing ability of the key areas affected by speckle noise in the image to improve the reconstruction accuracy, and finally output a high-quality three-dimensional ultrasound image to provide accurate image information for clinical diagnosis.
[0023] As an optimization, in step 1, the process of collecting two-dimensional ultrasound image sequence data is as follows:
[0024] S11: Use a computer, an ultrasound scanning device, an optical tracker, and an ultrasound probe to form an image acquisition system, where the computer is respectively connected to the ultrasound scanning device and the optical tracker, and the ultrasound scanning device is connected to the ultrasound probe;
[0025] S12: Use the ultrasound probe to scan the patient's target diagnosis and treatment area. During the scanning process, randomly use a linear scanning trajectory, a C-shaped scanning trajectory, or an S-shaped scanning trajectory for scanning operations. The ultrasound scanning device sends the scanned image to the computer. At the same time, during the scanning process, use the optical tracker to accurately position and track the ultrasound probe, and obtain the spatial position information of the ultrasound probe, and then send the spatial position information to the computer; during the scanning process, the ultrasound frames are recorded at a speed of 20fps. At the same time, the size of each frame of image is 480×640 pixels. At the same time, no speckle reduction processing is performed;
[0026] S13: The computer integrates the image data and the spatial position information according to the timing information of the spatial position information and the timing information of the image, and sorts the integrated data in chronological order to form two-dimensional ultrasound image sequence data; a number of two-dimensional ultrasound image sequence data are obtained through 1200 scanning processes, and the number of two-dimensional ultrasound image sequence data is stored in time to obtain an image data set.
[0027] As an optimization, in S24 of step two, the process of constructing the loss function is as follows:
[0028] S24-1: Measure the correlation between the predicted features and the true features based on the Pearson correlation coefficient, and use formula (1)
[0029] to calculate the Pearson correlation loss function corr(preds, labels);
[0030]
[0031] In the formula, Cov(pred i , label i ) is the covariance between the predicted value pred i and the true label label i , and σ(preds′) and σ(labels′) respectively represent the standard deviations of the predicted value and the true label;
[0032] S24-2: Combine the root mean square distance loss function and the Pearson correlation loss function corr(preds, labels), and use formula (2) to obtain the loss function L total ;
[0033] L total = L RMSD + λ(1 - corr(preds, labels)) (2);
[0034] In the formula, L RMSD is the root mean square distance loss function, and λ is the weight parameter that balances the root mean square distance loss function and the correlation loss function.
[0035] As an optimization, in step three, the process of encoding the context relationship between image frames by using the ultrasound sequence encoding module is as follows:
[0036] A1: For a given two-dimensional ultrasound image sequence, perform M samplings to obtain an ultrasound image sequence with a sequence length of m,
[0037] where m = 1, 2,..., M;
[0038] A2: For any pair of frame indices (i, j) in the two-dimensional ultrasound image sequence M, where i and j are not adjacent frames, i.e., i ≠ j - 1, when m ≠ M, the spatial transformation matrix T between the i-th frame and the j-th frame is predicted according to formula (3). i-j , the spatial transformation matrix T i-j is recursively predicted at the end of each sequence pair, using the context information from the ultrasound image sequence;
[0039] T i-j = f(S m , I (m-1) ; θ) (3);
[0040] In the formula, m represents the time step, S m represents the input information at time step m in the ultrasound sequence, I (m-1) represents the internal hidden state at time step m - 1, and θ is the parameter of the neural network f.
[0041] As an optimization, in step three, the process of multi-scale feature extraction of the image sequence data by using the wavelet transform convolution module is as follows:
[0042] B1: For the ultrasound image frame x[n], it is processed using the low-pass filter g[n] and the high-pass filter h[n], and the low-frequency component x L [n] is obtained according to formula (4), and the high-frequency component x H [n] is obtained according to formula (5). Through the above method, the ultrasound image frame x[n] is decomposed twice using the Haar wavelet transform to obtain the low-frequency component x L [n] and three high-frequency components x H [n];
[0043] x L [n]= ∑ k x[2n - k]·g[k] (4);
[0044] x H [n]= ∑ k x[2n - k]·h[k] (5);
[0045] In the formula, n is the position of the pixel, g[k] is the coefficient of the low-pass filter g[n], and h[k] is the coefficient of the high-pass filter h[n];
[0046] B2: Downsample the low-frequency component x L [n] and the three high-frequency components x H [n] to reduce the image resolution;
[0047] B3: In the low-frequency component x L [n] and the high-frequency component x HPerform depthwise convolution on [n] respectively to extract features;
[0048] B4: Perform inverse transformation on the low-frequency component x L [n] and the high-frequency component x H [n], and obtain the output Y;
[0049] Y = IWT(Conv(W, WT(X))) (6);
[0050] Where X is the input tensor, and W is the k×k depth kernel weight tensor that is four times the number of input channels of X;
[0051] B5: According to the cascade principle, perform cascade wavelet decomposition operation according to formula (7) and cascade convolution operation according to formula (8);
[0052]
[0053]
[0054] Where is the low-frequency component of the current layer, represents the three high-frequency components of the i-th level; represents the convolution output of the low-frequency component, represents the convolution output of the three high-frequency components of the i-th level;
[0055] B6: Accumulate the convolution outputs of different levels according to formula (9) to obtain the aggregated output Z after the i-th level (i) ;
[0056]
[0057] As an optimization, in step three, the process of using the hybrid attention module to enhance the recognition ability of the focusing ability of the key areas affected by speckle noise in the image is as follows:
[0058] C1: Use the CoordAtt module to capture global information through adaptive pooling operation and generate a spatial attention map, as shown in formula (10);
[0059]
[0060] Where F(i, j) represents the value of the feature map at position (i, j), Z1 and Z2 are the height and width of the feature map respectively, and σ is the activation function used to generate the final spatial attention map;
[0061] C2: Perform global average pooling operation on each channel of the input feature map, compress each channel of the feature map into a vector, so as to extract the global information of each channel, as shown in formula (11);
[0062]
[0063] where z c is the global description of channel c, and x i,j,c is the value of the c-th channel in the feature map at position (i, j), and H and W are the height and width of the feature map respectively;
[0064] C3: Generate the weight e of each channel through two fully connected layers. The first fully connected layer is used to reduce the number of channels by means of linear transformation, and the second fully connected layer is used to restore the number of channels to the original number by means of non-linear transformation, as shown in formula (12);
[0065] e = σ(W2·ReLU(W1·z)) (12);
[0066] where W1 and W2 are the weights of the two fully connected layers respectively, z is the vector after global pooling, e is the generated channel weight, and σ is the Sigmoid activation function;
[0067] C4: Multiply the generated channel weight e with the input feature map channel by channel according to formula (13) to complete the recalibration of the feature map and obtain the weighted output feature map x', so as to enhance the feature expression ability of each channel;
[0068] x' = x × e (13);
[0069] where x is the input feature map.
[0070] The present invention proposes a method for reconstructing a tracking-free three-dimensional ultrasound image based on deep learning. First, a large amount of two-dimensional ultrasound data is collected. Then, a deep learning model is constructed, which includes an ultrasound sequence encoding module, a Wavelet Transform Convolutions (WTConv) module, and a hybrid attention module (CoordAtt and SE modules). For the ultrasound sequence encoding module, a transformation matrix is predicted through a multi-task learning framework to model the spatial relationship between ultrasound image frames. By encoding the context relationship between image frames, the network can effectively capture the spatial transformation information in the sequence. For the wavelet transform convolution module, multi-scale features of the image are extracted through wavelet transform, effectively separating the high-frequency and low-frequency components of the image, reducing noise interference, and improving the discrimination ability of features. Performing convolution operations in the wavelet domain enables the network to better focus on features in different frequency bands, thereby effectively extracting multi-scale features. The low-frequency components help capture the macroscopic background and structural features of the image, such as the overall shape of the organ and large tissue regions; the high-frequency components focus on detailed information, such as tissue boundaries and textures. This multi-scale feature extraction ability enables the model to more comprehensively understand the image content when processing ultrasound images, not only paying attention to local details but also grasping the overall structure. At the same time, since wavelet transform preserves the spatial resolution to a certain extent, spatial operations (such as convolution) are more meaningful in the wavelet domain, enabling better utilization of the spatial information in the image, further enhancing the feature representation ability of the ultrasound image sequence, and laying a solid foundation for accurately reconstructing 3D ultrasound images. On this basis, through the combination of cascaded wavelet transform and small convolution kernels, efficient feature extraction is achieved, and the hybrid attention mechanism is integrated to focus on relevant features, significantly improving the reconstruction accuracy and efficiency of three-dimensional ultrasound images. For the hybrid attention module, an MBConv block network based on the EfficientNet architecture is adopted, and through innovative modifications to the depthwise separable convolution module and the attention mechanism, the performance of the network in the three-dimensional ultrasound image reconstruction task is improved. Thus, the coordinate attention (CoordAtt) and channel attention (SE) modules are combined, and through the dual attention mechanism of space and channel, the focusing ability and recognition ability of the network for key regions are enhanced, the convergence speed of the model is accelerated, the reconstruction accuracy is improved, and the overall reconstruction efficiency is enhanced. In addition, the present invention also designs a loss function that combines the root mean square distance (RMSD) loss and the correlation loss, ensuring that the model reduces the point-to-point error while maintaining the global structural consistency of the ultrasound image sequence. The three-dimensional ultrasound image reconstruction model obtained by these methods not only improves the three-dimensional reconstruction accuracy of ultrasound images but also optimizes the reconstruction efficiency, providing a more accurate and practical image analysis tool for clinical and scientific research.
[0071] Compared with the prior art, the present invention has the following advantages:
[0072] 1. Enhanced generalization ability: Through deep learning techniques, this method can effectively process ultrasound images from different devices and of different qualities, improving the application flexibility of the model on diverse data. This generalization ability enables the model to maintain stable performance in different clinical environments, especially in 3D ultrasound reconstruction, where it can adapt to different scanning trajectories and image features.
[0073] 2. Improved reconstruction accuracy: The 3D ultrasound reconstruction method proposed in this invention shows high accuracy in identifying key structures in ultrasound images. This accuracy is crucial for the accurate reconstruction of ultrasound images, especially in accurately estimating the probe trajectory and maintaining high reconstruction quality. This method can provide a more accurate spatial transformation matrix, thus significantly improving the reconstruction accuracy.
[0074] 3. Reduced dependence on high-quality devices: This method performs 3D ultrasound reconstruction in a sensorless manner, reducing the dependence on high-end devices and high-end positioning systems. This enables high-quality ultrasound image reconstruction and analysis in resource-constrained environments, reducing the need for additional devices or hardware, making this technology more economical and easy to deploy, especially in clinical scenarios where widespread use of ultrasound imaging is required.
[0075] In summary, the deep learning-based trackless free-form three-dimensional ultrasound image reconstruction method provided by this invention effectively solves the problem of determining the spatial position relationship between ultrasound image frames. At the same time, it effectively combines wavelet transform and hybrid attention mechanism, optimizing the multi-scale feature expression and learning efficiency of the image. This invention can significantly improve the efficiency and accuracy of three-dimensional ultrasound image reconstruction. It not only improves the adaptability and robustness of the model in processing ultrasound images under different conditions, but also enhances the accuracy and efficiency of reconstruction. In practical applications, even without additional hardware support, high-quality three-dimensional ultrasound image reconstruction can be achieved, providing more accurate diagnostic information for clinicians. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is a schematic diagram of image acquisition using the image acquisition system of this invention;
[0077] Figure 2 is a schematic diagram of ultrasound sequence encoding in this invention;
[0078] Figure 3 is a schematic diagram of wavelet transform convolution in this invention;
[0079] Figure 4 is the structure diagram of the hybrid attention module in this invention;
[0080] Figure 5 is the structure diagram of the three-dimensional ultrasound image reconstruction model in this invention. Detailed implementation manners
[0081] The present invention provides a three-dimensional ultrasound image reconstruction method based on spatial transformation. First, the ultrasound image sequence is effectively encoded, and the spatial transformation matrix between image frames is extracted by using context information; then, a wavelet transform convolution module is introduced for multi-scale feature extraction to enhance image details and structural information; next, a coordinate attention and channel attention module is combined to optimize the network's attention ability to key regions; finally, a loss function is designed to fuse the root mean square distance and correlation loss to ensure point-to-point error and global structural consistency. Thus, a three-dimensional ultrasound image reconstruction model is obtained, and the high-precision three-dimensional ultrasound image can be efficiently reconstructed by using the three-dimensional ultrasound image reconstruction model.
[0082] The following takes the three-dimensional ultrasound reconstruction of the arm part as an example and further illustrates the present invention in conjunction with the accompanying drawings.
[0083] The following will be combined with the attached Figures 1 to 5 Further description will be made on the specific implementation manners of the present invention.
[0084] The present invention provides a tracking-free free-form three-dimensional ultrasound image reconstruction method based on deep learning, including the following steps:
[0085] Step 1: Obtain an image data set;
[0086] The image acquisition system is used to scan the similar diagnosis and treatment areas of several patients with a non-linear scanning trajectory to obtain two-dimensional ultrasound image sequence data. The two-dimensional ultrasound image sequence data is stored to form an image data set, and is divided into a training set, a test set and a validation set according to a set ratio;
[0087] Step 2: Construct a three-dimensional ultrasound image reconstruction model;
[0088] S21: Construct an ultrasound sequence encoding module;
[0089] Using the frame indices (i, j) and the sequence length m as hyperparameters, the recurrent neural network f sequentially obtains image frames from the sequence based on the context relationship of the ultrasound sequence, and models the spatial relationship between ultrasound image frames by predicting the spatial transformation matrix between frames, obtaining an ultrasound sequence encoding module for subsequent frame sequence encoding, as Figure 2 shown;
[0090] The use of hyperparameters enables the encoding method to be adjusted and optimized according to different situations to adapt to different ultrasound sequence data and prediction requirements;
[0091] S22: Construct a wavelet transform convolution module;
[0092] A wavelet transform convolution module is constructed by combining the Haar wavelet transform with small kernel depth convolution; in the wavelet transform convolution module, the Haar wavelet transform is used to filter and downsample the input low-frequency and high-frequency content, then, small kernel depth convolution is performed on different frequency maps, and finally, the inverse wavelet transform is used to construct the output;
[0093] The wavelet transform is a powerful tool for signal processing and analysis. The present invention adopts the efficient and simple Haar wavelet transform. For a given image X, one layer of Haar wavelet transform in one spatial dimension (width or height) consists of depth convolutions with kernels of and , followed by a standard downsampling operation with a scaling factor of 2. When performing wavelet transform on a two-dimensional image, these operations need to be combined in two dimensions, that is, depth convolutions with a stride of 2 are performed using the following four groups of filters; the four groups of filters are respectively: where f LL is a low-pass filter, while f LH , f HL , f HH is a group of high-pass filters. For each input channel, the output of the convolution [X LL , X Lζ , X HL , X LH = Conv([f LL , f LH , f HL , f HH , X) has four channels, X LL is the low-frequency component of the image X, and X LH , X HL and X HH are its horizontal, vertical and diagonal high-frequency components respectively. Since these kernels form an orthonormal basis, the inverse wavelet transform (IWT) can be realized by transposed convolution, that is, X = Conv T ([f LL , f LH , f HL , f HH , [X LL , X LH , X HL , X HH ). Cascade wavelet decomposition can be realized by recursively decomposing the low-frequency component. Each level of decomposition increases the frequency resolution and reduces the low-frequency spatial resolution.
[0094] To alleviate the problem of the quadratic growth of the number of parameters when increasing the kernel size in the traditional convolutional layer, the present invention proposes a method of performing convolution in the wavelet domain. Wavelet transform convolution has significant advantages. It can accelerate model convergence. Thanks to the multi-scale characteristics of wavelet transform, the network can learn key image features in a shorter time, thus focusing on important information faster and reducing training time and resource consumption. In addition, by separating low-frequency and high-frequency information, redundant calculations can be reduced, the number of parameters can be decreased, and the training efficiency can be improved, enabling the model to more efficiently utilize computing resources when processing ultrasound images and improving the overall performance.
[0095] The present invention introduces a wavelet transform convolution (WTConv) module to replace the traditional convolutional layer, which can effectively improve the feature extraction ability of the network for ultrasound image sequences. This module replaces the traditional convolutional layer and can capture the low-frequency components and high-frequency components of ultrasound images at different scales, thereby enhancing the model's recognition ability for image details and structural features; the low-frequency components are mainly used to reflect the macroscopic background and structural features of the image, and the high-frequency components are mainly used to focus on the fine features of the image, such as tissue boundaries and textures, which helps to comprehensively understand the image content;
[0096] S23: Construct a hybrid attention module;
[0097] Based on the MBConv block network of the EfficientNet architecture, modifications are made. In the first MBConv block of each stage, the original SE module is replaced with a CoordAtt (coordinate attention) module. At the same time, the original SE (channel attention) module in the second MBConv block is retained. Furthermore, a hybrid attention module is constructed by combining the CoordAtt module and the SE module, as Figure 3 shown; in the hybrid attention module, the CoordAtt module is used to capture global information through adaptive pooling operations and generate a spatial attention map. The SE module, that is, the Squeeze-and-Excitation module, is a channel attention mechanism, and its core idea is to complete the learning of channel weights through the following operations: Squeeze operation; Excitation operation; Excitation mapping; The SE module explicitly models the relationships between channels in this way, allowing the network to dynamically emphasize useful features and suppress irrelevant features, thereby improving the network's representation ability. The SE (Squeeze-and-Excitation) module is used to enhance the expressive ability of the feature map by learning the weights between channels;
[0098] This design of the hybrid attention module allows the model to utilize the spatial focusing ability of CoordAtt in the initial block and continue to benefit from the channel-level recalibration ability of SE in subsequent blocks. The combination of the hybrid attention module utilizes spatial and channel-level information, enhancing the focusing ability and recognition ability for key image regions, enabling the model to more effectively identify and enhance key regions affected by speckle noise in ultrasound images, thereby improving the accuracy of 3D ultrasound image reconstruction, the convergence speed of the model, enhancing the reconstruction accuracy, and improving the overall reconstruction efficiency.
[0099] S24: Construct the loss function;
[0100] Define the loss function to optimize the accuracy and structural consistency of the model prediction. Specifically, construct the loss function by combining the weighted sum of the root mean square distance (RMSD) loss function and the Pearson correlation loss function;
[0101] S25: Obtain the 3D ultrasound image reconstruction model through training;
[0102] Construct a deep learning model by combining the ultrasound sequence encoding module, the wavelet transform convolution module, and the hybrid attention module, and use the training set to train the deep learning module. During the training process, continuously update the model parameters through the backpropagation algorithm to reduce the value of the loss function, and verify the results through the validation set to select the optimal parameters, thereby obtaining a trained deep learning model; use the test set to test the trained deep learning model to obtain the 3D ultrasound image reconstruction model;
[0103] Step three: Perform the reconstruction of the 3D ultrasound image;
[0104] Use the image acquisition system to scan the patient's diagnosis and treatment area with a non-linear scanning trajectory to obtain two-dimensional ultrasound image sequence data;
[0105] Use the two-dimensional ultrasound image sequence data as input data to input into the 3D ultrasound image reconstruction model for image reconstruction. During the image reconstruction process, use the ultrasound sequence encoding module to encode the context relationship between image frames to capture the context relationship in the ultrasound image sequence, use the wavelet transform convolution module to extract features from the image sequence data. In this process, extract multi-scale features by cascading wavelet transform (WT) and small kernel depth convolution to focus on features in different frequency bands while minimizing the number of trainable parameters, use the hybrid attention module to enhance the focusing ability and recognition ability for key regions affected by speckle noise in the image to improve the reconstruction accuracy, and finally output high-quality 3D ultrasound images. In this way, high-quality 3D ultrasound image reconstruction results can be obtained in an environment without additional hardware support to provide accurate image information for clinical diagnosis.
[0106] As a preference, in step one, the process of collecting two-dimensional ultrasonic image sequence data is as follows:
[0107] S11: An image acquisition system is composed of a computer, an ultrasonic scanning device, an optical tracker, and an ultrasonic probe. Among them, the computer is respectively connected to the ultrasonic scanning device and the optical tracker, and the ultrasonic scanning device is connected to the ultrasonic probe;
[0108] As a preference, the model number of the optical tracker is NDIPolaris Vicra, produced by NDI Company of Canada; the model number of the ultrasonic scanning device is an Ultrasonix machine with a curved probe (4DC7-3 / 40). In this way, two-dimensional ultrasonic image sequence data can be accurately collected;
[0109] S12: Use the ultrasonic probe to scan the target diagnosis and treatment area of the patient. During the scanning process, the ultrasonic probe randomly adopts a linear scanning trajectory, a C-shaped scanning trajectory, or an S-shaped scanning trajectory for scanning operations. The ultrasonic scanning device sends the scanned image to the computer. At the same time, during the scanning process, use the optical tracker to accurately position and track the ultrasonic probe, and obtain the spatial position information of the ultrasonic probe, and then send the spatial position information to the computer; during the scanning process, the ultrasonic frames are recorded at a speed of 20fps. In this way, the continuity of the image sequence and the capture of dynamic changes are ensured, providing sufficient time resolution for subsequent image analysis and processing. At the same time, the size of each frame of image is 480×640 pixels. In this way, sufficient detail information is retained. At the same time, no speckle reduction processing is performed. In this way, the original noise characteristics in the image are retained;
[0110] S13: The computer integrates the image data and the spatial position information according to the timing information of the spatial position information and the timing information of the image, and sorts the integrated data in chronological order to form two-dimensional ultrasonic image sequence data; a number of two-dimensional ultrasonic image sequence data are obtained through 1200 scanning processes, and the number of two-dimensional ultrasonic image sequence data is stored in time to obtain an image data set. Specifically, multiple patients can be used as volunteers, and each volunteer can contribute 24 scans. Thus, the diversity and representativeness of the data are ensured, enabling the model to learn more extensive features and rules from the data of different volunteers. The obtained image data can be stored in a.h5 file. This file format can effectively organize and store a large amount of image data and related information, facilitating subsequent data loading, processing, and analysis, and providing a convenient data access path for model training and verification.
[0111] As a preference, in S24 of step two, the process of constructing the loss function is as follows:
[0112] S24-1: Measure the correlation between the predicted features and the true features according to the Pearson correlation coefficient, and calculate the Pearson correlation loss function corr(preds, labels) through formula (1), that is, calculate the correlation loss by calculating the Pearson correlation coefficient between the predicted values and the true label feature points, ensuring that the model can not only minimize the point error,
[0113] but also maintain the global structural consistency of the ultrasound image sequence;
[0114]
[0115] wherein, Cov(pred i , label i ) is the covariance between the predicted value pred i and the true label label i , and σ(preds′) and σ(labels′) respectively represent the standard deviations of the predicted value and the true label;
[0116] S24-2: Combine the root mean square distance (RMSD) loss function and the Pearson correlation loss function corr(preds, labels), and obtain the loss function L using formula (2) total ; thus, incorporate the correlation loss into the total loss function, ensuring that the model maintains the global structural consistency while reducing the point error;
[0117] L total = L RMSD + λ(1 - corr(preds, labels)) (2);
[0118] wherein, L RMSD is the root mean square distance loss function, λ is the weight parameter for balancing the root mean square distance loss function and the correlation loss function, initially defaulting to 1. This design of the loss function helps the model to not only focus on reducing the local error during the training process, but also pay attention to maintaining the overall structure, thereby improving the accuracy and robustness of 3D ultrasound reconstruction.
[0119] As an optimization, in step three, the process of encoding the context relationship between image frames using the ultrasound sequence encoding module is as follows:
[0120] A1: Sample the given two-dimensional ultrasound image sequence M times to obtain an ultrasound image sequence with a sequence length of m,
[0121] wherein, m = 1, 2, …, M;
[0122] A2: For any pair of frame indices (i, j) in the two-dimensional ultrasound image sequence M, the available neighboring frame information includes {I m} ∈ ([1, i - 1] ∪ [j + 1, m]) and their relative positional relationships, where i and j are not adjacent frames, i.e., i ≠ j - 1. When m ≠ M, the spatial transformation matrix T between the i-th frame and the j-th frame is predicted according to formula (3) i-j , the spatial transformation matrix T i-j is recursively predicted at the end of each sequence pair. Using the context information from the ultrasound image sequence can provide important context information for spatial transformation prediction; the spatial transformation matrix T i-j contains the key information about the relative position change between two frames, such as relative displacement and rotation, etc., which is the core target of neural network prediction.
[0123] T i-j = f(S m , I (m-1) ; θ) (3);
[0124] In the formula, m represents the time step, which is used to clarify the specific moment during the ultrasound sequence processing. As m increases, the neural network orderly processes each image frame in the sequence and conducts corresponding prediction work. S m represents the input information at time step m in the ultrasound sequence, including the ultrasound image data at this moment or the feature information closely related to it, which is an indispensable key input for the neural network to make predictions. I (m-1) represents the internal hidden state at time step m - 1, which stores the key information extracted by the neural network when processing data at the previous time step and provides important context association for the prediction at the current time step. θ are the parameters of the neural network f, and these parameters are continuously adjusted and optimized during the model training process. Through algorithms such as backpropagation, the neural network can better fit the data, thereby achieving more accurate predictions;
[0125] Through the above input sequence encoding, when the neural network processes the image sequence, it can utilize the current, past, and future frames to provide conditional context information for spatial transformation prediction, greatly improving the encoding efficiency.
[0126] As an optimization, in step three, the process of using the wavelet transform convolution module to extract multi-scale features from the image sequence data is as follows:
[0127] B1: For the ultrasound image frame x[n], it is processed using the low-pass filter g[n] and the high-pass filter h[n] to achieve frequency domain separation; specifically, the low-frequency component x L [n] is obtained according to formula (4), and the high-frequency component x H[n], the ultrasonic image frame x[n] is decomposed twice using the Haar wavelet transform in the above manner to obtain the low-frequency component x L [n] and three high-frequency components x H [n]; this separation process helps reduce noise interference and emphasizes the key features in the image, providing more accurate feature information for 3D ultrasonic image reconstruction.
[0128] x L [n] = ∑ k x[2n - k]·g[k] (4);
[0129] x H [n] = ∑ k x[2n - k]·h[k] (5);
[0130] In the formula, n is the position of the pixel, g[k] is the coefficient of the low-pass filter g[n], and h[k] is the coefficient of the high-pass filter h[n];
[0131] B2: Downsample the low-frequency component x L [n] and the three high-frequency components x L [n] to reduce the image resolution. This process not only improves the efficiency of feature extraction but also reduces the computational complexity of the model;
[0132] B3: Perform small kernel depth convolution on the low-frequency component x L [n] and the high-frequency component x H [n] respectively to extract features;
[0133] B4: According to formula (6), perform inverse transformation on the low-frequency component x L [n] and the high-frequency component x H [n] and obtain the output Y;
[0134] Y = IWT(Conv(W, WT(X))) (6);
[0135] In the formula, X is the input tensor, and W is the k×k depth kernel weight tensor that is four times the number of input channels of X;
[0136] The above operations not only separate the convolution between frequency components but also allow smaller kernels to operate on a larger area of the original input, increasing the receptive field relative to the input.
[0137] B5: Through the cascade principle, perform cascade wavelet decomposition operations according to formula (7) and cascade convolution operations according to formula (8);
[0138]
[0139]
[0140] In the formula, is the low-frequency component of the current layer, represents the three high-frequency components of the i-th level; represents the convolutional output of the low-frequency component, represents the convolutional output of the three high-frequency components of the i-th level;
[0141] B6: To combine the outputs of different frequencies, taking advantage of the property that WT and IWT are linear operations, the convolutional outputs of different levels are accumulated according to formula (9) to obtain the aggregated output Z after the i-th level (i) ; Different from other methods, here, instead of performing separate normalization on each a channel scaling is used to weigh the contributions of each frequency component.
[0142]
[0143] As an optimization, in step three, the process of using the hybrid attention module to enhance the recognition ability of the focusing ability for key regions affected by speckle noise in the image is as follows:
[0144] C1: Use the CoordAtt module to capture global information through adaptive pooling operations and generate a spatial attention map, as shown in formula (10);
[0145]
[0146] In the formula, CoordAtt(Z,F) is to generate the final spatial attention map, F(i,j) represents the value of the feature map at position (i,j), Z1 and Z2 are the height and width of the feature map respectively, and σ is the activation function used to generate the final spatial attention map;
[0147] C2: Squeeze operation; perform global average pooling (Global AveragePooling, GAP) on each channel of the input feature map to compress each channel of the feature map into a vector, thereby extracting the global information of each channel, as shown in formula (11);
[0148]
[0149] In the formula, z c is the global description of channel c, x i,j,c is the value of the c-th channel in the feature map at position (i,j), and H and W are the height and width of the feature map respectively;
[0150] C3: Excitation operation; generate the weight e for each channel through two fully connected layers. The first fully connected layer is used to reduce the number of channels through linear transformation, and the second fully connected layer is used to restore the number of channels to the original number through non-linear transformation, and generate the weight for each channel through the Sigmoid activation function, as shown in formula (12);
[0151] e = σ(W2·ReLU(W1·z)) (12);
[0152] In the formula, W1 and W2 are the weights of the two fully connected layers respectively, z is the vector after global pooling, e is the generated channel weight, and σ is the Sigmoid activation function;
[0153] C4: Multiply the generated channel weight e with the input feature map channel by channel according to formula (13) to complete the recalibration of the feature map and obtain the weighted output feature map x′ to enhance the feature expression ability of each channel; this process can strengthen the response of important channels while suppressing unimportant channels.
[0154] x′ = x × e (13);
[0155] In the formula, x is the input feature map.
[0156] The present invention proposes a method for reconstructing a traceless free-form three-dimensional ultrasound image based on deep learning. First, a large amount of two-dimensional ultrasound data is collected. Then, a deep learning model is constructed, which includes an ultrasound sequence encoding module, a Wavelet Transform Convolutions (WTConv) module, and a hybrid attention module (CoordAtt and SE modules). For the ultrasound sequence encoding module, a transformation matrix is predicted through a multi-task learning framework to model the spatial relationship between ultrasound image frames. By encoding the context relationship between image frames, the network can effectively capture the spatial transformation information in the sequence. For the wavelet transform convolution module, multi-scale features of the image are extracted through wavelet transform, effectively separating the high-frequency and low-frequency components of the image, reducing noise interference, and improving the discrimination ability of features. Performing convolution operations in the wavelet domain enables the network to better focus on features in different frequency bands, thereby effectively extracting multi-scale features. The low-frequency components help capture the macroscopic background and structural features of the image, such as the overall shape of the organ and large tissue regions; the high-frequency components focus on detailed information, such as tissue boundaries and textures. This multi-scale feature extraction ability enables the model to more comprehensively understand the image content when processing ultrasound images, not only paying attention to local details but also grasping the overall structure. At the same time, since wavelet transform preserves the spatial resolution to a certain extent, spatial operations (such as convolution) are more meaningful in the wavelet domain, and the spatial information in the image can be better utilized, further enhancing the feature representation ability of the ultrasound image sequence, laying a solid foundation for accurately reconstructing 3D ultrasound images. On this basis, through the combination of cascaded wavelet transform and small convolutional kernels, efficient feature extraction is achieved, and a hybrid attention mechanism is integrated to focus on relevant features, significantly improving the reconstruction accuracy and efficiency of three-dimensional ultrasound images. For the hybrid attention module, an MBConv block network based on the EfficientNet architecture is adopted, and through innovative modifications to the depthwise separable convolution module and the attention mechanism, the performance of the network in the three-dimensional ultrasound image reconstruction task is improved. Thus, the coordinate attention (CoordAtt) and channel attention (SE) modules are combined. Through the dual attention mechanism of space and channel, the focusing ability and recognition ability of the network for key regions are enhanced, the convergence speed of the model is accelerated, the reconstruction accuracy is improved, and the overall reconstruction efficiency is enhanced. In addition, the present invention also designs a loss function that combines the root mean square distance (RMSD) loss and the correlation loss, ensuring that the model reduces the point-to-point error while maintaining the global structural consistency of the ultrasound image sequence. The three-dimensional ultrasound image reconstruction model obtained by these methods not only improves the three-dimensional reconstruction accuracy of ultrasound images but also optimizes the reconstruction efficiency, providing a more accurate and practical image analysis tool for clinical and scientific research.
[0157] The method for free-form three-dimensional ultrasound image reconstruction based on deep learning provided by the present invention effectively solves the problem of determining the spatial position relationship between ultrasound image frames. At the same time, it effectively combines wavelet transform and hybrid attention mechanism, optimizing the multi-scale feature expression and learning efficiency of the image. The present invention can significantly improve the efficiency and accuracy of three-dimensional ultrasound image reconstruction. It not only improves the adaptability and robustness of the model in processing ultrasound images under different conditions, but also enhances the accuracy and efficiency of reconstruction. In practical applications, even without additional hardware support, high-quality three-dimensional ultrasound image reconstruction can be achieved, providing more accurate diagnostic information for clinicians.
Claims
1. A method for reconstructing a free-form three-dimensional ultrasound image without tracking based on deep learning, characterized in that, Including the following steps: Step 1: Obtain an image dataset; Use an image acquisition system to scan similar diagnostic regions of several patients with a non-linear scanning trajectory to obtain two-dimensional ultrasound image sequence data. Store the two-dimensional ultrasound image sequence data to form an image dataset, and divide it into a training set, a test set, and a validation set according to a set ratio; Step 2: Construct a three-dimensional ultrasound image reconstruction model; S21: Construct an ultrasound sequence encoding module; Using the frame index (i, j) and the sequence length m as hyperparameters, make the recurrent neural network f sequentially obtain image frames from the sequence based on the context relationship of the ultrasound sequence, and model the spatial relationship between ultrasound image frames by predicting the spatial transformation matrix between frames, to obtain an ultrasound sequence encoding module for subsequent frame sequence encoding; S22: Construct a wavelet transform convolution module; Construct a wavelet transform convolution module by combining the Haar wavelet transform with small kernel depth convolution; in the wavelet transform convolution module, use the Haar wavelet transform to filter and downsample the input low-frequency and high-frequency content, then perform small kernel depth convolution on different frequency maps, and finally, use the inverse wavelet transform to construct the output; S23: Construct a hybrid attention module; Modify the MBConv block network based on the EfficientNet architecture. In the first MBConv block of each stage, replace the original SE module with a CoordAtt module, and at the same time, retain the original SE module in the second MBConv block, and then construct a hybrid attention module by combining the CoordAtt module and the SE module; in the hybrid attention module, the CoordAtt module is used to capture global information through adaptive pooling operations to generate a spatial attention map, and the SE module is used to enhance the expression ability of the feature map by learning the weights between channels. By combining spatial and channel-level information, enhance the focusing ability and recognition ability for key image regions; S24: Construct a loss function; Construct a loss function by combining the weighted sum of the root mean square distance loss function and the Pearson correlation loss function; S25: Obtain a three-dimensional ultrasound image reconstruction model through training; Combine the ultrasound sequence encoding module, the wavelet transform convolution module, and the hybrid attention module to construct a deep learning model, and use the training set to train the deep learning module. During the training process, continuously update the model parameters through the backpropagation algorithm to reduce the value of the loss function, and verify the results through the validation set to select the optimal parameters, and then obtain a trained deep learning model; use the test set to test the trained deep learning model to obtain a three-dimensional ultrasound image reconstruction model; Step 3: Reconstruct the three-dimensional ultrasound image; Use an image acquisition system to scan the diagnostic region of a patient with a non-linear scanning trajectory to obtain two-dimensional ultrasound image sequence data; The two-dimensional ultrasound image sequence data is input as input data into a three-dimensional ultrasound image reconstruction model for image reconstruction. During the image reconstruction process, an ultrasound sequence encoding module is used to encode the context relationship between image frames to capture the context relationship in the ultrasound image sequence, and a wavelet transform convolution module is used to extract features from the image sequence data. In this process, multi-scale feature extraction is performed by cascading wavelet transform and small kernel depth convolution to focus on features in different frequency bands. A hybrid attention module is used to enhance the recognition ability of the focusing ability of key regions affected by speckle noise in the image to improve the reconstruction accuracy. Finally, a high-quality three-dimensional ultrasound image is output to provide accurate image information for clinical diagnosis.
2. The method for reconstructing a traceless free-form three-dimensional ultrasound image based on deep learning according to claim 1, wherein In step one, the process of collecting two-dimensional ultrasound image sequence data is as follows: S11: An image acquisition system is composed of a computer, an ultrasound scanning device, an optical tracker, and an ultrasound probe. Among them, the computer is respectively connected to the ultrasound scanning device and the optical tracker, and the ultrasound scanning device is connected to the ultrasound probe; S12: The ultrasound probe is used to scan the target diagnosis and treatment area of the patient. During the scanning process, the ultrasound probe randomly uses a linear scanning trajectory, a C-shaped scanning trajectory, or an S-shaped scanning trajectory for scanning operations. The ultrasound scanning device sends the scanned image to the computer. At the same time, during the scanning process, the optical tracker is used to accurately position and track the ultrasound probe, and the spatial position information of the ultrasound probe is obtained, and then the spatial position information is sent to the computer; during the scanning process, the ultrasound frames are recorded at a speed of 20 fps. At the same time, the size of each frame of image is 480×640 pixels. At the same time, no speckle reduction processing is performed; S13: The computer integrates the image data and the spatial position information according to the timing information of the spatial position information and the timing information of the image, and sorts the integrated data in chronological order to form two-dimensional ultrasound image sequence data; a number of two-dimensional ultrasound image sequence data are obtained through 1200 scanning processes, and the number of two-dimensional ultrasound image sequence data is stored in time to obtain an image dataset.
3. A method for reconstructing a free-form three-dimensional ultrasound image without tracking based on deep learning according to claim 1, characterized in that In S24 of step two, the process of constructing the loss function is as follows: S24-1: The correlation between the predicted features and the true features is measured based on the Pearson correlation coefficient, and the Pearson correlation loss function corr(preds, labels) is calculated through formula (1); where Cov(pred i , label i ) is the covariance between the predicted value pred i and the true label label i , and σ(preds′) and σ(labels′) represent the standard deviations of the predicted value and the true label, respectively; S24-2: Combine the root mean square distance loss function and the Pearson correlation loss function corr(preds, labels), and obtain the loss function L using formula (2). total ; L total = L RMSD + λ(1 - corr(preds, labels))(2); Where L RMSD is the root distance loss function, and λ is the weight parameter that balances the root mean square distance loss function and the correlation loss function.
4. A method for reconstructing a free-form three-dimensional ultrasound image without tracking based on deep learning according to claim 1, wherein In step three, the process of encoding the context relationship between image frames by using the ultrasound sequence encoding module is as follows: A1: For a given two-dimensional ultrasound image sequence, M samplings are performed to obtain an ultrasound image sequence with a sequence length of m, where m = 1, 2,..., M; A2: For any pair of frame indices (i, j) in the two-dimensional ultrasound image sequence M, where i and j are not adjacent frames, i.e., i ≠ j - 1, when m ≠ M, predict the spatial transformation matrix T between the i-th frame and the j-th frame according to formula (3). i-j , the spatial transformation matrix T i-j is recursively predicted at the end of each sequence pair, using the context information from the ultrasound image sequence; T i-j = f(S m , I (m-1) ; θ)(3); where m represents the time step, S m represents the input information at time step m in the ultrasound sequence, I (m-1) represents the internal hidden state at time step m - 1, and θ are the parameters of the neural network f.
5. A method for reconstructing a free-form three-dimensional ultrasound image without tracking based on deep learning according to claim 1, characterized in that, In step three, the process of extracting multi-scale features from the image sequence data by using the wavelet transform convolution module is as follows: B1: For the ultrasonic image frame x[n], process it using the low-pass filter g[n] and the high-pass filter h[n], and obtain the low-frequency component x L [n] according to formula (4), and obtain the high-frequency component x H [n] according to formula (5). By the above method, decompose the ultrasonic image frame x[n] twice using the Haar wavelet transform to obtain the low-frequency component x L [n] and three high-frequency components x H [n]; x L [n] = ∑ k x[2n - k]·g[k] (4); x H [n] = ∑ k x[2n - k]·h[k] (5); In the formula, n is the position of the pixel, g[k] is the coefficient of the low-pass filter g[n], and h[k] is the coefficient of the high-pass filter h[n]; B2: Downsample the low-frequency component x L [n] and three high-frequency components x H [n] to reduce the image resolution; B3: On the low-frequency component x L [n] and the high-frequency component x H [n], perform small-core depth convolution respectively to extract features; B 4: Perform inverse transformation on the low-frequency component x L [n] and the high-frequency component x H [n], and obtain the output Y; Y = IWT(Conv(W, WT(X))) (6); In the formula, X is the input tensor, and W is a k×k depth kernel weight tensor that is four times the number of input channels of X; B5: According to the cascade principle, perform cascade wavelet decomposition operations according to formula (7) and cascade convolution operations according to formula (8). Wherein, is the low-frequency component of the current layer, represents the three high-frequency components of the i-th level; represents the convolution output of the low-frequency component, represents the convolution output of the three high-frequency components of the i-th level; B6: Accumulate the convolution outputs of different levels according to formula (9) to obtain the aggregated output Z after the i-th level (i) ; 6. A method for reconstructing a trackless free-form three-dimensional ultrasound image based on deep learning according to claim 1, characterized in that In step three, the process of using the hybrid attention module to enhance the recognition ability of the focusing ability of the key areas affected by speckle noise in the image is as follows: C1: Use the CoordAtt module to capture global information through adaptive pooling operations and generate a spatial attention map, as shown in formula (10); In the formula, F(i,j) represents the value of the feature map at position (i,j), Z1 and Z2 are the height and width of the feature map respectively, and σ is the activation function used to generate the final spatial attention map; C2: Perform global average pooling operations on each channel of the input feature map, compress each channel of the feature map into a vector, and thus extract the global information of each channel, as shown in formula (11); where z c is the global description of channel c, and x i,j,c is the value of the c-th channel in the feature map at position (i, j), where H and W are the height and width of the feature map respectively; C3: Generate the weight e of each channel through two fully connected layers. The first fully connected layer is used to reduce the number of channels through linear transformation, and the second fully connected layer is used to restore the number of channels to the original number through nonlinear transformation, as shown in formula (12); e = σ(W2·ReLU(W1·z)) (12); In the formula, W1 and W2 are the weights of the two fully connected layers respectively, z is the vector after global pooling, e is the generated channel weight, and σ is the Sigmoid activation function; C4: Multiply the generated channel weight e by the input feature map channel by channel according to formula (13) to complete the recalibration of the feature map and obtain the weighted output feature map x′ to enhance the feature expression ability of each channel; x′ = x×e(13); In the formula, x is the input feature map.
Citation Information
Cited By
Puncture method and system based on ultrasonic guidance
CN120616729A
Welding structure defect labeling method and defect recognition model training method
CN121391853A
Three-dimensional ultrasonic equipment integrating photoelectric sensor and motion sensor and reconstruction method
CN121533759A