Feature fusion method, image segmentation method, electronic equipment and storage medium
By aggregating deep feature maps and generating filters for processing, the feature inconsistency problem is solved, and the accuracy and model performance of nodule recognition in medical images are improved.
Patent Information
- Application Number
- CN202510361476.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing feature fusion method has feature inconsistency problems when dealing with nodule recognition tasks in medical images, resulting in deterioration in model performance and inaccurate boundary recognition.
By aggregating the deep feature map and the shallow feature map, low-pass filters and high-pass filters are generated, and the deep feature maps are smoothed and shallow feature maps are sharpened, and finally fused to generate smoother consistency features.
It reduces the inconsistency of overall features, improves the accuracy of boundary recognition, and enhances the overall performance of the model.
Smart Images

Figure CN120147807A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image feature processing, and in particular, to a feature fusion method, an image segmentation method, an electronic device, and a storage medium. Background Art
[0002] With the development of medical image analysis technology, in existing recognition methods, segmentation algorithms represented by Unet and detection algorithms represented by Yolo are widely used. When these algorithms process nodule recognition tasks in medical images, it is necessary to fuse the extracted features to more comprehensively evaluate the nodule area.
[0003] In related technologies, feature fusion usually upsamples high-level features by using nearest neighbor interpolation or bilinear interpolation, and then simply adds or splices them with low-level features. However, this simple fusion method has obvious limitations. Since there may be feature inconsistencies between low-level features and high-level features, that is, there may be different descriptions for different regions of the same object, this will cause the upsampled feature map to further amplify these inconsistencies, thereby affecting the overall performance of the model and resulting in inaccurate nodule recognition in medical images. Summary of the Invention
[0004] In view of this, the purpose of the embodiments of the present invention is to provide a feature fusion method, an image segmentation method, an electronic device, and a storage medium to at least partially improve the above problems.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, an embodiment of the present invention provides a feature fusion method, and the method includes: Obtain an extracted deep feature map and a shallow feature map; the resolution of the deep feature map is less than the resolution of the shallow feature map; Aggregate the deep feature map and the shallow feature map to obtain an aggregated feature map; Generate a low-pass filter according to the aggregated feature map, and use the low-pass filter to smooth the deep feature map to obtain a consistency feature map; Generate a high-pass filter according to the aggregated feature map, and use the high-pass filter to sharpen the shallow feature map to obtain a high-frequency feature map; Fuse the consistency feature map and the high-frequency feature map to obtain a fusion feature map.
[0006] Optionally, the aggregating the deep feature map and the shallow feature map to obtain an aggregated feature map includes: Upsample the deep feature map to the resolution of the shallow feature map to obtain an upsampled high-resolution deep feature map; Concatenate the high-resolution deep feature map and the shallow feature map to obtain an aggregated feature map.
[0007] Optionally, generating a low-pass filter according to the aggregated feature map and using the low-pass filter to smooth the deep feature map to obtain a consistent feature map includes: Extract the neighborhood feature relationship of each pixel position of the aggregated feature map according to a preset filter; the size of the neighborhood is the size of the preset filter; Normalize each of the neighborhood feature relationships to obtain a low-pass filter with the sum of corresponding weights equal to 1; Apply each of the low-pass filters to the high-resolution deep feature map through matrix multiplication to obtain a consistent feature map.
[0008] Optionally, the calculation formula for normalizing each of the neighborhood feature relationships is:
[0009] where, represents the feature relationship at the (i, j)th position belonging to the neighborhood at the (m, n) coordinate, represents the low-pass filter at the (m, n) coordinate; The formula for calculating the consistent feature map is:
[0010] where, represents the consistent feature map, represents the feature at the same position of the corresponding low-pass filter in the high-resolution deep feature map.
[0011] Optionally, generating a high-pass filter according to the aggregated feature map and using the high-pass filter to sharpen the shallow feature map to obtain a high-frequency feature map includes: Extract the neighborhood feature relationship of each pixel position of the aggregated feature map according to a preset filter; the size of the neighborhood is the size of the preset filter; Normalize each of the neighborhood feature relationships to obtain a low-pass filter with the sum of corresponding weights equal to 1; Subtract the low-pass filter from the unit kernel to obtain a high-pass filter; Apply each of the high-pass filters to the shallow feature map through matrix multiplication to obtain a high-frequency feature map.
[0012] Optionally, the method further includes: Adding the high-frequency feature map and the shallow feature map to obtain a high-frequency feature map that retains gradient information.
[0013] In a second aspect, an embodiment of the present invention provides a feature fusion method, the method including: Obtaining a plurality of extracted feature maps; wherein, the plurality of feature maps include a first feature map to an Nth feature map sorted in ascending order of resolution, and N is an integer greater than or equal to 2; Fusing the first feature map and the feature maps after the first feature map through the feature fusion method described in any one of claims 1-6 to obtain a fused feature map; Fusing the fused feature map and the feature maps after the second feature map through the feature fusion method described in any one of claims 1-6 to obtain a new fused feature map, until the fused feature map and the Nth feature map are fused through the feature fusion method described in any one of claims 1-6 to obtain a final feature fusion map.
[0014] In a third aspect, an embodiment of the present invention provides an image segmentation method, the method including: Performing feature extraction on the image to be segmented through a segmentation network to obtain a plurality of feature maps; Fusing each of the feature maps through the feature fusion method described in claim 7 to obtain a fused feature map; Inputting the fused feature map into a classifier of the segmentation network to obtain a segmentation result of the image to be segmented.
[0015] In a fourth aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, where the processor implements the method described in any one of the above when executing the program.
[0016] In a fifth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and the computer program implements the method described in any one of the above when executed by a processor.
[0017] A feature fusion method, an image segmentation method, an electronic device, and a storage medium provided by an embodiment of the present invention aggregate deep feature maps and shallow feature maps to obtain an aggregated feature map, then generate a low-pass filter and a high-pass filter according to the aggregated feature map, and then use the low-pass filter and the high-pass filter to perform smoothing and sharpening processing on the deep feature map and the shallow feature map respectively and then fuse them, so as to fuse and generate smoother consistent features, greatly reducing the inconsistency of the overall features and making the boundary recognition more accurate.
[0018] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically presents preferred embodiments and, in conjunction with the accompanying drawings, provides a detailed description as follows. Description of the Drawings
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0020] Figure 1 A schematic structural block diagram of an electronic device provided by an embodiment of the present invention; Figure 2 A schematic flowchart of a feature fusion method provided by an embodiment of the present invention; Figure 3 Another schematic flowchart of a feature fusion method provided by an embodiment of the present invention; Figure 4 A schematic flowchart of step S230 provided by an embodiment of the present invention; Figure 5 A schematic diagram of the neighborhood size at different pixel positions in an aggregated feature map provided by an embodiment of the present invention; Figure 6 A schematic flowchart of step S240 provided by an embodiment of the present invention; Figure 7 A schematic diagram of feature fusion provided by an embodiment of the present invention; Figure 8 A schematic flowchart of another feature fusion method provided by an embodiment of the present invention; Figure 9 A schematic diagram of multi-layer feature fusion provided by an embodiment of the present invention; Figure 10 A schematic flowchart of an image segmentation method provided by an embodiment of the present invention.
[0021] Icons: 100 - Electronic device; 101 - Memory; 102 - Communication interface; 103 - Processor; 104 - Communication bus. Detailed Embodiments
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0023] Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0024] It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0025] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0026] Currently, according to lymph node ultrasound imaging and combined with the professional judgment of physicians, it is the mainstream method for detecting lymph nodes at present. However, due to the diversity of ultrasound imaging and the complexity of lymph node structures, this method highly relies on the subjective experience of physicians, increasing the risks of misjudgment and missed diagnosis. In recent years, with the rapid development of artificial intelligence technology and computer vision models, medical diagnosis technologies using AI have made remarkable progress in the research of lymph node detection, and automatically identifying lymph nodes through models has gradually become the mainstream trend.
[0027] The current mainstream recognition methods include the segmentation algorithm represented by Unet and the detection algorithm represented by Yolo. In these algorithms, feature fusion is a very crucial step: the low-level features reflect more of the contour information and boundary information of the nodules, while the high-level features have stronger semantic information and reflect the potential patterns of the nodules themselves. Therefore, when making the final prediction, the low-level features and high-level features are usually fused to more comprehensively evaluate the nodule area. The existing feature fusion methods only upsample the high-level features through nearest neighbor interpolation or bilinear interpolation and simply add or splice them with the low-level features, which inevitably leads to the following problems: 1. There is a feature inconsistency between the high-level features and the low-level features, which is manifested in that there may be different descriptions for different regions of the same object. Therefore, the existing feature fusion methods will expand the inconsistent expression (the size of the inconsistent expression features will be expanded after upsampling), thus exacerbating the deterioration of the model performance; 2. The simple interpolation method will exacerbate the smoothing of the features, resulting in inaccurate boundary recognition.
[0028] Based on the above situation, the embodiments of the present invention provide a feature fusion method, an image segmentation method, an electronic device, and a storage medium. After aggregating the high-level feature map and the low-level feature map to obtain an aggregated feature map, a low-pass filter and a high-pass filter are generated according to the aggregated feature map, and then the low-pass filter and the high-pass filter are used to smooth and sharpen the high-level feature map and the low-level feature map respectively before fusion, so that more smooth and consistent features can be fused and generated, greatly reducing the inconsistency of the overall features and making the boundary recognition more accurate.
[0029] To implement the process steps and functions of the various examples of the present invention, please refer to Figure 1 , Figure 1 which is a schematic structural block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 may be a medical device, including a memory 101 and a processor 103. The memory 101 and the processor 103 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these components may be electrically connected to each other through one or more communication buses 104 or signal lines. The memory 101 may be used to store software programs and modules, and the processor 103 executes the software programs and modules stored in the memory 101 to perform various functional applications and data processing.
[0030] The electronic device 100 can be, but is not limited to, a personal computer (PC), a server, a distributed computer, and so on. It can be understood that the electronic device 100 is not limited to a physical server, and can also be a virtual machine on a physical server, a virtual machine built on a cloud platform, or other computers that can provide the same functions as the server or virtual machine. The operating system of the electronic device 100 can be, but is not limited to, the Windows system, the Linux system, and so on.
[0031] Among them, the memory 101 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), and so on.
[0032] The communication connection between the electronic device 100 and an external device is realized through at least one communication interface 102 (which can be wired or wireless).
[0033] The processor 103 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the embodiments of the present invention can be completed by the integrated logic circuit in the hardware of the processor 103 or by instructions in software form. The processor 103 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), and so on; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0034] It can be understood that Figure 1 The structure shown is only schematic, and the electronic device 100 may also include more or fewer components than those Figure 1 shown, or have a different configuration from that Figure 1 shown. Figure 1The components shown in [ID] can be implemented using hardware, software, or a combination thereof.
[0035] The following provides an exemplary description of the feature fusion method provided by the present invention. Figure 2 The flowchart of a feature fusion method provided by an embodiment of the present invention is shown in Figure 2 The execution subject of this method can be the above-mentioned Figure 1 shown electronic device 100. This method includes the following steps as Figure 2 described: S210: Obtain the extracted deep feature map and shallow feature map; the resolution of the deep feature map is smaller than that of the shallow feature map.
[0036] Among them, the deep feature map and the shallow feature map can be obtained by extracting ultrasonic images, etc. by a segmentation network or a detection network, etc. The resolution of the shallow features is larger, with better resolution and detailed information. The resolution of the deep features is smaller, with a more abstract form but rich semantic content.
[0037] S220: Aggregate the deep feature map and the shallow feature map to obtain an aggregated feature map.
[0038] S230: Generate a low-pass filter according to the aggregated feature map, and use the low-pass filter to smooth the deep feature map to obtain a consistency feature map.
[0039] S240: Generate a high-pass filter according to the aggregated feature map, and use the high-pass filter to sharpen the shallow feature map to obtain a high-frequency feature map.
[0040] S250: Fuse the consistency feature map and the high-frequency feature map to obtain a fused feature map.
[0041] The deep feature map and shallow feature map extracted by a segmentation network, a detection network, etc. are obtained. First, the deep feature map and the shallow feature map are combined in any way to generate a new aggregated feature map. Subsequently, based on the characteristics of the aggregated feature map, a low-pass filter and a high-pass filter are respectively generated. The low-pass filter is used to smooth the deep feature map to obtain a consistency feature map, which can better reflect the global consistency information in the deep feature map and reduce local noise interference. The high-pass filter is used to sharpen the shallow feature map to obtain a high-frequency feature map, which highlights the detailed information in the shallow feature map and enhances the local contrast of the image. Finally, the consistency feature map and the high-frequency feature map are combined to generate a final fused feature map. This fused feature map synthesizes the global semantic information of the deep feature map and the local detailed information of the shallow feature map, fuses to generate a smoother consistency feature, greatly reduces the inconsistency of the overall features, and can perform better in various tasks (such as image segmentation, object detection, super-resolution reconstruction, etc.).
[0042] In step S220, there are various ways to aggregate the deep feature map and the shallow feature map. They can be added, concatenated, etc. Exemplarily, in this embodiment, the deep feature map and the shallow feature map are concatenated to obtain the aggregated feature map. See Figure 3 , the above step S220 may include the following steps: S221: Upsample the deep feature map according to the resolution of the shallow feature map to obtain an upsampled high-resolution deep feature map.
[0043] S222: Concatenate the high-resolution deep feature map and the shallow feature map to obtain the aggregated feature map.
[0044] The simple aggregation of the deep feature map and the shallow feature map can be directly completed by linear interpolation. Since the size of the deep feature map is smaller than that of the shallow feature map, first, the deep feature map is upsampled. The upsampling can be performed by methods such as nearest interpolation or bilinear interpolation to obtain a high-resolution deep feature map with the same resolution as the shallow feature map. Subsequently, the high-resolution deep feature map and the shallow feature map are concatenated along the channel dimension to obtain the aggregated feature map. The steps of aggregating the deep feature map and the shallow feature map can be expressed by the following formula:
[0045] Among them, is the aggregated feature map, is the shallow feature map, is the deep feature map, resize represents the upsampling operation, and cat represents the concatenation operation.
[0046] In step S230, a low-pass filter is generated using the aggregated feature map to adaptively retain consistent feature information. Refer to Figure 4 , which may include the following steps: S231: According to a preset filter, extract the neighborhood feature relationship at each pixel position of the aggregated feature map; the size of the neighborhood is the size of the preset filter.
[0047] S232: Normalize each neighborhood feature relationship to obtain a low-pass filter with the sum of corresponding weights equal to 1.
[0048] S233: Apply each low-pass filter to the high-resolution deep feature map through matrix multiplication to obtain a consistent feature map.
[0049] For example, for a preset filter with a size of K, extract the neighborhood feature relationship at each pixel position of the aggregated feature map. This neighborhood feature relationship is a K*K window. Exemplarily, refer to Figure 5 , which shows the neighborhood sizes at different pixel positions. If K = 3, the size of the neighborhood feature relationship is 3*3. Taking the current pixel position as the center, there are 9 pixel points in the neighborhood at the middle position of the aggregated feature map, while there are only 4 pixel points in the neighborhoods at the four corner positions, and only 6 pixel points in the neighborhoods at the four side positions. Thus, the neighborhood feature relationship at each pixel position can be obtained. It can be represented by the following formula:
[0050] Where, is a filter of size K, is the aggregated feature map, is all the neighborhood feature relationships.
[0051] Then, normalize each neighborhood feature relationship to obtain a low-pass filter with the sum of corresponding weights equal to 1. For example, perform a softmax operation on a 3*3 neighborhood feature relationship to constrain the sum of weights within the filter to be 1, obtaining a low-pass filter. For example:
[0052] Where, the calculation formula for normalizing each neighborhood feature relationship can be:
[0053] Where, represents the feature relationship at the (i, j)th position belonging to the neighborhood at the (m, n) coordinate, represents the low-pass filter at the (m, n) coordinate.
[0054] Through the above operations, a low-pass filter at each pixel position of the aggregated feature map can be obtained. Its weight value reflects the correlation between this position and the features in the surrounding area. From the perspective of image information evaluation, high-frequency information usually appears at the junctions between objects and the background, and between objects, and has a relatively small correlation with the object contours, which is reflected as a relatively small weight value in the filter.
[0055] Finally, the low-pass filter is applied to the upsampled high-resolution deep feature map through matrix multiplication to obtain a consistency feature map. The formula for calculating this consistency feature map can be:
[0056] where represents the consistency feature map, represents the features at the same position of the corresponding low-pass filter in the high-resolution deep feature map.
[0057] Exemplarily, since the low-pass filter is obtained from the aggregated feature map, if the aggregated feature map is 20 * 20, there are 400 pixel points, and correspondingly 400 low-pass filters will be generated. Since the high-resolution deep feature map is upsampled from the shallow feature map, the high-resolution deep feature map is also 20 * 20. Take any low-pass filter, for example, with a size of 3 * 3, corresponding to 9 pixel points, find the features at the positions of 9 pixels at the same position in the high-resolution deep feature map, perform matrix multiplication, and then fuse the structures after matrix multiplication of each low-pass filter to obtain the consistency feature map.
[0058] In step S240, a high-pass filter is generated using the aggregated feature to extract the high-frequency information component of the shallow feature. Refer to Figure 6 , which may include the following steps: S241: Extract the neighborhood feature relationship at each pixel position of the aggregated feature map according to a preset filter; the size of the neighborhood is the size of the preset filter.
[0059] S242: Perform normalization processing on each neighborhood feature relationship to obtain a low-pass filter with the sum of corresponding weights equal to 1.
[0060] S243: Subtract the low-pass filter from the unit kernel to obtain the high-pass filter.
[0061] S244: Apply each high-pass filter to the shallow feature map through matrix multiplication to obtain a high-frequency feature map.
[0062] The overall process of steps S241 to S244 is similar to that of the above steps S231 to S233. The difference is that due to the low correlation of high-frequency information, the weights in its filter are smaller. Therefore, in order to obtain a high-pass kernel, the identity kernel is subtracted from the low-pass kernel, and finally the high-pass kernel is used to process the shallow features. Therefore, the generation method of the high-pass filter is to subtract the low-pass filter from the identity kernel to obtain the high-pass filter. The generation of the low-pass filter and the process of multiplying each high-pass filter by matrix multiplication on the shallow feature map will not be elaborated here. The entire process of steps S241 to S244 can be expressed by the following formula:
[0063]
[0064]
[0065] Among them, is a filter of size K, is the aggregated feature map, is all the neighborhood feature relationships, represents the feature relationship at the (i, j)th position in the neighborhood belonging to at the (m, n) coordinates, represents the high-pass filter at the (m, n) coordinates, represents the high-frequency feature map, represents the feature at the same position of the corresponding high-pass filter in the shallow feature map.
[0066] Furthermore, in order to retain the gradient information, the high-frequency feature map and the shallow feature map can also be added to obtain a high-frequency feature map that retains the gradient information.
[0067] Finally, the consistency feature map and the high-frequency feature map are fused to obtain the final fused feature map. This fusion process can be expressed by the following formula:
[0068] Among them, represents the fused feature map, represents the consistency feature map, represents the high-frequency feature map, represents the shallow feature map.
[0069] See Figure 7 , Figure 7 which is a schematic flow diagram of a feature fusion provided by an embodiment of the present invention. The shallow feature map y d and the deep feature map y n are fused. First, y n is upsampled to be the same size as y dHigh-resolution deep feature map y with the same resolution n+d , y n+d And y d Are concatenated to obtain the aggregated feature map y. Through the filter with the preset size K of y, a low-pass filter is generated. Subtracting the low-pass filter from the unit kernel obtains the high-pass filter. Then, the high-resolution deep feature map y n+d Is multiplied by the low-pass filter matrix to obtain the consistency feature map y f , The shallow feature map y d Is multiplied by the high-pass filter matrix to obtain the high-frequency feature map y g , y g And then added to y d To obtain y g+d Retain the gradient information. Finally, the added y g+d And y f Are concatenated to obtain the final fused feature map y z .
[0070] Based on the above feature fusion method, the embodiment of the present invention also provides another feature fusion method. See Figure 8 , The method includes the following steps: S310: Obtain multiple extracted feature maps; among them, the multiple feature maps include the first feature map to the Nth feature map sorted in ascending order of resolution, and N is an integer greater than or equal to 2.
[0071] S320: Fuse the first feature map and the feature maps after the first feature map through the feature fusion method described above to obtain the fused feature map.
[0072] S330: Fuse the fused feature map and the feature maps after the second feature map through the feature fusion method described above to obtain a new fused feature map, until the fused feature map and the Nth feature map are fused through the feature fusion method described above to obtain the final feature fusion map.
[0073] After obtaining the multiple extracted feature maps, through the above feature fusion method, the effective fusion of different information and features with different resolutions is completed from deep to shallow to obtain the semantic features for segmentation. Exemplarily, see Figure 9 , For example, the feature maps y1, y2, y3, and y4 from shallow to deep are obtained. First, y3 and y4 are fused to obtain yz3 with the same resolution as y3. Then, yz3 and y2 are fused to obtain yz2 with the same resolution as y2. Finally, yz2 and y1 are fused to obtain the final fused feature map yz1.
[0074] Furthermore, the embodiment of the present invention also provides an image segmentation method. See Figure 10 , The method includes the following steps: S410: Extract features from the image to be segmented through a segmentation network to obtain multiple feature maps.
[0075] S420: Fuse each feature map through a feature fusion method to obtain a fused feature map.
[0076] S430: Input the fused feature map into the classifier of the segmentation network to obtain the segmentation result of the image to be segmented.
[0077] Through feature extraction of the image to be segmented, fusion of each extracted feature through the above-mentioned feature fusion method, and finally inputting the fused feature map into the classifier, the segmentation result of the image to be segmented can be obtained. The above-mentioned feature fusion method can fuse and generate smoother consistent features, greatly reducing the inconsistency of the overall features, increasing the discriminability of the data, and thus promoting the performance improvement of the model segmentation effect.
[0078] In summary, a feature fusion method, an image segmentation method, an electronic device, and a storage medium provided by an embodiment of the present invention aggregate deep feature maps and shallow feature maps to obtain an aggregated feature map, then generate a low-pass filter and a high-pass filter according to the aggregated feature map, and then use the low-pass filter and the high-pass filter to smooth and sharpen the deep feature map and the shallow feature map respectively before fusing, so that smoother consistent features can be fused and generated, greatly reducing the inconsistency of the overall features, and making the boundary recognition more accurate.
[0079] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and a module, a program segment, or a part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0080] In addition, in each embodiment of the present invention, each functional module can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0081] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0082] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0083] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
Claims
1. A feature fusion method, characterized in that: The method comprises: Acquire the extracted deep feature map and shallow feature map; the resolution of the deep feature map is smaller than the resolution of the shallow feature map; Aggregating the deep feature map with the shallow feature map to obtain an aggregated feature map; Generate a low-pass filter according to the aggregated feature map, and use the low-pass filter to smooth the deep feature map to obtain a consistent feature map; Generate a high-pass filter according to the aggregated feature map, and use the high-pass filter to sharpen the shallow feature map to obtain a high-frequency feature map; The consistency feature map is fused with the high-frequency feature map to obtain a fused feature map.
2. The method according to claim 1, characterized in that The step of aggregating the deep feature map with the shallow feature map to obtain an aggregated feature map includes: Upsampling the deep feature map according to the resolution of the shallow feature map to obtain a high-resolution deep feature map after upsampling; The high-resolution deep feature map is concatenated with the shallow feature map to obtain an aggregated feature map.
3. The method according to claim 2, characterized in that The step of generating a low-pass filter according to the aggregated feature map and using the low-pass filter to smooth the deep feature map to obtain a consistent feature map includes: According to a preset filter, extracting a neighborhood feature relationship of each pixel position of the aggregate feature map; the size of the neighborhood is the size of the preset filter; Normalizing each of the neighborhood feature relationships to obtain a low-pass filter whose sum of corresponding weights is 1; Each of the low-pass filters is applied to the high-resolution deep feature map through matrix multiplication to obtain a consistency feature map.
4. The method according to claim 3, characterized in that The calculation formula for normalizing the neighborhood feature relationships is: in, Indicates that the coordinates (m, n) belong to The characteristic relationship of the (i, j)th location in the neighborhood, represents a low-pass filter at the (m, n) coordinate; The formula for calculating the consistency feature map is: in, represents the consistency feature map, Represents the features at the same position of the corresponding low-pass filter in the high-resolution deep feature map.
5. The method according to claim 1, characterized in that The step of generating a high-pass filter according to the aggregated feature map and using the high-pass filter to sharpen the shallow feature map to obtain a high-frequency feature map includes: According to a preset filter, extracting a neighborhood feature relationship of each pixel position of the aggregate feature map; the size of the neighborhood is the size of the preset filter; Normalizing each of the neighborhood feature relationships to obtain a low-pass filter whose sum of corresponding weights is 1; Subtracting the low-pass filter from the unit core to obtain a high-pass filter; Each of the high-pass filters is applied to the shallow feature map through matrix multiplication to obtain a high-frequency feature map.
6. The method according to claim 5, characterized in that The method further comprises: The high-frequency feature map is added to the shallow feature map to obtain a high-frequency feature map that retains gradient information.
7. A feature fusion method, characterized in that: The method comprises: Acquire multiple extracted feature maps; wherein the multiple feature maps include a first feature map to an Nth feature map sorted from small to large in terms of resolution, where N is an integer greater than or equal to 2; fusing the first feature map with a feature map subsequent to the first feature map by the feature fusion method according to any one of claims 1 to 6 to obtain a fused feature map; The fused feature map and the feature map after the second feature map are fused by the feature fusion method described in any one of claims 1 to 6 to obtain a new fused feature map, until the fused feature map is fused with the Nth feature map by the feature fusion method described in any one of claims 1 to 6 to obtain a final feature fusion map.
8. An image segmentation method, characterized in that: The method comprises: The image to be segmented is subjected to feature extraction through a segmentation network to obtain multiple feature maps; The feature maps are fused by the feature fusion method according to claim 7 to obtain a fused feature map; The fused feature map is input into the classifier of the segmentation network to obtain the segmentation result of the image to be segmented.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Medical image fusion method based on deconvolution network and guided filtering
CN111311529A
Image semantic segmentation method and system based on frequency perception feature fusion
CN116704180A
Neck metastatic lymph node identification method and device, electronic equipment and storage medium
CN119027394A
Appartus and method for quantifying lesion in biometric image
US20230274424A1