Image feature extraction method and device, and electronic equipment

By using a global-local feature fusion method, hybrid guided features are generated using Transformer and CNN models, which solves the problem of inaccurate portrait matting in complex backgrounds by deep learning methods and improves the accuracy and robustness of image segmentation.

CN121010774APending Publication Date: 2025-11-25WONDERSHARE TECH (HUNAN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510898340.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing deep learning methods struggle to accurately handle portrait matting in complex backgrounds, especially for fine details such as hair and edges, resulting in inaccurate matting results.

Method used

By fusing global and local features, a hybrid guiding feature with both semantic consistency and detail preservation is generated. The global contextual features of the image are extracted using the Transformer model and combined with the local features of the CNN model, and then fused using a dynamic weighting mechanism.

Benefits of technology

It improves the portrait matting effect in complex backgrounds, enhances the robustness of the model, and can better cope with interference such as noise and occlusion, thereby improving the accuracy and fineness of image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010774A_ABST
    Figure CN121010774A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image feature extraction method. The method comprises the following steps: extracting global context features of an image; extracting local features of the image; and fusing the global context feature and the local feature of the image to generate a hybrid guide feature of the image. According to the image feature extraction method provided by the embodiment of the invention, through interactive fusion of global-local features, mixed guide features with semantic consistency and detail retention capability are generated, so that a portrait matting effect under a complex background is improved. The embodiment of the invention further provides an image feature extraction device and electronic equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and more specifically, to a method, apparatus, and electronic device for extracting image features. Background Technology

[0002] With the rapid development of deep learning, image segmentation technology has made significant progress in the past few years. Early image segmentation methods mainly relied on pixel-based classification techniques, such as thresholding, region growing, and edge detection. However, these methods often performed poorly when faced with complex backgrounds, lighting variations, or changes in pose. Deep learning, especially the emergence of models such as Fully Convolutional Networks (FCN), U-Net, and DeepLab, has greatly improved the performance and accuracy of image segmentation, particularly in handling complex scenes rich in semantic information, demonstrating stronger robustness and fine-grained characteristics.

[0003] While these deep learning methods have improved detection accuracy to some extent, they still have some significant drawbacks. For example, CNNs (Convolutional Neural Networks) often struggle to accurately handle details, especially fine details like hair and edges, when performing portrait matting against complex backgrounds. CNNs may make false judgments in complex or irregular backgrounds, leading to inaccurate matting results. While Transformers have advantages in global modeling, they are often less refined than CNNs in handling local image details (such as fine hair strands). Summary of the Invention

[0004] To address the problems existing in the prior art, this application provides an image feature extraction method, apparatus, and electronic device. By fusion of global and local features, a hybrid guiding feature with both semantic consistency and detail preservation capabilities is generated, thereby improving the portrait cutout effect in complex backgrounds.

[0005] In a first aspect, embodiments of this application provide a method for extracting image features, including:

[0006] Extract the global contextual features of the image;

[0007] Extract local features from the image; and

[0008] The global contextual features and local features of the image are fused to generate a hybrid guiding feature of the image.

[0009] Further, the extraction of global contextual features of the image includes:

[0010] The Transformer model is used to extract the classification label features of the image; and

[0011] Global average pooling is performed on the extracted classification label features of the image to generate a global feature vector corresponding to the global context features.

[0012] Further, the extraction of local features of the image includes:

[0013] The image feature map is output using shallow convolutional layers of a CNN model; and

[0014] While maintaining the spatial resolution of the image feature map, generate a local feature tensor corresponding to the local features of the image.

[0015] Furthermore, the process of fusing the global contextual features and local features of the image to generate the hybrid guidance features of the image includes:

[0016] The global feature vector is mapped to the channel dimension of the local feature tensor through a fully connected layer;

[0017] Upsample the mapped global features to make their spatial size consistent with the local feature tensor; and

[0018] The image's hybrid guiding features are generated by fusing global and local features through a dynamic weighting mechanism.

[0019] Furthermore, the dynamic weighting mechanism includes:

[0020] Construct a channel attention module that takes the sum of global and local features as input to generate channel weights; and

[0021] The local features are weighted using the channel weights and then added to the upsampled global features.

[0022] Furthermore, after fusing global and local features through a dynamic weighting mechanism to generate the hybrid guiding features of the image, the method further includes:

[0023] The channel dimension is further compressed using convolutional layers.

[0024] Furthermore, after fusing global and local features through a dynamic weighting mechanism to generate the hybrid guiding features of the image, the method further includes:

[0025] Residual connections preserve the original local feature information.

[0026] Secondly, embodiments of this application also provide an image feature extraction apparatus, comprising:

[0027] A global feature extraction module is used to extract global contextual features of the image;

[0028] A local feature extraction module is used to extract local features of the image; and

[0029] The feature fusion module is used to fuse the global context features and local features of the image to generate a hybrid guiding feature of the image.

[0030] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to execute the program to implement the image feature extraction method according to the first aspect described above.

[0031] Fourthly, embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, the computer program being used to implement the image feature extraction method according to the first aspect described above.

[0032] Fifthly, embodiments of this application also provide a computer program product having a computer program stored thereon, the computer program being used to implement the image feature extraction method according to the first aspect described above.

[0033] The embodiments of this application bring the following beneficial effects:

[0034] In the image feature extraction method of this application embodiment, the global context features of the image are first extracted, then the local features of the image are extracted, and finally the global context features and local features of the image are fused to generate the hybrid guiding features of the image. The multi-scale information complementarity of the image feature extraction method provided in this application embodiment can avoid the limitations of a single scale, improve task adaptability, and the fused features can enhance the robustness of the model and cope with interference such as noise and occlusion, thereby improving the portrait matting effect in complex backgrounds. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0036] Figure 1 A schematic flowchart illustrating the image feature extraction method provided in this application embodiment;

[0037] Figure 2 A schematic diagram of the network architecture of the image feature extraction method provided in the embodiments of this application;

[0038] Figure 3A structural block diagram of the image feature extraction device provided in the embodiments of this application;

[0039] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0040] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of this application.

[0042] In the specification, claims, and accompanying drawings of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0043] Figure 1 and Figure 2 These are flowcharts and network architecture diagrams of an image feature extraction method according to an embodiment of this application. Figure 1 and Figure 2 As shown, the image feature extraction method of this application embodiment includes the following steps:

[0044] S101: Extract the global contextual features of the image;

[0045] S102: Extract local features of the image; and

[0046] S103: Fuse the global context features and local features of the image to generate a hybrid guidance feature of the image.

[0047] Specifically, global contextual feature extraction focuses on the overall semantics and spatial relationships of the image. It employs global pooling (such as average / max pooling) or self-attention mechanisms (such as the Transformer architecture) to model global dependencies and generate a global feature vector. This feature contains high-level semantic information such as scene category and object layout, making it suitable for tasks like image classification. Local feature extraction, on the other hand, focuses on image details and texture information. It extracts low-level features such as edges and textures through convolution operations, or generates local region features using region segmentation algorithms. This feature preserves detailed information such as object edge contours and local textures, which is crucial for tasks like object detection. Feature fusion combines global and local features through concatenation, weighting, or attention mechanisms to generate hybrid guided features. Concatenation fusion directly merges feature channels, weighted fusion dynamically adjusts the feature contribution ratio, and attention fusion adaptively optimizes feature weights. The fused features balance global semantics and local details, enhancing the model's ability to recognize complex scenes.

[0048] Therefore, in the image feature extraction method provided in this application embodiment, the global context features of the image are first extracted, then the local features of the image are extracted, and finally the global context features and local features of the image are fused to generate the hybrid guiding features of the image. The multi-scale information complementarity of the image feature extraction method provided in this application embodiment can avoid the limitations of a single scale, improve task adaptability, and the fused features can enhance the robustness of the model, cope with interference such as noise and occlusion, thereby improving the portrait matting effect in complex backgrounds.

[0049] Furthermore, in some embodiments of this application, the extraction of global contextual features of the image includes:

[0050] The Transformer model is used to extract the classification label features of the image; and

[0051] Global average pooling is performed on the extracted classification label features of the image to generate a global feature vector corresponding to the global context features.

[0052] Specifically, the image is segmented into non-overlapping blocks and serialized into tags, with learnable classification tokens added as global semantic carriers. A multi-layer Transformer encoder iteratively updates the tag features, utilizing a self-attention mechanism to model long-distance dependencies, allowing the classification tags to gradually integrate global contextual information. The final feature vector of the classification tags here contains global semantics such as scene category and overall object relationships. Next, global average pooling is performed on the classification tag feature vector (or the feature matrix of all tags), compressing the dimensionality by calculating the mean of feature channels to generate a fixed-length global feature vector. This process, through the collaboration of Transformer and global pooling, achieves efficient global contextual feature extraction, providing semantically rich and computationally friendly feature representations for subsequent tasks.

[0053] Further, the extraction of local features of the image includes:

[0054] The image feature map is output using shallow convolutional layers of a CNN model; and

[0055] While maintaining the spatial resolution of the image feature map, generate a local feature tensor corresponding to the local features of the image.

[0056] Specifically, the input image is processed using shallow convolutional layers (such as the first few convolutional blocks) of the CNN model. Shallow networks, with their smaller receptive fields, focus more on low-level visual features such as edges and textures. The convolutional operation extracts local patterns pixel-by-pixel through a sliding window, generating feature maps rich in detail. Next, pooling layers (such as max pooling or average pooling) or convolutions with a stride greater than 1 are avoided to prevent a decrease in the spatial resolution of the feature maps. Zero padding or adjustments to the convolutional kernel parameters ensure that the output feature map maintains spatial consistency with the input image. This process achieves accurate extraction of local image features, providing a detailed and spatially aligned feature representation for subsequent tasks.

[0057] Furthermore, in some embodiments of this application, the step of fusing the global context features and local features of the image to generate the hybrid guidance features of the image includes:

[0058] The global feature vector is mapped to the channel dimension of the local feature tensor through a fully connected layer;

[0059] Upsample the mapped global features to make their spatial size consistent with the local feature tensor; and

[0060] The image's hybrid guiding features are generated by fusing global and local features through a dynamic weighting mechanism.

[0061] Specifically, the global feature vector (such as the [CLS] feature of a Transformer) is first expanded to the channel dimension (such as C channels) of the local feature tensor through a fully connected layer, generating a tensor of, for example, 1×1×C, ensuring uniform channel dimensions. Next, the mapped global features are upsampled (e.g., by bilinear interpolation or transposed convolution) to make their spatial dimensions (height × width) consistent with the local feature tensor, achieving spatial dimension alignment and supporting pixel-by-pixel fusion. Finally, the weights (α and β, where α + β = 1) of the global and local features are dynamically adjusted through an attention mechanism (such as an SE module or a CBAM module) to generate a hybrid guided feature tensor. This process, through the collaborative fusion of global and local features, provides semantically rich and detail-accurate hybrid guided features for image understanding tasks, combining efficiency and adaptability.

[0062] Furthermore, in some embodiments of this application, the dynamic weighting mechanism includes:

[0063] Construct a channel attention module that takes the sum of global and local features as input to generate channel weights; and

[0064] The local features are weighted using the channel weights and then added to the upsampled global features.

[0065] That is, firstly, the global features (F) global ) and local features (F local ) Element-wise addition generates a fused feature map (F) sum ), as input to the attention module. Using an SE module structure, F sum Global average pooling is used to compress the spatial dimension, and then a channel weight vector (w∈RC) is generated through a fully connected layer and sigmoid activation, where C is the number of channels. The weight vector reflects the importance of each channel. Next, the channel weights w are combined with the local features F. local Channel-by-channel multiplication enhances local details (such as edges and textures) in key channels. Finally, the weighted local features are multiplied by the upsampled global features F. global Element-wise addition generates hybrid guided features that balance global semantics with local details. This process, through dynamic channel weighting, achieves collaborative optimization of global and local features, further enhancing the model's ability to represent complex scenes.

[0066] Furthermore, in some embodiments of this application, after fusing global and local features through a dynamic weighting mechanism to generate the hybrid guiding features of the image, the method further includes:

[0067] The channel dimension is further compressed using convolutional layers.

[0068] Specifically, after generating the hybrid guidance features through dynamic weighted fusion, the channel dimension is further compressed by a convolutional layer to optimize the feature expression efficiency. For example, a convolutional layer with a 1×1 convolutional kernel is used to compress the channel dimension of the hybrid guidance features. The 1×1 convolution reduces feature redundancy while retaining spatial information by adjusting the number of output channels (e.g., from C to C', where C' < C). The above process can further improve the model efficiency and flexibility by compressing the channel dimension through, for example, a 1×1 convolutional layer, while retaining the core information of the hybrid features, and providing a more concise feature input for subsequent tasks.

[0069] Further, in some embodiments of the present application, after fusing the global and local features through the dynamic weighting mechanism to generate the hybrid guidance features of the image, it further includes:

[0070] Retaining the original local feature information through a residual connection.

[0071] Specifically, after generating the hybrid guidance features through dynamic weighted fusion, integrating the original local feature information into the hybrid features through a residual connection can enhance the feature robustness. For example, adding the original local feature (F local ) and the hybrid guidance feature (F hybrid ) element-wise to generate the final feature (F final = F hybrid + F local ). This operation directly retains the original information of the local features and avoids the loss of details due to the fusion process.

[0072] Figure 3 is the structural block diagram of the image feature extraction device 200 according to an embodiment of the present application. As Figure 3 shown, the image feature extraction device 200 according to an embodiment of the present application includes: a global feature extraction module 210, a local feature extraction module 220, and a feature fusion module 230, where: <000017l>The global feature extraction module 210 is configured to extract the global context features of the image;

[0074] The local feature extraction module 220 is configured to extract the local features of the image; and

[0075] The feature fusion module 230 is configured to fuse the global context features and local features of the image to generate the hybrid guidance features of the image.

[0076] The image feature extraction device provided in this application first extracts the global context features of the image, then extracts the local features of the image, and finally fuses the global context features and local features of the image to generate a hybrid guiding feature of the image. The multi-scale information complementarity of the image feature extraction device provided in this application avoids the limitations of a single scale, improves task adaptability, and the fused features enhance the robustness of the model, enabling it to cope with noise, occlusion, and other interference, thereby improving the portrait matting effect against complex backgrounds.

[0077] It should be noted that the specific implementation of the image feature extraction device in this application embodiment is similar to the specific implementation of the image feature extraction method in this application embodiment. Please refer to the description in the method section for details, which will not be repeated here.

[0078] Figure 4 This is a schematic diagram of the structure of the electronic device 300 according to an embodiment of this application.

[0079] like Figure 4 As shown, the electronic device 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from the storage section 302 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The CPU 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0080] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. A removable medium 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 310 as needed so that computer programs read from it can be installed into storage section 308 as needed.

[0081] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the functions defined in the electronic device of this application.

[0082] It should be noted that the computer-readable medium shown in this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electronic device, apparatus, or device that is electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0083] In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used or combined with an electronic device, apparatus, or device by instructions. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use or combined with an electronic device, apparatus, or device by instructions. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.

[0084] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of processing and receiving devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based electronic device that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0085] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be housed in a processor, which executes the program to implement the image feature extraction method.

[0086] Extract the global contextual features of the image;

[0087] Extract local features from the image; and

[0088] The global contextual features and local features of the image are fused to generate a hybrid guiding feature of the image.

[0089] In another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable storage medium stores one or more programs, which, when used by one or more processors, execute the image feature extraction method described in this application:

[0090] Extract the global contextual features of the image;

[0091] Extract local features from the image; and

[0092] The global contextual features and local features of the image are fused to generate a hybrid guiding feature of the image.

[0093] In another aspect, this application also provides a computer program product, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer program product stores one or more programs, which, when used by one or more processors, execute the image feature extraction method described in this application:

[0094] Extract the global contextual features of the image;

[0095] Extract local features from the image; and

[0096] The global contextual features and local features of the image are fused to generate a hybrid guiding feature of the image.

[0097] The above description is merely a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural transformations made based on the content of this application's specification and drawings under the concept of this application, or direct / indirect applications in other related technical fields, are included within the patent protection scope of this application.

Claims

1. A method for extracting image features, characterized in that, include: Extract the global contextual features of the image; Extract local features from the image; and The global contextual features and local features of the image are fused to generate a hybrid guiding feature of the image.

2. The method according to claim 1, characterized in that, The extraction of global contextual features of the image includes: The Transformer model is used to extract the classification label features of the image; and Global average pooling is performed on the extracted classification label features of the image to generate a global feature vector corresponding to the global context features.

3. The method according to claim 2, characterized in that, The extraction of local features from the image includes: The image feature map is output using shallow convolutional layers of a CNN model; and While maintaining the spatial resolution of the image feature map, generate a local feature tensor corresponding to the local features of the image.

4. The method according to claim 3, characterized in that, The process of fusing global contextual features and local features of the image to generate hybrid guidance features of the image includes: The global feature vector is mapped to the channel dimension of the local feature tensor through a fully connected layer; Upsample the mapped global features to make their spatial size consistent with the local feature tensor; and The image's hybrid guiding features are generated by fusing global and local features through a dynamic weighting mechanism.

5. The method according to claim 4, characterized in that, The dynamic weighting mechanism includes: Construct a channel attention module that takes the sum of global and local features as input to generate channel weights; and The local features are weighted using the channel weights and then added to the upsampled global features.

6. The method according to claim 4, characterized in that, After generating the hybrid guiding features of the image by fusing global and local features through a dynamic weighting mechanism, the method further includes: The channel dimension is further compressed using convolutional layers.

7. The method according to claim 4, characterized in that, After generating the hybrid guiding features of the image by fusing global and local features through a dynamic weighting mechanism, the method further includes: Residual connections preserve the original local feature information.

8. An image feature extraction device, characterized in that, include: A global feature extraction module is used to extract global contextual features of the image; A local feature extraction module is used to extract local features of the image; and The feature fusion module is used to fuse the global context features and local features of the image to generate a hybrid guiding feature of the image.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to implement the image feature extraction method according to any one of claims 1-7 when executing the program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for implementing the image feature extraction method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Weak supervision crowd counting method, system, equipment and medium

    CN118298374A

  • Automatic matting method for adaptive feature extraction and semantic guidance

    CN119832253A

  • Scene understanding method and system for multi-level image feature extraction

    CN120070914A