Object processing method and device, electronic device, and storage medium

By utilizing the target feature map for overall fusion in a multi-scale network structure, the high latency and fixed fusion position problems caused by the feature fusion module in the existing technology are solved, and higher feature accuracy and network speed are achieved, which is suitable for a variety of task scenarios.

CN114821260BActive Publication Date: 2025-10-03GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210535929.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-17
Publication Date
2025-10-03
Estimated Expiration
2042-05-17

AI Technical Summary

Technical Problem

The high latency, low accuracy and fixed fusion position limitations caused by the feature fusion module in the existing multi-scale network structure affect the accuracy and processing speed of the network structure.

Method used

By obtaining multiple resolution feature maps of the object to be processed and using the target feature map for overall fusion, fragmented operations can be avoided, and unified fusion of multiple resolution feature maps can be achieved, thereby reducing network latency and improving feature accuracy.

Benefits of technology

It improves the accuracy of feature extraction and the precision of network structure, reduces processing delay, expands the scope of application and convenience, and adapts to various task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821260B_ABST
    Figure CN114821260B_ABST
Patent Text Reader

Abstract

The disclosed embodiments relate to an object processing method and apparatus, electronic device, and storage medium, and relate to the field of computer technology. The object processing method includes: obtaining an object to be processed, and fusing feature maps of the object to be processed at multiple resolutions in the network structure according to a target feature map of the network structure to obtain a fused feature map; transmitting the fused feature map to branches corresponding to the multiple resolutions to obtain an output feature map, and determining a target network structure based on the output feature map, so as to perform processing operations on the object to be processed through the target network structure. The technical solutions in the disclosed embodiments can improve the accuracy of features and improve the performance of the network structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an object processing method, an object processing device, an electronic device, and a computer-readable storage medium. Background Art

[0002] Multi-scale network structures can be widely used in various types of tasks, and multi-resolution branches and repeated fusion modules can be used in multi-scale network structures to obtain a larger field of view and multi-scale features.

[0003] In related technologies, each branch exchanges features through a fusion module. Alternatively, different feature fusion modules are used at the end of the network to exchange features. When each branch exchanges features through a fusion module, due to the large number of branches and the need for more fragmented operations, this results in higher latency, reduced feature accuracy, and impacts the accuracy and processing speed of the network structure. Furthermore, the location of feature fusion in related technologies is fixed, which has certain limitations.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention

[0005] The purpose of the present disclosure is to provide an object processing method and device, an electronic device, and a storage medium, thereby overcoming, at least to a certain extent, the problems of low feature extraction accuracy and poor network structure performance caused by the limitations and defects of related technologies.

[0006] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.

[0007] According to a first aspect of the present disclosure, there is provided an object processing method, comprising: obtaining an object to be processed, and fusing feature maps of multiple resolutions of the object to be processed in the network structure according to a target feature map of the network structure to obtain a fused feature map; transmitting the fused feature map to branches corresponding to the multiple resolutions to obtain an output feature map, and determining a target network structure based on the output feature map, so as to perform processing operations on the object to be processed through the target network structure.

[0008] According to a second aspect of the present disclosure, an object processing device is provided, comprising: a feature fusion module for acquiring an object to be processed, and fusing feature maps of multiple resolutions of the object to be processed in the network structure according to a target feature map of the network structure to obtain a fused feature map; a feature determination module for transmitting the fused feature map to branches corresponding to the multiple resolutions to obtain an output feature map, and determining a target network structure based on the output feature map, so as to perform processing operations on the object to be processed through the target network structure.

[0009] According to a third aspect of the present disclosure, there is provided an electronic device, including: a processor; and

[0010] A memory for storing executable instructions of the processor; wherein the processor is configured to execute the object processing method of the first aspect and its possible implementation manner by executing the executable instructions.

[0011] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the object processing method of the first aspect and its possible implementation methods are implemented.

[0012] In the object processing method, object processing device, electronic device and computer-readable storage medium provided in the embodiments of the present disclosure, on the one hand, according to the target feature map in the feature maps of multiple resolutions corresponding to the network structure, the feature maps of multiple resolutions are fused to obtain a fused feature map, avoiding the fragmented operation when the fusion module is used to fuse each branch in the related art, and instead performing an overall fusion operation, reducing the network delay caused by feature fusion, and enabling other separate features to be fused to each branch, thereby improving the accuracy of the extracted features and also improving the accuracy and processing speed of the network structure. On the other hand, as long as there are feature maps of multiple resolutions, the feature maps of multiple resolutions can be fused according to the target feature map of the network structure, avoiding the limitation of the related art that fusion can only be performed at a fixed position, increasing the scope of application and improving convenience, and being able to improve the effectiveness of fusion and increasing versatility.

[0013] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0015] Figure 1 A schematic diagram showing a system architecture to which the object processing method according to an embodiment of the present disclosure can be applied.

[0016] Figure 2 A schematic diagram schematically illustrates an object processing method in an embodiment of the present disclosure.

[0017] Figure 3 The schematic diagram schematically shows the structure of the network structure in the embodiment of the present disclosure.

[0018] Figure 4 The following schematically illustrates a flow chart of obtaining an output feature map in an embodiment of the present disclosure.

[0019] Figure 5 The following schematically illustrates the process flow of fusion in the embodiment of the present disclosure.

[0020] Figure 6 The following schematically illustrates the flow chart of sampling in an embodiment of the present disclosure.

[0021] Figure 7 A schematic diagram illustrating the determination of a branch feature graph for each branch in an embodiment of the present disclosure is schematically shown.

[0022] Figure 8 A schematic diagram schematically illustrates hybrid fusion in an embodiment of the present disclosure.

[0023] Figure 9 A block diagram schematically illustrates an object processing device in an embodiment of the present disclosure.

[0024] Figure 10 A block diagram schematically illustrates an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or that other methods, components, devices, steps, etc. may be employed. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.

[0026] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0027] In related technologies, the HRNet network structure uses multi-resolution branches and repeated fusion modules to capture a wider field of view and multi-scale features. HRNet has four branches corresponding to different resolutions, each retaining information at a different scale and exchanging features through a fusion module called Transition. While lightweight networks such as Fast-SCNN and BiSeNet lack the Transition fusion module, they do incorporate different feature fusion modules at the end of the network.

[0028] In a typical HR-Net, the Transition fusion module, while not requiring many floating point operations (FLOPs), requires a relatively large number of fragmented operations, which often result in higher I / O latency. Furthermore, because the network structure has branches corresponding to multiple resolutions, the latency associated with the fusion module increases exponentially with the number of branches. Furthermore, the fusion modules in HR-Net are manually designed and can only be placed in fixed locations. Networks like Fast-SCNN and BiSeNet only perform feature exchange at the end of the network structure. The placement and type of these modules should be optimized through NAS search to generate the optimal network structure.

[0029] To address the technical issues in related technologies, the present disclosure provides an object processing method that can be applied to scenarios where a network search is performed on a network structure to obtain a target network structure, and then a processing operation is performed on the target network structure. The processing operation can be performed on various types of tasks, such as dense image prediction tasks. Dense image prediction tasks can include, but are not limited to, semantic segmentation tasks, human pose recognition tasks, object detection tasks, and the like.

[0030] Figure 1 A schematic diagram showing a system architecture to which the object processing method and apparatus according to the embodiments of the present disclosure can be applied is shown.

[0031] like Figure 1As shown, the system architecture 100 may include a client 101 and a server 102. The client 101 may be an intelligent device, such as a smart phone, a computer, a tablet computer, a smart speaker, or the like. The client 101 obtains the object to be processed and sends the object to be processed to the server 102, so that the server 102 performs a processing operation on the object to be processed according to the target network structure. The object to be processed may include, for example, an image to be processed, a voice to be processed, a text to be processed, etc., which may be specifically determined according to the type corresponding to the processing operation. The server 102 may be a background system that provides object processing related services in the embodiment of the present disclosure, and may include a portable computer, a desktop computer, a smart phone, or a cluster formed by one or more electronic devices with computing functions, for fusing feature maps of multiple resolutions in the network structure to obtain the target network structure, and performing processing operations on the object to be processed sent by the client based on the target network structure. In addition, the client may also not need to send the object to be processed to the server, but may simply fuse the feature maps of multiple resolutions in the network structure by the client itself to obtain the target network structure, and perform processing operations on the object to be processed sent by the client based on the target network structure. For example, the processing object is subjected to semantic segmentation tasks, human posture recognition tasks, target detection tasks, etc., and which processing operation is performed is determined according to actual needs and actual application scenarios.

[0032] This object processing method can be applied to the application scenario of searching the network structure corresponding to the processing operation. Figure 1 As shown in , the client sends the object to be processed 101 to the server 102. The server 102 fuses the feature maps of multiple resolution branches in the network structure according to the target feature map of the network structure to obtain a fused feature map; transmits the fused feature map to multiple resolution branches to determine the output feature map, and performs a network structure search based on the output feature map to obtain the target network structure, and performs processing operations on the object to be processed sent by the client 101 based on the target network structure.

[0033] The server 102 may be the same as the client 101 , that is, both the client 101 and the server 102 are smart devices capable of performing computing functions, such as smart phones.

[0034] It should be noted that the object processing method provided in the embodiments of the present disclosure can be executed by the server 102. Accordingly, the object processing method can be set in the server 102 through a program or other means. The object processing method provided in the embodiments of the present disclosure can also be executed by the client 101. Accordingly, the object processing method can be set in the client 101 through a program or other means. In the embodiments of the present disclosure, the object processing method is executed by the end side represented by the client as an example.

[0035] Next, refer to Figure 2 The object processing method in the embodiment of the present disclosure is described in detail.

[0036] In step S210, an object to be processed is obtained, and feature maps of multiple resolution branches of the object to be processed in the network structure are fused according to a target feature map of the network structure to obtain a fused feature map.

[0037] In the embodiment of the present disclosure, the object to be processed can be determined according to actual needs, for example, it can be an image to be processed, etc., and the image to be processed can be an image of any resolution. The network structure can be a model used for the processing operation, and the network structure can be a multi-scale network structure. The multi-scale network structure refers to a network structure that uses images of multiple scales (image pyramids) as input and then fuses the results. The multi-scale network structure may include a parallel multi-branch network and a skip-layer connection network. In the embodiment of the present disclosure, the multi-scale network structure is described as a parallel multi-branch network. The parallel multi-branch network usually contains convolution kernels with different receptive fields, for example, it can contain 1×1 convolution, 3×3 convolution, 5×5 convolution and 3×3 maximum pooling. This structure extracts features from the four branches and fuses them, which are then used as the feature input for the next layer.

[0038] Figure 3 A schematic diagram of the network structure is shown schematically in FIG. Figure 3 As shown in , it mainly includes a basic structure 301 and a transition structure 302, and the transition structure is used to represent the fusion module. Figure 3 The network structure in has four stages, each stage consists of branches that extract features of different scales. Each branch unifies the scale of the extracted feature size by upsampling or downsampling, and then fuses the unified different scale features with each other. That is, the network structure has four branches corresponding to different resolutions, each branch retains information of different scales and exchanges features through the fusion module. The arrangement order between the branches corresponding to multiple resolutions can be determined according to the network structure, and for the same network structure, the multiple resolution branches it contains are fixed. For example, the starting point of the network structure can be used as the starting point to determine the branches corresponding to multiple different resolutions in sequence.

[0039] In an embodiment of the present disclosure, in the target stage, the network structure may include multiple resolution branches for performing feature extraction on the object to be processed to obtain features of multiple resolution branches, that is, feature maps of multiple resolutions. The resolutions of the feature maps obtained by the multiple resolution branches are different, but the models and other parameters corresponding to the multiple resolution branches are the same, and the input and output parameters of each resolution branch are the same. The target stage can be any stage of the network structure and is not specifically limited here. Convolution operations can be performed in parallel on the inputs of the object to be processed corresponding to multiple different resolution branches to extract feature maps of each resolution.

[0040] Figure 4 Figure A in FIG schematically shows a reference fusion method in the related art. Figure 4 Figure B in FIG schematically shows a schematic diagram of fusion based on the target feature map. Figure 4 As shown in Figure B in , the resolutions of multiple branches can include, but are not limited to, h / 4, h / 8, etc., and the dimensions corresponding to branches with different resolutions are also different, for example, the dimension of the h / 4 resolution branch is 32, the dimension of the h / 8 resolution branch is 64, etc. The features of multiple resolution branches (feature maps of multiple resolutions) can be further fused to obtain a fused feature map, which can be specifically expressed as formula (1):

[0041] Formula (1)

[0042] in, represents the fusion function, represents the fused feature map, are s feature maps with different resolutions.

[0043] Figure 5 The flowchart for fusion is shown schematically in FIG. Figure 5 As shown in , it mainly includes the following steps:

[0044] In step S510, the reference feature maps other than the target feature map in the feature maps of the multiple resolutions are scaled according to the attribute information of the target feature map to obtain a plurality of scaled feature maps.

[0045] In this step, the target feature map in the feature maps of multiple resolutions can be used as a benchmark to fuse the reference feature maps other than the target feature map. Among them, the branches corresponding to multiple different resolutions are arranged according to the order of the network structure. For the feature maps of multiple resolutions, the arrangement order can be the same as that of the branches corresponding to the multiple resolutions. Based on this, the feature map arranged at the last position in the feature maps of multiple resolutions can be used as the target feature map, and the remaining feature maps can be used as reference feature maps. For example, if the feature maps of multiple resolutions are represented as , then the feature map of the last resolution branch can be As the target feature map, all feature maps before the target feature map are As a reference feature map.

[0046] The attribute information of the target feature map refers to the scale of the target feature map. The scale can be used to describe the size of the target feature map. To express. Among them, Can be the width of the target feature map, It can be the height of the target feature map.

[0047] Based on this, the reference feature map can be scaled according to the attribute information of the target feature map. Exemplarily, a pooling operation can be performed on the reference feature map to perform scaling so that the attribute information of each reference feature map is the same as the attribute information of the target feature map. The pooling operation can be an average pooling operation, that is, a custom average pooling operation is performed on multiple reference feature maps of different resolutions to expand or reduce the scale of the reference feature map until the scale of the expanded or reduced reference feature map is the same as the scale of the target feature map. In addition, when performing the pooling operation, the dimension of the reference feature map can be adjusted, for example, its dimension can be adjusted by dimensionality reduction processing. Moreover, the feature map obtained by scaling the reference feature map according to the attribute information of the target feature map can be used as a scaled feature map.

[0048] It should be noted that since the resolution of the reference feature map is different, the scale of the reference resolution is also different. Therefore, the degree of scaling of each reference feature map according to the scale of the target feature map is also different. The scaling degree is specifically determined according to the difference between the scale of the reference feature map and the scale of the target feature map.

[0049] In step S520, a connection operation is performed on each scaled feature map to obtain a connection result, and a convolution operation is performed on the connection result to convert the feature maps of the multiple resolutions into the fused feature map.

[0050] In this step, a concatenation operation can be performed on the scaled feature maps corresponding to each resolution to obtain a concatenation result. For example, the concatenation operation can be performed by splicing the scaled feature maps according to the order of the branches corresponding to the multiple resolutions. If there are branches corresponding to s resolutions, the concatenation result can be obtained by splicing the scaled feature maps of s-1 resolutions, i.e., the number of scaled feature maps in the splicing result is the number of feature maps of the multiple resolutions minus one.

[0051] After obtaining the concatenation result, the concatenation result can be convolved to fuse feature maps of multiple resolutions based on the target feature map to generate a fused feature map. For example, the convolution operation can include a first convolution operation and a second convolution operation. The first convolution operation can be PWConv (Pointwise Convolution). The convolution kernel size of pointwise convolution is 1×1×M, where M is the number of channels in the previous layer. Therefore, pointwise convolution performs a weighted combination of the feature maps from the previous step in the depthwise direction to generate a new feature map, and the number of output feature maps is determined by the number of convolution kernels. The second convolution operation can be DWConv (Depth-wise Separable Convolution). The convolution kernel of depthwise separable convolution can be a 3×3 convolution kernel or other convolution kernel, depending on actual needs. In depthwise separable convolution, each convolution kernel is responsible for one channel, and each channel is convolved with only one convolution kernel. This process generates a feature map with the same number of channels as the input. For example, for a three-channel RGB image, a normal convolution kernel performs convolution on all three channels simultaneously. This means that after a single convolution, the output is a single number. However, with depthwise separable convolution, three convolutions are performed on each of the three channels, resulting in a single convolution that outputs three numbers.

[0052] Based on this, the connection result can be subjected to a first convolution operation to obtain a first convolution result; the first convolution result can be subjected to a second convolution operation to obtain a second convolution result, and the second convolution result can be determined as the fused feature map. Exemplarily, the connection result can be subjected to ordinary convolution processing through the first convolution operation to obtain the corresponding convolution result. Furthermore, depthwise separable convolution can be used to perform convolution processing on each channel of the first convolution result separately, thereby outputting the convolution result of each channel as the second convolution result. Based on the second convolution operation, it is possible to achieve the fusion of feature maps of multiple resolutions based on the target feature map, and obtain the second convolution results of the three channels as the fused feature map.

[0053] In the disclosed embodiments, feature maps of multiple resolutions of the network structure are fused based on the target feature map of the network structure, thus avoiding the fragmentation operation in related technologies. By fusing the reference feature map in the feature maps of multiple resolutions in the network structure, the network delay caused by feature fusion is reduced, and the accuracy and processing speed of the network structure are improved. Furthermore, because the features of other branches are integrated into the branches corresponding to each resolution, the accuracy and comprehensiveness of the fusion can be improved, thereby improving the precision of the acquired feature map.

[0054] Next, continue to refer to Figure 2As shown in , in step S220, the fused feature map is transmitted to the branches corresponding to the multiple resolutions to obtain an output feature map, and a target network structure is determined based on the output feature map to perform processing operations on the object to be processed through the target network structure.

[0055] In the embodiment of the present disclosure, after obtaining the fused feature map, the fused feature map can be transmitted to multiple branches corresponding to different resolutions. If there are branches corresponding to s resolutions, the fused feature map is separately transmitted to the branches corresponding to s resolutions.

[0056] For the branch corresponding to each resolution, the branch feature map of the branch corresponding to each resolution can be determined based on the feature map of each branch and the fusion feature map, and then the output feature map can be obtained. Figure 6 The flowchart for obtaining the output feature map is shown schematically in Figure 6 As shown in , it mainly includes the following steps:

[0057] In step S610, a convolution operation is performed on the fused feature map, and the convolution result is upsampled to obtain a sampling result;

[0058] In step S620, the branch feature map of each branch is determined according to the feature map of each branch and the sampling result, and the branch feature map of each branch is fused to obtain an output feature map; the branch feature map of each branch has the same attribute parameters as the feature map corresponding to the branch.

[0059] In the embodiment of the present disclosure, the convolution operation can be a point-by-point convolution PWConv. After performing point-by-point convolution on the fused feature maps corresponding to multiple channels, the convolution result can be determined, and the number of convolution results is the same as the number of convolution kernels. Furthermore, since the pooling operation on the reference feature map reduces the dimension of the reference feature map in the process of obtaining the fused feature map, in order to improve the accuracy, the convolution result can be upsampled to obtain the sampling result. The upsampling process is used to adjust the channel and resolution of the branch feature map of each branch so that the resolution, channel and feature map of each branch are consistent. Figure X The resolution and channel of si (the feature map input before fusion) are kept consistent, so as to ensure that the input features and output features of the branch corresponding to each resolution are consistent.

[0060] After the sampling result is determined through the upsampling operation, the feature map of the branch corresponding to each resolution can be fused with the sampling result to determine the branch feature map of the branch corresponding to each resolution. Figure 7As shown in , the feature map 701 of the branch corresponding to each resolution can be added to the sampling result 702 of the branch corresponding to each resolution to obtain a branch feature map 704 of the branch corresponding to each resolution. Exemplarily, the fusion of the feature map of the branch corresponding to each resolution and the sampling result can be performed by adding two matrices.

[0061] Continue to refer Figure 4 As shown in Figure B, the first branch has a resolution of h / 4 and a dimension of 32; the second branch has a resolution of h / 8 and a dimension of 64. For the first branch, its feature map has a resolution of h / 4 and a dimension of 32; for the second branch, its feature map has a resolution of h / 8 and a dimension of 64. Pooling can be performed on the feature maps of the first branch and the second branch based on the target feature map, and the scaled feature maps of the first and second branches can be concatenated to determine the concatenation result. Furthermore, point-by-point convolution and depth-wise separable convolution can be performed on the concatenation result to obtain a fused feature map. Next, the fused feature map can be split into the first and second branches. When splitting into each branch, point-by-point convolution can be performed on the fused feature map, and the convolution result can be upsampled to obtain a sampling result. The sampling result is fused with the feature map of each branch to obtain a branch feature map for each branch. For example, the branch feature map corresponding to the first branch has a resolution of h / 4 and a dimension of 32; the branch feature map corresponding to the second branch has a resolution of h / 8 and a dimension of 64. That is, the branch feature map of each branch has the same resolution as the feature map of its corresponding branch input.

[0062] Furthermore, the branch feature maps of the branches corresponding to each resolution can be fused to obtain the output feature map of the object to be processed at the target stage.

[0063] It should be noted that, in the embodiment of the present disclosure, the feature maps of the multiple resolutions may be fused after each network layer of the network structure; or the feature maps of the multiple resolutions may be fused at preset intervals. The preset interval refers to the fusion of feature maps of multiple resolutions at network layers with a fixed interval, and the preset interval may be, for example, 2 network layers, etc. That is, feature fusion may be performed every 2 network layers or every 3 network layers, which is not specifically limited here. Among them, the feature maps of multiple resolutions are fused after each network layer, and the accuracy of the obtained output feature map is greater than the accuracy of the fusion performed at the preset interval.

[0064] In addition, when there are branches corresponding to multiple resolutions, for feature maps of multiple resolutions, feature maps in a first range of the multiple resolutions are fused using the target feature map, and feature maps in a second range of the multiple resolutions are fused using a reference method. The first range can be a range where the number of branches is greater than a preset value, and the second range can be a range where the number of branches is less than the preset value. The preset value can be, for example, 2 or other numerical values, and is not specifically limited here. Since each branch corresponds to a resolution, the number of resolutions in the first range is greater than the number of resolutions in the second range. The reference method can be any fusion method, as long as it differs from fusion based on the target feature map. Therefore, when there are branches corresponding to multiple resolutions, it is also possible to fuse the range with fewer branches using the reference method and the range with more branches using the target feature map to improve accuracy. Alternatively, fusion can be performed entirely based on the target feature map. The choice of fusion method can be determined based on actual needs. The first range can be the first few layers of the network structure; the second range can be the last few layers of the network structure.

[0065] In the embodiment of the present disclosure, since the feature maps of branches corresponding to multiple different resolutions are connected, the branches corresponding to each resolution can fuse the feature maps of other branches, and each element in the output channel of the output feature map of the target stage receives contributions from all positions of all other input channels. The network structure starts from high resolution and gradually fuses low resolution information, and there is information exchange across resolutions at each fusion stage. By fusing the feature maps of branches corresponding to different resolutions, it is possible to more accurately express the features of high-resolution images, etc., and improve the accuracy and comprehensiveness of the obtained feature maps. In addition, since the same fusion operation can be performed on each resolution corresponding branch, the process of requiring more fragmented operations due to the large number of branches in the related art is avoided, the delay is reduced, and the accuracy and processing speed of the network structure are improved. Moreover, since the feature maps of multiple resolutions can be fused after each network layer of the network structure, the limitation of the related art that features can only be fused at fixed positions is avoided, and the scope of application and comprehensiveness are increased.

[0066] After obtaining the output feature map, a network structure search can be performed based on the output feature map to obtain the target network structure, and the target network structure can be used to perform processing operations on the object to be processed. In the embodiment of the present disclosure, after obtaining the output feature map of the target stage, the output feature map can be input into the next stage of the network structure for processing, thereby implementing a network structure search according to the search strategy to obtain the target network structure in the search space represented by the network structure, thereby determining the supernet structure. Furthermore, a processing operation can be performed on the object to be processed based on the supernet structure represented by the target network structure. Among them, the processing operation can be a target task, and the target task can be various types of tasks, such as classification tasks, detection tasks, segmentation tasks, human posture recognition tasks, etc., which can be determined according to the actual application scenario and actual needs.

[0067] In an embodiment of the present disclosure, a branch fusion module may be provided for executing the above-mentioned steps S210 and S220 to obtain an output feature map. The branch fusion module fuses the feature maps of branches corresponding to different resolutions in the network structure according to the target feature map of the network structure to obtain a fused feature map. The fused feature map is then transmitted to branches corresponding to multiple different resolutions to complete the network search and obtain the target network structure. Each layer of the network structure can be connected to a branch fusion module, and the number of branch fusion modules is positively correlated with the performance parameters of the target network structure. The performance parameters of the target network structure may include but are not limited to the accuracy and speed of the target network structure. That is, the more branch fusion modules there are, the higher the performance parameters of the target network structure.

[0068] It should be noted that the branch fusion module can be used alone or in combination with the reference fusion module. When used in combination, the branch fusion module can be placed at a location with a large number of branches in the network structure, while the reference fusion module can be placed at a location with a small number of branches in the network structure. Figure 8 As shown in , the network structure may include a reference fusion module 801 and a branch fusion module 802, the reference fusion module is used to fuse the feature maps of the second range in the multiple resolutions, and the branch fusion module is used to fuse the feature maps of the first range in the feature maps of the multiple resolutions.

[0069] It should be added that there will also be a zero module in the search space represented by the network structure. This module is parallel to the branch fusion module. When searching, once the zero module is selected, it can be considered that there is no need to execute the branch fusion module, that is, the zero module can be selected in the case of single resolution.

[0070] In some embodiments, the branch fusion module can be tested and verified based on a human posture recognition scenario. To avoid computational errors caused by different convolution operations, all convolution operations in the network structure can be uniformly set to DWConv convolution operations, or the network structure can be called DW-HRNet. When testing and verifying the branch fusion module, verification can be performed using two dimensions: branch fusion modules with different parameters and comparison with a reference fusion module.

[0071] First, the number of branch fusion modules can be tested and verified. To compare the impact of the number and location of branch fusion modules (EFMs) on the accuracy and speed of the network structure, an ablation experiment was conducted based on a human pose recognition dataset, and a comparative experiment with three different variables was designed. As shown in Table 1, the three variables are used to represent the performance parameters of the network structure, which may include but are not limited to complexity, human pose estimation metrics, and network speed. The human pose estimation metric can be the distance between the predicted key points and the annotated key points after head size normalization. "without" indicates that no branch fusion modules are included. The branch fusion module refers to the Efficient Fusion Module (EFM), and no reference fusion modules are included. "less" indicates a fixed spacing and a total of three EFM modules. "full" indicates that each layer is connected to an EFM module, for a total of eight EFM modules.

[0072]

[0073] As shown in Table 1, compared to a network without any fusion modules, the human pose estimation index of the network structure with three branch fusion modules improved by 0.93, and the human pose estimation index of the network structure with eight branch fusion modules improved by 1.34. Compared to a network without any fusion modules, the network complexity of the network structure with three branch fusion modules was reduced by 17.2%, and the network complexity of the network structure with eight branch fusion modules was reduced by 48.2%. Furthermore, the speed of the network structure with three branch fusion modules increased by 23.9% and 57.5%, respectively, while the speed of the network structure with eight branch fusion modules increased by 54% and 111.3%, respectively. Therefore, it can be concluded that the human pose estimation index, network complexity, and network speed are all positively correlated with the number of branch fusion modules.

[0074] Furthermore, the performance of the efficient fusion module and the reference fusion module can be compared through three different variables. In order to avoid calculation errors caused by different scenarios and different calculation methods, ablation experiments are also performed on the human posture recognition dataset. Referring to Table 2, the network structure containing 3 fusion modules and the network structure containing 3 branch fusion modules are compared, and the network structure containing 8 fusion modules and the network structure containing 8 branch fusion modules are compared. In Table 2, original represents the reference fusion module, the pooling Transition module represents the branch fusion module proposed in the embodiment of the present disclosure, DW-HRNet (full, original) represents the network structure containing 8 HRNet transition modules, and DW-HRNet (full, pooling) represents the network structure containing 8 branch fusion modules EFM proposed in the embodiment of the present disclosure.

[0075]

[0076] As can be seen from Table 2, compared with the reference fusion module with the same configuration, the branch fusion module provided by the embodiment of the present disclosure has lower computational complexity (-16.2%, -5.8%), higher human posture estimation index effect (+0.48, +0.2), and faster speed (34%, 25.4%).

[0077] In summary, the branch fusion module can improve the performance parameters of the network structure, and the number of branch fusion modules is positively correlated with the performance parameters of the network structure. That is, the more branch fusion modules there are, the higher the performance parameters of the network structure.

[0078] It should be noted that the number of convolution channels in the branch fusion module is a selectable searchable parameter that can directly participate in the network search process.

[0079] In the embodiments of the present disclosure, the feature fusion module is mostly used in the multi-scale network structure, and the multi-scale network architecture can be applied to dense image prediction tasks such as human posture recognition, semantic segmentation, and target detection. BiseNet, which is commonly used in semantic segmentation, also has a feature exchange module, and the feature fusion method provided by the embodiments of the present disclosure can be applied to the feature fusion process of BiseNet. In addition, the technical solution provided by the embodiments of the present disclosure can realize feature fusion on the server side, and perform network search of the multi-scale network structure based on the output feature map, and can also realize network search of the multi-scale network structure on the end side represented by the client, thereby improving universality and convenience. In addition, the required hardware cost is also reduced. The feature fusion method of the embodiment of the present disclosure can accelerate multi-scale feature fusion, and can achieve optimal configuration through network search, thereby improving the efficiency of network search.

[0080] The present disclosure provides an object processing device, referring to Figure 9 As shown in , the object processing device 900 may include:

[0081] A feature fusion module 901 is used to obtain an object to be processed and fuse feature maps of multiple resolutions of the object to be processed in the network structure according to a target feature map of the network structure to obtain a fused feature map;

[0082] The feature determination module 902 is used to transmit the fused feature map to the branches corresponding to the multiple resolutions to obtain an output feature map, and determine the target network structure based on the output feature map, so as to perform processing operations on the object to be processed through the target network structure.

[0083] In an exemplary embodiment of the present disclosure, the feature fusion module includes: a feature scaling module, configured to scale the reference feature maps, excluding the target feature map, in the feature maps of the multiple resolutions according to the attribute information of the target feature map, to obtain a plurality of scaled feature maps;

[0084] The feature conversion module is used to perform a connection operation on each scaled feature map to obtain a connection result, and perform a convolution operation on the connection result to convert the feature maps of the multiple resolutions into the fused feature map.

[0085] In an exemplary embodiment of the present disclosure, the feature scaling module includes: a pooling module for performing a pooling operation on the reference feature map according to the attribute information of the target feature map, so that the attribute information of each scaled feature map is the same as the attribute information of the target feature map.

[0086] In an exemplary embodiment of the present disclosure, the feature conversion module includes: a first convolution module, used to perform a first convolution operation on the connection result to obtain a first convolution result; a second convolution module, used to perform a second convolution operation on the first convolution result to obtain a second convolution result, and determine the second convolution result as the fusion feature map.

[0087] In an exemplary embodiment of the present disclosure, the feature determination module includes: a sampling module, configured to perform a convolution operation on the fused feature map and upsample the convolution result to obtain a sampling result; a branching module, configured to determine a branch feature map of each branch based on the feature map of each branch and the sampling result, and obtain the output feature map based on the branch feature map; the branch feature map of each branch has the same attribute parameters as the feature map corresponding to the branch;

[0088] In an exemplary embodiment of the present disclosure, the diversion module is configured to: fuse the feature map of each branch with the sampling result of each branch to obtain the branch feature map of each branch.

[0089] In an exemplary embodiment of the present disclosure, the device also includes: a hybrid fusion module, which is used to fuse the feature maps of a first range of the multiple resolutions through the target feature map, and to fuse the feature maps of a second range of the feature maps of the multiple resolutions through a reference manner; wherein the number of resolutions of the feature maps in the first range is greater than the number of resolutions of the feature maps in the second range.

[0090] It should be noted that the specific details of each module in the above-mentioned object processing device have been described in detail in the corresponding object processing method, and therefore will not be repeated here.

[0091] The exemplary embodiments of the present disclosure further provide an electronic device. The electronic device may be the aforementioned terminal 101 or server 102. Generally, the electronic device may include a processor and a memory, the memory being configured to store executable instructions of the processor, and the processor being configured to execute the aforementioned image denoising method by executing the executable instructions.

[0092] Below Figure 10 The structure of the electronic device is exemplarily described by taking the mobile terminal 1000 in FIG. 1 as an example. It should be understood by those skilled in the art that, in addition to the components specifically used for mobile purposes, Figure 10 The construction in can also be applied to fixed type equipment.

[0093] like Figure 10 As shown, the mobile terminal 1000 may specifically include: a processor 1001, a memory 1002, a bus 1003, a mobile communication module 1004, an antenna 1, a wireless communication module 1005, an antenna 2, a display screen 1006, a camera module 1007, an audio module 1008, a power module 1009 and a sensor module 1010.

[0094] Processor 1001 may include one or more processing units, for example, an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit). The image denoising method in this exemplary embodiment may be executed by an AP, a GPU, or a DSP. When the method involves neural network-related processing, it may be executed by an NPU. For example, the NPU may load neural network parameters and execute neural network-related algorithm instructions.

[0095] An encoder can encode (i.e., compress) an image or video to reduce the data size for easier storage or transmission. A decoder can decode (i.e., decompress) the encoded image or video data to restore the image or video data. Mobile terminal 1000 can support one or more encoders and decoders, such as image formats such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and BMP (Bitmap), and video formats such as MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).

[0096] The processor 1001 may be connected to the memory 1002 or other components via a bus 1003 .

[0097] Memory 1002 can be used to store computer-executable program code, which includes instructions. Processor 1001 executes various functional applications and data processing of mobile terminal 1000 by running the instructions stored in memory 1002. Memory 1002 can also store application data, such as images, videos, and other files.

[0098] The communication functions of mobile terminal 1000 are implemented through mobile communication module 1004, antenna 1, wireless communication module 1005, antenna 2, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1004 can provide 3G, 4G, and 5G mobile communication solutions for mobile terminal 1000. Wireless communication module 1005 can provide wireless communication solutions for mobile terminal 1000, such as wireless LAN, Bluetooth, and near-field communication.

[0099] The display screen 1006 is used to implement display functions, such as displaying a user interface, images, videos, etc. The camera module 1007 is used to implement shooting functions, such as shooting images, videos, etc. The audio module 1008 is used to implement audio functions, such as playing audio, collecting voice, etc. The power module 1009 is used to implement power management functions, such as charging the battery, powering the device, monitoring the battery status, etc. The sensor module 1010 may include one or more sensors for implementing corresponding sensing detection functions. For example, the sensor module 1010 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 1000 and output inertial sensing data.

[0100] It should be noted that a computer-readable storage medium is also provided in an embodiment of the present disclosure. The computer-readable storage medium may be included in the electronic device described in the above embodiment; or it may exist independently without being assembled into the electronic device.

[0101] Computer-readable storage media may be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0102] Computer-readable storage media can transmit, propagate, or transfer programs for use by or in conjunction with an instruction execution system, apparatus, or device. Program code contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wireline, optical cable, RF, or any suitable combination thereof.

[0103] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments.

[0104] Through the description of the above embodiments, it will be readily understood by those skilled in the art that the example embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes several instructions for enabling a computing device (such as a personal computer, server, terminal device, or network device) to execute the methods according to the embodiments of the present disclosure.

[0105] Furthermore, the figures above are merely illustrative of the processes included in the methods according to exemplary embodiments of the present disclosure and are not intended to be limiting. It is readily understood that the processes illustrated in the figures above do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0106] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0107] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing what is disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims. It should be understood that the present disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and that various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An object processing method, characterized in that: Applied to dense image prediction tasks, including: Obtaining an object to be processed, and fusing feature maps of multiple resolutions of the object to be processed in the network structure according to a target feature map of the network structure to obtain a fused feature map; the object to be processed includes an image to be processed, a voice to be processed, or a text to be processed; Transmitting the fused feature map to the branches corresponding to the multiple resolutions to obtain an output feature map, and determining a target network structure based on the output feature map, so as to perform a processing operation on the object to be processed through the target network structure; The step of transmitting the fused feature map to the branches corresponding to the multiple resolutions to obtain the output feature map includes: Transfer the fused feature map to multiple branches corresponding to different resolutions; For feature maps with multiple resolutions, the feature maps of a first range among the multiple resolutions are fused using the target feature map, and the feature maps of a second range among the feature maps of the multiple resolutions are fused using a reference method; the reference method is a method different from the fusion method based on the target feature map; The number of resolutions of the feature maps in the first range is greater than the number of resolutions of the feature maps in the second range.

2. The object processing method according to claim 1, characterized in that: The step of fusing feature maps of the object to be processed at multiple resolutions in the network structure according to the target feature map of the network structure to obtain a fused feature map includes: Scaling the reference feature maps other than the target feature map in the feature maps of the multiple resolutions according to the attribute information of the target feature map to obtain a plurality of scaled feature maps; A concatenation operation is performed on each scaled feature map to obtain a concatenation result, and a convolution operation is performed on the concatenation result to convert the feature maps of the multiple resolutions into the fused feature map.

3. The object processing method according to claim 2, characterized in that: Scaling the reference feature maps other than the target feature map in the feature maps of the multiple resolutions according to the attribute information of the target feature map to obtain a plurality of scaled feature maps includes: According to the attribute information of the target feature map, a pooling operation is performed on the reference feature map so that the attribute information of each scaled feature map is the same as the attribute information of the target feature map.

4. The object processing method according to claim 2, wherein: The performing a convolution operation on the connection result to convert the feature maps of the multiple resolutions into the fused feature map includes: Performing a first convolution operation on the connection result to obtain a first convolution result; Perform a second convolution operation on the first convolution result to obtain a second convolution result, and determine the second convolution result as the fusion feature map.

5. The object processing method according to claim 1, characterized in that: The step of transmitting the fused feature map to the branches corresponding to the multiple resolutions to obtain an output feature map includes: Performing a convolution operation on the fused feature map, and upsampling the convolution result to obtain a sampling result; The branch feature map of each branch is determined according to the feature map of each branch and the sampling result, and the output feature map is obtained according to the branch feature map; the branch feature map of each branch has the same attribute parameters as the feature map corresponding to the branch.

6. The object processing method according to claim 5, characterized in that: The determining of the branch feature graph of each branch according to the feature graph of each branch and the sampling result includes: The feature map of each branch is fused with the sampling result of each branch to obtain the branch feature map of each branch.

7. An object processing device, characterized in that: Applied to dense image prediction tasks, including: A feature fusion module is used to obtain an object to be processed and fuse the feature maps of multiple resolutions of the object to be processed in the network structure according to the target feature map of the network structure to obtain a fused feature map; the object to be processed includes an image to be processed, a voice to be processed, or a text to be processed; a feature determination module, configured to transmit the fused feature map to the branches corresponding to the multiple resolutions to obtain an output feature map, and determine a target network structure based on the output feature map, so as to perform a processing operation on the object to be processed through the target network structure; The step of transmitting the fused feature map to the branches corresponding to the multiple resolutions to obtain the output feature map includes: Transfer the fused feature map to multiple branches corresponding to different resolutions; For feature maps with multiple resolutions, the feature maps of a first range among the multiple resolutions are fused using the target feature map, and the feature maps of a second range among the feature maps of the multiple resolutions are fused using a reference method; the reference method is a method different from the fusion method based on the target feature map; The number of resolutions of the feature maps in the first range is greater than the number of resolutions of the feature maps in the second range.

8. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the object processing method according to any one of claims 1 to 6 by executing the executable instructions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the object processing method according to any one of claims 1 to 6 is implemented.