Target detection method based on multi-scale attention and network architecture search

By using multi-scale attention and network architecture search (NAS) technology, the feature extraction strategy is dynamically adjusted to solve the occlusion and blur problems of organ detection in medical images, improve the detection accuracy and robustness, and is suitable for CT or MRI images.

CN120747115AActive Publication Date: 2025-10-03HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511263029.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-10-03
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing technologies face the problem of incomplete target feature extraction when detecting organs in medical images, especially in CT or MRI images where occlusion and blurring lead to decreased detection performance.

Method used

A method based on multi-scale attention and network architecture search (NAS) is adopted. By introducing a deep guidance module and a multi-scale context fusion attention mechanism, the feature extraction strategy is dynamically adjusted, the network structure is optimized, and the target is accurately located.

Benefits of technology

It significantly improves the accuracy and robustness of target detection in complex scenarios, reduces computational costs, has strong adaptability, and is suitable for medical environments with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747115A_ABST
    Figure CN120747115A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method based on multi-scale attention and network architecture search, which can dynamically adjust a feature extraction strategy, optimize the network architecture design and improve the target detection efficiency by introducing a deep guidance adaptive feature extraction and multi-scale context fusion attention mechanism. Therefore, the accuracy and robustness of target detection are remarkably improved in a complex scene. Through a multi-scale situation fusion attention mechanism, a spatial relationship between image areas can be modeled on global and local levels, the ability of the model to understand a complex scene is enhanced, and the precision of target detection is further improved. By using the neural network architecture search technology, the optimal network architecture can be automatically explored, the limitation of the traditional manual design of the network architecture is avoided, the NAS framework dynamically optimizes the feature extraction strategy, the calculation cost is reduced to the greatest extent while the detection performance is improved, the method is more efficient and practical in practical application, and the method is suitable for popularization and application. And the adaptability and performance of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and relates to a method for target detection based on multi-scale attention and network architecture search. Background Art

[0002] Object detection methods based on network architecture search (NAS) aim to improve the accuracy and robustness of organ detection in medical images by automatically designing optimal neural network architectures. With the widespread application of deep learning in medical image analysis, traditional manually designed network architectures, while achieving some success, still face challenges in complex scenarios (such as blurred organ boundaries, lesion interference, or partial occlusion). In particular, in CT or MRI images, organs may be partially occluded by surrounding tissue, lesions, or other organs, resulting in incomplete target feature extraction and thus affecting detection performance. Therefore, utilizing network architecture search techniques to automatically explore network architectures with strong adaptability and excellent feature extraction capabilities is crucial for addressing occlusion issues and improving object detection accuracy. By introducing a depth-guided attention mechanism and a multi-scale feature fusion strategy, NAS can dynamically adjust the network structure and enhance the model's ability to learn features from occluded regions, thereby achieving more accurate target localization and recognition in complex medical images.

[0003] Since the efficient architecture can accurately detect and locate targets from a single CT or MRI image, this method provides a flexible and cost-effective solution for practice. It uses automated feature extraction and optimized network design to quickly identify target areas in complex images, reduce manual intervention, and adapt to resource-limited medical environments, achieving high-precision detection without expensive equipment. However, target detection still faces two major challenges: first, the target in a single image may be blurred due to imaging quality, noise, or occlusion, making it difficult for traditional methods to accurately extract features; second, the target feature requirements at different depth ranges are different, and traditional methods have difficulty taking into account the differences between nearby organs and distant lesions, affecting the detection effect.

[0004] To solve the above problems, we explore advanced deep learning and multi-scale feature fusion technologies to enhance the model's ability to recognize blurred targets and optimize the network architecture to improve detection accuracy and robustness, providing more reliable support for practice. Summary of the Invention

[0005] Based on this, the present invention aims to provide a method based on multi-scale contextual fusion attention and Neural Architecture Search (NAS) to address the existing problem of degraded detection performance caused by target blur in a single image and varying feature requirements at different depth ranges. By introducing a depth-guided module and a multi-scale contextual fusion attention mechanism, it is possible to dynamically adjust feature extraction strategies, optimize network structure, and accurately locate targets.

[0006] In a first aspect, an embodiment of the present application provides a target detection method based on multi-scale attention and network architecture search, comprising the following steps: S1: Obtain an input human body image, perform feature extraction on the human body image through the backbone network, and obtain a multi-scale feature map and visual features and depth features of the multi-scale feature map; S2: Inputting the visual features and the depth features into a multi-scale context fusion encoder to perform multi-scale context fusion to obtain a multi-scale context fusion result; S3: Processing the multi-scale context fusion result through a depth prediction module to generate a panoramic depth map; S4: Dividing the generated panoramic depth map into multiple depth ranges, performing feature extraction and search on the multiple depth ranges through multiple network search architectures to obtain multiple target search results; S5: Calculating prediction losses corresponding to the multiple target search results based on the MLP detection method; S6: Filter out the target search results searched by the network search architecture corresponding to the minimum prediction loss, and output them as the final target.

[0007] Preferably, the human body image is derived from a public dataset for endoscopic surgery visual analysis; the backbone network includes a ResNet-50 model and a lightweight depth predictor.

[0008] Preferably, step S1 includes: S11: performing feature extraction processing on the human body image through a ResNet-50 model to obtain a multi-scale feature map; S12: Filtering out a feature map with the highest level semantics in the multi-scale feature map as the visual feature; S13: unify the downsampling rates of the multi-scale feature maps by bilinear pooling, fuse the unified multi-scale feature maps by element-wise addition, and finally process the fused multi-scale feature maps by a convolutional layer to obtain the depth features.

[0009] Preferably, step S2 includes: S21: Divide and mark the human body image into blocks; S22: Modeling the global dependency relationships of the divided and labeled tiles using a Transformer-based multi-layer self-attention module to obtain the global dependency relationships of the human body image; S23: Based on the multi-scale context fusion encoder, feature aggregation is performed on the divided and marked blocks and their surrounding blocks to obtain a local context fusion result; The global dependency relationship and the local context fusion result constitute the multi-scale context fusion result.

[0010] Preferably, the loss function for calculating the loss in step S5 is: (1); in, is the loss function, is the focal loss of the target corresponding panoramic depth map in the target search results, is the number of target search results, The attribute information of the panoramic depth map corresponding to the target in the target search result, Search architecture for the web.

[0011] Preferably, It is expressed by the following formula: (2); in, Search architecture for the web, 、 is the balance coefficient, Searching for a network architecture The average precision of the corresponding target is Searching for a network architecture The maximum performance index corresponding to the target is: Searching for a network architecture The intersection-over-union ratio of the corresponding target is given below.

[0012] Compared with the prior art, the present invention has the following beneficial effects: By introducing deep guided adaptive feature extraction (DGAFE) and multi-scale contextual fusion attention mechanism, the present invention can dynamically adjust the feature extraction strategy and optimize the network architecture design, thereby significantly improving the accuracy and robustness of target detection in complex scenarios (such as target blur, occlusion, etc.); through the multi-scale contextual fusion attention mechanism, the present invention can model the spatial relationship between image regions at both the global and local levels, enhance the model's ability to understand complex scenes, and further improve the accuracy of target detection; using neural network architecture search (NAS) technology, the present invention can automatically explore the optimal network architecture, avoiding the limitations of traditional manually designed network architecture. The NAS framework dynamically optimizes the feature extraction strategy, while improving detection performance, it minimizes computational costs, making the method more efficient and practical in practical applications, and improving the adaptability and performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] A more complete understanding of the exemplary embodiments of the present invention can be obtained by referring to the following drawings. The drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present invention and do not constitute a limitation of the present invention. In the drawings, the same reference numerals generally represent the same components or steps.

[0014] Figure 1 A flowchart of a target detection method based on multi-scale attention and network architecture search provided in an embodiment of the present application. DETAILED DESCRIPTION

[0015] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0016] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0017] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0018] Reference Figure 1 This embodiment discloses a target detection method based on multi-scale attention and network architecture search, comprising the following steps: S1: Obtain an input human body image, perform feature extraction on the human body image through the backbone network, and obtain a multi-scale feature map and visual features and depth features of the multi-scale feature map; Specifically, this example uses the dataset Endovis2018, a public dataset for visual analysis of endoscopic surgery, which is mainly used to promote research in the field of computer-assisted surgery. It contains 8 surgical sequences, each sequence contains multiple frames of images, and the resolution of each frame is 1920x1080 or 1280x1024. Taking the image as input, our framework uses feature backbones such as ResNet50 and lightweight depth predictors to generate its visual and depth features respectively. Given an image , where H and W represent its height and width, and its multi-scale feature maps are obtained from the last three stages of ResNet-50 The downsampling ratio to the original size is The highest level that will have sufficient semantics Considered as visual features of the input image , The number of channels of the image input layer.

[0019] Deep features are obtained from the image through a lightweight depth predictor. First, the sizes of the three-level features are unified to the same 1 / 16 downsampling rate through bilinear pooling and fused together through element-wise addition. In this way, multi-scale visual appearances can be integrated and fine-grained patterns can be preserved for small-sized objects. Then, two 3×3 convolutional layers are applied to obtain the deep features of the input image. , is the number of channels of the input image.

[0020] S2: Inputting the visual features and the depth features into a multi-scale context fusion encoder to perform multi-scale context fusion to obtain a multi-scale context fusion result; Step S2 includes: S21: Divide and mark the human body image into blocks; S22: Modeling the global dependency relationships of the divided and labeled tiles using a Transformer-based multi-layer self-attention module to obtain the global dependency relationships of the human body image; S23: Based on the multi-scale context fusion encoder, feature aggregation is performed on the divided and marked blocks and their surrounding blocks to obtain a local context fusion result; The global dependency relationship and the local context fusion result constitute the multi-scale context fusion result.

[0021] The highest-level semantics refers to the most abstract and global semantic information captured by the last or deepest layer during feature extraction. The multi-scale feature maps of ResNet-50 are obtained from the last three stages. The input visual and depth features are passed through a multi-scale contextual fusion encoder for global dependency modeling and local optimization. While depth-guided feature extraction focuses on depth-specific adaptation, the multi-scale context-fused attention mechanism enhances the spatial relationships between image regions. This mechanism operates in two stages: global dependency modeling and local optimization.

[0022] (1) Global dependency modeling First, divide and label the image blocks: divide the input image I into P×P blocks, and each block is used as an independent label. is the label of block i. The sequence of blocks is represented as , these blocks are passed through a Transformer-based multi-layer self-attention module to model global dependencies, which is expressed as follows: (3); Among them, A() represents the self-attention operation based on the autotransformer, which adaptively captures the long-distance dependencies across token sequences. By utilizing the adaptive attention mechanism of the autotransformer, each block It effectively integrates global context while maintaining computational efficiency. Adaptively capturing long-range dependencies (a global attention mechanism combined with data-driven weight learning) allows the model to dynamically learn and model complex relationships between any two elements (such as pixels, objects, and words) in the input sequence, regardless of their distance in the sequence. This capability is one of the core advantages of the Transformer and is particularly suitable for tasks requiring global context understanding.

[0023] (2) Local context integration For local refinement, multi-scale context fusion attention is taken from the center patch of an image and its surrounding patches. Aggregate features, where representations are centered around a central block The set of adjacent blocks is expressed as: (4); in, Indicates that the image I is divided into a set of P×P blocks sorted in sequence; Represents the weight, which is calculated by the following formula: (5); Variance-based weight allocation: In order to ensure the robustness of feature fusion, variance-based weight allocation is introduced. The characteristic variance of is: (6); Among them, the variance Smaller patches are given higher weights during the fusion process, reflecting their consistency within the local context. This approach integrates a depth-guided module for adaptive feature extraction and a multi-scale context-fused attention mechanism for spatial relationship modeling. By combining depth perception with sophisticated attention-based feature fusion, the framework achieves robust performance in object detection.

[0024] S3: Process the multi-scale context fusion results through the depth prediction module to generate a panoramic depth map; Specifically, the depth distribution of foreground objects is predicted rather than a complete dense depth map, thereby reducing computational overhead and improving detection accuracy. Based on the visual and depth features generated in step S1, these features are then globally and locally encoded in step S2 to generate new visual and depth features. These are then input into the depth-guided decoder in step S3-S4 to perform a network architecture search on the features and output the best result.

[0025] S4: Divide the generated panoramic depth map into multiple depth ranges, perform feature extraction and search on the multiple depth ranges through multiple network search architectures, and obtain multiple target search results.

[0026] Specifically, accurate object detection in space requires precise depth information to distinguish objects at different distances. The depth-guided module addresses this challenge by dividing the input depth map (the depth features output from the encoder, which is the same as the original input image) into different ranges and applying customized feature extraction strategies in each range. These strategies are dynamically optimized through the NAS framework.

[0027] Depth prediction: The depth guidance module first uses the depth prediction module (the depth prediction module is a lightweight CNN module that predicts the discrete depth interval of the object, provides geometric priors for the Transformer, and outputs depth features (for the encoder) and foreground depth map) to generate a panoramic depth map D. Formally, given an input image , panoramic depth map is calculated as: ; in Represents the depth prediction model.

[0028] Depth range will Divided into K predefined depth ranges , the threshold is . The depth region after segmentation Defined as: ; in, is the pixel index of the image. Each range corresponds to an object at a different distance (such as near, medium, and far).

[0029] Personalized feature extraction based on NAS: For each segmentation depth range , using the NAS framework to search for the optimal feature extraction strategy. The NAS search space S contains configurations of transformer layers and attention mechanisms of different depths: ; in, Represents a candidate architecture (e.g., a shallow transformer or a deep transformer with multi-head self-attention).

[0030] The search goal is to maximize the performance index , mean average precision and intersection and comparison , while minimizing the computational cost. The optimization problem is stated as: (2); in, Indicates the architecture selected by NAS, is the balance coefficient.

[0031] Specifically, maximize the performance index Indicates the comprehensive performance of the model optimized by Neural Architecture Search (NAS) in the target detection task. It represents the average precision (AP) of all categories at different recall rates, and measures the comprehensive accuracy of the detector at different difficulty levels. It represents the ratio of the intersection and union of the predicted bounding box and the real box, which measures the positioning accuracy. By dynamically adjusting the depth-related architecture parameters θ and fusion strategy, we can directly optimize and , while constraining the computational cost (i.e. The depth guidance and multi-scale attention mechanism are key innovations, which respectively target the modeling of depth ambiguity and spatial relationships and jointly improve detection performance.

[0032] S5: Calculating prediction losses corresponding to the multiple target search results based on the MLP detection method; S6: Filter out the target search results searched by the network search architecture corresponding to the minimum prediction loss, and output them as the final target.

[0033] In order to correctly match each query with the ground-truth object, the loss of each query-label (query-Label represents the matching relationship between learnable object queries (Object Queries) and real labels (Ground Truth Labels), its core role is to solve the correspondence problem of "unordered prediction" and "ordered labeling", and realize true end-to-end detection without manual design of Anchor or NMS post-processing. It is the association mechanism between prediction results and real labels (labels), and its core role is to optimize the model's positioning and classification of objects through the attention mechanism) is calculated, and the Hungarian algorithm is used to find the global best match. The Hungarian algorithm is used to perform optimal one-to-one matching between learnable object queries (Object Queries) and real labels (Ground Truth Labels), ensuring that each real object is predicted by only one Query to avoid redundancy. After matching, Get Valid pairs, where represents the number of true objects. Then, the overall loss for training images is formulated as: (1); in, is the loss function, is the focal loss of the target corresponding panoramic depth map in the target search results, is the number of target search results, The attribute information of the panoramic depth map corresponding to the target in the target search result, Search architecture for the web.

[0034] Specifically, The focus loss of the panoramic depth map corresponding to the target in the target search results is actually obtained based on the set of all panoramic depth images generated in the target search results. Specifically, it focuses on training a sparse set of difficult samples (composed of all panoramic depth maps generated in the target search). By directly introducing two penalty factors based on the standard cross entropy loss model (this is based on existing technology and will not be described here), the weight of easy-to-classify samples is reduced, making the model more focused on difficult samples during training. As a result, the model converges faster. When the model converges, the focus loss of the image set can be obtained. Specifically, It is expressed by the following formula: ; in, is the probability of a positive sample in the sample set, is the negative sample probability of the sample set; where, For the control parameters, and 1- They are used to control the ratio of positive / negative samples, and their value range is [0, 1]. The value of can generally be selected through cross-validation; is the focus parameter, which ranges from [0, +∞). It reduces the weight of easy-to-classify samples, so that the model can focus more on difficult samples during training. When it is 0, is the cross entropy loss. The sample set consists of all the corresponding panoramic depth maps of the target in the target search results.

[0035] Compared with the prior art, the present invention has the following beneficial effects: By introducing deep guided adaptive feature extraction (DGAFE) and multi-scale contextual fusion attention mechanism, the present invention can dynamically adjust the feature extraction strategy and optimize the network architecture design, thereby significantly improving the accuracy and robustness of target detection in complex scenarios (such as target blur, occlusion, etc.); through the multi-scale contextual fusion attention mechanism, the present invention can model the spatial relationship between image regions at both the global and local levels, enhance the model's ability to understand complex scenes, and further improve the accuracy of target detection; using neural architecture search (NAS) technology, the present invention can automatically explore the optimal network architecture, avoiding the limitations of traditional manually designed network architecture. The NAS framework dynamically optimizes the feature extraction strategy, while improving detection performance, it minimizes the computational cost, making the method more efficient and practical in practical applications, and improving the adaptability and performance of the model.

[0036] It should be noted that the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the module, program segment or part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0037] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0038] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0039] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0040] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0041] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0042] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and they should all be included in the scope of the claims and description of the present application.

Claims

1. A target detection method based on multi-scale attention and network architecture search, characterized in that: The steps include: S1: Obtain an input human body image, perform feature extraction on the human body image through the backbone network, and obtain a multi-scale feature map and visual features and depth features of the multi-scale feature map; S2: Inputting the visual features and the depth features into a multi-scale context fusion encoder to perform multi-scale context fusion to obtain a multi-scale context fusion result; S3: Processing the multi-scale context fusion result through a depth prediction module to generate a panoramic depth map; S4: Dividing the generated panoramic depth map into multiple depth ranges, performing feature extraction and search on the multiple depth ranges through multiple network search architectures to obtain multiple target search results; S5: Calculating prediction losses corresponding to the multiple target search results based on the MLP detection method; S6: Filter out the target search results searched by the network search architecture corresponding to the minimum prediction loss, and output them as the final target.

2. The method according to claim 1, characterized in that The human body images are derived from a public dataset for endoscopic surgery visual analysis; the backbone network includes a ResNet-50 model and a lightweight depth predictor.

3. The method according to claim 2, characterized in that Step S1 includes: S11: performing feature extraction processing on the human body image through a ResNet-50 model to obtain a multi-scale feature map; S12: Filtering out a feature map with the highest level semantics in the multi-scale feature map as the visual feature; S13: unify the downsampling rates of the multi-scale feature maps by bilinear pooling, fuse the unified multi-scale feature maps by element-wise addition, and finally process the fused multi-scale feature maps by a convolutional layer to obtain the depth features.

4. The method according to claim 3, characterized in that Step S2 includes: S21: Divide and mark the human body image into blocks; S22: Modeling the global dependency relationships of the divided and labeled tiles using a Transformer-based multi-layer self-attention module to obtain the global dependency relationships of the human body image; S23: Based on the multi-scale context fusion encoder, feature aggregation is performed on the divided and marked blocks and their surrounding blocks to obtain a local context fusion result; The global dependency relationship and the local context fusion result constitute the multi-scale context fusion result.

5. The method according to claim 4, characterized in that The loss function for calculating the loss in step S5 is: (1); in, is the loss function, is the focal loss of the target corresponding panoramic depth map in the target search results, is the number of target search results, The attribute information of the panoramic depth map corresponding to the target in the target search result, Search architecture for the web.

6. The method according to claim 5, characterized in that It is expressed by the following formula: (2); in, Search architecture for the web, 、 is the balance coefficient, Searching for a network architecture The average precision of the corresponding target is Searching for a network architecture The maximum performance index corresponding to the target is: Searching for a network architecture The intersection-over-union ratio of the corresponding target is given below.

Citation Information

Patent Citations

  • Target tracking method and system based on multi-scale aggregation attention feature extraction network

    CN118429389A

  • Multi-target deep neural network architecture search method oriented to hybrid CNN-Transform architecture

    CN120562484A