Multi-scale local enhanced lightweight visual conversion model design method

By introducing multi-layer patch embedding layer, hollow conversion layer and recognition network into the visual conversion model, the problems of low accuracy and information loss in the recognition of target objects in the prior art are solved, and more accurate target recognition results are achieved.

CN120182776APending Publication Date: 2025-06-20BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510062166.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing visual transformation model has poor accuracy in the recognition process of target object, and there is a problem of multi-scale information and local information loss.

Method used

A multi-scale locally enhanced lightweight visual transformation model is designed, and the extraction ability of multi-scale information and local information is enhanced through multi-layer patch embedding layer, first hollow transformation layer and identification network. Specifically, it includes a multi-head hollow attention layer and a local enhanced feedforward network layer, which are used to extract and enhance feature information.

Benefits of technology

By enhancing the ability of visual transformation model to extract multi-scale information and local information, the accuracy of target recognition results in target visual tasks is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182776A_ABST
    Figure CN120182776A_ABST
Patent Text Reader

Abstract

The invention provides a multi-scale local enhancement lightweight visual conversion model design method, and relates to the technical field of visual task processing, and the method comprises the steps: constructing a multi-scale local enhancement lightweight visual conversion model based on a multi-layer patch embedding layer, a first hole conversion layer and a recognition network; the multi-layer patch embedding layer is used for acquiring a target visual task and determining a plurality of sub-visual tasks corresponding to the target visual task based on the target visual task; the first hole conversion layer is used for determining an annotated visual task corresponding to the target visual task based on each sub-visual task; the first cavity conversion layer comprises a multi-head cavity attention layer and a local enhancement feedforward network layer; and the identification network is used for determining a target identification result corresponding to the target visual task based on the labeled visual task corresponding to the target visual task. According to the technical scheme, the extraction capability of the visual conversion model on multi-scale information and local information can be enhanced, so that the target recognition result in the target visual task is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of visual task processing, and in particular to a design method for a lightweight visual transformation model with multi-scale local enhancement. Background Art

[0002] A visual transformation model is a deep learning model based on the Transformer architecture. The visual transformation model can learn powerful visual expression capabilities from visual tasks and has shown great potential in aspects such as visual task classification and object detection.

[0003] However, after the embedding layer of the existing visual transformation model divides the visual task into several sub-visual tasks of a fixed size, it directly extracts features and classifies the sub-visual tasks. However, this will cause the visual transformation model to be unable to fully utilize the overall information and information of different scales of the visual task, and the accuracy of the model is relatively low. Secondly, the existing visual transformation model has relatively weak ability to extract local detail information. In summary, the existing visual transformation model has the phenomenon of loss of multi-scale information and local information, resulting in inaccurate recognition of visual tasks. Summary of the Invention

[0004] The present invention provides a design method for a lightweight visual transformation model with multi-scale local enhancement, which is used to solve the defect that the accuracy of the visual transformation model is poor and there is loss of multi-scale information and local information in the process of recognizing the target object in the prior art. The technical solution of the present invention can enhance the ability of the visual transformation model to extract multi-scale information and local information, making the target recognition result in the target visual task more accurate.

[0005] The present invention provides a design method for a lightweight visual transformation model with multi-scale local enhancement, including the following steps.

[0006] Construct a multi-scale local enhancement lightweight visual transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and a recognition network; The multi-layer patch embedding layer is used to obtain the target visual task and determine a plurality of sub-visual tasks corresponding to the target visual task based on the target visual task; The first dilated transformation layer is used to determine the labeled visual task corresponding to the target visual task based on each sub-visual task; the first dilated transformation layer includes a multi-head dilated attention layer and a local enhancement feed-forward network layer; The recognition network is used to determine the target recognition result corresponding to the target visual task based on the labeled visual task corresponding to the target visual task.

[0007] According to a design method for a lightweight visual transformation model with multi-scale local enhancement provided by the present invention, the multi-head dilated attention layer is used for: Determine the multi-head splicing feature corresponding to the target visual task based on all the sub-visual tasks; The local enhanced feed-forward network layer is used for: Determine the labeled visual task corresponding to the target visual task based on the multi-head splicing feature.

[0008] According to a method for designing a lightweight visual transformation model with multi-scale local enhancement provided by the present invention, the multi-head dilated attention layer determines the multi-head splicing feature corresponding to the target visual task through the following formula: Wherein, The sub-visual task corresponding sub-feature, represents multi-head dilated attention, represents multi-head attention, represents the query corresponding to the sub-visual task, represents a preset query projection matrix, represents the key corresponding to the sub-visual task, represents a preset key projection matrix, represents the value corresponding to the sub-visual task, represents a preset value projection matrix, represents the preset dilation rate corresponding to the sub-visual task, represents the dilated key dilated by , represents the dilated value dilated by , represents an activation function, represents a scale criterion for enhancing gradient stability, represents the multi-head splicing feature, represents splicing processing, represents the total number of all the sub-visual tasks, represents a preset output projection matrix.

[0009] According to a method for designing a lightweight visual transformation model with multi-scale local enhancement provided by the present invention, the local enhanced feed-forward network layer determines the labeled visual task corresponding to the target visual task through the following formula: Wherein, represents the labeled visual task, and All represent convolution processing of represents a pre-set non-linear activation function represents depth convolution processing

[0010] According to a method for designing a lightweight vision transformation model with multi-scale local enhancement provided by the present invention, the recognition network includes a first overlapping downsampler, a second dilated transformation layer, a second overlapping downsampler, a first general transformation layer, a third overlapping downsampler, and a second general transformation layer; The first overlapping downsampler is used to determine the first visual task feature corresponding to the labeled visual task based on the labeled visual task; The second dilated transformation layer is used to determine the first labeled visual task corresponding to the first visual task feature based on the first visual task; The second overlapping downsampler is used to determine the second visual task feature corresponding to the first labeled visual task based on the first labeled visual task; The first general transformation layer is used to determine the second labeled visual task corresponding to the second visual task feature based on the second visual task; The third overlapping downsampler is used to determine the third visual task feature corresponding to the second labeled visual task based on the second labeled visual task; The second general transformation layer is used to determine the target recognition result corresponding to the target visual task based on the third visual task feature.

[0011] According to a method for designing a lightweight vision transformation model with multi-scale local enhancement provided by the present invention, the multi-scale local enhancement lightweight vision transformation model is deployed on the corresponding edge computing device in the following manner: Under the conditions of model parameter quantity constraint and model floating-point operation number constraint, randomly sample multiple candidate subnet architectures from a pre-constructed search space; Iteratively optimize all the candidate subnet architectures until the iteration round reaches a preset iteration round, and determine the candidate subnet architecture with the highest performance in the last iteration round as the best subnet architecture; Deploy the multi-scale local enhancement lightweight vision transformation model on the edge computing device corresponding to the multi-scale local enhancement lightweight vision transformation model based on the best subnet architecture.

[0012] According to a method for designing a lightweight vision transformation model with multi-scale local enhancement provided by the present invention, the iterative optimization of all the candidate subnet architectures includes: In each iteration round, for each of the candidate subnet architectures, determine the candidate subnet architecture weights corresponding to the candidate subnet architecture weights based on the pre-trained supernet weights; based on the validation visual task dataset and the candidate subnet architecture weights, determine the subnet architecture performance corresponding to the candidate subnet architecture; determine the preset number of candidate subnet architectures with the top subnet architecture performance as the target candidate subnet architectures; perform mutation processing and crossover processing on all the target candidate subnet architectures respectively to obtain new candidate subnet architectures in the next iteration round.

[0013] The present invention also provides a lightweight visual transformation model design device with multi-scale local enhancement, including the following modules: A design and construction module for constructing a multi-scale local enhancement lightweight visual transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and an identification network; The multi-layer patch embedding layer is used to obtain a target visual task and determine a plurality of sub-visual tasks corresponding to the target visual task based on the target visual task; The first dilated transformation layer is used to determine the annotated visual task corresponding to the target visual task based on each of the sub-visual tasks; the first dilated transformation layer includes a multi-head dilated attention layer and a local enhancement feed-forward network layer; The identification network is used to determine the target recognition result corresponding to the target visual task based on the annotated visual task corresponding to the target visual task.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the multi-scale local enhancement lightweight visual transformation model design method as described in any one of the above.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the multi-scale local enhancement lightweight visual transformation model design method as described in any one of the above.

[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the multi-scale local enhancement lightweight visual transformation model design method as described in any one of the above.

[0017] The multi-scale local enhancement lightweight visual transformation model design method provided by the present invention constructs a multi-scale local enhancement lightweight visual transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and a recognition network; the multi-layer patch embedding layer is used to obtain a target visual task and determine a plurality of sub-visual tasks corresponding to the target visual task based on the target visual task; the first dilated transformation layer is used to determine an annotated visual task corresponding to the target visual task based on each sub-visual task; the first dilated transformation layer includes a multi-head dilated attention layer and a local enhancement feed-forward network layer; the recognition network is used to determine a target recognition result corresponding to the target visual task based on the annotated visual task corresponding to the target visual task. The multi-scale local enhancement lightweight visual transformation model of the technical solution of the present invention includes a multi-head dilated attention layer and a local enhancement feed-forward network layer, which can enhance the ability of the visual transformation model to extract multi-scale information and local information, so that the target recognition result in the target visual task is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of the multi-scale local enhancement lightweight visual transformation model design method provided by the present invention.

[0020] Figure 2 It is a structural diagram of the multi-scale local enhancement lightweight visual transformation model provided by the present invention.

[0021] Figure 3 It is a structural diagram of the feed-forward network layer provided by the present invention.

[0022] Figure 4 It is a comparison diagram of the target recognition results provided by the present invention.

[0023] Figure 5 It is a structural diagram of the multi-scale local enhancement lightweight visual transformation model design device provided by the present invention.

[0024] Figure 6 It is a structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0026] In view of the above problems in the prior art, the present invention provides a design method for a lightweight vision transformation model with multi-scale local enhancement. Figure 1 It is a schematic flowchart of the design method for the lightweight vision transformation model with multi-scale local enhancement provided by the present invention. As Figure 1 shown, the method includes the following step 110.

[0027] Step 110: Construct a multi-scale local enhancement lightweight vision transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and a recognition network; The multi-layer patch embedding layer is used to obtain a target vision task and determine a plurality of sub-vision tasks corresponding to the target vision task based on the target vision task; The first dilated transformation layer is used to determine an annotated vision task corresponding to the target vision task based on each sub-vision task; the first dilated transformation layer includes a multi-head dilated attention layer and a local enhancement feed-forward network layer; The recognition network is used to determine a target recognition result corresponding to the target vision task based on the annotated vision task corresponding to the target vision task.

[0028] Specifically, a multi-scale local enhancement lightweight vision transformation model can be constructed based on a multi-layer patch embedding layer, a first dilated transformation layer, and a recognition network. The multi-layer patch embedding layer can sample a target vision task into a plurality of sub-vision tasks, thereby determining a plurality of sub-vision tasks corresponding to the target vision task. Figure 2 It is a schematic structural diagram of the lightweight vision transformation model with multi-scale local enhancement provided by the present invention. As Figure 2 shown, Multi-Layer Patch Embedding represents the multi-layer patch embedding layer, and the multi-layer patch embedding layer includes a plurality of 3x3 convolutional layers (3x3 Conv), stride represents the stride corresponding to the convolutional layer, and Sampled Embedding in Multi-Layer Patch Embedding represents the output of the multi-layer patch embedding layer, that is, a plurality of sub-vision tasks.

[0029] Furthermore, as Figure 2 shown, Dilate Transformer Block Denote the first dilated transformation layer, which includes a multi-head dilated attention layer (MHDA) and a hybrid feedforward neural network layer (Hybrid FFN). The Layer Norm in the Dilate TransformerBlock represents normalization processing.

[0030] Furthermore, the recognition network can determine the target recognition result corresponding to the target visual task based on the labeled visual task corresponding to the target visual task.

[0031] The method for designing a multi-scale locally enhanced lightweight visual transformation model provided by the present invention constructs a multi-scale locally enhanced lightweight visual transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and a recognition network. The multi-layer patch embedding layer is used to obtain the target visual task and determine multiple sub-visual tasks corresponding to the target visual task based on the target visual task. The first dilated transformation layer is used to determine the labeled visual task corresponding to the target visual task based on each sub-visual task. The first dilated transformation layer includes a multi-head dilated attention layer and a hybrid feedforward neural network layer. The recognition network is used to determine the target recognition result corresponding to the target visual task based on the labeled visual task corresponding to the target visual task. The multi-scale locally enhanced lightweight visual transformation model of the technical solution of the present invention includes a multi-head dilated attention layer and a hybrid feedforward neural network layer, which can enhance the ability of the visual transformation model to extract multi-scale information and local information, thereby making the target recognition result in the target visual task more accurate.

[0032] In one embodiment, the multi-head dilated attention layer is used for: Determining the multi-head concatenated feature corresponding to the target visual task based on all the sub-visual tasks; The hybrid feedforward neural network layer is used for: Determining the labeled visual task corresponding to the target visual task based on the multi-head concatenated feature.

[0033] Specifically, the multi-head dilated attention layer is used to determine the multi-head concatenated feature corresponding to the target visual task based on all sub-visual tasks. Optionally, the multi-scale locally enhanced lightweight visual transformation model may further include a normalization layer, and the normalization layer may also perform normalization processing on the multi-head concatenated feature and input the normalized multi-head concatenated feature into the hybrid feedforward neural network layer in the first dilated transformation layer. The hybrid feedforward neural network layer determines the labeled visual task corresponding to the target visual task based on the multi-head concatenated feature.

[0034] In the above embodiments, the multi-head dilated attention layer and the local enhancement feed-forward network layer in the first dilated transformation layer are used to extract multi-scale information and local information, thus avoiding as much as possible the loss of multi-scale information and local information in the labeled visual task corresponding to the target visual task.

[0035] In one embodiment, the multi-head dilated attention layer determines the multi-head concatenated feature corresponding to the target visual task through the following formula: Wherein, The sub-visual task corresponding sub-feature, represents multi-head dilated attention, represents multi-head attention, represents the query corresponding to the sub-visual task, represents a preset query projection matrix, represents the key corresponding to the sub-visual task, represents a preset key projection matrix, represents the value corresponding to the sub-visual task, represents a preset value projection matrix, represents the preset dilation rate corresponding to the sub-visual task, represents the dilated key dilated by , represents the dilated value dilated by , represents an activation function, represents a scale criterion for enhancing gradient stability, represents the multi-head concatenated feature, represents concatenation processing, represents the total number of all the sub-visual tasks, represents a preset output projection matrix.

[0036] Specifically, , , , , , represents the set of real number matrices, represents the row of the matrix, represents the column of the matrix, Denote the set of positive integers. The preset dilation rate corresponding to each sub-visual task is different. The preset query projection matrix, preset key projection matrix, preset value projection matrix, preset output projection matrix, and preset dilation rate can all be preset according to needs. The embodiments of the present invention do not make specific limitations here.

[0037] It should be noted that in the prior art, the multi-head attention mechanism (Multi-Head Self-Attention, MHSA) is adopted. Compared with the prior art, in the present invention, the keys and values corresponding to the sub-visual tasks are dilated by the dilation rate, the receptive field of the visual transformation model is improved, and thus the extraction of multi-scale information in the visual task can be ensured.

[0038] In the above embodiment, the keys and values corresponding to the sub-visual tasks are dilated by the dilation rate, so that the multi-head concatenated features corresponding to the target visual task can better reflect the multi-scale information of the target visual task.

[0039] In one embodiment, the local enhanced feed-forward network layer determines the labeled visual task corresponding to the target visual task through the following formula: where, denotes the labeled visual task, and both denote the convolution processing of, denotes a preset non-linear activation function, denotes the depth convolution processing.

[0040] Specifically, Figure 3 is the structural schematic diagram of the feed-forward network layer provided by the present invention. The structure of the local enhanced feed-forward network layer provided by the present invention is as shown in Figure 3 (a). AvgPooling represents the pooling operation, Linear represents the linear layer, RELU is an example of the non-linear activation function , and Sigmoid is also an example of the non-linear activation function , Batch Norm represents Batch normalization, Stride represents the stride corresponding to the convolutional layer. The structure of the traditional feed-forward network layer is as shown in Figure 3 (b). The local enhanced feed-forward network layer provided by the present invention introduces the Squeeze-and-Excitation Networks (SEN). Among them, can be preset according to needs. The embodiments of the present invention do not make specific limitations here. For example the rectified linear activation function (RELU) can be adopted, The h-swish activation function can also be adopted. The Sigmoid activation function can also be adopted.

[0041] In the above embodiments, the local enhancement feed-forward network layer in the present invention can introduce different activation functions. This local enhancement feed-forward network layer can extract local information in visual tasks more precisely, enabling the labeled visual task corresponding to the target visual task to accurately represent local information.

[0042] In one embodiment, the recognition network includes a first overlapping downsampler, a second dilated transformation layer, a second overlapping downsampler, a first general transformation layer, a third overlapping downsampler, and a second general transformation layer; The first overlapping downsampler is used to determine the first visual task feature corresponding to the labeled visual task based on the labeled visual task; The second dilated transformation layer is used to determine the first labeled visual task corresponding to the first visual task feature based on the first visual task; The second overlapping downsampler is used to determine the second visual task feature corresponding to the first labeled visual task based on the first labeled visual task; The first general transformation layer is used to determine the second labeled visual task corresponding to the second visual task feature based on the second visual task; The third overlapping downsampler is used to determine the third visual task feature corresponding to the second labeled visual task based on the second labeled visual task; The second general transformation layer is used to determine the target recognition result corresponding to the target visual task based on the third visual task feature.

[0043] Specifically, as Figure 2 shown, the multi-scale local enhancement lightweight visual transformation model further includes an overlapping downsampler Overlapping Downsampler, a second dilated transformation layer Dilate Transformer Block , a first general transformation layer General Transformer Block and a second general transformation layer General Transformer Block , General Transformer Block represents the structure of the general transformation layer, and Locality FFN represents the traditional feed-forward network layer. It is easy to understand that the structure of the second dilated transformation layer is the same as that of the first dilated transformation layer, and the structure of the first general transformation layer is also the same as that of the second general transformation layer.

[0044] In the above embodiments, the multi-scale local enhancement lightweight vision transformation model can better capture visual task information by using an overlapping downsampler. Through the combination of a dilated transformation layer and a general transformation layer, it can more accurately extract multi-scale information and local information of visual tasks.

[0045] In one embodiment, the multi-scale local enhancement lightweight vision transformation model is deployed on the corresponding edge computing device in the following manner: Under the conditions of model parameter quantity constraints and model floating-point operation number constraints, randomly sample multiple candidate subnet architectures from a pre-constructed search space; Iteratively optimize all the candidate subnet architectures until the number of iteration rounds reaches a preset number of iteration rounds, and determine the candidate subnet architecture with the highest performance in the subnet architecture of the last iteration round as the optimal subnet architecture; Deploy the multi-scale local enhancement lightweight vision transformation model on the edge computing device corresponding to the multi-scale local enhancement lightweight vision transformation model based on the optimal subnet architecture.

[0046] Specifically, the pre-constructed search space can be encoded as a supernet architecture. A supernet architecture includes multiple subnet architectures, and a subnet architecture includes multiple candidate blocks. The number of layers of the subnet architecture is the same as that of the multi-scale local enhancement lightweight vision transformation model. The subnet architecture and its corresponding subnet weights can be represented by the following formula: Among them, represents the sampling block of the th layer, represents the weight corresponding to the sampling block , represents the number of layers of the subnet architecture. The process of sampling to determine the candidate subnet architecture, that is, determining the candidate blocks in the subnet as sampling blocks, can be represented by the following formula: Among them, represents the th candidate block of the th layer, represents the weight corresponding to the candidate block , represents the total number of candidate blocks in the th layer. Among them, the weights and corresponding to any two candidate blocks in the same layer need to meet the following conditions: Further, after sampling multiple candidate subnet architectures under the constraints of the number of model parameters and the number of floating-point operations of the model, all candidate subnet architectures can be iteratively optimized until the number of iteration rounds reaches a preset number of iteration rounds, and the candidate subnet architecture with the highest subnet architecture performance in the last iteration round is determined as the optimal subnet architecture. The preset number of iteration rounds can be set as needed, and the embodiments of the present invention do not make specific limitations here. The subnet architecture performance can be determined by the embedding vector dimension of the multi-scale local enhancement lightweight vision transformation model, the number of heads of the multi-head dilated attention layer, the number of model stacking layers, and the expansion ratio of the local enhancement feed-forward network layer under the corresponding subnet architecture. Furthermore, the multi-scale local enhancement lightweight vision transformation model can be deployed on the edge computing device corresponding to the multi-scale local enhancement lightweight vision transformation model based on the optimal subnet architecture.

[0047] In the above embodiment, by seeking the optimal subnet architecture to deploy the multi-scale local enhancement lightweight vision transformation model, the deployment of the multi-scale local enhancement lightweight vision transformation model is made more lightweight, and the computing power resources can be utilized to the greatest extent while ensuring the accuracy.

[0048] In one embodiment, the iteratively optimizing all the candidate subnet architectures includes: In each iteration round, for each of the candidate subnet architectures, determining the candidate subnet architecture weight corresponding to the candidate subnet architecture weight based on the pre-trained supernet weight; determining the subnet architecture performance corresponding to the candidate subnet architecture based on the validation vision task dataset and the candidate subnet architecture weight; determining the preset number of candidate subnet architectures with the top-ranked subnet architecture performance as the target candidate subnet architectures; and performing mutation processing and crossover processing on all the target candidate subnet architectures respectively to obtain the new candidate subnet architectures in the next iteration round.

[0049] Specifically, all candidate subnet architectures can be iteratively optimized by the following method: In each iteration round, for each candidate subnet architecture, the candidate subnet architecture weight corresponding to the candidate subnet architecture weight can be determined based on the pre-trained supernet weight Determining the candidate subnet architecture weight corresponding to the candidate subnet architecture weight . Further, the subnet architecture performance corresponding to the candidate subnet architecture can be determined based on the validation vision task dataset and the candidate subnet architecture weight . Furthermore, all candidate subnet architectures can be sorted according to the subnet architecture performance, and the preset number of candidate subnet architectures with the top-ranked subnet architecture performance can be determined as the target candidate subnet architectures , and the preset number can be set as needed. For example, the preset number can be set to Mutate and crossover all target candidate subnet architectures respectively. The mutation process can be represented by the following formula: where represents the subnet architecture obtained after the mutation process, represents the deep mutation probability, represents the mutation probability for each layer, represents the number of mutations, represents the number of candidate subnet architectures in the current iteration round, represents the model parameter quantity constraint, represents the model floating-point operation count constraint.

[0050] The crossover process can be represented by the following formula: where represents the subnet architecture obtained after the crossover process.

[0051] The new candidate subnet architecture in the next iteration round can be represented by the following formula: where represents the current iteration round.

[0052] In the above embodiments, the advantages and disadvantages of the subnet architecture are judged by the subnet architecture performance. The search scope for optimization is expanded through crossover and mutation, reducing the possibility of local convergence and making the finally determined optimal subnet architecture more reasonable.

[0053] Exemplarily, ablation experiments can be conducted on the method of the present invention. First, after replacing the traditional feedforward network layer with the locally enhanced feedforward network layer provided by the present invention, the accuracy of the visual transformation model is improved by 1.87. After introducing the dilated transformation layer and combining the dilated transformation layer with the general transformation layer, the accuracy of the visual transformation model is improved by 1.74. In the case of adopting the model architecture as Figure 2 shown, the multi-scale locally enhanced lightweight visual transformation model provided by the present invention has an accuracy improvement of 3.75 compared with the traditional visual transformation model.

[0054] In addition, the target recognition results determined by the present invention and the recognition results determined by the prior art can be visually compared. Figure 4 is a schematic diagram of the comparison of the target recognition results provided by the present invention. As shown in Figure 4As shown, the first-line visual task is the original image, the second-line visual task is the recognition result determined by the prior art, and the third-line visual task is the target recognition result determined by the present invention. Obviously, the present invention performs better in locating the target object, has more continuous recognition of the visual task target, and pays more comprehensive and complete attention to the semantic region, reflecting the strong recognition ability of the multi-scale local enhancement lightweight vision transformation model.

[0055] The multi-scale local enhancement lightweight vision transformation model design device provided by the present invention will be described below. The multi-scale local enhancement lightweight vision transformation model design device described below can be correspondingly referred to the multi-scale local enhancement lightweight vision transformation model design method described above.

[0056] Figure 5 is a schematic structural diagram of the multi-scale local enhancement lightweight vision transformation model design device provided by the present invention, as Figure 5 shown, the multi-scale local enhancement lightweight vision transformation model design device 500 includes the following modules: A design and construction module 510, configured to construct a multi-scale local enhancement lightweight vision transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and a recognition network; The multi-layer patch embedding layer is used to obtain a target visual task and determine a plurality of sub-visual tasks corresponding to the target visual task based on the target visual task; The first dilated transformation layer is used to determine an annotated visual task corresponding to the target visual task based on each of the sub-visual tasks; the first dilated transformation layer includes a multi-head dilated attention layer and a local enhancement feed-forward network layer; The recognition network is used to determine a target recognition result corresponding to the target visual task based on the annotated visual task corresponding to the target visual task.

[0057] In one embodiment, the multi-head dilated attention layer is used to: Determine a multi-head concatenated feature corresponding to the target visual task based on all the sub-visual tasks; The local enhancement feed-forward network layer is used to: Determine an annotated visual task corresponding to the target visual task based on the multi-head concatenated feature.

[0058] In one embodiment, the multi-head dilated attention layer determines the multi-head concatenated feature corresponding to the target visual task through the following formula: where the th The corresponding sub - feature, represents the multi - head dilated attention, represents the multi - head attention, represents the query corresponding to the sub - visual task, represents the preset query projection matrix, represents the key corresponding to the sub - visual task, represents the preset key projection matrix, represents the value corresponding to the sub - visual task, represents the preset dilation rate corresponding to the sub - visual task, represents the dilated key dilated by represents the dilated value dilated by represents the activation function, represents the scale criterion for enhancing gradient stability, represents the multi - head concatenated feature, represents the concatenation process, represents the total number of all the sub - visual tasks,

[0059] In one embodiment, the local enhancement feed - forward network layer determines the labeled visual task corresponding to the target visual task through the following formula: where, represents the labeled visual task, and both represent the convolution process of represents the preset non - linear activation function, represents the depth convolution process.

[0060] In one embodiment, the recognition network includes a first overlapping down - sampler, a second dilated transformation layer, a second overlapping down - sampler, a first general transformation layer, a third overlapping down - sampler, and a second general transformation layer; The first overlapping down - sampler is used to determine the first visual task feature corresponding to the labeled visual task based on the labeled visual task; The second dilated transformation layer is used to determine the first labeled visual task corresponding to the first visual task feature based on the first visual task; The second overlapping downsampler is configured to determine second visual task features corresponding to the first labeled visual task based on the first labeled visual task; The first general conversion layer is configured to determine a second labeled visual task corresponding to the second visual task features based on the second visual task; The third overlapping downsampler is configured to determine third visual task features corresponding to the second labeled visual task based on the second labeled visual task; The second general conversion layer is configured to determine a target recognition result corresponding to the target visual task based on the third visual task features.

[0061] In one embodiment, the multi-scale local enhancement lightweight vision transformation model design device 500 further includes a deployment module, and specifically, the deployment module is configured to: Under the conditions of model parameter quantity constraints and model floating-point operation number constraints, randomly sample multiple candidate subnet architectures from a pre-constructed search space; Iteratively optimize all the candidate subnet architectures until the iteration round reaches a preset iteration round, and determine the candidate subnet architecture with the highest subnet architecture performance in the last iteration round as the optimal subnet architecture; Deploy the multi-scale local enhancement lightweight vision transformation model based on the optimal subnet architecture on the edge computing device corresponding to the multi-scale local enhancement lightweight vision transformation model.

[0062] In one embodiment, the deployment module is further specifically configured to: In each iteration round, for each of the candidate subnet architectures, determine the candidate subnet architecture weight corresponding to the candidate subnet architecture weight based on the pre-trained supernet weight; based on the validation visual task dataset and the candidate subnet architecture weight, determine the subnet architecture performance corresponding to the candidate subnet architecture; determine a preset number of candidate subnet architectures with the top-ranked subnet architecture performance as target candidate subnet architectures; perform mutation processing and crossover processing on all the target candidate subnet architectures respectively to obtain new candidate subnet architectures in the next iteration round.

[0063] The multi-scale local enhancement lightweight visual transformation model design device provided by the present invention constructs a multi-scale local enhancement lightweight visual transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and a recognition network; the multi-layer patch embedding layer is used to obtain a target visual task and determine multiple sub-visual tasks corresponding to the target visual task based on the target visual task; the first dilated transformation layer is used to determine an annotated visual task corresponding to the target visual task based on each sub-visual task; the first dilated transformation layer includes a multi-head dilated attention layer and a local enhancement feed-forward network layer; the recognition network is used to determine a target recognition result corresponding to the target visual task based on the annotated visual task corresponding to the target visual task. The multi-scale local enhancement lightweight visual transformation model of the technical solution of the present invention includes a multi-head dilated attention layer and a local enhancement feed-forward network layer, which can enhance the ability of the visual transformation model to extract multi-scale information and local information, so that the target recognition result in the target visual task is more accurate.

[0064] Figure 6 An example of the physical structure diagram of an electronic device is shown as Figure 6 shown. The electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete communication with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the multi-scale local enhancement lightweight visual transformation model design method, and the method includes: Constructing a multi-scale local enhancement lightweight visual transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and a recognition network; The multi-layer patch embedding layer is used to obtain a target visual task and determine multiple sub-visual tasks corresponding to the target visual task based on the target visual task; The first dilated transformation layer is used to determine an annotated visual task corresponding to the target visual task based on each sub-visual task; the first dilated transformation layer includes a multi-head dilated attention layer and a local enhancement feed-forward network layer; The recognition network is used to determine a target recognition result corresponding to the target visual task based on the annotated visual task corresponding to the target visual task.

[0065] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0066] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the multi-scale local enhancement lightweight vision transformation model design method provided by the above-mentioned various methods. The method includes: Constructing a multi-scale local enhancement lightweight vision transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and an identification network; The multi-layer patch embedding layer is used to obtain a target vision task and determine multiple sub-vision tasks corresponding to the target vision task based on the target vision task; The first dilated transformation layer is used to determine an annotated vision task corresponding to the target vision task based on each of the sub-vision tasks; the first dilated transformation layer includes a multi-head dilated attention layer and a local enhancement feed-forward network layer; The identification network is used to determine a target recognition result corresponding to the target vision task based on the annotated vision task corresponding to the target vision task.

[0067] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the multi-scale local enhancement lightweight vision transformation model design method provided by the above-mentioned various methods. The method includes: Constructing a multi-scale local enhancement lightweight vision transformation model based on a multi-layer patch embedding layer, a first dilated transformation layer, and an identification network; The multi-layer patch embedding layer is used to obtain a target vision task and determine multiple sub-vision tasks corresponding to the target vision task based on the target vision task; The first cavity conversion layer is used to determine the labeled visual task corresponding to the target visual task based on each of the sub-visual tasks; the first cavity conversion layer includes a multi-head cavity attention layer and a local enhancement feed-forward network layer; The recognition network is used to determine the target recognition result corresponding to the target visual task based on the labeled visual task corresponding to the target visual task.

[0068] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.

[0069] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A lightweight visual conversion model design method with multi-scale local enhancement, characterized in that: include: A multi-scale local enhancement lightweight visual transformation model is constructed based on multi-layer patch embedding layers, the first hole transformation layer and the recognition network; The multi-layer patch embedding layer is used to obtain a target visual task, and determine a plurality of sub-visual tasks corresponding to the target visual task based on the target visual task; The first hole conversion layer is used to determine the labeled visual task corresponding to the target visual task based on each of the sub-visual tasks; The first hole conversion layer includes a multi-head hole attention layer and a local enhanced feedforward network layer; The recognition network is used to determine a target recognition result corresponding to the target visual task based on the labeled visual task corresponding to the target visual task.

2. The method for designing a multi-scale locally enhanced lightweight visual conversion model according to claim 1 is characterized in that: The multi-headed atrous attention layer is used to: Determine the multi-head stitching features corresponding to the target visual task based on all the sub-visual tasks; The local enhanced feed-forward network layer is used to: Determine a labeled visual task corresponding to the target visual task based on the multi-head stitching features.

3. The method for designing a multi-scale locally enhanced lightweight visual conversion model according to claim 2 is characterized in that: The multi-head void attention layer determines the multi-head splicing features corresponding to the target visual task through the following formula: in, No. Sub-vision tasks The corresponding sub-features, Indicates the long empty attention, Indicates multiple attentions, Indicates The query corresponding to each sub-vision task is represents the preset query projection matrix, Indicates The keys corresponding to the sub-visual tasks are: represents the preset key projection matrix, Indicates The value corresponding to each sub-vision task is Represents the preset projection matrix, Indicates The preset void rate corresponding to each sub-vision task, Indicates that Expanded expansion key, Indicates that The expansion value of the expansion, represents the activation function, represents the scaling criterion used to enhance gradient stability, Indicates the multi-head splicing feature, Indicates splicing processing, represents the total number of all the sub-vision tasks, Represents the preset output projection matrix.

4. The method for designing a multi-scale locally enhanced lightweight visual conversion model according to claim 2, characterized in that: The local enhanced feed-forward network layer determines the labeled visual task corresponding to the target visual task by the following formula: in, represents the annotation vision task, and Both said The convolution process, represents the pre-set nonlinear activation function, Represents deep convolution processing.

5. The method for designing a multi-scale locally enhanced lightweight visual conversion model according to any one of claims 1 to 4, characterized in that: The recognition network includes a first overlapping downsampler, a second hole conversion layer, a second overlapping downsampler, a first universal conversion layer, a third overlapping downsampler and a second universal conversion layer; The first overlapping downsampler is used to determine a first visual task feature corresponding to the labeled visual task based on the labeled visual task; The second hole conversion layer is used to determine a first labeled visual task corresponding to the first visual task feature based on the first visual task; The second overlapping downsampler is used to determine a second visual task feature corresponding to the first labeled visual task based on the first labeled visual task; The first universal conversion layer is used to determine a second annotated visual task corresponding to the second visual task feature based on the second visual task; The third overlapping downsampler is used to determine a third visual task feature corresponding to the second labeled visual task based on the second labeled visual task; The second universal conversion layer is used to determine the target recognition result corresponding to the target visual task based on the third visual task feature.

6. The method for designing a lightweight visual conversion model with multi-scale local enhancement according to any one of claims 1 to 4, characterized in that: The multi-scale local enhanced lightweight visual conversion model is deployed on the corresponding edge computing device in the following way: Under the constraints of model parameter quantity and model floating point operation number, multiple candidate subnetwork architectures are randomly sampled from the pre-constructed search space; Iteratively optimize all the candidate subnet architectures until the number of iterations reaches a preset number of iterations, and determine the candidate subnet architecture with the highest subnet architecture performance in the last iteration as the optimal subnet architecture; Based on the optimal subnet architecture, the multi-scale local enhanced lightweight visual conversion model is deployed on an edge computing device corresponding to the multi-scale local enhanced lightweight visual conversion model.

7. The method for designing a multi-scale locally enhanced lightweight visual conversion model according to claim 6, characterized in that: The iterative optimization of all the candidate subnet architectures includes: In each iteration round, for each of the candidate subnet architectures, the candidate subnet architecture weight corresponding to the candidate subnet architecture weight is determined based on the pre-trained supernet weight; based on the verification visual task data set and the candidate subnet architecture weight, the subnet architecture performance corresponding to the candidate subnet architecture is determined; a preset number of candidate subnet architectures with top subnet architecture performance are determined as target candidate subnet architectures; all the target candidate subnet architectures are mutated and cross-processed respectively to obtain new candidate subnet architectures in the next iteration round.

8. A multi-scale local enhanced lightweight visual conversion model design device, characterized in that: include: Design a building block for constructing a multi-scale locally enhanced lightweight visual transformation model based on a multi-layer patch embedding layer, a first hole transformation layer, and a recognition network; The multi-layer patch embedding layer is used to obtain a target visual task, and determine a plurality of sub-visual tasks corresponding to the target visual task based on the target visual task; The first hole conversion layer is used to determine the labeled visual task corresponding to the target visual task based on each of the sub-visual tasks; The first hole conversion layer includes a multi-head hole attention layer and a local enhanced feedforward network layer; The recognition network is used to determine a target recognition result corresponding to the target visual task based on the labeled visual task corresponding to the target visual task.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the method for designing a lightweight visual conversion model with multi-scale local enhancement as described in any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for designing a lightweight visual conversion model with multi-scale local enhancement as described in any one of claims 1 to 7 is implemented.