Task execution method and device, electronic equipment and storage medium

By executing multiple convolution subtasks, mask extraction subtasks and mask processing subtasks in image processing, the Transformer model has solved the problem of high computational complexity when processing complex image structures, and achieved efficient dark-light image processing and resource conservation.

CN120086021APending Publication Date: 2025-06-03KUNWANG (SHANGHAI) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213526.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing Transformer model has high computational complexity when processing complex image structures, resulting in low efficiency and high hardware resource consumption, making it difficult to effectively deal with noise, detail loss and color distortion problems in dark light images.

Method used

By performing multiple convolution subtasks, mask extraction subtasks and mask processing subtasks, the processor is used to obtain rich feature information, and local features are enhanced through the mask mechanism to reduce the computational complexity.

Benefits of technology

It effectively reduces hardware resource overhead, improves the efficiency and effect of image processing, and can handle details and noise in dark light images more finely.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086021A_ABST
    Figure CN120086021A_ABST
Patent Text Reader

Abstract

The invention provides a task execution method, and relates to the technical field of artificial intelligence, in particular to the technical field of chips, deep learning, computer vision and natural language processing. According to the specific implementation scheme, a processor is used for executing multiple convolution sub-tasks of a target task according to a to-be-processed feature to obtain a first convolution feature and a second convolution feature, and the second convolution feature comprises multiple convolution sub-features; according to a plurality of convolution sub-features of the second convolution feature and at least one neighborhood sub-feature of each of the plurality of convolution sub-features, executing a mask extraction sub-task of the target task by using a processor to obtain a mask feature; according to the first convolution feature and the mask feature, utilizing a processor to execute a mask processing subtask of the target task to obtain a first feature subjected to mask processing; an output feature of the target task is determined using a processor according to the first masked feature. The invention further provides a task execution device, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies, and particularly to the fields of chips, deep learning, computer vision, and natural language processing technologies, and can be applied to the scenario of low-light image enhancement. More specifically, the present disclosure provides a task execution method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of artificial intelligence technologies, the application of attention mechanisms is increasing continuously. Among various attention mechanisms, the multi-head self-attention mechanism of the Transformer model has strong adaptive capabilities and strong global modeling capabilities. Summary of the Invention

[0003] The present disclosure provides a task execution method, apparatus, device, and storage medium.

[0004] According to one aspect of the present disclosure, there is provided a task execution method, which includes: according to the feature to be processed, using a processor to execute multiple convolutional subtasks of a target task to obtain multiple convolutional features, where the multiple convolutional features include a first convolutional feature and a second convolutional feature, and the second convolutional feature includes multiple convolutional sub-features; according to the multiple convolutional sub-features of the second convolutional feature and at least one neighborhood sub-feature of each of the multiple convolutional sub-features, using a processor to execute a mask extraction subtask of the target task to obtain a mask feature; according to the first convolutional feature and the mask feature, using a processor to execute a mask processing subtask of the target task to obtain a first mask-processed feature; and according to the first mask-processed feature, using a processor to determine an output feature of the target task.

[0005] According to another aspect of the present disclosure, there is provided a task execution apparatus, which includes: a first execution module, configured to use a processor to execute multiple convolutional subtasks of a target task according to the feature to be processed to obtain multiple convolutional features, where the multiple convolutional features include a first convolutional feature and a second convolutional feature, and the second convolutional feature includes multiple convolutional sub-features; a second execution module, configured to use a processor to execute a mask extraction subtask of the target task according to the multiple convolutional sub-features of the second convolutional feature and at least one neighborhood sub-feature of each of the multiple convolutional sub-features to obtain a mask feature; a third execution module, configured to use a processor to execute a mask processing subtask of the target task according to the first convolutional feature and the mask feature to obtain a first mask-processed feature; and a determination module, configured to use a processor to determine an output feature of the target task according to the first mask-processed feature.

[0006] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided according to the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method provided according to the present disclosure.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0011] Figure 1 is a flowchart of a task execution method according to an embodiment of the present disclosure;

[0012] Figure 2A is a schematic diagram of a target processing module according to an embodiment of the present disclosure;

[0013] Figure 2B is a schematic diagram of a convolutional sub-feature and a neighborhood sub-feature according to an embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram of a target network according to an embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram of a target sub-model according to an embodiment of the present disclosure;

[0016] Figure 5 is a schematic diagram of a deep learning model according to an embodiment of the present disclosure;

[0017] Figure 6 is a schematic diagram of a task execution device according to an embodiment of the present disclosure;

[0018] Figure 7 is a block diagram of an electronic device to which the task execution method can be applied according to an embodiment of the present disclosure. Detailed Embodiments

[0019] The exemplary embodiments of the present disclosure will be described below with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0020] Taking the application scenario of low-light image enhancement as an example, in a low-light environment, there is insufficient illumination. In addition, the performance of the photosensitive device of an image acquisition device (such as a camera) is limited, resulting in problems such as image noise interference, detail loss, and color distortion in the low-light image captured by the image acquisition device, thereby significantly reducing the availability of visual information in the low-light image.

[0021] Low-light image restoration can be performed based on the Transformer model. However, the computational complexity of the multi-head self-attention mechanism of the Transformer model is relatively high, almost quadratic in terms of time complexity, resulting in low efficiency of the Transformer model in processing complex image structures and also requiring more hardware resources for image processing using the Transformer model.

[0022] Thus, in order to reduce the hardware resource overhead, the present disclosure provides a task execution method, which will be described below.

[0023] Figure 1 It is a flowchart of a task execution method according to an embodiment of the present disclosure.

[0024] As Figure 1 shown, the method 100 may include operation S110 to operation S140.

[0025] In operation S110, according to the feature to be processed, a processor is used to execute multiple convolutional subtasks of a target task to obtain multiple convolutional features.

[0026] In an embodiment of the present disclosure, the feature to be processed may be an image feature to be processed, a text feature to be processed, an audio feature to be processed, or a multi-modal feature to be processed. The multi-modal feature to be processed includes features of one or more modalities among images, texts, and audios.

[0027] In the embodiments of the present disclosure, the target task includes a plurality of convolutional subtasks. The plurality of convolutional subtasks may include a first convolutional subtask and a second convolutional subtask. The sizes of the plurality of convolutional kernels respectively used for the plurality of convolutional subtasks are different. For example, the size of the first convolutional kernel used for the first convolutional subtask may be 1×1. The size of the second convolutional kernel used for the second convolutional subtask may be 3×3. The first convolutional subtask may perform convolution on the feature to be processed based on the first convolutional kernel. The second convolutional subtask may perform convolution on the feature to be processed based on the second convolutional kernel. When the processor executes the first convolutional subtask, a first convolutional feature can be obtained. When the processor executes the second convolutional task, a second convolutional feature can be obtained.

[0028] In operation S120, according to the multiple convolutional sub-features of the second convolutional feature and at least one neighborhood sub-feature of each of the multiple convolutional sub-features, the processor is used to execute the mask extraction subtask of the target task to obtain a mask feature.

[0029] In the embodiments of the present disclosure, in the second convolutional feature, one or more sub-features adjacent to the convolutional sub-feature can be used as one or more neighborhood sub-features. For example, the second convolutional feature includes a first convolutional sub-feature, a second convolutional sub-feature, and a third convolutional sub-feature. The first convolutional sub-feature is adjacent to the second convolutional sub-feature and is also adjacent to the third convolutional sub-feature. The second convolutional sub-feature and the third convolutional sub-feature can be used as two neighborhood sub-features of the first convolutional sub-feature.

[0030] In the embodiments of the present disclosure, the mask extraction subtask can extract a mask feature according to the second convolutional feature. In the process of extracting the mask feature, the mask sub-feature of the mask feature can be determined according to the feature difference information between the convolutional sub-feature and the domain sub-feature of the convolutional sub-feature.

[0031] In operation S130, according to the first convolutional feature and the mask feature, the processor is used to execute the mask processing subtask of the target task to obtain a first mask-processed feature.

[0032] In the embodiments of the present disclosure, the mask processing subtask can process the first convolutional feature based on the mask feature to obtain a first mask-processed feature.

[0033] In operation S140, according to the first mask-processed feature, the processor is used to determine the output feature of the target task.

[0034] In the embodiments of the present disclosure, various processes can be performed according to the mask-processed feature to obtain the output feature of the target task.

[0035] Through the embodiments of the present disclosure, multiple convolutional subtasks are executed according to the feature to be processed, and richer information can be obtained based on the feature to be processed. According to multiple convolutional sub-features of the second convolutional feature and the neighborhood sub-features of each of the multiple convolutional sub-features, a mask feature is obtained, and the long-range correlation of the sub-features in the second convolutional feature can be obtained based on the non-local prior. According to the mask feature and the first convolutional feature, a mask processing subtask is executed, which can increase local features and strengthen the detailed features of the data to be processed. Thus, based on the characteristics of non-local prior and convolutional processing, the global and local information of the feature to be processed can be fully utilized, and the computational complexity is effectively reduced. Compared with the multi-head self-attention mechanism, the hardware resource overhead required for the processor to execute multiple convolutional subtasks, mask extraction subtasks, and mask processing subtasks is significantly reduced, and the costs of image processing, text processing, and audio processing can be effectively reduced.

[0036] It can be understood that the method of the present disclosure has been described above, and the execution manner of the target task will be further described below.

[0037] In some embodiments, the target task may include multiple convolutional subtasks, mask extraction subtasks, mask processing subtasks, first activation subtasks, second activation subtasks, fusion subtasks, and feature integration subtasks.

[0038] For example, for inference based on a deep learning model, a user may write or use a code generation tool to generate multiple lines of code for the deep learning model. The deep learning model may include one or more target sub-models. The target sub-model may include one or more target networks. The target network may include one or more target processing modules. The target processing module may include one or more processing layers. By compiling the multiple lines of code, one or more target task sequences executed by the processor can be obtained. The processor may be a central processing unit (CPU) or an artificial intelligence processor. The artificial intelligence processor may be various processors such as a general-purpose graphics processing unit (GPGPU), a neural network processor (NPU), or a tensor processing unit (TPU). The target task sequence corresponds to the target sub-model, and when the processor executes the target task sequence, one or more functions of the target sub-model can be realized. The target task sequence may include one or more target sub-task sequences. The target sub-task sequence may correspond to the target network. When the processor executes the target sub-task sequence, one or more functions of the target network can be realized. The target sub-task sequence may include one or more target tasks. When the processor executes the target task, one or more functions of the target processing module can be realized. Next, it will be described in conjunction with Figure 2A The target processing module will be described.

[0039] Figure 2A It is a schematic diagram of a target processing module according to an embodiment of the present disclosure.

[0040] As Figure 2A shown, the target processing module MCC200 can implement a Masked Control Convolution (MCC) mechanism. The target processing module MCC200 may include multiple processing layers. The multiple processing layers may include a convolutional layer conv210, a convolutional layer conv220, a convolutional layer conv230, a mask extraction layer NLTV220, an activation layer a220, a mask processing layer mask-conv210, an activation layer a210, a fusion layer fuse200, and a feature integration layer c200.

[0041] Next, some embodiments of the above operation S110 will be described in conjunction with the convolutional layer conv210, the convolutional layer conv220, and the convolutional layer conv230.

[0042] In some embodiments, the feature to be processed may be an image feature to be processed. Inputting the feature to be processed into multiple convolutional layers respectively may obtain multiple convolutional features. As shown in Figure 2, the size of the first convolutional kernel of the convolutional layer conv210 may be 1×1. The first convolutional feature may be obtained by convolving the feature to be processed based on the first convolutional kernel. In one example, the first convolutional feature may be obtained through the following formula:

[0043] (Formula 1)

[0044] may be the first convolutional feature. may be the parameter of the convolutional layer conv210. may be the feature to be processed.

[0045] For another example, the size of the second convolutional kernel of the convolutional layer conv220 may be 3×3. The second convolutional feature may be obtained by convolving the feature to be processed based on the second convolutional kernel. In one example, the second convolutional feature may be obtained through the following formula:

[0046] (Formula 2)

[0047] may be the second convolutional feature. may be the parameter of the convolutional layer conv220.

[0048] For another example, the size of the third convolutional kernel of the convolutional layer conv230 may be 1×1. The third convolutional feature may be obtained by convolving the feature to be processed based on the third convolutional kernel. In one example, the third convolutional feature may be obtained through the following formula:

[0049] (Formula 3)

[0050] It can be the third convolution feature. It can be the parameter of the convolution layer conv230.

[0051] It can be understood that the convolution layer conv210, the convolution layer conv220, and the convolution layer conv230 respectively correspond to the first convolution subtask, the second convolution subtask, and the third convolution subtask to be executed by the processor. After the processor executes the first convolution subtask, the second convolution subtask, and the third convolution subtask respectively, the convolution layer conv210, the convolution layer conv220, and the convolution layer conv230 can respectively implement the processing of the to-be-processed feature, and obtain the first convolution feature, the second convolution feature, and the third convolution feature. The processor can execute multiple convolution subtasks serially or in parallel. Through the embodiments of the present disclosure, by respectively executing multiple convolution subtasks according to the to-be-processed feature and establishing multiple feature extraction branches, richer information can be obtained from the to-be-processed feature.

[0052] Next, some embodiments of the above operation S220 will be described in conjunction with the mask extraction layer NLTV220.

[0053] In some embodiments, the second convolution feature may include multiple convolution sub-features. Below will be combined with Figure 2B to describe the convolution sub-features of the present disclosure.

[0054] Figure 2B It is a schematic diagram of the convolution sub-feature and the neighborhood sub-feature according to an embodiment of the present disclosure.

[0055] As Figure 2B shown, the second convolution feature may include multiple convolution sub-features. The multiple convolution sub-features may include the convolution sub-feature F conv2_ij . The multiple convolution sub-features further include the following sub-features adjacent to the convolution sub-feature convolution sub-feature F conv2_ij : the convolution sub-feature F conv2_ij_11 , the convolution sub-feature F conv2_ij_12 , the convolution sub-feature F conv2_ij_13 , the convolution sub-feature F conv2_ij_21 , the convolution sub-feature F conv2_ij_23 , the convolution sub-feature F conv2_ij_31 , the convolution sub-feature F conv2_ij_32 , the convolution sub-feature F conv2_ij_33 . The multiple convolution sub-features adjacent to the convolution sub-feature convolution sub-feature F conv2_ij can be used as the neighborhood sub-features of the convolution sub-feature convolution sub-feature F conv2_ij . It can be understood that Figure 2B the convolution sub-feature F shownconv2_ij The number of neighborhood sub - features is 8. However, the present disclosure is not limited thereto, and the number of neighborhood sub - features of the convolutional sub - features can be less than 8. For example, multiple convolutional sub - features can be randomly determined as neighborhood sub - features from convolutional sub - feature F conv2_ij_11 to convolutional sub - feature F conv2_ij_33 . Also, the number of neighborhood sub - features of the convolutional sub - features can be greater than 8. Figure 2B The number of neighborhood sub - features shown is only an example.

[0056] In some embodiments, the mask extraction layer can determine multiple mask sub - features of the mask feature according to at least one neighborhood sub - feature of each of the multiple convolutional sub - features and the multiple convolutional self - sub - features. The mask sub - features can be determined according to at least one difference sub - feature for the convolutional sub - features, and the difference sub - feature can be obtained from the feature difference information between the convolutional sub - feature and the domain sub - feature of the convolutional sub - feature.

[0057] For example, the difference sub - feature is obtained from the first weight feature for the feature difference information and the feature difference information. In one example, the mask sub - feature can be determined by the following formula:

[0058] (Formula Four)

[0059] can be the mask sub - feature at the i - th row and j - th column in the mask feature. can be the convolutional sub - feature at the i - th row and j - th column. can be the convolutional sub - feature 's neighborhood sub - feature. can be the feature difference information. can be the first weight feature for the feature difference information. used to represent the set of domain sub - features of the convolutional sub - feature . can be the difference sub - feature. Through the embodiments of the present disclosure, through non - local total variation (NLTV), long - range correlations in the feature to be processed can be captured using non - local priors, the smoothness of the output feature can be enhanced, and noise can be suppressed. Taking the feature to be processed as the feature of the image to be processed as an example, the mask extraction subtask can capture long - range correlations in the low - light image, enhance the overall illumination smoothness, and suppress image noise.

[0060] In some embodiments, the first weight feature is obtained from the feature difference information processed by a preset function, the feature difference information processed by the preset function is obtained by processing the processed feature difference information using the preset function, and the processed feature difference information is obtained from the first preset weight control parameter and the feature difference information.

[0061] For example, various operations can be performed based on the feature difference information to obtain a feature difference operation result. The square of the feature difference operation result can be divided by the first processed preset weight control parameter to obtain the processed feature difference information. The first processed pre-identified weight control parameter can be obtained by multiplying the first preset value by the square of the first preset weight control parameter. The first preset value can be a value less than 0, for example, it can be -1. In one example, the first weighted feature can be determined by the following formula:

[0062] (Formula Five)

[0063] can be a preset function. can be the first preset weight control parameter. can be the feature difference operation result. Through the embodiments of the present disclosure, the first weighted feature is determined according to the feature difference information, which can further reduce the influence of noise in the data.

[0064] Accordingly, the second convolutional feature is input into the mask extraction layer NLTV220, and a mask feature can be obtained. It can be understood that the mask extraction layer NLTV220 corresponds to a mask extraction subtask to be executed by the processor. The mask extraction subtask is used to determine a plurality of mask sub-features according to a plurality of convolutional sub-features and at least one neighborhood sub-feature of each of the plurality of convolutional sub-features. When the processor executes the mask extraction subtask, the mask extraction layer NLTV220 can process the second convolutional feature to obtain a mask feature. Next, some embodiments of the above operation S130 will be described in conjunction with the activation layer a220 and the mask processing layer mask-conv210.

[0065] The mask feature is input into the activation layer a220, and an activated mask feature can be obtained. The activation function of the activation layer a220 can be the softmax function. In one example, the processed mask feature can be obtained by the following formula:

[0066] (Formula Six)

[0067] can be the mask feature. can be the activated mask feature.

[0068] It can be understood that the activation layer a220 corresponds to a first activation subtask to be executed by the processor. According to the mask feature, the processor can execute the first activation subtask to implement the processing of the mask feature by the activation layer a220 to obtain an activated mask feature.

[0069] In some embodiments, the first convolutional feature and the activated mask feature can be input into a mask processing layer to obtain a first masked processed feature. As Figure 2A shown, the mask processing layer can fuse the activated mask feature and the first depthwise separable convolutional feature into a first masked processed feature. The first depthwise separable convolutional feature is obtained by performing depthwise separable convolution on the first convolutional feature. The fusion method between the first depthwise separable convolutional feature and the activated mask feature can be multiplication. In one example, the first masked processed feature can be obtained through the following formula:

[0070] (Formula VII)

[0071] can be the first masked processed feature. can be the parameters of the depthwise separable convolution. The first depthwise separable convolutional feature and the activated mask feature can be fused.

[0072] It can be understood that the mask processing layer mask-conv210 corresponds to the mask processing subtask to be executed by the processor. The mask processing subtask can fuse the activated mask feature and the first depthwise separable convolutional feature into a first masked processed feature. When the processor executes the mask extraction subtask, it can implement the processing of the mask processing layer mask-conv210 on the first convolutional feature and the activated mask feature to obtain the first masked processed feature. Through the embodiments of the present disclosure, when the processor executes the mask processing task, it can enhance local features, effectively utilize the details of the data, make the processing effect more refined, and effectively optimize the feature expression. Taking the feature to be processed as the feature of the image to be processed as an example, when the processor executes the mask processing subtask, it can enhance the local features of the image, effectively utilize the information of the image edges and details, and process the image more finely. It can also dynamically adjust the degree of light enhancement in the low-light image according to the image content, adaptively process the information of different positions and channels, and thus effectively optimize the feature expression under low-light conditions.

[0073] Next, some embodiments of the above operation S140 will be further described in combination with the activation layer a210, the fusion layer fuse200, and the feature integration layer c200.

[0074] In some embodiments, when the first masked processed feature is input into the activation layer a210, an activation feature can be obtained. The activation function of a210 in the activation layer can be the HSwish function. In one example, the activation feature can be obtained through the following formula:

[0075] (Formula VIII)

[0076] It can be an activation feature. The HSwish function will be further described below.

[0077] For example, the activation function of the activation layer a210 can fuse the first masked feature and the processed rectified linear feature into an activation feature. The fusion method can be multiplication. The processed rectified linear feature is obtained according to the rectified linear feature and the first preset rectification parameter. The rectified linear feature is obtained by processing the second masked feature using the rectified linear function. The second masked feature is obtained according to the first masked feature and the second preset rectification parameter. In one example, dividing the rectified linear feature by the first preset rectification parameter can obtain the processed rectified linear feature. The first preset rectification parameter can be 6, for example. The second masked feature is obtained by adding the first masked feature and the second preset rectification parameter. The second preset rectification parameter can be 3. The activation feature can also be obtained through the following formula:

[0078] (Formula Nine)

[0079] It can be a rectified linear function.

[0080] Next, in some embodiments, inputting the activation feature and the third convolutional feature into the fusion layer can obtain a fusion feature. Inputting the fusion feature into the feature integration layer can obtain the output feature of the target processing module.

[0081] For example, the fusion layer fuse200 multiplies the activation feature and the third convolutional feature. In one example, the fusion feature can be obtained through the following formula :

[0082] (Formula Ten)

[0083] The feature integration layer can be a convolutional layer. The convolutional kernel of this convolutional layer can be a 1×1 convolutional kernel, which can perform convolution on the fusion feature to obtain the output feature seen by the target processing module.

[0084] It can be understood that the activation layer a210, the fusion layer fuse200, and the feature integration layer c200 respectively correspond to the second activation subtask, the fusion subtask, and the feature integration subtask to be executed by the processor. According to the first masked feature, the processor executes the second activation subtask to implement the processing of the first masked feature by the activation layer a210, obtaining an activated feature. According to the activated feature and the third convolutional feature, the processor can execute the fusion subtask to implement the processing of the activated feature and the third convolutional feature by the fusion layer fuse200, obtaining a fused feature. According to the fused feature, the processor can execute the feature integration task to implement the processing of the fused feature by the feature integration layer c200, obtaining the output feature of the target task. The output feature of the target task can be the output feature of the above-mentioned target processing module. In some other embodiments, the fused feature can also be used as the output feature of the target task.

[0085] It can be understood that the above has described the target processing module and the target task of the present disclosure. Next, the target network and the target subtask sequence corresponding to the target network will be described.

[0086] In some embodiments, the target task can be a task in the target subtask sequence. The target subtask sequence includes a first normalization task, a target task, and a first residual connection task. As described above, after compiling multiple lines of code of the target network, the target subtask sequence can be obtained. Next, the Figure 3 target network will be described.

[0087] Figure 3 is a schematic diagram of a target network according to an embodiment of the present disclosure.

[0088] As Figure 3 shown, the target network MCA300 can be a Masked Control Attention (MCA) network. The target network MCA300 can include multiple processing modules. As Figure 3 shown, the multiple processing modules can include a first normalization module LN301, a target processing module MCC300, and a first residual connection module RC301.

[0089] In some embodiments, the feature to be processed may be the first normalized feature. For example, by inputting the input feature of the target network into the first normalization module LN301, the first normalized feature can be obtained. The first normalization module LN301 can implement layer normalization. By inputting the first normalized feature into the target processing module MCC300, the output feature of the target processing module MCC300 can be obtained. By inputting the input feature of the target network MCA300 and the output feature of the target processing module MCC300 into the first residual connection module RC301, the first residual connection feature can be obtained. It can be understood that the above description of the target processing module MCC200 also applies to the target processing module MCC300, and details are not described herein again in the present disclosure.

[0090] It can be understood that the first normalization module LN301, the target processing module MCC300, and the first residual connection module RC301 respectively correspond to the first normalization task, the target task, and the first residual connection task to be executed by the processor. According to the input feature of the target subtask sequence, the processor executes the first normalization task to implement the processing of the input feature by the first normalization module LN301, and obtains the feature to be processed. According to the feature to be processed, the processor can execute the target task to implement the processing of the feature to be processed by the target processing module MCC300, and obtain the output feature of the target task. According to the output feature of the target task and the input feature of the target subtask sequence, the processor can execute the first residual connection task to implement the processing of the output feature of the target task and the input feature of the target subtask sequence by the first residual connection module RC301, and obtain the first residual connection feature. Through the embodiments of the present disclosure, when the processor executes the target task, different weights can be assigned to different regions based on the dynamic masking mechanism, global features can be efficiently extracted, long-range dependencies across regions can be captured, global context information can be obtained, and then the feature output by the target task is residually connected to the input feature of the target subtask sequence, which can enhance the stability and expressiveness of the feature.

[0091] Next, in some embodiments, according to the first residual connection feature, the processor can be used to determine the output feature of the target subtask sequence. The following will continue to be described in conjunction with Figure 3 for illustration.

[0092] As Figure 3As shown, the target network MCA300 may further include a second normalization module LN302, a depthwise separable convolution module DW300, and a second residual connection module RC302. Inputting the first residual connection feature into the second layer normalization module LN302 can obtain the second layer normalization feature. Inputting the second layer normalization feature into the depthwise separable convolution module DW300 can obtain the second depthwise separable convolution feature. Inputting the second depthwise separable convolution feature and the first residual connection feature into the second residual connection module RC302 can obtain the output feature of the target network MCA300.

[0093] It can be understood that the second normalization module LN302, the depthwise separable convolution module DW300, and the second residual connection module RC302 respectively correspond to the second normalization task, the depthwise separable convolution task, and the second residual connection task to be executed by the processor. According to the first residual connection feature, the processor executes the second normalization task to implement the processing of the first residual connection feature by the second normalization module LN302, obtaining the second normalization feature. According to the second normalization feature, the processor can execute the depthwise separable convolution task to implement the processing of the feature to be processed by the depthwise separable convolution module DW300, obtaining the second depthwise separable convolution feature. According to the second depthwise separable convolution feature and the first residual connection feature, the processor can execute the second residual connection task to implement the processing of the second depthwise separable convolution feature and the first residual connection feature by the second residual connection module RC302, obtaining the second residual connection feature. The second residual connection feature can be used as the output feature of the target subtask sequence. Through the embodiments of the present disclosure, when the processor executes the depthwise separable convolution task, it can fully obtain local structural information with relatively low computational resource overhead, extract fine-grained information in the first residual connection feature, capture tiny but crucial spatial information, and fully perceive complex information in the feature. Subsequently, when the processor executes the second residual connection task, it can retain global features while further obtaining local information, improving the robustness of task execution.

[0094] It can be understood that the target network and the target subtask sequence of the present disclosure have been described above. Next, the target task sequence and the target submodel will be described.

[0095] In some embodiments, the target task sequence includes an N-level target subtask sequence, which is sequentially executed by a processor. The input feature of the first-level target subtask sequence in the N-level target subtask sequence is the input feature of the target task sequence. The input feature of the (n + 1)-level target subtask sequence in the N-level target subtask sequence is the output feature of the n-level target subtask sequence in the N-level target subtask sequence. The output feature of the target task sequence is obtained based on the output feature of the N-level target subtask sequence in the N-level target subtask sequence. n is an integer greater than or equal to 1 and less than N. N is an integer greater than 1. As described above, after compiling multiple lines of code of the target submodel, a target task sequence can be obtained. The following will be combined with Figure 4 to illustrate the target submodel.

[0096] Figure 4 is a schematic diagram of a target submodel according to an embodiment of the present disclosure.

[0097] As Figure 4 shown, the target submodel MMCA400 can be a Multi-Masked Control Attention (MMCA) submodel. The target submodel MMCA400 can include an N-level target network. As Figure 4 shown, the N-level target network can include multiple target networks such as the first-level target network MCA401 and the second-level target network MCA402. The input feature of the first-level target network MCA401 can be the input feature of the target submodel MMCA400. The output feature of the first-level target network MCA401 can be used as the input feature of the second-level target network MCA402. It can be understood that the above description of the target network MCA300 also applies to the first-level target network MCA401 and the second-level target network MCA402, and the present disclosure will not elaborate herein.

[0098] It can be understood that after compiling the code, the N-level target network including the first-level target network MCA401 and the second-level target network MCA402 corresponds to the N-level target subtask sequence to be executed by the processor. According to the input feature of the target task sequence, the processor sequentially executes the N-level target subtask sequence to implement the processing of the input feature by the N-level target network and obtain the output feature of the N-level target subtask sequence. Next, based on the output feature of the N-level target subtask sequence, the output feature of the target task sequence can be obtained. The following will be combined with Figure 4 for further illustration.

[0099] As Figure 4As shown, the target sub-model may further include a convolutional module conv40 and a third residual connection module RC40. Inputting the output features of the Nth-level target network into the convolutional module conv40 can obtain a first execution result. Inputting the first execution result and the input features of the target sub-model MMCA400 into the third residual connection module RC40 can obtain the output features of the target sub-model MMCA400.

[0100] It can be understood that the convolutional module conv40 and the third residual connection module RC40 respectively correspond to the first task to be executed and the second task to be executed by the processor. The first task to be executed is used to perform a convolution operation according to the output features of the Nth-level target sub-task sequence, and the second task to be executed is used to perform a fusion operation according to the input features of the target task sequence and the first execution result of the first task to be executed. According to the output features of the Nth-level target network, the processor executes the first task to be executed to implement the processing of the output features of the Nth-level target network by the convolutional module conv40, obtaining a first execution result. According to the first execution result and the input features of the target task sequence, the processor can execute the second task to be executed to implement the fusion processing of the input features of the target task sequence and the first execution result by the third residual connection module RC40, obtaining a second execution result. The second execution result can be used as the output features of the target task sequence.

[0101] It can be understood that the above has described the target sub-model and the target task sequence of the present disclosure. Next, the deep learning model of the present disclosure and the corresponding M target task sequences will be described. M is an integer not less than 1.

[0102] In some embodiments, the above method may further include: according to the data to be processed, using the processor to execute at least one preprocessing task to obtain the input features of the first target task sequence among the M target task sequences. According to the output features of the Mth target task sequence among the M target task sequences, using the processor to execute at least one postprocessing task to obtain the processed data corresponding to the data to be processed. As described above, after compiling multiple lines of code of the deep learning model, M target task sequences, at least one preprocessing task, and at least one postprocessing task can be obtained.

[0103] Next, it will be combined with Figure 5 to describe the deep learning model of the present disclosure.

[0104] Figure 5 is a schematic diagram of a deep learning model according to an embodiment of the present disclosure.

[0105] As Figure 5As shown, the deep learning model NLAD500 may include M target sub-models, at least one preprocessing module, and at least one post-processing module. The deep learning model NLDA500 may be a Non-local dynamic adaptive (NLDA) model. The deep learning model NLAD500 may include the preprocessing module conv511, M target sub-models, the post-processing module conv521, the post-processing module conv522, and the post-processing module RC523. The preprocessing module conv511 may perform a convolution operation according to the data to be processed. The convolution kernel of the preprocessing module conv511 may be a 3×3 convolution kernel. The post-processing module conv521 may perform a convolution operation according to the output features of the M-th target sub-model. The convolution kernel of the post-processing module conv521 may be a 3×3 convolution kernel. The post-processing module conv522 may perform a convolution operation according to the first post-processing result output by the post-processing module conv521. The convolution kernel of the post-processing module conv522 may be a 1×1 convolution kernel. The post-processing module conv523 may perform a fusion operation according to the second post-processing result output by the post-processing module conv522 and the data to be processed.

[0106] In some embodiments, when M>1, inputting the input features of the m-th target sub-model into the (m + 1)-th target sub-model may obtain the output features of the (m + 1)-th target sub-model. For example, the data to be processed may be the image to be processed image500. The image to be processed image500 may be a low-light image. Inputting the image to be processed image500 into the preprocessing module conv511 may obtain the input features of the first target sub-model MMCA501. Inputting the input features of the first target sub-model MMCA501 into the first target sub-model MMCA501 may obtain the output features of the first target sub-model MMCA501, which are used as the input features of the second target sub-model MMCA502. Inputting the output features of the M-th target sub-model into the post-processing module conv521 may obtain the first post-processing result. Inputting the first post-processing result into the post-processing module conv522 may obtain the second post-processing result. Inputting the second post-processing result and the image to be processed image500 into the post-processing module RC523 may obtain the processed image image501 corresponding to the image to be processed image500.

[0107] It can be understood that the above description of the present disclosure is given by taking the data to be processed as an image to be processed as an example. However, the present disclosure is not limited thereto, and the data to be processed may also be text to be processed or audio to be processed, which will be described below.

[0108] In some embodiments, at least one preprocessing task includes a second preprocessing task and a third preprocessing task. The second preprocessing task is used to perform an embedding operation according to the data to be processed, and the third preprocessing task is used to perform a convolution operation according to the second preprocessing result of the second preprocessing task.

[0109] In some embodiments, according to the data to be processed, using a processor to execute at least one preprocessing task to obtain the input features of the first target task sequence among the M target task sequences may include: according to the data to be processed, using a processor to execute the second preprocessing task to obtain a second preprocessing result. For example, taking the data to be processed as the data to be processed text as an example, the processor may perform an embedding operation on the data to be processed text to obtain a second preprocessing result.

[0110] In some embodiments, according to the data to be processed, using a processor to execute at least one preprocessing task to obtain the input features of the first target task sequence among the M target task sequences may include: according to the second preprocessing result, using a processor to execute the third preprocessing task to obtain the input features of the first target task sequence. For example, the processor may perform a convolution operation on the second preprocessing result to obtain a third preprocessing result. The convolution kernel of this convolution operation may be a 3×3 convolution kernel.

[0111] It can be understood that the processing method of the data to be processed audio is the same as or similar to that of the data to be processed text, and the present disclosure will not elaborate herein.

[0112] It can be understood that some ways of obtaining the processed data corresponding to the data to be processed are described above. In some other embodiments, the data to be processed may be used as a training sample data and may have a label. According to one or more of the data to be processed, the processed data, and the label, the above-mentioned deep learning model may be trained. The following will be described.

[0113] In some embodiments, the above method may further include: according to the label of the data to be processed and the processed data, using a processor to execute at least one of multiple loss determination tasks to obtain at least one of multiple loss information. For example, the multiple loss determination tasks may include a first loss determination task. The first loss determination task may determine the first loss information according to the difference between the processed data and the label. In one example, the first loss information may be determined by the following formula:

[0114] (Formula XI)

[0115] may be the first loss information. may be the data to be processed. may be the processed data, may be the data to be processed. It can be the difference between the processed data and the label. It can be understood that if the data to be processed is the image to be processed, the processed data can be the processed image, and the first loss information can represent the pixel value difference between the processed image and the label.

[0116] In some embodiments, the above method may further include: according to at least one of the multiple loss information, using a processor to execute at least one of multiple parameter adjustment tasks. The multiple parameter adjustment tasks include a first parameter adjustment task, a second parameter adjustment task, and a third parameter adjustment task. The first reference adjustment task is used to adjust the parameters of at least one target task sequence among the M target task sequences. The second parameter adjustment task is used to adjust the parameters of at least one preprocessing task. The third parameter adjustment task is used to adjust the parameters of at least one postprocessing task.

[0117] For example, one or more parameters of the preprocessing module can be used as one or more parameters of the preprocessing task. One or more parameters of the postprocessing module can be used as one or more parameters of the postprocessing task. One or more parameters of the target submodel can be used as one or more parameters of the target task sequence. It can be understood that one or more parameters of the target submodel include one or more parameters of the target network, and one or more parameters of the target network include one or more parameters of the target processing module. One or more parameters of the target processing module include one or more parameters of the processing layer.

[0118] For another example, at least one of the multiple parameter adjustment tasks can be executed based on the gradient descent method.

[0119] It can be understood that the above first loss determination task is used as an example to illustrate the present disclosure. However, the present disclosure is not limited thereto. The multiple loss determination tasks may further include a second loss determination task, which will be described below.

[0120] In some embodiments, the second loss determination task may determine the second loss information according to multiple difference values. The difference value is determined according to the processed sub-data and the neighboring sub-data of the processed sub-data. For example, the processed data may include multiple processed sub-data. The neighboring sub-data of the processed sub-data can be the processed sub-data adjacent to the processed sub-data in the processed data. Taking the processed data as the processed image as an example, the processed sub-data can be the pixels in the processed image. The neighboring sub-data of the pixel can be the pixels adjacent to the pixel. The difference value can be the difference between the pixel values of different pixels.

[0121] In some embodiments, the second loss determination task may also determine the second loss information according to multiple weighted difference values. The weighted difference value is obtained according to the second weight feature and the difference value.

[0122] In one example, the second loss information can be determined by the following formula:

[0123] (Formula XII)

[0124] Taking the processed data as the processed image as an example, can be the pixel at the i-th row and j-th column of the processed image. can be the pixel at the k-th row and l-th column of the processed image, and can be the domain sub-data. can be the difference between the processed sub-data and the neighborhood sub-data of the processed sub-data. can be the second weight feature for the difference value. Based on the second loss information, the interference of noise in the data can be suppressed, and the texture structure of the image can be maintained.

[0125] In some embodiments, the second weight feature is obtained from the difference value processed by a preset function. The difference value processed by the preset function is obtained by processing the processed difference value using the preset function. The processed difference value is obtained based on the second preset weight control parameter and the difference value.

[0126] For example, various operations can be performed based on the difference value to obtain the data difference operation result. The square of the data difference operation result can be divided by the second processed preset weight control parameter to obtain the processed difference value. The second processed preset recognition weight control parameter can be obtained by multiplying the second preset value by the square of the second preset weight control parameter. The first preset value can be a value less than 0, for example, it can be -1. In one example, the first weight feature can be determined by the following formula:

[0127] (Formula XIII)

[0128] can be the preset function. can be the second preset weight control parameter. can be the data difference operation result.

[0129] In some embodiments, during the training of the deep learning model, an adaptive moment estimation (adam) optimizer with the first adaptive parameter (beta1) being 0.6 and the second adaptive parameter (beta2) being 0.66 can be used for parameter adjustment. And, the learning rate can be set to 1x10 -5 . During the training process, the batch size can be 256.

[0130] In some embodiments, multiple data augmentation strategies are applied during the training process to improve the generalization ability of the model, including random cropping, horizontal flipping, brightness adjustment, and noise injection, etc. For example, for random cropping, the cropping ratio is controlled within a certain range (such as 80%-100%) to ensure the integrity of data features. In addition, training can be carried out in stages. In the initial stage, a lower learning rate is used for model warm-up, and then the learning rate is gradually increased to a preset value (1x10 -5 ), and the learning rate is dynamically adjusted to accelerate convergence.

[0131] It can be understood that the method of the present disclosure has been described above, and the apparatus of the present disclosure will be described below.

[0132] Figure 6 is a schematic diagram of a task execution apparatus according to an embodiment of the present disclosure.

[0133] As Figure 6 shown, the apparatus 600 may include a first execution module 610, a second execution module 620, a third execution module 630, and a determination module 640.

[0134] The first execution module 610 is configured to execute multiple convolutional subtasks of a target task using a processor according to the feature to be processed, and obtain multiple convolutional features. The multiple convolutional features include a first convolutional feature and a second convolutional feature, and the second convolutional feature includes multiple convolutional sub-features.

[0135] The second execution module 620 is configured to execute a mask extraction subtask of the target task using a processor according to the multiple convolutional sub-features of the second convolutional feature and at least one neighborhood sub-feature of each of the multiple convolutional sub-features, and obtain a mask feature.

[0136] The third execution module 630 is configured to execute a mask processing subtask of the target task using a processor according to the first convolutional feature and the mask feature, and obtain a first mask-processed feature.

[0137] The first determination module 640 is configured to determine an output feature of the target task using a processor according to the first mask-processed feature.

[0138] In some embodiments, the mask feature includes multiple mask sub-features, and the mask extraction subtask is configured to determine the multiple mask sub-features according to the multiple convolutional sub-features and at least one neighborhood sub-feature of each of the multiple convolutional sub-features. The mask sub-feature is determined according to at least one difference sub-feature for the convolutional sub-feature, and the difference sub-feature is obtained according to the feature difference information between the convolutional sub-feature and the neighborhood sub-feature of the convolutional sub-feature.

[0139] In some embodiments, the differential sub-feature is obtained based on a first weight feature for feature difference information and the feature difference information. The first weight feature is obtained based on the feature difference information processed by a preset function. The feature difference information processed by the preset function is obtained by processing the processed feature difference information using the preset function. The processed feature difference information is obtained based on a first preset weight control parameter and the feature difference information.

[0140] In some embodiments, the target task includes a first activation subtask. The third execution module includes: a first execution sub-module, configured to execute the first activation subtask using a processor according to a mask feature to obtain an activated mask feature. A second execution sub-module, configured to execute a mask processing subtask using the processor according to the activated mask feature and a first convolutional feature to obtain a first masked processed feature.

[0141] In some embodiments, the mask processing subtask is configured to fuse the activated mask feature and a first depthwise separable convolutional feature into a first masked processed feature, and the first depthwise separable convolutional feature is obtained by performing depthwise separable convolution on the first convolutional feature.

[0142] In some embodiments, the target task includes a second activation subtask and a fusion subtask, and the multiple convolutional features further include a third convolutional feature. The first determination module includes: a third execution sub-module, configured to execute the second activation subtask using a processor according to the first masked processed feature to obtain an activation feature. A fourth execution sub-module, configured to execute the fusion subtask using the processor according to the activation feature and the third convolutional feature to obtain an output feature of the target task.

[0143] In some embodiments, the second activation subtask is configured to fuse the first masked processed feature and a processed rectified linear unit feature into an activation feature. The processed rectified linear unit feature is obtained based on a rectified linear unit feature and a first preset rectification parameter. The rectified linear unit feature is obtained by processing a second masked processed feature using a rectified linear unit function. The second masked processed feature is obtained based on the first masked processed feature and a second preset rectification parameter.

[0144] In some embodiments, the target task is a task in a target subtask sequence of a target task sequence. The target subtask sequence includes a first normalization task, the target task, and a first residual connection task. The feature to be processed of the target task is a first normalization feature, which is obtained by the processor performing the first normalization task according to the input feature of the target subtask sequence. The apparatus further includes: a fourth execution module, configured to perform the first residual connection task according to the input feature of the target subtask sequence and the output feature of the target task by using the processor to obtain a first residual connection feature; and a second determination module, configured to determine the output feature of the target subtask sequence according to the first residual connection feature by using the processor.

[0145] In some embodiments, the target subtask sequence further includes a second normalization task, a depthwise separable convolution task, and a second residual connection task. The second determination module includes: a fifth execution sub-module, configured to perform the second normalization task according to the first residual connection feature by using the processor to obtain a second normalization feature; a sixth execution sub-module, configured to perform the depthwise separable convolution task according to the second normalization feature by using the processor to obtain a second depthwise separable convolution feature; and a seventh execution sub-module, configured to perform the second residual connection task according to the second depthwise separable convolution feature and the first residual connection feature by using the processor to obtain a second residual connection feature as the output feature of the target subtask sequence.

[0146] In some embodiments, the target task sequence includes an N-level target subtask sequence, which is sequentially executed by the processor. The input feature of the first-level target subtask sequence in the N-level target subtask sequence is the input feature of the target task sequence. The input feature of the (n + 1)-th level target subtask sequence in the N-level target subtask sequence is the output feature of the n-th level target subtask sequence in the N-level target subtask sequence. The output feature of the target task sequence is obtained according to the output feature of the N-th level target subtask sequence in the N-level target subtask sequence, where n is an integer not less than 1 and less than N, and N is an integer greater than 1.

[0147] In some embodiments, the target task sequence further includes a first task to be executed and a second task to be executed. The first task to be executed is configured to perform a convolution operation according to the output feature of the N-th level target subtask sequence. The second task to be executed is configured to perform a fusion operation according to the input feature of the target task sequence and the first execution result of the first task to be executed. The output feature of the target task sequence is the second execution result of the second task to be executed.

[0148] In some embodiments, the number of target task sequences is M, where M is an integer not less than 1. The apparatus further includes: a fifth execution module, configured to execute at least one preprocessing task according to the data to be processed by using a processor, so as to obtain input features of the first target task sequence among the M target task sequences. A sixth execution module, configured to execute at least one postprocessing task according to output features of the Mth target task sequence among the M target task sequences by using a processor, so as to obtain processed data corresponding to the data to be processed.

[0149] In some embodiments, at least one preprocessing task includes a first preprocessing task, and the first preprocessing task is configured to perform a convolution operation according to the data to be processed. The fifth execution module includes: an eighth execution sub-module, configured to execute the first preprocessing task according to the data to be processed by using a processor, so as to obtain input features of the first target task sequence.

[0150] In some embodiments, at least one preprocessing task includes a second preprocessing task and a third preprocessing task. The second preprocessing task is configured to perform an embedding operation according to the data to be processed, and the third preprocessing task is configured to perform a convolution operation according to a second preprocessing result of the second preprocessing task. The fifth execution module includes: a ninth execution sub-module, configured to execute the second preprocessing task according to the data to be processed by using a processor, so as to obtain a second preprocessing result. A tenth execution sub-module, configured to execute the third preprocessing task according to the second preprocessing result by using a processor, so as to obtain input features of the first target task sequence.

[0151] In some embodiments, at least one postprocessor task includes a first postprocessing task, a second postprocessing task, and a third postprocessing task. The first postprocessing task is configured to perform a convolution operation according to output features of the Mth target task sequence. The second postprocessing task is configured to perform a convolution operation according to a first postprocessing result of the first postprocessing task. The third postprocessing task is configured to perform a fusion operation according to a second postprocessing result of the second postprocessing task and the data to be processed. The sixth execution module includes: an eleventh execution sub-module, configured to execute the first postprocessing task according to output features of the Mth target task sequence by using a processor, so as to obtain a first postprocessing result. A twelfth execution sub-module, configured to execute the second postprocessing task according to the first postprocessing result by using a processor, so as to obtain a second postprocessing result. A thirteenth execution sub-module, configured to execute the third postprocessing task according to the second postprocessing result and the data to be processed by using a processor, so as to obtain processed data.

[0152] In some embodiments, the apparatus further includes: a seventh execution module, configured to perform, by using a processor, at least one of a plurality of loss determination tasks based on a tag of data to be processed and processed data, to obtain at least one of a plurality of loss information. An eighth execution module, configured to perform, by using a processor, at least one of a plurality of parameter adjustment tasks based on at least one of the plurality of loss information. The plurality of parameter adjustment tasks include a first parameter adjustment task, a second parameter adjustment task, and a third parameter adjustment task. The first reference adjustment task is used to adjust parameters of at least one target task sequence among M target task sequences. The second parameter adjustment task is used to adjust parameters of at least one preprocessing task. The third parameter adjustment task is used to adjust parameters of at least one postprocessing task.

[0153] In some embodiments, the processed data includes a plurality of processed sub-data. The plurality of loss determination tasks include a first loss determination task and a second loss determination subtask. The plurality of loss information includes a first loss information and a second loss information. The first loss determination task is used to determine the first loss information according to a difference between the processed data and the tag. The second loss determination task is used to determine the second loss information according to a plurality of difference values, and the difference value is determined according to a difference between the processed sub-data and neighborhood sub-data of the processed sub-data.

[0154] In some embodiments, the second loss determination task is used to determine the second loss information according to a plurality of weighted difference values, and the weighted difference value is obtained according to a second weight feature and the difference value. The second weight feature is obtained by processing the difference value by using a preset function. The difference value processed by the preset function is obtained by processing the processed difference value by using the preset function. The processed difference value is obtained according to a second preset weight control parameter and the difference value.

[0155] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0156] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0157] Figure 7FIG. 0 shows a schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementations of the present disclosure described and / or claimed herein.

[0158] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0159] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0160] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the task execution method. For example, in some embodiments, the task execution method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the task execution method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the task execution method by any other suitable means (e.g., by means of firmware).

[0161] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0162] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes may be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0163] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0164] In order to provide interaction with a user, the systems and techniques described herein may be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)); and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other kinds of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0165] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: Local Area Network (LAN), Wide Area Network (WAN), and the Internet.

[0166] A computer system can include clients and servers. The clients and servers are generally far apart from each other and typically interact through a communication network. The client-server relationship is created by computer programs that run on the respective computers and have a client-server relationship with each other.

[0167] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0168] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A task execution method, comprising: According to the feature to be processed, a processor is used to execute a plurality of convolution subtasks of the target task to obtain a plurality of convolution features, wherein the plurality of convolution features include a first convolution feature and a second convolution feature, and the second convolution feature includes a plurality of convolution subfeatures; According to the multiple convolution sub-features of the second convolution feature and at least one neighborhood sub-feature of each of the multiple convolution sub-features, using the processor to perform the mask extraction subtask of the target task to obtain a mask feature; According to the first convolution feature and the mask feature, using the processor to perform the mask processing subtask of the target task to obtain a first masked feature; Determining, using the processor, output features of the target task based on the first masked features.

2. The method according to claim 1, wherein: The mask feature includes a plurality of mask sub-features, and the mask extraction subtask is used to determine the plurality of mask sub-features according to the plurality of convolution sub-features and at least one neighborhood sub-feature of each of the plurality of convolution sub-features. The mask sub-feature is determined based on at least one difference sub-feature used for the convolution sub-feature, and the difference sub-feature is obtained based on feature difference information between the convolution sub-feature and a neighborhood sub-feature of the convolution sub-feature.

3. The method according to claim 2, wherein: The difference sub-feature is obtained based on the first weight feature used for the feature difference information and the feature difference information, The first weight feature is obtained based on feature difference information processed by a preset function, the feature difference information processed by a preset function is obtained by processing the processed feature difference information using a preset function, and the processed feature difference information is obtained based on a first preset weight control parameter and the feature difference information.

4. The method according to claim 1, wherein: The target task includes a first activation subtask, The step of using the processor to perform the mask processing subtask of the target task according to the first convolution feature and the mask feature to obtain the first masked feature includes: According to the mask feature, using the processor to perform the first activation subtask to obtain an activated mask feature; The processor is used to perform the mask processing subtask according to the activated mask feature and the first convolution feature to obtain the first masked feature.

5. The method according to claim 4, wherein: The mask processing subtask is used to fuse the activated mask feature and the first depth-separable convolutional feature into the first masked feature, where the first depth-separable convolutional feature is obtained by performing a depth-separable convolution on the first convolutional feature.

6. The method according to claim 1, wherein: The target task includes a second activation subtask and a fusion subtask, and the plurality of convolutional features also includes a third convolutional feature. Determining the output feature of the target task by using the processor according to the first masked feature includes: executing the second activation subtask using the processor according to the first masked feature to obtain an activation feature; The processor is used to perform the fusion subtask according to the activation feature and the third convolution feature to obtain the output feature of the target task.

7. The method according to claim 6, wherein: The second activation subtask is used to fuse the first masked feature and the processed linear rectification feature into the activation feature, the processed linear rectification feature is obtained based on the linear rectification feature and the first preset rectification parameter, the linear rectification feature is obtained by processing the second masked feature using a linear rectification function, and the second masked feature is obtained based on the first masked feature and the second preset rectification parameter.

8. The method according to claim 1, wherein: The target task is a task in a target subtask sequence of a target task sequence, and the target subtask sequence includes a first normalization task, the target task, and a first residual connection task. The feature to be processed of the target task is a first normalized feature, and the first normalized feature is obtained by using the processor to perform the first normalized task according to the input feature of the target subtask sequence; The method further comprises: According to the input features of the target subtask sequence and the output features of the target task, using the processor to perform the first residual connection task to obtain the first residual connection feature; The processor is used to determine output features of the target subtask sequence according to the first residual connection features.

9. The method according to claim 8, wherein: The target subtask sequence also includes a second normalization task, a depth-separable convolution task, and a second residual connection task. Determining the output feature of the target subtask sequence by using the processor according to the first residual connection feature includes: According to the first residual connection feature, using the processor to perform the second normalization task to obtain a second normalized feature; According to the second normalized feature, using the processor to perform the depthwise separable convolution task to obtain a second depthwise separable convolution feature; According to the second depth-wise separable convolutional features and the first residual connection features, the processor is used to perform the second residual connection task to obtain a second residual connection feature as an output feature of the target subtask sequence.

10. The method according to claim 8, wherein: The target task sequence includes an N-level target subtask sequence, which is executed sequentially by the processor, the input features of the 1st-level target subtask sequence in the N-level target subtask sequence are the input features of the target task sequence, the input features of the n+1th-level target subtask sequence in the N-level target subtask sequence are the output features of the nth-level target subtask sequence in the N-level target subtask sequence, and the output features of the target task sequence are obtained based on the output features of the Nth-level target subtask sequence in the N-level target subtask sequence, n is an integer not less than 1 and less than N, and N is an integer greater than 1.

11. The method according to claim 9, wherein: The target task sequence also includes a first task to be executed and a second task to be executed, wherein the first task to be executed is used to perform a convolution operation according to the output features of the Nth level target subtask sequence, and the second task to be executed is used to perform a fusion operation according to the input features of the target task sequence and the first execution result of the first task to be executed, and the output feature of the target task sequence is the second execution result of the second task to be executed.

12. The method according to claim 8, wherein: The target task sequence is M, where M is an integer not less than 1. The method further comprises: According to the data to be processed, using the processor to perform at least one preprocessing task to obtain input features of the first target task sequence among the M target task sequences; According to the output characteristics of the Mth target task sequence among the M target task sequences, the processor is used to perform at least one post-processing task to obtain processed data corresponding to the data to be processed.

13. The method according to claim 12, wherein: The at least one preprocessing task includes a first preprocessing task, wherein the first preprocessing task is used to perform a convolution operation according to the data to be processed, The step of using the processor to perform at least one preprocessing task according to the data to be processed to obtain the input features of the first target task sequence among the M target task sequences includes: The processor is used to execute the first preprocessing task according to the data to be processed to obtain input features of the first target task sequence.

14. The method according to claim 12, wherein: The at least one preprocessing task includes a second preprocessing task and a third preprocessing task, wherein the second preprocessing task is used to perform an embedding operation according to the data to be processed, and the third preprocessing task is used to perform a convolution operation according to a second preprocessing result of the second preprocessing task, The step of using the processor to perform at least one preprocessing task according to the data to be processed to obtain the input features of the first target task sequence among the M target task sequences includes: According to the data to be processed, using the processor to perform the second preprocessing task to obtain the second preprocessing result; According to the second preprocessing result, the processor is used to perform a third preprocessing task to obtain input features of the first target task sequence.

15. The method according to claim 12, wherein: The at least one post-processor task includes a first post-processing task, a second post-processing task and a third post-processing task, wherein the first post-processing task is used to perform a convolution operation according to an output feature of the Mth target task sequence, the second post-processing task is used to perform a convolution operation according to a first post-processing result of the first post-processing task, and the third post-processing task is used to perform a fusion operation according to a second post-processing result of the second post-processing task and the data to be processed, The step of performing at least one post-processing task using the processor according to the output feature of the Mth target task sequence among the M target task sequences to obtain a processing result corresponding to the data to be processed includes: According to the output feature of the Mth target task sequence, using the processor to perform the first post-processing task to obtain the first post-processing result; According to the first post-processing result, using the processor to perform the second post-processing task to obtain the second post-processing result; The processor is used to perform the third post-processing task according to the second post-processing result and the data to be processed to obtain the processed data.

16. The method according to claim 12, further comprising: According to the label of the data to be processed and the processed data, using the processor to perform at least one of a plurality of loss determination tasks to obtain at least one of a plurality of loss information; Based on at least one of the multiple loss information, the processor is used to execute at least one of the multiple parameter adjustment tasks, wherein the multiple parameter adjustment tasks include a first parameter adjustment task, a second parameter adjustment task and a third parameter adjustment task, the first reference adjustment task is used to adjust the parameters of at least one target task sequence among the M target task sequences, the second parameter adjustment task is used to adjust the parameters of at least one pre-processing task, and the third parameter adjustment task is used to adjust the parameters of at least one post-processing task.

17. The method according to claim 16, wherein: The processed data includes a plurality of processed sub-data, The plurality of loss determination tasks include a first loss determination task and a second loss determination subtask, the plurality of loss information include first loss information and second loss information, The first loss determination task is used to determine the first loss information based on the difference between the processed data and the label, and the second loss determination task is used to determine the second loss information based on multiple difference values, and the difference values ​​are determined based on the difference between the processed sub-data and the neighborhood sub-data of the processed sub-data.

18. The method according to claim 17, wherein: The second loss determination task is used to determine the second loss information according to a plurality of weighted difference values, wherein the weighted difference values ​​are obtained according to the second weight feature and the difference value, The second weight feature is obtained according to the difference value processed by a preset function, the difference value processed by the preset function is obtained by processing the processed difference value using a preset function, and the processed difference value is obtained according to a second preset weight control parameter and the difference value.

19. A task execution device, comprising: A first execution module, configured to execute, according to the feature to be processed, a plurality of convolution subtasks of the target task using a processor to obtain a plurality of convolution features, wherein the plurality of convolution features include a first convolution feature and a second convolution feature, and the second convolution feature includes a plurality of convolution subfeatures; a second execution module, configured to use the processor to perform the mask extraction subtask of the target task according to the multiple convolution subfeatures of the second convolution feature and at least one neighborhood subfeature of each of the multiple convolution subfeatures, so as to obtain a mask feature; a third execution module, configured to use the processor to execute the mask processing subtask of the target task according to the first convolution feature and the mask feature, so as to obtain a first masked feature; The first determination module is used to determine the output feature of the target task using the processor according to the first masked feature.

20. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 18.

21. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 18.

22. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 18.