Image task processing method and device, electronic equipment and storage medium
By dividing the compressed image feature sequence and comprehensive processing of target image features in the image task processing model, the problem of information redundancy in the compression process of image feature sequences is solved, and more efficient image feature compression and model processing efficiency are achieved.
Patent Information
- Application Number
- CN202510185937.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, image feature sequences still have high information redundancy during the compression process, resulting in poor compression effect, unable to effectively reduce the image task processing cost of the model, and affecting the user experience.
By introducing a target task processing layer into the image task processing model, dividing and processing the image feature sequence to be compressed, multiple image feature subsequences are obtained, and target image features are determined from each subsequence, and the target number of target image features is obtained, and the target image feature sequence to be compressed is updated to reduce information redundancy.
It effectively reduces the information redundancy of the image feature sequence, improves the compression quality of the image feature sequence, reduces the calculation amount and processing cost of the model, and improves the user experience.
Smart Images

Figure CN120220148A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to an image task processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the continuous development of large language model technologies, large vision language models (LVLMs) integrate visual processing capabilities into large language models (LLMs) to achieve cross-modal information processing and interaction. Although large vision language models perform well in image task processing, the cost of using these models in actual applications is relatively high. To reduce the processing cost of large vision language models for image tasks, usually, during the process of the model processing image tasks, the amount of model operations can be reduced by compressing the image feature (Token) sequences in some model layers.
[0003] In related image feature compression technologies, usually only the image features with high correlation with the model output features in the image feature sequence are retained, and the remaining image features are cropped; however, there may be some redundant information in the image features with high correlation with the model output features, resulting in a still high information redundancy in the compressed image feature sequence, poor compression effect of the image feature sequence, inability to effectively reduce the image task processing cost of the model, and further affecting the user experience. Summary of the Invention
[0004] The present disclosure provides an image task processing method, apparatus, electronic device, and storage medium to at least solve the technical problems in the related art that the compressed image feature sequence still has a high information redundancy, resulting in a poor compression effect of the image feature sequence, inability to effectively reduce the image task processing cost of the model, and further affecting the user experience. The technical solution of the present disclosure is as follows:
[0005] According to the first aspect of the embodiments of the present disclosure, an image task processing method is provided, including:
[0006] During the process of inputting the image feature sequence corresponding to the image to be processed into an image task processing model for image task processing, performing a partitioning process on the image feature sequence to be compressed output by the target task processing layer to obtain a target number of image feature subsequences; the image task processing model includes at least two sequentially connected task processing layers, and the target task processing layer is one of the at least two task processing layers except the last task processing layer;
[0007] Determining a target image feature from each of the image feature subsequences, and comprehensively obtaining the target number of the target image features;
[0008] Based on the target number of the target image features, update the image feature sequence to be compressed, and obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer. The input feature sequence is used to improve the processing efficiency of the image task of the image task processing model, and the image task is used for visual content understanding of the image to be processed.
[0009] In an alternative embodiment, the dividing the image feature sequence to be compressed output by the target task processing layer to obtain the target number of image feature subsequences includes:
[0010] Determine the preset compression index data;
[0011] Based on the preset compression index data, determine the sequence division length, and the sequence division length is positively correlated with the preset compression index data;
[0012] Average-divide the image feature sequence to be compressed according to the sequence division length to obtain the target number of the image feature subsequences, and the sequence length of each image feature subsequence is the sequence division length.
[0013] In an alternative embodiment, each image feature subsequence includes at least two image features to be compressed. Determining one target image feature from each image feature subsequence and comprehensively obtaining the target number of the target image features includes:
[0014] Determine the respective importance index data corresponding to at least two image features to be compressed in each image feature subsequence;
[0015] Use the image feature to be compressed with the largest value of the corresponding importance index data among the at least two image features to be compressed as the target image feature in each image feature subsequence;
[0016] Use the target image features of the target number of image feature subsequences as the target number of the target image features.
[0017] In an alternative embodiment, before the step of updating the image feature sequence to be compressed based on the target number of the target image features to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer, the method further includes:
[0018] Determine an associated feature set associated with the target number of the target image features from a plurality of non-target image features; the plurality of non-target image features are the image features to be compressed in the image feature sequence to be compressed except the target number of the target image features;
[0019] Merge the associated features in the set of associated features into the target number of the target image features to obtain the target number of compressed image features;
[0020] The updating of the sequence of image features to be compressed based on the target number of the target image features to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer includes:
[0021] Update the sequence of image features to be compressed based on the target number of the compressed image features to obtain the input feature sequence.
[0022] In an alternative embodiment, the determining of the set of associated features associated with the target number of the target image features from a plurality of non-target image features includes:
[0023] Perform a correlation analysis on each of the non-target image features and the target number of the target image features respectively to obtain the target number of associated index data corresponding to each of the non-target image features, and each of the associated index data respectively corresponds to one of the target number of the target image features;
[0024] Use the associated index data with the largest corresponding value among the target number of the associated index data as the target associated index data corresponding to each of the non-target image features;
[0025] Add the non-target image features among the plurality of non-target image features whose corresponding target associated index data satisfies a preset associated index threshold to the set of associated features.
[0026] In an alternative embodiment, the merging of the associated features in the set of associated features into the target number of the target image features to obtain the target number of compressed image features includes:
[0027] Based on the target associated index data corresponding to each associated feature in the set of associated features, determine the target image feature to which each of the associated features is to be merged from the target number of the target image features;
[0028] Based on the target image feature to which each of the associated features is to be merged, determine the associated feature to be merged corresponding to each of the target image features among the target number of the target image features;
[0029] Based on the important index data of the associated feature to be merged corresponding to each of the target image features, determine the first merging weight corresponding to the associated feature to be merged;
[0030] Based on the first merging weight, merge the associated feature to be merged corresponding to each of the target image features into the corresponding target image feature to obtain the compressed image feature corresponding to each of the target image features;
[0031] Use the compressed image features corresponding to each of the target number of target image features as the target number of compressed image features.
[0032] In an optional embodiment, before updating the image feature sequence to be compressed based on the target number of the target image features to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer, the method further includes:
[0033] Perform feature merging processing on the non-associated feature set to obtain compensated image features; the non-associated feature set is a set of non-target image features among the multiple non-target image features excluding the associated feature set.
[0034] The updating the image feature sequence to be compressed based on the target number of the target image features to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer includes:
[0035] Update the image feature sequence to be compressed based on the target number of the compressed image features and the compensated image features to obtain the input feature sequence.
[0036] In an optional embodiment, the performing feature merging processing on the non-associated feature set to obtain compensated image features includes:
[0037] Determine the second merging weight corresponding to each non-associated feature based on the importance index data corresponding to each non-associated feature in the non-associated feature set.
[0038] Perform merging processing on each non-associated feature based on the second merging weight to obtain the compensated image features.
[0039] In an optional embodiment, when the at least two layers of task processing layers include at least three layers of task processing layers, the target task processing layer is any task processing layer in the target compression processing layer, and the target compression processing layer includes: at least one first task processing layer and at least one second task processing layer. The at least one first task processing layer is at least one task processing layer between the first task processing layer and the intermediate task processing layer in the at least three layers of task processing layers. The at least one second task processing layer is at least one task processing layer between the intermediate task processing layer and the second-to-last task processing layer in the at least three layers of task processing layers. The intermediate task processing layer is the task processing layer located in the middle layer number in the at least three layers of task processing layers.
[0040] According to a second aspect of the embodiments of the present disclosure, there is provided an image task processing device, including:
[0041] A feature sequence partitioning module, configured to perform partitioning processing on the image feature sequence to be compressed output by a target task processing layer during the process of inputting the image feature sequence corresponding to the image to be processed into an image task processing model for image task processing, so as to obtain a target number of image feature subsequences; the image task processing model includes: at least two sequentially connected task processing layers, and the target task processing layer is one of the at least two task processing layers except the last task processing layer;
[0042] A target image feature determination module, configured to perform determining one target image feature from each of the image feature subsequences, and comprehensively obtain the target number of the target image features;
[0043] A compressed feature update module, configured to perform updating the image feature sequence to be compressed based on the target number of the target image features, so as to obtain an input feature sequence for the subsequent task processing layer connected to the target task processing layer, and the input feature sequence is used to improve the processing efficiency of the image task of the image task processing model, and the image task is used for visual content understanding of the image to be processed.
[0044] In an optional embodiment, the feature sequence partitioning module includes:
[0045] A compression index determination unit, configured to perform determining preset compression index data;
[0046] A sequence partitioning length determination unit, configured to perform determining a sequence partitioning length based on the preset compression index data, and the sequence partitioning length is positively correlated with the preset compression index data;
[0047] A sequence partitioning unit, configured to perform evenly partitioning the image feature sequence to be compressed according to the sequence partitioning length, so as to obtain the target number of the image feature subsequences, and the sequence length of each image feature subsequence is the sequence partitioning length.
[0048] In an optional embodiment, the target image feature determination module includes:
[0049] An important index data determination unit, configured to perform determining important index data corresponding to at least two image features to be compressed in each of the image feature subsequences;
[0050] A subsequence target feature determination unit, configured to perform using the one image feature to be compressed with the largest numerical value of the corresponding important index data among the at least two image features to be compressed as the target image feature in each of the image feature subsequences;
[0051] A target image feature unit, configured to perform using the target image features of each of the target number of image feature subsequences as the target number of target image features.
[0052] In an optional embodiment, the apparatus further includes:
[0053] An associated feature set determination module, configured to perform determining an associated feature set associated with the target number of the target image features from a plurality of non-target image features; the plurality of non-target image features are the image features to be compressed in the image feature sequence to be compressed except for the target number of the target image features.
[0054] A first feature merging module, configured to perform merging the associated features in the associated feature set into the target number of the target image features to obtain the target number of compressed image features.
[0055] The compressed feature update module includes:
[0056] A first update unit, configured to perform updating the image feature sequence to be compressed based on the target number of the compressed image features to obtain the input feature sequence.
[0057] In an optional embodiment, the associated feature set determination module includes:
[0058] An associated index data determination unit, configured to perform performing a correlation analysis on each of the non-target image features and the target number of the target image features respectively to obtain the target number of associated index data corresponding to each of the non-target image features, and each of the associated index data respectively corresponds to one of the target number of the target image features.
[0059] A target associated index data determination unit, configured to perform using the associated index data with the largest corresponding value among the target number of the associated index data as the target associated index data corresponding to each of the non-target image features.
[0060] An associated feature screening unit, configured to perform adding the non-target image features among the plurality of non-target image features whose corresponding target associated index data meets a preset associated index threshold to the associated feature set.
[0061] In an optional embodiment, the first feature merging module includes:
[0062] A target image feature to be merged determination unit, configured to perform determining the target image feature to be merged for each of the associated features from the target number of the target image features based on the target associated index data corresponding to each of the associated features in the associated feature set.
[0063] The target number of target image features to be merged correlation feature determination unit, which is configured to execute the target image features corresponding to each of the target image features to be merged based on the correlation feature, and determine the correlation feature to be merged corresponding to each of the target image features;
[0064] The first merging weight determination unit, which is configured to execute the important index data based on the correlation feature to be merged corresponding to each of the target image features, and determine the first merging weight corresponding to the correlation feature to be merged;
[0065] The correlation feature merging unit, which is configured to execute the merging of the correlation feature to be merged corresponding to each of the target image features to the corresponding target image feature based on the first merging weight, and obtain the compressed image feature corresponding to each of the target image features;
[0066] The compressed image feature generation unit, which is configured to execute the compressed image features corresponding to the target number of target image features respectively as the target number of compressed image features.
[0067] In an alternative embodiment, the apparatus further includes:
[0068] The second feature merging module, which is configured to execute the feature merging process on the non-correlation feature set to obtain the compensated image feature; the non-correlation feature set is a set of non-target image features other than the correlation feature set among the plurality of non-target image features;
[0069] The compressed feature update module includes:
[0070] The second update unit, which is configured to execute the update of the image feature sequence to be compressed based on the target number of the compressed image features and the compensated image feature, and obtain the input feature sequence.
[0071] In an alternative embodiment, the second feature merging module includes:
[0072] The second merging weight determination unit, which is configured to execute the important index data based on each non-correlation feature corresponding to the non-correlation feature set, and determine the second merging weight corresponding to each non-correlation feature;
[0073] The non-correlation feature merging unit, which is configured to execute the merging process on each non-correlation feature based on the second merging weight, and obtain the compensated image feature.
[0074] In an optional embodiment, when the at least two task processing layers include at least three task processing layers, the target task processing layer is any one of the task processing layers in the target compression processing layer, and the target compression processing layer includes: at least one first task processing layer and at least one second task processing layer. The at least one first task processing layer is at least one task processing layer between the first task processing layer and the middle task processing layer among the at least three task processing layers, and the at least one second task processing layer is at least one task processing layer between the middle task processing layer and the second-to-last task processing layer among the at least three task processing layers. The middle task processing layer is the task processing layer located in the middle layer among the at least three task processing layers.
[0075] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the method described in any one of the image task processing methods of the embodiments of the present disclosure.
[0076] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in any one of the image task processing methods of the embodiments of the present disclosure.
[0077] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product containing instructions, when it runs on a computer, enabling the computer to execute the method described in any one of the image task processing methods of the embodiments of the present disclosure.
[0078] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0079] In this application, one task processing layer except the last task processing layer in an image task processing model (including at least two sequentially connected task processing layers) is used as the target task processing layer. The image task processing model processes an image task for visual content understanding of an image to be processed. Then, during the process of the image task processing model performing image task processing based on the image feature sequence corresponding to the image to be processed, the to-be-compressed image feature sequence output by the target task processing layer is divided to obtain a target number of image feature subsequences, and a target image feature is determined from each image feature subsequence, and a target number of target image features are comprehensively obtained. By separately selecting target image features from different image feature subsequences, the spatial continuity of the target image features is broken, the information redundancy of the selected target image features is effectively reduced, and based on the target number of target image features, the to-be-compressed image feature sequence is updated to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer. On the basis of realizing the compression of the image feature sequence, the information redundancy in the image feature sequence can be reduced, the compression quality of the image feature sequence can be improved. By inputting the high-quality compressed image feature sequence into the subsequent task processing layer, the subsequent task processing layer performs task processing based on the high-quality compressed image feature sequence, which can effectively reduce the calculation amount and image task processing cost of the image task processing model, improve the processing efficiency of the image task of the image task processing model, and further improve the user experience.
[0080] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Brief Description of the Drawings
[0081] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0082] Figure 1 is a schematic diagram of an application environment shown according to an exemplary embodiment;
[0083] Figure 2 is a flowchart of an image task processing method shown according to an exemplary embodiment;
[0084] Figure 3 is a structural diagram of an image task processing model shown according to an exemplary embodiment;
[0085] Figure 4 is a flowchart of determining a target image feature from each image feature subsequence and comprehensively obtaining a target number of target image features shown according to an exemplary embodiment;
[0086] Figure 5 It is a flowchart of another image task processing method provided according to an exemplary embodiment;
[0087] Figure 6 It is a flowchart of determining an associated feature set associated with a target number of target image features from a plurality of non-target image features according to an exemplary embodiment;
[0088] Figure 7 It is a flowchart of merging the associated features in the associated feature set into a target number of target image features to obtain a target number of compressed image features according to an exemplary embodiment;
[0089] Figure 8 A flowchart of another image task processing method shown according to an exemplary embodiment;
[0090] Figure 9 It is a block diagram of an image task processing device shown according to an exemplary embodiment;
[0091] Figure 10 It is a block diagram of an electronic device for image task processing shown according to an exemplary embodiment;
[0092] Figure 11 It is a block diagram of another electronic device for image task processing shown according to an exemplary embodiment. Detailed implementation manners
[0093] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0094] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0095] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.
[0096] First, some nouns or terms that appear in the process of describing the embodiments of the present disclosure are applicable to the following explanations:
[0097] Large Language Model (LLM): It can also be referred to as a large model, natural language model, or large-scale language model, etc. It refers to a natural language processing model with a large number of parameters and training data. The training process of the large language model usually adopts the unsupervised learning method, that is, the model is trained through a large-scale text corpus to learn the probability distribution and language rules of the language. During the training process, the large language model usually uses the Language Model as the objective function, that is, the model parameters are optimized by maximizing the prediction probability of the next word.
[0098] Multimodal Large Language Model (MLLM): Based on the LLM, it integrates other modal media data (such as images, videos, audio, etc.), enabling the model to process information of different modalities simultaneously, better understand and express semantics, thereby improving the effect and accuracy of applications.
[0099] Large Vision-Language Model (LVLM): It is a technology based on multimodal deep learning, which is a large language model that deeply integrates and processes different visual and language tasks (such as image caption generation, visual question answering, multimodal content search, etc.). By effectively combining the pre-trained LLM and visual models, the LVLM can understand and generate text and visual content simultaneously, thereby realizing cross-modal information processing and interaction, and improving the performance of multimodal tasks in complex scenarios.
[0100] Transformer: An encoder-decoder architecture based on the attention mechanism.
[0101] Token: In the LLM, Token represents the smallest information unit that the model can understand and generate. In different contexts, the definition of Token may vary, but usually they are the basic units for multimodal data analysis and processing. Tokens are assigned numerical values or identifiers, arranged in sequences or vectors, and are input into or output from the model, and are the language components of the model.
[0102] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application environment shown according to an exemplary embodiment. The application environment may include a terminal 100 and a server 200.
[0103] Specifically, the terminal 100 may include, but is not limited to, electronic devices such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, etc., or may also be software running on the above-mentioned electronic devices, such as application programs. Optionally, the operating system running on the electronic device may include, but is not limited to, Android system, IOS system, Linux, Windows, etc. In an alternative embodiment, the terminal 100 may send the to-be-processed image corresponding to the image task to the server 200 in response to an image task for visual content understanding of the to-be-processed image.
[0104] Specifically, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or may also be a cloud server providing cloud computing services. In an alternative embodiment, the server 200 may perform feature compression on the to-be-compressed image feature sequence output by the target task processing layer during the process of inputting the image feature sequence corresponding to the to-be-processed image into the image task processing model for image task processing, so as to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer, thereby improving the processing efficiency of the image task of the image task processing model.
[0105] In addition, it should be noted that Figure 1 The application environment shown is only one provided by the present disclosure. In actual applications, there may also be other application environments, for example, there may be more terminals.
[0106] In the embodiments of this specification, the above-mentioned terminal 100 and the server 200 may be directly or indirectly connected through wired or wireless communication methods, and the present disclosure does not limit this.
[0107] Figure 2 is a flowchart of an image task processing method shown according to an exemplary embodiment. As Figure 2 shown, the method may include the following steps:
[0108] In step S201, during the process of inputting the image feature sequence corresponding to the to-be-processed image into the image task processing model for image task processing, a partitioning process is performed on the to-be-compressed image feature sequence output by the target task processing layer to obtain a target number of image feature subsequences; the image task processing model includes at least two layers of task processing layers connected in sequence, and the target task processing layer is one of the at least two layers of task processing layers except the last layer of task processing layer.
[0109] In the embodiments of this specification, the image task processing model can be an artificial intelligence model for processing image tasks. Specifically, the image tasks can be multi-modal tasks related to visual content understanding. Exemplarily, the image tasks can include, but are not limited to: image / video content description tasks, image / video question answering tasks, image / video content search tasks, etc.
[0110] In a specific embodiment, the image to be processed can be the image required for the above image tasks, that is, the image tasks are used for visual content understanding of the image to be processed. Optionally, the image to be processed can be a single image or a sequence of image frames in a video. Specifically, the image feature sequence can be used to represent the image features of the image to be processed. The image feature sequence can be a sequence of basic units obtained by discretizing the image features of the image to be processed (which can be processed by the image task processing model). The image feature sequence can be an image Token sequence, and each image feature in the image feature sequence can refer to an image Token.
[0111] In a specific embodiment, the image task processing model can adopt an autoregressive model with visual content understanding ability. Optionally, the autoregressive model here can be based on the Transformer architecture. In an alternative embodiment, the image task processing model can be a model that only processes image modality data. Correspondingly, the input data of the image task processing model can be the image feature sequence corresponding to the image to be processed, and the output data of the image task processing model can be the processing result of the image task corresponding to the image to be processed. In another alternative embodiment, the image task processing model can be a model that processes multi-modal data, such as the large language model in a multi-modal large language model. Correspondingly, in addition to the image feature sequence corresponding to the image to be processed, the input data of the image task processing model can also include: data feature sequences corresponding to other modality data, for example, text feature sequences corresponding to text (i.e., text Token sequences), audio feature sequences corresponding to audio (i.e., audio Token sequences). Optionally, in order to enable the image feature sequence corresponding to the image to be processed and the data feature sequences corresponding to other modality data to be processed in the same image task processing model, the image feature sequence corresponding to the image to be processed and the data feature sequences corresponding to other modality data can be aligned to the same semantic space (i.e., represented in the same vector space), and then input into the image task processing model.
[0112] Exemplarily, taking a multimodal large language model as an example of a large vision language model, the large vision language model may include: an image encoder, a text encoder, a graphic-text feature alignment model, and an image task processing model; the image encoder may be used to extract image feature information from the image to be processed; the text encoder may be used to extract text feature information from the input text (for example, system prompt text, user instruction text); the graphic-text feature alignment model may be used to semantically align the image feature information with the text feature information, and discretize the aligned feature information to obtain an image feature sequence and a text feature sequence; the image task processing model may be used to perform image task processing based on the image feature sequence and the text feature sequence.
[0113] In a specific embodiment, the image task processing model may be sequentially connected by at least two identical task processing layers. Each task processing layer may be used to extract the global dependency relationship between each feature in the input feature sequence of itself, and perform feature mapping on the input feature sequence based on the global dependency relationship to obtain an output feature sequence, so as to input the output feature sequence into the next task processing layer. Optionally, when the image task processing model only processes image modality data, the output feature sequence may be an image feature sequence; when the image task processing model processes multimodal data, the output feature sequence may include: an image feature sequence and a data feature sequence corresponding to other modality data
[0114] In an alternative embodiment, the image task processing model may be a large language model, such as Figure 3 shown, each task processing layer may include: an attention sublayer (for example, a self-attention sublayer) and a feed-forward neural network sublayer. Specifically, the attention sublayer may be used to extract the attention metric data corresponding to each input feature in the input feature sequence. The attention metric data corresponding to each input feature may be used to measure the feature correlation between the target metric feature and the corresponding input feature. For example, the manifestation form of the attention metric data may be an attention score, where the target metric feature may be an input feature selected from the input feature sequence. In an alternative embodiment, the attention sublayer may adopt a multi-head attention mechanism; the feed-forward neural network sublayer may be used to perform feature mapping on the input feature sequence based on the attention metric data to obtain an output feature sequence.
[0115] In a specific embodiment, at least one task processing layer can be selected from the task processing layers of the image task processing model except the last task processing layer as the target compression processing layer. Correspondingly, the image feature sequence in the output feature sequence of the target task processing layer (any task processing layer in the target compression processing layer) is used as the image feature to be compressed, and the image feature sequence to be compressed output by the target task processing layer is compressed to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer, thereby improving the image task processing efficiency of the image task processing model by reducing the information redundancy of the image feature sequence. It can be understood that the output feature sequence of the non-target task processing layer (any task processing layer in the image task processing model except the target compression processing layer) is directly used as the input feature sequence of the subsequent task processing layer connected to the non-target task processing layer. Schematically, the image task processing model includes N task processing layers, N≥2, and the target compression processing layer can be at least one task processing layer between the first task processing layer and the N-1th task processing layer. Exemplarily, N = 10, and the target compression processing layer can include: the third task processing layer and the seventh task processing layer. Correspondingly, the image feature sequence to be compressed in the output feature sequence of the third task processing layer can be compressed to obtain the input feature sequence of the fourth task processing layer, and the image feature sequence to be compressed in the output feature sequence of the seventh task processing layer can be compressed to obtain the input feature sequence of the eighth task processing layer, while the output feature sequence of each task processing layer in the 1-2, 4-6, and 8-9 task processing layers is directly used as the input feature sequence of the subsequent task processing layer connected to itself.
[0116] In a specific embodiment, the image feature sequence to be compressed can be a sequence composed of image features output by the target task processing layer. In practical applications, since images present continuous values in the color space, the redundancy of the image feature sequence to be compressed is mainly reflected in the spatially continuous image features. Therefore, in the embodiments of the present disclosure, the image feature sequence to be compressed is divided into several image feature subsequences, and important target image features are selected from each image feature subsequence to break the spatial continuity of the target image features, thereby reducing the redundancy of the selected target image features.
[0117] In a specific embodiment, the above-mentioned division processing of the image feature sequence to be compressed output by the target task processing layer to obtain a target number of image feature subsequences may include:
[0118] In step S2011, preset compression index data is determined.
[0119] In step S2012, based on the preset compression index data, the sequence division length is determined, and the sequence division length is positively correlated with the preset compression index data.
[0120] Specifically, the preset compression index data can represent the feature compression degree required for the image feature sequence to be compressed. Generally, the preset compression index data is positively correlated with the sequence division length. The larger the preset compression index data, the larger the sequence division length. Correspondingly, the number of image feature subsequences obtained by dividing the image feature sequence to be compressed (i.e., the target number) is smaller. Optionally, the preset compression index data can be set in combination with the model calculation cost reduction requirements in practical applications. Schematically, the manifestation form of the preset compression index data can be a preset compression rate, and the value range of the preset compression rate can be expressed as (0, 1).
[0121] In an optional embodiment, determining the sequence division length based on the preset compression index data can be expressed by the following formula: R represents the preset compression index data, and L represents the sequence division length.
[0122] In step S2013, the image feature sequence to be compressed is evenly divided according to the sequence division length to obtain the target number of image feature subsequences, and the sequence length of each image feature subsequence is the sequence division length.
[0123] In a specific embodiment, the target number can be the number of image feature subsequences obtained by evenly dividing the image feature sequence to be compressed according to the sequence division length. Schematically, the target number can be expressed by the following formula:
[0124] Where N represents the sequence length of the image feature sequence to be compressed (i.e., the total number of image features to be compressed included in the image feature sequence to be compressed), and N t represents the target number.
[0125] Exemplarily, when the preset compression index data R = 75%, the length of each image feature subsequence is 4, which means that every 4 image features form an image feature subsequence, and 1 image feature is selected from every 4 image features as the target image feature.
[0126] In the above embodiments, based on a preset feature compression index, the sequence division length is determined, and the image feature sequence to be compressed is evenly divided according to the sequence division length to obtain a target number of image feature subsequences. By separately selecting target image features from different image feature subsequences, the spatial continuity of the target image features is broken, the information redundancy of the selected target image features is effectively reduced, the balance between feature importance and information redundancy in the process of screening target image features is achieved, and it is applicable to image processing tasks in different scenarios, thereby improving the compression quality of the subsequent image feature sequence.
[0127] In step S202, a target image feature is determined from each image feature subsequence, and a target number of target image features are obtained comprehensively.
[0128] Specifically, each image feature subsequence includes at least two image features to be compressed, and the target image feature in each image feature subsequence can be the most important one of the at least two image features to be compressed.
[0129] In a specific embodiment, as Figure 4 shown, each image feature subsequence includes at least two image features to be compressed. The above-mentioned determining a target image feature from each image feature subsequence and comprehensively obtaining a target number of target image features includes:
[0130] In step S2021, the important index data corresponding to each of the at least two image features to be compressed in each image feature subsequence is determined.
[0131] Specifically, the important index data corresponding to each image feature to be compressed can be used to represent the importance degree of each image feature to be compressed in the corresponding image feature subsequence. In an optional embodiment, the important index data may include, but is not limited to, attention index data, feature similarity index data, etc. Specifically, the attention index data corresponding to each image feature to be compressed can be used to measure the feature attention degree of the target metric feature to the corresponding image feature to be compressed, and the feature similarity index data corresponding to each image feature to be compressed can be used to measure the feature similarity between the target metric feature and the corresponding image feature to be compressed.
[0132] In a specific embodiment, the target metric feature can be a feature (i.e., Token) used to measure the importance of the features of the image to be compressed. Specifically, the target metric feature can be a feature selected from the output feature sequence of the target task processing layer. In an alternative embodiment, when the image processing task is a classification task, an initial category feature corresponding to the classification identifier (CLS) can be added in front of the input data of the image task processing model. Usually, the category feature marked by the classification identifier can be used to represent the global feature of the input data, so as to perform classification prediction on the input data. Correspondingly, the output feature sequence of the target task processing layer can include: the category feature (CLS Token) and the image feature sequence to be compressed. Taking the category feature as the target metric feature, the importance of each image feature to be compressed is measured by the category feature. In another alternative embodiment, when the image task processing model is a multimodal model, the output feature sequence of the target task processing layer can include: the text feature and the image feature sequence to be compressed. Taking the text feature as the target metric feature, the importance of each image feature to be compressed is measured by the text feature.
[0133] Schematically, the image task processing model is a large language model in a multimodal large language model. The input of the multimodal large language model can include: the above-mentioned image to be processed, the system prompt text (System Prompt) and the user instruction text (Instruction). The system prompt text can refer to the general prompt text for the system to control the behavior of the large language model, which is usually determined in the instruction tuning stage of the large language model. The user instruction text can refer to the instruction description text input by the user for the image to be processed. Correspondingly, the input data of the image task processing model can include: the above-mentioned image feature sequence, the initial system prompt feature corresponding to the system prompt text, the initial user instruction feature corresponding to the user instruction text, and the initial model output feature. In addition, when the image processing task corresponding to the image task processing model is a classification task, an initial category feature corresponding to the classification identifier (CLS) can also be added in front of the input data of the image task processing model. Correspondingly, the output feature sequence of the target task processing layer can include: the category feature (CLS Token), the image feature sequence to be compressed, the system prompt feature (System Prompt Token), the user instruction feature (Instruction Token), and the model output feature (OutputToken). The target metric feature can be one of the category feature, the system prompt feature, the user instruction feature, and the model output feature. Specifically, the selection of the target metric feature can be set in combination with the feature attention requirements corresponding to the image processing task in the actual application.
[0134] In an alternative embodiment, determining the respective importance metric data corresponding to at least two to-be-compressed image features in each image feature subsequence may include:
[0135] Determining the attention metric data of the target metric feature for each to-be-compressed image feature;
[0136] Taking the attention metric data corresponding to each to-be-compressed image feature as the importance metric data corresponding to each to-be-compressed image feature.
[0137] In an alternative embodiment, when the importance metric data is attention metric data, the attention metric data of the target metric feature for each to-be-compressed image feature may be expressed by the following formula:
[0138] where Q represents the target metric feature, K i represents the i-th to-be-compressed image feature in the to-be-compressed image feature sequence, N represents the sequence length of the to-be-compressed image feature sequence (i.e., the total number of to-be-compressed image features included in the to-be-compressed image feature sequence), D represents the feature dimension, and Softmax() represents the normalization process.
[0139] In another alternative embodiment, determining the respective importance metric data corresponding to at least two to-be-compressed image features in each image feature subsequence may further include:
[0140] Determining the feature similarity metric data between the target metric feature and each to-be-compressed image feature;
[0141] Taking the feature similarity metric data corresponding to each to-be-compressed image feature as the importance metric data corresponding to each to-be-compressed image feature.
[0142] In an alternative embodiment, the feature similarity metric data here may include, but is not limited to, cosine similarity, etc.
[0143] In the above embodiment, when the output feature sequence of the target task processing layer further includes: class features, system prompt features, user instruction features, and model output features, one of the class features, system prompt features, user instruction features, and model output features may be selected as the target metric feature, and the importance metric data corresponding to the to-be-compressed image features is calculated based on this target metric feature, so as to screen out the target image features in each image feature subsequence. By improving the flexibility of target metric feature selection, the flexibility of target image feature screening can be effectively improved to meet the feature compression requirements of different image processing tasks, thereby ensuring the user experience.
[0144] In step S2022, one of the at least two image features to be compressed with the largest value of the corresponding important index data is used as the target image feature in each image feature subsequence.
[0145] Optionally, when the important index data is attention index data, the target image feature in each image feature subsequence can be one of the image features to be compressed with the largest value of the attention index data in each image feature subsequence; when the important index data is feature similarity index data, the target image feature in each image feature subsequence can be one of the image features to be compressed with the largest value of the feature similarity index data in each image feature subsequence.
[0146] In step S2023, the target image features of the target number of image feature subsequences are used as the target number of target image features.
[0147] In the above embodiment, by dividing the image feature sequence to be compressed into several image feature subsequences and selecting one of the image features to be compressed with the largest value of the corresponding important index data in each image feature subsequence as the target image feature in each image feature subsequence, the target number of target image features in the image feature sequence to be compressed is screened out, breaking the spatial continuity between the selected target image features, effectively reducing the information redundancy of the target image features, achieving the balance between feature importance and information redundancy in the target image feature screening process, and further improving the compression quality of the subsequent image feature sequence.
[0148] In step S203, based on the target number of target image features, the image feature sequence to be compressed is updated to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer. The input feature sequence is used to improve the processing efficiency of the image task of the image task processing model, and the image task is used for the visual content understanding of the image to be processed.
[0149] Specifically, the image feature sequence to be compressed output by the target task processing layer can be replaced with the target number of target image features to obtain the input feature sequence of the subsequent task processing layer of the target task processing layer, so that the subsequent task processing layer performs tasks based on the high-quality compressed image feature sequence, which can effectively reduce the computational amount and the image task processing cost of the image task processing model and improve the processing efficiency of the image task of the image task processing model.
[0150] In the above embodiments, one task processing layer except the last task processing layer in the image task processing model (including at least two sequentially connected task processing layers) is used as the target task processing layer. The image task processing model processes the image task for visual content understanding of the image to be processed. Then, during the process of the image task processing model performing image task processing based on the image feature sequence corresponding to the image to be processed, the to-be-compressed image feature sequence output by the target task processing layer is partitioned to obtain a target number of image feature subsequences, and a target image feature is determined from each image feature subsequence, and a target number of target image features are comprehensively obtained. By separately selecting target image features from different image feature subsequences, the spatial continuity of the target image features is broken, the information redundancy of the selected target image features is effectively reduced, and based on the target number of target image features, the to-be-compressed image feature sequence is updated to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer. On the basis of realizing the compression of the image feature sequence, the information redundancy in the image feature sequence can be reduced, the compression quality of the image feature sequence can be improved. By inputting the high-quality compressed image feature sequence into the subsequent task processing layer, the subsequent task processing layer performs task processing based on the high-quality compressed image feature sequence, which can effectively reduce the calculation amount and image task processing cost of the image task processing model, improve the processing efficiency of the image task of the image task processing model, and further improve the user experience.
[0151] In an alternative embodiment, as Figure 5 shown, before updating the to-be-compressed image feature sequence based on the target number of target image features to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer, the above method further includes:
[0152] In step S204, an association feature set associated with the target number of target image features is determined from a plurality of non-target image features; the plurality of non-target image features are the to-be-compressed image features in the to-be-compressed image feature sequence except the target number of target image features.
[0153] In the embodiments of this specification, the plurality of non-target image features may be the to-be-compressed image features in the to-be-compressed image feature sequence except the target number of target image features. The association feature set may be a set of non-target image features screened from the plurality of non-target image features and associated with the target image features.
[0154] In a specific embodiment, as Figure 6 shown, determining the association feature set associated with the target number of target image features from the plurality of non-target image features includes:
[0155] In step S2041, correlation analysis is performed between each non-target image feature and the target number of target image features respectively to obtain the target number of associated index data corresponding to each non-target image feature, and each associated index data corresponds to one of the target number of target image features respectively.
[0156] In a specific embodiment, the associated index data between each non-target image feature and the target number of target image features is used to measure the feature correlation degree between each non-target image feature and the target number of target image features. Optionally, the associated index data may include, but is not limited to, cosine similarity, etc. The above-mentioned performing correlation analysis between each non-target image feature and the target number of target image features respectively to obtain the target number of associated index data corresponding to each non-target image feature may include: calculating the cosine similarity between each non-target image feature and the target number of target image features respectively.
[0157] In a specific embodiment, the target number of associated index data corresponding to each non-target image feature corresponds to one of the target number of target image features respectively, that is, each associated index data corresponding to each non-target image feature may be the associated index data between each non-target image feature and one target image feature, and each associated index data may be used to characterize the feature correlation degree between the corresponding non-target image feature and one target image feature.
[0158] In step S2042, the associated index data with the largest corresponding value among the target number of associated index data is used as the target associated index data corresponding to each non-target image feature.
[0159] Optionally, when the associated index data is cosine similarity, the target associated index data corresponding to each non-target image feature may be the cosine similarity with the largest corresponding value among the target number of cosine similarities corresponding to each non-target image feature.
[0160] In step S2043, the non-target image features among the multiple non-target image features whose corresponding target associated index data meets the preset associated index threshold are added to the associated feature set.
[0161] Specifically, the preset correlation index threshold can be the correlation index threshold for screening correlation features. The preset correlation index threshold can be set in combination with the information retention requirements for non-target image features in practical applications. Generally, the smaller the preset correlation index threshold, the more correlation features are screened out from the non-target image features, and the more non-target image features are retained subsequently (i.e., the non-target image features merged into the target image features); the larger the preset correlation index threshold, the fewer correlation features are screened out from the non-target image features, and the fewer non-target image features are retained subsequently. Schematically, the value range of the preset correlation index threshold can be (0, 1). For example, the preset correlation index threshold is taken as 0.7.
[0162] Schematically, a certain non-target image feature is represented as v g , the number of targets is represented as N t , and the number of target image features of the number of targets is represented as Correspondingly, the non-target image feature v g is respectively subjected to correlation analysis with the number of target image features of the number of targets, and the number of target correlation index data obtained can be represented as Take the correlation index data with the largest corresponding value in g as the target correlation index data S of the non-target image feature v max . In the case that the target correlation index data is greater than the preset correlation index threshold, the non-target image feature v g is added to the correlation feature set.
[0163] In the above embodiment, each non-target image feature is respectively subjected to correlation analysis with the number of target image features of the number of targets, the number of target correlation index data corresponding to each non-target image feature is obtained, and the correlation index data with the largest corresponding value in the number of target correlation index data is used as the target correlation index data corresponding to each non-target image feature, so that the non-target image features whose target correlation index data meets the preset correlation index threshold are added to the correlation feature set, which can improve the accuracy of screening correlation features in non-target image features and avoid damaging important information in the target image features and affecting the quality of image feature compression in the subsequent process of merging the correlation features into the target image features.
[0164] In step S205, the correlation features in the correlation feature set are merged into the number of target image features of the number of targets to obtain the number of compressed image features of the number of targets.
[0165] In the embodiments of this specification, the number of the above-mentioned compressed image features of the number of targets can be the number of image features of the number of targets obtained by merging the correlation features in the correlation feature set into the number of target image features of the number of targets.
[0166] In a specific embodiment, as Figure 7 shown, the merging of the associated features in the associated feature set into the target number of target image features to obtain the target number of compressed image features may include:
[0167] In step S2051, based on the target association index data corresponding to each associated feature in the associated feature set, the target image feature to be merged for each associated feature is determined from the target number of target image features.
[0168] Specifically, the target image feature to be merged for each associated feature may be the target image feature that each associated feature needs to be incorporated into. In a specific embodiment, the target image feature with the largest value of the association index data between each associated feature and the target number of target image features may be used as the target image feature to be merged for each associated feature, that is, the target image feature that is most similar to a certain associated feature among the target number of target image features is used as the target image feature to be merged for this associated feature, so as to minimize the damage to the information contained in the target image features caused by feature merging.
[0169] In step S2052, based on the target image feature to be merged for each associated feature, the associated feature to be merged corresponding to each target image feature among the target number of target image features is determined.
[0170] Schematically, if the target image feature to be merged for the associated feature a is the target image feature b, then the associated feature a can be called the associated feature to be merged corresponding to the target image feature b.
[0171] In step S2053, based on the important index data of the associated feature to be merged corresponding to each target image feature, the first merging weight corresponding to the associated feature to be merged is determined.
[0172] Specifically, the important index data corresponding to the associated feature to be merged is used to characterize the importance of the associated feature to be merged in the sequence of image features to be compressed. For specific details, reference can be made to the refinement content of step S2021, which will not be elaborated here. In a specific embodiment, the first merging weight corresponding to the associated feature to be merged is positively correlated with the important index data of the associated feature to be merged. Generally, the larger the important index data corresponding to the associated feature to be merged, the larger the corresponding first merging weight.
[0173] In an alternative embodiment, the important index data of the associated feature to be merged corresponding to each target image feature may be normalized to obtain the first merging weight of the associated feature to be merged corresponding to each target image feature. Schematically, the normalization process here may use the Softmax function.
[0174] In step S2054, based on the first merging weight, the associated features to be merged corresponding to each target image feature are merged into the corresponding target image feature to obtain the compressed image feature corresponding to each target image feature.
[0175] Specifically, based on the first merging weight of the associated features to be merged corresponding to each target image feature, the associated features to be merged corresponding to each target image feature are weighted and added to the corresponding target image feature to obtain the compressed image feature corresponding to each target image feature.
[0176] In step S2055, the compressed image features corresponding to each of the target number of target image features are used as the target number of compressed image features.
[0177] In an alternative embodiment, the correlation index data (cosine similarity) between the target number (N t ) of target image features and the N to-be-compressed image features in the to-be-compressed image feature sequence may also be calculated to obtain a correlation index matrix: Then, the target correlation index data (highest similarity) between each to-be-compressed image feature and the N t target image features is calculated, and the to-be-compressed image features with the highest similarity greater than a preset correlation index threshold (preset similarity threshold) among the N to-be-compressed image features are added to the positive set; a first matching matrix with the same dimension as the correlation index matrix is constructed If the cosine similarity between a to-be-compressed image feature and a certain target image feature is the preset similarity threshold of the to-be-compressed image feature, the element in the first matching matrix corresponding to the to-be-compressed image feature and the certain target image feature is set to 1, the elements in the first matching matrix corresponding to the to-be-compressed image feature and other target image features are set to 0, then the elements with 0 in the first matching matrix are replaced with negative infinity, and the non-zero elements in the first matching matrix are replaced with the important index data of the corresponding to-be-compressed image feature to obtain a second matching matrix:
[0178] M′ ij =-∞(1 - M ij ) + α j ×M ij
[0179] Then, the normalization function (e.g., Softmax function) is used to model the first merging weight for each row of matrix elements in the second matching matrix: where ρ represents the temperature coefficient; finally, the features in the to-be-compressed image feature sequence (V ∈ R N×D ) are merged to obtain the target number of compressed image features
[0180] V′ = W × V
[0181] Wherein, W represents the first merging weight matrix, V represents the image feature sequence to be compressed, and V′ represents the compressed image features of the target number.
[0182] In the above embodiments, based on the association index data between each association feature in the association feature set and the target number of target image features, the target image feature to be merged for each association feature is determined from the target number of target image features, and based on the target image features to be merged for each association feature, the association feature to be merged corresponding to each target image feature among the target number of target image features is determined, so as to merge each association feature in the association feature set into its most relevant target image feature, and then based on the importance index data of the association feature to be merged corresponding to each target image feature, the first merging weight corresponding to the association feature to be merged is determined, and thus based on the first merging weight, the association feature to be merged is merged into the corresponding target image feature, and the compressed image feature corresponding to each target image feature is obtained. By avoiding the introduction of noise information, finally, the target number of compressed image features are obtained from the compressed image features corresponding to the target number of target image features respectively, which can reduce the loss of important information in the association features while avoiding the destruction of important information in the target image features during feature merging on the basis of realizing image feature compression, thereby improving the compression quality of the image feature sequence, effectively reducing the calculation amount and image task processing cost of the image task processing model, and further improving the user experience.
[0183] Correspondingly, the above-mentioned updating the image feature sequence to be compressed based on the target number of target image features to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer includes:
[0184] In step S2031, the image feature sequence to be compressed is updated based on the target number of compressed image features to obtain the input feature sequence.
[0185] Specifically, the image feature sequence to be compressed output by the target task processing layer can be replaced with the target number of compressed image features to obtain the input feature sequence of the subsequent task processing layer of the target task processing layer, so that the subsequent task processing layer performs task processing based on the high-quality compressed image feature sequence, which can effectively reduce the calculation amount and image task processing cost of the image task processing model and improve the processing efficiency of the image task of the image task processing model.
[0186] In the above embodiments, by separately selecting target image features from different image feature subsequences to break the spatial continuity of the target image features, the information redundancy of the selected target image features is effectively reduced; then, an associated feature set (non-target image features related to the target image features) is screened out from multiple non-target image features of the image feature sequence to be compressed, and the associated features in the associated feature set are merged into the target number of target image features to obtain the target number of compressed image features. Based on the target number of compressed image features, the image feature sequence to be compressed is updated to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer. On the basis of realizing the compression of the image feature sequence, the information redundancy in the image feature sequence can be reduced, and at the same time, the loss of important feature information in the non-target image features can be reduced, improving the compression quality of the image feature sequence. By inputting the high-quality compressed image feature sequence into the subsequent task processing layer, the subsequent task processing layer can perform task processing based on the high-quality compressed image feature sequence, effectively reducing the computational amount and image task processing cost of the image task processing model, improving the processing efficiency of the image tasks of the image task processing model, and further improving the user experience.
[0187] In a specific embodiment, as Figure 8 shown, before updating the image feature sequence to be compressed based on the target number of target image features to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer, the above method further includes:
[0188] In step S206, the non-associated feature set is subjected to feature merging processing to obtain compensated image features; the non-associated feature set is a set of non-target image features other than the associated feature set among the multiple non-target image features.
[0189] Specifically, the non-associated feature set can be a set of non-target image features screened out from multiple non-target image features that are non-associated with the target image features. In an alternative embodiment, the non-associated feature set can be a set of non-target image features screened out from multiple non-target image features based on the association index data between each non-target image feature and the target number of target image features. Correspondingly, the associated features can be non-target image features corresponding to the target association index data that meet the preset association index threshold, and the non-associated features can be non-target image features corresponding to the target association index data that do not meet the preset association index threshold.
[0190] Specifically, the compensated image features can be information compensation features obtained by merging non-associated features (i.e., non-target image features that have not been merged into the target image features), and the compensated image features are used to represent the important information contained in the non-associated feature set.
[0191] Optionally, step S205 and step S206 can be executed in parallel, that is, the merging process of associated features and the merging process of non-associated features in the non-target image features are executed in parallel, thereby improving the compression efficiency of the image feature sequence.
[0192] In a specific embodiment, the above-mentioned feature merging process for the non-associated feature set to obtain the compensated image features includes:
[0193] In step S2061, based on the important index data corresponding to each non-associated feature in the non-associated feature set, the second merging weight corresponding to each non-associated feature is determined.
[0194] Specifically, the important index data corresponding to each non-associated feature is used to characterize the importance of each non-associated feature in the image feature sequence to be compressed. For specific details, refer to the refinement content of step S2021, which will not be elaborated here. In a specific embodiment, the second merging weight corresponding to each non-associated feature is positively correlated with the important index data corresponding to each non-associated feature. Generally, the larger the important index data corresponding to each non-associated feature, the larger the corresponding second merging weight.
[0195] In an alternative embodiment, the important index data corresponding to each non-associated feature can be normalized to obtain the second merging weight corresponding to each non-associated feature. Schematically, the normalization process here can use the Softmax function.
[0196] In step S2062, based on the second merging weight, each non-associated feature is merged to obtain the compensated image features.
[0197] Specifically, based on the second merging weight corresponding to each non-associated feature, each non-associated feature in the non-associated feature set is weighted and added to obtain the compensated image features.
[0198] It can be understood that since the number of targets corresponding to the target image features is usually much larger than 1, the increase in the model inference calculation amount brought by adding a compensated image feature can be ignored.
[0199] In the above embodiment, based on the important index data corresponding to each non-associated feature in the non-associated feature set, the second merging weight corresponding to each non-associated feature can be determined, and based on the second merging weight, the non-associated features in the non-associated feature set are merged to obtain the compensated image features. By improving the accuracy of the second merging weight corresponding to the non-associated features, the accuracy of the non-associated feature merging is improved, thereby effectively retaining the important information in the non-associated features.
[0200] Correspondingly, the above-mentioned method for updating the image feature sequence to be compressed based on the target number of target image features to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer includes:
[0201] In step S2032, based on the target number of compressed image features and compensation image features, the image feature sequence to be compressed is updated to obtain the input feature sequence.
[0202] Specifically, the image feature sequence to be compressed output by the target task processing layer can be replaced with the target number of compressed image features and compensation image features to obtain the input feature sequence of the subsequent task processing layer of the target task processing layer, so that the subsequent task processing layer performs task processing based on the high-quality compressed image feature sequence, which can effectively reduce the computational complexity and image task processing cost of the image task processing model and improve the processing efficiency of the image tasks of the image task processing model.
[0203] In the above embodiment, the non-target image features in the image feature sequence to be compressed that are not associated with the target image features are merged into a compensation image feature, and based on the target number of compressed image features and the compensation image feature, the image feature sequence to be compressed is updated to obtain the input feature sequence, which can avoid destroying the important information contained in the target image features and prevent the loss of important information in the non-associated features, thereby further improving the compression quality of the image feature sequence, reducing the computational complexity and image task processing cost of the image task processing model, and further improving the user experience.
[0204] In practical applications, the image task processing model under the autoregressive model architecture mainly includes two stages of inference calculation when processing image tasks, namely the prefill stage and the decode stage, and these two stages are executed serially.
[0205] The prefill stage includes the following processing: the image task processing model performs global dependency analysis on the input feature sequence (for example, the Token sequence), generates the attention information of the feature sequence, caches the attention information, and generates the Token of the first output content (for example, the output text) of the image task processing model. The context information (i.e., the attention information) inferred and calculated in the prefill stage is used in the decode stage. When the image task processing model includes N task processing layers, the prefill stage needs to sequentially execute the calculations of the N task processing layers to obtain the context information in each task processing layer respectively.
[0206] The decoding stage refers to the process in which the image task processing model actually generates the content to be output. In this stage, given the context information in the prefill stage, the model inference continues to run, and after passing through the calculations of N task processing layers again, the tokens of the subsequent content to be output are generated one by one. Each token of the output content needs to be obtained through the calculations of N task processing layers until the inference ends.
[0207] Since the prefill stage is mainly for globally analyzing the input feature sequence of the image task processing model, and the decoding stage mainly continues to run the model inference based on the context information in the prefill stage. When the image feature sequence in the input feature sequence has a high degree of redundancy, it usually affects the model calculation amount and model calculation cost in the prefill stage. Therefore, the image task processing method provided by the disclosed embodiments of the present application can be executed in the prefill stage of the image task processing model.
[0208] In practical applications, as the number of layers of the task processing layer deepens, the redundancy of the image features in the output feature sequence of the task processing layer will be higher. If image feature compression is only performed in the shallow task processing layer, the image features in the deep task processing layer still have redundancy. To further improve the compression efficiency of the image features, at least one shallow task processing layer and at least one deep task processing layer can be respectively selected for image feature compression, reducing the number of image features layer by layer.
[0209] In a specific embodiment, when at least two task processing layers include at least three task processing layers, the target task processing layer is any task processing layer in the target compression processing layer. The target compression processing layer includes: at least one first task processing layer and at least one second task processing layer. At least one first task processing layer is at least one task processing layer between the first task processing layer and the middle task processing layer in the at least three task processing layers. At least one second task processing layer is at least one task processing layer between the middle task processing layer and the penultimate task processing layer in the at least three task processing layers. The middle task processing layer is the task processing layer located in the middle layer number in the at least three task processing layers.
[0210] Specifically, at least one first task processing layer can be at least one shallow task processing layer in the image task processing model, and at least one second task processing layer can be at least one deep task processing layer in the image task processing model.
[0211] In the above embodiments, at least one shallow task processing layer can be selected between the first-layer task processing layer and the middle-layer task processing layer, and at least one deep task processing layer can be selected between the middle-layer task processing layer and the penultimate-layer task processing layer. Image feature compression is performed respectively in the shallow task processing layer and the deep task processing layer, reducing the number of image features layer by layer, further improving the compression efficiency of the image features, effectively reducing the computational amount and the image task processing cost of the image task processing model, and thus enhancing the user experience.
[0212] In an alternative embodiment, the image feature compression mechanism involved in the image task processing method provided in the embodiments of the present disclosure can also be configured into an image feature compression plug-in, and the image feature compression plug-in is controlled to execute between any two adjacent task processing layers in the image task processing model. The image feature compression plug-in is controlled to perform feature compression on the image feature sequence to be compressed in the output feature sequence of the previous task processing layer in the two adjacent layers, so as to obtain the input feature sequence of the subsequent task processing layer, thereby improving the image task processing efficiency of the image task processing model.
[0213] As can be seen from the technical solutions provided in the embodiments of this specification above, in this specification, one task processing layer except the last task processing layer in an image task processing model (including at least two sequentially connected task processing layers) is used as the target task processing layer. The image task processing model processes an image task for visual content understanding of a to-be-processed image. Then, during the process of the image task processing model performing image task processing based on the image feature sequence corresponding to the to-be-processed image, the to-be-compressed image feature sequence output by the target task processing layer is partitioned to obtain a target number of image feature subsequences, and a target image feature is determined from each image feature subsequence, and a target number of target image features are comprehensively obtained. By separately selecting target image features from different image feature subsequences, the spatial continuity of the target image features is broken, the information redundancy of the selected target image features is effectively reduced, and based on the target number of target image features, the to-be-compressed image feature sequence is updated to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer, which can reduce the information redundancy in the image feature sequence and improve the compression quality of the image feature sequence on the basis of realizing the compression of the image feature sequence; in addition, an associated feature set (non-target image features related to the target image features) can be screened out from multiple non-target image features of the to-be-compressed image feature sequence, and the associated features in the associated feature set are merged into the target number of target image features to obtain the target number of compressed image features, and based on the target number of compressed image features, the to-be-compressed image feature sequence is updated to obtain the input feature sequence of the subsequent task processing layer connected to the target task processing layer, which can reduce the loss of important feature information in the non-target image features while reducing the information redundancy in the image feature sequence; in addition, the non-target image features that are not associated with the target image features in the to-be-compressed image feature sequence can be merged into a compensation image feature, and based on the target number of compressed image features and the compensation image feature, the to-be-compressed image feature sequence is updated to obtain the input feature sequence, which can avoid destroying the important information contained in the target image features and prevent the loss of important information in the non-associated features. By inputting the high-quality compressed image feature sequence into the subsequent task processing layer, the subsequent task processing layer performs task processing based on the high-quality compressed image feature sequence, which can effectively reduce the calculation amount and the image task processing cost of the image task processing model and improve the processing efficiency of the image task of the image task processing model; in addition, image feature compression can be performed respectively in the shallow task processing layer and the deep task processing layer, and the number of image features is gradually reduced layer by layer, which can further improve the compression efficiency of the image feature sequence and the processing efficiency of the image task of the image task processing model, and thus improve the user experience.
[0214] Figure 9 is a block diagram of an image task processing device shown according to an exemplary embodiment. Refer to Figure 9, the device includes:
[0215] A feature sequence division module 910, configured to perform division processing on the image feature sequence to be compressed output by the target task processing layer during the process of inputting the image feature sequence corresponding to the image to be processed into the image task processing model for image task processing, so as to obtain a target number of image feature subsequences; the image task processing model includes: at least two task processing layers connected in sequence, and the target task processing layer is one of the at least two task processing layers except the last task processing layer;
[0216] A target image feature determination module 920, configured to perform determining one target image feature from each image feature subsequence and comprehensively obtaining a target number of target image features;
[0217] A compressed feature update module 930, configured to perform updating the image feature sequence to be compressed based on the target number of target image features to obtain an input feature sequence for the subsequent task processing layer connected to the target task processing layer, and the input feature sequence is used to improve the processing efficiency of the image task of the image task processing model, and the image task is used for visual content understanding of the image to be processed.
[0218] In an optional embodiment, the feature sequence division module 910 includes:
[0219] A compression index determination unit, configured to perform determining preset compression index data;
[0220] A sequence division length determination unit, configured to perform determining the sequence division length based on the preset compression index data, and the sequence division length is positively correlated with the preset compression index data;
[0221] A sequence division unit, configured to perform evenly dividing the image feature sequence to be compressed according to the sequence division length to obtain a target number of image feature subsequences, and the sequence length of each image feature subsequence is the sequence division length.
[0222] In an optional embodiment, the target image feature determination module 920 includes:
[0223] An important index data determination unit, configured to perform determining the important index data corresponding to at least two image features to be compressed in each image feature subsequence;
[0224] A subsequence target feature determination unit, configured to perform taking the image feature to be compressed with the largest numerical value of the corresponding important index data among the at least two image features to be compressed as the target image feature in each image feature subsequence;
[0225] A target image feature unit, configured to perform taking the target image features of each of the target number of image feature subsequences as the target number of target image features.
[0226] In an optional embodiment, the apparatus further includes:
[0227] An associated feature set determination module, configured to perform determining an associated feature set associated with the target number of target image features from a plurality of non-target image features; the plurality of non-target image features are the image features to be compressed other than the target number of target image features in the image feature sequence to be compressed;
[0228] A first feature merging module, configured to perform merging the associated features in the associated feature set into the target number of target image features to obtain the target number of compressed image features;
[0229] The compressed feature update module 930 includes:
[0230] A first update unit, configured to perform updating the image feature sequence to be compressed based on the target number of compressed image features to obtain an input feature sequence.
[0231] In an optional embodiment, the associated feature set determination module includes:
[0232] An associated index data determination unit, configured to perform performing a correlation analysis on each non-target image feature and the target number of target image features respectively to obtain the target number of associated index data corresponding to each non-target image feature, and each associated index data corresponds to one of the target number of target image features;
[0233] A target associated index data determination unit, configured to perform taking the associated index data with the largest corresponding value among the target number of associated index data as the target associated index data corresponding to each non-target image feature;
[0234] An associated feature screening unit, configured to perform adding the non-target image features among the plurality of non-target image features whose corresponding target associated index data satisfies a preset associated index threshold to the associated feature set.
[0235] In an optional embodiment, the first feature merging module includes:
[0236] A target image feature to be merged determination unit, configured to perform determining the target image feature to be merged for each associated feature from the target number of target image features based on the target associated index data corresponding to each associated feature in the associated feature set;
[0237] The target associated feature determination unit to be merged is configured to determine, based on the target image features to be merged for each associated feature, the associated features to be merged corresponding to each target image feature among the target number of target image features;
[0238] The first merging weight determination unit is configured to determine, based on the important index data of the associated features to be merged corresponding to each target image feature, the first merging weight corresponding to the associated features to be merged;
[0239] The associated feature merging unit is configured to merge, based on the first merging weight, the associated features to be merged corresponding to each target image feature into the corresponding target image feature to obtain the compressed image feature corresponding to each target image feature;
[0240] The compressed image feature generation unit is configured to use the compressed image features corresponding to the target number of target image features respectively as the target number of compressed image features.
[0241] In an optional embodiment, the apparatus further includes:
[0242] The second feature merging module is configured to perform feature merging processing on the non-associated feature set to obtain a compensated image feature; the non-associated feature set is a set of non-target image features other than the associated feature set among the multiple non-target image features;
[0243] The compressed feature update module 930 includes:
[0244] The second update unit is configured to update the image feature sequence to be compressed based on the target number of compressed image features and the compensated image feature to obtain an input feature sequence.
[0245] In an optional embodiment, the second feature merging module includes:
[0246] The second merging weight determination unit is configured to determine, based on the important index data of each non-associated feature in the non-associated feature set, the second merging weight corresponding to each non-associated feature;
[0247] The non-associated feature merging unit is configured to perform merging processing on each non-associated feature based on the second merging weight to obtain a compensated image feature.
[0248] In an alternative embodiment, when there are at least three layers in the at least two task processing layers, the target task processing layer is any one of the task processing layers in the target compression processing layer, and the target compression processing layer includes: at least one first task processing layer and at least one second task processing layer. The at least one first task processing layer is at least one task processing layer between the first task processing layer and the middle task processing layer among the at least three task processing layers. The at least one second task processing layer is at least one task processing layer between the middle task processing layer and the penultimate task processing layer among the at least three task processing layers. The middle task processing layer is the task processing layer located in the middle layer among the at least three task processing layers.
[0249] Regarding the device in the above embodiment, the specific manners in which each module performs operations have been described in detail in the embodiment related to the method, and will not be elaborated herein.
[0250] Figure 10 is a block diagram of an electronic device for image task processing shown according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as Figure 10 shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an image task processing method. The display screen of the electronic device may be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0251] Figure 11 is a block diagram of another electronic device for image task processing shown according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as Figure 11As shown. The electronic device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements an image task processing method.
[0252] Those skilled in the art can understand that Figure 10 or Figure 11 the structure shown in is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0253] In an exemplary embodiment, an electronic device is further provided, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the image task processing method in the exemplary embodiment of the present disclosure.
[0254] In an exemplary embodiment, a computer-readable storage medium is further provided. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the image task processing method in the exemplary embodiment of the present disclosure.
[0255] In an exemplary embodiment, a computer program product containing instructions is further provided. When it runs on a computer, the computer executes the image task processing method in the exemplary embodiment of the present disclosure.
[0256] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When this computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0257] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0258] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An image task processing method, characterized in that: The method comprises: In the process of inputting the image feature sequence corresponding to the image to be processed into the image task processing model for image task processing, the image feature sequence to be compressed output by the target task processing layer is divided and processed to obtain a target number of image feature subsequences; the image task processing model comprises: at least two task processing layers connected in sequence, the target task processing layer is a task processing layer other than the last task processing layer in the at least two task processing layers; Determine a target image feature from each of the image feature subsequences, and comprehensively obtain a target number of the target image features; Based on a target number of the target image features, the feature sequence of the image to be compressed is updated to obtain an input feature sequence of a subsequent task processing layer connected to the target task processing layer, wherein the input feature sequence is used to improve the processing efficiency of the image task of the image task processing model, and the image task is used to understand the visual content of the image to be processed.
2. The image task processing method according to claim 1, characterized in that: The image feature sequence to be compressed output by the target task processing layer is divided and processed to obtain a target number of image feature subsequences, including: Determine preset compression index data; Based on the preset compression index data, determining a sequence division length, wherein the sequence division length is positively correlated with the preset compression index data; The image feature sequence to be compressed is evenly divided according to the sequence division length to obtain a target number of the image feature subsequences, and the sequence length of each of the image feature subsequences is the sequence division length.
3. The image task processing method according to claim 1, characterized in that: Each of the image feature subsequences includes at least two image features to be compressed, and determining a target image feature from each of the image feature subsequences to obtain a target number of target image features includes: Determine important indicator data corresponding to at least two image features to be compressed in each of the image feature subsequences; The image feature to be compressed having the largest value of the corresponding important indicator data among the at least two image features to be compressed is used as the target image feature in each of the image feature subsequences; The target image features of each of the target number of image feature subsequences are used as the target number of target image features.
4. The image task processing method according to claim 1, characterized in that: Before updating the to-be-compressed image feature sequence based on the target number of the target image features to obtain an input feature sequence of a subsequent task processing layer connected to the target task processing layer, the method further includes: Determine, from a plurality of non-target image features, a set of associated features associated with a target number of the target image features; the plurality of non-target image features are image features to be compressed other than the target number of the target image features in the sequence of image features to be compressed; Merging the associated features in the associated feature set into a target number of target image features to obtain a target number of compressed image features; The updating of the to-be-compressed image feature sequence based on the target number of the target image features to obtain an input feature sequence of a subsequent task processing layer connected to the target task processing layer comprises: Based on the target number of compressed image features, the to-be-compressed image feature sequence is updated to obtain the input feature sequence.
5. The image task processing method according to claim 4, characterized in that: The step of determining a set of associated features associated with a target number of target image features from a plurality of non-target image features comprises: Perform correlation analysis on each of the non-target image features and a target number of the target image features to obtain a target number of correlation index data corresponding to each of the non-target image features, each of the correlation index data corresponding to one of the target number of the target image features; The associated index data with the largest corresponding value among the target number of associated index data is used as the target associated index data corresponding to each of the non-target image features; The non-target image features whose corresponding target association index data satisfy a preset association index threshold value among the plurality of the non-target image features are added to the association feature set.
6. The image task processing method according to claim 4, characterized in that: The step of merging the associated features in the associated feature set into a target number of target image features to obtain a target number of compressed image features comprises: Based on the target association index data corresponding to each association feature in the association feature set, determining the target image feature to be merged for each association feature from a target number of the target image features; Based on the target image feature to be merged with each of the associated features, determining the associated features to be merged corresponding to each of the target image features among a target number of the target image features; Determining a first merging weight corresponding to the associated features to be merged based on important indicator data of the associated features to be merged corresponding to each of the target image features; Based on the first merging weight, merging the to-be-merged associated features corresponding to each of the target image features into the corresponding target image features to obtain compressed image features corresponding to each of the target image features; The compressed image features corresponding to the target number of target image features are used as the target number of compressed image features.
7. The image task processing method according to claim 4, before updating the to-be-compressed image feature sequence based on the target number of the target image features to obtain an input feature sequence of a subsequent task processing layer connected to the target task processing layer, the method further comprises: Perform feature merging processing on the non-correlated feature set to obtain the compensating image features; The non-associated feature set is a set of non-target image features other than the associated feature set among the plurality of non-target image features; The updating of the to-be-compressed image feature sequence based on the target number of the target image features to obtain an input feature sequence of a subsequent task processing layer connected to the target task processing layer comprises: Based on a target number of the compressed image features and the compensated image features, the to-be-compressed image feature sequence is updated to obtain the input feature sequence.
8. The image task processing method according to claim 7, characterized in that: The step of performing feature merging processing on the non-correlated feature set to obtain the compensated image features comprises: Determine a second merging weight corresponding to each non-associated feature in the non-associated feature set based on the important indicator data corresponding to each non-associated feature in the non-associated feature set; Based on the second merging weight, each of the non-associated features is merged to obtain the compensated image feature.
9. The image task processing method according to any one of claims 1 to 8, characterized in that: In the case where the at least two task processing layers include at least three task processing layers, the target task processing layer is any task processing layer in the target compression processing layer, and the target compression processing layer includes: at least one first task processing layer and at least one second task processing layer, the at least one first task processing layer is at least one task processing layer between the first task processing layer and the middle task processing layer in the at least three task processing layers, the at least one second task processing layer is at least one task processing layer between the middle task processing layer and the penultimate task processing layer in the at least three task processing layers, and the middle task processing layer is a task processing layer located in the middle layer among the at least three task processing layers.
10. An image task processing device, characterized in that: The device comprises: The feature sequence division module is configured to perform division processing on the image feature sequence to be compressed output by the target task processing layer in the process of inputting the image feature sequence corresponding to the image to be processed into the image task processing model for image task processing, so as to obtain a target number of image feature subsequences; the image task processing model comprises: at least two task processing layers connected in sequence, the target task processing layer being a task processing layer other than the last task processing layer in the at least two task processing layers; a target image feature determination module, configured to determine a target image feature from each of the image feature subsequences, and comprehensively obtain a target number of the target image features; The compression feature update module is configured to execute an update of the feature sequence of the image to be compressed based on a target number of the target image features, and obtain an input feature sequence of a subsequent task processing layer connected to the target task processing layer, wherein the input feature sequence is used to improve the processing efficiency of the image task of the image task processing model, and the image task is used to understand the visual content of the image to be processed.
11. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image task processing method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the image task processing method as described in any one of claims 1 to 9.