Task processing method, apparatus, device, medium, and program product

CN122816893APending Publication Date: 2026-09-25CHINA MOBILE GROUP JIANGSU +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611049108.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]本申请实施例提供一种任务处理方法、装置、设备、介质及程序产品,可解决由于大模型的路由方法的路由匹配度较差所导致的大模型的计算资源浪费的问题

Benefits of technology

[0015]本申请实施例中,接收包含至少一个模态数据的待处理任务,对至少一个模态数据的模态重要性评分以及每个模态数据的词元的词语贡献度进行量化分析,得到至少一个模态重要性得分和至少一个词元贡献度集,并利用所述至少一个模态重要性得分和至少一个词元贡献度集来确定目标路由策略,相较于现有技术而言,由于将路由决策从单一词元维度的考量,扩展综合考虑词元以及模态来进行路由,因此,可改变原有仅依赖词元这一局部特征做决策的局限,使路由匹配度大幅提升,有效缓解了路由瓶颈与负载不均衡问题,降低了分布式部署中的节点通信开销,从而减少大模型进行数据处理的计算资源浪费。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816893A_ABST
    Figure CN122816893A_ABST
Patent Text Reader

Abstract

The application provides a task processing method and device, equipment, medium and program product, which are applied to the field of artificial intelligence technology. The method comprises the following steps: receiving a to-be-processed task; determining at least one modality importance score corresponding to the at least one modality data; quantifying the word contribution of the at least one modality data to obtain at least one word contribution set corresponding to the at least one modality data; determining a target routing strategy based on the at least one modality importance score and the at least one word contribution set; and calling the target model to process the to-be-processed task based on the target routing strategy. In the method, the original limitation of only relying on the local feature of the word to make decisions can be changed, the routing matching degree is greatly improved, the node communication overhead in distributed deployment is reduced, and the waste of computing resources for data processing of a large model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a task processing method, apparatus, device, medium, and program product. Background Technology

[0002] To accelerate the sparse training and inference of multimodal large models, the relevant technologies mainly reduce resource overhead through feature processing, model compression, dynamic routing, and system optimization.

[0003] For example, dynamic routing methods based on single lexical features select computational resources solely based on local lexical features. This routing decision-making analyzes from only a single dimension, neglecting parameters from other dimensions, which can easily lead to route exhaustion and load imbalance, as well as high cross-node communication overhead in distributed deployments. Therefore, due to the poor route matching accuracy of routing methods for large models, computational resources for data processing in large models are wasted. Summary of the Invention

[0004] This application provides a task processing method, apparatus, device, medium, and program product that can solve the problem of wasted computing resources in large models caused by poor routing matching in routing methods for large models.

[0005] In a first aspect, embodiments of this application provide a task processing method applied to an inference device on which a target model is deployed, the method comprising: Receive a task to be processed, the task to be processed includes at least one modal data, wherein different modal data in the at least one modal data correspond to different modalities, and the modality is used to indicate the data type of the corresponding modal data; Determine at least one modality importance score corresponding one-to-one with the at least one modality data. The modality importance score is used to indicate the degree of difference in the output results of the target model when processing the task to be processed before and after the corresponding modality data is masked in the target model. The value of the modality importance score is proportional to the magnitude of the difference. The at least one modal data is quantified by word contribution to obtain at least one word contribution set corresponding to the at least one modal data. The first word contribution set includes at least one word contribution corresponding to at least one word obtained by splitting the first modal data. The word contribution is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding word is masked in the target model. The value of the word contribution is proportional to the magnitude of the difference. The first word contribution set is any word contribution set in the at least one word contribution set. The first modal data is the modal data in the at least one modal data that corresponds to the first word contribution set. Based on the at least one modality importance score and the at least one lexical contribution set, a target routing strategy is determined, wherein the target routing strategy is used to indicate that the target computing resources in the target model are used to process the task to be processed; Based on the target routing strategy, the target model is invoked to process the task to be processed.

[0006] Optionally, determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set includes: Based on the at least one modality importance score, a first threshold, and a second threshold, at least one modality value level corresponding to the at least one modality data is determined. The modality value level includes a first value level, a second value level, and a third value level. Specifically, if the first modality importance score corresponding to the first modality data is less than or equal to the first threshold, the first modality value level corresponding to the first modality data is the first value level; if the first modality importance score is greater than the first threshold and less than the second threshold, the first modality value level is the second value level; and if the first modality importance score is greater than the second threshold, the first modality value level is the third value level. The target routing policy is determined based on the at least one modal value level, wherein the target routing policy includes: Mask the modal data of the first value level among the at least one modal value level; The target computing resources are used to process the modal data at the third value level; The target lexical is processed using the target computing resources. The target lexical is a lexical determined from at least one lexical based on the first lexical contribution set harmonics and a preset contribution threshold, when the first modality level is the second value level. The lexical contribution of the target lexical is greater than the preset contribution threshold.

[0007] Optionally, before determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: When the modality of the first modal data is a text modality, the contribution of the first word is determined based on the self-attention weight and similarity parameter of the first word. The first word is any one of the at least one word. The similarity parameter is the semantic similarity between the data after the first word is removed from the first modal data and the first modal data. When the modality of the first modal data is an image modality, the word contribution of the first word is determined based on the feature variance of the first word in the foreground region of the first modal data. The magnitude of the feature variance is proportional to the magnitude of the word contribution of the first word. The first word is an image obtained by splitting the foreground region of the first modal data. When the modality of the first modal data is video modality, frame extraction is performed on the first modal data to obtain multiple video frames, and multiple structural similarities corresponding to the multiple video frames are calculated one-to-one. The N video frames corresponding to the top N structural similarities with the largest structural similarity are determined as N target video frames. The foreground regions of the N target video frames are split into N word sets, and the at least one word includes all words in the N word sets. Based on the feature variance of the first word in the foreground region of the corresponding target video frame, the word contribution of the first word is determined, where N is a positive integer.

[0008] Optionally, determining at least one modality importance score corresponding one-to-one with the at least one modality data based on the at least one modality data includes: The feature norm, feature information entropy, and preset attention weight of the first modality data are obtained. The feature norm is the norm calculated by performing L2 norm on the feature matrix corresponding to the first modality data. The feature information entropy is used to indicate the information richness of the first modality data. The target feature is obtained by concatenating the feature norm, the feature information entropy, and the preset attention weight. The target features are input into a gating network for modal importance evaluation, and the modal importance score output by the gating network corresponding to the first modal data is obtained.

[0009] Optionally, the target model includes multiple expert computing resources; Before determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: At least one modal feature corresponding to the at least one modal data is obtained. When the modality of the first modal data is text, the modal feature of the first modal data is the mean and variance corresponding to the at least one word. When the modality of the first modal data is an image, the modal feature of the first modal data is the proportion of the foreground of the first modal data in the image corresponding to the first modal data. When the modality of the first modal data is video, the modal feature of the first modal data is the average proportion of the foreground of the video frame obtained by splitting the first modal data in the video frame obtained by splitting the first modal data. The at least one modal feature is input into the routing network for resource allocation, resulting in at least one expert computing resource output by the routing network that corresponds one-to-one with the at least one modal data. The target computing resource includes the at least one expert computing resource. The routing network is used to determine multiple current selection probabilities corresponding one-to-one with the multiple expert computing resources based on the at least one modal feature, and to determine the target selection probability of each expert computing resource based on the current selection probability and historical selection probability of each expert computing resource, and to output the at least one expert computing resource based on the target selection probability of each expert computing resource.

[0010] Optionally, after determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: Obtain a resource selection information set, which includes multiple resource selection information corresponding one-to-one with the multiple expert computing resources. The resource selection information includes: the number of times the corresponding expert computing resource is selected within a preset time period, and the modality of the data processed when the corresponding expert computing resource is selected. The preset time period includes the time when the target computing resource is determined. Based on the resource selection information set, a loss function value is determined. The loss function value is used to indicate the degree of concentration of computing resources output by the routing network within the preset time period. The magnitude of the loss function value is directly proportional to the magnitude of the concentration. The routing network is tuned based on the loss function value.

[0011] Secondly, embodiments of this application also provide a task processing apparatus applied to an inference device on which a target model is deployed, the apparatus comprising: A receiving module is used to receive a task to be processed, the task to be processed including at least one modal data, wherein different modal data in the at least one modal data correspond to different modalities, and the modality is used to indicate the data type of the corresponding modal data; The first determining module is used to determine at least one modal importance score corresponding one-to-one with the at least one modal data. The modal importance score is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding modal data is masked in the target model. The value of the modal importance score is proportional to the magnitude of the difference. A contribution metric module is used to metric the lexical contribution of the at least one modal data to obtain at least one lexical contribution set corresponding to the at least one modal data. The first lexical contribution set includes at least one lexical contribution corresponding to at least one lexical obtained by splitting the first modal data. The lexical contribution is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding lexical is masked in the target model. The value of the lexical contribution is proportional to the magnitude of the difference. The first lexical contribution set is any lexical contribution set in the at least one lexical contribution set. The first modal data is the modal data in the at least one modal data that corresponds to the first lexical contribution set. The second determining module is used to determine a target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, wherein the target routing strategy is used to indicate that the target computing resources in the target model are used to process the task to be processed; The processing module is used to process the task to be processed by calling the target model based on the target routing strategy.

[0012] Thirdly, embodiments of this application also provide an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the task processing method as described in the first aspect.

[0013] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the task processing method described in the first aspect.

[0014] Fifthly, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the task processing method as described in the first aspect.

[0015] In this embodiment, a task to be processed containing at least one modality data is received. The modality importance score of the at least one modality data and the word contribution of each modality data lexical are quantitatively analyzed to obtain at least one modality importance score and at least one lexical contribution set. The target routing strategy is then determined using the at least one modality importance score and at least one lexical contribution set. Compared to existing technologies, this approach expands the consideration of routing decisions from a single lexical dimension to a comprehensive consideration of both lexical and modality. Therefore, it overcomes the limitations of relying solely on the local feature of lexicals for decision-making, significantly improving routing matching accuracy. This effectively alleviates routing bottlenecks and load imbalances, reduces node communication overhead in distributed deployments, and thus reduces the waste of computational resources for processing large models. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of the task processing method provided in the embodiments of this application; Figure 2 This is a flowchart of the multimodal hierarchical dynamic routing mechanism provided in the embodiments of this application; Figure 3 This is a schematic diagram of the training and enhancement of the routing network provided in the embodiments of this application; Figure 4 This is a schematic diagram of the inference resource awareness adaptive strategy provided in an embodiment of this application; Figure 5 This is a structural diagram of a task processing apparatus provided in an embodiment of this application; Figure 6 This is a structural diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] This application provides a task processing method, apparatus, device, medium, and program product.

[0020] This application provides a task processing method applied to an inference device with a target model deployed. Figure 1 This is a flowchart of the task processing method provided in the embodiments of this application, such as... Figure 1 As shown, it includes the following steps: Step 101: Receive a task to be processed. The task to be processed includes at least one modal data, wherein different modal data in the at least one modal data correspond to different modalities, and the modality is used to indicate the data type of the corresponding modal data.

[0021] In this step, "modality" refers to data type, such as text, image, video, etc., and "target model" refers to a multimodal large model. The following example illustrates the scenario where the target model receives the task to be processed: Send a video to the target model with the prompt: Analyze how many people appear in the video; in this example, the task to be processed includes two modal data, one is the video (video modality) and the other is the prompt (text modality).

[0022] Step 102: Determine at least one modal importance score corresponding one-to-one with the at least one modal data. The modal importance score is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding modal data is masked in the target model. The value of the modal importance score is proportional to the magnitude of the difference.

[0023] In this step, using the example above as an example, we will explain the modal importance score: The task to be processed includes two modalities: a video and a cue word. Understandably, if neither modality is masked, the target model can analyze the video and output the number of people appearing in the video as required by the cue word. If the cue word is masked, the target model only analyzes the video, and may output various content about the video (possibly including the number of people appearing in the video). The difference in the output of the task before and after masking the cue word is relatively small. Conversely, if the video is masked, the target model cannot output the number of people appearing in the video under any circumstances; therefore, the difference in the output of the task before and after masking the video is very large. By comparison, in this example, the modal importance score corresponding to the video is greater than the modal importance score corresponding to the cue word.

[0024] Step 103 involves quantifying the lexical contribution of the at least one modal data to obtain at least one lexical contribution set corresponding to each of the at least one modal data. The first lexical contribution set includes at least one lexical contribution corresponding to each of the at least one lexical obtained from the splitting of the first modal data. The lexical contribution is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding lexical is masked in the target model. The value of the lexical contribution is proportional to the magnitude of the difference. The first lexical contribution set is any lexical contribution set in the at least one lexical contribution set. The first modal data is the modal data in the at least one modal data that corresponds to the first lexical contribution set.

[0025] In this step, each modal data can be broken down into at least one lexical unit. For example, "Please analyze how many characters appeared in the video" can be broken down into five lexical units: (analyze, video, appear, how many, characters). For each lexical unit, a lexical contribution analysis can be performed. After disabling "person," the remaining lexical units are (analysis, video, appearance, quantity). When the target model performs analysis, it cannot know what content in the video is being analyzed. Therefore, the output of the target model may or may not include people. After disabling "video," the remaining lexical units are (analysis, appearance, quantity, person). When the target model performs analysis, although the text data does not specify that it is analyzing video, the task to be analyzed includes video. The target model can understand that it is analyzing video, and the probability of the output of the target model changing is small. Through the above two examples, it can be seen that the contribution of the lexical units corresponding to "person" is greater than that of the lexical units corresponding to "video."

[0026] For each lexical segment extracted from the at least one modal data, the lexical contribution is quantified to obtain the at least one lexical contribution set.

[0027] Step 104: Based on the at least one modality importance score and the at least one lexical contribution set, determine a target routing strategy, wherein the target routing strategy is used to indicate that the target computing resources in the target model are used to process the task to be processed.

[0028] A comprehensive analysis is performed based on modality importance scores and lexical contribution sets. For example, for modality data with modality importance scores greater than a certain preset threshold, all modality data is retained. Further analysis of the lexical contribution set of this modality data is conducted, retaining lexical contributions greater than a certain preset threshold. Based on the retained lexical contributions from the retained modality data, suitable computational resources for processing these "retained lexical contributions" are determined from the target model and designated as target computational resources. Understandably, the more data the task retains that needs to be processed, the more target computational resources the target model needs to process that data; conversely, the less data retained that needs to be processed, the fewer target computational resources are needed to process the data, thus saving computational resources.

[0029] Step 105: Based on the target routing strategy, call the target model to process the task to be processed.

[0030] The target computing resources are used to process the lexical units retained in the modal data of the task to be processed, and the processing results are output.

[0031] In the task processing method of this application, a task to be processed containing at least one modality data is received. The modality importance score of the at least one modality data and the word contribution of each modality data lexical are quantitatively analyzed to obtain at least one modality importance score and at least one lexical contribution set. The target routing strategy is determined by using the at least one modality importance score and at least one lexical contribution set. Compared with the prior art, since the routing decision is expanded from a single lexical dimension to comprehensively consider both lexical and modality for routing, the limitation of relying only on the local feature of lexical for decision-making can be changed. This significantly improves the routing matching degree, effectively alleviates the routing bottleneck and load imbalance problem, reduces the node communication overhead in distributed deployment, and thus reduces the waste of computing resources for data processing of large models.

[0032] Optionally, determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set includes: Based on the at least one modality importance score, a first threshold, and a second threshold, at least one modality value level corresponding to the at least one modality data is determined. The modality value level includes a first value level, a second value level, and a third value level. Specifically, if the first modality importance score corresponding to the first modality data is less than or equal to the first threshold, the first modality value level corresponding to the first modality data is the first value level; if the first modality importance score is greater than the first threshold and less than the second threshold, the first modality value level is the second value level; and if the first modality importance score is greater than the second threshold, the first modality value level is the third value level. The target routing policy is determined based on the at least one modal value level, wherein the target routing policy includes: Mask the modal data of the first value level among the at least one modal value level; The target computing resources are used to process the modal data at the third value level; The target lexical is processed using the target computing resources. The target lexical is a lexical determined from at least one lexical based on the first lexical contribution set harmonics and a preset contribution threshold, when the first modality level is the second value level. The lexical contribution of the target lexical is greater than the preset contribution threshold.

[0033] like Figure 2 As shown, the steps in this embodiment can be understood as modality-level dynamic routing and token (i.e., lexical)-level dynamic routing. Through a two-layer collaborative strategy of filtering high-value modalities through modality-level routing and filtering core tokens through token-level routing, the accurate allocation of computing resources can be achieved.

[0034] Modal-level dynamic routing includes: Based on the modality importance score, the participation strategy is divided into three levels. If the score is higher than 0.8 (the second threshold), the modality is judged as a high-value modality (the third value level) and then participates in the subsequent calculation.

[0035] If the score is below 0.4 (the first threshold), and the modality is determined to be low-value (first value level), only the modality embedding layer is retained, and subsequent expert (target computing resource) calculations are blocked. Optionally, a feedback mechanism can be set so that if the task accuracy decreases by more than 1% after blocking, the importance score of the modality is automatically increased.

[0036] If the score is higher than 0.4 but lower than 0.8, the modality is judged as a medium-value modality (second value level), and then proceeds to the token-level route for further filtering of core tokens.

[0037] Token-level dynamic routing includes: Retain the target word and block other word elements; Alternatively, the sparsity ratio of tokens can be dynamically adjusted based on the modality importance score of the modality data. For high-value modalities, the most numerous tokens with the largest lexical contribution are retained, while for low-value modalities, more tokens with lower contribution are deleted to save resources.

[0038] In this embodiment, modality-level dynamic routing is first performed to determine the modal data that can proceed to subsequent computations. For data whose modality level is the second value level, token-level dynamic routing is further performed based on its lexical contribution set to determine the lexicals that can proceed to subsequent computations. Then, based on the determined data and lexicals that can proceed to subsequent computations, target computing resources suitable for processing this data are determined from the target model, resulting in a target routing strategy. This method filters out data, reducing the amount of data to be processed and the amount of computing resources to be scheduled, thereby achieving resource conservation.

[0039] Optionally, before determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: When the modality of the first modal data is a text modality, the contribution of the first word is determined based on the self-attention weight and similarity parameter of the first word. The first word is any one of the at least one word. The similarity parameter is the semantic similarity between the data after the first word is removed from the first modal data and the first modal data. When the modality of the first modal data is an image modality, the word contribution of the first word is determined based on the feature variance of the first word in the foreground region of the first modal data. The magnitude of the feature variance is proportional to the magnitude of the word contribution of the first word. The first word is an image obtained by splitting the foreground region of the first modal data. When the modality of the first modal data is video modality, frame extraction is performed on the first modal data to obtain multiple video frames, and multiple structural similarities corresponding to the multiple video frames are calculated one-to-one. The N video frames corresponding to the top N structural similarities with the largest structural similarity are determined as N target video frames. The foreground regions of the N target video frames are split into N word sets, and the at least one word includes all words in the N word sets. Based on the feature variance of the first word in the foreground region of the corresponding target video frame, the word contribution of the first word is determined, where N is a positive integer.

[0040] In this embodiment, for text modality, the self-attention weights of the tokens and the semantic similarity of the sentences after token removal are fused for calculation. For image modality, a lightweight object detection model, such as YOLO-Nano, is used to identify the foreground region of the image, and the feature variance of the labeled foreground patch is calculated; a higher variance indicates richer details. For video modality, the structural similarity (SSIM) of adjacent frames is calculated, keyframes with high SSIM are retained, and image-specific filtering logic is executed to obtain the contribution of video tokens. Different methods are used to determine token contribution for different modality types of modal data. Compared to using the same method to determine lexical contribution, the specific method is more adaptable to the characteristics of the corresponding modality data, thus improving the accuracy of the determined lexical contribution.

[0041] Optionally, determining at least one modality importance score corresponding one-to-one with the at least one modality data based on the at least one modality data includes: The feature norm, feature information entropy, and preset attention weight of the first modality data are obtained. The feature norm is the norm calculated by performing L2 norm on the feature matrix corresponding to the first modality data. The feature information entropy is used to indicate the information richness of the first modality data. The target feature is obtained by concatenating the feature norm, the feature information entropy, and the preset attention weight. The target features are input into a gating network for modal importance evaluation, and the modal importance score output by the gating network corresponding to the first modal data is obtained.

[0042] In this embodiment, for each modality data, the feature norm, feature information entropy, and preset attention weight are extracted. The feature norm is the feature norm obtained by calculating the L2 norm of the modality feature matrix. The feature information entropy is the feature information entropy of the feature matrix. The preset attention weight is the average attention weight of different modalities in the cross-modal attention layer.

[0043] A modality adaptive gating network (MAG) is trained using a 3-layer lightweight MLP (256 hidden layer dimensions, GELU activation function, 1 output dimension), with a total of less than 100,000 parameters. The core features obtained in Step 1 are standardized and concatenated before being input to obtain the modality importance score as the output. The training objective is to optimize the MAG parameters using the "task accuracy loss after modality masking" as a supervision signal, ensuring that the matching error between the output score and the actual contribution of the modality is less than 3%.

[0044] The method described in this embodiment can improve the accuracy of modal importance scores.

[0045] Optionally, the target model includes multiple expert computing resources; Before determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: At least one modal feature corresponding to the at least one modal data is obtained. When the modality of the first modal data is text, the modal feature of the first modal data is the mean and variance corresponding to the at least one word. When the modality of the first modal data is an image, the modal feature of the first modal data is the proportion of the foreground of the first modal data in the image corresponding to the first modal data. When the modality of the first modal data is video, the modal feature of the first modal data is the average proportion of the foreground of the video frame obtained by splitting the first modal data in the video frame obtained by splitting the first modal data. The at least one modal feature is input into the routing network for resource allocation, resulting in at least one expert computing resource output by the routing network that corresponds one-to-one with the at least one modal data. The target computing resource includes the at least one expert computing resource. The routing network is used to determine multiple current selection probabilities corresponding one-to-one with the multiple expert computing resources based on the at least one modal feature, and to determine the target selection probability of each expert computing resource based on the current selection probability and historical selection probability of each expert computing resource, and to output the at least one expert computing resource based on the target selection probability of each expert computing resource.

[0046] like Figure 3As shown, in this embodiment, a separate set of routing parameters is trained for each modality. The input to the routing network only contains the modal features of the current modality data. For text modality routing, the mean and variance of the text token embedding are used as input; for image modality routing, the feature entropy of the image patch and the foreground proportion are used as input; for video modality routing, the video modality data is first split into image modality data, and then for each split image modality data, the feature entropy of the image patch and the foreground proportion are used as input.

[0047] In the process of using routing network output experts to calculate resources, the routing selection probability of the determined target selection is calculated using the exponential moving average (EMA), with the formula as follows: ; in, Let be the historical choice probability of expert e under mode m in round t. This represents the current selection probability in the current round, to avoid sudden switching of routing decisions.

[0048] It should be noted that during the training of the routing network, the gradient of the network is dynamically pruned. The pruning threshold is 1.5 times the standard deviation of the current training round to filter out abnormally large gradients generated by noisy modes.

[0049] By using the above method, target computing resources suitable for processing the task to be processed can be selected, thereby improving resource utilization efficiency.

[0050] Optionally, after determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: Obtain a resource selection information set, which includes multiple resource selection information corresponding one-to-one with the multiple expert computing resources. The resource selection information includes: the number of times the corresponding expert computing resource is selected within a preset time period, and the modality of the data processed when the corresponding expert computing resource is selected. The preset time period includes the time when the target computing resource is determined. Based on the resource selection information set, a loss function value is determined. The loss function value is used to indicate the degree of concentration of computing resources output by the routing network within the preset time period. The magnitude of the loss function value is directly proportional to the magnitude of the concentration. The routing network is tuned based on the loss function value.

[0051] In this embodiment, the loss function of the routing network is determined as follows: ; Where M is the number of modes. To train the number of times expert e is selected under the intra-batch modality m. The global load is the number of times expert e is selected across all modalities within the training batch. The resource selection information includes... and α and β are the weighting coefficients for controlling load balancing; here, α = 0.6 and β = 0.4 are chosen.

[0052] Loss calculation will The routing network is trained by weighting and fusing the task loss.

[0053] By using the method described in this embodiment, we can avoid situations where some computing resources in the target model are concentrated and selected multiple times, while the probability and frequency of selection of some computing resources are low, thereby reducing idle expert computing resources and improving resource utilization.

[0054] In one optional implementation, the load percentage of each expert is monitored in real time, i.e., the load of a certain expert / the average load of all experts in the same modality. If the expert's load percentage is high (load percentage > 1.5), a penalty is imposed on the expert's selection probability by multiplying it by a coefficient less than 1 to guide the token to select other experts; otherwise, a high coefficient is used to reward experts to prevent them from being idle.

[0055] In an alternative implementation, such as Figure 4 As shown, before deploying the target model to a device for inference, it needs to be adapted to the device's environment, including the following methods: Step 1: Pre-inference resource exploration and SLA analysis Step 1: Before inference starts, obtain the core resource parameters of the device, including: computing power (GFLOPS), memory (GB), and bandwidth (Gbps).

[0056] Step 2: Perform SLA requirement analysis and extract key metrics from the inference task configuration, including: latency requirement (ms), throughput requirement (QPS), and accuracy requirement.

[0057] Step 2: Dynamically adjust the number of Top-k experts Step 1: Establish a mapping table for "Device computing power - SLA latency - k value", generated through offline calibration, as shown in the example below:

[0058] Step 2: Monitor latency in real time during the inference process, and count the average latency every 10 inferences. If the average latency is 10% higher than the upper SLA latency limit, decrease the value of k by 1 while maintaining accuracy monitoring; if the average latency is 20% lower than the lower SLA latency limit, increase the value of k by 1 to improve inference accuracy; constraint: after adjusting the value of k, the drop in accuracy shall not exceed 5% of the SLA accuracy requirement, and SLA accuracy is counted every 50 inferences.

[0059] Step 3: Determine the early stopping trigger condition according to the device status and adjust dynamically based on real-time accuracy. If the accuracy is 5% lower than the lower SLA accuracy limit, increase to increase the number of expert calculations; if the accuracy is 5% higher than the upper SLA accuracy limit, decrease it to reduce the number of expert calculations and improve the inference speed at the terminal.

[0060] In the task processing method of the present application, through "modal-level + token-level" routing, the importance of modalities is quantified and the contribution of modality-specific tokens is calculated, which can solve the problem of resource waste existing in the single routing of the prior art. The routing selection network is optimized with "modality-specific gating + gradient clipping EMA +双层负载损失", which can avoid the problem of poor routing accuracy in the prior art. By detecting device resources and SLA requirements and dynamically adjusting Top-k and early stopping thresholds, the method can adapt to the differences between edge devices and cloud. It can reuse training parameters and routing rules, optimize expert caching, and avoid the high cost of independent optimization in the prior art.

[0061] See Figure 5 , Figure 5 which is a structural diagram of a task processing apparatus provided by an embodiment of the present application, the task processing apparatus is applied to an inference device deployed with a target model. As shown in Figure 5 , the apparatus 500 comprises: a receiving module 501, configured to receive a to-be-processed task, the to-be-processed task comprising at least one piece of modal data, wherein different modal data in the at least one piece of modal data correspond to different modalities, and the modalities are used to indicate data types of the corresponding modal data; a first determining module 502, configured to determine at least one modality importance score corresponding one-to-one to the at least one piece of modal data, the modality importance score being used to indicate: the difference between results output by the target model when processing the to-be-processed task before and after the corresponding modal data is masked in the target model, and the value of the modality importance score is directly proportional to the degree of difference; The contribution metric module 503 is used to perform lexical contribution metric on the at least one modal data to obtain at least one lexical contribution set corresponding to the at least one modal data. The first lexical contribution set includes at least one lexical contribution corresponding to at least one lexical obtained by splitting the first modal data. The lexical contribution is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding lexical is masked in the target model. The value of the lexical contribution is proportional to the magnitude of the difference. The first lexical contribution set is any lexical contribution set in the at least one lexical contribution set. The first modal data is the modal data in the at least one modal data that corresponds to the first lexical contribution set. The second determining module 504 is used to determine a target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, wherein the target routing strategy is used to indicate that the target computing resources in the target model are used to process the task to be processed. The processing module 505 is used to process the task to be processed by calling the target model based on the target routing strategy.

[0062] Optionally, the second determining module 504 is also used for: Based on the at least one modality importance score, a first threshold, and a second threshold, at least one modality value level corresponding to the at least one modality data is determined. The modality value level includes a first value level, a second value level, and a third value level. Specifically, if the first modality importance score corresponding to the first modality data is less than or equal to the first threshold, the first modality value level corresponding to the first modality data is the first value level; if the first modality importance score is greater than the first threshold and less than the second threshold, the first modality value level is the second value level; and if the first modality importance score is greater than the second threshold, the first modality value level is the third value level. The target routing policy is determined based on the at least one modal value level, wherein the target routing policy includes: Mask the modal data of the first value level among the at least one modal value level; The target computing resources are used to process the modal data at the third value level; The target lexical is processed using the target computing resources. The target lexical is a lexical determined from at least one lexical based on the first lexical contribution set harmonics and a preset contribution threshold, when the first modality level is the second value level. The lexical contribution of the target lexical is greater than the preset contribution threshold.

[0063] Optionally, the device 500 further includes a third determining module for: When the modality of the first modal data is a text modality, the contribution of the first word is determined based on the self-attention weight and similarity parameter of the first word. The first word is any one of the at least one word. The similarity parameter is the semantic similarity between the data after the first word is removed from the first modal data and the first modal data. When the modality of the first modal data is an image modality, the word contribution of the first word is determined based on the feature variance of the first word in the foreground region of the first modal data. The magnitude of the feature variance is proportional to the magnitude of the word contribution of the first word. The first word is an image obtained by splitting the foreground region of the first modal data. When the modality of the first modal data is video modality, frame extraction is performed on the first modal data to obtain multiple video frames, and multiple structural similarities corresponding to the multiple video frames are calculated one-to-one. The N video frames corresponding to the top N structural similarities with the largest structural similarity are determined as N target video frames. The foreground regions of the N target video frames are split into N word sets, and the at least one word includes all words in the N word sets. Based on the feature variance of the first word in the foreground region of the corresponding target video frame, the word contribution of the first word is determined, where N is a positive integer.

[0064] Optionally, the device 500 further includes a fourth determining module for: The feature norm, feature information entropy, and preset attention weight of the first modality data are obtained. The feature norm is the norm calculated by performing L2 norm on the feature matrix corresponding to the first modality data. The feature information entropy is used to indicate the information richness of the first modality data. The target feature is obtained by concatenating the feature norm, the feature information entropy, and the preset attention weight. The target features are input into a gating network for modal importance evaluation, and the modal importance score output by the gating network corresponding to the first modal data is obtained.

[0065] Optionally, the target model includes multiple expert computing resources; the device 500 further includes a fifth determining module, used for: At least one modal feature corresponding to the at least one modal data is obtained. When the modality of the first modal data is text, the modal feature of the first modal data is the mean and variance corresponding to the at least one word. When the modality of the first modal data is an image, the modal feature of the first modal data is the proportion of the foreground of the first modal data in the image corresponding to the first modal data. When the modality of the first modal data is video, the modal feature of the first modal data is the average proportion of the foreground of the video frame obtained by splitting the first modal data in the video frame obtained by splitting the first modal data. The at least one modal feature is input into the routing network for resource allocation, resulting in at least one expert computing resource output by the routing network that corresponds one-to-one with the at least one modal data. The target computing resource includes the at least one expert computing resource. The routing network is used to determine multiple current selection probabilities corresponding one-to-one with the multiple expert computing resources based on the at least one modal feature, and to determine the target selection probability of each expert computing resource based on the current selection probability and historical selection probability of each expert computing resource, and to output the at least one expert computing resource based on the target selection probability of each expert computing resource.

[0066] Optionally, the device 500 also includes a training module for: Obtain a resource selection information set, which includes multiple resource selection information corresponding one-to-one with the multiple expert computing resources. The resource selection information includes: the number of times the corresponding expert computing resource is selected within a preset time period, and the modality of the data processed when the corresponding expert computing resource is selected. The preset time period includes the time when the target computing resource is determined. Based on the resource selection information set, a loss function value is determined. The loss function value is used to indicate the degree of concentration of computing resources output by the routing network within the preset time period. The magnitude of the loss function value is directly proportional to the magnitude of the concentration. The routing network is tuned based on the loss function value.

[0067] The task processing apparatus 500 of this application can implement all the steps of the above-described task processing method and achieve the same beneficial effects. To avoid repetition, it will not be described again.

[0068] This application also provides an electronic device. Since the principle by which the electronic device solves the problem is similar to the task processing method in this application, the implementation of this electronic device can refer to the implementation of the above-described task processing method; repeated details will not be elaborated further. Figure 6 As shown, the electronic device according to an embodiment of this application includes: a processor 600, configured to read a program from a memory 620 and execute the following processes: Receive a task to be processed, the task to be processed includes at least one modal data, wherein different modal data in the at least one modal data correspond to different modalities, and the modality is used to indicate the data type of the corresponding modal data; Determine at least one modality importance score corresponding one-to-one with the at least one modality data. The modality importance score is used to indicate the degree of difference in the output results of the target model when processing the task to be processed before and after the corresponding modality data is masked in the target model. The value of the modality importance score is proportional to the magnitude of the difference. The at least one modal data is quantified by word contribution to obtain at least one word contribution set corresponding to the at least one modal data. The first word contribution set includes at least one word contribution corresponding to at least one word obtained by splitting the first modal data. The word contribution is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding word is masked in the target model. The value of the word contribution is proportional to the magnitude of the difference. The first word contribution set is any word contribution set in the at least one word contribution set. The first modal data is the modal data in the at least one modal data that corresponds to the first word contribution set. Based on the at least one modality importance score and the at least one lexical contribution set, a target routing strategy is determined, wherein the target routing strategy is used to indicate that the target computing resources in the target model are used to process the task to be processed; Based on the target routing strategy, the target model is invoked to process the task to be processed.

[0069] Among them, Figure 6In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 600) and memory (memory 620). The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. A bus interface provides the interface. Processor 600 is responsible for managing the bus architecture and general processing, and memory 620 can store data used by processor 600 during operation.

[0070] Optionally, the processor 600 is configured to read the program from the memory 620 and execute the following processes: Based on the at least one modality importance score, a first threshold, and a second threshold, at least one modality value level corresponding to the at least one modality data is determined. The modality value level includes a first value level, a second value level, and a third value level. Specifically, if the first modality importance score corresponding to the first modality data is less than or equal to the first threshold, the first modality value level corresponding to the first modality data is the first value level; if the first modality importance score is greater than the first threshold and less than the second threshold, the first modality value level is the second value level; and if the first modality importance score is greater than the second threshold, the first modality value level is the third value level. The target routing policy is determined based on the at least one modal value level, wherein the target routing policy includes: Mask the modal data of the first value level among the at least one modal value level; The target computing resources are used to process the modal data at the third value level; The target lexical is processed using the target computing resources. The target lexical is a lexical determined from at least one lexical based on the first lexical contribution set harmonics and a preset contribution threshold, when the first modality level is the second value level. The lexical contribution of the target lexical is greater than the preset contribution threshold.

[0071] Optionally, the processor 600 is configured to read the program from the memory 620 and execute the following processes: When the modality of the first modal data is a text modality, the contribution of the first word is determined based on the self-attention weight and similarity parameter of the first word. The first word is any one of the at least one word. The similarity parameter is the semantic similarity between the data after the first word is removed from the first modal data and the first modal data. When the modality of the first modal data is an image modality, the word contribution of the first word is determined based on the feature variance of the first word in the foreground region of the first modal data. The magnitude of the feature variance is proportional to the magnitude of the word contribution of the first word. The first word is an image obtained by splitting the foreground region of the first modal data. When the modality of the first modal data is video modality, frame extraction is performed on the first modal data to obtain multiple video frames, and multiple structural similarities corresponding to the multiple video frames are calculated one-to-one. The N video frames corresponding to the top N structural similarities with the largest structural similarity are determined as N target video frames. The foreground regions of the N target video frames are split into N word sets, and the at least one word includes all words in the N word sets. Based on the feature variance of the first word in the foreground region of the corresponding target video frame, the word contribution of the first word is determined, where N is a positive integer.

[0072] Optionally, the processor 600 is configured to read the program from the memory 620 and execute the following processes: The feature norm, feature information entropy, and preset attention weight of the first modality data are obtained. The feature norm is the norm calculated by performing L2 norm on the feature matrix corresponding to the first modality data. The feature information entropy is used to indicate the information richness of the first modality data. The target feature is obtained by concatenating the feature norm, the feature information entropy, and the preset attention weight. The target features are input into a gating network for modal importance evaluation, and the modal importance score output by the gating network corresponding to the first modal data is obtained.

[0073] Optionally, the target model includes multiple expert computing resources; the processor 600 is used to read the program from the memory 620 and execute the following processes: At least one modal feature corresponding to the at least one modal data is obtained. When the modality of the first modal data is text, the modal feature of the first modal data is the mean and variance corresponding to the at least one word. When the modality of the first modal data is an image, the modal feature of the first modal data is the proportion of the foreground of the first modal data in the image corresponding to the first modal data. When the modality of the first modal data is video, the modal feature of the first modal data is the average proportion of the foreground of the video frame obtained by splitting the first modal data in the video frame obtained by splitting the first modal data. The at least one modal feature is input into the routing network for resource allocation, resulting in at least one expert computing resource output by the routing network that corresponds one-to-one with the at least one modal data. The target computing resource includes the at least one expert computing resource. The routing network is used to determine multiple current selection probabilities corresponding one-to-one with the multiple expert computing resources based on the at least one modal feature, and to determine the target selection probability of each expert computing resource based on the current selection probability and historical selection probability of each expert computing resource, and to output the at least one expert computing resource based on the target selection probability of each expert computing resource.

[0074] Optionally, the processor 600 is configured to read the program from the memory 620 and execute the following processes: Obtain a resource selection information set, which includes multiple resource selection information corresponding one-to-one with the multiple expert computing resources. The resource selection information includes: the number of times the corresponding expert computing resource is selected within a preset time period, and the modality of the data processed when the corresponding expert computing resource is selected. The preset time period includes the time when the target computing resource is determined. Based on the resource selection information set, a loss function value is determined. The loss function value is used to indicate the degree of concentration of computing resources output by the routing network within the preset time period. The magnitude of the loss function value is directly proportional to the magnitude of the concentration. The routing network is tuned based on the loss function value.

[0075] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described task processing method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0076] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the task processing method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.

[0077] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0078] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0079] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A task processing method, characterized in that, The method, applied to an inference device on which a target model is deployed, includes: Receive a task to be processed, the task to be processed includes at least one modal data, wherein different modal data in the at least one modal data correspond to different modalities, and the modality is used to indicate the data type of the corresponding modal data; Determine at least one modality importance score corresponding one-to-one with the at least one modality data. The modality importance score is used to indicate the degree of difference in the output results of the target model when processing the task to be processed before and after the corresponding modality data is masked in the target model. The value of the modality importance score is proportional to the magnitude of the difference. The at least one modal data is quantified by word contribution to obtain at least one word contribution set corresponding to the at least one modal data. The first word contribution set includes at least one word contribution corresponding to at least one word obtained by splitting the first modal data. The word contribution is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding word is masked in the target model. The value of the word contribution is proportional to the magnitude of the difference. The first word contribution set is any word contribution set in the at least one word contribution set. The first modal data is the modal data in the at least one modal data that corresponds to the first word contribution set. Based on the at least one modality importance score and the at least one lexical contribution set, a target routing strategy is determined, wherein the target routing strategy is used to indicate that the target computing resources in the target model are used to process the task to be processed; Based on the target routing strategy, the target model is invoked to process the task to be processed.

2. The method according to claim 1, characterized in that, The determination of the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set includes: Based on the at least one modality importance score, a first threshold, and a second threshold, at least one modality value level corresponding to the at least one modality data is determined. The modality value level includes a first value level, a second value level, and a third value level. Specifically, if the first modality importance score corresponding to the first modality data is less than or equal to the first threshold, the first modality value level corresponding to the first modality data is the first value level; if the first modality importance score is greater than the first threshold and less than the second threshold, the first modality value level is the second value level; and if the first modality importance score is greater than the second threshold, the first modality value level is the third value level. The target routing policy is determined based on the at least one modal value level, wherein the target routing policy includes: Mask the modal data of the first value level among the at least one modal value level; The modal data with the modal level of the third value level are processed using the target computing resources; The target lexical is processed using the target computing resources. The target lexical is a lexical determined from at least one lexical based on the first lexical contribution set harmonics and a preset contribution threshold, when the first modality level is the second value level. The lexical contribution of the target lexical is greater than the preset contribution threshold.

3. The method according to claim 1, characterized in that, Before determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: When the modality of the first modal data is a text modality, the contribution of the first word is determined based on the self-attention weight and similarity parameter of the first word. The first word is any one of the at least one word. The similarity parameter is the semantic similarity between the data after the first word is removed from the first modal data and the first modal data. When the modality of the first modal data is an image modality, the word contribution of the first word is determined based on the feature variance of the first word in the foreground region of the first modal data. The magnitude of the feature variance is proportional to the magnitude of the word contribution of the first word. The first word is an image obtained by splitting the foreground region of the first modal data. When the modality of the first modal data is video modality, frame extraction is performed on the first modal data to obtain multiple video frames, and multiple structural similarities corresponding to the multiple video frames are calculated one-to-one. The N video frames corresponding to the top N structural similarities with the largest structural similarity are determined as N target video frames. The foreground regions of the N target video frames are split into N word sets, and the at least one word includes all words in the N word sets. Based on the feature variance of the first word in the foreground region of the corresponding target video frame, the word contribution of the first word is determined, where N is a positive integer.

4. The method according to claim 1, characterized in that, The step of determining at least one modality importance score corresponding one-to-one with the at least one modality data, based on the at least one modality data, includes: The feature norm, feature information entropy, and preset attention weight of the first modality data are obtained. The feature norm is the norm calculated by performing L2 norm on the feature matrix corresponding to the first modality data. The feature information entropy is used to indicate the information richness of the first modality data. The target feature is obtained by concatenating the feature norm, the feature information entropy, and the preset attention weight. The target features are input into a gating network for modal importance evaluation, and the modal importance score output by the gating network corresponding to the first modal data is obtained.

5. The method according to any one of claims 1 to 4, characterized in that, The target model includes multiple expert computing resources; Before determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: At least one modal feature corresponding to the at least one modal data is obtained. When the modality of the first modal data is text, the modal feature of the first modal data is the mean and variance corresponding to the at least one word. When the modality of the first modal data is an image, the modal feature of the first modal data is the proportion of the foreground of the first modal data in the image corresponding to the first modal data. When the modality of the first modal data is video, the modal feature of the first modal data is the average proportion of the foreground of the video frame obtained by splitting the first modal data in the video frame obtained by splitting the first modal data. The at least one modal feature is input into the routing network for resource allocation, resulting in at least one expert computing resource output by the routing network that corresponds one-to-one with the at least one modal data. The target computing resource includes the at least one expert computing resource. The routing network is used to determine multiple current selection probabilities corresponding one-to-one with the multiple expert computing resources based on the at least one modal feature, and to determine the target selection probability of each expert computing resource based on the current selection probability and historical selection probability of each expert computing resource, and to output the at least one expert computing resource based on the target selection probability of each expert computing resource.

6. The method according to claim 5, characterized in that, After determining the target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, the method further includes: Obtain a resource selection information set, which includes multiple resource selection information corresponding one-to-one with the multiple expert computing resources. The resource selection information includes: the number of times the corresponding expert computing resource is selected within a preset time period, and the modality of the data processed when the corresponding expert computing resource is selected. The preset time period includes the time when the target computing resource is determined. Based on the resource selection information set, a loss function value is determined. The loss function value is used to indicate the degree of concentration of computing resources output by the routing network within the preset time period. The magnitude of the loss function value is directly proportional to the magnitude of the concentration. The routing network is tuned based on the loss function value.

7. A task processing device, characterized in that, An inference device for deploying a target model, the device comprising: A receiving module is used to receive a task to be processed, the task to be processed including at least one modal data, wherein different modal data in the at least one modal data correspond to different modalities, and the modality is used to indicate the data type of the corresponding modal data; The first determining module is used to determine at least one modal importance score corresponding one-to-one with the at least one modal data. The modal importance score is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding modal data is masked in the target model. The value of the modal importance score is proportional to the magnitude of the difference. A contribution metric module is used to metric the lexical contribution of the at least one modal data to obtain at least one lexical contribution set corresponding to the at least one modal data. The first lexical contribution set includes at least one lexical contribution corresponding to at least one lexical obtained by splitting the first modal data. The lexical contribution is used to indicate the degree of difference between the output results of the target model in processing the task to be processed before and after the corresponding lexical is masked in the target model. The value of the lexical contribution is proportional to the magnitude of the difference. The first lexical contribution set is any lexical contribution set in the at least one lexical contribution set. The first modal data is the modal data in the at least one modal data that corresponds to the first lexical contribution set. The second determining module is used to determine a target routing strategy based on the at least one modality importance score and the at least one lexical contribution set, wherein the target routing strategy is used to indicate that the target computing resources in the target model are used to process the task to be processed; The processing module is used to process the task to be processed by calling the target model based on the target routing strategy.

8. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the task processing method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the task processing method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the task processing method as described in any one of claims 1 to 6.