Artificial intelligence-based task processing method and device, computer equipment and medium

By using a dynamic feedback-driven multimodal fusion attention mechanism, the problem of fixed modality weights in traditional multimodal data processing is solved, enabling adaptive adjustment of modality weights and improving the robustness and output accuracy of multimodal tasks.

CN120850191APending Publication Date: 2025-10-28PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510716472.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In traditional multimodal data processing methods, modality weights are fixed and cannot be adaptively adjusted according to the dynamic characteristics of the input data. This makes it difficult to effectively capture complementary information between different modalities in complex tasks, and limits the accuracy of the output results.

Method used

A dynamic feedback-driven multimodal fusion attention mechanism based on artificial intelligence is adopted. The performance monitoring module generates feedback vectors, dynamically adjusts modality weights, and performs feature fusion and decision processing through the fusion layer and decision layer to form a closed-loop optimization system.

Benefits of technology

It significantly improves the robustness and adaptability of multimodal task processing, enhances the accuracy and reliability of output results, and can adaptively adjust modal weights to cope with dynamic changes in input data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850191A_ABST
    Figure CN120850191A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a task processing method and device based on artificial intelligence, computer equipment and a storage medium, and the method comprises the steps: receiving multi-modal data of a target task of a current batch; performing feature extraction on the multi-modal data to obtain multi-modal features; generating a feedback vector corresponding to the specified task of the previous batch based on a performance monitoring module; carrying out weight calculation on the multi-modal features based on the feedback vector through a fusion layer to obtain an attention weight; performing feature fusion processing on the multi-modal features based on the attention weight to obtain fused features; performing decision processing on the fusion features based on a decision layer, and generating an output result corresponding to the target task; and performing feedback processing based on the output result. In addition, the invention also relates to a block chain technology, and the output result can be stored in a block chain. The method can be applied to task processing scenes based on multi-modal data in the financial field and the medical field, and the accuracy of the generated output result can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and can be applied to fields such as fintech and digital healthcare, particularly to task processing methods, devices, computer equipment, and storage media based on artificial intelligence. Background Art

[0002] In the field of multimodal data processing, traditional fusion techniques mainly rely on static fusion methods (such as early fusion or late fusion), where modality weights are fixed during the training phase and cannot be adaptively adjusted according to the dynamic characteristics of the input data. This rigid fusion mechanism makes it difficult for the system to effectively capture complementary information between different modalities when processing complex multimodal tasks (such as image caption generation, video content understanding, and sentiment analysis), significantly limiting the accuracy of the output results. Specifically, static fusion methods typically assign fixed weights to modalities such as visual, audio, and text, ignoring the dynamic changes in modal contributions under different task scenarios.

[0003] For example, in customer risk assessment scenarios in the financial sector, traditional methods may only rely on static feature splicing based on textual modalities (such as questionnaires filled out by customers) and visual modalities (such as ID photos), without dynamically adjusting the weight of audio modalities (such as emotional fluctuations in the customer's voice). If a customer exhibits signs of anxiety or withholding information in their voice, the static fusion method may distort the risk assessment results by ignoring the potential risk signals of this modality, thereby affecting the accuracy of credit decisions.

[0004] Similar problems exist in the healthcare field. Taking disease assessment as an example, traditional static fusion methods may perform fixed-weighted analysis of a patient's medical images (such as X-rays and CT scans), electronic medical records, and some basic physiological indicators. However, when faced with certain complex conditions, the patient's verbal expressions (such as the way they describe symptoms and changes in tone of voice) may contain important diagnostic clues. If these modal information are always processed with fixed weights during disease assessment, dynamic information arising from changes in the patient's condition may be overlooked, such as the urgency of the patient's pain description or changes in emotional state, which may indicate the severity of the condition or the risk of potential complications. Static fusion methods cannot dynamically adjust modal weights according to individual patient differences and disease progression, which can easily lead to inaccurate assessment results, delays in treatment, and adverse effects on the patient's health.

[0005] Therefore, there is an urgent need to propose a dynamic multimodal fusion technology to improve the accuracy of multimodal task processing, thereby meeting the pressing needs of finance and other fields for intelligent decision-making systems. Summary of the Invention

[0006] The purpose of this application is to propose a task processing method, apparatus, computer device, and storage medium based on artificial intelligence, so as to solve the technical problem that the existing multimodal task processing methods based on static fusion methods have low accuracy of output results.

[0007] Firstly, an artificial intelligence-based task processing method is provided, including:

[0008] Receive multimodal data of the target task in the current batch; wherein, the multimodal data includes visual data, audio data, and text data;

[0009] Feature extraction is performed on the multimodal data to obtain corresponding multimodal features; wherein, the multimodal features include visual features, audio features, and text features;

[0010] The preset performance monitoring module generates a feedback vector corresponding to the specified tasks in the previous batch.

[0011] Through a preset fusion layer, the corresponding attention weights are obtained by weighting the multimodal features based on the feedback vector.

[0012] Based on the attention weights, the multimodal features are fused to obtain the corresponding fused features;

[0013] The fused features are processed based on a preset decision layer to generate an output result corresponding to the target task.

[0014] Feedback processing is performed based on the output results.

[0015] Secondly, an artificial intelligence-based task processing device is provided, comprising:

[0016] The receiving module is used to receive multimodal data of the target task in the current batch; wherein, the multimodal data includes visual data, audio data and text data;

[0017] The extraction module is used to extract features from the multimodal data to obtain corresponding multimodal features; wherein, the multimodal features include visual features, audio features, and text features;

[0018] The generation module is used to generate feedback vectors corresponding to the specified tasks in the previous batch based on the preset performance monitoring module.

[0019] The calculation module is used to calculate the corresponding attention weights by performing weight calculation on the multimodal features based on the feedback vector through a preset fusion layer;

[0020] The fusion module is used to perform feature fusion processing on the multimodal features based on the attention weights to obtain the corresponding fused features;

[0021] The decision module is used to perform decision processing on the fused features based on a preset decision layer, and generate an output result corresponding to the target task;

[0022] The feedback module is used to perform feedback processing based on the output results.

[0023] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based task processing method.

[0024] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described task processing method based on artificial intelligence.

[0025] In the above-mentioned scheme implemented by the AI-based task processing method, apparatus, computer equipment, and storage medium, the following steps are taken: First, multimodal data of the target task in the current batch is received; wherein, the multimodal data includes visual data, audio data, and text data; then, feature extraction is performed on the multimodal data to obtain corresponding multimodal features; wherein, the multimodal features include visual features, audio features, and text features; then, a feedback vector corresponding to the specified task in the previous batch is generated based on a preset performance monitoring module; subsequently, through a preset fusion layer, the multimodal features are weighted based on the feedback vector to obtain corresponding attention weights; and feature fusion processing is performed on the multimodal features based on the attention weights to obtain corresponding fused features; further, decision processing is performed on the fused features based on a preset decision layer to generate an output result corresponding to the target task; finally, feedback processing is performed based on the output result. This application receives multimodal data of the target task in the current batch, extracts features from the multimodal data to obtain corresponding multimodal features, generates a feedback vector corresponding to the specified task in the previous batch based on the performance monitoring module, calculates attention weights based on the feedback vectors using the fusion layer, performs feature fusion processing based on the attention weights to obtain fused features, and then performs decision processing on the fused features using the decision layer to generate output results corresponding to the target task. Finally, feedback processing is performed based on the output results. This application uses the combined use of the performance monitoring module, fusion layer, and decision layer to perform decision processing on the target task. It can achieve dynamic optimization of multimodal data fusion through real-time feedback, thus overcoming the limitations of traditional static fusion methods. It can adaptively adjust the attention weights of each modality according to changes in the quality of the input data and task requirements, significantly improving the robustness and adaptability of multimodal tasks in complex environments. This allows the subsequent decision processing using the fused features generated based on dynamically adjusted attention weights by the decision layer to effectively improve the accuracy and reliability of the generated output results. Attached Figure Description

[0026] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is an exemplary system architecture diagram to which this application can be applied;

[0028] Figure 2This is a flowchart of an embodiment of the AI-based task processing method according to this application;

[0029] Figure 3 This is a schematic diagram of a structure of an embodiment of the AI-based task processing device according to this application;

[0030] Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0032] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0034] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0035] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0036] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0037] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0038] It should be noted that the AI-based task processing method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the AI-based task processing device is generally located in the server / terminal device.

[0039] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0040] Continue to refer to Figure 2 The flowchart illustrates an embodiment of the AI-based task processing method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different needs. The AI-based task processing method provided in this application can be applied to any scenario requiring task processing, and thus can be applied to products in these scenarios, such as multimodal task processing scenarios based on multimodal data in the financial and medical fields. The AI-based task processing method includes the following steps:

[0041] Step S201: Receive multimodal data of the target task in the current batch; wherein the multimodal data includes visual data, audio data and text data.

[0042] In this embodiment, the artificial intelligence-based task processing method runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire multimodal data of the target task in the current batch via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future-developed wireless connection methods. The executing entity of this application is a data processing system, which can be simply referred to as the system. The system can receive raw data of three modalities related to the target task in the current batch input by the user—visual data (images / visions), audio data (waveforms / spectral graphs), and text data (natural language sequences).

[0043] This application proposes a dynamic feedback-driven multimodal fusion attention mechanism. Its core innovation lies in achieving dynamic optimization of multimodal data fusion through real-time performance feedback and online learning techniques. This mechanism overcomes the limitations of traditional static fusion methods, adaptively adjusting the weights of each modality based on changes in input data quality and task requirements, significantly improving the robustness and adaptability of the multimodal system.

[0044] The system's workflow includes: First, the system receives multimodal input data (visual, audio, and text) and generates feature vectors for each modality through a feature extraction layer. These feature vectors are then input into a dynamically feedback-driven fusion layer, the core component of which is an adaptive attention gate (AAG). The AAG dynamically adjusts the attention weights of each modality based on feedback signals provided by a performance monitoring module, generating a fused feature representation. The fused features are then passed to the decision layer to generate the final output. Simultaneously, the performance monitoring module evaluates the output quality and sends feedback signals back to the AAG, forming a closed-loop optimization system.

[0045] This application can be applied to various scenarios in the financial and medical fields, such as video content understanding, image description generation, and sentiment analysis. For example, in the financial field, the above-mentioned target task can correspond to the task scenario of video review for insurance claims. Data input includes: Visual modality: Videos or photos of the vehicle accident scene submitted by the user (e.g., collision location, license plate number). Audio modality: Voice recordings of the user describing the accident on-site (e.g., "The vehicle was rear-ended, and the rear bumper is damaged"). Text modality: Claim application form filled out by the user (e.g., accident time, location, and liability determination). Tasks to be processed (target task) include: Multimodal consistency verification: Verifying whether the vehicle damage in the video is consistent with the user's description (e.g., the audio mentions "damaged rear bumper," but no obvious damage is seen visually). Checking whether the liability determination in the claim form contradicts the video evidence (e.g., the user claims full responsibility, but the video shows the other party illegally changing lanes). Automatic damage assessment and fraud detection: Identifying damage types (e.g., scratches, dents) through the visual modality and combining vehicle model information from the text modality to automatically calculate repair costs. Potential fraud risks can be identified through audio sentiment analysis (such as a user's tense tone) and text semantic analysis (such as vague descriptions).

[0046] In the medical field, the above-mentioned target tasks correspond to the task scenario of telemedicine multimodal diagnosis. Data inputs include: Visual modality: photos / videos of wounds or rashes taken by the patient (e.g., degree of redness and swelling, ulcer morphology). Audio modality: audio recordings of the patient describing symptoms (e.g., "pain lasting for three days, worsening at night"). Text modality: electronic medical records (e.g., past medical history, allergy history) or symptom questionnaires completed by the patient. Tasks to be processed (target tasks) include: Multimodal symptom integration: identifying skin lesion types (e.g., eczema, shingles) through the visual modality, combining symptom descriptions from the audio (e.g., itching, pain) and allergy history from the text to assist doctors in diagnosis. Verifying consistency between the patient's description (audio) and visual evidence (e.g., rash morphology) (e.g., the patient claims "no external injury," but photos show obvious scratches). Urgency level assessment: automatically determining whether emergency referral is necessary based on the severity of the wound in the visual modality (e.g., amount of bleeding, signs of infection) and emotional analysis in the audio (e.g., the patient's anxious tone).

[0047] Step S202: Extract features from the multimodal data to obtain corresponding multimodal features; wherein, the multimodal features include visual features, audio features, and text features.

[0048] In this embodiment, the specific implementation process of extracting features from the multimodal data to obtain the corresponding multimodal features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here. Specifically, in the financial insurance field, the aforementioned text features may correspond to features of claims data, which may include data such as claims amount, claims event description, and customer claims risk. In the healthcare field, the aforementioned text features may correspond to features of medical data, which may include data such as personal health records, prescriptions, and examination reports.

[0049] Step S203: Generate a feedback vector corresponding to the specified tasks in the previous batch based on the preset performance monitoring module.

[0050] In this embodiment, the designated task refers to the task in the previous batch that is adjacent to the target task in the current batch, and the designated task has the same task type as the target task. Furthermore, the specific implementation process of generating the feedback vector corresponding to the designated task in the previous batch based on the preset performance monitoring module will be described in further detail in subsequent embodiments of this application, and will not be elaborated upon here.

[0051] Step S204: Through a preset fusion layer, the corresponding attention weights are obtained by weighting the multimodal features based on the feedback vector.

[0052] In this embodiment, the specific implementation process of calculating the corresponding attention weights by weighting the multimodal features based on the feedback vector through the preset fusion layer will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0053] Step S205: Perform feature fusion processing on the multimodal features based on the attention weights to obtain the corresponding fused features.

[0054] In this embodiment, the specific implementation process of performing feature fusion processing on the multimodal features based on the attention weights to obtain the corresponding fused features will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0055] Step S206: Based on a preset decision layer, perform decision processing on the fused features to generate an output result corresponding to the target task.

[0056] In this embodiment, the selection of the decision-making layer is not specifically limited and can be determined according to actual business needs. For example, a model such as GPT-3 can be used. The fused features can be passed to the decision-making layer, which will process the fused features based on a task processing strategy matching the target task and generate corresponding output results. For example, if the target task is a video content understanding task, the generated output result can be the final understanding result of the video content, such as the video's theme, key plot points, and character relationships.

[0057] Step S207: Perform feedback processing based on the output results.

[0058] In this embodiment, the specific implementation process of feedback processing based on the output result will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0059] This application receives multimodal data of the target task in the current batch, extracts features from the multimodal data to obtain corresponding multimodal features, generates a feedback vector corresponding to the specified task in the previous batch based on the performance monitoring module, calculates attention weights based on the feedback vectors using the fusion layer, performs feature fusion processing based on the attention weights to obtain fused features, and then performs decision processing on the fused features using the decision layer to generate output results corresponding to the target task. Finally, feedback processing is performed based on the output results. This application uses the combined use of the performance monitoring module, fusion layer, and decision layer to perform decision processing on the target task. It can achieve dynamic optimization of multimodal data fusion through real-time feedback, thus overcoming the limitations of traditional static fusion methods. It can adaptively adjust the attention weights of each modality according to changes in the quality of the input data and task requirements, significantly improving the robustness and adaptability of multimodal tasks in complex environments. This allows the subsequent decision processing using the fused features generated based on dynamically adjusted attention weights by the decision layer to effectively improve the accuracy and reliability of the generated output results.

[0060] In some alternative implementations, the closed-loop optimization system constructed in this application forms a feedback regulation loop through the interaction between the performance monitoring module and AAG. This design draws inspiration from the feedback mechanisms of biological sensory systems, such as internal model control, which can automatically adjust the processing strategy based on the output quality.

[0061] Specifically, the closed-loop optimization system forms a feedback regulation loop through the interaction between the performance monitoring module and AAG. This design draws inspiration from the feedback mechanisms of biological sensory systems, such as internal model control, which can automatically adjust the processing strategy based on the output quality. Specifically, after the system output is generated, the performance monitoring module evaluates its quality and generates a feedback signal f. t This signal is sent back to AAG to adjust the attention weights for each modality. For example, if the quality of the visual modality deteriorates (e.g., due to insufficient light causing image blurring), the performance monitoring module will detect the deterioration in output quality, generate a corresponding feedback signal, and cause AAG to reduce the weight of the visual modality and increase the weight of other modalities (such as audio and text) to ensure the quality of the system output.

[0062] This closed-loop optimization mechanism enables the system to adaptively respond to dynamic changes in input data without human intervention, greatly improving the system's robustness and reliability.

[0063] In some optional implementations, the performance monitoring module includes a metric evaluator and a feedback generator; step S203 includes the following steps:

[0064] The performance metrics corresponding to the specified task are calculated based on the metric evaluator.

[0065] In this embodiment, the performance monitoring module is a pre-built module responsible for evaluating the quality of the system output and generating a feedback vector f. t This performance monitoring module consists of two sub-components: Metric Evaluator and Feedback Generator.

[0066] Among them, the Metric Evaluator is used to calculate task-related performance metrics. t Examples of metrics include accuracy in classification tasks and BLEU scores (Bilingual Evaluation Replacement, used to assess the similarity between generated and reference text) in generation tasks. The performance monitoring module also receives external feedback signals. t Such as user ratings or suggestions for improvement.

[0067] Specifically, the performance metrics of the specified task can be calculated and processed based on the use of the aforementioned metric evaluator, and corresponding performance metrics can be generated.

[0068] Obtain the feedback signal corresponding to the specified task.

[0069] In this embodiment, the feedback signal may be a user rating or correction opinion related to the specified task received by the performance monitoring module from external input.

[0070] Obtain the preset feedback calculation formula.

[0071] In this embodiment, the above feedback calculation formula refers to the feedback vector f t The calculation formula is as follows:

[0072] f t =ReLU(W f ·Concat(Metric t Feedback t )+b f )

[0073] Here, Concat represents a vector concatenation operation, which concatenates performance metrics and external feedback (i.e., feedback information) into a single vector. W f and b f These are learnable parameters. ReLU (Rectified Linear Unit) is an activation function, defined as ReLU(x) = max(0,x), which sets negative values ​​to 0 and keeps positive values ​​unchanged.

[0074] Based on the feedback generator, the performance index and the feedback signal are calculated and processed using the feedback calculation formula to obtain the corresponding calculation results.

[0075] In this embodiment, the above-mentioned performance indicators and feedback signals can be substituted into the above-mentioned feedback calculation formula by using the above-mentioned feedback generator, and the calculation result can be used as the final feedback vector.

[0076] The calculation result is used as the feedback vector.

[0077] In this embodiment, to better understand the context of the input data and task requirements, the performance monitoring module uses a lightweight Transformer encoder to handle performance metrics and external feedback. The Transformer encoder is based on a multi-head self-attention mechanism, which can capture long-term dependencies in the input data. The calculation formula for the multi-head self-attention mechanism is as follows:

[0078] MultiHead(Q,K,V)=Concat(head1,…,head h W O

[0079] in, Q, K, and V represent the query, key, and value matrices, respectively. and W O It is a learnable parameter matrix. K T This represents the transpose of matrix K, which involves swapping the rows and columns of the matrix. `softmax` is a function that normalizes a vector, defined as... This makes the sum of all elements of the output vector equal to 1. x d represents the natural constant e (approximately 2.718) raised to the power of x. k It is the dimension of the key vector, used to scale the dot product and prevent gradient vanishing.

[0080] This application calculates performance metrics corresponding to a specified task based on the metric evaluator; then obtains feedback signals corresponding to the specified task; and obtains a preset feedback calculation formula. Subsequently, based on the feedback generator, the performance metrics and feedback signals are processed using the feedback calculation formula to obtain corresponding calculation results. These calculation results are then used as the feedback vector. This application calculates performance metrics corresponding to a specified task using the metric evaluator, obtains feedback signals corresponding to the specified task, and obtains a preset feedback calculation formula. Then, based on the feedback generator, the performance metrics and feedback signals are processed according to the feedback calculation formula, thereby achieving efficient and accurate calculation and generation of feedback vectors corresponding to the previous batch of specified tasks, effectively improving the generation efficiency and accuracy of feedback vectors. Furthermore, the use of the generated feedback vectors corresponding to the previous batch of specified tasks can automatically and intelligently adjust the attention weights of the multimodal features of the current batch of target tasks, enabling the system to adaptively respond to dynamic changes in input data, thereby improving the quality and robustness of the output results of the subsequently generated target tasks.

[0081] In some optional implementations of this embodiment, the fusion layer includes an adaptive attention gate; step S204 includes the following steps:

[0082] Invoke the weight generation formula corresponding to the feedback vector.

[0083] In this embodiment, the aforementioned fusion layer can also be referred to as a dynamically feedback-driven fusion layer, the core component of which is an adaptive attention gate (AAG). AAG is the core component for implementing dynamic multimodal fusion in this application; its function is to dynamically adjust the attention weights of each modality feature through real-time feedback signals. For each modality m (e.g., visual, audio, text), at time step t, its attention weight... Calculated using the following weight generation formula:

[0084]

[0085] in, f is the feature vector of modality m at time step t. For example, the feature vector of the visual modality might be the feature extracted by ResNet-152. t This is a feedback vector generated by the performance monitoring module, used to reflect the quality of the current system output. W m and U m It is a learnable weight matrix, b m This is the bias term. σ is the sigmoid activation function, ensuring that the attention weights are in the range [0,1]. The sigmoid function is defined as... Maps the input to the interval (0,1). · Represents the matrix multiplication operator.

[0086] Attention weight This reflects the importance of modality m in the current task. The higher the weight, the greater the contribution of that modality to the final fused features.

[0087] Obtain the specified modal features.

[0088] In this embodiment, the specified modal feature is any one of the visual feature, the audio feature, and the text feature.

[0089] Based on the adaptive attention gate, the weight generation formula is used to calculate and process the specified modal features and the feedback vector to obtain the corresponding generated weights.

[0090] In this embodiment, based on the use of the adaptive attention gate, the specified modal features and feedback vector can be substituted into the weight generation formula for calculation, and the generated weights can be used as the specified attention weights corresponding to the specified modal features.

[0091] The generated weights are used as the specified attention weights corresponding to the specified modal features.

[0092] This application achieves efficient and accurate calculation of the specified attention weights corresponding to the feedback vector by calling the weight generation formula corresponding to the feedback vector; then obtaining the specified modal features, wherein the specified modal features are any one of the visual features, audio features, and text features; subsequently, based on the adaptive attention gate, the weight generation formula is used to calculate and process the specified modal features and the feedback vector to obtain the corresponding generated weights; subsequently, the generated weights are used as the specified attention weights corresponding to the specified modal features. This application, by calling the weight generation formula corresponding to the feedback vector and obtaining the specified modal features, and then using the adaptive attention gate to calculate and process the specified modal features and the feedback vector according to the weight generation formula, achieves efficient and accurate calculation and generation of the specified attention weights corresponding to the specified modal features, and completes the dynamic adjustment of the attention weights of each modality based on the feedback vector, effectively improving the generation efficiency and accuracy of attention weights.

[0093] In some alternative implementations, step S205 includes the following steps:

[0094] Call the preset linear transformation function.

[0095] In this embodiment, the aforementioned linear transformation function is specifically a mode-specific linear transformation function, used to project the features of each mode onto a common representation space, defined as follows: in, It is a weight matrix. It is the bias vector. ∑ represents the summation symbol, which sums all elements in the set.

[0096] The multimodal features are projected based on the linear transformation function to obtain the corresponding first processed features.

[0097] In this embodiment, the multimodal features can be projected onto the shared semantic space using the above-described linear transformation function, and the projected features can be used as the first processing features.

[0098] The first processing feature is weighted and fused based on the attention weights to obtain the corresponding second processing feature.

[0099] In this embodiment, the projected features, i.e. the first processed features, can be weighted and summed according to the attention weights mentioned above, and the corresponding processing result (the second processed features) can be obtained as the final fused features.

[0100] Specifically, after obtaining the attention weights for each modality, AAG calculates the fused feature representation Z. t as follows:

[0101]

[0102] The second processing feature is used as the fusion feature.

[0103] This application achieves automatic and accurate feature fusion processing of multimodal features by invoking a preset linear transformation function, then projecting the multimodal features based on the linear transformation function to obtain a corresponding first processed feature; subsequently, it performs weighted fusion processing on the first processed feature based on attention weights to obtain a corresponding second processed feature; and finally, it uses the second processed feature as the fused feature. This application obtains a first processed feature by projecting the multimodal features using a linear transformation function, then performs weighted fusion processing on the first processed feature based on attention weights, and finally uses the obtained second processed feature as the final fused feature.

[0104] In some alternative implementations, step S207 includes the following steps:

[0105] Determine whether the output result is a multimodal output.

[0106] In this embodiment, if the target task is a modal sentiment classification task, it is necessary to further ensure that the output results of different modalities (such as text and images) are semantically consistent to avoid deviations in the overall output due to errors or noise in a single modality. This can be achieved by performing quantitative analysis on the output results to determine whether the output result belongs to multimodal output. For example, if the number of output results is greater than 1, the output result is determined to be multimodal output; if the number of output results is equal to 1, the output result is determined not to be multimodal output.

[0107] For example, the target task described above is a modal sentiment classification task. The decision layer generates multimodal neglected features based on the fused features. For example: text output: generating text describing the sentiment (e.g., "The user exhibits positive sentiment") through the Transformer decoder. Image output: generating sentiment labels for images (e.g., "positive") through the classifier.

[0108] If so, obtain the preset consistency verification strategy.

[0109] In this embodiment, the consistency verification strategy includes: using the alignment processing of the CLIP model to achieve consistency verification. CLIP (Contrastive Language–Image Pre-training) is a pre-trained multimodal model that can measure the semantic similarity between text and images. The specific implementation steps include: 1) Text encoding: Inputting the generated text description (e.g., "positive sentiment") into the CLIP text encoder to generate a text feature vector. 2) Image encoding: Inputting the original image into the CLIP image encoder to generate an image feature vector. 3) Similarity calculation: Calculating the cosine similarity between the text feature vector and the image feature vector, with a similarity range of [-1, 1]. The closer the value is to 1, the stronger the semantic consistency. 4) Setting a similarity threshold according to task requirements. The processing logic for consistency verification is: if the cosine similarity is greater than or equal to the similarity threshold, it indicates that the text and image are semantically consistent, the consistency verification is passed, and the result is directly output. If the cosine similarity is less than the similarity threshold, it indicates an inter-modal conflict, and the consistency check fails, requiring further processing: a. Conflicting modality identification: Compare the individual confidence scores of text and image. Select the modality with the higher confidence score as the final output (e.g., if text has a higher confidence score, the text result will be used). b. Feedback mechanism: Feedback the conflict information to the performance monitoring module for subsequent adjustment of the AAG weights (e.g., reducing the weight of the image modality).

[0110] The output result is validated based on the validation strategy to obtain the corresponding validation result.

[0111] In this embodiment, the consistency verification of the output result can be performed according to the implementation steps included in the strategy content of the above consistency verification strategy, and a corresponding consistency verification result can be generated. The consistency verification result includes either passing the consistency verification or failing the consistency verification.

[0112] Based on the consistency verification result, the output result is adjusted accordingly to obtain the target output result.

[0113] In this embodiment, the output result can be adjusted according to the consistency verification result processing logic to generate a unified output, which is then used as the target output result (e.g., "positive"). The consistency verification result and conflict handling logic can be recorded for subsequent system optimization and analysis. Furthermore, by feeding back the consistency verification result (e.g., conflict information) to the performance monitoring module to influence the AAG weight adjustment, the quality of the fused features can be optimized in subsequent iterations.

[0114] Feedback processing is performed based on the target output result.

[0115] In this embodiment, feedback processing of the target output result can be completed by sending the target output result to the user. The user can be a designated person who inputs the multimodal data of the target task in the current batch.

[0116] This application determines whether the output result is a multimodal output; if so, it obtains a preset consistency verification strategy; then, based on the consistency verification strategy, it performs consistency verification on the output result to obtain the corresponding consistency verification result; subsequently, based on the consistency verification result, it performs corresponding adjustment processing on the output result to obtain the target output result; and then, based on the target output result, it performs feedback processing. When this application detects that the output result generated by the decision layer is a multimodal output, it automatically and intelligently performs consistency verification on the output result based on the use of the consistency verification strategy, and performs corresponding adjustment processing on the output result based on the obtained consistency verification result, and then performs feedback processing based on the obtained target output result. This effectively ensures the semantic consistency of the target output result, avoids output deviation due to noise or errors in a single modality, and improves the reliability of the target output result.

[0117] In some optional implementations of this embodiment, step S202 includes the following steps:

[0118] Based on a preset visual modality feature extraction model, feature extraction is performed on the visual data to obtain visual features corresponding to the visual data.

[0119] In this embodiment, the input multimodal data is M = M v M a M t , of which M v M represents the visual modal input (i.e., visual data). a M represents the audio modal input (i.e., audio data). t This represents text modal input (i.e., text data). Specifically, the visual modal feature extraction model mentioned above can employ a pre-trained ResNet-152 model. The visual data M can be extracted based on the ResNet-152 model. v The global feature vector in the image is usually taken from the output of the last fully connected layer or the pooling result of the intermediate layers, and is used as the corresponding visual feature F. v ,in, Representing visual features, or visual feature vectors, R represents the set of real numbers, and d v It is a dimension of visual features.

[0120] Based on a preset audio modality feature extraction model, feature extraction is performed on the audio data to obtain audio features corresponding to the audio data.

[0121] In this embodiment, the audio modal feature extraction model can specifically employ VGGish or a pre-trained Wav2Vec2 model. The audio data M can be processed according to the selected audio modal feature extraction model. a Perform spectral feature extraction or temporal feature extraction, and use it as the corresponding audio feature F. a ,in, This represents audio features, or audio feature vectors, where R represents the set of real numbers, and d... a It is a dimension of audio features.

[0122] Based on a preset text modality feature extraction model, features are extracted from the text data to obtain the text features corresponding to the text data.

[0123] In this embodiment, the text modality feature extraction model can specifically employ pre-trained language models such as BERT or RoBERTa. The selected text modality feature extraction model can be used to extract the text data M. t The text is processed to generate context-sensitive word vectors, which are then subjected to average pooling to obtain the final text features F. t ,in, Representing text features, or text feature vectors, R represents the set of real numbers, and d t It is a dimension of text features.

[0124] In addition, layer normalization can be performed on each modal feature (visual features, audio features, and text features) to eliminate dimensional differences.

[0125] This application extracts features from visual data based on a preset visual modality feature extraction model to obtain visual features corresponding to the visual data; extracts features from audio data based on a preset audio modality feature extraction model to obtain audio features corresponding to the audio data; and extracts features from text data based on a preset text modality feature extraction model to obtain text features corresponding to the text data. By using the visual modality feature extraction model to extract features from visual data, the audio modality feature extraction model to extract features from audio data, and the text modality feature extraction model to extract features from text data, this application can efficiently and accurately extract corresponding visual, audio, and text features, improving the efficiency of multimodal feature extraction and ensuring the accuracy and standardization of the obtained multimodal feature data.

[0126] In some optional implementations of this embodiment, before step S203, the electronic device may further perform the following steps:

[0127] Obtain preset online learning strategies.

[0128] In this embodiment, the online learning strategy includes the following: To achieve real-time optimization, this application employs Online Gradient Descent (OGD) technology to update the parameters of the Adaptive Attention Gate (AAG) and the performance monitoring module in the fusion layer. Unlike traditional batch gradient descent, OGD updates the model parameters immediately upon the arrival of each sample, resulting in a faster response speed.

[0129] The parameter update formula is:

[0130]

[0131] Where, θ t These are the parameters for the current time step, including W in AAG. m U m b m And so on. η is the learning rate, which controls the step size for updating the parameters. It is a loss function The gradient of Z with respect to parameter θ. t It is a fusion feature representation, y t This is the target output. Loss function. Depending on the specific task, for example, classification tasks may use cross-entropy loss, defined as:

[0132]

[0133] Where C is the number of categories, y t,i It is a one-hot representation of the true label. It is the probability distribution predicted by the model, obtained by fusing features Z. t The input to the decision layer is obtained, where log represents the natural logarithm function.

[0134] In addition, to prevent overfitting and improve the model's generalization ability, an L2 regularization term is added to the loss function:

[0135]

[0136] in, It is a task-related loss function (such as cross-entropy loss), where λ is the regularization coefficient. The L2 norm square of the parameters is defined as the sum of the squares of all parameters.

[0137] This online learning mechanism enables the system to dynamically adjust parameters based on real-time feedback, eliminating the need to retrain the entire model and greatly improving the system's response speed and adaptability.

[0138] The performance monitoring module is updated based on the online learning strategy.

[0139] In this embodiment, the parameter update process corresponding to the performance monitoring module can be performed based on the strategy content of the above-mentioned online learning strategy.

[0140] The parameters of the fusion layer are updated based on the online learning strategy.

[0141] In this embodiment, parameter updates corresponding to the fusion layer can be performed based on the strategy content of the aforementioned online learning strategy. Specifically, by employing online gradient descent technology to achieve real-time parameter updates, the latency problem of traditional batch training methods is overcome, enabling the system to quickly adapt to changes in task requirements without retraining the entire model. In continuous learning tests, the computational efficiency of this scheme is 3.2 times higher than that of traditional methods.

[0142] This application utilizes an online learning strategy to automatically and intelligently update the parameters of the performance monitoring module and the fusion layer. This enables the system to quickly adapt to changes in task requirements and dynamically adjust parameters based on real-time feedback, without needing to retrain the entire model, thus greatly improving the system's response speed and adaptability.

[0143] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.

[0144] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0145] Furthermore, the task processing method of this application has the following advantages:

[0146] 1. By employing a closed-loop feedback mechanism to adjust the multimodal fusion strategy in real time, the system can automatically respond to dynamic changes in input data quality (such as the degradation of visual features due to insufficient ambient light), significantly improving the system's robustness in complex environments. Experiments show that in scenarios with fluctuating modal quality, this approach achieves 28% higher accuracy than traditional static fusion methods.

[0147] 2. Employing online gradient descent technology enables real-time parameter updates, overcoming the latency issues of traditional batch training methods. This allows the system to quickly adapt to changes in task requirements without retraining the entire model. In continuous learning tests, this approach demonstrates a 3.2-fold improvement in computational efficiency compared to traditional methods.

[0148] 3. Drawing on the feedback regulation principle in neuroscience (such as the internal model control mechanism), human-like adaptability to environmental changes was achieved through biomimetic design, breaking through the forward propagation limitations of traditional attention mechanisms, and constructing a complete closed-loop system that includes quality assessment, feedback generation, and weight adjustment.

[0149] 4. This application adopts a modular design, which can seamlessly integrate existing feature extractors (such as ResNet, VGGish, RoBERTa) and decision modules (such as GPT-3), and is suitable for various application scenarios such as video content understanding, image description generation, and sentiment analysis, and is compatible with mainstream deep learning frameworks.

[0150] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0151] It should be emphasized that, to further ensure the privacy and security of the above output results, the output results can also be stored in a blockchain node.

[0152] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0153] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0154] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When executed, the program can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0155] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0156] Further reference Figure 3 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of an artificial intelligence-based task processing device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0157] like Figure 3 As shown, the AI-based task processing device 300 described in this embodiment includes: a receiving module 301, an extraction module 302, a generation module 303, a calculation module 304, a fusion module 305, a decision-making module 306, and a feedback module 307. Wherein:

[0158] The receiving module 301 is used to receive multimodal data of the target task in the current batch; wherein, the multimodal data includes visual data, audio data and text data;

[0159] The extraction module 302 is used to extract features from the multimodal data to obtain corresponding multimodal features; wherein, the multimodal features include visual features, audio features, and text features;

[0160] The generation module 303 is used to generate a feedback vector corresponding to the specified tasks in the previous batch based on the preset performance monitoring module.

[0161] The calculation module 304 is used to calculate the corresponding attention weights by performing weight calculation on the multimodal features based on the feedback vector through a preset fusion layer.

[0162] The fusion module 305 is used to perform feature fusion processing on the multimodal features based on the attention weights to obtain the corresponding fused features;

[0163] The decision module 306 is used to perform decision processing on the fused features based on a preset decision layer, and generate an output result corresponding to the target task;

[0164] Feedback module 307 is used to perform feedback processing based on the output result.

[0165] In some optional implementations of this embodiment, the performance monitoring module includes an indicator evaluator and a feedback generator; the generation module 303 includes:

[0166] The first calculation submodule is used to calculate the performance index corresponding to the specified task based on the index evaluator;

[0167] The first acquisition submodule is used to acquire the feedback signal corresponding to the specified task;

[0168] The second acquisition submodule is used to acquire the preset feedback calculation formula;

[0169] The second calculation submodule is used to calculate and process the performance index and the feedback signal based on the feedback generator and the feedback calculation formula to obtain the corresponding calculation results.

[0170] The first determining submodule is used to use the calculation result as the feedback vector.

[0171] In some optional implementations of this embodiment, the fusion layer includes an adaptive attention gate; the calculation module 304 includes:

[0172] The first calling submodule is used to call the weight generation formula corresponding to the feedback vector;

[0173] The third acquisition submodule is used to acquire specified modal features; wherein, the specified modal features are any one of the visual features, the audio features, and the text features;

[0174] The third calculation submodule is used to calculate and process the specified modal features and the feedback vector based on the adaptive attention gate and the weight generation formula to obtain the corresponding generated weights.

[0175] The second determining submodule is used to use the generated weights as the specified attention weights corresponding to the specified modal features.

[0176] In some optional implementations of this embodiment, the fusion module 305 includes:

[0177] The second calling submodule is used to call the preset linear transformation function;

[0178] The projection submodule is used to project the multimodal features based on the linear transformation function to obtain the corresponding first processed features;

[0179] The fusion submodule is used to perform weighted fusion processing on the first processed feature based on the attention weight to obtain the corresponding second processed feature;

[0180] The third determining submodule is used to use the second processing feature as the fusion feature.

[0181] In some optional implementations of this embodiment, the feedback module 307 includes:

[0182] The judgment submodule is used to determine whether the output result is a multimodal output;

[0183] The fourth submodule is used to obtain the preset consistency verification strategy if the condition is met.

[0184] The verification submodule is used to perform consistency verification on the output result based on the consistency verification strategy, and obtain the corresponding consistency verification result.

[0185] The adjustment submodule is used to perform corresponding adjustment processing on the output result based on the consistency verification result to obtain the target output result;

[0186] The feedback submodule is used to perform feedback processing based on the target output result.

[0187] In some optional implementations of this embodiment, the extraction module 302 includes:

[0188] The first extraction submodule is used to extract features from the visual data based on a preset visual modality feature extraction model to obtain visual features corresponding to the visual data.

[0189] The second extraction submodule is used to extract features from the audio data based on a preset audio modality feature extraction model to obtain audio features corresponding to the audio data.

[0190] The third extraction submodule is used to extract features from text data based on a preset text modality feature extraction model to obtain text features corresponding to the text data.

[0191] In some optional implementations of this embodiment, the artificial intelligence-based task processing device further includes:

[0192] The acquisition module is used to acquire preset online learning strategies;

[0193] The first update module is used to update the parameters of the performance monitoring module based on the online learning strategy.

[0194] The second update module is used to update the parameters of the fusion layer based on the online learning strategy.

[0195] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0196] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0197] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0198] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for task processing methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0199] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the artificial intelligence-based task processing method.

[0200] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0201] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based task processing method described above.

[0202] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0203] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A task processing method based on artificial intelligence, characterized in that, Includes the following steps: Receive multimodal data of the target task in the current batch; wherein, the multimodal data includes visual data, audio data, and text data; Feature extraction is performed on the multimodal data to obtain corresponding multimodal features; wherein, the multimodal features include visual features, audio features, and text features; The preset performance monitoring module generates a feedback vector corresponding to the specified tasks in the previous batch. Through a preset fusion layer, the corresponding attention weights are obtained by weighting the multimodal features based on the feedback vector. Based on the attention weights, the multimodal features are fused to obtain the corresponding fused features; The fused features are processed based on a preset decision layer to generate an output result corresponding to the target task. Feedback processing is performed based on the output results.

2. The task processing method based on artificial intelligence according to claim 1, characterized in that, The performance monitoring module includes an indicator evaluator and a feedback generator; the step of generating a feedback vector corresponding to the specified tasks in the previous batch based on a preset performance monitoring module specifically includes: The performance metrics corresponding to the specified task are calculated based on the metric evaluator. Obtain the feedback signal corresponding to the specified task; Obtain the preset feedback calculation formula; Based on the feedback generator, the performance index and the feedback signal are calculated and processed using the feedback calculation formula to obtain the corresponding calculation results; The calculation result is used as the feedback vector.

3. The task processing method based on artificial intelligence according to claim 1, characterized in that, The fusion layer includes an adaptive attention gate; the step of calculating the corresponding attention weights by weighting the multimodal features based on the feedback vector through the preset fusion layer specifically includes: Call the weight generation formula corresponding to the feedback vector; Obtain a specified modal feature; wherein the specified modal feature is any one of the visual feature, the audio feature, and the text feature; Based on the adaptive attention gate, the weight generation formula is used to calculate and process the specified modal features and the feedback vector to obtain the corresponding generated weights; The generated weights are used as the specified attention weights corresponding to the specified modal features.

4. The task processing method based on artificial intelligence according to claim 1, characterized in that, The step of performing feature fusion processing on the multimodal features based on the attention weights to obtain the corresponding fused features specifically includes: Call the preset linear transformation function; The multimodal features are projected based on the linear transformation function to obtain the corresponding first processed features; The first processed feature is weighted and fused based on the attention weight to obtain the corresponding second processed feature; The second processing feature is used as the fusion feature.

5. The task processing method based on artificial intelligence according to claim 1, characterized in that, The step of performing feedback processing based on the output result specifically includes: Determine whether the output result is a multimodal output; If so, obtain the preset consistency verification strategy; Based on the consistency verification strategy, the output result is verified to obtain the corresponding consistency verification result. Based on the consistency verification result, the output result is adjusted accordingly to obtain the target output result; Feedback processing is performed based on the target output result.

6. The task processing method based on artificial intelligence according to claim 1, characterized in that, The step of extracting features from the multimodal data to obtain the corresponding multimodal features specifically includes: Based on a preset visual modality feature extraction model, feature extraction is performed on the visual data to obtain visual features corresponding to the visual data. Based on a preset audio modality feature extraction model, feature extraction is performed on the audio data to obtain audio features corresponding to the audio data; Based on a preset text modality feature extraction model, features are extracted from the text data to obtain the text features corresponding to the text data.

7. The task processing method based on artificial intelligence according to claim 1, characterized in that, Before the step of generating a feedback vector corresponding to the specified tasks in the previous batch based on the preset performance monitoring module, the method further includes: Obtain preset online learning strategies; The performance monitoring module is updated based on the online learning strategy. The parameters of the fusion layer are updated based on the online learning strategy.

8. A task processing device based on artificial intelligence, characterized in that, include: The receiving module is used to receive multimodal data of the target task in the current batch; wherein, the multimodal data includes visual data, audio data and text data; The extraction module is used to extract features from the multimodal data to obtain corresponding multimodal features; wherein, the multimodal features include visual features, audio features, and text features; The generation module is used to generate feedback vectors corresponding to the specified tasks in the previous batch based on the preset performance monitoring module. The calculation module is used to calculate the corresponding attention weights by performing weight calculation on the multimodal features based on the feedback vector through a preset fusion layer; The fusion module is used to perform feature fusion processing on the multimodal features based on the attention weights to obtain the corresponding fused features; The decision module is used to perform decision processing on the fused features based on a preset decision layer, and generate an output result corresponding to the target task; The feedback module is used to perform feedback processing based on the output results.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the task processing method based on artificial intelligence as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based task processing method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Emotional state evaluation system and method based on multiple modes

    CN119214657A

  • Method and device for detecting anomalies, corresponding computer program and non-transitory computer-readable medium

    US20220277225A1