Explainable cross-media agent decision path tracking method

By employing a decision attention mechanism and a visualized decision path, the problems of transparency and interpretability in cross-media decision-making are solved, thereby improving the transparency and accuracy of the decision-making process of cross-media intelligent agents.

CN120632598BActive Publication Date: 2025-12-05UNIVERSAL UBIQUITOUS TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511122200.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-12-05
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing technologies lack transparency and interpretability in cross-media decision-making processes, making it difficult to understand the decision-making basis and path. In particular, the feature fusion process in multimodal data scenarios suffers from insufficient interpretability and imperfect adaptive adjustment mechanisms.

Method used

The weights of the features to be decided are determined by a decision attention mechanism, the features are processed by a cross-modal data fusion model, and each step of the processing is displayed by a visual decision path. The weights are adjusted by combining sparse attention and interpretability constraints, data changes are monitored in real time, and model parameters are dynamically updated to optimize the decision path.

Benefits of technology

It significantly improves the interpretability and transparency of cross-media decision-making processes, enhances the accuracy and adaptability of decision-making paths, and strengthens users' trust and reliability in the agent's decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632598B_ABST
    Figure CN120632598B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides an interpretable cross-media intelligent agent decision path tracking method, which comprises the following steps: receiving to-be-decided data corresponding to a target task, extracting data features, and obtaining to-be-decided features; determining to-be-decided weights of the to-be-decided features through a decision attention mechanism, processing the to-be-decided features through a trained cross-modal data fusion model and the to-be-decided weights, obtaining fusion features, and determining a classification label of the target task based on the fusion features; recording a fusion process of obtaining the fusion features and a classification process of obtaining the classification label step by step according to a processing sequence, and determining the fusion process and the classification process as a processing process; synchronizing the processing process to a visual decision path step by step according to the processing sequence, and displaying input data, output data and processing methods corresponding to each processing process through the visual decision path. The method effectively solves the problems in decision path tracking and interpretation, and significantly improves the interpretability and transparency of the process of obtaining the classification label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, specifically to an interpretable cross-media intelligent agent decision path tracing method. Background Technology

[0002] With the widespread application of artificial intelligence technology in the cross-media field, the decision-making process in processing cross-media information has become increasingly complex and lacks transparency, forming a typical "black box" mechanism that makes it difficult for users and R&D personnel to understand the decision-making basis and path.

[0003] In existing technologies, weights can be assigned by calculating feature correlation based on the cross-modal fusion method of attention mechanism, the classification process can be displayed by decision tree visualization technology, and model parameters can be dynamically updated by online learning mechanism. However, the cross-modal fusion method based on attention mechanism does not establish weight interpretability constraints, the decision tree visualization method is difficult to be compatible with heterogeneous feature representations of multimodal data, and the online learning mechanism lacks the ability to compare and analyze historical decision paths. As a result, the feature fusion process in cross-media scenarios has limitations such as insufficient interpretability, lack of completeness of decision paths, and imperfect adaptive adjustment mechanism. Summary of the Invention

[0004] To address the problems in existing technologies, this application provides an interpretable cross-media intelligent agent decision path tracking method, which can effectively solve the shortcomings of traditional technologies in decision path tracking and interpretation, and significantly improve the interpretability and transparency of the process of obtaining the classification label of the target task based on the decision data of the target task.

[0005] To solve at least one of the above problems, this application provides the following technical solution:

[0006] Firstly, this application provides an interpretable cross-media intelligent agent decision-making path tracing method, including:

[0007] Receive the decision-making data corresponding to the target task, extract the data features of the decision-making data, and obtain the decision-making features;

[0008] The decision weights of the features to be decided are determined by the decision attention mechanism. The features to be decided are then processed by the trained cross-modal data fusion model and the decision weights to obtain fused features. The classification label of the target task is then determined based on the fused features.

[0009] The fusion process of obtaining fusion features and the classification process of obtaining classification labels are recorded step by step according to the processing order, and the fusion process and the classification process are defined as the processing process;

[0010] The processing steps are synchronized to the visualized decision path in the order of processing. The visualized decision path displays the input data, output data and processing method corresponding to each step of the processing. The visualized decision path is constructed by a graph structure. The edges of the graph structure represent the data flow, processing method and processing weight in the processing. Each node of the graph structure corresponds to each step of the processing. The processing weight is obtained based on the weight to be decided and the target task.

[0011] Furthermore, the data to be decided includes text data, image data, and audio data, and also includes: mapping the text data into low-dimensional semantic vectors through a pre-trained word vector model to obtain the decision features corresponding to the text data;

[0012] Local and global features of image data are extracted by convolutional neural networks, and these local and global features are determined as the decision features corresponding to the image data.

[0013] The audio data is processed by short-time Fourier transform to generate a spectrogram corresponding to the audio data, and the features of the spectrogram are extracted by a recurrent neural network to obtain the decision features corresponding to the audio data.

[0014] The features to be decided include text features, image features, and audio features.

[0015] Furthermore, the features to be decided include text features, image features, and audio features, and also include: determining the correlation between the text features, image features, and audio features through sparse attention in the decision attention mechanism, and generating an initial weight matrix based on the correlation.

[0016] Significance analysis was performed on the initial decision weight matrix using interpretability constraints to obtain the decision weights corresponding to the text features, image features, and audio features to be decided.

[0017] Furthermore, after determining the classification label for the target task based on the fusion features, the process also includes:

[0018] The system monitors the statistical characteristics of the data to be decided in real time. When the statistical characteristics change, it identifies historical data to be decided that have a data similarity exceeding a preset data similarity threshold with the updated data to be decided. It also receives the historical classification labels corresponding to the historical data to be decided. The statistical characteristics include at least one of variance, mean, and distribution.

[0019] By comparing historical category labels with the category labels corresponding to the current data to be decided, if the label similarity between the historical category labels and the current category labels is less than the preset label similarity, the visual decision path is adjusted, and the current category label is updated based on the adjusted visual decision path.

[0020] Furthermore, after determining the classification label for the target task based on the fusion features, the process also includes:

[0021] Determine the confidence level of the classification label. If the confidence level is less than the preset confidence threshold, backtrack the step of determining the decision weight of the feature to be decided through the decision attention mechanism during the fusion process to obtain the updated decision weight.

[0022] Based on the updated decision weights, the model parameters of the cross-modal data fusion model are adjusted through gradient backpropagation, and the processing of the corresponding nodes in the visualized decision path is updated synchronously based on the updated model parameters.

[0023] Furthermore, it also includes: generating a corresponding natural language description based on the output data of each node, and displaying the natural language description on the display screen of the terminal device;

[0024] If the change in the processing weight corresponding to an edge exceeds a preset change threshold, the visual decision path corresponding to the edge is highlighted, and the reason for the change is displayed.

[0025] Furthermore, after displaying the input data, output data, and processing method corresponding to each step of the process through a visual decision path, it also includes:

[0026] Receive correction operations on classification labels and transform these correction operations into reward signals in the visualized decision path;

[0027] The update strategy of the visualization decision path is adjusted based on the near-end strategy optimization method, and the cross-modal data fusion model is updated based on the adjusted update strategy when the accumulated reward signal reaches the preset reward threshold.

[0028] Secondly, this application provides an interpretable cross-media intelligent agent decision-making path tracking device, comprising:

[0029] The extraction module is used to receive the decision data corresponding to the target task, extract the data features of the decision data, and obtain the decision features.

[0030] The classification module is used to determine the decision weights of the features to be decided through a decision attention mechanism, process the features to be decided through a trained cross-modal data fusion model and decision weights to obtain fused features, and determine the classification label of the target task based on the fused features.

[0031] The processing module is used to record the fusion process of obtaining fusion features and the classification process of obtaining classification labels step by step according to the processing order, and to define the fusion process and the classification process as the processing process;

[0032] The decision module is used to synchronize the processing process step by step to the visual decision path according to the processing order. The visual decision path displays the input data, output data and processing method corresponding to each step of the processing process. The visual decision path is constructed by a graph structure. The edges of the graph structure represent the data flow, processing method and processing weight in the processing process. Each node of the graph structure corresponds to each step of the processing process. The processing weight is obtained based on the weight to be decided and the target task.

[0033] Thirdly, this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the interpretable cross-media intelligent agent decision path tracing method.

[0034] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the explained cross-media intelligent agent decision path tracing method.

[0035] Fifthly, this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the described interpretable cross-media intelligent agent decision path tracing method.

[0036] As can be seen from the above technical solution, this application provides an interpretable cross-media intelligent agent decision path tracking method. It innovatively receives decision-making data corresponding to the target task and extracts data features from the decision-making data to obtain decision-making features. It determines the decision-making weights of the decision-making features through a decision attention mechanism, processes the decision-making features using a trained cross-modal data fusion model and the decision-making weights to obtain fused features and determine the classification label corresponding to the target task. Simultaneously, it records the feature fusion process and the classification process in the processing order and identifies them as processing steps. Through a visual decision path, it displays the input data, output data, and processing method corresponding to each step of the processing step in the processing order. The visual decision path is... The graph structure is constructed so that the edges represent the data flow, processing method, and processing weight in the processing process. Each node in the graph structure corresponds to each processing step, and the processing weight is obtained based on the decision weight and the target task. By visualizing the decision path, the input data, output data, and processing method of each processing step can be dynamically displayed, which improves the interpretability of the decision process. By using the decision attention mechanism and the edge weights of the graph structure to quantify the impact of each processing step on the decision, the transparency and adaptability of decision path tracking are enhanced. This effectively solves the shortcomings of traditional technologies in decision path tracking and interpretation, and significantly improves the interpretability and transparency of the process of obtaining the classification label of the target task based on the decision data of the target task. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating an interpretable cross-media intelligent agent decision path tracing method according to an embodiment of this application.

[0039] Figure 2 This is a structural diagram of an interpretable cross-media intelligent agent decision path tracking device according to an embodiment of this application;

[0040] Figure 3 This is a schematic diagram of the structure of the electronic device in the embodiments of this application.

[0041] Figure label:

[0042] Electronic device 9600, central processing unit 9100, memory 9140, communication module 9110, input unit 9120, audio processor 9130, display 9160, power supply 9170, buffer memory 9141, application / function storage unit 9142, data storage unit 9143, driver storage unit 9144, antenna 9111, speaker 9131, microphone 9132. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0044] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.

[0045] In existing technologies, with the widespread application of artificial intelligence in cross-media fields, the decision-making process of intelligent agents when processing cross-media information has become increasingly complex. However, the decision-making process of most cross-media intelligent agents is currently like a "black box," making it difficult to understand the basis and path of their decisions. For example, in some multimedia content recommendation systems, an intelligent agent decides to recommend a video or image to a user, but cannot clearly explain why the content was selected. This is unacceptable in fields that require high transparency in decision-making, such as medical diagnosis (involving cross-media data such as images) and financial risk assessment (involving the analysis of multiple data types). Meanwhile, with the continuous growth of cross-media data and the diversification of application scenarios, the need to track and explain the decision-making path of intelligent agents is becoming increasingly urgent.

[0046] To effectively address the shortcomings of traditional technologies in decision path tracking and interpretation, and to significantly improve the interpretability and transparency of obtaining the classification label of the target task from the decision-making data, this application provides an embodiment of an interpretable cross-media intelligent agent decision path tracking method. See [link to embodiment]. Figure 1 The explainable cross-media intelligent agent decision path tracing method specifically includes the following:

[0047] Step S101: Receive the decision data corresponding to the target task, extract the data features of the decision data, and obtain the decision features.

[0048] Optionally, this embodiment receives decision-making data corresponding to the target task, wherein the decision-making data includes data of multiple media types, and may include at least two of text data, image data and audio data.

[0049] In addition, feature extraction is performed on the data to be decided to obtain the features to be decided. During the feature extraction process, different levels of detailed information can be captured by multi-scale feature extraction. For example, when extracting features of the image to be decided, features such as the edge and texture of the image can be extracted by multi-scale convolution kernels, so as to more comprehensively represent the content of the image data. The judgment ability of features can also be enhanced by contrastive learning, making the feature differences between different media types more significant.

[0050] This embodiment enables the transformation of decision-making data of different media types into highly representative and semantically rich decision-making data features, thereby improving the agent's ability to understand and analyze target tasks and thus improving the accuracy and reliability of decision-making.

[0051] Step S102: Determine the decision weights of the features to be decided through the decision attention mechanism, process the features to be decided through the trained cross-modal data fusion model and the decision weights to obtain fused features, and determine the classification label of the target task based on the fused features.

[0052] Optionally, this embodiment determines the decision weight of the feature to be decided through a decision attention mechanism. The decision attention mechanism is used to determine the relevance of different features to the target task decision. Based on the calculated relevance, a weight coefficient is assigned to each feature to be decided. The higher the weight coefficient, the more important the feature to be decided is in the decision-making process. For example, in the traffic condition prediction task of an intelligent transportation system, if the traffic congestion data obtained by traffic cameras is highly correlated with the traffic condition classification, the corresponding feature to be decided will be given a larger weight.

[0053] In addition, the weighted features to be decided are input into a trained cross-modal data fusion model for processing. The cross-modal data fusion model is used to effectively integrate the features to be decided to obtain fused features.

[0054] Furthermore, to determine the classification label of the target task based on the fusion features, the fusion features can be input into a trained classification model so that the classification model outputs the classification label corresponding to the target task according to the pattern and regularity of the fusion features. The classification model can be a vector machine, a deep neural network classifier, etc.

[0055] For example, in a news recommendation system, the news category of a news story can be determined based on fusion features, thereby achieving accurate classification and recommendation of the news story. The news category includes, but is not limited to, sports news, financial news, and entertainment news.

[0056] Furthermore, a dynamic weight adjustment strategy can be incorporated into the decision-making attention mechanism. This involves dynamically adjusting the weights of the features to be decided based on real-time changes in the target task or changes in data distribution. For example, in social media data analysis where trending topics change frequently, the popularity and sentiment of topics can be monitored in real time, and the weights of relevant features to be decided can be dynamically adjusted to adapt to new needs in social media data judgment.

[0057] In addition, feature interaction modules can be added to cross-modal data fusion models through multimodal feature interaction enhancement, enabling different decision features to complement and enhance each other, thereby further improving the expressive power of fused features.

[0058] This embodiment achieves the acquisition of decision-making features through a decision attention mechanism, enabling the agent to pay more attention to feature information closely related to the target task during the decision-making process, thereby improving the accuracy and efficiency of decision-making. The cross-modal data fusion model can effectively integrate the advantages of different decision-making features, enrich the basis for decision-making, provide more comprehensive decision support, and thus improve the accuracy of classification labels.

[0059] Step S103: Record the fusion process of obtaining fusion features and the classification process of obtaining classification labels step by step according to the processing order, and define the fusion process and the classification process as the processing process;

[0060] Optionally, in the feature fusion stage, this embodiment records each fusion step in detail, including but not limited to the input of different features to be decided, the fusion operation (such as weighted summation, splicing, etc.), and the output after fusion.

[0061] For example, in a multimodal learning effectiveness evaluation task of an intelligent education tutoring system, the process of fusing text learning content features, voice answer features, and image writing features is recorded, including the algorithm, parameter settings, and feature changes after each fusion operation.

[0062] During the classification label determination stage, information such as the input fusion features of the classification model, the classification decision logic (e.g., the activation function and decision boundary of the classification model), and the output classification label are recorded. For example, the forward propagation process after the fusion features are input into the classification model is recorded, including but not limited to the output results of each layer, the classification probability distribution, and the output classification label.

[0063] Furthermore, the fusion and classification processes of the aforementioned records are integrated to form a complete processing procedure. This processing procedure describes in detail each step of the agent's operation from processing the data to be decided to obtaining the decision result, including the specific data processing steps, the algorithms and models used, and the corresponding input and output data.

[0064] Furthermore, timestamps and version control mechanisms can be introduced when recording the processing steps to accurately track their evolution and changes. Simultaneously, distributed data recording can spread the processing records across multiple nodes, improving data reliability and availability.

[0065] This embodiment achieves transparency in the decision-making of the intelligent agent by recording the processing process in detail, making the decision-making process traceable and explainable, improving the user's understanding of the basis for the intelligent agent's decision, and facilitating problem localization and analysis when deviations or errors occur in the decision.

[0066] Step S104: Synchronize the processing process step by step to the visual decision path according to the processing order, and display the input data, output data and processing method corresponding to each step of the processing process through the visual decision path. The visual decision path is constructed by a graph structure. The edges of the graph structure represent the data flow, processing method and processing weight in the processing process. Each node of the graph structure corresponds to each step of the processing process, and the processing weight is obtained based on the weight to be decided and the target task.

[0067] Optionally, in this embodiment, the recorded processing procedure is gradually synchronized to the visualized decision path according to its actual processing order. The visualized decision path is constructed using a graph structure, where each node corresponds to a step in the processing procedure, such as feature extraction, feature fusion, classification decision, etc.

[0068] In this graph structure, nodes are connected by edges, which represent the data flow direction during the processing, i.e., the path through which data is passed from one processing step to the next.

[0069] Edges in a graph structure can also represent processing methods, that is, the specific processes by which data is processed and transformed between two steps. For example, an edge between feature extraction and feature fusion represents the way in which extracted text features, image features, and audio features are fused using a specific fusion algorithm (such as weighted fusion based on attention mechanisms).

[0070] The edges of the graph structure can also represent processing weights, which are determined based on the weights to be decided and the target task. The weights to be decided reflect the importance of features in the decision-making process, while the target task further influences the specific allocation of weights.

[0071] For example, in a text-based news recommendation task, the decision weight of text features is relatively high, and the processing weight of text features can be further increased depending on the target task (recommending news that is highly relevant to user interests).

[0072] In addition, by visualizing the decision path, the input data (such as raw text, images, audio data or feature data after preliminary processing), output data (such as extracted feature vectors, fused feature vectors or classification labels) and processing methods (such as word vector conversion, convolutional neural network extraction, attention mechanism fusion, etc.) corresponding to each processing step can be displayed intuitively.

[0073] For example, users can see in the visualization interface that the text data corresponding to the input decision data is processed by lexical analysis and word vector transformation to obtain decision text features, which are then passed as input to the next processing step; the image data corresponding to the decision data is processed by a convolutional neural network to obtain decision image features, and similarly, the audio data corresponding to the decision data is processed to obtain decision audio features; then the decision text features, decision image features, and decision audio features are fused through an attention mechanism to obtain fused features, which are then input into a trained classification model to obtain the corresponding classification label.

[0074] This embodiment makes complex decision-making processes more intuitive and easier to understand by visualizing the decision-making path, enhancing users' trust and acceptance of the agent's decisions. At the same time, by displaying the input data, output data, and processing methods of each step, it provides users with a means to deeply analyze and verify the decision-making process, which helps to discover potential problems and optimize the decision-making process. It reduces the difficulty for users to understand the decision-making process, increases users' trust in the classification labels, and provides a visual basis for optimizing the decision-making process, making it easier for users to make targeted improvement suggestions.

[0075] In some embodiments, the data to be decided includes text data, image data, and audio data;

[0076] Extract the data features of the data to be decided, and obtain the decision features, including:

[0077] By mapping text data into low-dimensional semantic vectors through a pre-trained word vector model, the decision features corresponding to the text data are obtained.

[0078] Local and global features of image data are extracted by convolutional neural networks, and these local and global features are determined as the decision features corresponding to the image data.

[0079] The audio data is processed by short-time Fourier transform to generate a spectrogram corresponding to the audio data, and the features of the spectrogram are extracted by a recurrent neural network to obtain the decision features corresponding to the audio data.

[0080] Optionally, in the cross-media intelligent agent decision-making process of this embodiment, the data to be decided usually includes multiple media types, including text data, image data and audio data. Among them, text data can be various forms of natural language information, such as news articles, social media posts, user comments, and medical record descriptions; image data includes photos, video frames, medical images, and surveillance footage; and audio data can be speech, music, and ambient sounds.

[0081] In addition, the text data in the received decision-making data undergoes preprocessing operations, including but not limited to word segmentation and stop word removal. Then, each word or phrase in the text data can be mapped to a low-dimensional vector space using a pre-trained word vector model (such as Word2Vec, GloVe, or BERT) to obtain the corresponding word vector. The word vector captures the semantic information and contextual relationships of the words.

[0082] The word vectors in the text are averaged, concatenated, or pooled to obtain a low-dimensional semantic vector that can represent the semantics of the entire text, which serves as the decision feature corresponding to the text data.

[0083] Furthermore, when using pre-trained word vector models, they can be fine-tuned by incorporating relevant domain knowledge. For example, in medical text analysis, a word vector model pre-trained on a large-scale medical literature can be used and fine-tuned according to the specific medical task to better capture the semantic features of medical terms. In addition, multimodal pre-trained models can be employed, such as models jointly trained with text and other media data like images and audio, to enhance the cross-media relevance and expressive power of text features.

[0084] In addition, for image data in the decision-making data, features can be extracted using convolutional neural networks (CNNs). CNNs include multiple convolutional and pooling layers, which can automatically learn local features (such as edges, textures, shapes, etc.) and global features (such as overall layout and object categories) in the image.

[0085] For example, using the classic VGG network, its multiple convolutional layers can progressively extract local features of an image, while fully connected layers can extract global features.

[0086] The local and global features extracted by the convolutional neural network are integrated to determine the decision features corresponding to the image data.

[0087] Furthermore, based on convolutional neural networks, attention mechanisms can be introduced to make the network focus more on local regions in the image relevant to the target task, thereby enhancing the expressive power of local features. In addition, multi-scale feature extraction methods can be employed, using convolutional kernels of different sizes or performing feature extraction on image pyramids of different scales to capture feature information at different scales in the image, improving the comprehensive expressive power of image details and overall information.

[0088] Furthermore, for audio data, preprocessing is performed on the audio data in the decision-making data. This preprocessing includes, but is not limited to, sampling and normalization. Then, the audio signal corresponding to each preprocessed audio data can be converted from the time domain to the frequency domain using Short Time Fourier Transform (STFT) to generate a spectrogram of the audio data. The spectrogram can intuitively show the energy distribution of the audio signal at different times and frequencies.

[0089] Using a spectrogram as input, features are extracted from the spectrogram using a recurrent neural network (RNN) or its variants (such as LSTM and GRU) to obtain the decision features corresponding to the audio data. The RNN can capture time-series information in the spectrogram, learning the patterns and rules of audio signal changes over time. For example, in speech recognition tasks, extracting features from the spectrogram using a RNN can reveal the speech content information contained in the speech signal, thereby achieving speech recognition.

[0090] Furthermore, an architecture combining convolutional neural networks and recurrent neural networks can be used. First, convolutional neural networks are used to extract local features from the spectrogram, and then recurrent neural networks are used to capture time-series information, thereby further improving the expressive power of audio features.

[0091] This embodiment achieves feature extraction from text data, image data, and audio data separately. By using pre-trained word vector models, convolutional neural networks, and short-time Fourier transform combined with recurrent neural networks, high-quality decision features corresponding to the data to be decided can be obtained. This fully expresses the semantic information, visual content, and audio signal features in the data to be decided, enabling the fusion of cross-media data to be carried out in a unified feature space, and better discovering the correlation and complementary information between the data to be decided.

[0092] In some embodiments, the features to be decided include text features to be decided, image features to be decided, and audio features to be decided;

[0093] The decision weights of the features to be decided are determined through a decision attention mechanism, including:

[0094] The correlation between the text features, image features, and audio features to be decided is determined by sparse attention in the decision attention mechanism, and an initial decision weight matrix is ​​generated based on the correlation.

[0095] Significance analysis was performed on the initial decision weight matrix using interpretability constraints to obtain the decision weights corresponding to the text features, image features, and audio features to be decided.

[0096] Optionally, in this embodiment, the decision-making features obtained from text data are determined as decision-making text features, the decision-making features obtained from image data are determined as decision-making image features, and the decision-making features obtained from audio data are determined as decision-making audio features.

[0097] The correlation between text features, image features, and audio features to be decided is determined by sparse attention in the decision attention mechanism. The decision attention mechanism is a method for modeling the relationship between features, which can highlight features related to the decision objective. Sparse attention introduces sparsity constraints to make the attention weight distribution more concentrated, thereby more accurately capturing the correlation between features to be decided.

[0098] Specifically, the sparse attention mechanism determines the degree of mutual influence among the text features, image features, and audio features to be decided by calculating the similarity or correlation between them.

[0099] For example, in sentiment analysis tasks, there is a close correlation between the emotional words expressed in a text and intonation features, and this correlation can be quantified through sparse attention mechanisms.

[0100] In addition, an initial decision weight matrix is ​​generated based on the correlation between the text features, image features, and audio features to be decided. Each element in the initial decision weight matrix represents the correlation strength between the corresponding decision features, reflecting the potential role of each decision feature in the decision-making process.

[0101] For example, in the initial decision weight matrix, the text features and audio features have relatively high weight values, indicating that they have a strong synergistic effect in expressing emotions.

[0102] Furthermore, imposing interpretability constraints on the initial decision weight matrix can increase the interpretability of the decision-making process. These constraints require that the weight distribution in the initial decision weight matrix can be reasonably interpreted and is consistent with the decision logic.

[0103] Significance analysis is used to identify the importance of different features to the decision and adjust the weight matrix accordingly.

[0104] Furthermore, after significance analysis, the decision weights corresponding to the text features, image features, and audio features to be decided are obtained. The decision weights reflect the importance of the features to be decided in the decision-making process and can provide users with personalized interpretations.

[0105] For example, in medical diagnostic systems, saliency analysis reveals that tumor edge texture features in medical images have a high weight, while family medical history features in medical record texts have a relatively low weight, which can provide doctors with clear diagnostic basis and decision-making reference.

[0106] Furthermore, a multi-task learning framework can be introduced to simultaneously optimize interpretability constraints with multiple related tasks (such as feature selection and classification tasks). For example, in intelligent transportation systems, where traffic condition prediction and traffic incident detection tasks are performed concurrently, multi-task learning can ensure that the initial decision weight matrix meets interpretability requirements in both tasks and improves feature utilization efficiency.

[0107] In addition, saliency analysis can be aided by visualization, which can intuitively display the initial decision weight matrix in the form of heatmaps, etc., to help users quickly understand the importance distribution of features.

[0108] This embodiment achieves the accurate determination of the correlation between the text features, image features, and audio features to be decided through a combination of sparse attention mechanism and interpretability constraints. Based on this, reasonable decision weights are generated, enabling the agent to pay more attention to features closely related to the target task during the decision-making process, thereby improving the accuracy and reliability of the decision. At the same time, by applying interpretability constraints and performing significance analysis, the distribution of feature weights in the decision-making process has a clear interpretable basis, which improves the transparency of the agent's decision-making basis for users and enhances their trust in the agent's decision results.

[0109] In some embodiments, after determining the classification label of the target task based on the fusion features, the method further includes:

[0110] The system monitors the statistical characteristics of the data to be decided in real time. When the statistical characteristics change, it identifies historical data to be decided that have a data similarity exceeding a preset data similarity threshold with the updated data to be decided. It also receives the historical classification labels corresponding to the historical data to be decided. The statistical characteristics include at least one of variance, mean, and distribution.

[0111] By comparing historical category labels with the category labels corresponding to the current data to be decided, if the label similarity between the historical category labels and the current category labels is less than the preset label similarity, the visual decision path is adjusted, and the current category label is updated based on the adjusted visual decision path.

[0112] Optionally, after the agent completes the classification label determination for the target task, this embodiment enters the continuous monitoring stage, which monitors the statistical characteristics of the data to be decided in real time, and tracks and analyzes the statistical characteristics. The statistical characteristics include, but are not limited to, variance, mean and distribution. The statistical characteristics are used to reflect the stability and changing trend of the data to be decided.

[0113] For example, in financial risk assessment scenarios, real-time monitoring of market data fluctuations (variance), average trading volume (mean), and the overall distribution of data is used to determine whether the market environment has changed significantly.

[0114] Furthermore, statistical features can be updated using incremental statistical calculations, enabling efficient updates of statistical feature values ​​even as the data stream is continuously updated, without requiring recalculation of the entire dataset.

[0115] In addition, when a change in statistical characteristics is detected, historical decision-making data with a high degree of similarity to the currently updated decision-making data is identified from historical data records. High similarity means exceeding a preset data similarity threshold.

[0116] The similarity can be determined by various algorithms, such as Euclidean distance, cosine similarity, or Jaccard similarity coefficient. The preset data similarity threshold can be set according to the actual application scenario and experience, and is used to filter out historical data that is sufficiently similar to the current data to be decided.

[0117] In addition, the historical classification labels obtained from the historical data to be decided are compared with the classification labels corresponding to the current data to be decided to determine the degree of consistency between the two. This can be achieved by calculating label similarity. The method for determining label similarity can be selected based on the type and structure of the label. For example, for one-hot encoded classification labels, simple matching similarity can be used; for classification labels with hierarchical structures, similarity can be determined by path-based or set-based methods.

[0118] Furthermore, label similarity calculations can be enhanced through semantic analysis of labels, by analyzing their semantic meaning. For example, pre-trained language models (such as BERT) can be used to semantically encode textual labels, and then the cosine similarity between semantic vectors can be calculated to obtain more accurate label similarity. In addition, incorporating contextual information about the classification labels (such as task background and label generation time) can provide a comprehensive evaluation of label similarity, thereby improving the reliability of the comparison results.

[0119] Furthermore, when the similarity between historical and current classification labels is detected to be lower than a preset label similarity threshold, the visualization decision path is adjusted to optimize the agent's decision logic and enable it to better adapt to data changes. Adjustments may include, but are not limited to, adding or deleting nodes and edges in the graph structure, modifying edge weights, or introducing new processing steps.

[0120] Furthermore, by using reinforcement learning-based decision path adjustment methods, the performance of the adjusted decision path (e.g., classification accuracy, decision consistency) can be used as a reward signal, and the adjustment method can be optimized through trial and error learning. Meta-learning techniques can also be combined to quickly adapt to different types of decision path adjustment needs, improving the efficiency and effectiveness of the adjustment.

[0121] This embodiment enables real-time monitoring of the statistical characteristics of the data to be decided, allowing the agent to promptly perceive changes in the data and quickly adjust its decision path and classification labels. This ensures the agent maintains accuracy and consistency in its decisions when faced with changes in data distribution or concept drift. Furthermore, by comparing historical and current classification labels, potential decision biases can be identified and corrected by adjusting the visualized decision path, improving the reliability and consistency of the agent's decisions and reducing erroneous decisions caused by data changes or model limitations.

[0122] In some embodiments, after determining the classification label of the target task based on the fusion features, the method further includes:

[0123] Determine the confidence level of the classification label. If the confidence level is less than the preset confidence threshold, backtrack the step of determining the decision weight of the feature to be decided through the decision attention mechanism during the fusion process to obtain the updated decision weight.

[0124] Based on the updated decision weights, the model parameters of the cross-modal data fusion model are adjusted through gradient backpropagation, and the processing of the corresponding nodes in the visualized decision path is updated synchronously based on the updated model parameters.

[0125] Optionally, after determining the classification label of the target task based on the fusion features, this embodiment determines the confidence level of the classification label, where the confidence level reflects the degree of certainty about the classification result and is usually a value between 0 and 1.

[0126] For example, in an image recognition task, if the classification label of an image to be decided is "cat" and the confidence level is 0.95, it means that the current classification model has a high degree of confidence in this classification result.

[0127] The method for determining the confidence level can depend on the classification model used. For a softmax regression model, the confidence level can be the probability value of the classification result; for a support vector machine model, the confidence level can be determined by the distance from the classification boundary.

[0128] Furthermore, the confidence level of the classification label can be determined by combining the prediction results of multiple classification models through ensemble learning, or by combining prior knowledge and model prediction results through Bayesian methods to calculate the posterior probability of the classification label as the confidence level.

[0129] In addition, when the confidence level of the classification label is less than the preset confidence threshold, the process goes back to the feature fusion stage, re-examines the decision weights determined by the decision attention mechanism, identifies potential weight setting issues that may affect the reliability of the classification results, and makes corresponding adjustments.

[0130] The backtracking process includes reassessing the correlation between the text features, image features, and audio features determined by sparse attention in the decision attention mechanism. For example, it may be found that the correlation between the text features and image features is overestimated or underestimated, causing the fused features to fail to accurately reflect the actual decision requirements.

[0131] Based on the reassessed correlation, a new initial decision weight matrix is ​​generated, and significance analysis is performed again through interpretability constraints to obtain updated decision weights. For example, the weights for audio features can be increased, while the weights for image features can be decreased to better reflect the decision logic of the target task.

[0132] Furthermore, after obtaining the updated decision weights, the cross-modal data fusion model is further optimized using the updated decision weights. Among these optimizations, gradient backpropagation is a commonly used neural network training method. By calculating the gradient of the loss function with respect to the model parameters, the model parameters are updated through backpropagation to minimize the prediction error.

[0133] The gradient can be calculated and the model parameters updated based on the difference between the current classification label and the true label (e.g., calculated using the cross-entropy loss function) and the updated decision weights.

[0134] Furthermore, distributed gradient descent can be used to distribute computational tasks across multiple computing nodes, accelerating the model parameter update process. It can also be combined with second-order optimization methods (such as Newton's method or its approximations) to further improve the efficiency and convergence speed of model parameter updates based on gradient backpropagation.

[0135] In addition, after the model parameters of the cross-modal data fusion model are updated, the processing procedures of the corresponding nodes in the visualization decision path are updated synchronously. Each node in the visualization decision path represents a processing step, which includes, but is not limited to, feature extraction, feature fusion, and classification decision. The update of model parameters may affect the processing steps, such as the weight allocation in the feature fusion process and the boundary of the classification decision.

[0136] Based on the updated model parameters, adjust the display information of relevant nodes in the visualized decision path. For example, update the display of the decision weights for feature fusion nodes, or adjust the visualization of the decision boundary for classification decision nodes, so that the visualized decision path can accurately reflect the current state and decision logic of the cross-modal data fusion model.

[0137] Furthermore, if the confidence level is less than a preset confidence threshold, subgraphs in the graph structure can be replaced. For example, a weighted fusion subgraph can be replaced with an attention-based subgraph. The graph structure can also be layered, and updates can be performed step by step according to the layer to reduce computational overhead. For example, feature extraction nodes at the bottom layer of the graph structure can be updated first, and classification nodes at the top layer of the graph structure can be adjusted after the feature extraction nodes have been updated.

[0138] Furthermore, based on the confidence level of the current classification label, the weights of the edges in the graph structure can be adjusted in reverse. For example, if the confidence level of the updated classification label exceeds the preset true confidence level after adjusting the weights to be decided during feature fusion, the weights of the edges corresponding to the weights to be decided can be adjusted. The weights in the graph structure can also be reallocated when the task requirements of the target task change. For example, when the task requirements of the target task are updated from accuracy priority to interpretability priority, the weights of the edges corresponding to the paths that highlight interpretability can be reallocated. For example, the weights of text feature nodes can be increased.

[0139] This embodiment enables the correction of potential decision-making errors and the optimization of the feature fusion process and model parameters of the cross-modal data fusion model by determining the confidence level of the classification label and backtracking to update the decision weights. This improves the accuracy and confidence level of the classification label corresponding to the target task. It can also trigger the backtracking and update mechanism when the confidence level of the classification label is low, so that the agent can maintain good performance in the ever-changing decision-making data environment and adapt to new decision-making needs.

[0140] In some embodiments, a visual decision path is used to display the input data, output data, and processing method corresponding to each step of the process, including:

[0141] A corresponding natural language description is generated based on the output data of each node, and the natural language description is displayed on the display screen of the terminal device.

[0142] If the change in the processing weight corresponding to an edge exceeds a preset change threshold, the visual decision path corresponding to the edge is highlighted, and the reason for the change is displayed.

[0143] Optionally, in the process of constructing a visual decision path, this embodiment generates an intuitive natural language description for the output data corresponding to each node. The natural language description can explain in detail the output data of the node and its meaning.

[0144] For example, a feature extraction node can describe the basic statistical information (such as quantity, range, extraction method, etc.) of the extracted feature types and features to be decided. The generation of natural language descriptions can be based on predefined templates and rules, and the content can be dynamically filled in by combining the actual output data of the node.

[0145] In addition, the generated natural language description can be sent and displayed on the terminal device's screen. In web applications, it can be formatted using HTML and CSS, while in mobile applications, it can be presented using native UI components.

[0146] Furthermore, natural language generation (NLG) can be used to automatically generate richer and more flexible descriptive text based on the output data of nodes using deep learning models (such as the GPT architecture). At the same time, combined with speech synthesis technology, the natural language description can be converted into speech playback so that users can obtain information when it is inconvenient to look at the screen.

[0147] In addition, the system continuously monitors the changes in the processing weights of each edge in the visualized decision-making path. When the change in the processing weight of a certain edge in the graph structure exceeds a preset threshold, a highlighting mechanism is triggered. This highlighting can be achieved by changing the edge's color, increasing its thickness, or adding animation effects, making the corresponding portion of the visualized decision-making path prominently displayed in the user interface.

[0148] Simultaneously, the reasons for the changes in the edge processing weights can be analyzed and displayed to the user in a concise and clear manner. This analysis involves tracing back relevant data and processing steps, such as examining factors that might cause weight changes, including variations in the distribution of input data, updates to model parameters, or external feedback signals. For example, if the weight of an edge increases due to changes in data distribution, an explanation such as "Processing weight increased by 20% because the variance of the input data increased, strengthening the model's dependence on this path" can be displayed.

[0149] Furthermore, machine learning models (such as decision trees or Bayesian networks) can be used to analyze the multidimensional reasons for weight changes, providing a more accurate and comprehensive explanation.

[0150] This embodiment enables users to intuitively understand the decision-making process of an intelligent agent through natural language description and highlighting of changes, reducing confusion and doubt about the decision-making process, and quickly capturing changes in factors affecting the decision-making results, thereby improving the transparency of the decision-making process.

[0151] In some embodiments, after displaying the input data, output data, and processing method corresponding to each step of the process through a visual decision path, the method further includes:

[0152] Receive correction operations on classification labels and transform these correction operations into reward signals in the visualized decision path;

[0153] The update strategy of the visualization decision path is adjusted based on the near-end strategy optimization method, and the cross-modal data fusion model is updated based on the adjusted update strategy when the accumulated reward signal reaches the preset reward threshold.

[0154] Optionally, after visually displaying the classification results and decision-making process to the user, this embodiment can enable a user interaction mode, allowing users to modify the classification labels based on their own professional knowledge or practical experience.

[0155] When a user performs a correction operation, the correction operation is captured in real time and converted into a reward signal. That is, if the user's corrected classification label is inconsistent with the classification label originally output by the agent, it is regarded as feedback of the agent's decision-making bias and converted into a negative reward signal; otherwise, it is converted into a positive reward signal.

[0156] The conversion algorithm can be designed based on reinforcement learning theory, taking into account factors such as the frequency and magnitude of user corrections and historical correction records, to ensure the rationality and guidance of the reward signal.

[0157] For example, if a user frequently and significantly modifies the category label of a certain category, it can be determined that the agent has a large deviation in its decision-making regarding that category label, and a strong negative reward signal can be given to prompt the agent to focus on adjusting the relevant decision-making strategy.

[0158] Furthermore, a user reputation mechanism can be incorporated, assigning different weights to users based on their historical correction accuracy, professional qualifications, and other information, thereby adjusting the strength of their correction operations as reward signals.

[0159] For example, senior doctors' corrective actions carry higher weight than those of ordinary doctors, resulting in stronger reward signals that significantly influence the agent's decision-making strategies. Furthermore, natural language processing techniques can be combined to mine annotations or explanations added by users during corrections, further enriching the semantic meaning of the reward signals. This allows the agent to more accurately understand user intent and optimize decision-making logic.

[0160] In addition, the update strategy of the visualized decision path can be adjusted by the Proximal Policy Optimization (PPO) algorithm. The PPO algorithm ensures the stability of the policy update by truncating the update step size of the policy search, while taking into account a certain degree of exploration, and avoids frequent and large fluctuations in the agent's decision strategy due to the randomness of user correction operations.

[0161] In the initial stage of adjustment, the initial update strategy can be set based on default parameters, including key parameters such as learning rate and discount factor. As user correction operations are continuously transformed into reward signals and accumulated, the cumulative reward value is calculated in real time. When the cumulative reward value reaches the preset reward threshold, it is determined that the current user feedback is rich enough and statistically significant, triggering the model update process.

[0162] During the model update process, the reward signal can be backpropagated based on the adjusted update strategy to optimize the parameters of the cross-modal data fusion model. That is, for the feature extraction, feature fusion and classification decision processes in the cross-modal data fusion model, the gradients are calculated and targeted adjustments are made respectively.

[0163] For example, if a user frequently corrects classification errors dominated by a specific modality of data, they can focus on optimizing the feature extraction and fusion parameters of that modality of data, strengthen the ability of the cross-modal data fusion model to capture and express the features of that modality of data, and at the same time adjust the classification decision boundary to improve classification accuracy.

[0164] Furthermore, multi-agent reinforcement learning technology can be integrated, allowing multiple users to simultaneously correct the agent's decision-making results. The correction operations of multiple users are transformed into independent reward signals, and then the reward signals are integrated through an intelligent negotiation mechanism to generate a comprehensive update strategy. Meta-learning methods can also be combined to enable the agent to quickly adapt to the correction habits and preferences of different user groups, further improving the efficiency and effectiveness of cross-modal data fusion model updates.

[0165] Furthermore, the update strategy may include at least one of the following: adding nodes, adding edges, deleting redundant nodes, deleting redundant edges, and modifying node parameters. Adding nodes and adding edges may be done when new data to be decided appears or when the target task requires new processing steps. For example, in medical diagnosis, a new "pathology report text analysis" node may be added and connected to the existing "image feature extraction" node.

[0166] Deleting redundant nodes and edges can be done by evaluating the contribution of nodes and / or edges in the graph structure to the classification label of the target task, and deleting nodes and edges that have not been used for a long time or have a low contribution. For example, if the weight of the feature to be decided corresponding to the audio data approaches 0 in multiple decisions, then the relevant nodes and / or edges are deleted.

[0167] Modifying node parameters can be done by adjusting the processing algorithm parameters inside the nodes of the graph structure based on the new distribution of the data to be decided, such as updating the weights of convolutional kernels or the weight allocation of attention mechanisms.

[0168] This embodiment realizes the ability to accurately optimize cross-modal data fusion models by converting user correction operations into reward signals and adjusting update strategies according to the PPO algorithm. This provides users with a more intuitive and convenient way to participate in the intelligent agent decision-making optimization process, enhancing users' trust and reliance on the intelligent agent decision-making system.

[0169] To effectively address the shortcomings of traditional technologies in decision path tracking and interpretation, and to significantly improve the interpretability and transparency of obtaining the classification label of the target task from the decision-making data, this application provides an embodiment of an interpretable cross-media intelligent agent decision path tracking device for implementing all or part of the aforementioned interpretable cross-media intelligent agent decision path tracking method. See [link to embodiment]. Figure 2 The interpretable cross-media intelligent agent decision path tracking device specifically includes the following:

[0170] Extraction module 10 is used to receive decision data corresponding to the target task, extract data features of the decision data, and obtain decision features;

[0171] The classification module 20 is used to determine the decision weights of the features to be decided through a decision attention mechanism, process the features to be decided through a trained cross-modal data fusion model and decision weights to obtain fused features, and determine the classification label of the target task based on the fused features.

[0172] The processing module 30 is used to record the fusion process of obtaining fusion features and the classification process of obtaining classification labels step by step according to the processing order, and to define the fusion process and the classification process as the processing process;

[0173] The decision module 40 is used to synchronize the processing process step by step to the visual decision path according to the processing order, and to display the input data, output data and processing method corresponding to each step of the processing process through the visual decision path. The visual decision path is constructed by a graph structure. The edges of the graph structure represent the data flow direction, processing method and processing weight in the processing process. Each node of the graph structure corresponds to each step of the processing process, and the processing weight is obtained based on the weight to be decided and the target task.

[0174] As described above, the interpretable cross-media intelligent agent decision path tracking device provided in this application embodiment can innovatively receive decision-making data corresponding to a target task and extract data features from the decision-making data to obtain decision-making features. It then determines the decision-making weights of the decision-making features through a decision attention mechanism, processes the decision-making features using a trained cross-modal data fusion model and the decision-making weights to obtain fused features and determine the classification label corresponding to the target task. Simultaneously, it records the feature fusion process and the classification process in the processing order and identifies them as processing steps. Finally, it displays the input data, output data, and processing method corresponding to each step of the processing step in the processing order through a visual decision path. The visual decision path is... The graph structure is constructed to represent the data flow, processing method, and processing weight in the processing process. Each node in the graph structure corresponds to each processing step, and the processing weight is obtained based on the decision weight and the target task. By visualizing the decision path, the input data, output data, and processing method of each processing step can be dynamically displayed, which improves the interpretability of the decision process. The decision attention mechanism and the edge weight of the graph structure quantify the impact of each processing step on the decision, which enhances the transparency and adaptability of decision path tracking. It can effectively solve the shortcomings of traditional technologies in decision path tracking and interpretation, and significantly improve the interpretability and transparency of the process of obtaining the classification label of the target task based on the decision data of the target task.

[0175] From a hardware perspective, in order to effectively address the shortcomings of traditional technologies in decision path tracking and interpretation, and significantly improve the interpretability and transparency of obtaining the classification label of the target task based on the decision data of the target task, this application provides an embodiment of an electronic device for implementing all or part of the aforementioned interpretable cross-media intelligent agent decision path tracking method. The electronic device specifically includes the following components:

[0176] The system comprises a processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to implement information transmission between an interpretable cross-media intelligent agent decision path tracing device and core business systems, user terminals, and related databases and other related devices; the logic controller can be a desktop computer, tablet computer, or mobile terminal, etc., and this embodiment is not limited to these. In this embodiment, the logic controller can be implemented with reference to the embodiments of the interpretable cross-media intelligent agent decision path tracing method and the embodiment of the interpretable cross-media intelligent agent decision path tracing device, the contents of which are incorporated herein, and repeated details will not be described again.

[0177] It is understood that the user terminal may include smartphones, tablet computers, network set-top boxes, portable computers, desktop computers, personal digital assistants (PDAs), in-vehicle devices, smart wearable devices, etc. Among these, the smart wearable devices may include smart glasses, smartwatches, smart bracelets, etc.

[0178] In practical applications, parts of the interpretable cross-media intelligent agent decision path tracing method can be executed on the electronic device side as described above, or all operations can be completed in the client device. The choice can be made based on the processing power of the client device and the limitations of the user's usage scenario. This application does not impose any limitations on this. If all operations are completed in the client device, the client device may further include a processor.

[0179] The aforementioned client device may have a communication module (i.e., a communication unit) that can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side; in other implementation scenarios, it may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, a server cluster consisting of multiple servers, or a distributed server structure.

[0180] Figure 3 This is a schematic block diagram illustrating the system configuration of the electronic device 9600 according to an embodiment of this application. Figure 3 As shown, the electronic device 9600 may include a central processing unit 9100 and a memory 9140; the memory 9140 is coupled to the central processing unit 9100. It is worth noting that... Figure 3 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.

[0181] In one embodiment, the interpretable cross-media agent decision path tracing method functionality can be integrated into the central processing unit 9100. The central processing unit 9100 can be configured to perform the following control:

[0182] Step S101: Receive the decision data corresponding to the target task, extract the data features of the decision data, and obtain the decision features;

[0183] Step S102: Determine the decision weights of the features to be decided through the decision attention mechanism, process the features to be decided through the trained cross-modal data fusion model and the decision weights to obtain fused features, and determine the classification label of the target task based on the fused features;

[0184] Step S103: Record the fusion process of obtaining fusion features and the classification process of obtaining classification labels step by step according to the processing order, and define the fusion process and the classification process as the processing process;

[0185] Step S104: Synchronize the processing process step by step to the visual decision path according to the processing order, and display the input data, output data and processing method corresponding to each step of the processing process through the visual decision path. The visual decision path is constructed by a graph structure. The edges of the graph structure represent the data flow, processing method and processing weight in the processing process. Each node of the graph structure corresponds to each step of the processing process, and the processing weight is obtained based on the weight to be decided and the target task.

[0186] As described above, the electronic device provided in this application innovatively receives decision-making data corresponding to a target task and extracts data features from the decision-making data to obtain decision-making features. It then determines the decision-making weights of the decision-making features through a decision attention mechanism, processes the decision-making features using a trained cross-modal data fusion model and the decision-making weights to obtain fused features and determine the classification label corresponding to the target task. Simultaneously, it records the feature fusion process and the classification process in the processing order and identifies them as processing steps. A visual decision path displays the input data, output data, and processing method corresponding to each step of the processing process in the processing order. The visual decision path is constructed using a graph structure. The edges of the graph structure represent the data flow, processing method, and processing weight in the processing process. Each node in the graph structure corresponds to each processing step. The processing weight is obtained based on the decision weight and the target task. By visualizing the decision path, the input data, output data, and processing method of each processing step can be dynamically displayed, which improves the interpretability of the decision process. By quantifying the impact of each processing step on the decision through the decision attention mechanism and the edge weight of the graph structure, the transparency and adaptability of decision path tracking are enhanced. It can effectively solve the shortcomings of traditional technologies in decision path tracking and interpretation, and significantly improve the interpretability and transparency of the process of obtaining the classification label of the target task based on the decision data of the target task.

[0187] In another embodiment, an interpretable cross-media agent decision path tracing device can be configured separately from the central processing unit 9100. For example, an interpretable cross-media agent decision path tracing device can be configured as a chip connected to the central processing unit 9100, and the interpretable cross-media agent decision path tracing method function can be implemented through the control of the central processing unit.

[0188] like Figure 3As shown, the electronic device 9600 may further include: a communication module 9110, an input unit 9120, an audio processor 9130, a display 9160, and a power supply 9170. It is worth noting that the electronic device 9600 does not necessarily need to include these components. Figure 3 All components shown; in addition, the electronic device 9600 may also include Figure 3 For components not shown, please refer to existing technologies.

[0189] like Figure 3 As shown, the central processing unit 9100, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives inputs and controls the operation of various components of the electronic device 9600.

[0190] The memory 9140 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned failure-related information, and also store a program for executing that information. The central processing unit 9100 may execute the program stored in the memory 9140 to perform information storage or processing, etc.

[0191] Input unit 9120 provides input to central processing unit 9100. Input unit 9120 may be, for example, a keypad or touch input device. Power supply 9170 provides power to electronic device 9600. Display 9160 displays images and text. Display may be, for example, an LCD display, but is not limited thereto.

[0192] The memory 9140 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 9140 can also be some other type of device. The memory 9140 includes a buffer memory 9141 (sometimes referred to as a buffer). The memory 9140 may include an application / function storage unit 9142 for storing application programs and function programs or processes for executing the operation of the electronic device 9600 via the central processing unit 9100.

[0193] The memory 9140 may also include a data storage unit 9143 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 9144 of the memory 9140 may include various drivers for the electronic device for communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).

[0194] The communication module 9110 is a transmitter / receiver that sends and receives signals via the antenna 9111. The communication module 9110 (transmitter / receiver) is coupled to the central processing unit 9100 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.

[0195] Based on different communication technologies, multiple communication modules 9110 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module 9110 (transmitter / receiver) is also coupled to a speaker 9131 and a microphone 9132 via an audio processor 9130 to provide audio output via the speaker 9131 and receive audio input from the microphone 9132, thereby realizing typical telecommunications functions. The audio processor 9130 may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processor 9130 is coupled to a central processing unit 9100, enabling on-device recording via the microphone 9132 and on-device playback of stored audio via the speaker 9131.

[0196] Embodiments of this application also provide a computer-readable storage medium capable of implementing all steps of the interpretable cross-media intelligent agent decision path tracing method with a server or client as the execution subject in the above embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the interpretable cross-media intelligent agent decision path tracing method with a server or client as the execution subject in the above embodiments. For example, when the processor executes the computer program, it implements the following steps:

[0197] Step S101: Receive the decision data corresponding to the target task, extract the data features of the decision data, and obtain the decision features;

[0198] Step S102: Determine the decision weights of the features to be decided through the decision attention mechanism, process the features to be decided through the trained cross-modal data fusion model and the decision weights to obtain fused features, and determine the classification label of the target task based on the fused features;

[0199] Step S103: Record the fusion process of obtaining fusion features and the classification process of obtaining classification labels step by step according to the processing order, and define the fusion process and the classification process as the processing process;

[0200] Step S104: Synchronize the processing process step by step to the visual decision path according to the processing order, and display the input data, output data and processing method corresponding to each step of the processing process through the visual decision path. The visual decision path is constructed by a graph structure. The edges of the graph structure represent the data flow, processing method and processing weight in the processing process. Each node of the graph structure corresponds to each step of the processing process, and the processing weight is obtained based on the weight to be decided and the target task.

[0201] As described above, the computer-readable storage medium provided in this application innovatively receives decision-making data corresponding to a target task and extracts data features from the decision-making data to obtain decision-making features. It then determines the decision-making weights of the decision-making features through a decision attention mechanism, processes the decision-making features using a trained cross-modal data fusion model and the decision-making weights to obtain fused features and determine the classification label corresponding to the target task. Simultaneously, it records the feature fusion process and the classification process in the processing order and identifies them as processing steps. A visual decision path displays the input data, output data, and processing method corresponding to each step of the processing process in the processing order. The visual decision path is constructed using a graph structure. The graph structure is constructed so that the edges represent the data flow, processing method, and processing weight in the processing process. Each node in the graph structure corresponds to each processing step, and the processing weight is obtained based on the decision weight and the target task. By visualizing the decision path, the input data, output data, and processing method of each processing step can be dynamically displayed, which improves the interpretability of the decision process. By quantifying the impact of each processing step on the decision through the decision attention mechanism and the edge weight of the graph structure, the transparency and adaptability of decision path tracking are enhanced. It can effectively solve the shortcomings of traditional technologies in decision path tracking and interpretation, and significantly improve the interpretability and transparency of the process of obtaining the classification label of the target task based on the decision data of the target task.

[0202] Embodiments of this application also provide a computer program product capable of implementing all steps of the interpretable cross-media intelligent agent decision path tracing method with the execution subject being a server or client in the above embodiments. When executed by a processor, the computer program / instructions implement the steps of the interpretable cross-media intelligent agent decision path tracing method. For example, the computer program / instructions implement the following steps:

[0203] Step S101: Receive the decision data corresponding to the target task, extract the data features of the decision data, and obtain the decision features;

[0204] Step S102: Determine the decision weights of the features to be decided through the decision attention mechanism, process the features to be decided through the trained cross-modal data fusion model and the decision weights to obtain fused features, and determine the classification label of the target task based on the fused features;

[0205] Step S103: Record the fusion process of obtaining fusion features and the classification process of obtaining classification labels step by step according to the processing order, and define the fusion process and the classification process as the processing process;

[0206] Step S104: Synchronize the processing process step by step to the visual decision path according to the processing order, and display the input data, output data and processing method corresponding to each step of the processing process through the visual decision path. The visual decision path is constructed by a graph structure. The edges of the graph structure represent the data flow, processing method and processing weight in the processing process. Each node of the graph structure corresponds to each step of the processing process, and the processing weight is obtained based on the weight to be decided and the target task.

[0207] As described above, the computer program product provided in this application innovatively receives decision-making data corresponding to a target task and extracts data features from the decision-making data to obtain decision-making features. It then determines the decision-making weights of the decision-making features through a decision attention mechanism, processes the decision-making features using a trained cross-modal data fusion model and the decision-making weights to obtain fused features and determine the classification label corresponding to the target task. Simultaneously, it records the feature fusion process and the classification process in the processing order and identifies them as processing steps. A visual decision path displays the input data, output data, and processing method corresponding to each step of the processing process in the processing order. The visual decision path is constructed using a graph structure. The edges of the graph structure represent the data flow, processing method, and processing weight in the processing process. Each node in the graph structure corresponds to each processing step, and the processing weight is obtained based on the decision weight and the target task. By visualizing the decision path, the input data, output data, and processing method of each processing step can be dynamically displayed, improving the interpretability of the decision process. By quantifying the impact of each processing step on the decision through the decision attention mechanism and the edge weight of the graph structure, the transparency and adaptability of decision path tracking are enhanced. This effectively solves the shortcomings of traditional technologies in decision path tracking and interpretation, and significantly improves the interpretability and transparency of the process of obtaining the classification label of the target task based on the decision data of the target task.

[0208] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0209] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0210] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0211] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0212] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. An explainable cross-media agent decision path tracking method, characterized in that, The method comprises: receiving to-be-decided data corresponding to a target task, extracting data features of the to-be-decided data to obtain to-be-decided features, the to-be-decided features comprising to-be-decided text features, to-be-decided image features and to-be-decided audio features; determining the correlation between the to-be-decided text features, the to-be-decided image features and the to-be-decided audio features through sparse attention in the decision attention mechanism, generating an initial to-be-decided weight matrix based on the correlation, performing significance analysis on the initial to-be-decided weight matrix through explainability constraints to obtain to-be-decided weights corresponding to the to-be-decided text features, the to-be-decided image features and the to-be-decided audio features respectively, processing the to-be-decided features through a trained cross-modal data fusion model and the to-be-decided weights to obtain fusion features, and determining a classification label of the target task based on the fusion features; recording the fusion process of obtaining the fusion features and the classification process of obtaining the classification label step by step according to a processing order, and determining the fusion process and the classification process as a processing process; synchronizing the processing process to a visual decision path step by step according to the processing order, and displaying input data, output data and processing methods corresponding to each step of the processing process through the visual decision path, wherein the visual decision path is obtained through a graph structure, edges of the graph structure represent data flow, processing methods and processing weights in the processing process, each node of the graph structure corresponds to each step of the processing process, and the processing weights are based on to-be-decided weights and the target task.

2. The method of claim 1, wherein, The to-be-decided data comprises text data, image data and audio data; The extraction of the data features of the to-be-decided data to obtain to-be-decided features comprises: mapping the text data into a low-dimensional semantic vector through a pre-trained word vector model to obtain to-be-decided features corresponding to the text data; extracting local features and global features of the image data through a convolutional neural network, and determining the local features and the global features as to-be-decided features corresponding to the image data; processing the audio data through a short-time Fourier transform to generate a frequency spectrum graph corresponding to the audio data, and extracting features of the frequency spectrum graph through a recurrent neural network to obtain to-be-decided features corresponding to the audio data.

3. The method of claim 1, wherein, After determining the classification label of the target task based on the fusion features, the method further comprises: monitoring statistical features of the to-be-decided data in real time, determining historical to-be-decided data with a data similarity to updated to-be-decided data exceeding a preset data similarity threshold in the case that the statistical features change, and receiving a historical classification label corresponding to the historical to-be-decided data, wherein the statistical features comprise at least one of variance, mean value and distribution; comparing the historical classification label with the classification label corresponding to the current to-be-decided data, and adjusting the visual decision path in the case that a label similarity between the historical classification label and the current classification label is less than a preset label similarity, and updating the current classification label based on the adjusted visual decision path.

4. The method of claim 1, wherein, After the classification label of the target task is determined based on the fusion feature, the method further includes: determining a confidence of the classification label, and in a case where the confidence is less than a preset confidence threshold, backtracking a step of determining a to-be-decided weight of the to-be-decided feature by a decision attention mechanism in the fusion process to obtain an updated to-be-decided weight; based on the updated to-be-decided weight, adjusting model parameters of the cross-modal data fusion model by a gradient back propagation manner, and synchronously updating the processing process of the corresponding node in the visual decision path based on the updated model parameters.

5. The method of claim 1, wherein, The displaying of the input data, the output data and the processing manner corresponding to each step of the processing process through the visual decision path includes: generating a corresponding natural language description based on the output data corresponding to each node, and displaying the natural language description on a display screen of a terminal device; in a case where a change amount of the processing weight corresponding to the edge exceeds a preset change threshold, highlighting the visual decision path corresponding to the edge, and displaying a change reason of the change amount.

6. The method of claim 1, wherein, After the displaying of the input data, the output data and the processing manner corresponding to each step of the processing process through the visual decision path, the method further includes: receiving a correction operation on the classification label, converting the correction operation into a reward signal in the visual decision path; adjusting an update strategy of the visual decision path based on a proximal policy optimization manner, and in a case where the accumulated reward signal reaches a preset reward threshold, updating the cross-modal data fusion model based on the adjusted update strategy.

7. An explainable cross-media agent decision path tracing apparatus, characterized by, The device includes: an extraction module configured to receive to-be-decided data corresponding to a target task, extract data features of the to-be-decided data to obtain to-be-decided features, the to-be-decided features including to-be-decided text features, to-be-decided image features and to-be-decided audio features; a classification module configured to determine an association degree between the to-be-decided text features, the to-be-decided image features and the to-be-decided audio features through sparse attention in the decision attention mechanism, generate an initial to-be-decided weight matrix based on the association degree, perform saliency analysis on the initial to-be-decided weight matrix through an explainability constraint to obtain to-be-decided weights corresponding to the to-be-decided text features, the to-be-decided image features and the to-be-decided audio features respectively, process the to-be-decided features through a trained cross-modal data fusion model and the to-be-decided weights to obtain fusion features, and determine a classification label of the target task based on the fusion features; a processing module configured to record a fusion process of obtaining the fusion features and a classification process of obtaining the classification label step by step according to a processing order, and determine the fusion process and the classification process as a processing process. A decision module is configured to synchronize the processing procedure into a visual decision path according to the processing sequence, and display the input data, output data and processing mode corresponding to each step of the processing procedure through the visual decision path, wherein the visual decision path is constructed by a graph structure, edges of the graph structure represent data flow direction, processing mode and processing weight in the processing procedure, each node of the graph structure corresponds to each step of the processing procedure, and the processing weight is based on a to-be-decided weight and the target task.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the interpretable cross-media intelligent agent decision path tracking method of any one of claims 1 to 6 when executing the program.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the steps of the interpretable cross-media intelligent agent decision path tracking method of any one of claims 1 to 6 when executed by the processor.

Citation Information

Patent Citations

  • Structural recognition model training method and device, model structure recognition method and device and medium

    CN117609870A

  • Object classification method, classification model training method, related equipment and program product

    CN120257024A