User feedback classification method and electronic device

By utilizing metadata from user feedback in enterprise after-sales and customer service scenarios to determine the target environment domain, suppressing causal relationships in multimodal features, and combining multiple agents to process single-modal features, this approach solves the problem of poor robustness of single-modal deep learning models in new environments, achieving highly accurate and efficient user feedback classification.

CN121093097BActive Publication Date: 2026-01-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511553422.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-01-27
Estimated Expiration
2045-10-28

AI Technical Summary

Technical Problem

In enterprise after-sales and customer service scenarios, existing technologies show that single-modal deep learning models have poor classification accuracy for user feedback content in new and complex environments, are difficult to process multimodal information, and cannot fully extract effective information from user feedback.

Method used

The target environment domain is determined by using metadata based on user feedback content. Environmental features are used to suppress causal relationships in multimodal features. Multiple agents are combined to process single-modal features separately, thereby achieving the decoding and prediction of causal features.

Benefits of technology

It improves the accuracy and robustness of user feedback classification, maintains high classification accuracy in new environments, and enhances classification efficiency through the collaborative work of multiple agents, while also tracking the source of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121093097B_ABST
    Figure CN121093097B_ABST
Patent Text Reader

Abstract

The application provides a user feedback classification method and an electronic device, and relates to the technical field of data processing. The method comprises the following steps: in response to receiving user feedback content, determining a target environment domain matched with the user feedback content from a plurality of candidate environment domains based on metadata of the user feedback content; performing suppression processing on information related to the target environment domain in multi-modal features of the user feedback content based on environment characteristics of the target environment domain to obtain causal features, the suppression processing being used for weakening the causal relationship between the information and a to-be-predicted category, and the causal features representing features in the multi-modal features that have a causal relationship with the to-be-predicted category; inputting a plurality of single-modal features obtained by decoding the causal features into a plurality of agents respectively to obtain a plurality of prediction features, the plurality of single-modal features being matched with a plurality of modal types of the multi-modal features respectively; and obtaining a target prediction category of the user feedback content based on the plurality of prediction features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more specifically, to a user feedback classification method and an electronic device. Background Technology

[0002] In enterprise after-sales and customer service scenarios, the classification of user feedback mainly relies on a single-modal deep learning model. It can achieve high accuracy when processing data similar to the training samples. However, when faced with user feedback in new environments and complex scenarios, the classification accuracy drops rapidly and the robustness is poor.

[0003] Furthermore, user feedback often contains multimodal information, such as questions described in natural language or images of questions uploaded by users. The aforementioned classification methods are unable to handle multimodal input and cannot fully extract the effective information from user feedback. Summary of the Invention

[0004] In view of the above problems, this application provides a user feedback classification method and electronic device to improve the accuracy and robustness of user feedback classification.

[0005] According to a first aspect of this application, a user feedback classification method is provided, the method comprising: in response to receiving user feedback content, determining a target environment domain matching the user feedback content from multiple candidate environment domains based on the metadata of the user feedback content, wherein each of the multiple candidate environment domains has environment features corresponding to the metadata; based on the environment features of the target environment domain, performing suppression processing on information related to the target environment domain in the multimodal features of the user feedback content, obtaining causal features, wherein the suppression processing is used to weaken the causal relationship between the information and the category to be predicted, and the causal features characterize the features in the multimodal features that have a causal relationship with the category to be predicted; inputting multiple single-modal features obtained by decoding the causal features into multiple agents respectively, obtaining multiple prediction features, wherein the multiple single-modal features are respectively matched with multiple modality types of the multimodal features, and the prediction features indicate the candidate prediction categories for the user feedback content; and obtaining the target prediction category of the user feedback content based on the multiple prediction features.

[0006] A second aspect of this application provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0007] According to the user feedback classification method and electronic device provided in this application, the environmental information contained in the metadata provided by the user is used to determine the environmental features corresponding to the user feedback content, providing accurate guidance for subsequent suppression processing; the multimodal features of the user feedback content are used to comprehensively understand the after-sales issues and other information in the user feedback; at the same time, the environmental features are used to suppress the information related to the target environment domain in the multimodal features, reducing the impact of environmental changes on causal feature extraction, so that it can still maintain a high accuracy in new environments and improve the robustness of user feedback classification; and multiple agents are used to process the multiple single-modal features obtained by decoding, which improves classification efficiency and facilitates the tracking of error sources. Attached Figure Description

[0008] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:

[0009] Figure 1 The illustrations depict application scenarios of user feedback classification methods, apparatus, devices, storage media, and program products according to embodiments of this application.

[0010] Figure 2 A flowchart illustrating a user feedback classification method according to an embodiment of this application is shown schematically.

[0011] Figure 3 An exemplary flowchart illustrates a process according to an embodiment of this application, in which multiple single-modal features obtained by decoding causal features are input into multiple agents to obtain multiple predicted features.

[0012] Figure 4 An exemplary flowchart illustrates a process for determining multiple weights corresponding to multiple modal types based on multiple correlations between a target environment domain and multiple modal types, according to an embodiment of this application.

[0013] Figure 5 A schematic diagram illustrating the structure of a user feedback classification device according to an embodiment of this application is shown; and

[0014] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing a user feedback classification method according to an embodiment of this application. Detailed Implementation

[0015] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.

[0016] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0017] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0018] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0019] The embodiments of this application provide a user feedback classification method that combines causal reasoning with multi-agent collaborative work to extract cross-environment invariant causal features from multimodal feedback provided by users, thereby improving the accuracy and robustness of user feedback classification.

[0020] Figure 1 The diagram illustrates an application scenario of the user feedback classification method according to an embodiment of this application.

[0021] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0022] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0023] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0024] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0025] It should be noted that the user feedback classification method provided in this application embodiment can generally be executed by server 105. Correspondingly, the user feedback classification device provided in this application embodiment can generally be located in server 105. The user feedback classification method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the user feedback classification device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0026] It should be understood that Figure 1 The number of terminal devices, networks 104, and servers 105 shown is merely illustrative. Any number of terminal devices, networks, and servers can be included depending on implementation needs.

[0027] In the technical solution of this application, all user information involved is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of related data all comply with relevant laws, regulations and standards, and necessary confidentiality measures have been taken, and do not violate public order and good morals.

[0028] The following will be based on Figure 1 The described scene, through Figures 2-4 The user feedback classification method of this application embodiment will be described in detail.

[0029] It should be noted that the sequence numbers of the operations in the following methods are for descriptive purposes only and should not be considered as indicating the execution order of the operations. Unless explicitly stated otherwise, the method does not need to be executed in the exact order shown.

[0030] Figure 2 A flowchart illustrating a user feedback classification method according to an embodiment of this application is shown schematically. Figure 3 An exemplary flowchart illustrates a process according to an embodiment of this application, in which multiple single-modal features obtained by decoding causal features are input into multiple agents to obtain multiple predicted features. Figure 4 An exemplary flowchart is shown, illustrating a process for determining multiple weights 480 corresponding to multiple modal types based on multiple correlations between a target environment domain and multiple modal types, according to an embodiment of this application.

[0031] Combination Figures 2-4 The user feedback classification method in this embodiment includes operations S210 to S240.

[0032] In operation S210, in response to receiving user feedback content, a target environment domain that matches the user feedback content is determined from multiple candidate environment domains based on the metadata of the user feedback content.

[0033] According to embodiments of this application, metadata refers to environmental conditions input by the user to guide user feedback classification, such as at least one of region, version, language, and model; multiple candidate environment domains each have environmental features corresponding to the metadata, and multiple candidate environment domains correspond to multiple environmental features. A candidate environment domain refers to the environmental distribution corresponding to environmental conditions in a real scene, and the environmental features can be environmental embedding vectors corresponding to the candidate environment domains, used to learn semantic information in environmental conditions and guide causal reasoning.

[0034] For example, based on metadata and preset mapping relationships, multiple candidate environment domains can be searched to obtain the target environment domain that matches the user feedback content. The mapping relationship indicates the correspondence between multiple metadata and multiple candidate environment domains.

[0035] Based on this, the metadata of user input can be used directly to reflect the environment in which the user feedback content is located, without relying on surface features such as image background, scene vocabulary, and feedback tone to learn environmental conditions, thereby reducing the bias caused by relying on the above surface features in user feedback classification.

[0036] In operation S220, based on the environmental features 420 of the target environment domain, information related to the target environment domain in the multimodal features of user feedback content is suppressed to obtain causal features.

[0037] According to embodiments of this application, multimodal features include multiple submodal features 410 with different modal types, used to comprehensively reflect after-sales issues reported by users through features of different modalities; suppression processing is used to weaken the causal relationship between information and the category to be predicted, reducing the interference of information related to the target environment domain in the multimodal features on the classification of user feedback; causal features characterize features in the multimodal features that have a causal relationship with the category to be predicted, and subsequent classification is performed based on causal features in subsequent operations to filter out false interference related to the environment.

[0038] Specifically, the causal relationship between information related to the target environment domain and the category to be predicted refers to the situation where information related to the target environment domain in multimodal features is the cause and the category to be predicted is the effect. The category to be predicted is obtained based on the information related to the target environment domain in multimodal features. This relationship changes with environmental changes, is not robust, and is not suitable for handling new environments. In contrast, causal features represent features in multimodal features that have a causal relationship with the category to be predicted and are not affected by environmental changes.

[0039] For example, environmental features and multiple submodal features 410 can be mapped to the same target dimension. Using environmental features as query vectors, and multiple submodal features 410 as multiple key vectors and multiple value vectors respectively, attention mechanism 450 is applied to obtain multiple output weights of multiple submodal features 410. Multiple output weights are used to perform weighted fusion of multiple value vectors to obtain causal features. This reduces the output weights of multimodal features that are related to the target environment domain, and increases the output weights of features in multimodal features that have a causal relationship with the category to be predicted.

[0040] In operation S230, multiple single-modal features obtained from decoding causal features are input into multiple agents to obtain multiple prediction features.

[0041] According to embodiments of this application, multiple unimodal features are respectively matched with multiple modal types of multimodal features, multiple agents correspond to multiple modal types of multimodal features, and prediction features indicate candidate prediction categories for user feedback content.

[0042] For example, if the modality type of the single-modal feature is image, the corresponding image agent is invoked to process the single-modal feature and obtain the corresponding predicted feature. The image agent can be a multi-layer convolutional network or a visual transformer. If the modality type of the single-modal feature is text, the corresponding text agent is invoked to process the single-modal feature. The text agent can be a Bidirectional Encoder Representations from Transformers (BERT) model. If the modality type of the single-modal feature is speech, the corresponding speech agent is invoked to process the single-modal feature. The speech agent can be a speech recognition model. If the modality type of the single-modal feature is video, the corresponding video agent is invoked to process the single-modal feature. The video agent can be a 3D convolutional model or a sequence model, etc.

[0043] In operation S240, the target predicted category of the user feedback content is obtained based on multiple prediction features.

[0044] For example, the prediction features include candidate prediction categories and corresponding classification probabilities for user feedback content. The average classification probabilities of the same candidate prediction category among multiple prediction features are calculated to obtain the average classification probability corresponding to each candidate prediction category and sorted in descending order. The candidate prediction category with the highest average classification probability is taken as the target prediction category.

[0045] Through operations S210 to S240, the multimodal features of user feedback content are used to comprehensively understand information such as after-sales issues reported by users. The environmental information contained in the metadata provided by users is used to determine the target environmental domain and environmental features corresponding to the user feedback content. This facilitates the determination of information related to the target environmental domain in the multimodal features during subsequent suppression processing, reducing the impact of environmental changes on causal feature extraction and enabling it to maintain high accuracy in new environments, thereby improving the robustness of user feedback classification. Furthermore, multiple agents process the decoded single-modal features separately to improve classification efficiency. At the same time, in the event of classification anomalies, the source of error can be traced based on the deviation between multiple predicted features.

[0046] Multimodal features include multiple submodal features 410 with different modal types. Each submodal feature includes a single modal feature corresponding to its modality type and relevant environmental domain information. In some embodiments, to effectively utilize the user's multimodal input and comprehensively reflect after-sales issues using data from different modalities, the method further includes:

[0047] First, based on the multiple modal information corresponding to multiple modal types in the user feedback content, the corresponding multiple feature extraction models are called.

[0048] According to embodiments of this application, the modal type can be text, voice, image, audio, or video.

[0049] For example, the multiple modal types include text, speech, image, audio, and video, and the multiple modal information includes text information, speech information, image information, audio information, and video information.

[0050] Then, multiple feature extraction models are used to encode the features of multiple modal information to obtain multiple submodal features 410.

[0051] According to an embodiment of this application, multiple feature extraction models process modal information of the same modal type based on their matched modal types to obtain multiple submodal features 410.

[0052] For example, a pre-trained residual network (ResNet) is used to encode image information features to obtain image submodal features. The text information is feature-encoded using a pre-trained BERT model to obtain text submodal features. The speech information is converted into text using a speech recognition model (Whisper), and then the speech submodal features are extracted using the BERT model. The video information is sampled frame by frame per second, and features are extracted using a ResNet model. These features are then aggregated into video submodal features using temporal average pooling. , Represents the space of real numbers. Indicates the feature dimension.

[0053] To match the target environment domain in real time based on user input and improve the accuracy of target environment domain matching, thus providing accurate environmental conditions for subsequent causal inference; in some embodiments, the target environment domain is determined based on the similarity between metadata and multiple candidate environment domains. Operation S210 includes:

[0054] First, based on the metadata of the user feedback content, a matching vector is obtained.

[0055] According to embodiments of this application, the metadata of user feedback content includes at least one of the following: device model, language, region, and sales channel.

[0056] For example, the metadata of the user feedback content includes the region and device model, where the region is a certain city and the device model is v1. The metadata is one-hot encoded to obtain the matching vector [1, 0, 0, 1].

[0057] Then, multiple similarities are calculated between the vector to be matched and multiple candidate environment domains to determine the target environment domain based on multiple similarities.

[0058] According to an embodiment of this application, multiple similarities between the vector to be matched and multiple candidate environment domains are calculated and sorted in descending order, and the candidate environment domain with the highest similarity to the vector to be matched is taken as the target environment domain.

[0059] For example, multiple candidate environment domains are multiple clusters obtained by clustering multiple metadata samples. These metadata samples include device model, language, region, and sales channel, etc. First, based on the mean of the same type of metadata samples in each cluster, the cluster center vector of each cluster is calculated. Then, multiple Euclidean distances between the vector to be matched and the multiple cluster center vectors are calculated and sorted in ascending order; the smaller the Euclidean distance, the higher the similarity. The candidate environment domain with the smallest Euclidean distance is selected as the target environment domain.

[0060] Each candidate environment domain has relatively consistent distribution characteristics, and there may be common changing information between different candidate environment domains. In order to clarify all environmental features involved in user feedback classification, ensure that the feature extractor learns sufficiently during training, clearly distinguish between environmental features and causal features, and provide conditions for causal invariance learning; before executing the user feedback classification method of this application embodiment, in some embodiments, multiple environmental features are obtained using the following method:

[0061] First, based on clustering processing, multiple metadata samples of multiple user feedback content samples are divided into multiple clusters to serve as multiple candidate environment domains.

[0062] According to embodiments of this application, multiple metadata samples include device model, language, region, and sales channel, to comprehensively cover the environmental conditions involved in user feedback classification.

[0063] For example, firstly, multiple metadata samples are one-hot encoded to obtain multiple metadata vectors, which are then assembled into a vector matrix; then, a clustering algorithm is used to cluster the vector matrix to obtain multiple clusters, with each cluster representing a candidate environment domain; the clustering algorithm can be the K-means algorithm.

[0064] For example, the environment domain label of each candidate environment domain is uniquely encoded as an integer for grouping, candidate environment domains Represented as:

[0065] ;

[0066] In the formula, This represents the i-th candidate environment domain label.

[0067] Then, multiple multimodal sample features of multiple user feedback content samples are obtained, as well as multiple initial environmental features corresponding to multiple candidate environmental domains.

[0068] According to an embodiment of this application, multiple user feedback content samples correspond to multiple candidate environment domains, and each candidate environment domain has an initial environment feature that matches it. The initial environment feature is generated through random initialization.

[0069] For example, for each candidate environment domain, a corresponding embedding vector matrix is ​​randomly generated, and the embedding vector matrix is ​​input into the embedding layer for processing to map it as initial environment features, which are learnable embedding vectors.

[0070] Then, based on multiple initial environmental features, the information related to the corresponding candidate environmental domain in multiple multimodal sample features is suppressed to obtain multiple causal sample features.

[0071] According to an embodiment of this application, the corresponding candidate environment domain is determined based on the metadata sample in the user feedback content sample, ensuring that the corresponding candidate environment domain matches the user feedback content sample and ensuring the effectiveness of training.

[0072] For example, multiple multimodal sample features and corresponding initial environmental features are concatenated and input into a feature extractor for causal inference to obtain multiple causal sample features.

[0073] Then, multiple classifiers are used to process the features of multiple causal samples respectively to obtain multiple classification prediction results.

[0074] According to embodiments of this application, multiple classifiers correspond to multiple candidate environment domains. Data from different candidate environment domains are classified and predicted using a classifier that matches the candidate environment domain, so as to provide classification prediction results corresponding to different candidate environment domains.

[0075] Then, based on the multiple deviations between multiple classification prediction results and multiple classification labels of multiple user feedback content samples, the invariant loss during suppression is determined.

[0076] According to embodiments of this application, the invariant loss indicates the prediction accuracy of each candidate environment domain and the difference in prediction accuracy among multiple candidate environment domains. The invariant loss function designed based on the invariant learning idea includes a loss term for minimizing the prediction bias in each candidate environment domain and a regularization term for minimizing the loss bias between different candidate environment domains.

[0077] For example, the invariant loss function Represented as:

[0078]

[0079] In the formula, This indicates the deviation between the classification prediction result and the classification label. Indicates feature extractor, Indicates candidate environment domain The corresponding classifier, Indicates candidate environment domain Multimodal feature representation under the following conditions Indicates candidate environment domain Category tags corresponding to user feedback content samples, This represents a hyperparameter used to balance prediction accuracy and the strength of invariance constraints. This represents the gradient of the loss calculated for each candidate environment domain corresponding to the classifier at a scalar value of 1.0.

[0080] Then, multiple initial environmental features are adjusted to reduce invariant losses, resulting in multiple environmental features.

[0081] According to embodiments of this application, the embedding layer and feature extractor are trained with the goal of minimizing invariant loss. During iterative training, the parameters of the feature extractor and embedding layer are updated. The initial environmental features are input into the updated embedding layer for adjustment, and the adjusted initial environmental features are used for the next training iteration to reduce the invariant loss, until a preset termination condition is reached, resulting in multiple corresponding environmental features. , Representing dimensions, such as =16-dimensional, 32-dimensional, etc.

[0082] During the training of the embedding layer and feature extractor, environmental features can fully learn the semantic information corresponding to the candidate environmental domains they represent, providing environmental information for the subsequent dynamic scheduling of the agent.

[0083] To uncover invariant features shared across environmental domains in multimodal features and eliminate spurious correlations related to environmental variables, a trained feature extractor is used to process the multimodal features and environmental features of the target environmental domain 420; according to an embodiment of this application, operation S220 includes:

[0084] First, based on the environmental characteristics 420 of the target environment domain, suppression parameters that match the target environment domain are determined from multiple sets of suppression parameters.

[0085] According to embodiments of this application, the suppression parameter indicates the degree of filtering of environmental noise relevant to the target environment domain.

[0086] For example, the suppression parameter can be a model parameter related to the target environment domain retrieved from the trained model parameters based on the environmental features 420 of the target environment domain by a trained feature extractor.

[0087] Then, based on the suppression parameter, the multimodal features are processed to obtain causal features.

[0088] According to embodiments of this application, a pre-trained feature extractor is used to process multimodal features based on suppression parameters.

[0089] For example, firstly, the image submodal features are... Text submodal features Speech submodal features Video submodal features Environmental characteristics corresponding to the target environment domain By concatenating the features, a multimodal feature representation is obtained. Then, the multimodal features are represented. The input is processed by a pre-trained feature extractor to obtain causal features; here, the feature extractor can be a convolutional neural network or a transformer model.

[0090] To suppress weakly correlated features that are only effective in specific environments during multimodal feature extraction, and to make the model focus on features with a direct causal relationship to the category to be predicted, a feature extractor is trained based on an invariant loss function to improve the generalization ability of the feature extractor to data in new environments. In some embodiments, multiple sets of suppression parameters are obtained using the following methods:

[0091] First, multiple metadata samples based on multiple user feedback content samples are clustered to obtain multiple candidate environment domains.

[0092] According to embodiments of this application, multiple metadata samples include device model, version model, language, region, and sales channel, etc.

[0093] For example, firstly, multiple metadata samples are one-hot encoded to obtain multiple metadata vectors, which are then assembled into a vector matrix; then, the K-means algorithm is used to cluster the vector matrix to obtain multiple clusters, with each cluster representing a candidate environment domain.

[0094] Then, based on multiple environmental features corresponding to multiple candidate environmental domains, the information related to the corresponding candidate environmental domains in the multiple multimodal sample features of multiple user feedback content samples is suppressed to obtain multiple causal sample features.

[0095] According to an embodiment of this application, the corresponding candidate environment domain is determined based on the metadata sample in the user feedback content sample.

[0096] For example, multiple multimodal sample features and corresponding environmental features are concatenated and then input into a feature extractor for processing to obtain multiple causal sample features. The feature extractor can use a convolutional neural network or a transformer model.

[0097] Then, multiple classifiers are used to process the features of multiple causal samples respectively to obtain multiple classification prediction results.

[0098] According to embodiments of this application, multiple classifiers correspond to multiple candidate environment domains. These multiple classifiers are used only during the training phase to assist in training and calculate invariant loss. Fully connected layers can be used for the classifiers.

[0099] Then, based on the multiple deviations between multiple classification prediction results and multiple classification labels of multiple user feedback content samples, the invariant loss during suppression is determined.

[0100] According to embodiments of this application, the invariant loss indicates the prediction accuracy of each candidate environment domain and the difference in prediction accuracy among multiple candidate environment domains.

[0101] For example, the invariant loss function Represented as:

[0102]

[0103] In the formula, This represents the loss function, which is the deviation between the classification prediction result and the classification label. Indicates feature extractor, Indicates candidate environment domain The corresponding classifier, Indicates candidate environment domain Multimodal feature representation under the following conditions Indicates candidate environment domain Category tags corresponding to user feedback content samples, This represents a hyperparameter used to balance prediction accuracy and the strength of invariance constraints. This represents the gradient of the loss calculated for each candidate environment domain corresponding to the classifier at a scalar value of 1.0.

[0104] Then, multiple sets of initial suppression parameters corresponding to multiple candidate environment domains are adjusted to reduce invariant loss, resulting in multiple sets of suppression parameters.

[0105] According to embodiments of this application, multiple sets of initial suppression parameters corresponding to multiple candidate environment domains are adjusted to indicate the degree of suppression of information related to the corresponding candidate environment domain in multiple multimodal sample features after updating.

[0106] For example, the feature extractor is trained with the goal of minimizing the invariant loss. During the iterative training process, the model parameters of the feature extractor are updated, and the updated model parameters are used for the next training to reduce the invariant loss until a preset termination condition is reached, and the final model parameters are obtained. The model parameters include the suppression parameters corresponding to each candidate environment domain.

[0107] Based on this, environment invariance constraints are embedded into the feature extractor during its training process. This allows the feature extractor to fully exploit the invariant features shared between candidate environment domains while processing input features, and suppress weakly correlated features that are only effective in specific candidate environment domains. As a result, user feedback classification has stable predictive ability across different environment domains.

[0108] To ensure that the obtained single-modal features maximize the content with a direct causal relationship to the category to be predicted under the corresponding modality type, and that multiple single-modal features reside in the same causal feature space, in some embodiments, a decoder matching the modality type of the multimodal features is invoked to process the same causal feature. The method further includes:

[0109] First, based on the multiple modality types of multimodal features, multiple decoders are invoked.

[0110] According to embodiments of this application, the modal type includes at least one of text, speech, image, audio, and video.

[0111] Then, the causal features are processed by multiple decoders to obtain multiple single-modal features.

[0112] According to an embodiment of this application, multiple single-modal features are obtained by decoding the same causal feature. The multiple single-modal features have a direct causal relationship with the category to be predicted. For each decoder that matches the modality type, a training set is constructed using the causal sample features and the corresponding single-modal sample features to train the decoder.

[0113] For example, for each modality type, single-modal features Represented as:

[0114] ;

[0115] In the formula, This indicates the decoder that matches the modality type. Indicates causal characteristics.

[0116] Here, the decoder can be a multilayer perceptron trained using data of the corresponding modality type.

[0117] To assist the decoder in converting causal features into single-modal features of different modal types without destroying the core information of causal features, and to avoid the decoder's output deviating from the causal relationship due to environmental changes, in some embodiments, the causal features and the environmental features 420 of the target environment domain are respectively input into multiple trained decoders to obtain multiple single-modal features.

[0118] According to an embodiment of this application, firstly, causal features and environmental features are mapped to the same target dimension and concatenated along the feature dimension direction; then, the concatenated features are input into a trained decoder that matches the modality type for processing to obtain the corresponding single-modality features.

[0119] For example, when training the decoder, a training set is constructed based on the obtained causal sample features, environmental features, and single-modal sample features. With the goal of minimizing cross-entropy loss, the decoder of the corresponding modality type is trained iteratively until a preset stopping condition is reached.

[0120] For example, for each modality type, the single-modal feature is represented as:

[0121] .

[0122] To deeply explore the complementarity between features corresponding to multiple modalities and improve classification accuracy and robustness, a dedicated agent is designed for each modality. These agents work in parallel, each capable of processing its own preferred information type, and combining... Figure 3 Operation S230 includes:

[0123] First, multiple agents are invoked based on multiple modality types matched by multiple single-modal features.

[0124] According to embodiments of this application, the modal type can be text, voice, image, or video.

[0125] Then, multiple agents classify and predict multiple single-modal features to obtain multiple predicted features.

[0126] According to embodiments of this application, the predicted features include multiple candidate predicted categories and corresponding classification probabilities.

[0127] For example, an image agent is used to process the single-modal features of image type to obtain prediction feature_1, which is a probability vector corresponding to multiple candidate prediction categories; a text agent is used to process the single-modal features of text type to obtain prediction feature_2; a speech agent is used to process the single-modal features of speech type to obtain prediction feature_3; and a video agent is used to process the single-modal features of video type to obtain prediction feature_4.

[0128] For example, the image agent is represented as:

[0129] ;

[0130] ;

[0131] In the formula, Represents the predicted features that match the image. This indicates that the classification probability of the first candidate predicted category is 0.92, the classification probability of the second candidate predicted category is 0.05, and the classification probability of the third candidate predicted category is 0.03. This represents an image-based intelligent agent.

[0132] Here, the image agent can also be an image classifier, consisting of a linear layer and a normalization layer connected in sequence.

[0133] To improve the robustness and accuracy of user feedback classification by focusing more on features robust to the current environment and effective information in the current model type, operation S240 includes the following in some embodiments:

[0134] First, based on the multiple correlations between the target environment domain and multiple modal types, multiple weights 480 corresponding to multiple modal types are determined.

[0135] For example, multiple weights 480 corresponding to multiple modality types can be determined and normalized by calculating multiple Pearson correlations between the target environment domain and multiple modality types.

[0136] Then, multiple predicted features are weighted and fused using multiple weights of 480 to obtain fused features.

[0137] For example, multiple weights 480 corresponding to multiple modality types are multiplied by the predicted features of the corresponding modality types and then summed to obtain fused features.

[0138] Then, classification is performed based on the fused features to obtain the target predicted category.

[0139] For example, a trained fine-grained classifier is used to process the fused features to obtain the target predicted category. The fine-grained classifier can be a multi-layer fully connected network, etc.

[0140] The outputs of multiple agents are dynamically weighted based on the current input modality type and identified environmental features. This flexibly fuses the outputs of agents from various modalities while avoiding the erroneous activation of modalities that do not exist, thereby achieving optimal prediction for the current situation. Figure 4 Multiple predictive features are weighted and fused using multiple weights of 480 to obtain the fused features, which include:

[0141] First, a query matrix 440 is obtained based on the environmental features 420 of the target environment domain, and multiple key matrices 430 are obtained based on multiple submodal features 410.

[0142] According to an embodiment of this application, query matrix 440 represents the target environment domain's requirement for modal types, and key matrix represents the attributes of the corresponding submodal features.

[0143] For example, environmental features are mapped to the target dimension as a query matrix 440 through linear transformation, and multiple submodal features 410 are mapped to the target dimension as multiple key matrices 430 and multiple value matrices, respectively, through linear transformation.

[0144] Then, attention mechanism 450 is applied to multiple key matrices 430 and query matrices 440 to obtain multiple attention score vectors 460.

[0145] According to embodiments of this application, the attention score vector indicates the relevance between the target environment domain and the corresponding modality type.

[0146] For example, the i-th attention score vector Represented as:

[0147] ;

[0148] In the formula, This indicates a query for matrix 440. This represents the transpose of the i-th key matrix. This represents the dimension of the key matrix.

[0149] Then, based on the modality mask vector 470, the multiple attention score vectors 460 are normalized to obtain multiple weights 480.

[0150] According to an embodiment of this application, multiple vector values ​​in the modality mask vector 470 are used to characterize the existence state of multiple modality types.

[0151] For example, the modal mask vector 470 can be represented as Where 4 is the total number of modes, for example, This indicates that the input consists only of images and text. If the mask is 1, the original value of the fraction vector is used; if the mask is 0, it is set to a very large negative value, so that after normalization it will be infinitely close to 0. This ensures that the system can work normally when some modalities are missing, avoiding the assignment of weights to missing modalities.

[0152] For example, weight Represented as:

[0153] ;

[0154] Fusion features Represented as:

[0155] ;

[0156] In the formula, This indicates attention normalization. Indicates candidate environment domain Lower mode type Attention score vector after masking This represents the predicted features that match the modality type. Represents a set of modal types.

[0157] Based on this, the output of each agent can be flexibly adjusted by using multiple weights 480. For example, if the user feedback content is weak text, the output weight of the text agent can be increased, thereby achieving the optimal prediction for the current situation.

[0158] To improve the accuracy of secondary classification, a multi-layer fully connected network and The final target prediction category is output. According to the embodiments of this application, the target prediction category is obtained by classifying based on the fusion features. The process includes: first, mapping the fusion features to the hidden feature space through a linear transformation to obtain the first hidden feature.

[0159] According to embodiments of this application, fused features indicate multiple predicted categories for user feedback content, and the hidden feature space indicates the degree of correlation between the multiple predicted categories.

[0160] Specifically, the fused features are input into the hidden layer for processing, and the fused features are subjected to linear transformation.

[0161] For example, the first hidden feature Represented as:

[0162] ;

[0163] In the formula, Indicates the first weight. This indicates the first bias term.

[0164] Then, information unrelated to the target prediction category is filtered out from the first hidden feature through nonlinear activation processing to obtain the second hidden feature.

[0165] Specifically, the first hidden feature is processed by the ReLU activation function to filter out negative values ​​in the first hidden feature in order to suppress irrelevant or negative features.

[0166] For example, the second hidden feature Represented as:

[0167] .

[0168] Then, the second hidden feature is mapped to the target predicted category through a linear transformation.

[0169] Specifically, the second hidden feature is input into the hidden layer for processing, and the output of the hidden layer is normalized by attention to obtain the target prediction category.

[0170] For example, target prediction category Represented as:

[0171] ;

[0172] In the formula, Indicates the second weight. This indicates the second bias term.

[0173] For example, if the user feedback content is multimodal information related to after-sales issues, the target prediction category can be "battery swelling" or "screen cracking", which can be used to interface with the enterprise service system to realize the automatic generation of after-sales work orders; for example, based on the classification results, corresponding repair, customer service or software support work orders can be generated to improve after-sales processing efficiency and customer satisfaction.

[0174] In some embodiments, an invariant loss function is used to train the embedding layer and feature extractor, a cross-entropy loss function is used to train multiple decoders and multiple agents, and a multi-layer fully connected network is trained using the cross-entropy loss function. The function, the overall loss function of the user feedback method described in this embodiment. Represented as:

[0175] ;

[0176] In the formula, , Indicates hyperparameters, Represents the cross-entropy loss function. Indicates the predicted feature label, This indicates the predicted category label.

[0177] Based on the above-described user feedback classification method, this application also provides a user feedback classification device. The following will combine... Figure 5 The device is described in detail.

[0178] Figure 5 A schematic block diagram of a user feedback classification device according to an embodiment of this application is shown.

[0179] like Figure 5 As shown, the user feedback classification device 500 in this embodiment includes an acquisition module 510, a causal reasoning module 520, a multi-agent collaboration module 530, and a classification prediction module 540.

[0180] The acquisition module 510 is used to, in response to receiving user feedback content, determine a target environment domain that matches the user feedback content from multiple candidate environment domains based on the metadata of the user feedback content. Each of the multiple candidate environment domains has environment characteristics corresponding to the metadata. In one embodiment, the acquisition module 510 can be used to perform the operation S210 described above, which will not be repeated here.

[0181] The causal reasoning module 520 is used to suppress information related to the target environment domain in the multimodal features of user feedback content based on the environmental features 420 of the target environment domain, thereby obtaining causal features. The suppression processing is used to weaken the causal relationship between information and the category to be predicted. The causal features characterize the features in the multimodal features that have a causal relationship with the category to be predicted. In one embodiment, the causal reasoning module 520 can be used to perform the operation S220 described above, which will not be repeated here.

[0182] The multi-agent collaboration module 530 is used to input multiple single-modal features obtained from decoding causal features into multiple agents respectively, resulting in multiple predicted features. Each single-modal feature is matched with multiple modality types of the multi-modal features, and the predicted features indicate the candidate prediction category for the user feedback content. In one embodiment, the multi-agent collaboration module 530 can be used to perform the operation S230 described above, which will not be repeated here.

[0183] The classification prediction module 540 is used to obtain the target predicted category of the user feedback content based on multiple prediction features. In one embodiment, the classification prediction module 540 can be used to perform the operation S240 described above, which will not be repeated here.

[0184] According to an embodiment of this application, the acquisition module 510 includes an acquisition submodule and a similarity calculation submodule. The acquisition submodule is used to obtain a vector to be matched based on the metadata of the user feedback content. The metadata of the user feedback content includes at least one of device model, language, region and sales channel. The similarity calculation submodule is used to calculate multiple similarities between the vector to be matched and multiple candidate environment domains, so as to determine the target environment domain based on the multiple similarities.

[0185] According to an embodiment of this application, the causal reasoning module 520 includes a retrieval submodule and a feature extraction submodule. The retrieval submodule is used to determine the suppression parameter that matches the target environment domain from multiple sets of suppression parameters based on the environmental features 420 of the target environment domain. The suppression parameter indicates the degree of filtering of environmental noise related to the target environment domain. The feature extraction submodule is used to process multimodal features based on the suppression parameter to obtain causal features.

[0186] According to an embodiment of this application, the multi-agent collaboration module 530 includes a calling submodule and a classification prediction submodule. The calling submodule is used to call multiple agents based on multiple modal types matched by multiple single-modal features. The modal types include at least one of text, speech, image, audio and video. The classification prediction submodule is used to perform classification prediction on multiple single-modal features by multiple agents respectively to obtain multiple predicted features. The predicted features include multiple candidate predicted categories and corresponding classification probabilities.

[0187] According to an embodiment of this application, the classification prediction module 540 includes a relevance calculation submodule, a fusion submodule, and a prediction submodule. The relevance calculation submodule is used to determine multiple weights 480 corresponding to multiple modal types based on multiple relevances between the target environment domain and multiple modal types. The fusion submodule is used to perform weighted fusion of multiple prediction features using multiple weights 480 to obtain fused features. The prediction submodule is used to perform classification based on the fused features to obtain the target prediction category.

[0188] According to an embodiment of this application, the device further includes a feature extraction module, which is used to call multiple decoders based on multiple modal types of multimodal features. The modal types include at least one of text, speech, image, audio and video. The causal features are processed by the multiple decoders respectively to obtain multiple single-modal features, and the multiple single-modal features have a direct causal relationship with the category to be predicted.

[0189] According to an embodiment of this application, the device further includes a decoding module, which is used to invoke multiple agents based on multiple modal types matched by multiple single-modal features. The modal types include at least one of text, speech, image, audio and video. The multiple agents perform classification prediction on the multiple single-modal features respectively to obtain multiple predicted features, which include multiple candidate predicted categories and corresponding classification probabilities.

[0190] According to embodiments of this application, any multiple modules among the acquisition module 510, causal reasoning module 520, multi-agent collaboration module 530, and classification prediction module 540 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least some of the functions of one or more of these modules can be combined with at least some of the functions of other modules and implemented in one module. According to embodiments of this application, at least one of the acquisition module 510, causal reasoning module 520, multi-agent collaboration module 530, and classification prediction module 540 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the acquisition module 510, the causal reasoning module 520, the multi-agent collaboration module 530, and the classification prediction module 540 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.

[0191] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing a user feedback classification method according to an embodiment of this application.

[0192] like Figure 6 As shown, an electronic device 600 according to an embodiment of this application includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 606 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor, and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.

[0193] RAM 603 stores various programs and data required for the operation of electronic device 600. Processor 601, ROM 602, and RAM 603 are interconnected via bus 604. Processor 601 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.

[0194] According to embodiments of this application, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to a bus 604. The electronic device 600 may also include one or more of the following components connected to the input / output (I / O) interface 605: an input section 608 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 606 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 606 as needed.

[0195] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.

[0196] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603 described above.

[0197] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to enable the computer system to implement the user feedback classification method provided in the embodiments of this application.

[0198] When the computer program is executed by the processor 601, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0199] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 609, and / or installed from the removable medium 611. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0200] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0201] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0202] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0203] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.

[0204] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.

Claims

1. A user feedback classification method, characterized in that, The method includes: In response to receiving user feedback content, based on the metadata of the user feedback content, a target environment domain that matches the user feedback content is determined from multiple candidate environment domains, each of the multiple candidate environment domains having environment characteristics corresponding to the metadata; Based on the environmental characteristics of the target environment domain, the information related to the target environment domain in the multimodal features of the user feedback content is suppressed to obtain causal features. The suppression process is used to weaken the causal relationship between the information and the category to be predicted. The causal features characterize the features in the multimodal features that have a causal relationship with the category to be predicted. Multiple single-modal features obtained by decoding the causal features are input into multiple agents to obtain multiple prediction features. The multiple single-modal features are matched with multiple modality types of the multimodal features, and the prediction features indicate the candidate prediction categories for the user feedback content. The target predicted category of the user feedback content is obtained based on the multiple predictive features; The environmental features mentioned above are obtained using the following methods: Based on clustering processing, multiple metadata samples of multiple user feedback content samples are divided into multiple clusters to serve as multiple candidate environment domains. The multiple metadata samples include device model, language, region, and sales channel. Obtain multiple multimodal sample features of the multiple user feedback content samples, and multiple initial environmental features corresponding to the multiple candidate environment domains respectively; Based on the multiple initial environmental features, the information related to the corresponding candidate environmental domain in the multiple multimodal sample features is suppressed to obtain multiple causal sample features. The corresponding candidate environmental domain is determined based on the metadata samples in the user feedback content samples. Multiple classifiers are used to process the features of the multiple causal samples respectively to obtain multiple classification prediction results, and the multiple classifiers correspond to the multiple candidate environment domains; Based on the multiple deviations between the multiple classification prediction results and the multiple classification labels of the multiple user feedback content samples, the invariant loss during suppression processing is determined, and the invariant loss indicates the difference in prediction accuracy among the multiple candidate environment domains. The plurality of initial environmental features are adjusted to reduce the invariant loss, thereby obtaining the plurality of environmental features.

2. The method according to claim 1, characterized in that, The target prediction category for the user feedback content obtained based on the multiple prediction features includes: Based on the multiple correlations between the target environment domain and the multiple modal types, multiple weights corresponding to the multiple modal types are determined; The multiple predicted features are weighted and fused using the multiple weights to obtain the fused features; The target prediction category is obtained by classifying based on the fused features.

3. The method according to claim 2, characterized in that, The multimodal features include multiple submodal features with different modality types. Each submodal feature includes a single modality feature corresponding to the modality type and corresponding environment domain-related information. Determining multiple weights corresponding to the multiple modality types based on multiple correlations between the target environment domain and the multiple modality types includes: A query matrix is ​​obtained based on the environmental features of the target environment domain, and multiple key matrices are obtained based on the multiple submodal features; Attention mechanisms are applied to the multiple key matrices and the query matrix to obtain multiple attention score vectors, which indicate the relevance between the target environment domain and the corresponding modality type. The multiple attention score vectors are normalized based on the modality mask vector to obtain the multiple weights. The multiple vector values ​​in the modality mask vector are used to characterize the existence state of the multiple modality types. The multiple submodal features are obtained using the following method: Based on multiple modal information corresponding to multiple modal types in the user feedback content, multiple corresponding feature extraction models are invoked, wherein the modal types include at least one of text, speech, image, audio and video; The multiple feature extraction models are used to encode the multiple modal information to obtain the multiple submodal features.

4. The method according to claim 2, characterized in that, The classification based on the fused features to obtain the predicted target category includes: The fused features are mapped to the hidden feature space by a linear transformation to obtain the first hidden features. The fused features indicate multiple predicted categories for the user feedback content, and the hidden feature space indicates the degree of correlation between the multiple predicted categories. The second hidden feature is obtained by filtering out information unrelated to the target prediction category from the first hidden feature through nonlinear activation processing; The second hidden feature is mapped to the target predicted category through a linear transformation.

5. The method according to claim 1, characterized in that, The step of determining the target environment domain that matches the user feedback content from multiple candidate environment domains based on the metadata of the user feedback content includes: Based on the metadata of the user feedback content, a matching vector is obtained, wherein the metadata of the user feedback content includes at least one of the following: device model, language, region, and sales channel; Calculate multiple similarities between the vector to be matched and the multiple candidate environment domains, and determine the target environment domain based on the multiple similarities.

6. The method according to claim 1, characterized in that, Each of the multiple candidate environment domains has a suppression parameter. Based on the environmental features of the target environment domain, the information related to the target environment domain in the multimodal features of the user feedback content is suppressed to obtain causal features, including: Based on the environmental characteristics of the target environment domain, a suppression parameter matching the target environment domain is determined from multiple sets of suppression parameters. The suppression parameter indicates the degree of filtering of environmental noise related to the target environment domain. The multimodal features are processed based on the suppression parameters to obtain the causal features.

7. The method according to any one of claims 1 to 6, characterized in that, Each of the multiple candidate environment domains has a suppression parameter, and the multiple sets of suppression parameters are obtained using the following method: Clustering is performed on multiple metadata samples based on multiple user feedback content samples to obtain multiple candidate environment domains. The multiple metadata samples include device model, language, region, and sales channel. Based on the multiple environmental features corresponding to the multiple candidate environmental domains, the information related to the corresponding candidate environmental domain in the multiple multimodal sample features of the multiple user feedback content samples is suppressed to obtain multiple causal sample features. The corresponding candidate environmental domain is determined based on the metadata samples in the user feedback content samples. Multiple classifiers are used to process the features of the multiple causal samples respectively to obtain multiple classification prediction results, and the multiple classifiers correspond to the multiple candidate environment domains; Based on the multiple deviations between the multiple classification prediction results and the multiple classification labels of the multiple user feedback content samples, the invariant loss during suppression processing is determined, and the invariant loss indicates the difference in prediction accuracy among the multiple candidate environment domains. The multiple sets of initial suppression parameters corresponding to the multiple candidate environment domains are adjusted to reduce the invariant loss, thus obtaining the multiple sets of suppression parameters.

8. The method according to claim 1, characterized in that, The process of decoding multiple single-modal features based on the causal features and inputting them into multiple agents to obtain multiple predicted features includes: Based on the multiple modality types matched by the multiple single-modal features, multiple intelligent agents are invoked, wherein the modality types include at least one of text, speech, image, audio and video; The multiple agents classify and predict the multiple single-modal features respectively to obtain the multiple predicted features, which include multiple candidate predicted categories and corresponding classification probabilities; The multiple single-modal features are obtained using the following method: Based on the multiple modality types of the aforementioned multimodal features, multiple decoders are invoked; The causal features are processed by the multiple decoders to obtain the multiple single-modal features, which have a direct causal relationship with the category to be predicted.

9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Social media sentiment analysis method and system based on multi-modal feature fusion

    CN112508077A

  • Cross-domain voice classification method and device based on feature decoupling and multi-task learning

    CN120452429A