A Multimodal Learning Automatic Complaint Content Analysis and Classification Method and System

Through multimodal learning, automated complaint content analysis and classification methods, the problems of low processing efficiency and inflexible response in the existing technology are solved, and fast and accurate complaint content classification and personalized processing strategies are realized, which improves customer satisfaction and system adaptability.

CN119669867BActive Publication Date: 2025-07-22JIANGSU HUCHUAN TECH CO LTD

Patent Information

Application Number
CN202510187948.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-22
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing complaint handling methods have problems such as inefficient processing efficiency and inflexible response, which is difficult to effectively integrate and utilize multimodal data, and lack dynamic learning and adaptability, resulting in insufficient classification accuracy and response timeliness.

Method used

Multimodal learning automated complaint content analysis and classification method is adopted, and deep learning models are constructed for real-time classification through data preprocessing, feature extraction and fusion, and the processing strategy is optimized by combining self-attention mechanisms and dynamic decision engines.

Benefits of technology

It realizes fast and accurate classification of complaint content, automatically triggers emergency response, optimizes resource allocation, and improves the intelligence level of processing strategies and customer satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669867B_ABST
    Figure CN119669867B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal learning automated complaint content analysis and classification method and system, which relates to the technical field of semantic analysis, and includes: collecting multi-modal complaint data and preprocessing the collected multi-modal complaint data; extracting the features of the preprocessed multi-modal complaint data and fusing the features of different modal data; using the fused features to construct a complaint content classification model, performing model training, evaluating and optimizing the trained model; and deploying the optimized model for real-time classification. The multi-modal learning automated complaint content analysis and classification method provided by the present invention classifies complaint content quickly and accurately, and proposes targeted processing strategies according to the main categories, emotional tendencies and specific topics of complaints. Automatically trigger an emergency response mechanism to ensure that problems are processed immediately and ensure the effective utilization of resources. Learn from historical data, continuously optimize decision rules, and improve the intelligence level of processing strategies and customer satisfaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of semantic analysis, and specifically to a multi-modal learning automated complaint content analysis and classification method and system. Background Art

[0002] In current multi-modal complaint handling systems, the main challenge is how to efficiently analyze and respond to complex complaints covering text, images, audio, and contextual data. Traditional processing methods often rely on analyzing these different types of data in isolation, lacking effective mechanisms to integrate and utilize the potential associations between these data, which may lead to the neglect of important connections in the processing, affecting the accuracy of classification and the timeliness of response.

[0003] In addition, existing automated complaint handling systems usually make decisions using fixed rules, which lack flexibility and cannot self-adjust according to real-time data or historical processing effects. Such systems lacking dynamic learning and adaptation capabilities perform poorly in handling sudden or complex complaint situations, unable to provide personalized processing strategies and difficult to meet the growing customer service expectations.

[0004] Therefore, there is an urgent need for a multi-modal learning automated complaint content analysis and classification method to optimize the analysis process of complaint handling, provide more intelligent and personalized technical solutions, thereby improving processing efficiency and customer satisfaction. Summary of the Invention

[0005] In view of the above existing problems, the present invention is proposed.

[0006] Therefore, the technical problems solved by the present invention are: the existing complaint handling methods have problems of low processing efficiency and inflexible response, and the optimization problems of how to comprehensively utilize multi-modal data and dynamically optimize processing strategies.

[0007] To solve the above technical problems, the present invention provides the following technical solutions: a multi-modal learning automated complaint content analysis and classification method, including:

[0008] Collect multi-modal complaint data and preprocess the collected multi-modal complaint data;

[0009] Extract the features of the preprocessed multi-modal complaint data and fuse the features of different modal data;

[0010] Use the fused features to construct a complaint content classification model, conduct model training, evaluate and optimize the trained model;

[0011] Deploy the optimized model for real-time classification.

[0012] As a preferred solution of the multi-modal learning automated complaint content analysis and classification method described in the present invention, wherein: the multi-modal complaint data includes complaint data and context data;

[0013] The complaint data includes text data, image data, and audio data;

[0014] The preprocessing includes data cleaning, data standardization, format conversion, and data integration.

[0015] As a preferred solution of the multi-modal learning automated complaint content analysis and classification method described in the present invention, wherein: extracting the features of the preprocessed multi-modal complaint data includes constructing a multi-input feature extraction model, processing text, image, audio, and context data through four branches, extracting the features of each modality, and forming a preliminary feature representation;

[0016] Introduce a feature interaction layer, apply the self-attention mechanism to optimize the feature relationship, and generate a comprehensive joint feature representation through the feature interaction mechanism;

[0017] Based on the joint feature representation generated by the feature interaction layer, design a context-aware fusion strategy, dynamically adjust the feature fusion weights, and optimize the fusion effect.

[0018] As a preferred solution of the multi-modal learning automated complaint content analysis and classification method described in the present invention, wherein: extracting the features of each modality includes using the BERT model to encode the text data, extract semantic features, and convert them into text feature vectors;

[0019] Use the pre-trained ResNet model to encode the image data, extract the visual features of the image, and convert them into image feature vectors;

[0020] Use the MFCC technology to extract the spectral features of the audio, and use a convolutional neural network to encode the extracted audio features and convert them into audio feature vectors;

[0021] Convert the context data into numerical features through the embedding layer, generate context feature vectors, and integrate them with the features of other modalities within the model.

[0022] As a preferred solution of the multi-modal learning automated complaint content analysis and classification method described in the present invention, wherein: the generated joint feature representation includes weighting the feature vectors extracted from text, image, audio, and context data;

[0023] By calculating the similarity between each feature vector and other feature vectors, the self-attention mechanism automatically adjusts and optimizes the weights and associations between different features;

[0024] After the self-attention mechanism processes, the feature interaction layer fuses and interactively learns all the weighted feature vectors, captures the potential correlations between modalities, and combines and optimizes the features of each modality through the feature interaction mechanism;

[0025] When the feature interaction layer finishes processing, a joint feature representation is generated;

[0026] The optimization fusion strategy includes generating a dynamically weighted adjusted joint feature representation based on the joint feature representation and according to the context information analyzed in real time;

[0027] By training the optimized joint feature representation, using the patterns in the historical data and the information feedback in real time, the feature fusion strategy is further adjusted to generate the finally optimized joint feature representation.

[0028] As a preferred solution of the multi-modal learning automated complaint content analysis and classification method described in the present invention, wherein: constructing the complaint content classification model includes using the finally optimized joint feature representation as input to construct a multi-output deep learning model. The underlying feature extraction layer of the model first shares the input joint feature representation and performs preliminary feature processing through a multi-layer neural network;

[0029] The model is divided into three parallel task branches, and the task branches include a main classification task branch, a sentiment analysis task branch, and a topic recognition task branch;

[0030] The main classification task branch uses the features further processed from the primary feature layer, conducts in-depth learning through a multi-layer fully connected neural network, and finally outputs the main category of the complaint;

[0031] The sentiment analysis task branch starts from the shared primary feature layer, adds neural layers to capture emotion features, and the neural layers focus on understanding the sentiment tendency, and finally outputs the sentiment tendency of the complaint;

[0032] The topic recognition task branch starts from the shared primary feature layer, adds a text analysis layer to process text data, identifies and classifies the key topics in the text, and finally outputs the specific topic of the complaint;

[0033] In the training stage of the model, a comprehensive loss function is defined to simultaneously consider the task effects of main classification, sentiment analysis, and topic recognition, and a multi-task learning strategy is used to optimize the branches;

[0034] The weights and biases of the model are adjusted through the gradient descent algorithm, and the Adam optimization algorithm is used for dynamic learning rate adjustment;

[0035] After each training cycle ends, an independent validation set is used to evaluate the performance of the model, and hyperparameters are adjusted according to the evaluation results.

[0036] As a preferred solution of the multi-modal learning automated complaint content analysis and classification method of the present invention, wherein: the real-time classification includes, after the model is deployed, receiving new complaint data in real time, preprocessing the data and extracting joint feature representations, and inputting them into the model for classification. The model simultaneously outputs the main category, sentiment tendency and specific theme of the complaint according to the input features;

[0037] The extracted features are input into the trained deep learning model, and the main category, sentiment tendency and specific theme of the complaint are output synchronously;

[0038] The output result is then sent to the dynamic decision engine. The dynamic decision engine constructs a decision model based on a decision tree and generates processing strategies for various situations;

[0039] If the main category is a security issue, trigger an emergency response mechanism, start an automated monitoring and instant reporting process, monitor the system status in real time, and automatically send emergency notifications to relevant departments and management to ensure instant handling;

[0040] If the main category is product quality or service feedback, enter the sentiment analysis process;

[0041] For negative sentiment, introduce emotion recognition technology to further analyze the urgency and severity of the complaint, and adjust the response priority and resource allocation;

[0042] For positive or neutral sentiment, conduct satisfaction prediction analysis, predict the customer's satisfaction with the current processing progress, and then advance to the theme recognition process;

[0043] According to the result of theme recognition, for a refund request, trigger a financial process, predict the potential refund risk based on the customer's historical data and purchase pattern, and adjust the inventory and pricing strategies accordingly;

[0044] For the theme of product suggestions, integrate customer feedback and generate a report.

[0045] Another object of the present invention is to provide a multi-modal learning automated complaint content analysis and classification system, which can solve the problems of insufficient information utilization in the existing single data source processing and inflexible response in the fixed rule decision system by constructing a multi-modal learning automated complaint content analysis and classification system.

[0046] To solve the above technical problems, the present invention provides the following technical solutions: A multi-modal learning automated complaint content analysis and classification system, comprising: a data collection module, a data fusion module, a model training module, and a classification prediction module; the data collection module is used to collect multi-modal complaint data and preprocess the collected multi-modal complaint data; the data fusion module is used to extract the features of the preprocessed multi-modal complaint data and fuse the features of different modal data; the model training module is used to perform model training using the fused features, evaluate and optimize the trained model; the classification prediction module is used to deploy the optimized model for real-time classification.

[0047] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned multi-modal learning automated complaint content analysis and classification method are implemented.

[0048] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned multi-modal learning automated complaint content analysis and classification method are implemented.

[0049] The beneficial effects of the present invention: The multi-modal learning automated complaint content analysis and classification method provided by the present invention, after receiving new complaint data in real time, quickly and accurately classifies the complaint content, and proposes targeted processing strategies according to the main categories, sentiment tendencies, and specific topics of the complaints. Automatically trigger an emergency response mechanism to ensure that problems are handled immediately. Adjust the processing priority and resource allocation according to sentiment analysis to ensure the effective utilization of resources. Through a context-aware dynamic decision-making engine, it can learn from historical data and continuously optimize decision rules, improving the intelligent level of processing strategies and customer satisfaction. Description of the Drawings

[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is the overall flowchart of a multi-modal learning automated complaint content analysis and classification method provided by an embodiment of the present invention.

[0052] Figure 2 It is the overall structure diagram of a multi-modal learning automated complaint content analysis and classification system provided by the second embodiment of the present invention. Detailed Embodiments

[0053] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0054] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0055] Embodiment 1

[0056] Referring to Figure 1 , for an embodiment of the present invention, a multi-modal learning automated complaint content analysis and classification method is provided, including:

[0057] Collect multi-modal complaint data and preprocess the collected multi-modal complaint data;

[0058] Extract the features of the preprocessed multi-modal complaint data and fuse the features of different modal data;

[0059] Use the fused features to construct a complaint content classification model, perform model training, evaluate, and optimize the trained model;

[0060] Deploy the optimized model for real-time classification.

[0061] The multi-modal complaint data includes complaint data and context data;

[0062] The complaint data includes text data, image data, and audio data;

[0063] The preprocessing includes data cleaning, data standardization, format conversion, and data integration.

[0064] The context data includes device type, error code, timestamp, geographical location information, and user historical data.

[0065] The device type and error code collect technical parameters such as the device model and error code related to the complaint.

[0066] The timestamp and geographical location information collect the specific time and location of the complaint occurrence to provide the environmental background of the event.

[0067] The user historical data includes the user's past complaint records, service usage history, and any relevant user behavior patterns.

[0068] Text cleaning removes HTML tags, URL links, and non-printable characters, and uses regular expressions to clean special symbols and numbers to purify the text data and make it more suitable for subsequent NLP processing.

[0069] Image cleaning automatically detects and repairs damaged image files, removes background noise, and improves image quality to prepare for subsequent image feature extraction.

[0070] Audio cleaning filters out background noise and echoes, normalizes the volume, and clarifies the audio data to provide accurate input for audio feature extraction.

[0071] Unified formatting processing ensures the consistency of all data in time and space, which helps to consider spatio-temporal factors in subsequent analysis.

[0072] The features of the preprocessed multi-modal complaint data extraction include constructing a multi-input feature extraction model, processing text, image, audio, and context data through four branches, extracting features of each modality, and forming a preliminary feature representation;

[0073] Introduce a feature interaction layer, apply the self-attention mechanism to optimize feature relationships, and generate a comprehensive joint feature representation through the feature interaction mechanism;

[0074] Based on the joint feature representation generated by the feature interaction layer, design a context-aware fusion strategy, dynamically adjust the feature fusion weights, and optimize the fusion effect.

[0075] For the feature extraction of multi-modal data, by constructing a multi-input feature extraction model and using the self-attention mechanism and the feature interaction layer, the relevance and interaction between different modality features are optimized, effectively improving the model's understanding and processing ability of various data types, making the features extracted from text, image, audio, and context data more accurate and relevant. The dynamic feature fusion strategy further strengthens the response ability to real-time data changes, thus achieving more accurate data analysis and classification in practical applications.

[0076] The extraction of features of each modality includes using the BERT model to encode text data, extract semantic features, and convert them into text feature vectors;

[0077] Using the pre-trained ResNet model to encode image data, extract the visual features of the image, and convert them into image feature vectors;

[0078] Using MFCC technology to extract the spectral features of audio, and using a convolutional neural network to encode the extracted audio features and convert them into audio feature vectors;

[0079] Convert the situational data into numerical features through an embedding layer, generate a situational feature vector, and integrate it with features of other modalities within the model.

[0080] By applying BERT, ResNet, and MFCC technologies, text, image, and audio data are deeply processed. The application of these advanced models significantly improves the ability to extract deep-level features from complex data. Combining the self-attention mechanism and the feature interaction layer optimizes the extraction efficiency of each modality's features, enhances the model's capture of subtle relationships between different data, and makes the final feature representation more comprehensive and refined.

[0081] The generated joint feature representation includes weighting the feature vectors extracted from text, image, audio, and situational data;

[0082] By calculating the similarity between each feature vector and other feature vectors, the self-attention mechanism automatically adjusts and optimizes the weights and associations between different features, ensuring that the comprehensive feature vector can fully reflect the interaction of multimodal data and enhancing the correlation between features;

[0083] After the self-attention mechanism processes, the feature interaction layer fuses and interactively learns all the weighted feature vectors, captures the potential associations between modalities, and combines and optimizes the features of each modality through the feature interaction mechanism;

[0084] When the feature interaction layer finishes processing, a joint feature representation is generated, generating a more informative joint feature representation and enhancing the expression ability and analysis effect of the overall features;

[0085] The optimization fusion strategy includes, based on the joint feature representation, generating a dynamically weighted adjusted joint feature representation according to the context information analyzed in real time;

[0086] By training the optimized joint feature representation, using the patterns in historical data and the information feedback in real time, further adjust the feature fusion strategy to generate the final optimized joint feature representation.

[0087] The generated joint feature representation is further processed by the self-attention mechanism, strengthening the interaction and information fusion between features. By calculating the similarity between feature vectors and optimizing the associations between them, this step not only inherits the detailed features of the previous model but also improves the expression ability and classification accuracy of the overall features through the composite operation of the feature interaction layer. The adjustment of dynamic weights enables the model to flexibly adjust according to context information, enhancing the system's adaptability to different situations.

[0088] The construction of the complaint content classification model includes using the finally optimized joint feature representation as input to construct a multi-output deep learning model. The underlying feature extraction layer of the model first shares the input joint feature representation and performs preliminary feature processing through a multi-layer neural network;

[0089] Primary features refer to the basic features directly extracted from the joint feature representation that have not been optimized for any specific analysis task. Primary features provide a general feature basis for further task-specific processing.

[0090] The model is divided into three parallel task branches, each branch for a specific analysis task. The task branches include the main classification task branch, the sentiment analysis task branch, and the topic recognition task branch;

[0091] The main classification task branch uses the features further processed from the primary feature layer and conducts in-depth learning through a multi-layer fully connected neural network, and finally outputs the main category of the complaint;

[0092] The sentiment analysis task branch starts from the shared primary feature layer, adds specialized neural layers to capture emotion features. The neural layers focus on understanding the sentiment tendency, and finally output the sentiment tendency of the complaint;

[0093] The topic recognition task branch starts from the shared primary feature layer, adds a text analysis layer to process text data, identifies and classifies the key topics in the text, and finally outputs the specific topic of the complaint;

[0094] In the training stage of the model, a comprehensive loss function is defined to simultaneously consider the task effects of main classification, sentiment analysis, and topic recognition, and a multi-task learning strategy is used to optimize the branches;

[0095] The weights and biases of the model are adjusted through the gradient descent algorithm, and the Adam optimization algorithm is used for dynamic learning rate adjustment;

[0096] After each training cycle, an independent validation set is used to evaluate the performance of the model, and hyperparameters are adjusted according to the evaluation results;

[0097] The adjustments include the adjustment of the learning rate, the configuration of the number of neural network layers, etc., to ensure that the model has good generalization ability on each task.

[0098] After the model training is completed, an unseen dataset is used to evaluate the classification accuracy and robustness of the model on actual complaint data to ensure that the model can effectively perform the classification task in the actual environment.

[0099] The multi-output deep learning model uses the finally optimized joint feature representation as input, effectively achieving multi-task learning. This model structure allows each task branch, such as main classification, sentiment analysis, and topic recognition, to share the primary feature layer and then be independently optimized. This architecture significantly improves the processing efficiency and accuracy. The application of the multi-task learning strategy not only optimizes resource allocation but also ensures the balanced performance of the model on each task through the adjustment of the comprehensive loss function, significantly improving the classification accuracy and response speed in practical applications.

[0100] The real-time classification mentioned above includes, after the model is deployed, receiving new complaint data in real time, preprocessing the data and extracting the joint feature representation, and inputting it into the model for classification. The model simultaneously outputs the main category, sentiment tendency, and specific topic of the complaint according to the input features;

[0101] The main categories include product quality, service feedback, and safety issues.

[0102] The sentiment tendencies include positive, negative, and neutral.

[0103] The specific topics include refund requests, product suggestion reports.

[0104] The extracted features are input into the trained deep learning model, and the main category, sentiment tendency, and specific topic of the complaint are output synchronously;

[0105] The output results are then sent to the dynamic decision engine. The dynamic decision engine constructs a decision model based on decision trees and generates processing strategies for various situations;

[0106] If the main category is a safety issue, trigger the emergency response mechanism, start the automated monitoring and instant reporting process, monitor the system status in real time, and automatically send emergency notifications to relevant departments and management to ensure instant handling;

[0107] If the main category is product quality or service feedback, then enter the sentiment analysis process;

[0108] For negative sentiment, introduce emotion recognition technology to further analyze the urgency and severity of the complaint, and adjust the response priority and resource allocation;

[0109] For positive or neutral sentiment, conduct satisfaction prediction analysis, predict the customer's satisfaction with the current processing progress, and then advance to the topic recognition process;

[0110] According to the results of topic recognition, for refund requests, trigger the financial process, predict the potential refund risk based on the customer's historical data and purchase patterns, and adjust the inventory and pricing strategies accordingly;

[0111] For the theme of product suggestions, integrate customer feedback and generate reports for the reference of the product development team. At the same time, initiate innovation incentives to encourage employees to innovate based on customer suggestions.

[0112] The decision engine optimizes the decision-making path according to the above classifications and historical data, and adjusts the strategy in real time to maximize processing efficiency and customer satisfaction.

[0113] Introduce an adaptive learning mechanism. The decision engine continuously learns from the results of each complaint handling, and adjusts its decision-making algorithm to cope with changing customer needs and situational changes.

[0114] According to the output of the decision engine, automatically trigger corresponding back-end workflows, such as customer follow-up, refund processing, or product quality inspection.

[0115] The workflow realizes seamless execution of operations through integration with enterprise resource planning (ERP) systems and customer relationship management (CRM) systems, ensuring data consistency and operational efficiency.

[0116] Embodiment 2

[0117] Refer to Figure 2 , an embodiment of the present invention provides a multi-modal learning automated complaint content analysis and classification system, including:

[0118] A data collection module, a data fusion module, a model training module, and a classification prediction module;

[0119] The data collection module is used to collect multi-modal complaint data and preprocess the collected multi-modal complaint data;

[0120] The data fusion module is used to extract the features of the preprocessed multi-modal complaint data and fuse the features of different modal data;

[0121] The model training module is used to build a complaint content classification model using the fused features, perform model training, evaluate and optimize the trained model;

[0122] The classification prediction module is used to deploy the optimized model for real-time classification.

[0123] Embodiment 3

[0124] An embodiment of the present invention, which is different from the previous two embodiments:

[0125] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes of various kinds.

[0126] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a predefined sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0127] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), fiber optic devices, and portable compact disc read-only memories (CDROMs). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or, if necessary, other appropriate processing, and then stored in a computer memory.

[0128] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0129] Example 4

[0130] An embodiment of the present invention provides a method for analyzing and classifying multimodal learning automated complaint content, including:

[0131] Collect multimodal complaint data and preprocess the collected multimodal complaint data;

[0132] Extract the features of the preprocessed multimodal complaint data and fuse the features of different modal data;

[0133] Use the fused features to construct a complaint content classification model, perform model training, evaluate and optimize the trained model;

[0134] Deploy the optimized model for real-time classification.

[0135] Multimodal complaint data , including text data , image data , audio data and context data .

[0136] For each data modality, preprocessing is performed. The preprocessing of text data is expressed as:

[0137] ,

[0138] where represents the cleaned text data, and clean is a function for removing noise such as HTML tags, URL links, non-printable characters, etc. The specific implementation is as follows:

[0139] ,

[0140] where RemoveHTML removes HTML tags and RemoveNoise removes noise.

[0141] The preprocessing of image data is expressed as:

[0142] ,

[0143] Among them, represents the preprocessed image data, and preprocess is a function that includes operations such as image cropping, resizing, and normalization. The specific implementation is as follows:

[0144] ,

[0145] Among them, Crop crops the image, Resize adjusts the image size, and Normalize performs normalization.

[0146] The preprocessing of audio data is represented as:

[0147] ,

[0148] Among them, represents the denoised audio data, is a function used to filter out background noise and echo. The specific implementation is as follows:

[0149]

[0150] Among them, Filter filters out background noise.

[0151] The preprocessing of context data is represented as:

[0152] ,

[0153] Among them, represents the standardized context data, and normalize is a function used to uniformly convert all timestamps to the UTC format and standardize the geographical location data to the longitude and latitude format, etc. The specific implementation is as follows:

[0154] ,

[0155] Among them, StandardizeTime unifies the time format, and StandardizeGeo standardizes the geographical location information.

[0156] Extract the features of each data modality through a multi-input feature extraction model.

[0157] The text feature extraction uses the BERT model and is represented as:

[0158] ,

[0159] Among them, represents the extracted text features, is a function that encodes text data using the BERT model to extract deep semantic features. The specific implementation is as follows:

[0160] ,

[0161] Among them, Tokenize tokenizes the text data, and TransformerEncoder encodes the tokenized data.

[0162] For image feature extraction, the ResNet model is used, expressed as:

[0163] ,

[0164] Here represents the extracted image features, is a function that encodes image data using a pre-trained ResNet model to extract high-level visual features. The specific implementation is as follows:

[0165] ,

[0166] Among them, ConvLayers represents the convolutional layers in the ResNet model.

[0167] For audio feature extraction, MFCC and CNN are used, expressed as:

[0168] ,

[0169] Among them, represents the extracted audio features, is a function that extracts the spectral features of audio; is a function that encodes the extracted audio features using a convolutional neural network. The specific implementation is as follows:

[0170] ,

[0171] Among them, MFCC extracts the spectral features of audio, and ConvLayers represents the convolutional layers.

[0172] For context feature extraction, it is expressed as:

[0173] ,

[0174] Among them, represents the extracted context features, Embedding is a function that converts context data into numerical features through an embedding layer. The specific implementation is as follows:

[0175] ,

[0176] Among them, DenseLayer represents the fully connected layer.

[0177] The self-attention mechanism and the feature interaction layer are used to optimize and fuse the extracted features.

[0178] The self-attention mechanism calculates the attention weights, expressed as:

[0179] ,

[0180] Among them, represents the feature and the attention weight between is the feature and the similarity score between, calculated by as follows. The specific implementation is as follows:

[0181] ,

[0182] Among them, and are the weight matrices of the query and the key respectively, is the scaling factor.

[0183] The weighted fusion feature is expressed as:

[0184] ,

[0185] Among them, represents the weighted feature vector, obtained by weighted summation of all feature vectors as follows.

[0186] The features of all modalities are fused to generate a joint feature representation, expressed as:

[0187] ,

[0188] Here represents the joint feature representation, and concat is a function that concatenates the feature vectors of all modalities.

[0189] A multi-output deep learning model is constructed using the joint feature representation, including three task branches: main classification, sentiment analysis, and topic recognition.

[0190] The main classification task branch is expressed as:

[0191] ,

[0192] Among them, Represents the output of the main classification task. Softmax is an activation function. Represents the fully connected layer, and the specific implementation is as follows:

[0193] ,

[0194] Among them, and are the weights and biases of the fully connected layer respectively.

[0195] The sentiment analysis task branch is represented as:

[0196] ,

[0197] Among them, Represents the output of the sentiment analysis task. LSTM is a long short-term memory network used to capture sequence information, and the specific implementation is as follows:

[0198] ,

[0199] Among them, and are the weights and biases of the fully connected layer respectively.

[0200] The topic recognition task branch is represented as:

[0201] ,

[0202] Among them, Represents the output of the topic recognition task. Transformer is a model based on the attention mechanism used to process text data, and the specific implementation is as follows:

[0203] Among them, and are the weights and biases of the fully connected layer respectively.

[0204] Define the comprehensive loss function and optimize the main classification, sentiment analysis, and topic recognition tasks simultaneously:

[0205] ,

[0206] Among them, Represents the comprehensive loss function, and represent the losses of the main classification, sentiment analysis, and topic recognition respectively, and are weight coefficients, and the specific implementation is as follows:

[0207] ,

[0208] Optimize the weights and biases of the model through the gradient descent algorithm, and use the Adam optimization algorithm to dynamically adjust the learning rate, which is expressed as:

[0209] ,

[0210] where, represents the parameters of the model, represents the learning rate, represents the gradient of the loss function with respect to the parameters.

[0211] After the model is deployed, it receives new complaint data in real time for classification and generates a handling strategy according to the classification results.

[0212] Example 5

[0213] This is an embodiment of the present invention, which provides a multi-modal learning automated complaint content analysis and classification method. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.

[0214] Use a multi-modal complaint data set containing text, images and audio. Each sample includes a customer complaint (text description), relevant pictures (such as product photos), and audio recordings (customer voice complaints). The data set will be labeled as including main categories, sentiment tendencies and specific topics.

[0215] Perform normalization processing on the data of all modalities, including word segmentation and denoising of text, size unification and normalization of images, sampling rate unification and background noise reduction of audio.

[0216] The method of the present invention uses an optimized deep learning model, integrating the self-attention mechanism and the multi-task learning strategy, to automatically classify and respond to complaints.

[0217] Traditional methods use single-modal data processing and a decision-making system based on fixed rules to handle complaints.

[0218] Divide the data set into a training set, a validation set and a test set. Use the training set to train the multi-modal deep learning model and the traditional text classification model of the present invention. Run the two methods on the test set and record the time taken by each method to process each complaint. Evaluate the accuracy of each method in complaint classification and user satisfaction. The experimental results are shown in Table 1.

[0219] Table 1 Comparison table of experimental results

[0220] ,

[0221] The method of the present invention adopts the comprehensive utilization of multi-modal data and the introduction of self-attention mechanism, enabling the model to understand and classify complex complaint content at a deeper level. Such improvement significantly increases the classification accuracy. When dealing with complaints containing complex emotions and diverse information, the method of the present invention can more accurately capture key information, thus making more appropriate classification judgments.

[0222] Through automated feature extraction and real-time data processing, the response time is shortened. The introduction of a dynamic decision-making engine further accelerates the decision-making process. By analyzing the context information of complaints in real time and making rapid responses, the total time from receiving a complaint to making a response is reduced. The processing efficiency is improved, and the customer satisfaction is also enhanced.

[0223] The improvement of user satisfaction proves the effectiveness of the method of the present invention in practical applications. Through accurate classification, rapid response, and personalized processing strategies designed for different complaint scenarios, efficient and high-quality customer service is ensured.

[0224] The method of the present invention first extracts rich features from text, images, audio, and context data through a multi-input feature extraction model, uses the self-attention mechanism to significantly enhance the recognition of associations between features, and then adjusts the feature weights through a dynamic fusion strategy to adapt to different real-time contexts. This process is more accurate and flexible than traditional static feature processing methods. Thanks to the highly optimized joint feature representation, the constructed multi-output deep learning model can simultaneously handle classification, sentiment analysis, and topic recognition, significantly improving the efficiency and accuracy of data processing.

[0225] In addition, the introduced dynamic decision-making engine can automatically generate processing strategies based on real-time classification results to achieve rapid response. Especially in emergency situations, it can automatically trigger necessary operations, thus greatly enhancing the timeliness and accuracy of complaint handling.

[0226] The full process automation from data collection to decision output is realized, improving the operation efficiency, reducing the labor cost, enhancing the ability to handle complex situations, enabling enterprises to more effectively manage and respond to user complaints, and improving customer satisfaction and enterprise operation efficiency.

[0227] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A multi-modal learning automated complaint content analysis and classification method, characterized in that, Including: Collecting multi-modal complaint data and preprocessing the collected multi-modal complaint data; Extracting the features of the preprocessed multi-modal complaint data and fusing the features of different modal data; Using the fused features to build a complaint content classification model, conducting model training, evaluating and optimizing the trained model; Deploying the optimized model for real-time classification; The extracting the features of the preprocessed multi-modal complaint data includes building a multi-input feature extraction model, processing text, image, audio, and context data through four branches, extracting the features of each modality, and forming a preliminary feature representation; Introducing a feature interaction layer, applying the self-attention mechanism to optimize the feature relationships, and generating a comprehensive joint feature representation through the feature interaction mechanism; Based on the joint feature representation generated by the feature interaction layer, designing a context-aware fusion strategy, dynamically adjusting the feature fusion weights, and optimizing the fusion effect; The generated joint feature representation includes performing weighted processing on the feature vectors extracted from text, image, audio, and context data; By calculating the similarity between each feature vector and other feature vectors, the self-attention mechanism automatically adjusts and optimizes the weights and associations between different features; After being processed by the self-attention mechanism, the feature interaction layer fuses and interactively learns all the weighted feature vectors, captures the potential associations between modalities, and combines and optimizes the features of each modality through the feature interaction mechanism; When the feature interaction layer finishes processing, a joint feature representation is generated; The optimized fusion strategy includes, based on the joint feature representation, generating a dynamically weighted adjusted joint feature representation according to the context information analyzed in real time; By training the optimized joint feature representation, using the patterns in the historical data and the information feedback in real time, further adjusting the feature fusion strategy, and generating a finally optimized joint feature representation; The building the complaint content classification model includes using the finally optimized joint feature representation as the input to build a multi-output deep learning model. The underlying feature extraction layer of the model first shares the input joint feature representation and conducts preliminary feature processing through a multi-layer neural network; The model is divided into three parallel task branches, and the task branches include a main classification task branch, a sentiment analysis task branch, and a topic recognition task branch; The main classification task branch uses the features further processed from the primary feature layer, conducts in-depth learning through a multi-layer fully connected neural network, and finally outputs the main category of the complaint; The sentiment analysis task branch starts from the shared primary feature layer, adds neural layers to capture emotion features. The neural layers focus on understanding the sentiment tendency and finally output the sentiment tendency of the complaint; The topic recognition task branch starts from the shared primary feature layer, adds a text analysis layer to process text data, identifies and classifies the key topics in the text, and finally outputs the specific topic of the complaint; In the training stage of the model, defining a comprehensive loss function to simultaneously consider the task effects of main classification, sentiment analysis, and topic recognition, and using a multi-task learning strategy to optimize the branches; Adjusting the weights and biases of the model through the gradient descent algorithm and using the Adam optimization algorithm for dynamic learning rate adjustment; After each training cycle, the performance of the model is evaluated using an independent validation set, and hyperparameters are adjusted based on the evaluation results; The real-time classification includes, after the model is deployed, receiving new complaint data in real time, preprocessing the data and extracting joint feature representations, and inputting them into the model for classification. The model simultaneously outputs the main category, sentiment tendency, and specific topic of the complaint according to the input features; The extracted features are input into a trained deep learning model, and the main category, sentiment tendency, and specific topic of the complaint are output synchronously; The output results are then sent to a dynamic decision engine, which constructs a decision model based on a decision tree and generates processing strategies for various scenarios; If the main category is a safety issue, an emergency response mechanism is triggered, and an automated monitoring and instant reporting process is started. The system status is monitored in real time, and an emergency notice is automatically sent to relevant departments and management to ensure instant handling; If the main category is product quality or service feedback, it enters the sentiment analysis process; For negative sentiment, emotion recognition technology is introduced to further analyze the urgency and severity of the complaint, and the response priority and resource allocation are adjusted; For positive or neutral sentiment, satisfaction prediction analysis is carried out. After predicting the customer's satisfaction with the current processing progress, it advances to the topic recognition process; According to the results of the topic recognition, for refund requests, a financial process is triggered, and the potential refund risk is predicted based on the customer's historical data and purchase patterns, and the inventory and pricing strategies are adjusted accordingly; For the topic of product suggestions, integrate customer feedback and generate a report.

2. The multimodal learning automated complaint content analysis and classification method according to claim 1, characterized in that: The multi-modal complaint data includes complaint data and context data; The complaint data includes text data, image data, and audio data; The preprocessing includes data cleaning, data standardization, format conversion, and data integration.

3. The multimodal learning automated complaint content analysis and classification method according to claim 2, wherein: The extraction of features from each modality includes using the BERT model to encode the text data, extract semantic features, and convert them into text feature vectors; Using a pre-trained ResNet model to encode the image data, extract the visual features of the image, and convert them into image feature vectors; Using MFCC technology to extract the spectral features of the audio, and using a convolutional neural network to encode the extracted audio features and convert them into audio feature vectors; The context data is converted into numerical features through an embedding layer to generate context feature vectors, which are integrated with the features of other modalities within the model.

4. A system adopting the multi-modal learning automated complaint content analysis and classification method as described in any one of claims 1 to 3, characterized in that, It includes: A data collection module, a data fusion module, a model training module, and a classification prediction module; The data collection module is used to collect multi-modal complaint data and preprocess the collected multi-modal complaint data; The data fusion module is used to extract the features of the preprocessed multi-modal complaint data and fuse the features of different modality data; The model training module is used to construct a complaint content classification model using the fused features, perform model training, evaluate, and optimize the trained model; The classification prediction module is used to deploy the optimized model for real-time classification.

5. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the multi-modal learning automated complaint content analysis and classification method according to any one of claims 1 to 3.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the multi-modal learning automated complaint content analysis and classification method described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Robot emotion recognition method and system based on AIGC and storage medium

    CN118626966A

  • Customer complaint intelligent management method and device and electronic equipment

    CN118967148A

Cited By

  • Multi-modal cultural heritage data set classification optimization method

    CN121167621A