Intelligent customer service intention layered identification method oriented to multi-modal interaction

Through topological data analysis and causal inference, high-order semantic correlation characteristics and causal relationship diagrams of multimodal data are constructed, combined with self-supervised learning to optimize feature fusion, the problem of insufficient intention recognition accuracy and generalization capabilities of intelligent customer service systems in multimodal interaction scenarios is solved, and more efficient multimodal data fusion and intention recognition are achieved.

CN120296458AActive Publication Date: 2025-07-11KEXUN JIALIAN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510776286.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The existing intelligent customer service system is difficult to effectively integrate multimodal data in multimodal interaction scenarios, and lacks modeling of causal relationships between modal features, resulting in a decrease in intention recognition accuracy and insufficient generalization ability of model, especially under the condition of few samples.

Method used

Topological data analysis, causal inference, self-supervised contrast learning and spectral clustering methods are used to construct high-order semantic correlation features and explicit causal relationship diagrams of multimodal data, combined with cross-modal self-supervised consistency regularization method, feature fusion and model training are optimized to achieve accurate hierarchical recognition of intentions.

Benefits of technology

It significantly improves the depth and accuracy of multimodal information fusion, enhances the robustness and generalization capabilities of the model, improves the accuracy of intention recognition and user service experience, and is suitable for various practical multimodal service scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296458A_ABST
    Figure CN120296458A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent customer service intention layered recognition method oriented to multi-modal interaction. The method comprises the following steps: acquiring and preprocessing three modal data of a text, a voice and an image; constructing a high-dimensional topological structure of modal data by using a topological data analysis method, and carrying out feature fusion; a causal relationship graph of modal features is constructed through a structural equation model, and causal contribution significant features are screened; a cross-modal self-supervision consistency regularization method is adopted, and an intention recognition model is trained under the condition of extremely few supervision data; and realizing multi-level layered classification output of the intention based on a spectral clustering method. The intention recognition accuracy and generalization performance of the intelligent customer service system are remarkably improved, and the method is suitable for complex multi-modal interaction scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and natural language processing, and in particular to a hierarchical recognition method for intelligent customer service intentions for multimodal interaction. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, intelligent customer service systems have been widely used in multiple service industries such as communications, e-commerce, finance, transportation, and medical care. Intelligent customer service systems use artificial intelligence and natural language processing (NLP) technology to identify and understand user intentions, achieve automated responses and services, greatly improve service efficiency, and effectively reduce labor costs.

[0003] Early intelligent customer service systems mainly relied on rule-based matching methods, which manually formulated a large number of rules to identify user intent, such as keyword matching and sentence template recognition. Although such methods have certain effects in specific scenarios, they are gradually replaced by more generalized data-driven methods due to high rule maintenance costs, limited generalization capabilities, and inability to effectively handle undefined intents.

[0004] The intent recognition technology currently widely used in intelligent customer service is mainly based on machine learning and deep learning models. Traditional machine learning methods such as support vector machines (SVM) and naive Bayes (NB) usually require manual feature design and have limited recognition capabilities for complex intents and context-sensitive issues. In recent years, deep learning methods represented by deep neural networks, such as convolutional neural networks (CNN), long short-term memory networks (LSTM), and attention mechanisms (Attention) models, have been widely used in intent recognition tasks and have achieved significant improvements. These methods automatically extract features, can better capture the semantic information in user input, and effectively improve the accuracy of intent recognition.

[0005] However, with the diversification of the ways in which users interact with intelligent customer service, the interaction scenarios have gradually expanded from single text or voice input to multimodal forms such as text, voice, image and even video. Multimodal interaction methods are more in line with real scenarios and provide richer semantic information. But at the same time, traditional intent recognition methods usually only extract and classify features for data of a single modality, lack an effective fusion mechanism for multimodal data, and it is difficult to fully utilize the complementary information between modalities. For example, existing multimodal fusion methods such as simple feature concatenation, weighted summation, average pooling, etc., although they can play a certain role in some cases, their fusion mechanism is relatively simple, and it is difficult to deeply explore the high-order intrinsic correlations between modalities, which can easily cause redundancy of modal information or loss of effective information.

[0006] In addition, existing intelligent customer service intent recognition systems usually lack a clear modeling of the causal relationships between modal features, making it difficult to accurately evaluate the contribution of each modal information to the user's true intent. This results in a significant decrease in the intent recognition accuracy and poor robustness of the model in complex or scenarios with incomplete modal information. At the same time, the training of most deep learning models requires a large amount of labeled data. However, in practical applications, especially in some emerging business scenarios, labeled data is often extremely scarce or costly. The lack of effective self-supervised or weakly-supervised mechanisms leads to insufficient generalization performance of the model and makes it difficult to handle new service scenarios or few-shot problems.

[0007] To address the above problems, in recent years, some methods based on self-supervised learning and transfer learning have been proposed to achieve rapid model transfer and generalization with a small amount of labeled data. Although these methods have improved the model training cost problem to a certain extent, due to the lack of an efficient consistency constraint mechanism between cross-modal features, the design of multi-modal data augmentation strategies and consistency loss functions is still not perfect, resulting in certain limitations in the performance improvement in a multi-modal environment.

[0008] Although the existing technologies have achieved certain development results, in actual complex multi-modal interaction scenarios, there are still the following obvious deficiencies: The multi-modal information fusion method is relatively simple, making it difficult to effectively mine the high-order semantic associations between modalities; It fails to effectively model and utilize the causal relationships between modal features; The model has poor generalization ability, especially under the condition of a small amount of labeled data, making it difficult to maintain high intent recognition accuracy and robustness.

[0009] Therefore, how to provide an intelligent customer service intent hierarchical recognition method for multi-modal interaction is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0010] An object of the present invention is to propose an intelligent customer service intent hierarchical recognition method for multi-modal interaction. The present invention combines topological data analysis (TDA), causal inference, self-supervised contrast learning, and spectral clustering methods, and effectively solves the problems of insufficient accuracy and poor generalization ability of user intent recognition in complex multi-modal interaction scenarios of the intelligent customer service system through the following technical measures: First, collect three-modal data of text, voice, and images of user interaction, and perform preprocessing operations of denoising, normalization, and feature standardization on each modal data respectively to obtain clear and standardized multi-modal data features; Second, based on the persistent homology algorithm in topological data analysis, construct simplicial complexes for the feature spaces of each processed modal data respectively, and calculate chain groups and boundary operators to obtain Betti numbers, and use the topological feature vectors composed of Betti numbers to effectively fuse the high-order semantic association features between different modalities; Next, adopt the method of structural equation model (SEM) to construct an explicit causal relationship graph between the fused multi-modal features, and calculate the causal effect coefficients of each modal feature through maximum likelihood estimation (MLE), further clarify the specific causal contribution degree of each modal feature to the user intent, and then eliminate the modal features with lower causal contribution degree to obtain an optimized multi-modal fusion feature vector; Then, based on the optimized fusion feature vector, adopt a cross-modal self-supervised consistency regularization method, and construct a cross-modal self-supervised contrast learning loss function for model training through three specific data augmentation strategies of random masking, random noise superposition, and random projection, so that the intent recognition model can be effectively trained in the case of unsupervised or extremely few supervised data, and the generalization ability of the model is improved; Finally, use the trained intelligent customer service intent recognition model above, through the Laplacian eigenmap algorithm in spectral clustering, map the feature vector to a low-dimensional space for dimensionality reduction to calculate the Euclidean distance, and perform clustering operations through distance metrics to achieve accurate multi-level hierarchical recognition and classification output of user intent.The intelligent customer service intent hierarchical recognition method for multi-modal interaction proposed by the present invention has the following distinct technical advantages: By using topological data analysis method, the high-order correlation information between modalities is effectively mined and utilized, significantly enhancing the depth and accuracy of multi-modal information fusion; By using structural equation model, the causal contribution degree of each modality data feature to the user intent is clearly evaluated, avoiding the negative impact of redundant or interfering features on the model accuracy; By using cross-modal self-supervised consistency regularization method and specific data augmentation strategies, the dependence of the model on a large amount of labeled data is effectively reduced, significantly improving the generalization performance and robustness of the model in the environment of less labeled data; By using spectral clustering method, the precise hierarchical classification output of the intent is realized, further enhancing the refinement degree and practical applicability of the intent recognition result of the intelligent customer service system. The present invention effectively solves the problems of insufficient intent recognition accuracy and generalization ability of the intelligent customer service system in complex multi-modal interaction scenarios, significantly improves the application reliability and user service experience of the intelligent customer service system, is applicable to various actual multi-modal service scenarios, and has good practical value.

[0011] An intelligent customer service intent hierarchical recognition method for multi-modal interaction according to an embodiment of the present invention is characterized by comprising the following steps: S1. Obtain user interaction data, which simultaneously includes three modalities of text, voice, and image, and perform preprocessing operations of denoising, normalization, and feature standardization on each modality data respectively; S2. For the modality data preprocessed in step S1, use the persistent homology algorithm in topological data analysis to construct the high-dimensional topological structure between each modality data, clearly quantify the high-order correlation between modalities through Betti numbers, form topological feature representations, and fuse the topological features of each modality to obtain a preliminarily fused topological feature vector; S3. Based on the topological feature vector obtained in step S2, adopt a causal inference method based on structural equation model to construct an explicit causal relationship graph between multi-modal data features, quantitatively calculate the causal contribution weights of each modality feature to the user intent recognition, eliminate the features without causal effect or negative causal effect, and only retain the modality features with significant causal contribution for secondary fusion to form an optimized multi-modal fusion feature vector; S4. Taking the optimized feature vector obtained in step S3 as the input, adopt a cross-modal self-supervised consistency regularization method based on contrast learning, generate training sample pairs through data augmentation methods of random masking, random noise superposition, and random projection, construct a cross-modal self-supervised contrast learning loss function to optimize the model parameters, and realize the effective training of the intelligent customer service intent recognition model in the case of unsupervised or extremely few supervised data; S5. Utilize the intent recognition model trained in step S4 to achieve precise multi-level hierarchical recognition and classification output of user intents in the multi-modal fusion feature space through an intent distance metric algorithm based on spectral clustering.

[0012] Optionally, the S1 includes: S1-1. Collect the original interaction data of the three modalities of text, speech, and image of the user, which are respectively represented as , and ; S1-2. Perform text denoising processing on the text modality data to obtain clean text data after removing stop words, special characters, and meaningless words; S1-3. Perform audio denoising on the speech modality data and convert it to the frequency domain using the short-time Fourier transform method to obtain the frequency domain signal , where represents the time frame, represents the frequency, and then use spectral subtraction to suppress background noise to obtain the denoised speech data ; S1-4. Perform image denoising processing on the image modality data using the Gaussian filtering method, and the filtering operation expression is: ; In the formula, is the two-dimensional Gaussian function, represents the pixel coordinates of the image, is the standard deviation of the Gaussian kernel, and the denoised image data is obtained; S1-5. Respectively perform normalization processing on the denoised data obtained in steps S1-2 to S1-4 to obtain the normalized data, which are respectively represented as , , ; S1-6. Perform feature standardization processing on the normalized data, and calculate the standardized features through the z-score standardization method: ; In the formula, represents the standardized feature, represents the normalized input feature, represents the mean of the input feature, represents the standard deviation of the input feature, and the standardized features are respectively obtained as , , .

[0013] Optionally, S2 includes: S2-1, constructing corresponding modal feature spaces based on the standardized feature data obtained in step S1; S2-2. Use the continuous homology algorithm to construct a simplicial complex for each modal feature space, and define the chain group and boundary operator as: ; In the formula, represents the k-dimensional chain group, represents the boundary operator; S2-3. Using chain groups and boundary operators, calculate the k-dimensional Betti number of each modal feature space: ; In the formula, For a closed chain group, For the border group; S2-4, construct the topological eigenvectors using the Betti numbers of each modal eigenspace; S2-5. Forming a preliminary fused topological feature vector based on topological feature vector splicing and fusion .

[0014] Optionally, the S3 includes: S3-1, using the topological feature vector obtained in step S2 as input, constructing a causal relationship diagram between features using the SEM method; S3-2. Based on the causal relationship diagram, establish the structural equation expression: ; In the formula, For Node The corresponding modal features are: Representation Node The parent node set of For Node To Node The causal effect coefficient, is subject to the mean of 0 and the variance of Normally distributed residual term; S3-3, the maximum likelihood estimation (MLE) method is used to calculate the causal effect coefficient; S3-4. Calculate the total causal contribution of modal features: ; In the formula, Representation Node The collection of child nodes; S3-5. Setting thresholds , remove Form an optimized fusion feature vector with the characteristics .

[0015] Optionally, the S4 includes: S4-1. Using the optimized multi-modal fusion feature vector obtained in step S3 as the initial input, generate training sample pairs for cross-modal self-supervised contrast learning through data augmentation methods such as random masking, random noise superposition, and random projection , where and respectively represent the feature vectors of two different views generated by random data augmentation for the th sample; S4-2. Define a cross-modal self-supervised contrast learning loss function , and its specific formula is: ; In the formula: represents the number of sample pairs within each training batch; represents the cosine similarity between the vectors , ; is the temperature hyperparameter used to control the training difficulty, and its value range is from 0.05 to 0.2; is the indicator function, which takes the value of 1 when , otherwise it takes the value of 0; S4-3. Use the cross-modal self-supervised contrast learning loss function defined in step S4-2 to optimize the model parameters and obtain the final intelligent customer service intent recognition model through the random gradient descent method.

[0016] Optionally, the S4 includes: S5-1. Using the intelligent customer service intent recognition model trained in step S4, obtain the low-dimensional embedding representation vector of the input feature vector , which is implemented by the Laplacian eigenmap algorithm in spectral clustering, and its specific calculation process is as follows: First, construct the similarity matrix between the feature vectors, and the matrix element is defined as: ; In the formula, is the width parameter of the Gaussian kernel function; Secondly, calculate the degree matrix of the matrix , the diagonal elements are defined as: ; Then calculate the normalized graph Laplacian matrix : ; Finally, solve the eigenvalue decomposition problem: ; In the formula, represents the feature vector, Indicates the eigenvalue, take the corresponding minimum The eigenvectors corresponding to the non-zero eigenvalues ​​constitute a low-dimensional embedding representation vector ; S5-2, based on the low-dimensional embedding representation vector obtained in step S5-1 , calculate the distance metric of each user intent feature, specifically using the Euclidean distance formula: ; S5-3. Based on distance measurement indicators , perform spectral clustering on user intent and divide it into intent clustering subsets at different levels , to achieve accurate multi-level hierarchical recognition and classification output of user intentions.

[0017] Optionally, the voice modality data The preprocessing includes: Firstly, the mel-frequency cepstral coefficient method is used to extract the spectral features of speech; Then, based on the convolutional neural network, the time-frequency feature representation of the speech spectrum feature is further extracted to obtain the standardized speech feature data .

[0018] Optionally, the determination method in S3-5 is specifically: Using cross-validation method, for different thresholds Calculate the recognition accuracy of the model respectively , and finally select the optimal threshold that makes the model accuracy reach the maximum value ,Right now: ; Optionally, the intelligent customer service intention recognition system of the method includes: The data acquisition and preprocessing module is used to acquire and preprocess the user's text, voice and image interaction data to obtain standardized modal feature data; A topological feature fusion module is used to perform topological data analysis on the preprocessed data and obtain a fused topological feature vector; A causal inference module for performing causal relationship analysis between multi-modal features, calculating the causal contribution degree of features, and performing optimized fusion; An intent model training module for training an intelligent customer service intent recognition model based on a cross-modal self-supervised contrastive learning method; An intent recognition and output module for achieving precise multi-level hierarchical recognition and classification output of user intents based on a spectral clustering method.

[0019] The beneficial effects of the present invention are: (1) By using the persistent homology algorithm in topological data analysis method to construct the high-dimensional topological structure between multi-modal data, and performing modal fusion with Betti numbers as topological features, the present invention effectively captures and fuses the high-order semantic associations of multi-modal data, effectively improves the accuracy and depth of multi-modal data fusion, and enhances the accuracy and stability of intent recognition in complex interaction scenarios of intelligent customer service systems.

[0020] (2) By using the structural equation model (SEM) to explicitly construct the causal relationship graph between modal features, and quantitatively calculating the causal contribution weights of each modal feature based on the maximum likelihood estimation method, the present invention can achieve effective feature selection and optimized fusion, significantly improves the pertinence and accuracy of modal feature fusion, and shows better robustness and generalization ability in actual complex multi-modal interaction environments.

[0021] (3) In the training of the intent recognition model, by using the cross-modal self-supervised consistency regularization method, specifically constructing a cross-modal self-supervised contrastive learning loss function by using data augmentation strategies such as random masking, random noise superposition, and random projection, the present invention effectively solves the strong dependence problem of the prior art on a large amount of labeled data, breaks through the bottleneck of poor model generalization ability in the prior art, realizes the effective training of the model with few or no labeled data and high generalization performance, and thus effectively improves the technical level and actual engineering application ability of the intelligent customer service system in emerging or data-scarce service scenarios.

[0022] (4) By using the Laplacian eigenmap algorithm in the spectral clustering method to map the optimized fusion feature vector to a low-dimensional space and then perform clustering, the present invention realizes the precise multi-level hierarchical recognition and classification output of user intents, significantly improves the fine-grained and refined degree of user intent recognition, and shows better service adaptability and user experience effect in actual multi-modal interaction service application scenarios. Description of the Drawings

[0023] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention.

[0024] For example, in the drawings: Figure 1 It is the overall flowchart of a method for hierarchical recognition of intelligent customer service intentions for multimodal interaction proposed by the present invention; Figure 2 It is the data preprocessing flowchart of a method for hierarchical recognition of intelligent customer service intentions for multimodal interaction proposed by the present invention; Figure 3 It is the schematic diagram of the topological data analysis and feature fusion process of a method for hierarchical recognition of intelligent customer service intentions for multimodal interaction proposed by the present invention; Figure 4 It is the schematic diagram of the causal inference and feature optimization process based on the structural equation model of a method for hierarchical recognition of intelligent customer service intentions for multimodal interaction proposed by the present invention; Figure 5 It is the schematic diagram of the cross-modal self-supervised contrast learning process of a method for hierarchical recognition of intelligent customer service intentions for multimodal interaction proposed by the present invention; Figure 6 It is the schematic diagram of the intention hierarchical recognition based on spectral clustering of a method for hierarchical recognition of intelligent customer service intentions for multimodal interaction proposed by the present invention; Figure 7 It is the schematic diagram of the structure of the intelligent customer service intention recognition system of a method for hierarchical recognition of intelligent customer service intentions for multimodal interaction proposed by the present invention. Detailed implementation manners

[0025] Regarding the multimodal data preprocessing described in step S1, it should be further clarified that: In the preprocessing process of text modal data, the stop word list needs to adopt the Chinese standard stop word list, which contains 120 common stop words; in the preprocessing of speech modal data, the window function of the short-time Fourier transform (STFT) adopts the Hamming window, the window length is 25 ms, the frame shift step length is 10 ms, and the background noise in the spectral subtraction method is estimated by the spectral average value of the first 5 frames of the signal; the two-dimensional Gaussian filtering formula used in the preprocessing of image modal data is: ; where the filter kernel size is fixed at 5×5 pixels, and the standard deviation = 1.0, zero padding is used for the filter edge, and the pixel value normalization range is limited within the interval [0,1]. When constructing the topological structure by the persistent homology algorithm in step S2, the distance threshold used for calculating the topological feature vector is fixed at 0.5, and the Euclidean distance is used to calculate the distance between features. The specific calculation formula is: ; Among them, the vector dimension n is determined according to specific modal features; the calculation of Betti numbers is limited to 0-dimensional and 1-dimensional features, and the boundary operator uses modulo 2 addition for operation. When constructing the structural equation model in step S3, the initial structural equation graph is assumed to be a fully connected directed acyclic graph, and the structural equations between nodes are expressed as: ; Among them, pa(i) represents the set of parent nodes of node i, and the causal contribution weight threshold is set to 0.05 through cross-validation, and features less than this threshold are deleted. The specific settings of the random data augmentation operation for cross-modal self-supervised consistency regularization in step S4 are as follows: the random masking ratio is fixed at 20%, the standard deviation of random Gaussian noise is fixed at 0.01, and the reduced dimension of the random projection operation is fixed at 10% of the original dimension; the cross-modal self-supervised contrast learning loss function is defined as: ; Among them, the temperature hyperparameter is fixed at 0.1. In step S5, the Gaussian kernel function is used to construct the similarity matrix in the spectral clustering method, and the specific definition is: ; Among them, the Gaussian kernel width parameter is fixed at 1.0; the number of spectral clustering categories is fixed at 5, and the standard k-means algorithm is used to implement clustering, the number of iterations is set to 200 times, and the convergence criterion threshold is fixed at .

[0026] Through topological data analysis, structural equation models, and cross-modal self-supervised consistency regularization methods, the present invention can effectively capture and fuse the high-order correlations and explicit causal relationships between multi-modal data, significantly improving the accuracy of intention recognition and the generalization performance of the model. At the same time, using spectral clustering for hierarchical intention output enhances the refinement degree and response efficiency of user intention classification, effectively overcoming the problems of insufficient fusion mechanisms and limited generalization performance in the prior art.

[0027] Supplementary description is made for the technical details of the data preprocessing described in step S1: The original interactive data of the text modality is specifically collected from the Chinese natural language text input by the user; the original interactive data of the speech modality is collected using a PCM audio device with a sampling rate of 16 kHz and a quantization bit number of 16 bits; the original interactive data of the image modality An RGB image sensor with a fixed resolution of 1920×1080 pixels is used for acquisition. The stop word list used for denoising text data in S1-2 contains a total of 120 Chinese stop words, and the special character removal set includes 20 characters in total, including punctuation marks and symbols. In step S1-3, the short-time Fourier transform (STFT) in voice data denoising uses a Hamming window, with a window length of 25 ms and a frame shift step of 10 ms. The background noise spectrum used in spectral subtraction is accurately calculated from the average spectrum of the first 5 frames of the audio signal. In S1-4, the two-dimensional Gaussian filter in image denoising uses a filter kernel with a fixed size of 5×5 pixels, and the standard deviation is set to 1.0. Zero-padding method is used for image edge processing, and the pixel values of the filtered image data are accurately normalized within the range of [0,1]. The z-score normalization formula used for feature standardization in step S1-6 is specifically expressed as: ; where is the mean of all feature data, is the standard deviation of the feature data, and the feature mean and standard deviation are accurately calculated to 4 decimal places.

[0028] By strictly and meticulously performing preprocessing of data denoising, normalization, and feature standardization on text, voice, and image modal data, the present invention enables different modal data to be input into the subsequent analysis process with a consistent scale and high quality, effectively improving the reliability and accuracy of data fusion, significantly enhancing the stability and robustness of intent recognition in the multi-modal intelligent customer service system, and overcoming the problem of insufficient fusion accuracy caused by inconsistent data quality in the prior art.

[0029] Regarding the details of data analysis in step S2, in specific implementation, the dimensions of each modal feature space should be clearly defined respectively: The dimension of the text modal feature space is 300 dimensions, the dimension of the voice modal feature space is 13 dimensions, and the dimension of the image modal feature space is 1024 dimensions. The method for constructing the simplicial complex used in the persistent homology algorithm is the Vietoris-Rips method, in which the distance between data points is strictly calculated using the Euclidean distance, and the calculation formula is: ; where , are any two data points within the modal feature space, and the dimension is the dimension of the above-mentioned each modal feature space. The threshold distance for constructing the simplicial complex in the Vietoris-Rips method is set to 0.5. The chain group and the boundary operator The calculation is implemented strictly according to the modulo-2 addition rule. The calculation of Betti numbers is strictly limited to 0-dimensional and 1-dimensional Betti numbers. Among them, the 0-dimensional Betti number represents the number of connected components, and the 1-dimensional Betti number represents the number of loop structures. The Betti number feature vectors obtained from each modality are specifically represented as: ; That is, it is formed by sequentially splicing the 0-dimensional and 1-dimensional Betti numbers of each modality in the above order.

[0030] The present invention uses the persistent homology algorithm to quantitatively analyze the topological structure of the feature space of multi-modal data, so as to accurately construct high-order topological feature vectors, effectively reveal the deep correlation relationships of each modality data at the topological level, significantly improve the accuracy of subsequent feature fusion and intelligent customer service intention recognition, and overcome the problems of redundant feature information and missing important correlation information caused by existing simple splicing or weighted fusion methods.

[0031] Regarding the construction process of the causal relationship diagram of the structural equation model (SEM) in step S3, the following technical details should be clearly supplemented: When constructing the initial causal relationship diagram, a fully connected directed acyclic graph is formed between all topological feature nodes, and the number of nodes is consistent with the dimension of the topological feature vector obtained in step S2. In the present invention, the causal effect coefficient between each node is solved by using the maximum likelihood estimation (MLE) method. The specific logarithmic likelihood function for maximizing the observed data is expressed as: ; where , respectively represent the eigenvalue of node and node , in the th observation sample, represents the total number of observation samples, and the variance of the residual term of each node is initially set to 1.0 during the model solving process, and the specific value of the residual variance of each node is calculated through the MLE iteration method. When calculating the total causal contribution degree of the modality features, the sub-node set of the node is determined according to the final causal relationship diagram structure, and the causal contribution threshold is accurately determined to be 0.05 through cross-validation. All features with a causal contribution degree lower than 0.05 are deleted from the fusion feature vector, and the remaining features form the optimized fusion feature vector

[0032] The present invention precisely constructs the clear causal relationships among various topological features through a structural equation model, strictly adopts the maximum likelihood estimation method to calculate the causal effect coefficients between nodes, quantitatively evaluates the actual contributions of various modal features to the user intention recognition task, accurately screens out effective features, effectively avoids the interference of redundant and negative impact features, significantly enhances the generalization ability and robustness of the model, and effectively overcomes the problem of poor intention recognition accuracy caused by inaccurate causal relationship analysis in the prior art.

[0033] Regarding the cross-modal self-supervised contrastive learning training details described in step S4, the specific implementation process is specified as follows: The data augmentation method of random masking specifically adopts randomly masking 20% of the elements in each sample feature vector, and the masking positions are randomly selected uniformly; the method of adding random noise is to add Gaussian noise with an independent and identically distributed mean of 0 and a fixed standard deviation of 0.01 to each element of the feature vector; the data augmentation method of random projection is specifically to reduce the original feature dimension to 90% of the original dimension by using a random Gaussian projection matrix, where the elements of the random projection matrix are independently drawn from the standard normal distribution The cross-modal self-supervised contrastive learning loss function is specifically defined as: ; where the number of batch samples is specifically fixed at 64, the temperature hyperparameter is fixed at 0.1, the model parameter optimization algorithm clearly adopts the stochastic gradient descent method, where the initial learning rate is fixed at 0.001, the number of iterations in the optimization process is specifically set to 1000 times, and the gradient convergence threshold is fixed at 10^{-6}.

[0034] The present invention effectively trains an intention recognition model for multi-modal fusion features under the condition of unsupervised or extremely little supervised data through a cross-modal self-supervised consistency regularization method combined with a specific data augmentation strategy, significantly improves the generalization performance and recognition accuracy of the model, breaks through the problem that the generalization ability of the prior art is limited by the dependence on a large amount of labeled data, and has stronger practical application ability.

[0035] Regarding the specific implementation details of the spectral clustering method in step S5, it is further specified as follows: In spectral clustering, the similarity matrix is constructed using a Gaussian kernel function, and the specific value of the kernel width parameter is clearly fixed at 1.0, and the dimension of the input optimized fusion feature vector is specifically 128-dimensional. The degree matrix is a diagonal matrix, and its diagonal elements are specifically obtained by summing the elements of each row of the similarity matrix , that is: ; Normalized Laplacian matrix The specific calculation is as follows: ; Solve the eigenvalue decomposition problem: ; Specifically, the eigenvalue decomposition method is adopted. After the eigenvalues are sorted from small to large, the eigenvectors corresponding to the first 5 non-zero eigenvalues are taken as the low-dimensional embedding representation vectors . The distance index of the intention eigenvector is clearly calculated using the Euclidean distance formula: ; For the specific clustering operation of spectral clustering, the standard k-means algorithm is adopted. Among them, the number of clustering categories is specifically fixed at 5, the maximum number of iterations is clearly set at 200 times, and the convergence threshold is fixed at .

[0036] The present invention effectively realizes the multi-level hierarchical accurate recognition of user intentions through the spectral clustering method. By using the similarity matrix constructed by the Gaussian kernel function and Laplacian eigenmap, the high-dimensional fusion features are mapped to the low-dimensional space to improve the stability and accuracy of clustering, significantly enhancing the refinement degree of the intention recognition results of the intelligent customer service system, significantly improving the actual applicability and user experience effect of the system, and effectively overcoming the deficiency that the existing methods cannot achieve accurate and efficient hierarchical recognition.

[0037] The specific implementation details for the preprocessing of voice modality data are further clarified as follows: In the specific calculation process of Mel Frequency Cepstral Coefficients (MFCC), 26 Mel-scale filters are adopted. The time-frequency domain representation of the audio signal is converted through the Short-Time Fourier Transform (STFT) and then passed through the Mel filter bank. Subsequently, the logarithmic energy value is taken, and the first 13 cepstral coefficients are extracted using the Discrete Cosine Transform (DCT) as the voice spectrum features. The subsequent Convolutional Neural Network (CNN) structure is specifically a network containing two one-dimensional convolutional layers. The first convolutional layer contains 32 convolutional kernels, each with a size of 3 and a stride of 1; the second convolutional layer contains 64 convolutional kernels, each with a size of 5 and the same stride of 1. After each convolutional layer, the ReLU activation function is adopted, and a max-pooling operation with a kernel size of 2 is performed. Finally, through a fully connected layer, the features after convolutional pooling are mapped into a standardized voice feature vector with a dimension of 128, forming the data .

[0038] ​The present invention finely extracts the spectral and time-frequency feature representations of speech modality data through the method of combining Mel-frequency cepstral coefficients with a convolutional neural network, fully captures the key features of the speech signal in the frequency and time dimensions, effectively improves the accuracy and stability of the speech feature representation, significantly enhances the support ability of the speech modality data for subsequent multi-modal fusion and intention recognition tasks, and overcomes the problems of single speech feature extraction method and limited generalization ability in the prior art.

[0039] Regarding the feature selection threshold in step S3-5 The specific implementation details of the determination method are further clarified as follows: The ten-fold cross-validation method is adopted for the threshold selection process, where the value range of the threshold is strictly set from 0.01 to 0.1, with an interval step size of 0.01. The recognition accuracy of the model is calculated on the cross-validation data set for each threshold . The data division of each cross-validation is precisely 90% for the training set and 10% for the validation set. The feature selection threshold is finally clearly determined as the specific value corresponding to the maximum average accuracy of the ten-fold cross-validation: ; The present invention determines the feature selection threshold in the structural equation model through an accurate cross-validation method, can objectively and accurately evaluate the influence of model feature selection on the recognition performance under different thresholds, significantly optimizes the scientificity and effectiveness of feature selection, avoids the errors and performance fluctuations caused by artificially setting the threshold, effectively improves the generalization ability of the model and the accuracy of intention recognition, and overcomes the problem of lack of accurate criteria in the prior art feature selection method.

[0040] Regarding the intelligent customer service intention recognition system, the specific implementation details of each module are further clarified as follows: In the data acquisition and preprocessing module, for text data, 120 stop words from the standard Chinese stop word list are used for denoising; for voice data acquisition, a PCM audio device with a sampling rate of 16 kHz and a quantization bit depth of 16 bits is adopted. For preprocessing, the short-time Fourier transform (STFT) has a window length of 25 ms and a frame step size of 10 ms, and the spectral subtraction method uses the average spectrum of the first 5 frames of the audio as the background noise estimation standard; for image data acquisition, an RGB image sensor with 1920×1080 pixels is used, and for preprocessing, a 5×5 pixel two-dimensional Gaussian filter kernel with a standard deviation of 1.0 is used and normalized to the interval [0,1]; for all data feature standardization, the z-score method is uniformly adopted. The topological feature fusion module strictly constructs a simplicial complex based on the Vietoris-Rips method, the threshold distance is set to 0.5, and the calculation of feature vectors is limited to 0-dimensional and 1-dimensional Betti numbers. The causal inference module uses a structural equation model (SEM), initially set as a fully connected directed acyclic graph (DAG), and the causal effect coefficient is strictly calculated by the maximum likelihood estimation (MLE) method, and the causal contribution threshold is strictly determined to be 0.05 through cross-validation. The intent model training module adopts cross-modal self-supervised contrastive learning, where the random masking ratio is 20%, the standard deviation of random Gaussian noise is 0.01, the random projection dimension is reduced to 10% of the original dimension, the temperature parameter in the contrastive learning loss function is fixed at 0.1, and the number of samples in each training batch is 64. The intent recognition and output module is based on the spectral clustering method, constructs a similarity matrix using a Gaussian kernel function with a width parameter of 1.0, the number of clustering categories is fixed at 5, and the k-means algorithm is iterated 200 times, and the convergence criterion threshold is set to 10^{-4}.

[0041] The intelligent customer service intent recognition system proposed by the present invention realizes efficient and accurate multi-modal data fusion and intent recognition through efficient data preprocessing, topological data analysis, clear causal relationship modeling, and cross-modal self-supervised learning strategies, significantly improves the generalization ability and intent recognition fineness of the system in a complex multi-modal interaction environment, overcomes the problems of simple fusion methods, blind feature selection, and lack of robustness in the prior art, and has significant technical advantages and practical application value.

[0042] Example 1:

[0043] An online intelligent consultation platform of a certain hospital has introduced an intelligent customer service intention hierarchical recognition method for multi-modal interaction of the present invention to improve the accuracy of understanding patients' consultation intentions. In this scenario, patients often interact with the system using voice, text, and images simultaneously. For example, a patient may describe symptoms by voice, while entering text in the dialog box to supplement the condition information, or upload relevant image materials such as photos of the affected area. The multi-modal input provides rich information for the system, but also brings challenges in understanding the true intentions of users. Traditional single-modal intention recognition models are prone to misjudgment when facing such complex inputs, and cannot capture the dialogue context in a timely manner, resulting in delayed answers or irrelevant responses. For example, a patient may vaguely state "always feeling uncomfortable recently" without specifying the specific symptoms; or a patient describes a rash situation but does not upload a photo. In these cases, ordinary intelligent consultation systems may have difficulty understanding the questions that patients want to consult in a timely and accurate manner.

[0044] This embodiment optimizes the above problems through the method of the present invention. In terms of system architecture, the intelligent customer service first uses speech recognition technology to transcribe the patient's voice description into text, and extracts the emotions and key information in the voice (such as intonation can reflect the degree of pain). At the same time, for the images uploaded by the patient (such as photos of wounds or rashes), the system uses an image recognition model to analyze and extract the feature information related to medical intentions (such as the affected area shown in the image, the severity, or possible disease characteristics). Then, the platform fuses the multi-modal data to form a unified semantic representation. Based on the intention hierarchical recognition model of the present invention, the system analyzes the user's intention at two levels: the first level first identifies the high-level intention categories consulted by the patient (such as "symptom consultation", "appointment registration", "report interpretation", etc.), and the second level then determines the detailed intention in combination with specific semantic details (for example, after identifying "symptom consultation", further judge whether it is a consultation about "skin itching" or "cough and fever" and other specific problems). This hierarchical recognition process makes full use of the information contained in voice, text, and images, and can grasp patients' needs more comprehensively than relying solely on single-text analysis. It is worth mentioning that when the patient provides an image, the system can corroborate the image content with the text description to improve the confidence in intention judgment. For example, when the voice and text description mentions a rash, and the image analysis result also points to skin redness, the method of the present invention will comprehensively determine it as the intention of "consulting skin symptoms", greatly improving the recognition accuracy.

[0045] In the case of vague expressions or incomplete information from patients, the multi-modal intention hierarchical recognition in this embodiment demonstrates strong robustness. For vague inquiries, the system first classifies the question into a general category through the first-layer intention recognition. Then, if the confidence level is insufficient, it will initiate a clarification question in the conversation to guide the patient to provide more information. For example, when the patient only says "I'm a little uncomfortable in the stomach", the system initially classifies it as the intention of "symptom consultation - digestive category". However, due to the general description, the system then asks: "Do you feel stomachache or diarrhea? How long has the symptom lasted?" This strategy of multi-round interaction is effectively integrated into the method of the present invention and is part of the intention recognition process. By gradually refining the user's intention, the system can handle the patient's initial vague description and reduce the occurrence of misunderstandings. At the same time, if information in a certain modality is missing, this method can also make intelligent compensation: when the user does not provide an image, the system does not stop the recognition process but relies more on the content of the language text and context reasoning to judge the intention; on the contrary, if the voice input is unclear due to environmental noise, the system focuses on referring to the text information subsequently input by the user. In this way, even if there is a lack of a certain modality or incomplete information, the intention hierarchical recognition model can still make up for it through other modality information sources, thus maintaining a high accuracy in understanding the user's intention.

[0046] The method of the present invention also greatly improves the intention understanding ability in multi-round consultation dialogues. In the past, online consultation systems often had problems with poor context tracking in long conversations, resulting in the failure to accurately capture the true intentions of patients in multi-round consultations. In this embodiment, the intelligent customer service dynamically adjusts the understanding of the user's intention in each round of reply by using the intention hierarchical model. For example, the patient first asks "Do I need to go to the hospital for this rash?", and the system recognizes that the intention is "symptom consultation - medical advice". When the patient then asks "Then which department should I register for?", the system can realize that the user's concern has changed from "whether to seek medical treatment" to "department guidance for seeing a doctor". Due to the previous context, the system identifies the new question as the intention of "medical guidance" in the second round and associates it with the rash scenario discussed in the previous round, so as to give a coherent and correct answer. Throughout the process, the model of the present invention realizes the accurate understanding of multi-round questions by maintaining the semantic context and hierarchical intention labels of the conversation, avoiding context loss or misjudgment.

[0047] After deploying the multi-modal intention hierarchical recognition method of the present invention, the performance indicators of the hospital's intelligent consultation platform have been significantly improved. Table 1, the improved data of medical consultation intention recognition, presents the comparison of the main indicators before and after deployment. It can be seen that the intention recognition accuracy rate has increased from 85% before deployment to 95% after deployment, showing a substantial improvement. This benefits from the more comprehensive understanding brought by multi-modal information fusion and the effective reduction of errors caused by ambiguity through the hierarchical recognition strategy. At the same time, the average response time has been shortened from 5.2 seconds to 3.1 seconds, which means that the system can give targeted responses faster. On the one hand, because this method improves the accuracy of the initial understanding and reduces the number of rounds of repeated clarification, on the other hand, multi-modal parallel processing optimizes the response efficiency. The misinterpretation rate has dropped from the original 12% to 2.5%, indicating a significant reduction in the situation of misinterpreting the user's intention. The reduction of the misinterpretation rate directly improves the user's consultation experience and reduces the frustration caused by answering off-topic. From the above data, it can be seen that the application of the method of the present invention has brought obvious beneficial effects to the medical consultation scenario: it not only improves the accurate grasp of the patient's intention by the intelligent customer service, but also speeds up the response speed, reflecting high engineering practical value.

[0048] Next, Table 1 summarizes the specific numerical changes of each indicator in this embodiment: Table 1 Improved data of medical consultation intention recognition

[0049] It can be intuitively seen from Table 1 the performance improvement brought by the method of the present invention. The substantial increase in the intention recognition accuracy rate indicates that the system more reliably understands the patient's consultation intention, which is directly related to the reduction of understanding deviation and misinterpretation by multi-modal information fusion. The shortening of the response time reflects the improvement of system efficiency - patients receive useful replies faster, making the consultation process smoother. The significant reduction in the misinterpretation rate is particularly crucial, which means that the proportion of incorrect answers or failure to understand the patient's meaning has decreased significantly, reducing the patient's need for repeated explanations and the resulting potential dissatisfaction. Overall, this embodiment proves that the intention hierarchical recognition method for multi-modal interaction can effectively improve the accuracy and reliability of the dialogue system in the field of medical consultation, providing more accurate and efficient intelligent consultation services for patients.

[0050] Example 2:

[0051] In this embodiment, the multimodal interaction intention hierarchical recognition method of the present invention is deployed in the intelligent customer service system of a large e-commerce platform to optimize the understanding and response of user consultations. User consultations on e-commerce platforms usually have multimodal interaction characteristics: users may describe the problems encountered by voice (for example, voice consultation of product information or order status), while entering text in the dialog box (such as providing order numbers or specific demands), and uploading relevant pictures (such as invoices, photos of product defects, etc.) to assist in explaining the situation. Faced with such rich input information, if each modality is processed independently, traditional customer service systems often find it difficult to grasp the core demands of users in a timely and accurate manner, and may even make misjudgments due to inconsistent multi-source information.

[0052] After the method of the present invention is introduced, the system first recognizes and transcribes the user's voice content, and analyzes the emotions and key points in the voice. For example, if the user's tone is anxious, it may indicate that the problem is urgent. The text input part directly extracts key fields, such as order number, product name, etc. For images uploaded by users, the system uses image recognition and OCR technology to extract useful information: if the image is an invoice or an order screenshot, the transaction data such as order number and date are extracted; if it is a product photo, the product type or defective part is identified. Subsequently, the intelligent customer service integrates the information of the three modes of voice, text, and image to represent the user's current intention state in a unified semantic space. Based on the intention hierarchical model, the system also performs two-level intention recognition: the first layer determines the main category of the user's help, such as "product consultation", "logistics inquiry", "return and exchange" or "account / payment problem"; the second layer further subdivides the specific intention after locking the major category. For example, when the first-layer recognition result is "return and exchange", the system will combine the context to determine whether the user's specific intention is "apply for a return", "check the progress of the return", or "understand the return policy", etc., and adopt different response strategies accordingly. Compared with the previous practice of simply mapping a sentence to a fixed intention, this hierarchical recognition makes full use of the complementary advantages of different modal information, making the judgment more accurate and more context-adaptive. Especially when the user provides multiple pieces of evidence, this method can cross-verify the information: for example, the user says "I want to return the goods" and the text also provides the corresponding order number. The system will associate the intention in the voice with the order details provided by the text, accurately identify the user's intention as "initiate an order return", and locate the goods and orders involved. The fusion of multimodal information here reduces the errors that may occur in a single modality and significantly improves the recognition accuracy and reliability.

[0053] This embodiment focuses on demonstrating the system's ability to handle modal redundancy and conflicts. When there is redundancy in the user's multiple input modal contents, the system regards it as a favorable factor for information verification. For example, when the user describes a problem through voice and repeats some key information (such as the order number) in the text, the system will compare the consistency of the two inputs. If there are slight differences between the voice recognition result and the text content, the system can discover and correct these differences through analysis, thus avoiding misunderstandings caused by voice recognition errors. When the modal information is inconsistent or even in conflict, the method of the present invention can make a strategic fusion judgment: first determine the user's main intention, and at the same time not ignore the meaning contained in the secondary information. For example, a user says "I want to return the goods because there is a problem with this mobile phone" by voice, but the uploaded picture is a purchase invoice. Apparently, the voice intention points to returning the goods, while the picture content is a transaction voucher, and the two do not match in literal information. A traditional system may be confused about whether the user wants to handle the return or query the invoice, while the intention hierarchical recognition of this method will first clarify the main intention as "return request" based on the voice and text, and then regard the invoice in the picture as an auxiliary means to extract order and product information to assist in completing the return process, rather than misclassifying it as another "invoice problem" intention. Thus, it can be seen that even if the information provided by each modality does not match on the surface, the system can still correctly interpret the user's true needs through intention hierarchy and multi-modal fusion. If different modalities convey multiple independent requests, this method can also identify the implicit composite intention and distinguish and process it hierarchically. For example, when the user simultaneously makes two requests of "inquiring about product performance" and "requesting an invoice", the system will judge at the first level that the user has multiple intentions, which are respectively classified into two major categories of "product consultation" and "after-sales service", and then give answers separately in the dialogue, thus realizing the ability to answer multiple questions with one question.

[0054] Through the above mechanism, the e-commerce customer service system also demonstrates higher intelligence and flexibility during multi-round interactions. Users' questions are often continuous and related. For example, they first ask "Does this product have a warranty?" and then immediately ask "Then how do I apply for a return?" after getting an answer. The method of the present invention enables the system to track such intention changes: the first question is recognized as "Product Inquiry - Warranty Policy", and the second is recognized as "Returns and Exchanges - Return Process". The system realizes that these are the consecutive requirements of the same user in the same conversation, automatically switches the context from product information to return transactions, and at the same time retains the previous product-related information to answer related details. This ability to track intentions in multi-round conversations stems from the dynamic adjustment of the context by the intention hierarchical model, ensuring that each round of questions can be accurately mapped to the corresponding intention category without losing direction due to topic changes. Traditional customer service robots may require users to repeat information or directly transfer the conversation to human handling in such scenarios. However, after applying the method of the present invention, the system can independently handle more complex multi-round conversations, continuously understand different requests from users and make correct responses.

[0055] After actual deployment and testing on the e-commerce platform, the method of the present invention has significantly improved various performance indicators of the customer service system. Table 2 below, the improved data of customer service intention recognition performance on the e-commerce platform, lists the comparison of the main indicators before and after the improvement. It can be seen that the accuracy of hierarchical intention recognition has increased from the original 88% to 96%, indicating that the system's judgment of user intentions is more accurate. This improvement in accuracy directly reduces the situation of answering irrelevant questions and also reduces the frequency of manual intervention for correction. The user satisfaction score has also increased significantly (for example, from 80% to 94%). Users generally feedback that the new intelligent customer service can "understand" their requests better and the responses are more in line with their needs. At the same time, the average response time has been shortened from 4.5 seconds to 2.7 seconds, indicating that the time users wait for an answer after asking questions has been significantly reduced. The faster response stems from the higher recognition efficiency and process optimization of this method - the system can instantly extract key information from multi-modal inputs and quickly retrieve answers, reducing intermediate processing steps. The response speedup not only makes users feel that the service is more efficient but also further improves their satisfaction. Finally, the user complaint rate has dropped from 7.2% before deployment to 1.5%, a significant decrease. This means that user dissatisfaction and complaints caused by misunderstandings or inappropriate responses have been greatly reduced, fully demonstrating the reliability and superiority of the method of the present invention in practical applications. Generally speaking, through the data in Table 2, it can be clearly seen that the method of hierarchical intention recognition for multi-modal interaction brings higher accuracy, faster response speed and better user evaluation to the e-commerce customer service system, showing extremely high engineering practical value and commercial value.

[0056] Table 2 Improved data of customer service intention recognition performance on the e-commerce platform

[0057] It can be directly seen from Table 2 that the method of the present invention has significantly improved the effect in the e-commerce customer service scenario. The improvement in the intent recognition accuracy means that the intelligent customer service can almost accurately understand the requirements of customers without deviation, and can make correct judgments even in the face of complex information with coexisting voice, text, and images. This ensures that the customer's questions are accurately answered and reduces the repeated communication caused by misunderstandings. The substantial increase in user satisfaction reflects the recognition and satisfaction of customers with the new customer service system. This is because the system provides more accurate answers and efficient services, making customers feel smooth communication and their problems are truly solved. The reduction in the average response time further improves the user experience - a faster response makes customers feel that the service is efficient and they are valued. The decrease in the complaint rate indicates that almost all customer complaints caused by the misunderstanding or ineffective answers of the robot in the past have been eliminated, which reflects the reliability of the method of the present invention in the actual business environment. In summary, Example 2 proves that the multi-modal interaction intent hierarchical recognition method can effectively integrate multi-source information such as voice, text, and images in the e-commerce customer service field, accurately understand customer intent and respond in a timely manner, not only comprehensively improving the service quality indicators, but also enhancing the trust and satisfaction of users with the platform customer service.

Claims

1. An intelligent customer service intention hierarchical recognition method for multi-modal interaction, characterized in that It includes the following steps: S1. Obtain user interaction data, which simultaneously contains three modalities of text, speech, and images, and perform preprocessing operations of denoising, normalization, and feature standardization on each modality data respectively; S2. For the modality data preprocessed in step S1, use the persistent homology algorithm in topological data analysis to construct a high-dimensional topological structure among the modality data, clearly quantify the higher-order correlation between modalities through Betti numbers, form a topological feature representation, and fuse the topological features of each modality to obtain a preliminarily fused topological feature vector; S3. Based on the topological feature vector obtained in step S2, adopt a causal inference method based on the structural equation model to construct an explicit causal relationship graph among multi-modal data features, quantitatively calculate the causal contribution weights of each modality feature to user intention recognition, eliminate features with no causal effect or negative causal effect, and only retain modality features with significant causal contributions for secondary fusion to form an optimized multi-modal fusion feature vector; S4. Using the optimized feature vector obtained in step S3 as input, adopt a cross-modal self-supervised consistency regularization method based on contrast learning, generate training sample pairs through data augmentation methods such as random masking, random noise superposition, and random projection, construct a cross-modal self-supervised contrast learning loss function to optimize model parameters, and achieve effective training of the intelligent customer service intention recognition model in the case of unsupervised or extremely few supervised data; S5. Use the intention recognition model trained in step S4 to achieve precise multi-level hierarchical recognition and classification output of user intentions in the multi-modal fusion feature space through an intention distance metric algorithm based on spectral clustering.

2. The method according to claim 1, wherein Step S1 specifically includes the following sub-steps: S1-1: Collect the original interaction data of three modalities, namely text, speech, and image, of the user, and represent them as , and ; S1-2. For the text modal data perform text denoising processing to obtain clean text data after removing stop words, special characters, and meaningless words ; S1-3. For the voice modality data Perform audio denoising, convert it to the frequency domain using the short-time Fourier transform method, and obtain the frequency domain signal , where represents the time frame represents the frequency. Subsequently, use spectral subtraction to suppress background noise and obtain the denoised voice data ; S1-4. For the image modality data Perform image denoising processing using the Gaussian filtering method. The filtering operation expression is: ; In the formula, is a two-dimensional Gaussian function, represents the pixel coordinates of the image, is the standard deviation of the Gaussian kernel, and the denoised image data is obtained; S1-5. Normalize the denoised data obtained in steps S1-2 to S1-4 respectively to obtain the normalized data, which are respectively represented as , , ; S1-6. Perform feature standardization processing on the normalized data, and calculate the standardized features through the z-score standardization method: ; In the formula, represents the standardized feature, represents the normalized input feature, represents the mean of the input feature, represents the standard deviation of the input feature, and the standardized features are respectively obtained as , , .

3. The method according to claim 1, wherein Step S2 specifically includes the following sub-steps: S2-1. Respectively construct corresponding modality feature spaces based on the standardized feature data obtained in step S1; S2-2. Use the persistent homology algorithm to perform simplicial complex construction on each modality feature space respectively, and define the chain group and boundary operator as: ; In the formula, represents the k-dimensional chain group, represents the boundary operator; S2-3. Use the chain group and boundary operator to calculate the k-dimensional Betti number of each modality feature space; ; wherein, is a cycle group, is a boundary group; S2-4. Construct a topological feature vector with the Betti numbers of each modality feature space; S2-5. Form a preliminarily fused topological feature vector by splicing and fusing topological feature vectors .

4. The method according to claim 1, characterized in that, Step S3 specifically includes the following sub-steps: S3-1. Using the topological feature vector obtained in step S2 as input, use the SEM method to construct a causal relationship graph among features; S3-2. Based on the causal relationship graph, establish a structural equation expression; ; In the formula, is the modal feature corresponding to the node , represents the set of parent nodes of node , is the causal effect coefficient from node to node , is the residual term that follows a normal distribution with a mean of 0 and a variance of ; S3-3. Adopt the maximum likelihood estimation (MLE) method to calculate the causal effect coefficient; S3-4. Calculate the total causal contribution degree of modality features; ; In the formula, represents the set of child nodes of the node; S3-5. Set the threshold , eliminate features below to form an optimized fusion feature vector .

5. The method according to claim 1, characterized in that Step S4 specifically includes the following sub-steps: S4-1. Using the optimized multi-modal fusion feature vector obtained in step S3 as the initial input, training sample pairs for cross-modal self-supervised contrastive learning are generated through data augmentation methods such as random masking, random noise addition, and random projection , where and respectively represent the feature vectors of two different views generated through random data augmentation for the th sample; S4-2. Define the cross-modal self-supervised contrastive learning loss function , and its specific formula is as follows: ; In the formula: Indicates the number of sample pairs within each training batch; Indicates the cosine similarity and ; is a temperature hyperparameter used to control the training difficulty, with a value range of 0.05 to 0.2; is an indicator function that takes the value 1 when and 0 otherwise; S4-3. Use the cross-modal self-supervised contrastive learning loss function defined in step S4-2 Optimize the model parameters and obtain the final intelligent customer service intention recognition model through training by the stochastic gradient descent method.

6. The method according to claim 1, wherein Step S5 specifically includes the following sub-steps: S5-1. Obtain the input feature vector using the intelligent customer service intent recognition model trained in step S4 to obtain the low-dimensional embedded representation vector , which is implemented using the Laplacian eigenmap algorithm in spectral clustering. The specific calculation process is as follows: First, construct the similarity matrix between feature vectors , where the matrix element is defined as: ; In the formula, is the width parameter of the Gaussian kernel function; Secondly, calculate the matrix 's degree matrix , and the diagonal elements are defined as: ; Then calculate the normalized graph Laplacian matrix : ; Finally, solve the eigenvalue decomposition problem: ; In the formula, represents the eigenvector, represents the eigenvalue, and the eigenvectors corresponding to the first non-zero eigenvalues in ascending order are taken to form the low-dimensional embedded representation vector ; S5-2. Low-dimensional embedded representation vectors obtained based on step S5-1 , calculate the distance metric index for each user intention feature, specifically using the Euclidean distance formula: ; S5-3. According to the distance metric , perform spectral clustering operation on the user intention and divide it into intention clustering subsets at different levels , and achieve accurate multi-level hierarchical recognition and classification output of the user intention.

7. The method according to claim 2, wherein The voice modality data The preprocessing specifically includes: First, use the Mel Frequency Cepstral Coefficient method to extract the spectral features of speech; Then, when further extracting the time-frequency feature representation from the speech spectral features based on the convolutional neural network, standardized speech feature data is obtained .

8. The method according to claim 4, characterized in that The determination method of the feature selection threshold in step S3-5 is specifically as follows: Using the cross-validation method, for different thresholds calculate the recognition accuracy of the model respectively , and finally select the optimal threshold corresponding to the maximum accuracy of the model , that is: 。 9. An intelligent customer service intention recognition system for implementing the method described in claim 1, characterized in that, It includes: A data acquisition and preprocessing module, which is used to obtain and preprocess the text, speech, and image interaction data of users to obtain standardized modality feature data; A topological feature fusion module for performing topological data analysis on the preprocessed data and obtaining a fused topological feature vector; A causal inference module for performing causal relationship analysis between multimodal features, calculating the causal contribution degree of features, and performing optimized fusion; An intent model training module for training an intelligent customer service intent recognition model based on a cross-modal self-supervised contrast learning method; An intent recognition and output module for achieving accurate multi-level hierarchical recognition and classification output of user intents based on a spectral clustering method.

Citation Information

Patent Citations

  • Multi-modal intention recognition method and system based on comparative learning

    CN116304984A

  • Human-computer interaction method based on intelligent data analysis

    CN118349147A

  • Alzheimer disease classification method and system based on multi-view spatio-temporal topology fusion

    CN118899076A

  • Multi-modal question answering method and device in customer service application scene

    CN119204220A

  • Transformer life prediction method and system based on machine learning and Internet of Things

    CN119441993A

Cited By

  • Multi-modal intention recognition method and device, computer equipment and storage medium

    CN121074897A

  • Multimodal intent recognition method and apparatus, computer device, and storage medium

    CN121074897B