Customer intention analysis method, device and storage medium based on call content
By using multi-model speech recognition and deep learning natural language processing technology in customer intention analysis, the problems of accuracy and efficiency in the existing technology are solved, and more accurate recognition and more efficient analysis of customer intentions are achieved.
Patent Information
- Application Number
- CN202510363374.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The prior art has problems of accuracy and inefficiency in customer intention analysis, especially when dealing with complex language expressions and diverse call contexts.
By extracting the original audio data from the call recording system, noise suppression processing is performed, and speech recognition is performed using multiple speech recognition models. Then, natural language processing technology based on deep learning is used to perform semantic analysis and core semantic integration, and multiple core semantic points of speech recognition results are fused to achieve a deep understanding of audio content.
It improves the accuracy and efficiency of customer intention analysis, can more accurately identify customer intention types, and supports the company's customer service and marketing decisions.
Smart Images

Figure CN119920238B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of call data analysis, and more specifically, to a method, device and storage medium for analyzing customer intention based on call content. Background Art
[0002] Efficient and accurate understanding of customer intentions is crucial for a company's sales, customer service, and market strategy formulation. Among many communication methods, telephone calls have become one of the important channels for companies to interact with customers due to their immediacy and efficiency. However, after a large amount of call data is generated, how to accurately analyze customer intentions has become a major challenge for companies. Traditional customer intention analysis methods often rely on manual screening and analysis of call records. This method not only consumes a lot of manpower and time costs, but the analysis results are easily affected by subjective factors, and the accuracy is difficult to guarantee.
[0003] With the continuous advancement of information technology, some automated data analysis methods have gradually been applied to customer intention analysis based on call content. For example, based on speech recognition and natural language processing technology, the call content is converted into text, and the customer's intention is determined by keyword recognition and analysis of the text. However, this method has obvious limitations. It can only recognize pre-set keywords, and lacks the ability to understand complex language expressions and semantics, and cannot accurately grasp the customer's true intention. In addition, the recognition accuracy and adaptability of a single speech recognition model are often limited when facing different call backgrounds and diverse accents, resulting in low accuracy and reliability in customer intention recognition.
[0004] Therefore, an optimized method, device and storage medium for analyzing customer intention based on call content are desired. Summary of the invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide a customer intention analysis method, device and storage medium based on call content, which extracts the original audio data from the call recording system, performs noise suppression processing on it, and then uses multiple speech recognition models to perform speech recognition on the audio data to obtain multiple speech recognition results. Then, by introducing natural language processing technology based on deep learning, multiple speech recognition results are semantically parsed and core semantics are clustered to fuse the core semantic points of multiple speech recognition results, so as to achieve a deep understanding of the audio content, and then perform customer intention recognition on this basis to intelligently determine the customer's intention type. In this way, through the fusion analysis of multi-model speech recognition results, the accuracy and efficiency of customer intention analysis can be effectively improved, providing strong support for the company's customer service, marketing decisions, etc.
[0006] According to one aspect of the present application, a method for analyzing customer intention based on call content is provided, which includes:
[0007] Get the raw audio data from the call recording system;
[0008] Performing noise suppression on the original audio data to obtain enhanced audio data;
[0009] Inputting the enhanced audio data into a multi-model speech recognition component to obtain first to Nth model speech recognition results;
[0010] Performing core semantic clustering on the first to Nth model speech recognition results to obtain semantic core aggregation coding features of the speech recognition results, wherein the core semantic clustering on the first to Nth model speech recognition results comprises: performing dynamic clustering learning based on semantic feature significance perception on the first to Nth model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results;
[0011] The customer intention type is determined based on the semantic core aggregation coding features of the speech recognition result.
[0012] According to another aspect of the present application, a device for analyzing customer intention based on call content is provided, comprising:
[0013] The original audio data acquisition module is used to acquire the original audio data from the call recording system;
[0014] An original audio data noise suppression module, used for performing noise suppression on the original audio data to obtain enhanced audio data;
[0015] A multi-model speech recognition module, used for inputting the enhanced audio data into a multi-model speech recognition component to obtain first to Nth model speech recognition results;
[0016] A core semantic clustering module, used for performing core semantic clustering on the first to Nth model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results, wherein the core semantic clustering on the first to Nth model speech recognition results comprises: performing dynamic clustering learning based on semantic feature significance perception on the first to Nth model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results;
[0017] The customer intention type determination module is used to determine the customer intention type based on the semantic core aggregation coding features of the speech recognition result.
[0018] According to another aspect of the present application, a storage medium is provided, on which computer program instructions are stored. When the computer executable program is executed by a processor, the processor executes the customer intention analysis method based on call content as described above.
[0019] Compared with the prior art, the customer intention analysis method, device and storage medium based on call content provided by the present application extracts the original audio data from the call recording system, performs noise suppression processing on it, and then uses multiple speech recognition models to perform speech recognition on the audio data to obtain multiple speech recognition results. Then, by introducing natural language processing technology based on deep learning, multiple speech recognition results are semantically parsed and core semantics are clustered to fuse the core semantic points of multiple speech recognition results, so as to achieve a deep understanding of the audio content, and then perform customer intention recognition on this basis to intelligently determine the customer's intention type. In this way, through the fusion analysis of multi-model speech recognition results, the accuracy and efficiency of customer intention analysis can be effectively improved, providing strong support for the company's customer service, marketing decisions, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0021] Figure 1 The present invention is a flowchart of a method for analyzing customer intention based on call content according to an embodiment of the present application.
[0022] Figure 2 Schematic diagram of data flow of a method for analyzing customer intention based on call content according to an embodiment of the present application.
[0023] Figure 3 This is a flowchart of sub-step S4 of the customer intention analysis method based on call content according to an embodiment of the present application.
[0024] Figure 4 This is a flowchart of sub-step S42 of the customer intention analysis method based on call content according to an embodiment of the present application.
[0025] Figure 5 This is a flowchart of sub-step S422 of the customer intention analysis method based on call content according to an embodiment of the present application.
[0026] Figure 6 4 is a block diagram of a device for analyzing customer intention based on call content according to an embodiment of the present application. DETAILED DESCRIPTION
[0027] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0028] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.
[0029] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. Meanwhile, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.
[0030] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.
[0031] It is worth noting that in this application, all actions to obtain data are carried out in compliance with the relevant data protection laws and policies of the country where the data is located, and with the authorization given by the owner of the corresponding device.
[0032] In response to the technical problems described in the above background technology, this application proposes a customer intention analysis method based on call content, which extracts the original audio data from the call recording system, performs noise suppression processing on it, and then uses multiple speech recognition models to perform speech recognition on the audio data to obtain multiple speech recognition results. Then, by introducing natural language processing technology based on deep learning, multiple speech recognition results are semantically parsed and core semantics are clustered to fuse the core semantic points of multiple speech recognition results, so as to achieve a deep understanding of the audio content, and then perform customer intention recognition on this basis to intelligently determine the customer's intention type. In this way, through the fusion analysis of multi-model speech recognition results, the accuracy and efficiency of customer intention analysis can be effectively improved, providing strong support for the company's customer service, marketing decisions, etc.
[0033] Figure 1The present invention is a flowchart of a method for analyzing customer intention based on call content according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the method for analyzing customer intention based on call content according to an embodiment of the present application. Figure 1 and Figure 2 As shown, the customer intention analysis method based on call content includes the steps of: S1, obtaining original audio data from a call recording system; S2, performing noise suppression on the original audio data to obtain enhanced audio data; S3, inputting the enhanced audio data into a multi-model speech recognition component to obtain first to N-th model speech recognition results; S4, performing core semantic clustering on the first to N-th model speech recognition results to obtain semantic core aggregation coding features of speech recognition results, wherein performing core semantic clustering on the first to N-th model speech recognition results includes: performing dynamic clustering learning based on semantic feature significance perception on the first to N-th model speech recognition results to obtain semantic core aggregation coding features of the speech recognition results; S5, determining the customer intention type based on the semantic core aggregation coding features of the speech recognition results.
[0034] In the above-mentioned customer intention analysis method based on call content, the step S1 obtains the original audio data from the call recording system. Specifically, the call recording system is usually integrated into the information technology architecture of the enterprise customer service center (such as a call center), and records the call content between customers and service personnel in real time through a telephone gateway, CTI (computer telephone integration) or a cloud communication platform (such as Zoom, Tencent Cloud Call Center). In modern business operations, telephone communication between enterprises and customers is extremely frequent. As an important tool for recording the communication content between enterprises and customers, the call recording system can provide a rich data foundation for customer intention analysis and enterprise strategy formulation.
[0035] Specifically, when designing the system interface for extracting raw audio data, the diversity and complexity of the call recording system must be fully considered. In the specific implementation process, the data acquisition system needs to be compatible with multiple communication protocol standards, including the E1 / T1 digital trunk interface in the traditional PSTN network, the SIP / RTP streaming media protocol in the IP network, and the emerging WebRTC real-time communication protocol. Through the protocol conversion middleware, seamless access of signals of different formats is achieved: for analog lines, high-precision analog-to-digital conversion chips are used to complete the digital processing of voice signals; for IP signaling, SIP header information is parsed to identify the identity of the call participants, and the media stream is captured in real time through the RTP protocol. The system has a built-in adaptive protocol detection module that can automatically switch the acquisition mode according to the channel characteristics to ensure compatibility between different terminals such as VoIP soft phones, IP phones and mobile APPs.
[0036] The call process involves multiple audio streams such as the caller, the called party, and the agent's equipment, which require multi-channel synchronous recording through a hardware mixer. Each sound source channel is equipped with an independent signal processing link, including an anti-aliasing filter, an automatic gain control (AGC) circuit, and a noise threshold detection module to suppress background noise and optimize signal quality. In the digital domain, time division multiplexing (TDM) technology is used to arrange multiple audio signals in fixed time slots to form structured digital audio frames. For multi-party conference scenarios, the voice streams of all participants are dynamically tracked through the SIP media negotiation mechanism, and the synchronization of multiple audio channels is achieved based on the RTP timestamp.
[0037] In the face of call recordings of different sources and types, in order to facilitate subsequent processing, the collected data needs to be converted into a unified standard format. Common audio file formats such as WAV and MP3 have their own characteristics, but in this scenario, it is particularly important to choose a format that maintains high sound quality and effectively compresses the file size. Generally, the WAV format is given priority as an intermediate format due to its lossless characteristics. Although its file size is relatively large, it can ensure that the accuracy of speech recognition is not affected. For those recording files that are directly stored in a compressed format, it is necessary to use a special decoding tool to convert them to WAV format so that the subsequent processing steps can proceed smoothly. In this process, attention should also be paid to retaining the metadata associated with each call recording, such as call timestamps, participant identifiers, etc., which are important for accurately understanding the context of the call.
[0038] The collected raw audio data needs to be sent to the back-end processing system through a highly reliable transmission channel. The system adopts a dual-plane transmission architecture: the main path is based on enterprise dedicated lines or MPLS VPNs, providing a low-latency, high-bandwidth transmission environment; the backup path uses a public network encrypted tunnel to achieve redundant backup, combined with the new transport layer technology of the QUIC protocol, effectively reducing the impact of network jitter on voice streams. Each audio data packet contains a complete timestamp, sequence number and checksum information. The receiving end reorganizes the out-of-order data packets through a sliding window algorithm and uses forward error correction (FEC) technology to correct potential transmission errors. At the storage level, the system builds a distributed object storage cluster, and the audio data is written to multiple nodes in the form of fragments. The RAID6+ erasure code technology is used to ensure data reliability, and the hot and cold data are hierarchically managed to optimize storage efficiency. Considering data privacy and security issues, the call recording system also needs to comply with strict laws, regulations and industry standards. In the process of obtaining raw audio data, it is crucial to protect customers' personal information from being leaked. Therefore, a series of measures must be taken to strengthen data security, such as encrypted transmission and access rights control.
[0039] In the above-mentioned customer intention analysis method based on call content, the step S2 suppresses noise on the original audio data to obtain enhanced audio data. In a specific example of the present application, the step S2 includes: using spectral subtraction to process the original audio data to obtain the enhanced audio data. It should be understood that in actual call scenarios, various types of environmental noise are inevitably mixed into the original audio. For example, the customer is in a noisy shopping mall, a multi-person discussion environment in an office, or there is electromagnetic interference in the telephone line itself, which may cause the voice signal to be blurred, seriously affecting the performance of the voice recognition model. Therefore, in order to improve the quality of audio data and eliminate the negative impact of noise on subsequent voice recognition, the present application uses spectral subtraction to suppress noise interference in the original audio data. Specifically, the spectral subtraction method is based on the important assumption that the statistical characteristics of noise are relatively stable within a certain period of time, and uses a specific algorithm (such as a method based on statistical averaging) to estimate the spectral characteristics of noise. Then, the spectrum of the noisy speech is subtracted from the estimated noise spectrum in the frequency domain to obtain an approximate estimate of the pure speech spectrum. Finally, the estimation result in the frequency domain is converted back to the time domain by inverse transformation to obtain enhanced audio data. Through spectral subtraction processing, the interference of environmental noise, line noise, etc. in the original audio can be greatly reduced, making the speech part clearer and more distinguishable, which helps to improve the accuracy of subsequent speech recognition.
[0040] Specifically, first, for the input raw audio data, the short-time Fourier transform is used to convert the audio signal in the time domain into a frequency domain representation. The short-time Fourier transform applies a series of window functions to the signal and performs a Fourier transform on the signal segment within each window, thereby providing frequency resolution while retaining time information. It is crucial to choose an appropriate window size. A larger window size helps to improve frequency resolution but may reduce time resolution; conversely, a smaller window size has the opposite effect. In the specific implementation process, selecting a window length between 20ms and 40ms can better balance the relationship between the two. In addition, the choice of window functions, such as Hanning window, Hamming window or Blackman window, etc., should also be considered. These window functions can reduce spectral leakage and improve the accuracy of spectral estimation.
[0041] Then the noise spectrum needs to be estimated. A common method is to collect background noise samples before the call starts or during silent intervals, and use these samples to build a noise spectrum model. The specific approach is to calculate the average power spectral density of all frames in these silent segments as a preliminary estimate of the noise spectrum. However, this method relies on the existence of a sufficiently long silent period during the call, which may not be applicable to continuous speech. To this end, an adaptive algorithm can also be used in combination to dynamically adjust the noise spectrum estimate based on the spectral characteristics of the current frame. For example, when it is detected that the energy of the current frame is significantly lower than that of the adjacent frame, it can be regarded as a potential noise frame, and the noise spectrum estimate is updated accordingly. This adaptive method based on statistical characteristics can more flexibly respond to the changing noise environment, ensuring that a high noise reduction effect can be maintained even in complex backgrounds.
[0042] In the specific implementation process, in order to further improve the effect of spectral subtraction, a smoothing mechanism can also be introduced. Since the spectrum within a single frame may fluctuate greatly, directly applying spectral subtraction may cause the output spectrum to be not smooth enough, thus affecting the auditory experience. To this end, smoothing can be implemented between adjacent frames, for example, by averaging the spectrum estimation values of the previous and next frames, or using a recursive filter to correct the spectrum of the current frame, so that the output spectrum is more coherent and natural. In addition, considering that different frequency bands have different effects on speech clarity, different spectral subtraction parameters can be applied to different frequency bands according to frequency characteristics, focusing on protecting the high-frequency and low-frequency parts that are critical to speech recognition, while appropriately relaxing the requirements for the intermediate frequency bands to achieve the best overall noise reduction effect.
[0043] In the above-mentioned customer intention analysis method based on call content, the step S3 inputs the enhanced audio data into the multi-model speech recognition component to obtain the first to Nth model speech recognition results. It should be understood that, considering the inherent limitations of a single speech recognition model, in the face of complex and diverse call scenarios, such as the accents of customers in different regions, a large number of industry jargons or newly coined words mixed in calls, and various special voice habits (speaking too fast, too slow, abnormal intonation, etc.), a single model is difficult to fully adapt and accurately recognize. To address this problem, the present application uses multiple speech recognition models to process the enhanced audio data in parallel, and obtains diversified recognition results from different model perspectives, so as to make full use of the differences between different models in speech feature extraction, acoustic model training and language model adaptation, and complement each other, thereby improving the accuracy and comprehensiveness of overall speech recognition. Specifically, the multi-model speech recognition component integrates a variety of speech recognition models built based on different algorithm frameworks and large-scale training data. For example, some speech recognition models focus more on learning the pronunciation features of standard Mandarin during acoustic model training, and have excellent recognition effects on standard accents; while other models may incorporate a large amount of industry-specific text data in language model training, and have more advantages in recognizing industry terms; and some models have better speech recognition effects for customers in specific regions through targeted training of dialects and accents in different regions. In this application, when the enhanced audio data is input into the multi-model speech recognition component, each speech recognition model in the component extracts and decodes the voice signal in the audio based on its own unique training experience and algorithm logic, and finally outputs its own recognition text, thereby obtaining the first to Nth model speech recognition results. Through this multi-model parallel processing strategy, when facing complex and changeable call scenarios, the advantages of each model can be combined to improve the accuracy and robustness of speech recognition.
[0044] In the above-mentioned customer intention analysis method based on call content, the step S4 performs core semantic clustering on the first to Nth model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results, wherein the core semantic clustering on the first to Nth model speech recognition results includes: performing dynamic clustering learning based on semantic feature significance perception on the first to Nth model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results. Figure 3 FIG. 4 is a flowchart of sub-step S4 of the method for analyzing customer intention based on call content according to an embodiment of the present application. Figure 3As shown, the step S4 includes the steps of: S41, using a semantic embedding encoder based on a Bert model to perform semantic embedding encoding on each model speech recognition result in the first to N model speech recognition results to obtain a semantic embedding encoding vector of the first to N model speech recognition results; S42, performing core semantic clustering learning on the semantic embedding encoding vector of the first to N model speech recognition results to obtain a speech recognition result semantic core aggregation encoding vector as the speech recognition result semantic core aggregation encoding feature.
[0045] Specifically, the step S41 uses a semantic embedding encoder based on the Bert model to perform semantic embedding encoding on each model speech recognition result in the first to N model speech recognition results to obtain a semantic embedding encoding vector of the first to N model speech recognition results. It should be understood that since the traditional semantic analysis method based on keyword matching cannot capture the complex semantic relationships and context dependencies in the text, it is difficult to effectively distinguish texts with similar semantics but different expressions, which may lead to misjudgment of customer intentions. In this regard, the present application introduces a pre-trained Bert model as a semantic embedding encoder to perform semantic embedding encoding on each model speech recognition result in the first to N model speech recognition results, so as to map the speech recognition text of each model into a vector representation of a high-dimensional semantic space. The Bert model has strong semantic understanding capabilities through large-scale unsupervised pre-training. Based on the bidirectional Transformer architecture, it can effectively capture the long-range dependencies between words, understand complex semantic phenomena such as transitions and reference resolution in the text, and then generate high-quality semantic feature representations with context dependence, and obtain the semantic embedding coding vectors of the first to Nth model speech recognition results, thereby providing a comparable mathematical representation for the semantic aggregation of subsequent multi-model speech recognition results.
[0046] Specifically, the step S42 performs core semantic clustering learning on the semantic embedding coding vectors of the speech recognition results of the first to Nth models to obtain the semantic core aggregation coding vector of the speech recognition result as the semantic core aggregation coding feature of the speech recognition result. Specifically, in the actual scenario of customer intention analysis, the outputs of multiple speech recognition models often have problems of semantic redundancy, local contradictions or noise interference. For example, when the content of the call is "I want to cancel the package, but the signal is really good", model A may accurately transcribe it as "cancel the package", while model B may mistakenly identify "cancel" as "renewal" due to noise interference, resulting in differences in subsequent semantic understanding. That is, the heterogeneity of multi-model speech recognition results will lead to the fragmentation of semantic space. Therefore, in order to extract the most representative core semantic information from multiple model speech recognition results, the present application further performs core semantic clustering on the semantic embedding coding vectors of the first to Nth model speech recognition results, by mapping the semantic embedding coding vectors of multiple model speech recognition results to a unified spectral space, in order to utilize feature saliency modeling and clustering learning, to mine potential semantic commonalities, to extract stable core semantic expressions, thereby eliminating redundancy and error interference, and finally generating a highly generalized and consistent speech recognition result semantic core aggregation coding vector to support intent classification. Among them, Figure 4 FIG. 4 is a flowchart of sub-step S42 of the method for analyzing customer intention based on call content according to an embodiment of the present application. Figure 4 As shown, the step S42 includes the steps of: S421, based on the semantic independence between the semantic embedding coding vectors of each model speech recognition result in the semantic embedding coding vectors of the first to N model speech recognition results, performing feature significance modulation on the semantic embedding coding vectors of the first to N model speech recognition results to obtain the first to N independence modulated model speech recognition result semantic embedding coding vectors; S422, performing feature dynamic clustering analysis on the semantic embedding coding vectors of the first to N independence modulated model speech recognition results to obtain the semantic core aggregation coding vector of the speech recognition result.
[0047] More specifically, the step S421 includes: first, calculating the semantic space independence description operator between any two model speech recognition result semantic embedding coding vectors among the first to N model speech recognition result semantic embedding coding vectors based on the semantic projection matrix to obtain a model speech recognition result semantic independence spectral space coding matrix composed of multiple model speech recognition result semantic subspace independence description operators, which is expressed by the formula:
[0048]
[0049] in, represents the semantic embedding coding vector of the speech recognition results of the first to N models, , , , and They represent the first, second, and third semantic embedding encoding vectors of the speech recognition results of the first to Nth models, respectively. , and The semantic embedding encoding vector of the model speech recognition result, is the number of semantic embedding coding vectors of the model speech recognition results, represents the calculation of the second norm, and Represent different semantic projection matrices, represents the transpose of a vector, represents vector multiplication, express and The semantic subspace independence of the model speech recognition results is described by operators.
[0050] That is, when different speech recognition models decode the same audio clip, the underlying semantic space structure may be significantly different. This modality specificity leads to direct splicing or average fusion, which will destroy the intrinsic structure of the semantic features and dilute the key intent signal. Therefore, the present application projects the semantic embedding coding vector of the model speech recognition result into different semantic subspaces and calculates the semantic space independence description operator between any two semantic embedding coding vectors of the model speech recognition result. This can quantify the unique semantic subspace features captured by each model and their interdependencies, and then generate a semantic independence spectral space coding matrix of the model speech recognition result that can characterize the synergistic advantages of multiple models.
[0051] Then, the semantic independence spectral space coding matrix of the model speech recognition result is activated based on the activation function to obtain the semantic independence spectral space coding feature matrix of the model speech recognition result, which is expressed as follows:
[0052] in, represents the sigmoid activation function, , , and Respectively and , and , and , and The semantic subspace independence description operator of the model speech recognition results between The spectral space encoding feature matrix representing the semantic independence of the model's speech recognition results.
[0053] That is, the present application reconstructs the topological structure of the feature space by applying dynamic threshold constraints to each element of the semantic independence spectral space encoding matrix of the model speech recognition result through spectral space activation, and compresses the eigenvalue range with the help of the normalization characteristics of Sigmoid, so that the elements in the matrix can more accurately reflect the semantic differences and mutual independence between different models, thereby making the generated model speech recognition result semantic independence spectral space encoding feature matrix have stronger discrimination ability, thereby enhancing the robustness and accuracy of subsequent feature clustering analysis.
[0054] Finally, each model speech recognition result semantic embedding coding vector in the first to Nth model speech recognition result semantic embedding coding vector and the model speech recognition result semantic independence spectral space coding feature matrix are input into the semantic significance modulation module to obtain the first to Nth independence modulation model speech recognition result semantic embedding coding vector, which is expressed by the formula:
[0055] in, for The characteristic scale value of Indicates Semantic embedding coding vector of speech recognition results of independence modulation model.
[0056] That is, the present application performs semantic saliency modulation on the semantic embedding coding vectors of each model speech recognition result based on the semantic independence spectral-space coding feature matrix of the model speech recognition result, so that the semantic embedding coding vectors of each model speech recognition result, while maintaining their original semantic features, further consider the semantic differences and mutual independence between other models, and dynamically adjust their own semantic feature distribution accordingly, so as to highlight their own unique semantic contribution and reduce the interference of redundant information, thereby improving the semantic expression ability of each model speech recognition result in multi-model collaborative work, and providing clearer and more discriminative input for subsequent feature dynamic clustering analysis, so as to enhance the adaptability to complex semantic structures.
[0057] Figure 5 FIG. 4 is a flowchart of sub-step S422 of the method for analyzing customer intention based on call content according to an embodiment of the present application. Figure 5As shown, the step S422 includes the steps of: S4221, based on the feature distribution of each independence modulation model speech recognition result semantic embedding coding vector in the first to Nth independence modulation model speech recognition result semantic embedding coding vector, calculating the semantic stability factor of each independence modulation model speech recognition result semantic embedding coding vector to obtain the semantic stability factor of the first to Nth model speech recognition results; S4222, performing Softmax-based normalization processing on the semantic stability factor of the first to Nth model speech recognition result to obtain the semantic dynamic aggregation weight coefficient of the first to Nth model speech recognition result; S4223, based on the semantic dynamic aggregation weight coefficient of the first to Nth model speech recognition result, performing weighted aggregation on the semantic embedding coding vector of the first to Nth independence modulation model speech recognition result to obtain the semantic core aggregation coding vector of the speech recognition result.
[0058] That is, after independence modulation of the semantic embedding coding vectors of each model speech recognition result, in order to further utilize the complementarity of multi-model speech recognition results and extract the core semantic information of the original audio, the present application determines the semantic stability and importance of each independent modulation model speech recognition result semantic embedding coding vector by performing feature distribution analysis on the semantic embedding coding vectors of the speech recognition results of each independent modulation model, realizes dynamic weight allocation, and based on this, performs weighted aggregation on the semantic embedding coding vectors of the speech recognition results of the first to Nth independent modulation model, thereby generating a speech recognition result semantic core aggregated coding vector that can accurately reflect the core semantics of the original audio.
[0059] In particular, in a preferred example of the present application, the step S4221 includes: first, performing a covariance-orthogonal composite transformation on the semantic embedding coding vector of the independence modulation model speech recognition result to obtain an offset factor, which is expressed as:
[0060]
[0061] in, Represents the semantic alignment space matrix of the model speech recognition results, represents the identity matrix, express The corresponding semantic embedding coding vector of the speech recognition result of the spatial alignment independence modulation model, represents vector subtraction, Represents the semantic covariance matrix of the model speech recognition results, represents the scaling factor, represents the vector inner product, Indicates the offset factor.
[0062] Here, due to the differences in feature extraction of diverse accents and call backgrounds by different models, the feature distribution of the semantic embedding coding vectors of the first to Nth independence modulation model speech recognition results often presents a dynamic mismatch phenomenon from the cluster center to the edge, which ultimately affects the reliability of multi-model semantic fusion. Therefore, the present application further calculates the corresponding offset factor by performing a covariance-orthogonal composite transformation on the semantic embedding coding vector of the independence modulation model speech recognition result, so as to perform downward alignment through the offset factor. Specifically, the semantic alignment space matrix of the model speech recognition result is first obtained by decomposing the orthogonal constraint of the unit matrix, and the semantic embedding coding vector of the independence modulation model speech recognition result is projected into a unified orthogonal basis space to maintain the metric topological relationship between the original features. Subsequently, the covariance matrix of the semantic embedding coding vector of the independence modulation model speech recognition result is introduced to represent the geometric structure of the feature distribution, and the local geometric adaptation from the source domain (aggregation core) to the target domain (dispersion boundary) is reconstructed by combining the joint optimization of the mean direction representation and the variance representation, so as to achieve downward spatial alignment, so as to eliminate the dynamic cross-domain mismatch caused by emphasizing the modal significance relationship in the clustering process.
[0063] Then, the characteristic kurtosis of the semantic embedding coding vector of the independence modulation model speech recognition result is calculated, divided by two and then added to the offset factor to obtain the semantic stability factor of the model speech recognition result, which is expressed as follows:
[0064] in, Represents the feature mean of the semantic embedding coding vector of the speech recognition result of the independence modulation model, The characteristic standard deviation of the semantic embedding coding vector representing the speech recognition result of the independence modulation model, represents the calculation of expected value, represents the calculation of characteristic kurtosis, express Semantic stability factor of model speech recognition results.
[0065] That is, by calibrating the feature kurtosis to construct a robust feature representation, we can suppress the amplification effect of extreme distribution noise while retaining key semantic patterns. Then, by introducing the global benchmark after spectral space alignment, i.e., the offset factor, we force the local statistical characteristics to be aligned to the global optimization target, thereby eliminating the distribution offset caused by differences in noise suppression strength or speech recognition confidence between models. Through the above calculation method, we generate the semantic stability factor of the model speech recognition result, optimize the geometric consistency of the semantic embedding space, and make the feature distribution after multi-model fusion closer to the real semantic manifold.
[0066] In a specific example of the present application, step S4222 is expressed by the formula:
[0067] in, represents the normalization function, express The semantic dynamic aggregation weight coefficient of the model speech recognition results.
[0068] That is, through the normalization operation, the local statistical characteristics between models are further transformed into global optimization variables, the semantic dynamic aggregation weight coefficients of the model speech recognition results are generated, and the redundancy advantages of multiple models are transformed into geometrically robust semantic representations, laying the foundation for accurately extracting the core patterns of customer intentions.
[0069] In a specific example of the present application, step S4223 is expressed by the formula:
[0070] in, Represents the semantic core aggregate encoding vector of speech recognition results.
[0071] That is, the semantic core aggregate coding vector of the speech recognition result generated by weighted aggregation retains the complementary semantic information of each model while forcing the overall distribution to move closer to the feature manifold of the high-stability model, thereby solving the problems of noise amplification and semantic dilution caused by traditional direct splicing or average aggregation.
[0072] In the above-mentioned customer intention analysis method based on call content, the step S5 determines the customer intention type based on the semantic core aggregation coding features of the speech recognition result. In a specific example of the present application, the step S5 includes: inputting the semantic core aggregation coding vector of the speech recognition result into the intention recognition engine based on the classifier to obtain the customer intention type label. Here, the semantic core aggregation coding vector of the speech recognition result has been stripped of the noise caused by speech recognition errors after multi-model fusion, and the core features of the customer intention (such as emotional tendency, keywords, context logic) are retained. Based on this, the classifier (such as a multi-layer perceptron MLP) learns the nonlinear decision boundary from high-dimensional semantic coding features to customer intention type labels (such as "complaint", "consultation", "business handling") in the training stage through supervised learning, and establishes a corresponding relationship between complex semantic expressions and limited business labels; in the actual application stage, the classifier receives the semantic core aggregation coding vector of the speech recognition result to be classified, and maps the deep semantic features to the predefined customer intention type labels in the business system by performing multi-level feature learning on it, so as to quickly and accurately identify the customer's intention. This deep learning-based intent recognition method can not only effectively overcome the limitations of a single speech recognition model, but also accurately capture the customer's true intentions in a complex and changeable call environment, support downstream tasks such as automated customer service ticket allocation and real-time speech recommendation, so as to improve the intelligence level of customer service and customer satisfaction.
[0073] In the specific implementation process, after obtaining the customer's intention type, such as "complaint", "inquiry" or "business processing", the system immediately starts a series of preset response mechanisms to ensure that the customer's demands can be responded to quickly and accurately. For situations marked as "complaint", the system will automatically trigger a specially designed complaint handling process. This process may include generating a detailed complaint record, recording all the questions and dissatisfaction raised by the customer, and automatically generating a work order and assigning it to the appropriate customer service representative for follow-up. In addition, the system can also provide customer service personnel with suggestive solutions based on historical data and experience in handling similar cases, helping them to solve customer problems more quickly and effectively. In this way, not only the speed of problem solving is improved, but also the customer's trust in the quality of corporate services is enhanced.
[0074] When the system identifies the customer's intention type as "consultation", it will adopt different response strategies. In this case, the first task is to quickly locate the information that the customer is concerned about and provide answers in the most direct way. To this end, the system can integrate an intelligent knowledge base search function to automatically retrieve relevant information and present it to customer service personnel based on the keywords or topics in the customer's questions. This instant knowledge support greatly shortens the time to find answers, allowing customer service personnel to give customers satisfactory answers in a short time. Not only that, for some common questions, the system can also pre-set standard answer templates to further speed up the response speed and improve work efficiency. At the same time, in order to better meet the personalized needs of customers, the system can customize and recommend related additional information or product suggestions based on the customer's historical interaction records, thereby increasing sales opportunities and promoting business growth.
[0075] For customer intentions in the "business processing" category, the system needs to prepare a more detailed service plan. This type of intention often involves specific product or service operation processes, so special attention needs to be paid to details and service quality. After confirming the customer's intention, the system will guide the customer to complete the corresponding business process, whether it is filling out an application form online or guiding the customer to complete the necessary identity verification steps. In order to ensure a smooth process, the system may provide real-time help guides or video tutorials to help customers easily understand the requirements of each step. In addition, throughout the entire processing process, the system will continue to monitor progress and promptly remind customers of tasks that are about to expire or matters that need attention to avoid delays or errors due to negligence.
[0076] In summary, the customer intention analysis method based on call content based on the embodiment of the present application is explained, which extracts the original audio data from the call recording system, performs noise suppression processing on it, and then uses multiple speech recognition models to perform speech recognition on the audio data to obtain multiple speech recognition results. Then, by introducing natural language processing technology based on deep learning, multiple speech recognition results are semantically parsed and core semantics are clustered to fuse the core semantic points of multiple speech recognition results, so as to achieve a deep understanding of the audio content, and then perform customer intention recognition on this basis to intelligently determine the customer's intention type. In this way, through the fusion analysis of multi-model speech recognition results, the accuracy and efficiency of customer intention analysis can be effectively improved, providing strong support for the company's customer service, marketing decisions, etc.
[0077] Furthermore, a device for analyzing customer intention based on call content is also provided.
[0078] Figure 6 FIG. 1 is a block diagram of a device for analyzing customer intention based on call content according to an embodiment of the present application. Figure 6As shown, according to the embodiment of the present application, the customer intention analysis device 100 based on call content includes: an original audio data acquisition module 110, which is used to acquire original audio data from a call recording system; an original audio data noise suppression module 120, which is used to perform noise suppression on the original audio data to obtain enhanced audio data; a multi-model speech recognition module 130, which is used to input the enhanced audio data into a multi-model speech recognition component to obtain first to N-th model speech recognition results; a core semantic clustering module 140, which is used to perform core semantic clustering on the first to N-th model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results, wherein the core semantic clustering on the first to N-th model speech recognition results includes: performing dynamic clustering learning based on semantic feature significance perception on the first to N-th model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results; a customer intention type determination module 150, which is used to determine the customer intention type based on the semantic core aggregation coding features of the speech recognition results.
[0079] Here, those skilled in the art can understand that the specific operations of each module in the above-mentioned customer intention analysis device based on call content have been described in the above reference. Figures 1 to 5 The description of the customer intention analysis method based on call content has been introduced in detail, and therefore, its repeated description will be omitted.
[0080] In addition, an embodiment of the present application may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes the steps in the functions of the customer intention analysis method based on call content according to various embodiments of the present application described in this specification.
[0081] The computer readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0082] The basic principle of the present invention is described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, strengths, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and facilitation of understanding, rather than limitation, and the above details do not limit the present invention to being implemented by adopting the above specific details.
[0083] In the above embodiments, the description of each embodiment has its own emphasis. For the parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0084] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference to a figure in a claim should not be considered as limiting the claim to which it relates.
[0085] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.
[0086] Finally, it should be noted that the above description has been given for the purpose of illustration and description. In addition, the above embodiments are only used to illustrate the technical solution of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A method for analyzing customer intention based on call content, characterized in that: include: Get the raw audio data from the call recording system; Performing noise suppression on the original audio data to obtain enhanced audio data; Inputting the enhanced audio data into a multi-model speech recognition component to obtain first to Nth model speech recognition results; Performing core semantic clustering on the speech recognition results of the first to Nth models to obtain semantic core aggregation coding features of the speech recognition results; Determining the customer's intention type based on the semantic core aggregate coding features of the speech recognition result; The core semantic clustering of the first to Nth model speech recognition results includes: performing dynamic clustering learning based on semantic feature significance perception on the first to Nth model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results, specifically including: Using a semantic embedding encoder based on a Bert model, semantic embedding encoding is performed on each of the first to N-th model speech recognition results to obtain semantic embedding encoding vectors of the first to N-th model speech recognition results; Based on the semantic independence between the semantic embedding coding vectors of each model speech recognition result in the first to N model speech recognition result semantic embedding coding vectors, the first to N model speech recognition result semantic embedding coding vectors are subjected to feature significance modulation to obtain first to N independence modulated model speech recognition result semantic embedding coding vectors; A feature dynamic clustering analysis is performed on the semantic embedding coding vectors of the first to Nth independence modulation model speech recognition results to obtain the semantic core aggregation coding vector of the speech recognition results.
2. The method for analyzing customer intention based on call content according to claim 1, characterized in that: Performing noise suppression on the original audio data to obtain enhanced audio data, comprising: The original audio data is processed using spectral subtraction to obtain the enhanced audio data.
3. The method for analyzing customer intention based on call content according to claim 2, characterized in that: Based on the semantic independence between the semantic embedding coding vectors of each model speech recognition result in the first to N model speech recognition result semantic embedding coding vectors, the first to N model speech recognition result semantic embedding coding vectors are subjected to feature significance modulation to obtain first to N independence modulated model speech recognition result semantic embedding coding vectors, including: Calculating the semantic space independence description operator between any two model speech recognition result semantic embedding coding vectors among the first to Nth model speech recognition result semantic embedding coding vectors based on the semantic projection matrix to obtain a model speech recognition result semantic independence spectral space coding matrix composed of multiple model speech recognition result semantic subspace independence description operators; Performing spectral space activation based on an activation function on the semantic independence spectral space coding matrix of the model speech recognition result to obtain a semantic independence spectral space coding feature matrix of the model speech recognition result; The semantic embedding coding vectors of each model speech recognition result in the first to Nth model speech recognition result semantic embedding coding vectors and the semantic independence spectral space coding feature matrix of the model speech recognition result are input into the semantic significance modulation module to obtain the semantic embedding coding vectors of the first to Nth independence modulation model speech recognition results.
4. The method for analyzing customer intention based on call content according to claim 3, characterized in that: Performing feature dynamic clustering analysis on the semantic embedding coding vectors of the first to Nth independence modulation model speech recognition results to obtain the semantic core aggregation coding vector of the speech recognition result, including: Based on the feature distribution of each independence modulation model speech recognition result semantic embedding coding vector in the first to Nth independence modulation model speech recognition result semantic embedding coding vectors, calculating the semantic stability factor of each independence modulation model speech recognition result semantic embedding coding vector to obtain the first to Nth model speech recognition result semantic stability factors; Performing a Softmax-based normalization process on the semantic stability factors of the first to Nth model speech recognition results to obtain semantic dynamic aggregation weight coefficients of the first to Nth model speech recognition results; Based on the semantic dynamic aggregation weight coefficients of the first to Nth model speech recognition results, the semantic embedded coding vectors of the first to Nth independence modulation model speech recognition results are weightedly aggregated to obtain the semantic core aggregation coding vector of the speech recognition result.
5. The method for analyzing customer intention based on call content according to claim 4, characterized in that: Based on the feature distribution of each independence modulation model speech recognition result semantic embedding coding vector in the first to Nth independence modulation model speech recognition result semantic embedding coding vector, calculating the semantic stability factor of each independence modulation model speech recognition result semantic embedding coding vector to obtain the first to Nth model speech recognition result semantic stability factor, including: Performing a covariance-orthogonal composite transformation on the semantic embedding coding vector of the independence modulation model speech recognition result to obtain an offset factor; The characteristic kurtosis of the semantic embedding coding vector of the independence modulation model speech recognition result is calculated, divided by two and then added to the offset factor to obtain the semantic stability factor of the model speech recognition result.
6. The method for analyzing customer intention based on call content according to claim 5, characterized in that: Based on the semantic core aggregation coding features of the speech recognition result, the customer intention type is determined, including: The semantic core aggregate coding vector of the speech recognition result is input into the classifier-based intention recognition engine to obtain the customer intention type label.
7. A customer intention analysis device based on call content, characterized in that: include: The original audio data acquisition module is used to acquire the original audio data from the call recording system; An original audio data noise suppression module, used for performing noise suppression on the original audio data to obtain enhanced audio data; A multi-model speech recognition module, used for inputting the enhanced audio data into a multi-model speech recognition component to obtain first to Nth model speech recognition results; A core semantics clustering module, used for performing core semantics clustering on the speech recognition results of the first to Nth models to obtain semantic core aggregation coding features of the speech recognition results; A customer intention type determination module, used to determine the customer intention type based on the semantic core aggregation coding features of the speech recognition result; The core semantic clustering of the first to Nth model speech recognition results includes: performing dynamic clustering learning based on semantic feature significance perception on the first to Nth model speech recognition results to obtain the semantic core aggregation coding features of the speech recognition results, specifically including: Using a semantic embedding encoder based on a Bert model, semantic embedding encoding is performed on each of the first to N-th model speech recognition results to obtain semantic embedding encoding vectors of the first to N-th model speech recognition results; Based on the semantic independence between the semantic embedding coding vectors of each model speech recognition result in the first to N model speech recognition result semantic embedding coding vectors, the first to N model speech recognition result semantic embedding coding vectors are subjected to feature significance modulation to obtain first to N independence modulated model speech recognition result semantic embedding coding vectors; A feature dynamic clustering analysis is performed on the semantic embedding coding vectors of the first to Nth independence modulation model speech recognition results to obtain the semantic core aggregation coding vector of the speech recognition results.
8. A storage medium storing a computer program, wherein the computer program is used to execute the customer intention analysis method based on call content as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Digital human interactive dialogue method and system
CN117423338A
Client tag determination method and device based on call content, and storage medium
CN118377909A