A car purchase willingness evaluation method, system, device and medium
By analyzing the spectrum and semantic pairing of audio dialogues during car purchase intention assessment, an emotion fluctuation curve is constructed to identify the psychological state of car purchase. This solves the problem of inaccurate car purchase intention assessment in existing technologies and enables dynamic assessment and refined analysis of car purchase intention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING SHUNSHI SICHENG TECH CO LTD
- Filing Date
- 2025-09-04
- Publication Date
- 2026-08-04
AI Technical Summary
In existing technologies, car purchase intention assessments cannot accurately track customers' dynamic emotional changes during the decision-making process, resulting in assessment results that cannot distinguish between customers' true feelings and superficial words, thus affecting the accuracy of the assessment.
By acquiring audio recordings of conversations between salespeople and customers, converting them into time-stamped text, and performing spectral analysis, an emotional fluctuation curve for the target customer is constructed. This is combined with semantic content for temporal matching to generate an emotional feature sequence, identify a car-buying psychological state vector, and construct a probability curve for intention decay, thereby achieving dynamic assessment of car-buying intention.
It dynamically reflects the trajectory of customer emotional changes, quantifies the psychological state of car purchase, improves the accuracy and comprehensiveness of car purchase intention assessment, overcomes the problems of single information dimension and rough assessment in existing technologies, and improves the accuracy and robustness of emotion recognition.
Smart Images

Figure CN121190105B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital signal processing technology, specifically to a method, system, device, and medium for assessing car purchase intention. Background Technology
[0002] With the continuous development of the automobile consumer market, the quality of sales service during the car purchase process is receiving increasing attention. In the actual sales process, sales personnel need to promptly grasp the customer's purchasing intentions in order to provide more targeted services and improve the efficiency of closing the deal.
[0003] In existing technologies, the audio of the conversation is converted into text, and then a series of car-related keywords, such as "price," "model," and "discount," are set in the converted text. By statistically analyzing the frequency and position of the keywords, the customer's level of interest in purchasing the car can be preliminarily determined.
[0004] However, existing technologies only process the text content after speech conversion, which can only tell you what the customer said. The customer's true emotions are often revealed through tone of voice rather than words. This lack of information makes it impossible to distinguish between the customer's true feelings and superficial words in the car purchase intention assessment results, which affects the accuracy of the car purchase intention assessment. Summary of the Invention
[0005] This application provides a method, system, device, and medium for assessing car purchase intention, which addresses the technical problem of inaccurate car purchase intention assessment due to the inability to track dynamic emotional changes in customers during the decision-making process.
[0006] The first aspect of this application provides a method for assessing car purchase intention, the method comprising: The process involves acquiring audio recordings of conversations between sales personnel and target customers, converting the audio into time-stamped text; performing spectral analysis on the audio to determine the acoustic characteristics of the target customer during the conversation, and constructing an emotional fluctuation curve based on these acoustic characteristics; temporally pairing the text and the emotional fluctuation curve to generate an emotional feature sequence, which includes the target customer's semantic content and the emotional intensity value corresponding to the timestamps of the semantic content in the emotional fluctuation curve; identifying one or more target time periods in which the target customer expresses their purchase decision based on the emotional feature sequence, and determining the target customer's car-buying psychological state vector for each target time period, which includes multiple decision-making psychological states and the probability value of the target customer being in each decision-making psychological state; determining the target customer's willingness decay coefficient based on the car-buying psychological state vector, and inputting the willingness decay coefficient into a preset exponential decay function to obtain a willingness decay probability curve, which characterizes the changing trend of the target customer's car-buying willingness over time; and analyzing the willingness decay probability curve to obtain the target customer's car-buying willingness assessment result.
[0007] Optionally, spectral analysis is performed on the dialogue audio to determine the acoustic variation characteristics of the target customer during the dialogue. Based on the acoustic variation characteristics, an emotion fluctuation curve of the target customer is constructed. Specifically, this includes: detecting speech pauses in the dialogue audio and identifying the speech segment between two adjacent speech pauses as the target speech segment; differentiating the speakers in the target speech segment to obtain the target customer's speech segment; identifying neutral semantic segments in the target customer's speech segment, where neutral semantic segments are speech segments that meet preset semantic requirements; performing harmonic analysis on the neutral semantic segments to obtain pitch variation curves; performing formant tracking analysis on the neutral semantic segments to obtain spectral envelope features; performing periodic analysis on the neutral semantic segments to obtain voice jitter rate and voice flicker rate; combining the pitch variation curve, spectral envelope features, voice jitter rate, and voice flicker rate to obtain the acoustic baseline features of the target customer; and determining the acoustic variation characteristics of non-neutral semantic segments based on the acoustic baseline features and constructing an emotion fluctuation curve, where non-neutral semantic segments are the target customer's speech segments other than neutral semantic segments.
[0008] Optionally, based on acoustic baseline features, the acoustic variation characteristics of non-neutral semantic segments are determined, and an emotion fluctuation curve is constructed. Specifically, this includes: analyzing the target acoustic features of the non-neutral semantic segments, calculating the deviation between the target acoustic features and the acoustic baseline features to obtain acoustic variation characteristics; calculating the acoustic semantic matching degree for each non-neutral semantic segment, whereby the acoustic semantic matching degree characterizes the consistency between the semantic features and acoustic features of the target customer in the non-neutral semantic segments; weighted summing of the acoustic variation characteristics and the acoustic semantic matching degree to obtain the emotion intensity value of the non-neutral semantic segment; arranging the emotion intensity values in chronological order according to the timestamp information of the non-neutral semantic segments to generate a discrete emotion intensity sequence; smoothing the discrete emotion intensity sequence, and using an interpolation algorithm to connect the emotion intensity values at adjacent time points to obtain the emotion fluctuation curve.
[0009] Optionally, the acoustic semantic matching degree of each non-neutral semantic segment is calculated, specifically including: performing semantic analysis on the non-neutral semantic segments, obtaining a semantic feature sequence based on the semantic features of each preset target category word in each non-neutral semantic segment, the preset target category words including sentiment words, negation words, and transition words; analyzing the sentiment polarity of sentiment words in the non-neutral semantic segments to obtain semantic tendency values, where a positive semantic tendency value is used to characterize the positive degree of sentiment words, and a negative semantic tendency value is used to characterize the negative degree of sentiment words; identifying negation words and transition words that have a preset association with sentiment words, calculating the influence value of negation words and transition words on the semantic tendency value based on a preset semantic analysis model, and correcting the semantic tendency value based on the influence value to obtain the target semantic tendency value; inputting the target acoustic features and the target semantic tendency value into a preset matching degree evaluation model to obtain the acoustic semantic matching degree.
[0010] Optionally, based on the emotional feature sequence, identify one or more target time periods in which the target customer expresses their purchase decision, and determine the target customer's car purchase psychological state vector in each target time period. Specifically, this includes: analyzing the semantic content in the emotional feature sequence, identifying target word groups in the semantic content, and determining the target time period in which the target customer expresses their purchase decision based on the time distribution characteristics of the target word groups. The target word groups include price words, configuration words, and purchase intention words in the dialogue text; determining the target customer's expression feature vector in each target time period, including the frequency of use of interjections, the degree of repetition of words, and the duration of response delay; combining the expression feature vector with the emotional intensity value in the target time period to form a state feature vector; inputting the state feature vector into a preset psychological state classification model to obtain the target customer's decision-making psychological state in the target time period; statistically analyzing the frequency and duration of each decision-making psychological state of the target customer in the target time period, weighting and summing the frequency and duration to obtain a comprehensive score for the decision-making psychological state, normalizing the comprehensive score to obtain the probability value of the decision-making psychological state, and constructing a car purchase psychological state vector based on the decision-making psychological state and the corresponding probability value.
[0011] Optionally, the intention decay coefficient of the target customer is determined based on the car purchase psychological state vector, specifically including: determining the decision psychological state with a probability value greater than a preset probability value in the psychological state vector as the target decision psychological state; and calculating the intention decay coefficient by weighting the preset influence weight of each target decision psychological state on the car purchase intention and the probability value of each target decision psychological state.
[0012] Optionally, the purchase intention decay probability curve is analyzed to obtain the target customer's car purchase intention assessment result. Specifically, this includes: determining a first time window and a second time window in the purchase intention decay probability curve; the first time window is the time required for the purchase intention decay probability value to decay from the initial purchase intention decay probability value to a first preset warning threshold; the second time window is the time required for the purchase intention decay probability value to decay from the initial purchase intention decay probability value to a second preset warning threshold; the first preset warning threshold is greater than the second preset warning threshold; the ratio of the duration of the first time window to the second time window is determined as the purchase intention continuation coefficient; when the purchase intention continuation coefficient is greater than the first preset continuation threshold, the car purchase intention assessment result is a first purchase intention level; when the purchase intention continuation coefficient is less than the first preset continuation threshold but greater than or equal to the second preset continuation threshold, the car purchase intention assessment result is a second purchase intention level; when the purchase intention continuation coefficient is less than or equal to the second preset continuation threshold, the target customer's car purchase intention assessment result is a third purchase intention level; the purchase intention at the first purchase intention level is higher than the second purchase intention level, and the purchase intention at the second purchase intention level is higher than the third purchase intention level.
[0013] A second aspect of this application provides a vehicle purchase intention assessment system, comprising: The system comprises the following modules: an acquisition module for acquiring audio recordings of conversations between sales personnel and target customers, and converting the audio into time-stamped text; a first analysis module for performing spectral analysis on the audio recordings to determine the acoustic characteristics of the target customer during the conversation, and constructing an emotional fluctuation curve based on these characteristics; a pairing module for temporally pairing the text and the emotional fluctuation curve to generate an emotional feature sequence, which includes the semantic content of the target customer and the emotional intensity value corresponding to the timestamp of the semantic content in the emotional fluctuation curve; a first determination module for identifying one or more target time periods in which the target customer expresses their purchase decision based on the emotional feature sequence, and determining the target customer's car-buying psychological state vector for each target time period, which includes multiple decision-making psychological states and the probability value of the target customer being in each decision-making psychological state; a second determination module for determining the target customer's willingness decay coefficient based on the car-buying psychological state vector, and inputting the willingness decay coefficient into a preset exponential decay function to obtain a willingness decay probability curve, which characterizes the changing trend of the target customer's car-buying willingness over time; and a second analysis module for analyzing the willingness decay probability curve to obtain the target customer's car-buying willingness assessment result.
[0014] A third aspect of this application provides an electronic device including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, and both the user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method described in any of the foregoing descriptions.
[0015] A fourth aspect of this application provides a computer-readable storage medium storing instructions that, when executed, perform the method described in any of the preceding descriptions.
[0016] In summary, one or more technical solutions provided in this application have at least the following technical effects or advantages: 1. By combining semantic information from dialogue audio with acoustic emotion features, a customer's emotion fluctuation curve is constructed and temporally paired with text semantics. This not only overcomes the limitations of existing technologies that rely solely on keywords and static emotion classification, resulting in limited information dimensions and coarse evaluation, but also dynamically reflects the trajectory of the customer's emotional changes during the dialogue. Furthermore, by identifying emotional and semantic features within key time periods, a customer's car-buying psychological state vector is constructed, quantifying the probability distribution of customers across different psychological states. This achieves a fine-grained shift from "whether or not interested" to "degree of interest and psychological attitude." Based on this, an intention decay coefficient is introduced, and an intention decay probability curve is constructed, enabling mathematical modeling and prediction of changes in customer car-buying intentions. This effectively fills the gap in traditional methods' inability to handle dynamic changes in intentions, improving the accuracy and comprehensiveness of assessing customer car-buying intentions.
[0017] 2. By performing refined segmentation and acoustic analysis of customer speech, the accuracy and robustness of emotion recognition are improved. Specifically, compared to existing technologies that directly extract features from the entire speech segment while ignoring differences in speech structure and semantic context, this approach first divides the dialogue audio into semantically complete target speech segments based on speech pauses. Then, by distinguishing the speaker, the target customer's speech is accurately extracted, avoiding misjudgments caused by salesperson interference. Furthermore, by identifying semantically neutral segments and establishing acoustic baselines based on these segments (including multi-dimensional acoustic parameters such as pitch variation, formant spectral envelope, voice jitter rate, and flicker rate), a reference model conforming to individual customer speech habits is constructed, effectively eliminating fundamental acoustic biases caused by individual differences. Based on this baseline, acoustic variation analysis is then performed on non-neutral semantic segments, enabling more accurate capture of the degree of emotional deviation and the construction of an emotion curve reflecting true psychological fluctuations. Compared to traditional methods that suffer from poor recognition of emotional changes and are susceptible to interference from speech differences, this invention achieves dynamic emotional recognition based on the individual, greatly improving the authenticity, sensitivity, and personalization of emotional fluctuation curves, and further providing high-quality input for the refined evaluation of car purchase intentions.
[0018] 3. By jointly modeling the semantic content and expressive behavior in the emotional feature sequence, the granularity and accuracy of identifying customers' psychological states when purchasing a car are significantly improved. Compared to existing technologies that roughly judge customer intentions based solely on keyword frequency and sentiment classification, this invention first identifies target word groups in the semantic content and, combined with their distribution characteristics on the time axis, accurately locates the key time periods in which customers express their purchase decision intentions, thus avoiding the intent dilution problem caused by global average analysis. Furthermore, within these target time periods, this invention not only extracts the customer's semantic content but also introduces non-semantic expressive features such as the frequency of use of interjections, the degree of repetition of utterances, and the duration of response delays to form an expressive feature vector. This vector is then combined with the corresponding emotional intensity value to construct a state feature vector, comprehensively reflecting the customer's psychological state performance within a specific time period. By inputting the state feature vector into a preset psychological state classification model, this invention can identify multiple car-purchase-related psychological states and further quantify the credibility of each psychological state and normalize it into a probability value through a weighted scoring method based on frequency and duration, forming a car-purchase psychological state vector with probability distribution characteristics. This approach not only enhances the hierarchy and interpretability of psychological state recognition but also enables dynamic quantitative modeling of customer psychological states. Compared to the static judgment results of traditional methods, psychological state vectors can more realistically reflect the emotional and cognitive fluctuations of customers during the car purchase decision-making process, providing actionable psychological data support for subsequent car purchase intention assessment and sales strategy formulation. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the system architecture of an embodiment of a car purchase intention assessment method or a car purchase intention assessment system applied in this application. Figure 2 This is a flowchart illustrating a method for assessing car purchase intention in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a car purchase intention assessment system according to an embodiment of this application; Figure 4 This is a schematic diagram of the structure of the electronic device in the embodiments of this application.
[0020] Explanation of reference numerals in the attached drawings: 301, acquisition module; 302, first analysis module; 303, pairing module; 304, first determination module; 305, second determination module; 306, second analysis module; 401, processor; 402, communication bus; 403, user interface; 404, network interface; 405, memory. Detailed Implementation
[0021] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0022] Figure 1 An exemplary system architecture 100 is shown, which can be applied to an embodiment of a car purchase intention assessment method or a car purchase intention assessment system of this application.
[0023] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables. Network 104 is used to establish communication links between terminal devices and the server, supporting the transmission of various data types, including audio, text, and emotional data. This network can use various wired or wireless communication methods, including but not limited to Wi-Fi, cellular mobile networks (such as 4G and 5G), Ethernet, wired local area networks, or fiber optic communication. Server 105, as the core processing unit of the system, undertakes the entire process from data reception and model inference to the output of car purchase intention assessment results. Server 105 can be a single physical server or a distributed server cluster deployed in the cloud.
[0024] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as model training applications, video recognition applications, web browser applications, and social media platform software. It should be understood that... Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. In particular, if the target data does not need to be obtained remotely, the above system architecture may exclude the network and include only terminal devices or servers.
[0025] Figure 2 This is a flowchart illustrating a method for assessing car purchase intention in an embodiment of this application.
[0026] Please see Figure 2 This application provides a method for assessing car purchase intention, the method comprising: S201. Obtain audio of conversations between sales personnel and target customers, and convert the audio into text with timestamps. In step S201, complete audio data of the conversation between the salesperson and the target customer is collected using a recording device. This audio data forms the foundational information source for subsequent intention assessment processes. Its collection aims to obtain the target customer's vocal behavior characteristics and semantic content in a real-world communication environment, thereby providing raw data support for emotion change analysis, car purchase psychological state modeling, and car purchase intention trend prediction. Since car purchase behavior is often accompanied by complex psychological games and emotional fluctuations, directly extracting these changes from the voice signal is a crucial starting point for achieving high-precision car purchase intention assessment.
[0027] After acquiring the dialogue audio, the original audio is converted into timestamped dialogue text using a speech recognition module. An automatic speech recognition system built with deep learning algorithms can be employed, such as a hybrid model based on convolutional neural networks and long short-term memory networks. This model first preprocesses the audio signal, including noise reduction, framing, and windowing. Then, it extracts Mel-frequency cepstral coefficients as feature input and further decodes the audio using a joint acoustic and language model, outputting precise text content word by word. Each segment of text is labeled with its start and end timestamps within the audio. The purpose of timestamping is to provide a matching basis for subsequent emotion fluctuation curve alignment and emotion feature sequence construction, ensuring that each semantic content can find its corresponding emotion feature at the acoustic level. Specifically, timestamps unify the speech signal and text semantics in the time dimension, ensuring that subsequent emotion change analysis remains context-aware and accurately pinpointing the customer's emotional state when expressing a specific intention. In this embodiment, the generated timestamped dialogue text not only contains semantic information but also serves as a bridge for subsequent emotion recognition and psychological state modeling, ensuring an accurate and traceable mapping between the original audio and semantic analysis.
[0028] S202. Perform spectral analysis on the dialogue audio to determine the acoustic change characteristics of the target customer during the dialogue, and construct the emotional fluctuation curve of the target customer based on the acoustic change characteristics. Step S202, in order to extract key acoustic indicators from the speech signal that effectively reflect the emotional changes of the target customer, requires detailed spectral analysis of the audio of the dialogue between the salesperson and the target customer. Step S202 aims to identify the acoustic change characteristics of the target customer during the communication process, and then construct a dynamic emotional fluctuation curve over time, providing a foundation for subsequent emotional feature sequence construction and psychological state analysis. To ensure the high accuracy and individual adaptability of the extracted emotional features, S202 is further refined into multiple processing steps, including speech segmentation, speaker differentiation, neutral segment identification and its acoustic baseline feature extraction, and acoustic change analysis of non-neutral segments, gradually establishing a dynamic curve reflecting the customer's true emotional fluctuations, thereby achieving a deep perception and modeling of the customer's psychological state when purchasing a car. This may include steps S2021-S2025: S2021. Detect speech pauses in dialogue audio and identify the speech segment between two adjacent speech pauses as the target speech segment. Speech signal processing is performed on the acquired complete dialogue audio to detect speech pauses. The purpose of setting speech pauses is to divide the continuous speech stream into smaller structured units, facilitating subsequent speaker differentiation, semantic analysis, and acoustic feature extraction. Speech pauses are typically characterized by a significant drop in sound energy below the background noise level, lasting for more than a preset time threshold, followed by a silent segment. The preset time threshold is set by comprehensively considering the physiological patterns of speech behavior, the rhythm of language expression, and the statistical characteristics of a large corpus of data, ensuring that the identification of speech pauses conforms to human dialogue habits while possessing good algorithmic adaptability. In this embodiment, short-time energy and short-time zero-crossing rate are used as detection indicators. A sliding window is used to analyze the entire dialogue audio frame by frame, identifying regions where the short-time energy is below the set threshold and the zero-crossing rate changes little, and these regions are marked as speech pauses.
[0029] After detecting speech pauses, the speech signal content between two adjacent pauses is extracted and defined as the target speech segment. This segmentation method ensures that the target speech segment has a natural semantic boundary in time, usually corresponding to a complete semantic expression or a continuous output of speech intonation, with high speech integrity and emotional expressiveness. By dividing long-duration audio into multiple target speech segments, not only is the data complexity of subsequent processing significantly reduced, but the local accuracy of speaker recognition and emotion change recognition is also improved. In addition, since the existence of speech pauses is usually closely related to the speaker's thinking, emotional fluctuations, or semantic segmentation, this step provides a natural segmentation reference in the time dimension for emotion fluctuation modeling, which helps to accurately locate the start and end points and the duration of emotion changes when constructing emotion fluctuation curves in the subsequent process.
[0030] S2022. Speaker differentiation is performed on the target speech segment to obtain the target customer speech segment; The main purpose of speaker differentiation is to eliminate the interference of salespersons' voices on the emotion analysis results and ensure that the construction of the emotion fluctuation curve accurately reflects the acoustic change characteristics of the target customer in the actual dialogue scenario.
[0031] In the specific implementation process, the target speech segment is first processed using speaker detection and speaker clustering techniques. To achieve speaker differentiation, the system needs to construct a speaker voiceprint model for the target customer in advance. A voiceprint model is an individual identification model based on acoustic features, similar to a fingerprint, used to uniquely identify the speech features of a specific speaker. This model extracts low-level acoustic features (such as Mel-frequency cepstral coefficients (MFCC), speech formants, speech energy distribution, etc.) from known speech samples of the target customer and combines them with high-dimensional embedding representation methods such as x-vector or i-vector to generate a speaker vector representation for the customer.
[0032] When processing each target speech segment, the system extracts its voiceprint vector and compares it with the target customer's speaker voiceprint model. If the similarity exceeds a preset threshold, the speech segment is identified as belonging to the target customer and retained; otherwise, it is identified as belonging to a non-target customer and discarded. The preset threshold is typically determined by training and validating the speaker recognition model on a labeled speech dataset containing multiple speakers, and using ROC curve analysis based on the similarity distribution characteristics to select the optimal trade-off between recognition accuracy and false recognition rate. To further improve recognition accuracy, a speaker discrimination model based on a dual-channel neural network can be introduced. This model provides dual inputs to the current speech segment and the customer's voiceprint, and uses a similarity function trained by the neural network for precise discrimination.
[0033] Through the speaker differentiation process described above, the target customer's speech content can be accurately extracted in multi-turn dialogues, providing a clean and controllable data foundation for subsequent neutral semantic segment recognition, acoustic baseline feature extraction, and emotion fluctuation curve construction. In emotion analysis tasks, analyzing only the target customer's speech information is a prerequisite for ensuring that the results are individually targeted and representative of their psychological state. S2023. Identify neutral semantic segments in the target customer's speech segments. Neutral semantic segments are speech segments that meet preset semantic requirements. The preset semantic requirements refer to emotional confidence levels below a set threshold and semantic types conforming to a preset set of neutral semantic rules. In other words, neutral semantic segments refer to speech statements that do not exhibit a clear emotional bias. These typically include responsive language (such as "I understand," "Okay," "I get it"), descriptive statements (such as "My home is a bit far from here"), and information confirmations (such as "You mean ××"), etc., whose semantic content has a low emotional expression density within the conversational context. By extracting acoustic features only from these types of semantic segments, acoustic feature shifts caused by customer emotional fluctuations can be effectively reduced, thereby ensuring that the extracted acoustic baseline features accurately represent the natural pronunciation patterns of the target customer's speech in an emotionally neutral state.
[0034] In the specific implementation process, the target customer's speech segments are first parsed. A natural language processing module is then used to perform syntactic analysis and sentiment semantic determination on the text content after speech recognition. This module, based on a classification model that integrates bidirectional encoder representation or sentiment dictionary fusion, performs semantic sentiment labeling on the text corresponding to each speech segment, outputting a sentiment label and confidence score. If the sentiment confidence score of a semantic segment is lower than a set threshold, and its semantic type conforms to a preset set of neutral semantic rules, it is classified as a neutral semantic segment. This set of neutral semantic rules can be constructed through expert corpus annotation and semantic template extraction, covering common semantic patterns such as objective description, information confirmation, and descriptive observation.
[0035] The identified neutral semantic segments correspond one-to-one with speech segments in the time dimension. The corresponding audio signals are precisely located from the target customer's speech segments using timestamp information. These neutral semantic segments are then used as input for the next step, acoustic baseline feature extraction. Step S2023 filters out speech samples that do not contain significant emotional expression at the semantic level, effectively suppressing the interference of subjective emotions on acoustic features. This provides high-quality data input for building a stable and personalized acoustic baseline and lays an accurate comparative foundation for subsequent identification of abnormal acoustic changes in non-neutral semantic segments. This strategy ensures that the construction of the emotion fluctuation curve is based on the individual's natural speech characteristics, significantly improving the adaptability and accuracy of the overall emotion recognition system.
[0036] S2024. Perform harmonic analysis on the neutral semantic segment to obtain the pitch change curve, perform formant tracking analysis on the neutral semantic segment to obtain the spectral envelope characteristics, perform periodic analysis on the neutral semantic segment to obtain the sound jitter rate and sound flicker rate, and combine the pitch change curve, spectral envelope characteristics, sound jitter rate and sound flicker rate to obtain the acoustic baseline characteristics of the target customer. The acoustic baseline features represent the natural speech patterns of the target customer without significant emotional fluctuations. By comparing the acoustic features of other speech segments with these features, abnormal acoustic changes driven by emotions can be effectively identified, thereby enabling accurate modeling of the target customer's emotional fluctuations.
[0037] In the specific implementation process, harmonic analysis is performed on the audio signal of neutral semantic segments. The core of harmonic analysis lies in extracting the pitch variation curve. Pitch refers to the fundamental frequency of speech, which is usually closely related to the vocal cord vibration rate and is one of the most direct acoustic indicators of emotional expression. By using the autocorrelation function method or the modified Kipstral method, the fundamental frequency of short-time speech frames of neutral semantic segments is estimated to obtain the pitch curve that changes over time, reflecting the vocal cord vibration characteristics and intonation fluctuation patterns of the customer in a stable semantic context.
[0038] Formant tracking analysis is performed on neutral semantic segments to extract spectral envelope features. Formants, also known as "formanteaus," are frequency-enhancing regions determined by the shape of the vocal tract, typically reflecting the configuration of the oral cavity and tongue position during articulation. By analyzing the speech signal using a linear predictive coding method, the frequencies and bandwidths of the first few major formants are extracted, and their changes over time are tracked to obtain a spectral envelope feature sequence. This feature reflects the stability of the articulation path and tongue position adjustment pattern in the client's neutral semantic expression, and is an important vocal tract feature basis for judging emotional deviations.
[0039] Periodic analysis was performed on neutral semantic segments to extract jitter and shimmer. Jitter refers to the small, irregular variation in the length of a period between adjacent syllables, typically used to measure the stability of vocal cord vibration; shimmer refers to the small, irregular variation in the amplitude between syllables, used to assess the consistency of vocal intensity. By analyzing the periodic structure of speech and calculating the average difference ratio between all syllables, jitter and shimmer can be obtained. These two parameters should exhibit low and stable values in a neutral state. Significant increases in non-neutral semantic segments often indicate psychological changes such as emotional excitement, tension, or anxiety.
[0040] Finally, the extracted pitch variation curves, spectral envelope features, sound jitter rate, and sound flicker rate are fused to construct the acoustic baseline features of the target customer. Feature fusion can employ statistical descriptive methods, such as calculating the mean, standard deviation, skewness, and kurtosis for each feature, or it can construct a feature vector sequence through time-series modeling as a stable reference for the individual customer's acoustic state.
[0041] Through the above processing flow, it is possible to achieve comprehensive acoustic feature modeling of target customers under emotionally neutral conditions, providing a solid comparison benchmark for identifying acoustic deviations in non-neutral semantic segments, thereby effectively improving the individual accuracy and dynamic response capability of emotional fluctuation curve construction.
[0042] S2025. Based on the acoustic baseline features, determine the acoustic variation features of non-neutral semantic segments and construct an emotion fluctuation curve. Non-neutral semantic segments are target customer speech segments other than neutral semantic segments.
[0043] In step S2025, the system performs in-depth analysis of non-neutral semantic segments based on the established acoustic baseline features. This aims to identify deviations in acoustic features caused by emotional changes in the target customer during actual conversations, and to construct an emotion fluctuation curve that dynamically reflects the customer's emotional state over time. Non-neutral semantic segments are the target customer's speech segments other than neutral semantic segments. Since non-neutral semantic segments often carry strong subjective emotional expression, their acoustic features may show significant changes compared to the acoustic baseline in terms of pitch, formants, and jitter. Therefore, by systematically comparing the acoustic features of non-neutral semantic segments with the acoustic baseline features, the system can accurately quantify the customer's emotional intensity level in different semantic expression processes and visualize the overall emotional change trend through time series modeling. This may include steps S20251-S20255: S20251. Analyze the target acoustic features of non-neutral semantic segments, calculate the deviation between the target acoustic features and the acoustic baseline features, and obtain the acoustic variation features. In the specific implementation process, for each non-neutral semantic segment, its complete target acoustic features are re-extracted. These target acoustic features include pitch variation curves, spectral envelope features, voice jitter rate, and voice flicker rate. The pitch variation curve is obtained by performing a short-time Fourier transform on the speech signal and then using a fundamental frequency estimation algorithm to extract the trajectory of the fundamental frequency changing over time. The spectral envelope features are obtained by fitting the spectral shape of the speech signal using a linear predictive coding method and extracting the frequencies and bandwidths of the first few formants to characterize changes in vocal tract configuration. The voice jitter rate and voice flicker rate are obtained by statistically calculating the period length and amplitude of continuous sound cycles, respectively, reflecting the periodicity and energy stability of vocal cord vibration.
[0044] The target acoustic features are then compared dimension-by-dimensionally with the acoustic baseline features. The acoustic baseline features are obtained through long-term statistical modeling of acoustic features extracted from neutral semantic segments, typically using mean, standard deviation, maximum, and minimum values to form an individualized benchmark. During the comparison, the system employs multi-dimensional distance metrics, such as Euclidean or Mahalanobis distance, to calculate the feature difference between the current target acoustic features and the baseline features. For temporal features such as pitch variation curves and spectral envelope curves, Dynamic Time Warping (DTW) algorithms are used for alignment before calculating the difference metric to ensure temporal consistency of feature comparison in scenarios with varying speech rates.
[0045] After completing the above comparison, the deviation values between the features of each dimension are combined to form the acoustic variation features of the non-neutral semantic segment. The acoustic variation features not only reflect the overall deviation of the current customer's pronunciation state from its natural speech state, but also retain the deviation information in specific acoustic dimensions, which is helpful for subsequent multi-dimensional emotion pattern recognition and emotion intensity modeling.
[0046] S20252. Calculate the acoustic semantic matching degree for each non-neutral semantic segment. The acoustic semantic matching degree is used to characterize the consistency between the semantic features and acoustic features of the target customer in the non-neutral semantic segment. Step S20252 calculates the acoustic-semantic matching degree to determine the degree of consistency between the semantic content and acoustic emotional expression of the target customer during the expression process. The expression of emotional state is not only reflected in semantic content but also significantly in speech features. When semantics and acoustic features are inconsistent, it often indicates that the speaker is concealing emotions, experiencing psychological repression, or having underlying emotional conflict. Therefore, this step is a crucial prerequisite for accurate emotional intensity calculation. The calculation process of acoustic-semantic matching degree not only considers the emotional tendency of the semantic content but also comprehensively analyzes the synchronization degree between semantic modification relationships and acoustic emotional signals, thereby constructing a comprehensive index that more closely reflects the true emotional state. The process may include the following steps: performing semantic analysis on non-neutral semantic segments; obtaining a semantic feature sequence based on the semantic features of each preset target category word in each non-neutral semantic segment, where the preset target category words include sentiment words, negation words, and transition words; analyzing the sentiment polarity of sentiment words in non-neutral semantic segments to obtain semantic tendency values, where a positive semantic tendency value represents the positivity of the sentiment word, and a negative semantic tendency value represents the negativity of the sentiment word; identifying negation words and transition words that have a preset association with sentiment words; calculating the influence of negation words and transition words on the semantic tendency value based on a preset semantic analysis model; and correcting the semantic tendency value based on the influence value to obtain the target semantic tendency value; and inputting the target acoustic features and the target semantic tendency value into a preset matching degree evaluation model to obtain the acoustic semantic matching degree.
[0047] In its implementation, the system performs semantic analysis on non-neutral semantic segments to obtain their semantic feature sequences. This process includes word segmentation, part-of-speech tagging, and dependency parsing of the speech-to-text text, thereby identifying words in preset target categories, including sentiment words, negation words, and transition words. Sentiment words express subjective attitudes, such as "like," "disappointed," and "satisfied," and are the core basis for semantic emotion judgment. Negation words, such as "no," "not," and "unable," can directionally modify sentiment words. Transition words, such as "but," "however," and "although," often guide changes in semantic focus or emotional reversals. The system constructs a dual-channel fusion model combining dictionary rules and deep learning to model the semantic features of each target word and form a structured semantic feature sequence to express the emotional composition and structural relationships of the current semantic content. For example, if the sentence is "I was originally very satisfied, but later the service attitude disappointed me," the system will identify "satisfied" as a positive sentiment word, "disappointed" as a negative sentiment word, and "but" as a transition word, forming a complete semantic analysis path.
[0048] The system then performs sentiment polarity analysis on the identified sentiment words, calculating preliminary semantic tendency values. Sentiment polarity analysis is based on a combination of sentiment lexicon scoring and contextual neural network modeling. The sentiment lexicon provides static sentiment intensity scores, while the contextual neural network model dynamically adjusts the intensity of sentiment expression according to the context. The final output semantic tendency value represents the direction and intensity of emotion numerically; positive values indicate a positive sentiment tendency, and negative values indicate a negative sentiment tendency. A larger absolute value indicates a more pronounced sentiment expression. This indicator provides a quantitative reference for subsequent acoustic matching. For example, in the statement "I am very satisfied," the sentiment word "satisfied" will be assigned a high positive semantic tendency value, such as +0.85.
[0049] To further improve the accuracy of semantic sentiment analysis, it is necessary to identify negation words and transition words that have semantic modification relationships with sentiment words, and to correct the sentiment polarity. This process relies on a pre-set semantic analysis model, which is based on a graph neural network or a context modeling framework based on a Transformer structure. This model can identify the dependency paths between sentiment words and their modifiers and calculate the influence of negation words and transition words on the original semantic tendency value. For example, the negation word "not" appearing in "I am not satisfied" will reverse the polarity of "satisfied"; the transition word "but" may guide the semantic focus to shift, increasing the importance of the following sentiment word. Finally, a weighted adjustment method is used to obtain the target semantic tendency value, which is closer to the customer's actual semantic sentiment expression. For example, in "I am very satisfied, but the service attitude is not good," the system will combine the relationship between the preceding and following sentiment words and transition words, reduce the weight of "satisfied," and increase the weight of the negative emotion represented by "the service attitude is not good," forming a target semantic tendency value close to -0.7.
[0050] After calculating the target semantic tendency value, the system inputs this value along with the target acoustic features of the current non-neutral semantic segment into a preset matching evaluation model. The preset matching evaluation model is a multimodal matching neural network structure. The input includes acoustic feature vectors and a semantic tendency value scalar. The model calculates the degree of matching between the two through a multi-layered fusion structure and an attention mechanism. A higher matching degree indicates a high degree of consistency between the customer's semantic emotional content and the actual speech emotional signal, making the emotional expression authentic and credible. A lower matching degree suggests a potential emotional conflict or pretense between semantics and acoustics. For example, if a customer says "It's okay," the semantic tendency may be close to neutral or slightly positive, but acoustic features such as a sudden increase in pitch, faster speech rate, and enhanced spectral jitter indicate obvious tension or anger. In this case, the matching model will output a low consistency score (e.g., 0.2), representing a significant deviation between semantics and acoustics.
[0051] S20253. The acoustic change features and acoustic semantic matching degree are weighted and summed to obtain the emotional intensity value of the non-neutral semantic segment. Emotion intensity, as a key indicator for quantifying the degree of emotional expression in a specific speech segment, is directly used to construct an emotion fluctuation curve reflecting the overall dynamic changes in a customer's emotions. Since emotions are not only reflected at the acoustic level but are also modulated by semantic intent, relying solely on acoustic variation features may lead to misjudgments, such as identifying excited but semantically neutral utterances as expressing strong emotions. To improve the accuracy and reasonableness of emotion recognition, acoustic-semantic matching degree needs to be introduced as a moderating factor to provide context-assisted correction for the authenticity of acoustic variations, thereby achieving more robust emotion intensity modeling.
[0052] In practice, the system first receives acoustic variation features and acoustic semantic matching degree from the preprocessing module. Acoustic variation features represent the degree of deviation of the acoustic parameters of the current non-neutral semantic segment from the acoustic baseline features. They are usually represented in the form of a multi-dimensional vector, including dimensions such as pitch shift intensity, spectral structure change, and speech periodic instability. The acoustic semantic matching degree is a scalar with a value ranging from 0 to 1. It is used to measure the consistency between the semantic content and acoustic features in the segment. The closer the value is to 1, the more consistent the acoustic performance is with the semantic emotional tendency, and the more realistic and credible the emotional expression is.
[0053] The system uses a set of adjustable weighting coefficients to fuse the two factors. The basic form of the weighted sum is: Emotion Intensity Value = α × Acoustic Change Intensity + β × Acoustic-Semantic Matching Degree, where α and β are weighting coefficients that control the proportion of influence of acoustic change and semantic consistency in the emotion intensity assessment, respectively. The weighting coefficients can be automatically learned through supervised training during system initialization or manually set according to business needs. For example, in customer service scenarios, to avoid misjudging a customer's verbal agitation but positive semantic expression, the system can appropriately increase the weight of semantic matching degree, making the prediction more consistent with the real communication context.
[0054] During the fusion process, the system first performs dimensionality normalization on the acoustic variation features, compressing the multidimensional acoustic difference vector into a single variation intensity index. Common methods include principal component analysis (PCA) or a weighted synthesis model based on training data. The normalized acoustic variation intensity and the matching degree are input together into a weighted summation module, which executes the aforementioned weighting function and outputs a numerical result representing the emotional intensity of the non-neutral semantic segment. This value can be directly used to construct an emotion fluctuation curve or as input to an emotion classification model to determine the current segment's emotional level or type.
[0055] Through the above steps, the system can distinguish the authenticity of emotional expressions by integrating semantic information while preserving the sensitivity of acoustic signals. This effectively avoids misjudgments caused by inconsistencies between semantic content and acoustic performance, improving the accuracy and stability of the emotion recognition model in practical applications. The emotion intensity value, as a quantitative indicator of emotion expression after the fusion of high-dimensional acoustic and semantic information, is the basic data unit for constructing individualized emotion fluctuation curves.
[0056] S20254. Based on the timestamp information of non-neutral semantic segments, arrange the emotion intensity values in chronological order to generate a discrete emotion intensity sequence. In the specific implementation process, the system first performs timestamp parsing on each non-neutral semantic segment. The timestamp information typically originates from the framing and semantic annotation process of the speech signal in the speech recognition system, recording the start and end times of each semantic segment in the original speech stream. Based on this timestamp information, the system establishes a mapping index, associating each emotion intensity value with its corresponding timestamp, forming a one-to-one correspondence record containing time and emotion intensity. This record can be structured as key-value pairs with timestamps, or as a two-dimensional array, where the first dimension is the time series and the second dimension is the corresponding emotion intensity value, ensuring the data's orderliness in the time dimension.
[0057] Subsequently, the system sorts the emotion intensity values of all non-neutral semantic segments in ascending order according to the time axis, forming a complete discrete emotion intensity sequence. This sequence, with time as the main axis, expresses the changes in the customer's emotion intensity throughout the voice interaction process in the form of discrete points. It not only preserves the emotion expression intensity at each moment, but can also be used to observe the trend, fluctuation rate, and abnormal peaks of emotion changes.
[0058] S20255. Smooth the discrete emotion intensity sequence and use an interpolation algorithm to connect the emotion intensity values at adjacent time points to obtain the emotion fluctuation curve.
[0059] Emotional expression in actual conversations presents different patterns such as gradual enhancement, smooth transition, or sudden intensity. If it is represented only by discrete points, it is difficult to capture the potential emotional continuation or transition intention of customers between adjacent semantic segments. Therefore, it is necessary to introduce interpolation modeling technology to fit the emotional trajectory in time sequence.
[0060] In the specific implementation process, the system traverses the discrete emotion intensity sequence obtained through step S20254, identifies the time interval between two adjacent time points, and generates a fitting curve within the interval according to the set interpolation strategy to fill the emotional gap between non-neutral semantic segments. The interpolation algorithm used can select commonly used mathematical models according to the actual scenario, such as linear interpolation, cubic spline interpolation, or Gaussian kernel smoothing interpolation.
[0061] Linear interpolation is the most basic implementation method. Its principle is to construct a straight line with a constant slope between two known emotional intensity points, which is suitable for contexts where emotional changes are relatively stable. Cubic spline interpolation constructs a series of third-order polynomial curves, so that each curve segment has the continuity of the first and second derivatives at the endpoints, thereby generating a more natural and gentle emotional transition curve, which is suitable for handling voice interaction scenarios with large emotional fluctuations and complex changing trends. Gaussian kernel smoothing interpolation uses each discrete emotional intensity point as the center weight, and generates a smooth curve shape by applying Gaussian function weights to adjacent points and performing weighted averaging. It is suitable for suppressing outlier interference at emotional abrupt change points and enhancing the overall stability and readability of the emotional curve.
[0062] The system automatically selects the optimal interpolation algorithm based on preset parameters or model recommendations, and performs interpolation processing on the entire sequence, ultimately generating a continuous function curve with time as the horizontal axis and emotion intensity as the vertical axis, i.e., an emotion fluctuation curve. This curve can not only be used to display the customer's emotional dynamics in the visualization interface, but also serve as the input basis for subsequent modules to detect emotion inflection points, predict emotion trends, and analyze behavioral intentions.
[0063] S203. Pair the dialogue text with the emotion fluctuation curve in time to generate an emotion feature sequence. The emotion feature sequence includes the semantic content of the target customer and the emotion intensity value of the timestamp corresponding to the semantic content in the emotion fluctuation curve. Because customers' language content and emotional state are often closely coupled when expressing their car purchase needs, and this coupling relationship changes dynamically over time, it is necessary to accurately align the semantic information of the text with the emotional fluctuation curve in the time dimension in order to truly restore the degree of emotional expression of customers in specific semantic segments.
[0064] In practice, the system first acquires the timestamped dialogue text. This text is transcribed from the original audio dialogue using a speech recognition system, with start and end times annotated in each semantic segment, thus forming time-based structured semantic data. On the other hand, the system receives the emotion fluctuation curve constructed in the preceding steps. This curve has undergone acoustic feature extraction, emotion intensity calculation, and interpolation smoothing to generate an emotion intensity function continuously distributed across the entire dialogue timeline. This emotion intensity function uses time as the independent variable and emotion intensity as the dependent variable, enabling it to return the customer's real-time emotional state at any given moment.
[0065] The system then performs a temporal pairing operation, specifically by traversing the timestamp interval of each semantic segment in the dialogue text and retrieving the emotional intensity value within the start and end time of that interval from the emotional fluctuation curve. The system can employ strategies such as sample averaging, maximum value extraction, or median filtering to calculate representative emotional intensity values within the corresponding time range for each semantic segment. Sample averaging, by sampling the emotional curve at equal intervals within that time interval and taking the average, is suitable for scenarios with stable intonation. Maximum value extraction is suitable for capturing emotional peaks in semantic segments, used to identify emotional outburst segments. Median filtering can suppress sudden abnormal fluctuations and enhance the stability of emotional intensity. After the above calculations, the system associates each semantic segment with its corresponding emotional intensity value, forming an emotional feature unit.
[0066] The structure of an emotion feature unit includes: semantic content (i.e., the textual information currently expressed by the customer), a timestamp (marking the start and end times of the semantic segment), and an emotion intensity value (reflecting the customer's emotional activity level during that time period). The system combines all feature units according to the chronological order of the semantic segments in the dialogue to construct a complete emotion feature sequence.
[0067] For example, the customer's voice contains the following three semantic segments and their timestamps: "I think this car looks good" [00:05–00:10]; "However, the price seems a bit high" [00:11–00:16]; "I would consider it if there were a discount" [00:17–00:22]. The system will extract the emotional intensity value for each time period from the emotional fluctuation curve, assuming they are 0.35, 0.67, and 0.60 respectively. The final emotional feature sequence is as follows: ("I think this car looks good", 00:05–00:10, 0.35); ("However, the price seems a bit high", 00:11–00:16, 0.67); ("I would consider it if there were a discount", 00:17–00:22, 0.60).
[0068] Through this sequence, the system not only preserves the customer's complete semantic expression but also simultaneously records the intensity of the emotional state in each expression, providing a precise temporal fusion data foundation for subsequent inferences about the customer's psychological state changes in key car-buying expressions. This processing method greatly enhances the interpretability and predictability of emotion recognition in car-buying intention modeling, enabling the system to understand "how it was said" while recognizing "what was said," achieving deep collaborative modeling of semantics and emotion.
[0069] S204. Identify one or more target time periods when the target customer expresses a purchase decision based on the emotional feature sequence, and determine the target customer's car purchase psychological state vector in each target time period. The car purchase psychological state vector includes multiple decision psychological states and the probability value of the target customer being in each decision psychological state. Step S204 involves in-depth analysis of the emotional feature sequence to identify key time periods during which the target customer expresses their car purchase decision intentions in the dialogue. Within each key time period, multidimensional information reflecting the customer's psychological state is extracted, ultimately constructing a car purchase psychological state vector that characterizes the dynamics of the customer's decision-making process. This vector not only reveals the customer's psychological decision-making patterns within specific time periods but also provides structured input for further assessing the changing trends of the customer's car purchase intentions. Since customers' car purchase decisions are often accompanied by simultaneous changes in specific semantic content and emotional responses, and psychological states have phased and probabilistic characteristics, the system needs to integrate semantic, behavioral, and emotional multimodal information to establish an interpretable and quantifiable psychological state modeling mechanism. The construction of a car purchase psychological state vector may include the following steps: analyzing the semantic content in the emotional feature sequence, identifying target word groups in the semantic content, and determining the target time period for the target customer to express their purchase decision based on the time distribution characteristics of the target word groups. The target word groups include price words, configuration words, and purchase intention words in the dialogue text; determining the expression feature vector of the target customer in each target time period, including the frequency of use of interjections, the degree of repetition of words, and the duration of response delay; combining the expression feature vector with the emotional intensity value in the target time period to form a state feature vector; inputting the state feature vector into a preset psychological state classification model to obtain the decision-making psychological state of the target customer in the target time period; statistically analyzing the frequency and duration of each decision-making psychological state of the target customer in the target time period, weighting and summing the frequency and duration to obtain a comprehensive score for the decision-making psychological state, normalizing the comprehensive score to obtain the probability value of the decision-making psychological state, and constructing a car purchase psychological state vector based on the decision-making psychological state and the corresponding probability value.
[0070] In its implementation, the system first analyzes the semantic content of the emotional feature sequence to identify target word groups, thus determining semantic fragments highly relevant to the customer's car-buying behavior in the dialogue. Target word groups refer to words or phrases with key semantic implications in expressing car-buying intent, primarily including price terms (such as "how much," "too expensive," "are there any discounts?"), configuration terms (such as "navigation," "engine," "interior"), and purchase intention terms (such as "want to buy," "think about it," "can place an order"). The system extracts all target word groups by constructing a domain dictionary and a context-nested recognition model, and statistically analyzes their temporal distribution in the dialogue text. If multiple target word groups appear concentratedly within a time period, and the corresponding emotional intensity value is significantly higher than the average level of the dialogue, then that time period is considered the target time period for expressing a car-buying decision. For example, when the system detects the statement "If the price could be lower, I would buy it," it marks the time period in which it appears as a potential decision-making window.
[0071] After identifying the target time period, the system further extracts customer expression feature vectors within that time period to characterize the customer's psychological performance at the level of vocal behavior. The expression feature vectors include three key dimensions: frequency of use of interjections, degree of repetition of utterances, and response delay duration. The frequency of use of interjections reflects whether the customer tends to hesitate, emphasize, or emotionally embellish their expression; common interjections include "um," "ah," "actually," and "that's right." The degree of repetition of utterances refers to the frequency with which the customer repeatedly uses synonyms or sentence structures within the same semantic unit, reflecting their hesitation or emphasis during the decision-making process. Response delay duration refers to the time interval between the salesperson's question and the customer's response; a longer interval usually indicates that the customer is making psychological trade-offs. The system quantifies these features through voice frame-level alignment and keyword statistics mechanisms, forming a structured vector of expressive behavior.
[0072] The system fuses the expression feature vector with the emotion intensity value extracted during the time period to form a multi-dimensional state feature vector, thereby unifying the modeling of semantic behavioral information and emotional change information. This state feature vector serves as input to a pre-defined psychological state classification model, which is a deep classification model based on a multilayer perceptron (MLP) or dual-channel attention network structure. This model outputs the customer's car-buying psychological state within the current time period based on the input state feature vector. Car-buying psychological states include multiple categories such as hesitation, interest, resistance, acceptance, and rejection. The system labels the time period based on the model's prediction results, thus obtaining the customer's psychological state label for that time period.
[0073] To further quantify the proportion of each psychological state among customers within the target time period, the system statistically analyzes the frequency of occurrence and corresponding duration of each psychological state, and then weights and merges these two data points to generate a comprehensive score for each psychological state. Frequency of occurrence measures the activity level of the state within the target time period, while duration reflects the stability of the state's performance. The system uses a weighted summation method to generate the final score. In this embodiment, to further quantify the proportion of each psychological state of the customer within the target time period, the system statistically analyzes the frequency and duration of each psychological state within that time period and generates a comprehensive score for that psychological state through a weighted summation method. The weighted model adopts a linear fusion function in the form of: Comprehensive Score = Frequency × First Weight + Duration × Second Weight. To ensure that the fusion result reflects both the stability of the psychological state over time and its activity level in behavioral performance, the system determines the weight allocation of the frequency and duration items based on business needs or model training results. For example, under empirical rules, the first weight can be set to 0.4 and the second weight to 0.6 to enhance the impact of persistence on the total score. In scenarios with historical data, the weights can be optimized through supervised learning to better reflect actual customer behavior patterns. All comprehensive scores are then normalized to obtain the probability value of the customer being in each psychological state within the target time period. By combining all psychological states with their corresponding probability values, the system constructs a complete vector of car-buying psychological states.
[0074] S205. Determine the willingness decay coefficient of the target customer based on the car purchase psychological state vector, and input the willingness decay coefficient into the preset exponential decay function to obtain the willingness decay probability curve. The willingness decay probability curve is used to characterize the changing trend of the target customer's car purchase willingness over time. In this embodiment, the core objective of step S205 is to further quantify the changing trend of the customer's car purchase intention throughout the entire dialogue process based on the car purchase psychological state vector constructed in the previous steps. Specifically, this is achieved by calculating the intention decay coefficient and using this coefficient as an input variable, which is then substituted into a preset exponential decay function to generate an intention decay probability curve reflecting the decreasing trend of the customer's car purchase intention over time. Since the customer's car purchase decision psychology is dynamically volatile, and their actual car purchase intention may show different degrees of decreasing or maintaining trends when expressing different psychological states such as hesitation, rejection, or interest, it is necessary to introduce the intention decay coefficient as an intermediary variable to quantify the change in intention, thereby establishing a mathematical mapping model of behavioral trends. Through this curve, the system can model the intensity change of the customer's car purchase intention in the time dimension, thereby providing predictive support for subsequent sales strategy optimization. This may include the following steps: determining the decision psychological state with a probability value greater than a preset probability value in the psychological state vector as the target decision psychological state; and calculating the intention decay coefficient by weighting the preset influence weight of each target decision psychological state on the car purchase intention and the probability value of each target decision psychological state.
[0075] In its implementation, the system first analyzes the psychological state vector for car purchase, identifying the most influential psychological states as the target decision-making psychological states. The judgment is based on the probability value of each state in the psychological state vector. If the probability value of a psychological state exceeds a preset probability threshold (e.g., 0.3 or 0.4), then that state is considered to have actual dominant value in influencing the customer's car purchase intention within the current time period, and is thus included in subsequent weight calculations. This threshold can be set based on historical customer behavior data or automatically optimized through model parameter tuning to ensure the representativeness of the identified target decision-making psychological states.
[0076] After screening the target psychological states, the system assigns a preset influence weight to each target decision-making psychological state. This weight represents the degree of positive or negative impact of the psychological state on the customer's car purchase intention. For example, the weight of the "interest" state might be -0.2, indicating that it has a maintaining or even enhancing effect on the car purchase intention; the weight of the "hesitation" state might be +0.3, indicating that it has a certain degree of attenuation effect on the intention; and the weight of the "rejection" state might be +0.7, meaning a strong attenuation tendency. The system calculates a weighted average of the probability value of each target psychological state and its corresponding influence weight, using the following expression: Intention Attenuation Coefficient = Σ(State Probability Value × State Influence Weight). The weighted summation process uniformly measures the influence of all representative psychological states on the customer's car purchase intention within the current time period, generating a continuous numerical intention attenuation coefficient. The larger the value, the more significant the downward trend in the customer's intention; conversely, the smaller the value, the more stable or slightly enhanced the intention. The system inputs the willingness decay coefficient as a preset exponential decay function. The exponential decay function typically takes the form: P(t) = e^(-λt), where λ is the willingness decay coefficient, t is time, and the independent variable represents the duration of time after the customer's interaction with the salesperson. The output value P(t) represents the probability of purchasing intention at any given time point. Through this function, the system can obtain a complete willingness decay probability curve, thus depicting the decreasing trend of a customer's purchase intention over time. This curve can not only be used to determine whether a customer still has the potential to purchase a car, but also to identify the willingness threshold, assisting sales personnel in deciding whether to follow up promptly or change their communication strategy.
[0077] In summary, step S205 effectively achieves a continuous transformation from psychological state recognition to behavioral intention modeling by filtering and weighting significant psychological states in the car purchase psychological state vector, generating a willingness decay coefficient, and further generating a visualized car purchase willingness decay curve through exponential function mapping.
[0078] S206. Analyze the probability curve of intention decay to obtain the car purchase intention assessment results of the target customers.
[0079] Since the probability curve of purchase intention decay reflects the dynamic trend of target customers' car purchase intention changing over time, the system can further determine the persistence and decay rate of customer intention by identifying key time nodes and rates of change in the curve, and classify different levels of purchase intention accordingly. To achieve structured and standardized evaluation results, the system introduces a purchase intention persistence coefficient as a core discriminant. This coefficient is derived by comparing the reach time window lengths of different preset warning thresholds, quantifying the tolerance time for customers to transition from a high-intention state to a low-intention state, thus providing actionable intention level labels for sales strategy formulation. The process may include the following steps: determining a first time window and a second time window in the willingness decay probability curve, wherein the first time window is the time required for the willingness decay probability value to decay from the initial willingness decay probability value to a first preset warning threshold, and the second time window is the time required for the willingness decay probability value to decay from the initial willingness decay probability value to a second preset warning threshold, wherein the first preset warning threshold is greater than the second preset warning threshold; determining the ratio of the duration of the first time window to the second time window as the willingness continuation coefficient; when the willingness continuation coefficient is greater than the first preset continuation threshold, the car purchase willingness assessment result is a first willingness level; when the willingness continuation coefficient is less than the first preset continuation threshold but greater than or equal to the second preset continuation threshold, the car purchase willingness assessment result is a second willingness level, wherein the first preset continuation threshold is greater than the second preset continuation threshold; when the willingness continuation coefficient is less than or equal to the second preset continuation threshold, the target customer's car purchase willingness assessment result is a third willingness level, wherein the car purchase willingness of the first willingness level is higher than the second willingness level, and the car purchase willingness of the second willingness level is higher than the third willingness level.
[0080] In practical implementation, the system iterates through the willingness decay probability curve to determine two key time windows: the first time window and the second time window. The first time window refers to the time elapsed from the initial willingness value at the start of the curve until the curve declines to the first preset warning threshold. The second time window is the time elapsed from the initial willingness value to the second preset warning threshold, which is the final willingness threshold for the target customer. The first preset warning threshold is an intermediate target value higher than the final willingness threshold, representing the critical point where customer willingness begins to decline significantly but has not yet completely churned. It is used to identify potential churn risks in advance. The first and second preset warning thresholds can be objectively determined based on statistical analysis of historical customer data or preset by business experts based on industry experience. The division of these two time windows helps the system distinguish between the initial changes in customer willingness decline and the final tendency to abandon the business, thus depicting the rate of decline in customer willingness in a more granular way.
[0081] After obtaining these two time windows, the system calculates the ratio of the first time window's duration to the second time window's duration to obtain the intention continuation coefficient. This coefficient reflects the relationship between the customer's persistence in the early stages of intention decline and the speed of intention decay. If the first time window is significantly shorter than the second time window, it indicates that the customer's intention wavers quickly, belonging to a rapid decline pattern, and the intention continuation coefficient is small. If the two time windows are close in duration or the first window accounts for a larger proportion, it indicates that the customer's intention decline is relatively slow, the intention is maintained for a longer period, and the continuation coefficient is large.
[0082] The system compares the customer's intention continuation coefficient with a preset continuation threshold to output a customer's car purchase intention assessment level. If the continuation coefficient is greater than the first preset continuation threshold, the assessment result is level one, indicating that the customer's car purchase intention is strong and relatively stable. If the continuation coefficient is less than the first preset continuation threshold but greater than or equal to the second preset continuation threshold, the assessment is level two. Here, the first preset continuation threshold being greater than the second preset continuation threshold indicates that the customer's intention exists but is showing some degree of wavering. If the continuation coefficient is less than the second preset continuation threshold, it is classified as level three, indicating that the customer's intention is declining rapidly and the likelihood of purchasing a car is low. The first and second preset continuation thresholds can be obtained through clustering or quantile analysis of historical customer intention continuation coefficients and actual car purchase results. Alternatively, they can be heuristically set by business experts based on customer behavior experience in the initial stage and dynamically optimized through data iteration during system operation. This grading method maps continuous trends in intention changes to discrete level labels, enhancing the readability and business feasibility of the analysis results.
[0083] In summary, by performing window analysis and ratio calculation on the probability curve of intention decay, the system can effectively extract the rate of change characteristics of customers' car purchase intention, and perform standardized level classification through the intention continuation coefficient, realizing the transformation from dynamic trend modeling to static evaluation labels, and providing sales personnel with accurate and operable customer intention evaluation results.
[0084] Please see Figure 3 This is a schematic diagram of a car purchase intention assessment system provided in an embodiment of this application. The car purchase intention assessment system 300 specifically includes: The acquisition module 301 is used to acquire the audio of the conversation between the salesperson and the target customer, and convert the audio into text with timestamps. The first analysis module 302 is used to perform spectral analysis on the audio to determine the acoustic change characteristics of the target customer during the conversation, and construct the target customer's emotional fluctuation curve based on the acoustic change characteristics. The matching module 303 is used to perform temporal matching between the text and the emotional fluctuation curve to generate an emotional feature sequence, which includes the semantic content of the target customer and the emotional intensity value of the timestamp corresponding to the semantic content in the emotional fluctuation curve. The first determination module 304 is used to identify the target customer based on the emotional feature sequence. The system expresses one or more target time periods for the purchase decision and determines the target customer's car purchase psychological state vector in each target time period. The car purchase psychological state vector includes multiple decision psychological states and the probability value of the target customer being in each decision psychological state. The second determination module 305 is used to determine the target customer's willingness decay coefficient based on the car purchase psychological state vector, and inputs the willingness decay coefficient into a preset exponential decay function to obtain a willingness decay probability curve. The willingness decay probability curve is used to characterize the changing trend of the target customer's car purchase willingness over time. The second analysis module 306 is used to analyze the willingness decay probability curve to obtain the target customer's car purchase willingness assessment result.
[0085] Optionally, the first analysis module 302 is specifically used for: detecting speech pauses in the dialogue audio, identifying the speech segment between two adjacent speech pauses as the target speech segment; distinguishing the speaker in the target speech segment to obtain the target customer speech segment; identifying neutral semantic segments in the target customer speech segment, where the neutral semantic segment is a speech segment that meets preset semantic requirements; performing harmonic analysis on the neutral semantic segment to obtain a pitch change curve; performing formant tracking analysis on the neutral semantic segment to obtain spectral envelope features; performing periodic analysis on the neutral semantic segment to obtain voice jitter rate and voice flicker rate; combining the pitch change curve, spectral envelope features, voice jitter rate, and voice flicker rate to obtain the acoustic baseline features of the target customer; determining the acoustic change features of non-neutral semantic segments based on the acoustic baseline features, and constructing an emotion fluctuation curve, where the non-neutral semantic segments are the target customer speech segments other than the neutral semantic segments.
[0086] Optionally, the first analysis module 302 is further specifically used for: analyzing the target acoustic features of non-neutral semantic segments, calculating the deviation between the target acoustic features and the acoustic baseline features to obtain acoustic variation features; calculating the acoustic semantic matching degree of each non-neutral semantic segment, the acoustic semantic matching degree being used to characterize the consistency between the semantic features and acoustic features of the target customer in the non-neutral semantic segment; weighting and summing the acoustic variation features and the acoustic semantic matching degree to obtain the emotional intensity value of the non-neutral semantic segment; arranging the emotional intensity values in chronological order according to the timestamp information of the non-neutral semantic segment to generate a discrete emotional intensity sequence; smoothing the discrete emotional intensity sequence, and using an interpolation algorithm to connect the emotional intensity values at adjacent time points to obtain an emotional fluctuation curve.
[0087] Optionally, the first analysis module 302 is further specifically used for: performing semantic analysis on non-neutral semantic segments; obtaining a semantic feature sequence based on the semantic features of each preset target category word in each non-neutral semantic segment, wherein the preset target category words include sentiment words, negation words, and transition words; analyzing the sentiment polarity of sentiment words in non-neutral semantic segments to obtain semantic tendency values, wherein a positive semantic tendency value is used to characterize the positive degree of sentiment words, and a negative semantic tendency value is used to characterize the negative degree of sentiment words; identifying negation words and transition words that have a preset association with sentiment words; calculating the influence value of negation words and transition words on the semantic tendency value based on a preset semantic analysis model; and correcting the semantic tendency value based on the influence value to obtain the target semantic tendency value; and inputting the target acoustic features and the target semantic tendency value into a preset matching degree evaluation model to obtain the acoustic semantic matching degree.
[0088] Optionally, the first determining module 304 is specifically used for: analyzing the semantic content in the emotional feature sequence, determining the target word groups in the semantic content, and determining the target time period for the target customer to express a purchase decision based on the time distribution characteristics of the target word groups. The target word groups include price words, configuration words, and purchase intention words in the dialogue text; determining the expression feature vector of the target customer in each target time period, the feature vector including the frequency of use of interjections, the degree of repetition of speech, and the response delay duration; combining the expression feature vector with the emotional intensity value in the target time period to form a state feature vector; inputting the state feature vector into a preset psychological state classification model to obtain the decision-making psychological state of the target customer in the target time period; statistically analyzing the occurrence frequency and duration of each decision-making psychological state of the target customer in the target time period, weighting and summing the occurrence frequency and duration to obtain a comprehensive score of the decision-making psychological state, normalizing the comprehensive score to obtain the probability value of the decision-making psychological state, and constructing a car purchase psychological state vector based on the decision-making psychological state and the probability value corresponding to the decision-making psychological state.
[0089] Optionally, the second determining module 305 is specifically used to: determine the decision psychological state with a probability value greater than a preset probability value in the psychological state vector as the target decision psychological state; and calculate the intention decay coefficient by weighting the preset influence weight of each target decision psychological state on the car purchase intention and the probability value of each target decision psychological state.
[0090] Optionally, the second analysis module 306 is specifically used for: determining a first time window and a second time window in the willingness decay probability curve, wherein the first time window is the time required for the willingness decay probability value to decay from the initial willingness decay probability value to the first preset warning threshold, and the second time window is the time required for the willingness decay probability value to decay from the initial willingness decay probability value to the second preset warning threshold, and the first preset warning threshold is greater than the second preset warning threshold; determining the ratio of the duration of the first time window to the second time window as the willingness continuation coefficient; when the willingness continuation coefficient is greater than the first preset continuation threshold, the car purchase willingness assessment result is the first willingness level; when the willingness continuation coefficient is less than the first preset continuation threshold but greater than or equal to the second preset continuation threshold, the car purchase willingness assessment result is the second willingness level, and the first preset continuation threshold is greater than the second preset continuation threshold; when the willingness continuation coefficient is less than or equal to the second preset continuation threshold, the target customer's car purchase willingness assessment result is the third willingness level, and the car purchase willingness of the first willingness level is higher than the second willingness level, and the car purchase willingness of the second willingness level is higher than the third willingness level.
[0091] It should be noted that the device provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions.
[0092] This embodiment also discloses an electronic device, as shown in the reference. Figure 4The electronic device may include: at least one processor 401, at least one communication bus 402, a user interface 403, a network interface 404, and at least one memory 405. The communication bus 402 is used to enable communication between these components. The user interface 403 may include a display screen or a camera; optionally, the user interface 403 may also include a standard wired interface or a wireless interface. The network interface 404 may optionally include a standard wired interface or a wireless interface. The processor 401 may include one or more processing cores. The processor 401 connects to various parts of the server using various interfaces and lines, and performs various functions of the server and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 405, and by calling data stored in the memory 405. Optionally, the processor 401 may be implemented using at least one hardware form selected from digital signal processing, field-programmable gate arrays, and programmable logic arrays. The memory 405 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 405 may include a non-transitory computer-readable medium. The memory 405 can be used to store instructions, programs, code, code sets, or instruction sets.
[0093] The foregoing description is merely an exemplary embodiment of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Other embodiments of this disclosure will be readily apparent to those skilled in the art upon consideration of the disclosure in this specification. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.
Claims
1. A method for assessing car purchase intention, characterized in that, The method includes: Acquire audio recordings of conversations between sales personnel and target customers, and convert the audio recordings into time-stamped text. Spectral analysis is performed on the dialogue audio to determine the acoustic change characteristics of the target customer during the dialogue, and an emotion fluctuation curve of the target customer is constructed based on the acoustic change characteristics. The dialogue text is matched with the emotion fluctuation curve in time sequence to generate an emotion feature sequence. The emotion feature sequence includes the semantic content of the target customer and the emotion intensity value of the timestamp corresponding to the semantic content in the emotion fluctuation curve. Based on the emotional feature sequence, identify one or more target time periods in which the target customer expresses a purchase decision, and determine the car purchase psychological state vector of the target customer in each target time period. The car purchase psychological state vector includes multiple decision psychological states and the probability value of the target customer being in each decision psychological state. Based on the car purchase psychological state vector, the intention decay coefficient of the target customer is determined, and the intention decay coefficient is input into a preset exponential decay function to obtain the intention decay probability curve. The intention decay probability curve is used to characterize the changing trend of the target customer's car purchase intention over time. By analyzing the probability curve of the decline in willingness, the car purchase intention assessment results of the target customer are obtained; The step of performing spectral analysis on the dialogue audio to determine the acoustic change characteristics of the target customer during the dialogue, and constructing the emotional fluctuation curve of the target customer based on the acoustic change characteristics, specifically includes: detecting speech pauses in the dialogue audio, identifying the speech segment between two adjacent speech pauses as the target speech segment; distinguishing the speaker of the target speech segment to obtain the target customer speech segment; identifying neutral semantic segments in the target customer speech segment, the neutral semantic segment being a speech segment that meets preset semantic requirements; performing harmonic analysis on the neutral semantic segment to obtain a pitch change curve; performing formant tracking analysis on the neutral semantic segment to obtain spectral envelope features; performing periodic analysis on the neutral semantic segment to obtain voice jitter rate and voice flicker rate; combining the pitch change curve, the spectral envelope features, the voice jitter rate, and the voice flicker rate to obtain the acoustic baseline features of the target customer; determining the acoustic change characteristics of non-neutral semantic segments based on the acoustic baseline features, and constructing the emotional fluctuation curve, the non-neutral semantic segments being the target customer speech segments other than the neutral semantic segments; The process of determining the acoustic variation features of non-neutral semantic segments based on the acoustic baseline features and constructing the emotion fluctuation curve specifically includes: analyzing the target acoustic features of the non-neutral semantic segments, calculating the deviation between the target acoustic features and the acoustic baseline features to obtain the acoustic variation features; calculating the acoustic semantic matching degree of each non-neutral semantic segment, whereby the acoustic semantic matching degree characterizes the consistency between the semantic features and acoustic features of the target customer in the non-neutral semantic segments; weighting and summing the acoustic variation features and the acoustic semantic matching degree to obtain the emotion intensity value of the non-neutral semantic segment; arranging the emotion intensity values in chronological order according to the timestamp information of the non-neutral semantic segments to generate a discrete emotion intensity sequence; smoothing the discrete emotion intensity sequence and using an interpolation algorithm to connect the emotion intensity values at adjacent time points to obtain the emotion fluctuation curve.
2. The method according to claim 1, characterized in that, The calculation of the acoustic semantic matching degree for each of the non-neutral semantic segments specifically includes: Semantic analysis is performed on the non-neutral semantic segments, and a semantic feature sequence is obtained based on the semantic features of each preset target category word in each non-neutral semantic segment. The preset target category words include sentiment words, negation words, and transition words. The emotional polarity of the emotional words in the non-neutral semantic segment is analyzed to obtain a semantic tendency value. When the semantic tendency value is positive, it is used to characterize the positive degree of the emotional word, and when the semantic tendency value is negative, it is used to characterize the negative degree of the emotional word. Identify the negative words and transition words that have a preset association with the sentiment words, calculate the influence value of the negative words and transition words on the semantic tendency value based on the preset semantic analysis model, and correct the semantic tendency value based on the influence value to obtain the target semantic tendency value; The target acoustic features and target semantic tendency values are input into a preset matching degree evaluation model to obtain the acoustic semantic matching degree.
3. The method according to claim 1, characterized in that, The step of identifying one or more target time periods in which the target customer expresses a purchase decision based on the emotional feature sequence, and determining the target customer's car purchase psychological state vector in each target time period, specifically includes: The semantic content in the emotional feature sequence is analyzed to determine the target word groups in the semantic content. Based on the time distribution characteristics of the target word groups, the target time period for the target customer to express the purchase decision is determined. The target word groups include price words, configuration words and purchase intention words in the dialogue text. Determine the expression feature vector of the target customer in each target time period, the feature vector including the frequency of use of interjections, the degree of repetition of speech, and the duration of response delay; The expression feature vector is combined with the emotion intensity value in the target time period to form a state feature vector; The state feature vector is input into a preset psychological state classification model to obtain the decision-making psychological state of the target customer in the target time period. The frequency and duration of each decision-making psychological state of the target customer within the target time period are statistically analyzed. The frequency and duration are weighted and summed to obtain a comprehensive score for the decision-making psychological state. The comprehensive score is normalized to obtain a probability value for the decision-making psychological state. The car purchase psychological state vector is constructed based on the decision-making psychological state and the probability value corresponding to the decision-making psychological state.
4. The method according to claim 1, characterized in that, The determination of the target customer's willingness decay coefficient based on the car-buying psychological state vector specifically includes: The decision-making psychological state whose probability value in the psychological state vector is greater than a preset probability value is determined as the target decision-making psychological state; The intention decay coefficient is obtained by weighting the preset influence weight of each target decision psychological state on the car purchase intention and the probability value of each target decision psychological state.
5. The method according to claim 1, characterized in that, The analysis of the intention decay probability curve yields the car purchase intention assessment results for the target customer, specifically including: A first time window and a second time window are determined in the willingness decay probability curve. The first time window is the time required for the willingness decay probability value to decay from the initial willingness decay probability value to the first preset warning threshold. The second time window is the time required for the willingness decay probability value to decay from the initial willingness decay probability value to the second preset warning threshold. The first preset warning threshold is greater than the second preset warning threshold. The ratio of the duration of the first time window to that of the second time window is determined as the willingness continuation coefficient; When the intention continuation coefficient is greater than the first preset continuation threshold, the car purchase intention assessment result is the first intention level; When the intention continuation coefficient is less than the first preset continuation threshold and greater than or equal to the second preset continuation threshold, the car purchase intention assessment result is the second intention level, and the first preset continuation threshold is greater than the second preset continuation threshold; When the intention continuation coefficient is less than or equal to the second preset continuation threshold, the target customer's car purchase intention assessment result is the third intention level. The car purchase intention of the first intention level is higher than the second intention level, and the car purchase intention of the second intention level is higher than the third intention level.
6. A car purchase intention assessment system, characterized in that, include: The acquisition module is used to acquire audio of conversations between sales personnel and target customers, and convert the audio into text with timestamps. The first analysis module is used to perform spectral analysis on the dialogue audio, determine the acoustic change characteristics of the target customer during the dialogue, and construct the emotional fluctuation curve of the target customer based on the acoustic change characteristics. The matching module is used to perform time-series matching between the dialogue text and the emotion fluctuation curve to generate an emotion feature sequence. The emotion feature sequence includes the semantic content of the target customer and the emotion intensity value of the timestamp corresponding to the semantic content in the emotion fluctuation curve. The first determining module is used to identify one or more target time periods in which the target customer expresses a purchase decision based on the emotional feature sequence, and to determine the car purchase psychological state vector of the target customer in each target time period. The car purchase psychological state vector includes multiple decision psychological states and the probability value of the target customer being in each decision psychological state. The second determining module is used to determine the intention decay coefficient of the target customer based on the car purchase psychological state vector, and input the intention decay coefficient into a preset exponential decay function to obtain an intention decay probability curve. The intention decay probability curve is used to characterize the changing trend of the target customer's car purchase intention over time. The second analysis module is used to analyze the intention decay probability curve to obtain the car purchase intention assessment result of the target customer; The first analysis module is specifically used to detect speech pauses in the dialogue audio, determine the speech segment between two adjacent speech pauses as the target speech segment; perform speaker differentiation on the target speech segment to obtain the target customer speech segment; identify neutral semantic segments in the target customer speech segment, wherein the neutral semantic segments are speech segments that meet preset semantic requirements; perform harmonic analysis on the neutral semantic segments to obtain pitch variation curves; perform formant tracking analysis on the neutral semantic segments to obtain spectral envelope features; perform periodic analysis on the neutral semantic segments to obtain voice jitter rate and voice flicker rate; and combine the pitch variation curve, the spectral envelope features, the voice jitter rate, and the voice flicker rate to obtain the acoustic baseline features of the target customer. Based on the acoustic baseline features, the acoustic variation features of non-neutral semantic segments are determined, and the emotion fluctuation curve is constructed. The non-neutral semantic segments are the target customer speech segments other than the neutral semantic segments. The first analysis module is further specifically used to analyze the target acoustic features of the non-neutral semantic segment, calculate the deviation between the target acoustic features and the acoustic baseline features, and obtain the acoustic change features; Calculate the acoustic semantic matching degree for each of the non-neutral semantic segments, wherein the acoustic semantic matching degree is used to characterize the consistency between the semantic features and acoustic features of the target customer in the non-neutral semantic segments; The acoustic change features and the acoustic semantic matching degree are weighted and summed to obtain the emotion intensity value of the non-neutral semantic segment; based on the timestamp information of the non-neutral semantic segment, the emotion intensity values are arranged in chronological order to generate a discrete emotion intensity sequence; the discrete emotion intensity sequence is smoothed, and an interpolation algorithm is used to connect the emotion intensity values at adjacent time points to obtain the emotion fluctuation curve.
7. An electronic device, characterized in that, include: One or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-5.
8. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-5.