Sales agent voice dialogue method
By using temporal context frame association parsing and emotional dynamic spectrum quantitative modeling, combined with hierarchical feature denoising reconstruction and product attribute knowledge graph, customized response text is generated and emotional tone is adapted, solving the problem of lost context and emotional information in intelligent voice dialogue systems, and achieving accurate identification of user needs and natural adaptation of responses.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING DECK SMART TECH CO LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, intelligent voice dialogue systems have failed to effectively mine the dynamic features of context and emotional state in speech in scenarios such as e-commerce, financial insurance, and automobile sales. This results in responses that are difficult to match the user's real needs and emotional inclinations, leading to biases in demand recognition and insufficient adaptability of responses.
By parsing temporal context frames and quantizing emotional dynamic spectrum, temporal context frame features and emotional dynamic spectrum features are generated. Combined with hierarchical feature denoising and reconstruction and intent posterior probability matching, a scenario-based demand positioning feature topology is generated. The product attribute knowledge graph is called to perform weighted association mapping of product attribute features to generate customized recommendation response text. Finally, real-time response speech stream is generated through text-to-speech conversion and emotional tone adaptation modulation.
It achieves accurate identification of user needs and natural adaptation of responses, improves communication efficiency and user experience, and solves the problems of need identification deviation and response mismatch caused by the loss of context and emotional information in traditional solutions.
Smart Images

Figure CN121963707A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent voice interaction technology, and more specifically, to a sales agent voice dialogue method. Background Technology
[0002] In highly interactive sales scenarios such as e-commerce, financial insurance, and automobile sales, intelligent voice dialogue systems have become a key carrier connecting enterprises and users. Their core requirement is to accurately capture users' dynamic needs and output appropriate response content, thereby improving communication efficiency and sales conversion, and meeting the dual demands of enterprises for digital service upgrades and convenient user interaction.
[0003] In existing technologies, a typical sales agent voice dialogue scheme converts user speech into text information using automatic speech recognition technology, then constructs an intent recognition library based on preset keyword matching rules, and finally calls a fixed product information knowledge base to generate response text and output it through speech synthesis technology. This scheme first performs simple semantic translation of the speech signal, then matches user intent according to static rules, and finally feeds back fixed response content in a standardized speech form.
[0004] However, the existing technical solution only focuses on the semantic parsing and static intent matching of speech signals, without effectively mining and quantifying the dynamic features of the context and emotional state in the speech. This results in the generated response content being difficult to match the user's real needs and emotional tendencies, leading to problems such as demand identification bias and insufficient response adaptability, which restricts the improvement of service quality and commercial value of the sales intelligent agent voice dialogue system. Summary of the Invention
[0005] To address the aforementioned technical problems, this application provides a sales agent voice dialogue method to at least alleviate the aforementioned technical problems.
[0006] The technical solutions provided in this application are as follows: A sales agent voice dialogue method includes: Step 1, performing temporal context frame association parsing and emotional dynamic spectrum quantization modeling on real-time voice data streams representing user needs to generate temporal context frame features and emotional dynamic spectrum features, thereby forming a dynamic feature construct of user needs; Step 2, performing hierarchical feature denoising reconstruction and intent posterior probability matching on the dynamic feature construct of user needs to determine the user's core need feature basis vector, and performing scene feature anchoring mapping based on the user's core need feature basis vector to assign scene feature identifier codes and generate a scene-based need positioning feature topology; Step 3, calling a product attribute knowledge graph to perform weighted association mapping of product attribute features and recommendation semantic logic generation on the scene-based need positioning feature topology to generate customized recommendation response text; Step 4, based on the customized recommendation response text, fusing temporal context frame features and emotional dynamic spectrum features to perform text-to-speech conversion and emotional tone adaptation modulation to generate a real-time response voice stream, wherein the real-time response voice stream includes real recommendation content matched with product attributes to match user needs.
[0007] The technical solution provided in this application has the following technical advantages: Step 1 performs temporal context frame association parsing and emotional dynamic spectrum quantization modeling on the real-time speech data stream representing user needs, generating temporal context frame features, emotional dynamic spectrum features, and a dynamic feature construct of user needs. This process overcomes the limitations of traditional solutions that only perform semantic translation of speech. It captures the dynamic changes in context within speech through temporal context frame association parsing and mines user emotional state-related features through emotional dynamic spectrum quantization modeling. The user need dynamic feature construct formed by fusing these two approaches fully preserves the contextual and emotional information in the speech, laying a data foundation for accurate user need identification and solving the problem of need identification deviation caused by the loss of key contextual and emotional information in traditional solutions. Step 2, based on the dynamic feature construct of user needs, generates a contextualized need localization feature topology through hierarchical feature denoising and reconstruction, intent posterior probability matching, and scene feature anchoring mapping. Hierarchical feature denoising and reconstruction filters out irrelevant interference information, optimizes feature quality, and makes subsequent intent matching more targeted. Intent posterior probability matching determines the user's core need feature basis vector through quantitative calculation, which is more scientific and accurate than traditional static keyword matching. Scene feature anchoring mapping associates core needs with specific sales scenarios. The generated scenario-based need positioning feature topology clarifies the user's need in a specific scenario, providing clear guidance for subsequent accurate recommendations and effectively improving the shortcomings of traditional solutions in terms of rigid intent recognition and poor scenario adaptability. Step 3 calls the product attribute knowledge graph to perform weighted association mapping of product attribute features and generate recommendation semantic logic on the scenario-based need positioning feature topology to obtain customized recommendation response text. The product attribute knowledge graph provides rich product attribute data support for accurate matching. Weighted association mapping can highlight the correspondence between key needs and product attributes. Combining the scenario-based need positioning feature topology to construct recommendation semantic logic and organize text, the generated response text not only fits the product attributes but also accurately matches the user's scenario-based needs, avoiding the problem of traditional solution response text being out of touch with user needs and improving the targeting of the response. Step 4, based on customized recommended response text, integrates temporal contextual frame features and emotional dynamic spectrum features to perform text-to-speech conversion and emotional intonation adaptation modulation, generating a real-time response speech stream. This step reuses the temporal contextual frame features and emotional dynamic spectrum features generated in Step 1, enabling speech conversion to not only complete the text-to-speech transformation but also adapt to the user's contextual rhythm and emotional state. The generated real-time response speech stream is more natural and adaptable, solving the problems of mechanical and stiff speech output and lack of emotional adaptation in traditional solutions, thus improving the user communication experience.
[0008] This application utilizes temporal context frame association parsing and emotional dynamic spectrum quantization modeling, combined with hierarchical denoising and probability matching, to fully capture contextual, emotional, and core demand information in speech, making demand recognition more comprehensive and accurate. In terms of response adaptability, this application generates customized text through scenario-based demand localization, and then integrates contextual and emotional features for speech modulation, achieving dual adaptation of text content and speech form, significantly improving the appropriateness of responses. Regarding overall communication effectiveness, this application, through the integration of context and emotion throughout the entire process, makes responses more accurate and natural, effectively improving the user communication experience. Attached Figure Description
[0009] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.
[0010] Figure 1 This is a flowchart of a sales agent voice dialogue method according to an embodiment of the present invention. Detailed Implementation
[0011] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. Various aspects are provided by way of explanation and not limitation of the invention. Indeed, those skilled in the art will recognize that modifications and variations can be made to the invention without departing from its scope or spirit. For example, a feature represented or described as part of one embodiment may be used in another embodiment to produce yet another embodiment. Therefore, it is desirable that the invention encompass such modifications and variations falling within the scope of the appended claims and their equivalents.
[0012] In the description of this invention, the terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," and "bottom," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and do not require the invention to be constructed and operated in a specific orientation; therefore, they should not be construed as limitations on the invention. The terms "connected," "linked," and "set up" used in this invention should be interpreted broadly. For example, they can refer to a fixed connection or a detachable connection; they can refer to a direct connection or an indirect connection through intermediate components. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.
[0013] like Figure 1As shown, this invention provides a voice dialogue method for a sales agent, comprising: Step 1, performing temporal context frame association parsing and emotional dynamic spectrum quantization modeling on real-time voice data streams representing user needs to generate temporal context frame features and emotional dynamic spectrum features, thereby forming a dynamic feature construct of user needs; Step 2, performing hierarchical feature denoising reconstruction and intent posterior probability matching on the dynamic feature construct of user needs to determine the user's core need feature basis vector, and performing scene feature anchoring mapping based on the user's core need feature basis vector to assign scene feature identifier codes and generate a scene-based need positioning feature topology; Step 3, calling a product attribute knowledge graph to perform weighted association mapping of product attribute features and recommendation semantic logic generation on the scene-based need positioning feature topology to generate customized recommendation response text; Step 4, based on the customized recommendation response text, fusing temporal context frame features and emotional dynamic spectrum features to perform text-to-speech conversion and emotional tone adaptation modulation to generate a real-time response voice stream, wherein the real-time response voice stream includes real recommendation content matched with product attributes to match user needs.
[0014] Optionally, step 1 includes: step 11, performing frame segmentation and feature pre-extraction on the real-time speech data stream representing user needs to generate a speech frame feature sequence; step 12, performing adjacent frame contextual association analysis and cross-frame semantic mapping on the speech frame feature sequence to generate temporal contextual frame features; step 13, performing emotional dimension decomposition and dynamic amplitude modeling on the speech frame feature sequence to generate emotional dynamic spectrum features, and fusing the temporal contextual frame features and emotional dynamic spectrum features to form a dynamic feature profile of user needs. Optionally, step 11 includes: step 111, performing fixed-length frame segmentation on the real-time speech data stream representing user needs to generate a speech frame sequence; step 112, extracting Mel-frequency cepstral coefficients for each speech frame in the speech frame sequence to generate single-frame features; step 113, sequentially integrating all single-frame features to generate a speech frame feature sequence.
[0015] Preferably, the specific implementation process of step 111 is as follows: Based on the noise characteristics, speech activity features, and speech rhythm features of the sales scenario, a frame segmentation adaptation parameter system is constructed for the real-time speech data stream representing user needs, so as to generate scene-adaptive frame segmentation control parameters; specifically, taking the real-time speech data stream representing user needs as the processing object, combined with the typical environmental characteristics of the sales scenario (the background noise in the store is mostly low-frequency interference of 500-1000Hz, which is easy to superimpose with the low-frequency components of speech), the speech activity features of the user's needs expression (the effective speech energy is usually higher than the noise energy by a certain range, and has distinguishability) and speech syllables Based on the characteristics of speech patterns (rapid pronunciation of keywords, high syllable density, and smooth, continuous descriptive sentences), a frame segmentation adaptation parameter system for sales scenarios is constructed, including preprocessing parameters, speech segmentation determination parameters, and frame length allocation parameters. The preprocessing parameters cover a bandpass filter frequency range encompassing the core frequency band of human voice (300-3400Hz, which contains the main pronunciation components of Chinese speech), and a noise suppression strength coefficient (0.6-0.8, determined based on statistical analysis of noise intensity in sales scenarios; too high a coefficient can easily lead to speech distortion, while too low a coefficient cannot effectively reduce noise) that balances noise reduction effect and speech fidelity. The speech segmentation determination parameter package... This includes short-time energy threshold coefficients (1.5-2.0, adapted to the energy ratio of speech and noise in sales scenarios) and zero-crossing rate threshold coefficients (1.2-1.4, matching the syllable switching frequency of Chinese speech). Frame length allocation parameters are designed for different speech segment characteristics: high-frequency changing segment frame lengths (10-15 milliseconds, matching the time dimension features of rapid keyword pronunciation) for capturing keyword details; stable segment frame lengths (20-25 milliseconds, balancing sentence integrity and feature extraction efficiency) for preserving descriptive sentence features; and invalid segment frame lengths (30-40 milliseconds, reducing the processing of silence and noise segments) for reducing ineffective computation. (overhead), and finally generate scene-adaptive frame segmentation control parameters; preferably, in the specific technical implementation of step 111, multi-stage preprocessing and segmentation judgment are performed on the real-time voice data stream representing user needs based on the frame segmentation control parameters to generate voice segmentation results; specifically, firstly, bandpass filtering is performed on the real-time voice data stream representing user needs, and low-frequency noise and high-frequency interference are filtered through the 300-3400Hz bandpass filter in the frame segmentation control parameters, and interference components in non-human voice frequency bands are removed to generate a preliminary filtered voice data stream; then, based on the noise suppression strength coefficient (0.6-0.5) in the frame segmentation control parameters.8) Perform spectral subtraction on the initial filtered speech data stream. By subtracting the corresponding coefficient of the noise spectrum from the speech spectrum, the background noise and speech signal are further separated to generate a denoised speech data stream. Then, perform short-time energy calculation on the denoised speech data stream. Use a 10-millisecond sliding window (balancing time resolution and computational efficiency) to calculate the sum of squares of the signal within the window, quantifying the speech frame energy intensity to generate a short-time energy sequence. Simultaneously, perform zero-crossing rate calculation on the denoised speech data stream, counting the number of positive and negative signal alternations within the window to reflect the speech syllable switching characteristics, to generate a zero-crossing rate sequence. Finally, compare the short-time energy sequence with "global average energy × short-time energy threshold coefficient in frame segmentation control parameters," and compare the zero-crossing rate sequence with "..." The global average zero-crossing rate is compared with the zero-crossing rate threshold coefficient in the frame segmentation control parameters. Regions that meet both comparison conditions are determined as valid speech segments, and the remaining regions are determined as silent or noisy segments to generate speech segmentation results. This process solves the segmentation error problem that is prone to occur when the traditional fixed threshold is used when the scene noise fluctuates by customizing the scene parameters, thus improving the robustness of speech segmentation. Preferably, in a scenario, when step 111 is specifically implemented, differential frame segmentation and sequence normalization are performed based on the speech segmentation results and frame segmentation control parameters to generate a temporally normalized speech frame sequence. Specifically, the energy change rate of the valid speech segments in the speech segmentation results is first calculated, and the energy difference between adjacent sliding windows and the energy of the previous window is used to calculate the energy change rate. The ratio of quantities is used to quantify the drastic changes in speech energy to generate an energy change rate sequence. Based on this sequence, regions with a change rate higher than 0.3-0.4 are labeled as high-frequency change segments (corresponding to the pronunciation regions of keywords such as "pixels," "battery life," and "price," which have high information density), while regions with a change rate lower than 0.3 are labeled as stable segments (corresponding to descriptive phrases such as "I need" and "suitable for home use," which have semantic coherence), generating speech segment classification results. Then, based on frame segmentation control parameters and the speech segment classification results, high-frequency change segments are segmented at a fixed length of 10-15 milliseconds to ensure no loss of keyword details, while stable segments are segmented at a fixed length of 20-25 milliseconds to balance features. To balance integrity and data volume, silent and noisy segments are segmented into fixed-length frames of 30-40 milliseconds to reduce the cost of processing invalid data. A 25%-30% frame overlap rate is also set (to avoid losing frame boundary information and improve feature continuity) to generate a set of segmented speech frames. Finally, the segmented speech frames are arranged chronologically, and each frame is assigned a time sequence number (identifying its position in the original speech), frame duration (recording the frame segment length), and frame energy (characterizing the speech intensity within the frame) to generate a temporally regular speech frame sequence. This sequence, through differentiated frame length design, achieves a balance between capturing keyword details and preserving descriptive sentence features in sales scenarios, laying the foundation for subsequent feature extraction. Preferably, the specific implementation process of step 112 is as follows: Based on the speech pronunciation characteristics of the sales scenario, the distribution of high-frequency words, and the requirements for feature extraction, a dedicated Mel-Frequency Cepstral Coefficients (MFCC, a speech feature extraction method that simulates the characteristics of human hearing) extraction parameter set is constructed for the time-ordered speech frame sequence to generate scenario-specific MFCC extraction control parameters; specifically, taking the time-ordered speech frame sequence as the processing object, combined with the speech pronunciation characteristics of the sales scenario (high-frequency components of keyword consonants are easily attenuated, such as the initial consonant pronunciation energy of words like "fast" and "clear" concentrated in 2000-3000Hz, which are easily masked by noise), the frequency distribution of high-frequency word pronunciation (the core frequency of product attribute words and scenario description words is concentrated in 1000-2000Hz, such as the main vowel pronunciation frequency of words like "pixel," "battery life," and "capacity"), the requirements for feature condensation and temporal change capture, and the requirements for keyword feature enhancement, a sales scenario-specific MFCC extraction parameter set including preprocessing parameters, filter bank parameters, transform parameters, and enhancement parameters is constructed, where the preprocessing parameters are... Processing parameters include a pre-emphasis coefficient (0.93-0.97, to compensate for high-frequency consonant attenuation through first-order high-pass filtering) to enhance the amplitude of high-frequency components, and an energy normalization range (0-1, to unify speech frames with different energy levels to the same magnitude, improving feature comparability) to eliminate energy differences caused by the distance between the user and the device. Filter bank parameters include the total number of filters (24-32, optimized based on the speech spectrum distribution in sales scenarios), one filter every 50-80Hz in the key frequency range (1000-2000Hz) (to improve frequency resolution in the key range), and one filter every 100-150Hz in the non-key ranges (300-1000Hz, 2000-3400Hz) (to balance resolution and computational load). Transformation parameters include 12th-16th order Discrete Cosine Transforms (DCTs) to condense core features. Cosine Transform (DCT, which removes correlations between feature dimensions) and first- and second-order difference coefficients (reflecting the trend of feature changes between speech frames) are used to capture temporal changes. Enhancement parameters include feature matching threshold (0.7, determined based on the statistical accuracy of keyword feature library matching) and weight gain coefficient (1.1-1.3, appropriately amplifying the contribution of keyword features) to generate scenario-based MFCC extraction control parameters. Preferably, in the specific technical implementation of step 112, multi-stage feature extraction is performed on the temporally regular speech frame sequence based on the MFCC extraction control parameters to generate a basic cepstral coefficient sequence. Specifically, the pre-emphasis coefficient (0.93-0.93) in the MFCC extraction control parameters is first applied to each speech frame in the speech frame sequence.97) First-order high-pass filtering is performed to compensate for the amplitude attenuation of high-frequency components of keyword consonants, improving the recognizability of keywords such as "pixel" and "fast charging" to generate pre-emphasized speech frames. Minimum-maximum normalization is then performed on the pre-emphasized speech frames, mapping the frame energy to the 0-1 energy normalization range in the MFCC extraction control parameters to eliminate energy fluctuations caused by differences in user distance and vocal intensity, thus generating energy-normalized speech frames. Based on the filter bank parameters in the MFCC extraction control parameters, a non-uniformly distributed Mel filter bank is constructed (the filter center frequencies are distributed according to the Mel scale to simulate the differences in human ear sensitivity to different frequencies). The spectrum of the energy-normalized speech frames is filtered, focusing on preserving... This process retains the characteristic information of the frequency region where keywords are concentrated, suppresses interference from non-critical frequency components, and generates a Mel-filtered energy sequence. A natural logarithm operation is then performed on the Mel-filtered energy sequence to compress the spectral dynamic range and improve robustness to noise. Simultaneously, multiplication operations are converted to addition operations to simplify subsequent processing, generating a logarithmic energy sequence. Finally, an MFCC extraction process is performed on the logarithmic energy sequence to extract the 12th-16th order DCT transform from the control parameters, removing correlations between spectral dimensions and extracting the core coefficients that best represent the essence of speech, thus generating a basic cepstral coefficient sequence. This process, through scenario-based filter bank design, improves the feature recognition of core vocabulary in sales scenarios, distinguishing it from the filtering methods used in traditional MFCC extraction. A general design with uniformly distributed components; preferably, in a scenario, when step 112 is specifically implemented, feature enhancement and dimension concatenation are performed based on the basic cepstral coefficient sequence and MFCC-extracted control parameters to generate scene-enhanced single-frame features; specifically, first-order difference calculation is performed on the basic cepstral coefficient sequence, and the difference between adjacent cepstral coefficients is used to capture short-term temporal changes in speech frames, reflecting the dynamic features of keyword pronunciation, to generate a first-order difference coefficient sequence; second-order difference calculation is performed on the first-order difference coefficient sequence, and the difference between adjacent first-order difference coefficients is used to capture long-term temporal trends in speech frames, reflecting the prosodic features at the sentence level, to generate a second-order difference coefficient sequence; In the sales scenario keyword feature library (containing standard cepstral coefficient templates of core words such as "pixel", "battery life", and "price", which are generated by statistical modeling of a large number of sales scenario voice samples), the cosine similarity between the basic cepstral coefficient sequence and each standard template is calculated to quantify the matching degree between the current frame and the keyword, so as to generate a feature matching degree sequence. The feature matching degree sequence is compared with the feature matching threshold (0.7) in the MFCC extraction control parameters. For the basic cepstral coefficient sequence, first-order difference coefficient sequence and second-order difference coefficient sequence with a matching degree higher than the threshold, each element is multiplied by the weight gain coefficient (1.1-1) in the MFCC extraction control parameters.3) Amplify the contribution of keyword features to generate an enhanced coefficient set; concatenate the basic cepstral coefficients (representing static speech features), first-order difference coefficients (representing short-term dynamic features), and second-order difference coefficients (representing long-term dynamic features) in the enhanced coefficient set according to their dimensions to form a 36-48 dimensional feature vector. Each dimension corresponds to a specific acoustic feature indicator (e.g., the first dimension corresponds to the fundamental frequency feature of speech, the fifth dimension corresponds to the keyword consonant feature, etc.) to generate scene-enhanced single-frame features. These scene-enhanced single-frame features combine static acoustic features and dynamic temporal features, improving the accuracy of representing user needs in sales scenarios. Preferably, the specific implementation process of step 113 is as follows: Based on the requirements of sales scenario feature saliency, semantic importance, and temporal consistency, a temporal integration parameter system is constructed for the scene-enhanced single-frame features to generate scene-adaptive feature integration control parameters; specifically, taking the scene-enhanced single-frame features as the processing object, and combining the requirements of sales scenario feature saliency (effective demand feature variance is higher than noise feature, possessing distinguishability), speech frame semantic importance features (effective speech segment core region has higher energy, keywords are concentrated), and feature temporal consistency requirements (speech features change continuously over time, and temporal correlation must be maintained), a temporal feature integration parameter system is constructed, including feature selection parameters, weight allocation parameters, and sequence construction parameters, wherein the features... The selection parameters include a feature variance threshold coefficient (0.3-0.4, determined based on the variance ratio of effective features to noise features in the sales scenario), weight allocation parameters including a high weight coefficient (0.8-1.0, assigned to feature frames in the keyword set), a medium weight coefficient (0.5-0.7, assigned to feature frames in descriptive statements), and a low weight coefficient (0.1-0.2, assigned to feature frames in silent and noisy segments), and sequence construction parameters including a sequence dimension normalization range (36-48 dimensions, matching the feature dimensions of a single frame in scene enhancement) and a time axis alignment precision (millisecond level, ensuring temporal accuracy), to generate scene-adaptive feature integration control parameters; preferably, in the specific technical implementation of step 113, based on the feature integration control parameters and The speech segmentation results are used to filter and weight scene-enhanced single-frame features to generate a set of effective features with weighted labels. Specifically, the variance of all scene-enhanced single-frame features is first calculated to obtain the variance value of each scene-enhanced single-frame feature (variance reflects the discriminative power of the feature; the higher the variance, the better the feature can distinguish different speech content), thus generating a single-frame feature variance sequence. The global average variance of the single-frame feature variance sequence is calculated, and based on the feature variance threshold coefficient (0.3-0.4) in the feature integration control parameters, a feature filtering threshold (global average variance × feature variance threshold coefficient) is determined to distinguish effective features from noise features. Each variance value in the single-frame feature variance sequence is compared with the feature filtering threshold, and the variance values are retained. Scene-enhanced single-frame features with variance values above a threshold (these features have high discriminative power) are selected, while invalid features with variance values below a threshold (such as features corresponding to silent frames and noisy frames) are removed to generate an effective feature set. Based on the speech segmentation results and frame energy information obtained in step 111, weights are assigned to each scene-enhanced single-frame feature in the effective feature set: scene-enhanced single-frame features in the effective speech segments with energy values 1.2-1.5 times higher than the global average energy (features that are likely to correspond to keywords or core needs) are assigned high weight coefficients (0.8-1.0) in the feature integration control parameters, while the remaining scene-enhanced single-frame features in the effective speech segments (features corresponding to descriptive statements) are assigned medium weight coefficients (0.5-0.0) in the feature integration control parameters.7) For invalid segments that were not removed, assign low weight coefficients (0.1-0.2) to the feature integration control parameters to generate a set of effective features with weights. This weight allocation avoids the problem of key demand information being diluted by invalid features in the sales scenario. Preferably, in a scenario, when implementing step 113, temporal integration and dimensional regularization are performed based on the feature integration control parameters, the temporally regularized speech frame sequence, and the set of effective features with weights to generate a speech frame feature sequence. Specifically, based on the timeline order of the speech frame sequence generated in step 111, the effective features with weights are first... The feature set is sorted to ensure temporal consistency (restoring the temporal order of the speech and ensuring semantic coherence) to generate a temporally ordered feature set. For each effective feature with a weight label in the temporally ordered feature set, the values are updated according to the method of "scene enhancement single-frame feature dimension value × corresponding weight coefficient" to strengthen the influence of core features and generate a weighted fusion feature set. The weighted fusion feature set is then concatenated dimensionally to construct a two-dimensional feature matrix. The row dimensions of this matrix correspond to the temporal order of the features (each row represents a weighted fusion feature at a time point, and the row index corresponds one-to-one with the temporal number of the speech frame). The column dimensions correspond to the feature dimensions (36-48 dimensions, each column corresponding to an acoustic feature index). The elements at the intersection of rows and columns represent the specific values of the weighted fusion features (comprehensively reflecting the semantic contribution of the corresponding time point and the corresponding feature dimension). Based on the sequence dimension normalization range (36-48 dimensions) in the feature integration control parameters, the two-dimensional feature matrix undergoes dimension normalization: if the feature dimension is less than 36, it is padded to 36 dimensions using zero-padding (ensuring feature dimension uniformity); if the feature dimension is greater than 48, principal component analysis is used (extracting the core dimensions that best represent the semantics and removing redundant information). The first 48 core features are extracted to generate a normalized feature matrix. Finally, the normalized feature matrix is expanded row-wise to generate a temporally coherent, dimensionally normalized, and key feature-highlighted speech frame feature sequence. This speech frame feature sequence fully preserves the temporal structure and core acoustic features of the user's demand speech, providing accurate and efficient feature input for the subsequent user demand analysis module in sales scenarios. Through scenario-based parameter customization, keyword feature enhancement, and weight allocation design, it exhibits significant technical differences from traditional speech feature extraction methods, solving the problems of poor scenario adaptability and lack of emphasis on key demand information in traditional methods.
[0016] Optionally, step 12 includes: step 121, extracting contextual elements from individual frame features in the speech frame feature sequence to generate intra-frame contextual features; step 122, performing semantic correlation calculation and feature fusion on the intra-frame contextual features of adjacent frames to generate inter-frame correlation features; and step 123, performing cross-frame temporal mapping and feature integration on the inter-frame correlation features to generate temporal contextual frame features.
[0017] Preferably, the specific implementation process of step 121 is as follows: taking a single frame feature in the speech frame feature sequence as the processing object, based on the sales scenario-specific contextual element system and feature dimension mapping rules, the single frame feature is subjected to structured contextual element extraction and integration to generate intra-frame contextual features; specifically, the single frame feature is the 36-48 dimension scene-enhanced single frame feature output in step 11. Combining the characteristics of voice interaction in the sales scenario (user needs mostly revolve around "product attributes - functional demands - decision-making tendencies", such as "large storage + smooth gaming + moderate price"), four types of scenario-based contextual elements are designed: core attribute elements (corresponding to product core parameter keywords such as "storage", "screen", "battery", etc.), functional scenario elements (corresponding to usage scenario descriptions such as "gaming", "taking photos", "office", etc.), demand intensity elements (corresponding to emotional association features such as voice energy, pronunciation speed, etc.), and semantic association elements (corresponding to logical features connecting previous and subsequent frames). A contextual element-feature dimension mapping matrix is constructed, where the row dimension of the matrix is the four types of contextual elements, and the column dimension is the single frame feature. The system comprises 36-48 dimensions, with core attribute elements mapped to dimensions 6-14 (keyword consonant feature concentration area), functional scene elements mapped to dimensions 15-22 (semantic description feature area), demand intensity elements mapped to dimensions 23-27 (energy and speech rate feature area), and semantic association elements mapped to dimensions 28-32 (inter-frame connection feature area). Dimensional analysis and element extraction are performed on individual frame features. Core attribute elements are extracted using keyword matching (based on a product parameter feature library for sales scenarios; if a match is successful, the feature value is retained; otherwise, it is set to 0). Functional scene elements are encoded using semantic tags (matching a keyword library of scenarios and labeling it with 1-5 levels of scenario relevance). Demand intensity elements are integrated using normalization (mapping energy and speech rate features to the 0-1 range and fusing them with a weight of 0.7:0.3). Semantic association elements are extracted using trend prediction (predicting connection probability based on frame temporal position). This ultimately generates 32-38 dimensional intra-frame contextual features. This feature, through multi-dimensional contextual structured integration, differs from traditional single-feature extraction and achieves a comprehensive representation of intra-frame semantics. Preferably, the specific implementation process of step 122 is as follows: taking intra-frame context features as the processing object, based on the sales scenario semantic association priority rules and dynamic window mechanism, hierarchical association calculation and adaptive fusion are performed on the intra-frame context features of adjacent frames to generate inter-frame association features; specifically, based on the sales scenario sentence length statistics (Chinese short sentences contain an average of 3-5 speech frames), a dynamic association window is set (default 3 frames, adjusted to 2 frames if a semantic pause is detected). The window contains the intra-frame context features of the current frame and the adjacent frames before and after; hierarchical association calculation is performed on the intra-frame context features in the window according to the element type. The matching degree of core attribute elements is calculated using cosine similarity (quantifying the consistency of product parameter description), the overlap degree of functional scenario elements is calculated using Jaccard coefficient (quantifying the association of usage scenarios), and the coherence of demand intensity elements is calculated using Pearson correlation coefficient (quantifying the trend of emotional change). The semantic association elements are conditionally probabilistically calculated to determine the degree of coherence (quantifying logical consistency). Based on the semantic association priority in the sales scenario (core attributes > functional scenarios > demand intensity > semantic association), dynamic weight coefficients are assigned (core attributes 0.4-0.5, functional scenarios 0.3-0.35, demand intensity 0.1-0.15, semantic association 0.05-0.1), and weighted summation is used to obtain the inter-frame semantic association degree. A association degree threshold is set (0.65-0.75, determined based on sales scenario voice sample statistics). For strongly associated frames with an association degree higher than the threshold, a weighted average fusion is used (the higher the association degree, the greater the weight). For weakly associated frames with an association degree lower than the threshold, feature splicing fusion is used (preserving their respective contextual characteristics) to generate 48-52 dimensional inter-frame association features. These features are calculated through dynamic windows and hierarchical association, solving the problem of traditional fixed window fusion ignoring semantic boundaries. Preferably, in the specific technical implementation of step 122, the inter-frame correlation features are taken as the processing object. Based on the sales scenario redundant feature judgment rules and core feature enhancement strategies, the inter-frame correlation features are redundantly eliminated and optimized to generate optimized inter-frame correlation features. Specifically, variance analysis is performed on the initially generated inter-frame correlation features to calculate the variance value of each dimension feature. Core dimensions with variance values higher than the threshold (0.35-0.45, determined based on the variance statistics of effective correlation features) are retained, and redundant dimensions with variance values lower than the threshold (such as duplicate product attributes) are eliminated. (Description and undifferentiated scene annotation); For core dimensional features, weighted enhancement is performed based on their corresponding semantic relevance. Dimensions with a relevance higher than 0.75 are multiplied by an enhancement coefficient of 1.25-1.35 (to amplify the contribution of core semantics), while dimensions with a relevance lower than 0.55 are multiplied by a weakening coefficient of 0.75-0.85 (to reduce secondary semantic interference). Finally, a feature selection algorithm (based on the semantic importance ranking of sales scenarios) is used to retain 42-46 core features, generating optimized inter-frame association features. This process improves the semantic recognition and computational efficiency of the features. Preferably, the specific implementation process of step 123 is as follows: taking the optimized inter-frame association features as the processing object, and based on the sales scenario temporal context modeling rules and attention mechanism cross-frame mapping algorithm, the inter-frame association features are globally integrated temporally to generate temporal context frame features; specifically, a cross-frame temporal mapping matrix is first constructed, with the matrix row dimension being 42-46 dimensions of the optimized inter-frame association features, and the column dimension being the temporal step size (5-7 steps based on the temporal characteristics of the sales scenario sentences, covering complete semantic units), and the matrix elements representing the semantic contribution of the corresponding dimension features at the corresponding temporal step size; temporal position encoding is performed on the inter-frame association features, and temporal weights of 0.15-0.9 are assigned according to the temporal order of the frames in the sentence (the weight of core semantic frames is higher than that of edge frames); let A sales scenario-specific attention mechanism is designed to calculate the attention weights of each inter-frame correlation feature and other temporal position features. The weight calculation is combined with scene semantic rules (the attention weight of product attribute features is higher than that of scene description features, and the attention weight of core requirement frames is higher than that of auxiliary description frames). An attention temperature parameter is set (1.1-1.3, which controls the concentration of weight distribution). Based on the attention weights, adaptive weighted integration is performed to fuse inter-frame correlation features from different temporal positions into a unified feature vector. Finally, normalization processing is performed (mapping to the 0-1 interval) to generate 55-60 dimensional temporal context frame features. This feature fully preserves cross-frame temporal context information, realizes global semantic representation of voice requirements in the sales scenario, and provides a high-quality feature foundation for subsequent requirement analysis.
[0018] Optionally, step 13 includes: step 131, performing emotion-related feature separation on the speech frame feature sequence to generate original emotion features; step 132, performing multi-dimensional decomposition and amplitude change tracking on the original emotion features to generate dynamic emotion amplitude data; step 133, performing spectral modeling on the dynamic emotion amplitude data to generate dynamic emotion spectrum features, and fusing temporal context frame features and dynamic emotion spectrum features to form a dynamic feature profile of user needs.
[0019] Preferably, the specific implementation process of step 131 is as follows: taking the speech frame feature sequence as the processing object, based on the sales scenario emotional association feature system and separation matrix, the speech frame feature sequence is accurately separated into emotional and semantic features to generate the original emotional features; specifically, the speech frame feature sequence is the 36-48 dimension scene-enhanced single-frame feature sequence output in step 11. Combining the characteristics of user emotional expression in the sales scenario (when users consult about products, emotions are mostly reflected through changes in speech speed, volume, and tone, such as urgent needs corresponding to fast speech speed and high volume), three types of core emotional association features in the sales scenario are defined: speech speed association features (corresponding to speech frame time interval and syllable density), volume association features (corresponding to speech energy amplitude and energy fluctuation), and tone association features (corresponding to frequency change trend and fundamental frequency shift). An emotional-semantic feature separation matrix is constructed, with the matrix row dimension being the speech frame feature sequence. The 36-48 dimensions of the feature are labeled as "emotional features" and "semantic features". Among them, speech rate is associated with dimensions 8-13 (frame time interval feature area), volume is associated with dimensions 14-19 (energy feature area), and tone is associated with dimensions 20-25 (frequency feature area). The remaining dimensions are labeled as semantic features. Based on the separation matrix, the speech frame feature sequence is separated frame by frame. The feature projection separation algorithm (projecting the feature vector to the emotional feature subspace) is used to retain the dimension values of the three types of emotional features and remove the interference of the semantic feature dimension. The separated features are normalized (mapped to the 0-1 interval to eliminate the amplitude difference between frames) and finally the 18-22 dimension original emotional features are generated. This feature is defined and separated in a contextualized emotional dimension, which is different from the traditional generalized emotional extraction and realizes the targeted capture of emotional features in the sales scenario. Preferably, the specific implementation process of step 132 is as follows: taking the original emotional features as the processing object, based on multi-dimensional emotional amplitude tracking rules and dynamic windows, the original emotional features are decomposed in layers and tracked in real time amplitude changes to generate emotional dynamic amplitude data; specifically, the original emotional features are first decomposed into three sub-feature sequences according to three types of emotional association features: speech rate sub-feature sequence, volume sub-feature sequence, and tone sub-feature sequence, each sub-feature sequence having 6-8 dimensions; a dynamic tracking window is set for each sub-feature sequence (based on the emotional change cycle statistics of sales scenario sentences, the window size is set to 3-5 frames to ensure complete capture of emotional fluctuations), and amplitude change tracking is performed within the window: the speech rate sub-feature sequence uses time difference tracking. The algorithm (calculating the difference in syllable intervals between adjacent frames to quantify the rate of change in speech rate) uses energy amplitude difference tracking for the volume sub-feature sequence (calculating the absolute difference in energy amplitude between adjacent frames to reflect volume fluctuations) and frequency offset tracking for the intonation sub-feature sequence (calculating the offset of the fundamental frequency between adjacent frames to reflect intonation changes). The tracking results are then timestamped (marking amplitude change nodes according to frame temporal position) and encoded with amplitude levels (mapping the rate of change to levels 1-5, with level 5 corresponding to the most drastic change). Finally, four-dimensional emotional dynamic amplitude data containing "sub-feature type - temporal position - amplitude change rate - level encoding" is generated. This data, through multi-dimensional decomposition and real-time tracking, solves the problem that traditional static features cannot reflect dynamic emotional changes. Preferably, in the specific technical implementation of step 132, emotional dynamic amplitude data is used as the processing object. Based on the sales scenario emotional anomaly judgment rules and smoothing strategies, the emotional dynamic amplitude data is anomaly corrected and trend optimized to generate optimized emotional dynamic amplitude data. Specifically, anomaly thresholds for sales scenario emotional amplitudes are first set (speech rate change threshold 0.3-0.4, volume amplitude difference threshold 0.25-0.35, tone offset threshold 0.2-0.3, determined based on statistical analysis of normal consultation voice samples), and the emotional dynamic amplitude data is processed step by step. Node anomaly detection: if the rate of change of a node exceeds a threshold, it is determined to be an anomaly (mostly caused by environmental noise or unintentional pauses). A neighborhood interpolation correction algorithm (based on interpolation supplementation of normal data in the three frames before and after the anomaly node) is used to correct the anomaly value. The corrected data is then smoothed by a moving average (window size of 2 frames to weaken the interference of small fluctuations). Finally, the amplitude change trend features (rising, stable, falling) are extracted and bound to the amplitude data to generate optimized emotional dynamic amplitude data. This process improves the reliability and trend recognition of the emotional dynamic amplitude data. Preferably, the specific implementation process of step 133 is as follows: taking the optimized emotional dynamic amplitude data as the processing object, based on the spectrogram modeling rules and cross-feature fusion strategy, a two-dimensional spectrogram model is performed on the emotional dynamic amplitude data to generate emotional dynamic spectrum features, which are then fused with temporal context frame features to form a dynamic feature profile of user needs. Specifically, an emotional dynamic spectrum modeling matrix is first constructed, where the row dimension is the temporal step size (determined based on the number of frames in the emotional dynamic amplitude data, usually 10-15 frames), the column dimension is the three types of emotional sub-features (speech rate, volume, and tone), and the matrix elements are the amplitude change rates in the optimized emotional dynamic amplitude data. Based on the modeling matrix, a two-dimensional emotional dynamic spectrum feature is generated, where the row direction represents the time series (reflecting the changes in emotion during the consultation process), and the column direction represents the emotional dimensions (reflecting the collaborative changes of different emotional dimensions). Feature quantization is then performed on the spectrogram features. (Transform the two-dimensional matrix into a 30-35 dimensional vector, extracting the mean, variance, and peak value from each row) to generate a 30-35 dimensional emotional dynamic spectrum feature. A weighted fusion strategy is used to fuse the emotional dynamic spectrum feature with the 55-60 dimensional temporal context frame feature generated in step 12. Based on the emotional-semantic association weights in the sales scenario (emotional feature weight 0.3-0.4, semantic feature weight 0.6-0.7, semantics as primary and emotion as secondary), a dimensional weighted summation is performed. The fused features are then regularized (70-80 dimensional core features are retained through principal component analysis), ultimately generating a 70-80 dimensional user demand dynamic feature construct. This construct achieves the visualization of emotional dynamics through spectral modeling. Combined with a cross-feature fusion mechanism, it realizes a deep binding between user demand semantics and emotional dynamics, providing comprehensive feature support for subsequent accurate demand analysis.
[0020] Optionally, step 2 includes: step 21, performing multi-level noise filtering and feature dimension reconstruction on the dynamic feature construct of user needs to generate a denoised feature construct; step 22, performing intent feature library comparison and posterior probability calculation on the denoised feature construct to determine the user's core need feature basis vector; step 23, performing scene feature library anchoring matching and identifier encoding on the user's core need feature basis vector to assign scene feature identifier codes and generate a scene-based need positioning feature topology accordingly. Optionally, step 21 includes: step 211, performing Gaussian filtering on the dynamic feature construct of user needs to generate an initial filtered feature construct; step 212, performing principal component analysis dimensionality reduction and feature reconstruction on the initial filtered feature construct to generate a dimension-optimized construct; step 213, performing median filtering on the dimension-optimized construct to generate a denoised feature construct.
[0021] Preferably, the specific implementation process of step 211 is as follows: taking the dynamic feature profile of user needs as the processing object, based on the noise distribution characteristics of the sales scenario and the dynamic Gaussian kernel function, the dynamic feature profile of user needs is subjected to targeted Gaussian filtering to generate the initial filtered feature profile; specifically, the dynamic feature profile of user needs is the 70-80 dimension feature vector output in step 13. Combining the sources of noise in the sales scenario (mainly residual environmental noise, emotional feature fluctuation interference, and dimensional redundancy noise in the fusion process), the noise distribution characteristics are analyzed: residual environmental noise is distributed at high frequency randomly (concentrated in the edge dimension of the emotional dynamic spectrum feature), emotional fluctuation interference is distributed locally by impulse (concentrated in the speech rate and volume related feature dimensions), and redundant noise is distributed uniformly with low amplitude (dispersed in various dimensions); a dynamic Gaussian kernel function specific to the sales scenario is designed, and the standard deviation σ of the kernel function is dynamically adjusted based on the feature dimension type, and the semantic core dimension (corresponding to product attributes and functional requirements of 25-30) is used. The filter strength is set to σ=0.8-1.0 for the first dimension (18-22 corresponding to speech rate, volume, and tone) (weak filtering strength, preserving core semantic details), σ=1.2-1.4 for the sentiment association dimension (corresponding to speech rate, volume, and tone dimensions 18-22) (medium filtering strength, smoothing emotional fluctuation noise), and σ=1.5-1.7 for the remaining redundant dimensions (27-28 dimensions) (strong filtering strength, eliminating redundant interference). Based on the dynamic kernel function, the feature construct is subjected to dimension-wise Gaussian filtering. A weighted average filtering algorithm is used (core dimension weight 0.7-0.8, non-core dimension weight 0.2-0.3). Neighborhood weighted smoothing is performed on the feature values of each dimension to suppress high-frequency noise and impulse interference. The amplitude calibration is performed on the filtered features (maintaining the semantic-emotion association ratio of the original features). Finally, a 70-80 dimension initial filtered feature construct is generated. This feature, through the dynamic Gaussian kernel function and dimension-differentiated filtering, is different from the traditional fixed-parameter Gaussian filtering, achieving a balance between noise removal and core information preservation. Preferably, the specific implementation process of step 212 is as follows: taking the initial filtered feature construct as the processing object, based on the importance weight of sales scenario features and the principal component analysis model, the initial filtered feature construct is precisely dimensionality reduced and features reconstructed to generate a dimension-optimized construct; specifically, the feature importance of the initial filtered feature construct is first evaluated, and combined with the logic of expressing sales scenario needs (semantic features are more important than emotional features, and core attribute features are more important than auxiliary descriptive features), feature weights are assigned: product attribute and functional requirement related dimensions (25-30 dimensions) weight 0.4-0.5, emotional dynamic spectrum related dimensions (18-22 dimensions) weight 0.3-0.4, and other fusion dimensions weight 0.1-0.2; a weighted principal component analysis model is constructed, and the feature vectors are projected onto the principal components of the weighted covariance matrix. In the dimensionality reduction space, the variance contribution rate of each principal component is calculated, and a cumulative variance contribution rate threshold of 0.85-0.9 is set (determined statistically based on the requirement to retain features in the sales scenario to ensure that core information is not lost). The top 45-50 principal components are selected (covering more than 90% of the core information). Based on the selected principal components, feature reconstruction is performed, using the principal component inverse mapping algorithm (mapping the principal component space vector back to the original feature space), preserving the semantic-sentiment association structure of the core dimensions. The reconstructed features are normalized (mapped to the 0-1 interval to eliminate amplitude offset during the dimensionality reduction process), generating a 45-50 dimension optimized conformation. This conformation solves the problem of core semantic information loss in the traditional dimensionality reduction process through weighted principal component analysis and precise dimensionality reduction, while reducing the complexity of subsequent processing. Preferably, in the specific technical implementation of step 212, the dimension-optimized construct is taken as the processing object. Based on the sales scenario feature error correction rules, the dimension-optimized construct is reconstructed to correct errors and enhance features to generate an optimized dimension-optimized construct. Specifically, the dimension-wise error value (the difference between the reconstructed feature value and the original initial filter feature value) in the principal component reconstruction process is calculated, and an error threshold of 0.15-0.2 is set (determined based on the accuracy requirements of the sales scenario features). For dimensions with error values higher than the threshold (mostly core semantic dimensions), a weighted fusion of the original initial filter feature value and the reconstructed feature value is used for correction (original value weight 0.6-0.7, reconstructed value weight 0.3-0.4). Core feature enhancement processing is performed on the corrected dimension-optimized construct. Based on the feature importance weight, the core dimension feature values such as product attributes and functional requirements are multiplied by an enhancement coefficient of 1.1-1.2 (to amplify the recognition of core information), while the non-core dimension feature values remain unchanged. Finally, a 45-50 dimension-optimized construct is generated. This processing improves the accuracy and recognition of the features after dimensionality reduction through error correction and core enhancement. Preferably, the specific implementation process of step 213 is as follows: taking the dimension-optimized construct as the processing object, based on the hierarchical characteristics of sales scenario features and the adaptive median window, the dimension-optimized construct is subjected to secondary median filtering to generate a denoised feature construct; specifically, the dimension-optimized construct is divided into three layers according to feature type: core semantic layer (25-30 dimensions, product attributes, functional requirements features), emotional association layer (18-22 dimensions, speech rate, volume, tone features), and fusion transition layer (2-3 dimensions, cross-feature fusion and connection features); an adaptive median window is designed for different layers, the core semantic layer is set with a window size of 3 (small window, retaining semantic details and removing isolated impulse noise), and the emotional association layer is set with a window size of 3. A medium window (size 5) is used to smooth residual emotional fluctuations, while a large window (size 7) is set for the fusion transition layer to completely eliminate redundant noise during fusion. Layer-by-layer median filtering is performed on the feature construct based on the hierarchical window, using a sorted median algorithm (sorting feature values within the window and taking the median as the output). This focuses on eliminating impulse noise not completely eliminated by Gaussian filtering and minor errors introduced during reconstruction. Boundary calibration is performed on the filtered features (ensuring that feature values in each dimension are within the 0-1 range), ultimately generating a 45-50 dimensional denoised feature construct. This feature, through hierarchical adaptive median processing, achieves a balance between secondary denoising and core feature protection, providing a high-precision, low-noise feature foundation for subsequent user requirement analysis.
[0022] Optionally, step 22 includes: step 221, constructing an intent feature library, the intent feature library containing feature templates corresponding to preset demand intents; step 222, performing cosine similarity calculation on the denoised feature constructs and each feature template to obtain multiple similarity values; step 223, normalizing the multiple similarity values to convert them into posterior probabilities, and filtering the feature template corresponding to the maximum posterior probability to determine the user's core demand feature basis vector.
[0023] Preferably, the specific implementation process of step 221 is as follows: Based on the classification of high-frequency demand intents in sales scenarios and the rules for constructing feature templates, an intent feature library containing feature templates corresponding to preset demand intents is constructed; specifically, combined with the statistics of user consultation data in sales scenarios (user needs are concentrated in four core intents: product selection, function consultation, price comparison, and after-sales consultation), 15-20 subdivision intents are further subdivided (such as "camera phone selection", "game performance consultation", "cost-effectiveness comparison", etc.), each subdivision intent corresponds to one type of preset demand intent; for each preset demand intent, a large number of sales scenario voice samples are collected, and 45-50 dimensional denoised feature constructs of the samples are extracted (consistent with the feature dimensions output in step 21), and a feature template generation matrix is constructed, with the row dimension of the matrix being the number of samples (5 samples are collected for each intent). The dataset consists of 00-800 samples with 45-50 column dimensions. A weighted average template generation algorithm is used, which calculates the weighted average of sample features based on the importance weights of sales scenario intent features (core semantic layer weight 0.6-0.7, emotional association layer weight 0.2-0.3, and fusion transition layer weight 0.1) to generate a standard feature template corresponding to each preset demand intent. All feature templates are normalized (mapped to the 0-1 interval), and intent labels and confidence thresholds are marked (based on sample matching accuracy statistics, a minimum matching confidence of 0.7-0.8 is set). Finally, an intent feature library containing 15-20 feature templates is constructed. This feature library, through scenario-based subdivision of intent definition and weighted template generation, differs from traditional generalized intent libraries and achieves accurate representation of demand intent. Preferably, the specific implementation process of step 222 is as follows: taking the denoised feature construct and each feature template in the intent feature library as the processing objects, the differential similarity is calculated based on the sales scenario feature dimension weight and hierarchical cosine similarity algorithm to obtain multiple similarity values; specifically, the denoised feature construct is the 45-50 dimension feature vector output in step 21, and the feature template is 15-20 45-50 dimension standard templates in the intent feature library; combined with the sales scenario demand matching logic (the matching priority of core semantic features is higher than that of sentiment features), the denoised feature construct and feature template are both split into three layers: core semantic layer (25-30 dimensions), sentiment association layer (18-22 dimensions), and fusion transition layer. (2-3 dimensions), assign layered weight coefficients (core semantic layer weight 0.6-0.7, sentiment association layer weight 0.2-0.3, fusion transition layer weight 0.1); perform cosine similarity calculation (quantify the consistency of the angle between vector spaces) on each feature layer to obtain three layered similarity values: core semantic similarity, sentiment association similarity, and fusion transition similarity; use a weighted summation algorithm to fuse the three layered similarity values according to the weight coefficients to obtain the comprehensive similarity value between each feature template and the denoised feature construct, and finally output 15-20 corresponding similarity values. This calculation method solves the problem of ignoring the differences in feature importance in traditional single similarity calculation by using layered weighted processing, thus improving matching accuracy. Preferably, in the specific technical implementation of step 222, the multiple similarity values obtained initially are used as the processing object. Based on the sales scenario intent association rules and similarity correction strategies, the similarity values are optimized and corrected to generate accurate similarity values. Specifically, the characteristics of sales scenario intent association are analyzed (some intents have strong associations, such as "gaming phone selection" and "performance consultation" having semantic overlap), an intent association matrix is constructed, and strongly associated intent pairs are marked (such as "camera phone selection" and "imaging function consultation" being strongly associated); the initial similarity values are then correlated. The correction process involves adjusting the similarity value of a certain intent to a coefficient of 0.9-0.95 if the similarity value is higher than 0.6 (to reduce interference from related intents). Simultaneously, a similarity threshold is set (similarity values below 0.3 are set to 0, considered as having no matching possibility). The corrected similarity values are then sorted (from high to low), retaining the top 5-8 high similarity values and marking the remaining similarity values as invalid, ultimately generating 5-8 precise similarity values. This process, through association correction and threshold filtering, reduces the probability of false matches between similar intents. Preferably, the specific implementation process of step 223 is as follows: Multiple precise similarity values are processed as objects. Based on a probability normalization algorithm and confidence screening rules, the similarity values are converted into posterior probabilities, and the optimal feature template is screened to determine the user's core demand feature basis vector. Specifically, a softmax probability normalization algorithm is used to map 5-8 precise similarity values to the 0-1 interval, converting them into corresponding posterior probabilities (the sum of all posterior probabilities is 1), quantifying the matching confidence of each feature template; a posterior probability screening threshold is set (0.7 based on the matching accuracy requirements of the sales scenario). If a feature template with a posterior probability higher than the threshold exists, the feature template corresponding to the maximum posterior probability is selected. If all posterior probabilities are lower than the threshold, an intent completion mechanism is triggered (marked as "fuzzy intent", which can be improved through interactive follow-up questions). The selected optimal feature template is used as the user's core demand feature basis vector. This vector retains the standard feature structure and semantic-emotional association characteristics of the preset demand intent, providing accurate feature basis for subsequent demand response and product recommendation. This process ensures the reliability of core demand positioning through probabilistic transformation and confidence screening.
[0024] Optionally, step 23 includes: step 231, constructing a scene feature library, wherein the scene feature library contains feature identifiers and encoding rules for preset scenes; step 232, anchoring and matching the user core demand feature basis vector with the scene features in the scene feature library to determine the target scene; step 233, assigning scene feature identifier codes according to the encoding rules corresponding to the target scene, fusing the scene feature identifier codes with the user core demand feature basis vector, and generating a scene-based demand positioning feature topology accordingly.
[0025] Preferably, the specific implementation process of step 231 is as follows: Based on the sales scenario business scenario classification system and feature identifier construction rules, a scenario feature library containing feature identifiers and encoding rules for preset scenarios is constructed; specifically, combined with the sales scenario business process (the entire link from user consultation to transaction includes four main scenarios: online consultation, offline experience, telephone communication, and after-sales follow-up), 8-12 sub-scenarios are further subdivided (such as "online consultation on e-commerce platforms", "in-store real device experience consultation", "after-sales fault reporting", etc.), with each sub-scenarios corresponding to a type of preset scenario; for each preset scenario, core feature identifiers of the scenario are extracted (including three types: interaction channel features, scenario function features, and user identity features). Interaction channel features correspond to channel attributes such as "online / offline / telephone", scenario function features correspond to scenario uses such as "consultation / experience / reporting", and user identity features correspond to identity tags such as "new user / old user / potential user"; a scenario feature identifier matrix is constructed, with the matrix having 8-12 preset scenarios in the row dimension and the three types of feature identifiers in the column dimension. Each matrix element represents the specific value of the feature identifier for that scenario (e.g., "online consultation on e-commerce platforms" corresponds to the channel feature "online", the functional feature "consultation", and the identity feature "potential user"). A scenario coding rule system is designed, employing a hierarchical coding approach. Channel features occupy 2 binary bits (e.g., online = 01, offline = 10, telephone = 11), functional features occupy 2 binary bits (e.g., consultation = 01, experience = 10, repair = 11), and identity features occupy 2 binary bits (e.g., new user = 01, old user = 10, potential user = 11), combining to generate a 6-bit scenario feature identifier code. All preset scenario feature identifiers and coding rules are associated and stored, with scenario priority weights marked (consultation scenarios weight 0.6-0.7, experience scenarios weight 0.2-0.3, after-sales scenarios weight 0.1). Finally, a scenario feature library containing 8-12 preset scenarios is constructed. This feature library, through hierarchical scenario classification and structured coding, differs from traditional single-dimensional scenario libraries, achieving a comprehensive representation of scenario information. Preferably, the specific implementation process of step 232 is as follows: taking the user core demand feature basis vector and the scene features in the scene feature library as the processing objects, and based on the core anchoring dimension of the sales scene and the two-stage matching algorithm, the two are accurately anchored and matched to determine the target scene; specifically, the user core demand feature basis vector is the 45-50 dimension feature vector output in step 22, where the core semantic layer includes the scene association dimension (the 33rd-35th dimension, corresponding to channel, function, and identity related features); the scene association dimension features are extracted from the user core demand feature basis vector to generate a 3-dimensional scene matching vector (corresponding to the channel, function, and identity feature values respectively); the first stage performs coarse matching, and the scene matching vector is matched with the feature identifiers of all preset scenes in the scene feature library by keywords (e.g., if the channel feature in the vector is "online", then all are matched). In the first stage, scenarios identified as "online" are selected, and 3-5 candidate scenarios are chosen. The second stage involves fine-tuning, scoring the candidate scenarios based on scenario priority weights and feature similarity. Channel feature matching scores are weighted at 0.4, functional features at 0.3, and identity features at 0.3. A perfect match scores 1 point, a partial match at 0.5, and no match at 0. The overall matching score for each candidate scenario is calculated. A matching score threshold is set (based on scenario matching accuracy statistics, a minimum threshold of 0.7-0.8 is set), and scenarios with the highest overall matching score above the threshold are selected as target scenarios. If all candidate scenario scores are below the threshold, a scenario completion mechanism is triggered (defaulting to matching "general consultation scenario"). This matching method, through two-stage anchoring and weighted scoring, solves the problem of traditional single-match ignoring scenario priority. Preferably, in the specific technical implementation of step 232, the initially determined target scenario is used as the processing object. Based on the sales scenario association rules and matching result verification strategy, the target scenario is subjected to secondary verification and optimization adjustment to generate the final target scenario. Specifically, a scenario association matrix is constructed, and strongly associated scenario pairs are labeled. If the difference between the comprehensive matching score and the second highest score of the initially determined target scenario is less than 0.1 (considered as fuzzy matching), adjustments are made based on the scenario association rules, prioritizing the selection of scenarios that better match the user's core needs and intentions. The adjusted target scenario is then verified for matching results. The scenario association dimension of the feature identifier of the scenario and the user's core needs feature basis vector is extracted for consistency verification. If the verification pass rate is higher than 0.8 (i.e., at least two of the three feature identifiers are completely matched), the scenario is confirmed as the final target scenario. If the verification pass rate is lower than 0.8, the fine matching process is re-executed (the weight coefficients are adjusted and the score is re-evaluated) until a final target scenario that meets the requirements is determined. This process improves the matching reliability of the target scenario through association adjustment and consistency verification. Preferably, the specific implementation process of step 233 is as follows: taking the final target scene and the user's core demand feature basis vector as the processing objects, based on the scene feature encoding rules and the topology fusion algorithm, a scene feature identifier code is assigned and feature fusion is performed to generate a scene-based demand positioning feature topology; specifically, according to the corresponding encoding rules in the scene feature library, a 6-bit binary scene feature identifier code is generated and converted into a 2-dimensional decimal feature vector; a feature fusion topology matrix is constructed, with the row dimension of the matrix being 45-50 dimensions of the user's core demand feature basis vector and the column dimension being the 2-dimensional scene feature vector. The first part of the matrix is the user's core demand feature basis vector, and the second part is the scene feature vector, forming a 47-52 dimensional fusion feature. The algorithm employs a weighted topology fusion method, based on the fusion weights of sales scenario features (core demand feature weight 0.8-0.9, scenario feature weight 0.1-0.2). Dimensional correlation modeling is performed on the fused feature vectors, establishing a strong correlation between the core semantic layer and the scenario feature vectors (correlation coefficient 0.7-0.8), and a weak correlation between the emotional correlation layer and the fusion transition layer (correlation coefficient 0.2-0.3). The topology matrix is normalized (mapped to the 0-1 interval) to generate a scenario-based demand positioning feature topology containing a demand-scenario correlation structure. This topology, through dimensional correlation modeling, achieves a deep binding between core user needs and target scenarios, providing accurate feature support for subsequent scenario-based product recommendations and demand responses.
[0026] Optionally, step 3 includes: Step 31, extracting demand attribute dimensions and quantifying features from the scenario-based demand positioning feature topology to generate a demand attribute feature set; Step 32, calling the product attribute knowledge graph to perform weighted association calculation and matching degree ranking on the demand attribute feature set and product attribute features to generate a set of matching products; Step 33, based on the attribute information of the set of matching products and the scenario-based demand positioning feature topology, constructing recommendation semantic logic and organizing text to generate customized recommendation response text. Optionally, step 31 includes: Step 311, parsing attribute dimensions from the scenario-based demand positioning feature topology to determine demand attribute categories; Step 312, extracting features corresponding to each demand attribute category to generate original attribute features; Step 313, performing numerical quantization processing on the original attribute features to generate a set of demand attribute features.
[0027] Preferably, the specific implementation process of step 311 is as follows: taking the scenario-based demand positioning feature topology as the processing object, based on the sales scenario demand attribute classification system and dynamic dimension mapping rules, the feature topology is analyzed for attribute dimensions to determine the demand attribute categories; specifically, the scenario-based demand positioning feature topology is a 47-52 dimension fusion topology matrix (row dimensions are feature dimensions, column dimensions are association strength), combined with the core attributes of user needs in the sales scenario (five categories: product parameter attributes, functional demand attributes, scenario adaptation attributes, emotional tendency attributes, and decision weight attributes), an attribute dimension mapping library is constructed. The library contains the feature dimension range and association judgment rules corresponding to each type of attribute: product parameter attributes correspond to dimensions 1-15 of the topology matrix (core semantic layer product attribute features), and functional demand attributes correspond to dimensions 16-25 (core semantic layer functional features). The scene adaptation attribute corresponds to dimensions 26-30 (the intersection of scene-related features and the core semantic layer), the sentiment tendency attribute corresponds to dimensions 31-38 (sentiment-related layer features), and the decision weight attribute corresponds to dimensions 39-47 (weight distribution of core demand features). A dimension-by-dimensional association strength analysis is performed on the feature topology to calculate the association matching degree between each dimension and the five attribute categories (based on the association coefficient of the column dimensions of the topology matrix). If the matching degree between a dimension and an attribute category is higher than 0.65 (association judgment threshold), then that dimension is assigned to the corresponding attribute category. The completeness of the assigned attribute dimensions is verified. If any dimension is missing, the association dimensions are supplemented from the fusion transition layer. Finally, the five demand attribute categories and their corresponding feature dimensions are determined. This parsing method, through dynamic association matching, differs from traditional fixed-dimensional division and achieves accurate positioning of attribute categories.
[0028] Preferably, the specific implementation process of step 312 is as follows: taking the scenario-based demand positioning feature topology and the determined demand attribute categories as the processing objects, based on the sales scenario attribute feature extraction priority and hierarchical extraction algorithm, the features corresponding to each attribute category are extracted in a targeted manner to generate the original attribute features; specifically, the extraction priority weights are assigned according to the importance of demand attributes (product parameter attributes > functional appeal attributes > decision weight attributes > scenario adaptation attributes > emotional tendency attributes) (0.35, 0.25, 0.2, 0.1, 0.1 respectively); for the product parameter attribute category, threshold filtering extraction is adopted (retaining feature values with a correlation strength higher than 0.7 in the topology matrix, and setting the rest to 0), highlighting the core product parameter features; for the functional appeal attribute category, a threshold filtering extraction is adopted. Semantic association extraction is used (based on a sales scenario functional keyword feature library; if a match is successful, the complete feature vector is retained; otherwise, only the amplitude feature is retained). For decision weight attribute categories, weight normalization extraction is used (the corresponding dimension feature values are mapped to the 0-1 interval and integrated according to priority weights). For scenario adaptation attribute categories, cross-association extraction is used (combining the target scenario feature identifier code to extract feature dimensions strongly associated with the scenario). For sentiment attribute categories, amplitude peak extraction is used (retaining the feature values and location information corresponding to the sentiment fluctuation peaks). The features extracted from the five categories are organized into 10-12 dimensional vectors and combined to generate 50-60 dimensional original attribute features. This extraction method ensures the integrity of core attribute features through hierarchical priority processing.
[0029] Preferably, in the specific technical implementation of step 312, the original attribute features are taken as the processing object. Based on the sales scenario attribute feature redundancy judgment rules and purification strategies, the original attribute features are redundancy-eliminating and feature purification to generate optimized original attribute features. Specifically, variance analysis is performed on the original features corresponding to each type of attribute to calculate the redundancy correlation coefficient between dimensions. If the correlation coefficient between two dimensions is higher than 0.7 (redundancy judgment threshold), the dimension with larger variance is retained (retaining core information). For the cross-dimensional features of product parameter attributes and functional appeal attributes, weighted fusion purification is adopted (the fusion coefficient is allocated according to the priority weight of the two types of attributes). For abnormal amplitude features in the sentiment tendency attribute (exceeding the mean ± 2 times the standard deviation), neighborhood mean correction is adopted (based on the smoothing of feature values of the preceding and following dimensions). Dimension regularization is performed on each type of purified attribute feature (unified to 8-10 dimensions) to ensure the balance of dimensions of various attribute features, and finally, 40-50 dimension optimized original attribute features are generated. This processing improves the purity and reliability of the original attribute features through redundancy elimination and anomaly correction.
[0030] Preferably, the specific implementation process of step 313 is as follows: taking the optimized original attribute features as the processing object, based on the sales scenario attribute feature quantization rules and multimodal quantization algorithms, the original attribute features are classified and numerically quantized to generate a demand attribute feature set; specifically, for feature types of different attribute categories, a dedicated quantization scheme is designed: product parameter attributes adopt interval mapping quantization (the feature value is divided into 1-5 levels according to the product parameter level, such as "pixel" feature value above 0.8 corresponding to level 5), functional demand attributes adopt binary encoding quantization (if the functional keyword is matched, it is quantized as 1, otherwise it is 0, forming a binary feature vector), decision weight attributes adopt probability distribution quantization (the feature value is converted into a probability density value, satisfying the probability sum of 1), and scenario adaptation attributes adopt similarity. Quantization (calculating the matching degree with the standard features of the target scene, quantized into a value between 0 and 1). The sentiment tendency attribute is quantized using intensity-level grading (divided into three levels: weak, medium, and strong, quantized to 0.3, 0.6, and 0.9 respectively based on the feature amplitude). The dimensions of each type of attribute feature are unified (each attribute is quantized into 8 dimensions), and combined according to the original attribute category order to generate a 40-dimensional standard feature vector. The feature vector is then normalized as a whole (mapped to the 0-1 interval), and attribute category labels and quantization confidence scores are labeled (based on quantization accuracy statistics, with a confidence score of not less than 0.8). Finally, a set of requirement attribute features is generated. This feature set, through multimodal quantization, achieves a unified and computable representation of different types of attribute features, providing structured data support for subsequent product matching and requirement response.
[0031] Optionally, step 32 includes: step 321, calling the product attribute knowledge graph to extract the attribute features of the products and constructing a product feature index library; step 322, performing weighted association calculation on the demand attribute feature set and the features in the product feature index library to obtain the matching score of each product; step 323, sorting the matching scores in descending order and selecting the products with the highest ranking to form a matching product set.
[0032] Preferably, the specific implementation process of step 321 is as follows: using a product attribute knowledge graph as the data foundation, and based on the sales scenario product feature extraction rules and index construction algorithm, the attribute features of the product are extracted and a product feature index library is constructed. Specifically, the product attribute knowledge graph contains full-dimensional attribute information of the product (covering five major categories of nodes, including product parameters, functional characteristics, scenario adaptation, price range, and user reviews, as well as the relationships between nodes). Combining the five categories of demand attributes determined in step 31, product feature extraction mapping rules are designed: product parameter features correspond to 20-25 core attributes (such as processor model, memory capacity, pixel value, etc.) of the "product parameters" node in the knowledge graph; functional characteristic features correspond to 15-18 key functions (such as shooting mode, battery life, intelligent interaction, etc.) of the "functional characteristics" node; and scenario adaptation features correspond to 8-10 adaptation scenarios (such as office, gaming, outdoor) of the "scenario adaptation" node. The emotional association features correspond to the emotional tendency statistics of the "user evaluation" node (such as satisfaction, frequency of positive keywords), and the decision-adaptation features correspond to the quantitative attributes of the associated nodes such as "price range" and "cost-effectiveness". For each product in the knowledge graph, the above features are extracted according to the mapping rules and converted into a 40-dimensional feature vector with the same dimensions as the demand attribute feature set (each attribute corresponds to 8 dimensions, which are the same as the feature structure output in step 31). A hierarchical indexing algorithm is used to build a first-level index according to product categories (such as mobile phones, computers, home appliances, etc.) and a second-level index according to core demand attributes (such as product parameters, price). Each product feature vector is labeled with an index label and an update timestamp (supporting real-time updates of product features). Finally, a product feature index library containing all categories of products is constructed. This index library, through deep analysis of the knowledge graph and scenario-based feature mapping, is different from the traditional product feature library and achieves accurate alignment between product features and demand attributes.
[0033] Preferably, the specific implementation process of step 322 is as follows: taking the demand attribute feature set and the product features in the product feature index as the processing objects, and based on the sales scenario attribute hierarchy weight and multi-dimensional association calculation algorithm, differential weighted association calculation is performed on the two to obtain the matching score of each product; specifically, the demand attribute feature set is the 40-dimensional feature vector output in step 31, and the product features are the 40-dimensional feature vectors of each product in the index, which are split into five sub-feature vectors (8 dimensions per category) according to five attribute categories; combined with the sales scenario user decision-making logic (the weight of product parameters and functional requirements is higher than that of other attributes), attribute hierarchy weighting coefficients are assigned: product parameter attribute weight 0.35, functional requirement attribute weight 0.3, scenario adaptation attribute weight 0.15, and sentiment tendency attribute weight 0.15. The weights are assigned as follows: Attribute weight 0.1 and decision weight 0.1. For each attribute's corresponding demand sub-feature and product sub-feature, cosine similarity calculation (quantifying vector matching degree) is performed to obtain five attribute-level similarity values. A weighted summation algorithm is used to fuse the five-level similarity values by weighted coefficients. Simultaneously, a scenario adaptation correction factor is introduced (based on the target scenario determined in step 23, if the matching degree between the product scenario adaptation feature and the target scenario is higher than 0.8, the fusion result is multiplied by a correction coefficient of 1.05-1.1). This yields the comprehensive matching score for each product and the demand attribute feature set. Finally, the matching scores for all products in the product feature index are output. This calculation method, through attribute-level weighting and scenario correction, solves the problem of traditional single-dimensional matching ignoring demand priority.
[0034] Preferably, in the specific technical implementation of step 322, the initially obtained matching scores of each product are used as the processing object. Based on the sales scenario outlier filtering rules and matching score optimization strategies, the matching scores are corrected and optimized to generate accurate matching scores. Specifically, an outlier threshold for sales scenario matching scores is set (scores below 0.2 or above 0.95 are considered outliers; below 0.2 may indicate missing product features, and above 0.95 may indicate feature overfitting). For scores below 0.2, the missing features are supplemented by querying the product feature index library and then recalculated (if supplementation is not possible, it is marked as an invalid match). For scores above 0.95, the product features are introduced... Attribute integrity verification (verifying whether product features include core demand attributes; deducting 0.05-0.1 points if missing); simultaneously, based on product inventory status and sales volume, a dynamic adjustment factor is introduced: products with sufficient inventory are multiplied by a factor of 1.02-1.03, best-selling products (top 20% of sales in the past 30 days) are multiplied by a factor of 1.01-1.02, and products with tight inventory are multiplied by a factor of 0.98-0.99; the adjusted matching score is normalized (mapped to the 0-1 range) to generate an accurate matching score. This process, through anomaly correction and dynamic adjustment, improves the reliability and practicality of the matching score.
[0035] Preferably, the specific implementation process of step 323 is as follows: taking the accurate matching score of each product as the processing object, based on the sales scenario adaptive sorting rules and dynamic threshold filtering strategy, the matching scores are sorted in descending order and a set of matching products is selected; specifically, all products are sorted from high to low according to their accurate matching scores to generate a product sorting list; an adaptive dynamic threshold filtering mechanism is designed, and the threshold is dynamically adjusted based on the confidence of the demand attribute feature set and the total number of products: if the confidence of the demand attribute feature set is higher than 0.85 (the quantified confidence level marked in step 31), the threshold is set to 0.7-0.75 (filtering high-matching products); if the confidence level is between 0.7 and 0.85, the threshold is set to 0.6-0.7 (balancing matching degree and selection diversity). If the confidence level is below 0.7, set the threshold to 0.5-0.6 (to broaden the selection range); simultaneously, set a maximum selection limit (10-15 items to avoid selecting too many). If the number of items above the threshold exceeds the limit, select the top 10-15 items by score. If the number of items below the threshold is less than 5, lower the threshold by 0.05-0.1 and re-select (ensuring at least 5 candidate items). For the selected items, label the matching score, core matching attributes (attribute categories that highly match the needs), and difference attributes (attributes that differ from the needs). Finally, a matching item set is formed. This set, through dynamic thresholds and sorting strategies, achieves a balance between the accuracy and practicality of item matching, providing a high-quality candidate pool for subsequent item recommendations.
[0036] Optionally, step 33 includes: step 331, extracting scenario-adaptive semantic rules from the scenario-based demand positioning feature topology and extracting core attribute information of products from the matching product set; step 332, constructing a recommendation semantic logic framework based on the scenario-adaptive semantic rules; and step 333, filling the core attribute information of products into the semantic logic framework and optimizing the sentence organization to generate customized recommendation response text.
[0037] Preferably, the specific implementation process of step 331 is as follows: taking the scenario-based demand positioning feature topology and the matching product set as the processing objects, based on the sales scenario semantic rule extraction algorithm and the core attribute filtering strategy, the scenario adaptation semantic rules and product core attribute information are extracted; specifically, the scenario-based demand positioning feature topology is a 47-52 dimension fusion topology matrix (including scenario association dimension, core semantic dimension, etc.), focusing on the 26th-30th dimension features corresponding to the scenario adaptation attributes in the topology matrix, combined with the target scenario tags determined in step 23 (such as "e-commerce platform online consultation" "game scenario selection"), the semantic rule parsing algorithm is used to extract the core semantic elements of scenario adaptation: scenario interaction method (such as online consultation - concise text, offline experience - rich details), user demand tendency (such as product selection - parameter comparison, function consultation - operation instructions), emotional adaptation style (such as urgent needs - direct response, comparative needs - objective analysis), and these elements are combined into scenario adaptation semantic rules (such as "online consultation"). In the scenario of selecting gaming phones: concisely present core performance parameters, highlight game compatibility advantages, and use a proactive tone); for each product in the matching product set, based on the core matching attributes and difference attributes marked in step 32, filter the product's core attribute information: product parameter category (3-5 core parameters, such as processor, memory, pixels, with priority given to attributes that highly match user needs), functional feature category (2-3 key functions, such as game mode, fast charging technology, which fit the scenario adaptation requirements), scenario adaptation category (1-2 core adaptation points, such as "long-term gaming battery life" and "multi-tasking office work"), and decision support category (1 core advantage, such as cost-effectiveness and best-selling ranking, determined based on the product's dynamic adjustment factors). Each attribute information is labeled with the attribute name, quantitative value, and adaptation description, ultimately obtaining structured scenario adaptation semantic rules and product core attribute information. This extraction method, through precise alignment of scenario and product attributes, differs from traditional generalized extraction and achieves scenario-based adaptation of rules and information.
[0038] Preferably, the specific implementation process of step 332 is as follows: Based on the scene adaptation semantic rules, a recommendation semantic logic framework is constructed based on the sales scene recommendation logic modeling rules and the framework dynamic generation algorithm; specifically, the scene adaptation semantic rules include three core elements: interaction mode, demand tendency, and emotional style. Combined with the sales scene recommendation expression logic (a five-segment structure of "scene greeting - demand response - product recommendation - adaptation description - decision guidance"), the framework dynamic generation rules are designed: the framework simplicity is determined according to the interaction mode (the online consultation framework includes 3-4 modules, and the offline experience framework includes 5-6 modules), the module weight is determined according to the demand tendency (the weight of the selection demand - product comparison module is 0.4, and the weight of the functional demand - functional description module is 0.4), and the tone template is determined according to the emotional style (urgent demand - short sentence template, comparison demand - point template); a hierarchical framework construction algorithm is used to build a first-level framework (a five-segment core structure), and the second-level sub-modules are filled based on the dynamic generation rules: scene The framework includes a greeting module (containing scenario-based greeting templates, such as "Online consultation: Hello! For your gaming phone selection needs, we recommend the following products~"), a demand response module (containing demand-precise response templates, such as "The following products can fully meet your concerns about game performance and battery life"), a product recommendation module (containing multi-product arrangement templates, such as single product - detailed introduction, multiple products - comparison presentation), an adaptation instruction module (containing attribute adaptation association templates, such as "XX parameters can support XX scenario needs"), and a decision guidance module (containing scenario-based guidance templates, such as online consultation - "Click to view details", offline experience - "You can experience the actual device in-store"). Semantic association rules (such as a one-to-one correspondence between the product recommendation module and the adaptation instruction module) and consistent tone requirements (such as maintaining a professional and friendly tone throughout) are annotated for each framework module. Finally, a recommendation semantic logic framework adapted to the target scenario is generated. This framework, through dynamic adaptation of scenario rules, differs from a fixed template framework, achieving personalized matching of recommendation logic.
[0039] Preferably, in the specific technical implementation of step 332, the initially constructed recommendation semantic logic framework is used as the processing object. Based on the sales scenario demand intent enhancement rules and framework optimization strategies, the logic framework is further optimized to generate a precise semantic logic framework. Specifically, combined with the intent tags corresponding to the user core demand feature basis vectors determined in step 22 (such as "game performance consultation" and "cost-effectiveness comparison"), the weights of the framework modules are adjusted: if the intent is "parameter comparison", a product comparison sub-module (such as "core parameter comparison table") is added, and the weight is increased to 0.5; if the intent is "function consultation", a function description sub-module (such as "function operation scenario example") is expanded. The weight is increased to 0.45; at the same time, sentiment adaptation optimization is introduced: based on the sentiment dynamic spectrum features in step 13, if the user sentiment is "forced to cut", the framework hierarchy is simplified (secondary modules are removed) to highlight the core recommendation content; if the sentiment is "strong hesitation", a user evaluation sub-module is added to enhance decision-making confidence; the optimized framework is checked for logical coherence (ensuring natural transitions between modules and close semantic connections), and if there are logical gaps, connecting sub-modules are added (such as "Based on your needs, the core advantages of this product are reflected in:"). Finally, a precise semantic logic framework is generated. This process improves the accuracy and persuasiveness of the framework through dual optimization of demand intent and sentiment adaptation.
[0040] Preferably, the specific implementation process of step 333 is as follows: taking the core attribute information of the product and the precise semantic logic framework as the processing objects, based on the natural language organization algorithm and text optimization strategy, the attribute information is filled in and the sentence organization is optimized to generate customized recommendation response text; specifically, according to the module order of the precise semantic logic framework, the core attribute information of the product is filled into the corresponding sub-modules one by one: the product recommendation module is filled with the product name and core parameters, the adaptation description module is filled with the related description of attributes and needs, and the decision guidance module is filled with scenario-based guidance; the scenario-based language organization algorithm is used to adjust the expression style according to semantic rules: online consultation scenarios use concise short sentences (single sentence length not exceeding 15 characters), offline experience scenarios use detailed descriptions (supplementing attribute details and usage scenarios), and emotional urgency scenarios use active sentence structures (such as "This product..."). "This product fully meets your needs!" The comparison of the required scenarios uses objective sentence structures (such as "This product is superior to similar products in XX parameters"). The filled text undergoes three layers of optimization: semantic accuracy optimization (verifying that attribute information matches the actual product and ensuring the adaptation description is not exaggerated), sentence fluency optimization (adjusting word order and adding conjunctions such as "not only that," "more importantly"), and natural expression optimization. For multi-product recommendation text, it is sorted by matching score (high-matching products are given priority in detailed descriptions) and presented in a point-by-point or segmented manner (avoiding information clutter). Finally, a customized recommendation response text that meets scenario adaptation requirements, is logically clear, and naturally expressed is generated. Through precise attribute filling and multi-layered language optimization, this text achieves efficient transmission of product information and user needs, improving the acceptability and conversion potential of the recommendation response.
[0041] Optionally, step 4 includes: step 41, performing semantic segmentation and sentiment annotation on the customized recommendation response text to generate sentiment-annotated response text; step 42, based on the sentiment-annotated response text, fusing temporal context frame features and sentiment dynamic spectrum features to perform text-to-speech conversion and basic intonation generation to generate an initial speech stream; step 43, adjusting the sentiment intonation amplitude and adapting the context rhythm of the initial speech stream to generate a real-time response speech stream. Optionally, step 41 includes: step 411, performing punctuation recognition and semantic segmentation on the customized recommendation response text to generate independent sentence units; step 412, determining the sentiment tendency of each independent sentence unit and annotating it with a corresponding sentiment tag; step 413, integrating the sentiment-tagged independent sentence units to generate sentiment-annotated response text.
[0042] Preferably, the specific implementation process of step 411 is as follows: taking the customized recommendation response text as the processing object, based on the punctuation recognition algorithm and semantic segmentation rules for sales scenarios, the text is subjected to punctuation recognition and semantic segmentation to generate independent sentence units; specifically, the customized recommendation response text is the text from step 33. The output natural language text (including modules such as scene greetings, product recommendations, and adaptation instructions, with a sentence style that fits the sales scenario) first constructs a punctuation recognition library specifically for the sales scenario, covering terminating punctuation (period, exclamation mark, question mark), pause punctuation (comma, pause mark, semicolon), and special punctuation (quotation mark, parentheses, dash). Each punctuation mark is labeled with recognition features and scene adaptation rules (e.g., exclamation marks are often used for emotional emphasis in sales text). A multi-feature fusion recognition algorithm is used to scan the text character by character, combining position, contextual semantics, and scene usage features to accurately identify various punctuation marks (recognition accuracy is no less than 0.98). Based on the recognition results, semantic segmentation rules are designed: terminating punctuation is the main segmentation node (periods and exclamation marks are segmented first, and question marks are judged based on semantics to determine whether they form an independent sentence), pause punctuation is used as an auxiliary reference (if the semantics are closely related after a comma, it is not segmented), and the content of special punctuation is merged with the main sentence into one unit. The integrity of the segmented sentences is checked (each unit is semantically complete, and the length is controlled between 5-30 characters). If semantic breaks exist, they are supplemented and connected to ultimately generate 8-15 independent sentence units. This segmentation method, through scenario-based rule adaptation, differs from traditional pure punctuation segmentation and achieves semantic integrity and structural rationality of sentence units.
[0043] Preferably, the specific implementation process of step 412 is as follows: taking each independent sentence unit as the processing object, based on the sales scenario sentiment tendency judgment model and multi-dimensional feature extraction strategy, the sentiment tendency of the sentence unit is judged and the corresponding sentiment label is marked; specifically, combined with the sentiment expression characteristics of sales scenario recommendation text, it is divided into positive category (enthusiastic recommendation, advantage emphasis, etc.), neutral category (parameter description, factual statement, etc.), and guiding category (decision guidance, action suggestion, etc.); constructing a sales scenario sentiment tendency judgment model, extracting lexical features (based on the sales scenario sentiment dictionary), sentence structure features (weight allocation of exclamatory sentences, declarative sentences, and imperative sentences), and semantic association features (combined with scenario adaptation semantic rules); calculating three types of feature scores for each sentence unit: lexical feature matching score, sentence structure feature adaptation score, and semantic association score, with the total score = lexical feature score × 0.4 + sentence structure feature score × 0.3 + semantic association score × 0.3; judging the sentiment tendency based on the total score: 0.7 points and above are marked "positive", 0.4-0.7 points are marked "neutral", and 0.4 points are marked "insensitive". The following are labeled with "Guidance"; at the same time, the emotional intensity (strong, medium, weak) is also labeled. Finally, each independent sentence unit is labeled with a combination of "emotional type - emotional intensity". This judgment model, through scenario-based feature adaptation, is different from general sentiment judgment models and achieves accurate identification of emotional tendencies in sales scenarios.
[0044] Preferably, in the specific technical implementation of step 412, the independent sentence units labeled with sentiment tags are taken as the processing objects. Based on the sales scenario sentiment consistency verification rules and tag optimization strategies, the sentiment tags are optimized and adjusted a second time. Specifically, combined with the sentiment adaptation style in step 33, the label distribution is checked for consistency (the proportion of positive tags meets the scenario requirements, and the guiding tags are concentrated in the second half of the text). If a sentence unit's tag conflicts with the scenario, the score is recalculated and adjusted. If the guiding tags are scattered, the guiding tags in the non-second half are adjusted to neutral. For sentence units with close semantic connections, the sentiment tags are ensured to be consistent to avoid sentiment breaks. Finally, independent sentence units with accurate and reasonably distributed sentiment tags are generated. This optimization process improves the reliability of sentiment labeling through scenario consistency and logical coherence verification.
[0045] Preferably, the specific implementation process of step 413 is as follows: taking independent sentence units with sentiment tags as the processing object, and integrating sentence units based on sales scenario text integration rules and tag visualization strategies to generate sentiment-annotated response text; specifically, a hierarchical integration algorithm is used in the original order, with sentence units of the same sentiment type arranged continuously to avoid frequent switching; the original semantics and punctuation are preserved during the integration process, and contextualized connecting words are added to ensure text fluency; sentiment tags are visualized (placed at the end of the sentence unit, marked with square brackets) without affecting readability; the integrated text is verified as a whole (the tag distribution is reasonable, the sentence connection is natural, and there is no logical conflict), and finally, a sentiment-annotated response text with clear structure, accurate sentiment annotation, and fluent expression is generated. This text achieves clear transmission of sentiment information through sentiment hierarchical integration and visualization annotation, providing data support for subsequent sentiment-based interaction optimization.
[0046] Optionally, step 42 includes: step 421, performing text-to-pinyin processing on the sentiment-annotated response text and annotating prosodic pause information to generate a pinyin prosodic sequence; step 422, fusing temporal context frame features to determine speech rhythm parameters and fusing sentiment dynamic spectrum features to determine basic intonation parameters; step 423, performing speech synthesis based on the pinyin prosodic sequence, speech rhythm parameters and basic intonation parameters to generate an initial speech stream.
[0047] Preferably, the specific implementation process of step 421 is as follows: taking the sentiment-annotated response text as the processing object, based on the sales scenario text-to-pinyin conversion rules and prosodic pause annotation algorithm, the text is converted to pinyin and prosodic pause information is annotated to generate a pinyin prosodic sequence; specifically, the sentiment-annotated response text is the structured text output in step 41 (containing "sentiment type-sentiment intensity" combined tags and scenario-based sentence units). First, a sales scenario-specific text-to-pinyin mapping library is constructed, covering standard pinyin annotations for commonly used Chinese characters, sales scenario professional terms, numbers, and units. For polyphonic characters, a unique pinyin is determined based on the semantics of the sales scenario; a character-by-character pinyin conversion algorithm is used to perform pinyin conversion on each character in the text, retaining tone information and neutral tone markers (marked as 0), generating a basic pinyin sequence; combined with the sentiment tags and punctuation marks in the text, a sales scenario prosodic pause annotation rule is designed: for sentence units with sentiment tags of "positive-strong", the internal pause interval is shortened. (Pace between words: 0.05-0.08 seconds; pause between phrases: 0.1-0.15 seconds), with the final pause extended to 0.3-0.4 seconds; for "neutral-neutral" sentence units, maintain even pause intervals (0.08-0.1 seconds between words, 0.15-0.2 seconds between phrases, 0.2-0.3 seconds at the end of the sentence); for "introductory-strong" sentence units, extend the final pause to 0.4-0.5 seconds, and add a short pause (0.1-0.1 seconds) after introductory words (such as "suggest" or "may"). (12 seconds); a first-level pause (0.3-0.5 seconds) is marked after terminating punctuation marks, a second-level pause (0.15-0.25 seconds) is marked after pause-type punctuation marks, and there is no pause within special punctuation marks; pause duration markers and rhythmic stress markers are added to the basic pinyin sequence according to the annotation rules, and finally a pinyin rhythmic sequence containing pinyin, tone, pause duration, and stress information is generated. This sequence is different from the traditional general pinyin sequence through contextualized emotion adaptation annotation, and achieves precise binding of rhythm and emotion.
[0048] Preferably, the specific implementation process of step 422 is as follows: taking temporal context frame features and emotional dynamic spectrum features as core inputs, and based on the sales scenario parameter mapping rules and dual feature fusion algorithm, the speech rhythm parameters and basic intonation parameters are determined respectively; specifically, the temporal context frame features are derived from the speech signal temporal analysis results of step 12 (frame-level features divided according to a fixed time window, each frame contains temporal features such as short-time energy and zero-crossing rate, frame length 20-30 milliseconds, frame shift 10-15 milliseconds), and the emotional dynamic spectrum features are the emotional feature matrix output in step 13 (row dimension is frame order). The column dimension represents the quantified value of emotional intensity, reflecting the temporal changes in speech emotion. For determining the speech rhythm parameters, a long short-term memory network is used to construct a temporal feature modeling structure. Inputting temporal context frame features, the system captures inter-frame temporal dependencies through hierarchical processing of the input layer, hidden layer (containing 128-256 neurons), and output layer, outputting a frame-level speech rate coefficient (0.8-1.2 times the baseline speech rate). This coefficient is adjusted based on the sales scenario's sentence type (greeting, declarative, guiding sentence): greeting speech rate coefficient 1.0-1.1, declarative sentence 0.9-1.0, guiding sentence... 1.1-1.2; Combining the speech rate coefficient with frame length and frame shift, speech rhythm parameters including speech rate, pause interval, and sentence length are generated; For the determination of basic intonation parameters, the peak value and slope of emotional intensity in the emotional dynamic spectrum features are extracted to construct an emotional-intonation mapping table: for positive emotional intensity peak values above 0.8, the fundamental frequency (F0) range is 200-250Hz, and the slope of change above 0.6 corresponds to an intonation fluctuation amplitude of 0.3-0.4; for neutral emotional intensity peak values of 0.4-0.6, the fundamental frequency is 180-200Hz, and the slope of change is 0.3-0.5. The pitch fluctuation range is 0.15-0.25; the peak value of the guiding emotion intensity is 0.5-0.7, corresponding to the fundamental frequency of 190-220Hz, and the slope of change is 0.4-0.6, corresponding to the pitch fluctuation range of 0.2-0.3. A weighted fusion algorithm is adopted to fuse the emotional dynamic spectrum features with the sentence emotion label weights (0.7 for strong emotion and 0.5 for medium emotion) to generate basic pitch parameters that include fundamental frequency, pitch fluctuation, and tone change. This parameter modeling method, through the deep fusion of temporal and emotional features, is different from single feature modeling and achieves a high degree of adaptation between parameters and scene emotion.
[0049] Preferably, in the specific technical implementation of step 422, the initially generated speech rhythm parameters and basic intonation parameters are used as processing objects. Based on the sales scenario parameter consistency verification rules and dynamic optimization strategies, the two types of parameters are collaboratively optimized and adjusted. Specifically, a parameter collaborative verification matrix is constructed, with the row dimension being the sentence unit number, the column dimension being the rhythm parameters (speech rate, pauses) and the intonation parameters (fundamental frequency, fluctuations), and the intersection element being the parameter fit degree (0-1 range). If the fit degree of the rhythm parameters and intonation parameters of a certain sentence unit is lower than 0.65 (such as the "strongly positive" sentence having a slow speech rate and a low fundamental frequency), then it is adjusted based on the scenario fit rules: the speech rate coefficient is increased by 0.1-0. 2. Simultaneously increase the base frequency by 10-20Hz; combine the target scene labels from step 23 to optimize parameters: online consultation scene parameters are biased towards a brisk pace (speech rate coefficient 1.0-1.1, base frequency 200-220Hz), while offline experience scene parameters are biased towards a steady pace (speech rate coefficient 0.9-1.0, base frequency 180-200Hz); perform temporal smoothing on the optimized parameters, and use a moving average algorithm (window size 3-5 frames) to eliminate parameter mutations and ensure continuous changes in speech rate and intonation; finally, generate collaboratively adapted speech rhythm parameters and basic intonation parameters. This optimization process improves the rationality and consistency of parameters through parameter collaborative verification and scene adaptation adjustment.
[0050] Preferably, the specific implementation process of step 423 is as follows: taking the Pinyin prosodic sequence, optimized speech rhythm parameters, and basic intonation parameters as input, speech synthesis is performed based on the sales scenario speech synthesis model and multi-element collaborative synthesis algorithm to generate an initial speech stream; specifically, the Pinyin prosodic sequence provides the pronunciation basis (Pinyin, tone, pause), the speech rhythm parameters control the time dimension (speech rate, pause interval), and the basic intonation parameters control the frequency dimension (fundamental frequency, fluctuation); a sales scenario-specific speech synthesis model is constructed, which includes a three-layer structure: an acoustic feature generation layer, a prosodic adaptation layer, and a speech waveform generation layer; the acoustic feature generation layer inputs three types of parameters, focuses on core parameters (emotion-related parameters weight 0.6, prosodic parameters weight 0.3, and rhythm parameters weight 0.1) through an attention mechanism, and generates frame-level Mel-frequency spectral features (40-80 dimensions, with row dimension being the frame number and column dimension being the spectral amplitude, reflecting the speech's...). The frequency distribution layer and the prosody adaptation layer are based on the prosody templates of the sales scenario speech library (containing 50-100 hours of real-person speech samples from sales scenarios). The prosody calibration of the Mel spectrum features makes the synthesized speech fit the expression habits of the sales scenario. The speech waveform generation layer uses a waveform prediction algorithm to convert the calibrated Mel spectrum features into a continuous speech waveform signal (sampling rate 16kHz-48kHz, bit depth 16-24 bits). Real-time collaborative verification is performed during the synthesis process to ensure the accuracy of pinyin pronunciation (no misreading or omission), rhythm parameter matching degree (speech rate and pauses meet the requirements of the scenario), and intonation parameter adaptation degree (emotion and intonation are consistent). Finally, an initial speech stream with a duration of 10-30 seconds, natural emotional expression, and clear semantic transmission is generated. This speech stream has undergone multi-parameter collaborative synthesis and scenario-based model calibration, which is different from the general speech synthesis results and achieves a high degree of unity of emotion, prosody and semantics in the sales scenario.
[0051] Optionally, step 43 includes: step 431, extracting intonation amplitude adjustment coefficients from emotional dynamic spectrum features and extracting rhythm calibration parameters from temporal context frame features; step 432, adjusting the intonation intensity of the initial speech stream frame by frame based on the intonation amplitude adjustment coefficients; and step 433, performing rhythm synchronization calibration on the intonation-adjusted speech stream based on the rhythm calibration parameters to generate a real-time response speech stream.
[0052] Preferably, the specific implementation process of step 431 is as follows: taking the emotional dynamic spectrum features and temporal context frame features as the processing objects, based on the sales scenario parameter extraction algorithm and frame-level feature parsing strategy, the intonation amplitude adjustment coefficient and rhythm calibration parameters are extracted; specifically, the emotional dynamic spectrum features are the emotional feature matrix output in step 13 (the row dimension is the frame number, the column dimension is the emotional intensity quantification value, and the frame length is 20-30 milliseconds), focusing on the temporal change trend of emotional intensity in the matrix, using the frame-level amplitude parsing algorithm, calculating the ratio of the emotional intensity of each frame to the global average emotional intensity, combined with step 4 The sentence sentiment label strength (strong sentiment weight 0.7, medium sentiment weight 0.5) generates a frame-level intonation amplitude adjustment coefficient: when the frame sentiment intensity is 1.2 times higher than the global mean and the corresponding sentence sentiment label is "positive-strong", the adjustment coefficient is set to 1.1-1.3 (strengthening intonation intensity); when the frame sentiment intensity is 0.8-1.2 times the global mean and the corresponding sentence sentiment label is "neutral-medium", the adjustment coefficient is set to 0.95-1.05 (maintaining stable intonation); when the frame sentiment intensity is less than 0.8 times the global mean and the corresponding sentence sentiment label is "guiding-strong", the adjustment coefficient is set to 0.95-1.05. When the tone is raised, the adjustment coefficient is set to 1.0-1.1 (to moderately increase the tone and enhance guidance). The adjustment coefficient is limited to a range of 0.9-1.3 to avoid tone distortion. The temporal context frame features are the frame-level temporal features output in step 12 (including short-time energy, zero-crossing rate, etc.). The normalized value of the short-time energy of each frame is extracted and the energy difference between adjacent frames is extracted. Combined with the speech rhythm parameters (speech rate, pause interval) from step 42, rhythm calibration parameters are generated: including a frame-level speech rate correction coefficient (based on the normalized value of short-time energy; the higher the energy, the larger the correction coefficient, ranging from 0.9-1.2, reflecting the speech rate). The extraction method involves several steps: matching the intensity of speech with the speech rate; calibrating the pause boundary (based on the energy difference between adjacent frames; a difference exceeding a threshold of 0.6 is considered a pause boundary, with a calibration value set at 0.1-0.2 seconds to correct pause intervals); and adjusting the frame shift compensation value (dynamically adjusted based on the speech rate coefficient; the faster the speech rate, the larger the frame shift compensation value, ranging from 5-10 milliseconds to ensure smooth inter-frame transitions). This results in frame-level aligned intonation amplitude adjustment coefficients and rhythm calibration parameters. This extraction method, through frame-level refined analysis, differs from traditional segment-level parameter extraction, achieving precise matching between adjustment parameters and speech frames.
[0053] Preferably, the specific implementation process of step 432 is as follows: taking the initial speech stream and the pitch amplitude adjustment coefficient as the processing objects, the pitch intensity of the initial speech stream is adjusted frame by frame based on the sales scenario pitch frame-by-frame adjustment algorithm and emotional intensity adaptation strategy; specifically, the initial speech stream is the continuous speech waveform signal (sampling rate 16kHz-48kHz, bit depth 16-24 bits) output in step 42. First, the initial speech stream is divided into frames with a frame length of 20-30 milliseconds and a frame shift of 10-15 milliseconds to obtain frame-level speech signals corresponding one-to-one with the pitch amplitude adjustment coefficient; pitch intensity analysis is performed on each frame of speech signal to extract the maximum, minimum, and average values of the fundamental frequency (F0) within the frame, and an adjustment model is constructed based on the pitch amplitude adjustment coefficient: if the adjustment coefficient is greater than 1.0, the maximum and average values of the fundamental frequency are increased proportionally (e.g., when the coefficient is 1.2, the maximum value of the fundamental frequency is increased by 20%, and the average value is increased by 15%), while maintaining... The minimum fundamental frequency is kept constant to avoid distortion; if the adjustment coefficient is equal to 1.0, the fundamental frequency parameter remains unchanged; if the adjustment coefficient is less than 1.0, the maximum and average fundamental frequencies are reduced proportionally (e.g., when the coefficient is 0.95, the maximum fundamental frequency is reduced by 5% and the average by 3%) to ensure that the intonation intensity is consistent with the emotional dynamic spectrum characteristics; for the frames corresponding to the core words of the "positive-strong" sentence units (e.g., "super strong" and "preferred"), additional intonation peak enhancement is added (the fundamental frequency peak is further increased by 5%-8%) to highlight the emotional emphasis; during the adjustment process, inter-frame intonation smoothing is performed, and a linear interpolation algorithm is used to correct the fundamental frequency difference between adjacent frames (the maximum difference does not exceed 30Hz) to avoid abrupt changes in intonation; the frame-level speech signals adjusted frame by frame are spliced together in the original order to generate the intonation-adjusted speech stream. This adjustment method, through precise frame-level adaptation and emotional emphasis enhancement, differs from traditional overall intonation adjustment and achieves refined transmission of emotional expression.
[0054] Preferably, in the specific technical implementation of step 432, the speech stream after tone adjustment is taken as the processing object. Based on the tone rationality verification rules and anomaly correction strategies for sales scenarios, the speech stream after tone adjustment is anomaly calibrated to generate an optimized tone-adjusted speech stream. Specifically, a reasonable tone range library for sales scenarios is constructed (fundamental frequency range 150-300Hz, tone fluctuation amplitude 0.1-0.5, conforming to the speech expression habits of adults in sales scenarios). The speech stream after tone adjustment is verified frame by frame: if the fundamental frequency of a certain frame exceeds the reasonable range (below 150Hz or above 300Hz), the fundamental frequency is corrected to the range boundary value (below 150Hz is corrected to 150Hz, above 300Hz is corrected to 300Hz); if the fundamental frequency of a certain frame exceeds the reasonable range ... If the fluctuation range exceeds 0.5, it is compressed to within 0.5 proportionally (compression ratio = 0.5 / actual fluctuation range); combined with the target scenario label optimization in step 23: the upper limit of the reasonable range of tone in online consultation scenarios is increased by 10% (330Hz) to adapt to the need for brisk expression; the lower limit of the reasonable range of tone in offline experience scenarios is decreased by 10% (135Hz) to adapt to the need for steady expression; the overall tone consistency is checked on the calibrated speech stream, and the tone standard deviation within the sentence unit is calculated (the reasonable standard deviation does not exceed 40Hz). If it exceeds this, global smoothing is performed; finally, an optimized tone-adjusted speech stream with natural tone and no abnormal distortion is generated. This calibration process improves the reliability of tone adjustment through scenario-based range adaptation and anomaly correction.
[0055] Preferably, the specific implementation process of step 433 is as follows: taking the optimized intonation-adjusted speech stream and rhythm calibration parameters as the processing objects, based on the sales scenario rhythm synchronization calibration algorithm and multi-parameter collaborative strategy, the intonation-adjusted speech stream is subjected to rhythm synchronization calibration to generate a real-time response speech stream; specifically, the optimized intonation-adjusted speech stream is a continuous speech signal after frame-level splicing, and the rhythm calibration parameters include frame-level speech rate correction coefficient, pause boundary calibration value, and frame shift compensation value; firstly, the overall speech rate of the speech stream is adjusted based on the speech rate correction coefficient: the speech stream is divided into sentence units, and the playback duration of each sentence unit is scaled according to the speech rate correction coefficient to ensure that the speech rate is consistent with the energy change of the temporal context frame features; for the pause position corresponding to the pause boundary calibration value, the pause interval is adjusted: if the original pause interval is lower than the calibration value, it is extended to the calibration value (e.g., if the original pause is 0.2 seconds and the calibration value is 0.3 seconds, it is extended by 0.1 seconds); if the original pause interval is higher than the calibration value, it is shortened to the calibration value (e.g., if the original pause is 0.4 seconds and the calibration value is 0.3 seconds, it is shortened by 0.1 seconds), adapting to the sales scenario. Scene interaction rhythm; Frame shift compensation value adjustment for inter-frame transitions: For sentence units with faster speech rate (speech rate correction coefficient ≥ 1.1), increase the frame shift compensation value (8-10 milliseconds) to reduce inter-frame overlap; For sentence units with slower speech rate (speech rate correction coefficient ≤ 0.95), decrease the frame shift compensation value (5-7 milliseconds) to enhance inter-frame coherence; Rhythm-intonation co-verification is performed during calibration, and a co-verification matrix is constructed (row dimension is sentence unit number, column dimension is speech rate, pause, and intonation), with the intersection element representing the co-adaptation degree. (≥0.7 is considered acceptable). If the coordination and adaptation of a certain sentence unit is lower than 0.7, the speech rate correction coefficient and intonation adjustment coefficient are adjusted simultaneously. The calibrated speech stream is then spliced and standardized in format (uniform sampling rate 16kHz, bit depth 16-bit, mono) to generate a real-time response speech stream with a duration of 10-30 seconds, natural emotion, smooth rhythm, and clear semantics. This speech stream, through deep coordination of rhythm and intonation, is different from traditional single-dimensional calibration and achieves naturalness and adaptability of speech expression in sales scenario interactions.
[0056] The above are merely preferred embodiments of the present invention and are not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A sales agent voice dialogue method, characterized in that, include: Step 1: Perform temporal context frame association parsing and emotional dynamic spectrum quantization modeling on the real-time voice data stream representing user needs to generate temporal context frame features and emotional dynamic spectrum features, and form a dynamic feature profile of user needs based on these features. Step 2: Perform hierarchical feature denoising and reconstruction and intent posterior probability matching on the dynamic feature profile of user needs to determine the user core need feature basis vector. Based on the user core need feature basis vector, perform scene feature anchoring mapping to assign scene feature identification codes and generate scene-based need positioning feature topology accordingly. Step 3: Call the product attribute knowledge graph to perform weighted association mapping of product attribute features and generate recommendation semantic logic on the feature topology of scenario-based demand positioning, and generate customized recommendation response text; Step 4: Based on the customized recommended response text, integrate temporal context frame features and emotional dynamic spectrum features to perform text-to-speech conversion and emotional tone adaptation modulation to generate a real-time response speech stream. The real-time response speech stream includes real recommended content that matches product attributes to meet user needs.
2. The sales agent voice dialogue method according to claim 1, characterized in that, Step 1 includes: Step 11: Perform frame segmentation and feature pre-extraction on the real-time voice data stream representing user needs to generate a voice frame feature sequence; Step 12: Perform adjacent frame context association analysis and cross-frame semantic mapping on the speech frame feature sequence to generate temporal context frame features; Step 13: Decompose the emotional dimension and model the dynamic amplitude of the speech frame feature sequence to generate emotional dynamic spectrum features. Integrate the temporal context frame features and emotional dynamic spectrum features to form a dynamic feature profile of user needs.
3. The sales agent voice dialogue method according to claim 2, characterized in that, Step 13 includes: Step 131: Perform emotion-related feature separation on the speech frame feature sequence to generate original emotion features; Step 132: Decompose the original emotional features in multiple dimensions and track the amplitude changes to generate dynamic emotional amplitude data; Step 133: Perform spectral modeling on the emotional dynamic amplitude data to generate emotional dynamic spectrum features, and integrate temporal context frame features with emotional dynamic spectrum features to form a dynamic feature profile of user needs.
4. The sales agent voice dialogue method according to claim 1, characterized in that, Step 2 includes: Step 21: Perform multi-level noise filtering and feature dimension reconstruction on the dynamic feature profile of user needs to generate a denoised feature profile; Step 22: Perform intent feature library comparison and posterior probability calculation on the denoised feature construct to determine the user's core demand feature basis vector; Step 23: Perform scene feature library anchoring matching and identification encoding on the user core demand feature basis vectors to assign scene feature identification codes and generate scene-based demand positioning feature topology accordingly.
5. The sales agent voice dialogue method according to claim 4, characterized in that, Step 22 includes: Step 221: Construct an intent feature library, which contains feature templates corresponding to preset demand intents; Step 222: Calculate the cosine similarity between the denoised feature construct and each feature template to obtain multiple similarity values; Step 223: Normalize multiple similarity values to convert them into posterior probabilities, and select the feature template corresponding to the maximum posterior probability to determine the user's core demand feature basis vector.
6. The sales agent voice dialogue method according to claim 4, characterized in that, Step 23 includes: Step 231: Construct a scene feature library, which includes feature identifiers and encoding rules for preset scenes; Step 232: Anchor matching is performed between the user's core demand feature basis vector and the scene features in the scene feature library to determine the target scene; Step 233: Assign scene feature identification codes according to the coding rules corresponding to the target scene, merge the scene feature identification codes with the user core demand feature basis vectors, and generate scene-based demand positioning feature topology accordingly.
7. The sales agent voice dialogue method according to claim 1, characterized in that, Step 3 includes: Step 31: Extract the requirement attribute dimensions and quantify the features of the scenario-based requirement location feature topology to generate a requirement attribute feature set; Step 32: Call the product attribute knowledge graph to perform weighted association calculation and matching degree ranking on the demand attribute feature set and product attribute features to generate a set of matching products; Step 33: Based on the attribute information of the matching product set and the topology of the contextualized demand location features, construct the recommendation semantic logic and organize the text to generate customized recommendation response text.
8. The sales agent voice dialogue method according to claim 7, characterized in that, Step 32 includes: Step 321: Call the product attribute knowledge graph to extract the product attribute features and build a product feature index library; Step 322: Perform a weighted correlation calculation between the demand attribute feature set and the features in the product feature index library to obtain the matching score for each product; Step 323: Sort the matching scores in descending order and select the top-ranked products to form a set of matching products.
9. The sales agent voice dialogue method according to claim 1, characterized in that, Step 4 includes: Step 41: Perform semantic sentence segmentation and sentiment annotation on the customized recommendation response text to generate sentiment-annotated response text; Step 42: Based on the sentiment-annotated response text, integrate temporal context frame features and sentiment dynamic spectrum features to perform text-to-speech conversion and basic intonation generation to generate an initial speech stream; Step 43: Adjust the emotional tone amplitude and adapt the context rhythm of the initial speech stream to generate a real-time response speech stream.
10. The sales agent voice dialogue method according to claim 9, characterized in that, Step 43 includes: Step 431: Extract the intonation amplitude adjustment coefficient from the emotional dynamic spectrum features and extract the rhythm calibration parameter from the temporal context frame features; Step 432: Adjust the pitch intensity of the initial speech stream frame by frame based on the pitch amplitude adjustment coefficient; Step 433: Perform rhythm synchronization calibration on the intonation-adjusted speech stream based on the rhythm calibration parameters to generate a real-time response speech stream.