An AI voice interaction method and system
By obtaining customer information and preference records from the historical sales interaction database of automotive OEMs, and using customer classification rules and acoustic physical parameters to generate personalized speech synthesis, the problem of the inability to accurately match personalized needs in the existing system is solved, and the naturalness and emotional expressiveness are improved, thus enhancing the customer experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2026-04-07
AI Technical Summary
Existing intelligent car sales systems cannot accurately match the personalized expression needs at different stages of the car purchase process, and the voice output is mechanical and stiff, affecting the user experience.
By obtaining customer type information and car purchase preference records from the historical sales interaction database of automobile OEMs, customer type identifiers and preference feature vectors are generated using preset customer classification rules. Combined with car sales script templates, car sales script text containing vehicle function keywords is generated. Speech synthesis is driven by acoustic physical parameters to achieve the representation and multimodal display of speech prosody characteristics.
It enables refined classification of customer needs and personalized voice interaction, improving the targeting and effectiveness of sales communication, enhancing the naturalness and emotional expressiveness of voice output, and increasing customer immersion and trust.
Smart Images

Figure CN121075312B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine learning, and particularly relates to an AI voice interaction method and system. BACKGROUND
[0002] With the continuous upgrading of intelligent automobile sales services, customers' demand for personalized car buying experience is increasing. In the actual sales process, sales personnel need to quickly adjust the communication strategy according to the background characteristics of customers.
[0003] Currently, some intelligent sales assistance systems attempt to introduce a dialogue recommendation mechanism based on big data analysis and natural language processing technology. Generally, by analyzing customers' historical purchase records and interaction behaviors, a general customer portrait model is constructed, and preliminary dialogue suggestions are generated in combination with fixed templates or simple text generation algorithms. However, these existing solutions still have many deficiencies, for example, the customer classification method is relatively rough, which cannot accurately match the personalized expression needs at different car buying stages, resulting in the generated dialogue content lacking pertinence; the voice synthesis module often uses uniform prosody parameter settings, which is difficult to adjust the voice rhythm and emotional color, causing the voice output to be mechanical and harsh, affecting the user experience, etc. SUMMARY
[0004] The present application provides an AI voice interaction method and system to solve the problems in the prior art that cannot accurately match the personalized expression needs at different car buying stages, the voice output is mechanical and harsh, and the user experience is insufficient.
[0005] In a first aspect, the present application provides an AI voice interaction method, comprising:
[0006] obtaining customer type information and customer car buying preference records from a historical sales interaction database of an automobile manufacturer;
[0007] classifying the customer type information using a preset customer classification rule to generate a customer type identifier, the customer type identifier including a first-purchase type, a replacement type, and an additional purchase type;
[0008] analyzing the customer car buying preference records to generate a preference feature vector, combining the customer type identifier and a preset car sales dialogue template to generate a car sales dialogue text containing vehicle function keywords;
[0009] processing the car sales dialogue text to obtain acoustic physical parameters representing voice prosody characteristics;
[0010] generating synthesized speech associated with the vehicle function keywords using the acoustic physical parameters;
[0011] In the sales process, the synthesized voice is played, and a three-dimensional display of a car function corresponding to the car sales script text is triggered synchronously.
[0012] Optionally, the customer type information is classified by using a preset customer classification rule to generate a customer type identifier, the customer type identifier including a first-purchase type, a replacement type, and an additional-purchase type, and the customer type identifier includes:
[0013] The purchase frequency and consultation text data are extracted from the customer type information.
[0014] The purchase frequency and a preset threshold value are matched by using a main matching rule in the preset customer classification rule to generate an initial customer type identifier, the initial customer type identifier including a first-purchase type, a replacement type, and an additional-purchase type.
[0015] Key attention features in the consultation text data are detected by using an auxiliary rule in the preset customer classification rule.
[0016] The initial customer type identifier is corrected according to the key attention features to generate a customer type identifier.
[0017] Optionally, the customer car purchase preference record is analyzed to generate a preference feature vector, the customer type identifier and a preset car sales script template are combined to generate a car sales script text containing vehicle function keywords, and the car sales script text includes:
[0018] The target script template is selected from the preset car sales script template according to the customer type identifier.
[0019] The customer car purchase preference record is analyzed to generate a preference feature vector, the preference feature vector including a function attention point and a weight of the function attention point.
[0020] The function attention point and a placeholder in the target script template are matched to generate an initial script text.
[0021] The vehicle function keywords are selected from a preset car sales term library according to the weight of the function attention point, the keyword placeholder in the initial script text is replaced with the vehicle function keywords, and a car sales script text containing vehicle function keywords is obtained.
[0022] Optionally, the car sales script text is processed to obtain acoustic physical parameters representing prosodic characteristics of speech, and the acoustic physical parameters include:
[0023] The car sales script text is divided according to the vehicle function keywords to obtain a plurality of function unit segments.
[0024] Functional unit segments containing the vehicle function keywords are marked as strong rhythmic segments, and functional unit segments that do not contain the vehicle function keywords are marked as weak rhythmic segments.
[0025] Using a pre-defined car sales rhythm rule library, determine the high-intensity rhythm parameters corresponding to the strong rhythm segment and the low-intensity rhythm parameters corresponding to the weak rhythm segment;
[0026] By combining all high-intensity prosodic parameters and low-intensity prosodic parameters, acoustic physical parameters characterizing the prosodic properties of speech are obtained.
[0027] Optionally, all high-intensity prosodic parameters and low-intensity prosodic parameters are combined to obtain acoustic physical parameters characterizing the prosodic properties of speech, including:
[0028] Based on the temporal order of the functional unit segments in the car sales script text, assign a positional weight coefficient to each functional unit segment;
[0029] Based on the preset car sales weight allocation rules, a first proportional coefficient is assigned to the strong rhythmic segment, and a second proportional coefficient is assigned to the weak rhythmic segment.
[0030] Each high-intensity prosodic parameter is multiplied by its corresponding position weight coefficient and the first proportional coefficient to obtain multiple enhanced prosodic parameters;
[0031] Each low-intensity prosodic parameter is multiplied by its corresponding position weight coefficient and the second proportional coefficient to obtain multiple weakened prosodic parameters;
[0032] According to the original temporal order of the functional unit segments, all enhancement parameters and all weakening parameters are combined to obtain acoustic physical parameters that characterize the prosodic properties of speech.
[0033] Optionally, using the acoustic physical parameters, synthesized speech associated with the vehicle function keywords is generated, including:
[0034] The acoustic physical parameters are converted into speech synthesis control parameters, which include pitch adjustment values and speech rate adjustment values.
[0035] Based on the text content of all functional unit segments, generate corresponding initial speech waveforms, and adjust all initial speech waveforms using the pitch adjustment value and speech rate adjustment value to generate intermediate speech waveforms corresponding to each functional unit segment. The physical duration of each intermediate speech waveform is used as the actual duration.
[0036] Calculate the target duration of each functional unit segment based on its text length and prosody type;
[0037] Based on the actual duration and the target duration, all intermediate speech waveforms are subjected to time-domain scaling to generate corresponding aligned speech waveforms.
[0038] According to the original temporal sequence of the functional unit segments, all aligned speech waveforms are spliced together to generate synthesized speech.
[0039] Optionally, based on the actual duration and the target duration, all intermediate speech waveforms are subjected to temporal scaling to generate corresponding aligned speech waveforms, including:
[0040] Based on the vehicle function keyword type of each functional unit segment, a preset automotive terminology pronunciation benchmark table is queried to determine the syllable density weight coefficient.
[0041] The ratio of the actual duration to the target duration is used as the basic scaling factor. The syllable density weight coefficient is multiplied by the basic scaling factor to obtain the dynamic scaling parameters corresponding to each functional unit segment.
[0042] Based on all dynamic scaling parameters, the intermediate speech waveform corresponding to each functional unit segment is segmented and scaled to obtain the corresponding aligned speech waveform, so that the actual duration of each functional unit segment matches the corresponding target duration.
[0043] Secondly, this application provides an AI voice interaction system, including:
[0044] The acquisition module is used to retrieve customer type information and customer car purchase preference records from the historical sales interaction database of the car manufacturer;
[0045] The classification module is used to classify the customer type information using preset customer classification rules and generate customer type identifiers, which include first-time purchase type, replacement type and additional purchase type.
[0046] The parsing module is used to parse the customer's car purchase preference record, generate a preference feature vector, and combine it with the customer type identifier and the preset car sales script template to generate car sales script text containing vehicle function keywords.
[0047] The processing module is used to process the car sales script text to obtain acoustic physical parameters that characterize the prosodic features of speech.
[0048] The association module is used to generate synthesized speech associated with the vehicle function keywords using the acoustic physical parameters;
[0049] The display module plays the synthesized voice during the sales process and simultaneously triggers a 3D display of the car's functions corresponding to the car sales script text.
[0050] Thirdly, this application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement an AI voice interaction method as described in the first aspect above.
[0051] Fourthly, this application provides a computer storage medium storing a computer program, which, when executed by a computer, implements an AI voice interaction method as described in the first aspect.
[0052] This application retrieves customer type information and customer purchase preference records from the historical sales interaction database of automotive OEMs; classifies the customer type information using preset customer classification rules to generate customer type identifiers, including first-time buyers, replacement buyers, and additional purchase buyers; parses the customer purchase preference records to generate preference feature vectors, and combines the customer type identifiers with preset automotive sales script templates to generate automotive sales script text containing vehicle function keywords; processes the automotive sales script text to obtain acoustic physical parameters characterizing speech prosody; uses the acoustic physical parameters to generate synthesized speech associated with the vehicle function keywords; during the sales process, plays the synthesized speech and simultaneously triggers a 3D display of automotive functions corresponding to the automotive sales script text. The technical solution provided by this application systematically collects historical data, providing a data foundation for subsequent customer classification and script generation, enhancing the understanding of customer needs. It achieves refined customer classification, improving the matching degree between sales scripts and actual customer needs. It transforms structured data into executable personalized script content, improving the targeting and effectiveness of sales communication. By analyzing text semantics and context, key prosodic features of speech expression are extracted, laying the foundation for generating natural speech. Semantic consistency between speech output and vehicle function display content is achieved, enhancing the emotional tone and naturalness of speech expression. A multimodal interactive environment is constructed to improve customer immersion and the professionalism of sales presentations. Specifically, this application achieves a high degree of matching between speech output and content semantics through refined modeling of speech prosody and time alignment optimization, effectively avoiding problems such as uneven speech rate and abrupt rhythm in traditional speech synthesis, thereby improving the naturalness and emotional expressiveness of speech expression. It also solves the technical bottlenecks of existing systems such as mechanical speech and lack of context adaptability, making AI voice interaction closer to the expression style of real sales personnel, enhancing the immersion and trust level of customer dialogue experience.
[0053] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0055] Figure 1 A flowchart of an AI voice interaction method provided in this application is shown;
[0056] Figure 2 A schematic diagram of the structure of an AI voice interaction system provided in this application is shown;
[0057] Figure 3 A schematic diagram of the structure of a computing device provided in this application is shown. Detailed Implementation
[0058] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0059] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.
[0060] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0061] In response to the increasing demand for personalized experiences from customers in the current smart car sales scenario, traditional script recommendation systems fall short in terms of customer segmentation granularity, script generation accuracy, and naturalness of voice interaction. While existing technologies can build customer profiles and generate basic script content based on historical data, their customer classification methods are relatively general, making it difficult to differentiate the expressive needs at different stages of the car purchase process, resulting in a lack of targeted content generation. Furthermore, speech synthesis modules typically use fixed prosodic parameters, failing to dynamically adjust the rhythm and emotional tone of the voice based on vehicle function keywords, resulting in stiff and mechanical speech output that negatively impacts user experience and sales communication effectiveness. To address this, this application proposes an AI voice interaction method. This method achieves accurate matching of script content through a refined customer classification mechanism combined with multi-dimensional preference feature modeling. It also introduces acoustic physical parameter-driven speech synthesis technology to make the voice output more human-like and context-adaptive. Simultaneously, it integrates 3D functional displays to form a multimodal sales demonstration, thereby effectively improving the intelligence level of the sales process and enhancing customer immersion.
[0062] Figure 1 A flowchart of an AI voice interaction method is provided as an embodiment of this application, such as... Figure 1 As shown, the method includes:
[0063] Step 101: Obtain customer type information and customer car purchase preference records from the historical sales interaction database of the car manufacturer.
[0064] In this step, the historical sales interaction database refers to the distributed relational database deployed by the automaker, which stores the transaction history of dealer customer interactions, including consultation record tables, test drive feedback tables, and order configuration tables. Customer type information refers to the set of structured attributes extracted from the database, including the number of purchases, consultation type, and historical vehicle models. Customer purchase preference records refer to the set of semi-structured behavioral data, including consultation question text, configuration selection records, and attention tags.
[0065] In this embodiment of the application, structured customer type information and semi-structured customer car purchase preference records are extracted by accessing the historical sales interaction database dedicated to the car manufacturer.
[0066] Step 102: Classify the customer type information using preset customer classification rules and generate customer type identifiers, including first-time purchase type, replacement type, and additional purchase type.
[0067] In this step, the preset customer classification rules refer to the hierarchical decision rule set. The customer type identifier refers to the classification result label, including first-time purchaser, replacement purchaser, and add-on purchaser.
[0068] In this embodiment, a preset customer classification rule is loaded into memory. The preset customer classification rule consists of two layers: a main rule and a secondary rule. The main rule is based on a purchase frequency threshold, and the secondary rule is based on keywords in the consultation text. First, the main rule is run. If there are no purchases, the customer is marked as a first-time purchaser; if there are no purchases, the customer is marked as a replacement purchaser; if there are purchases, the customer is marked as an upgrade purchaser, thus obtaining an initial identifier. Second, the secondary rule is run. If there are replacement keywords of the same level in the consultation text, the customer is marked as a replacement purchaser; if there are keywords indicating an upgrade, the customer is marked as an upgrade purchaser, thus obtaining a customer type identifier.
[0069] Step 103: Analyze the customer's car purchase preference record to generate a preference feature vector. Combine the customer type identifier and the preset car sales script template to generate a car sales script text containing vehicle function keywords.
[0070] In this step, the preference feature vector is numerically represented to characterize customer preferences, with dimensions including focus on power performance, intensity of demand for intelligent driving features, and weighting of energy efficiency preferences. Pre-set car sales script templates refer to a text library with placeholders, grouped by customer type. The car sales script text is then customized into a sales script.
[0071] In this embodiment, based on the customer type identifier, a target template is selected from the preset car sales script templates, the customer's car purchase preference record is parsed, and three types of preference feature vectors, namely power performance, intelligent driving and energy efficiency, are extracted. The three types of preference feature vectors are matched with the placeholders of the target template, and then the keywords are sorted according to their weights to generate car sales script text containing vehicle function keywords.
[0072] Step 104: Process the car sales script text to obtain acoustic physical parameters that characterize the prosodic features of the speech.
[0073] In this step, acoustic physical parameters refer to a sequence of timing parameters, with each time point including the fundamental frequency value, duration scaling factor, and energy gain value.
[0074] In this embodiment, the car sales text is segmented into functional unit fragments according to vehicle function keywords; fragments containing technical parameters are analyzed and marked as strong prosodic fragments, and descriptive statements are marked as weak prosodic fragments; a preset car sales prosodic rule library is queried to assign high-intensity parameters to strong fragments and low-intensity parameters to weak fragments; all parameters are aggregated by temporal position weight to obtain acoustic physical parameters characterizing the prosodic characteristics of speech.
[0075] Step 105: Using the acoustic physical parameters, generate synthesized speech associated with the vehicle function keywords.
[0076] In this step, the synthesized speech refers to the audio stream labeled with vehicle function keywords.
[0077] In this embodiment, acoustic physical parameters are converted into speech synthesis control parameters. An initial speech waveform is generated based on functional unit segments, and the pitch and speech rate are adjusted. The target duration of each functional unit segment is calculated for differentiated temporal scaling processing. Finally, synthesized speech is generated by splicing the segments in time sequence to ensure accurate embedding of acoustic tags associated with vehicle function keywords.
[0078] Step 106: During the sales process, play the synthesized voice and simultaneously trigger a 3D display of the car's functions corresponding to the car sales script text.
[0079] In this embodiment, synthesized speech is transmitted to the display terminal, and vehicle function keywords and timestamps in the car sales script are parsed simultaneously. When the playback progress matches the keyword timestamp, the corresponding component in the 3D car model is automatically displayed in 3D.
[0080] This application embodiment generates customized sales scripts by dynamically matching customer classification rules and script templates, and combines prosodic parameter control and multimodal synchronization technology to achieve personalized voice interaction in the car sales scenario.
[0081] This application provides a specific embodiment. Step 102 involves classifying the customer type information using preset customer classification rules to generate customer type identifiers. The customer type identifiers include first-time purchase type, replacement type, and add-on purchase type. Specifically, this includes the following steps:
[0082] Step 201: Extract the number of purchases and consultation text data from the customer type information.
[0083] In this step, "purchase count" refers to the cumulative number of vehicle purchase transactions completed by the customer within the OEM's dealer network, stored as an integer. "Consultation text data" refers to the text records of conversations generated by the customer through online customer service or in-store consultations, including descriptions of vehicle configuration requirements and purchase intentions.
[0084] In this embodiment of the application, the number of purchases and the consultation text data string field are extracted by parsing the data structure of customer type information to obtain the number of purchases and consultation text data.
[0085] Step 202: Using the main matching rule in the preset customer classification rules, match the number of purchases with the preset threshold to generate an initial customer type identifier, which includes first-time purchase type, replacement type and additional purchase type.
[0086] In this step, the main matching rule refers to the classification logic based on a purchase frequency threshold. The preset threshold refers to the critical value set in the main rule. The initial customer type identifier refers to the preliminary classification result generated by the main rule matching.
[0087] In this embodiment of the application, the main matching rule in the customer classification rules is loaded. This rule defines the mapping logic between the number of purchases and a preset threshold. When there are no purchases, it matches the first purchase type; when there is one purchase, it matches the replacement type; and when there are multiple purchases, it matches the additional purchase type. The number of purchases is matched with the threshold by a comparison operator, and the initial customer type identifier is output. For example, when the number of purchases is 1, the rule engine determines and generates a replacement type identifier.
[0088] Step 203: Detect key attention features in the consultation text data using the auxiliary rules in the preset customer classification rules.
[0089] In this step, auxiliary rules refer to correction rules based on text keywords.
[0090] In this embodiment of the application, auxiliary rules in the preset customer classification rules are invoked to scan consultation text data and obtain key attention features; for example, when replacement keywords of the same level of vehicle model are detected, such as old car trade-in or model upgrade, a replacement feature mark is output; when power system upgrade keywords are detected, such as adding a motor or improving range, an additional purchase feature mark is output; these marks are the key attention features.
[0091] Step 204: Based on the key features of interest, modify the initial customer type identifier to generate a new customer type identifier.
[0092] In this step, key features of interest refer to intent markers extracted from the consultation text, including features reflecting replacement demand for same-class models and features reflecting demand for additional purchases of powertrain upgrades.
[0093] In this embodiment, a correction mechanism is triggered based on the type of the key feature of interest. If the key feature of interest is a replacement feature, the initial customer type identifier is overwritten as a replacement feature; if it is an add-on feature, it is overwritten as an add-on feature; if there is no feature, the initial identifier is retained and a generated customer type identifier is output.
[0094] This application's embodiments reduce the error rate of traditional single-purchase-time classification by employing a two-level verification mechanism of primary and secondary rules. When a customer's historical behavior conflicts with their current intent, the secondary rules can correct the classification results in real time, ensuring that subsequent sales scripts accurately match the customer's actual needs and avoiding misallocation of sales resources.
[0095] This application provides a specific embodiment. Step 103 involves parsing the customer's car purchase preference record to generate a preference feature vector. This vector, combined with the customer type identifier and a preset car sales script template, generates car sales script text containing vehicle function keywords. Specifically, this includes the following steps:
[0096] Step 301: Select the target sales script template from the preset car sales script templates based on the customer type identifier.
[0097] In this step, the target sales script template refers to a predefined text framework based on customer type, including the financial solution module for first-time purchase templates, the residual value assessment module for replacement templates, and the multi-vehicle collaboration module for add-on purchase templates.
[0098] In this embodiment, by querying a preset car sales script template, the customer type identifier is used as the index key. When the customer type identifier is a first-time purchaser, the first-time purchase template is called; when it is a replacement purchaser, the replacement template is called; and when it is an additional purchaser, the additional purchase template is called, so as to determine the target script template.
[0099] Step 302: Analyze the customer's car purchase preference record to generate a preference feature vector, which includes functional concerns and the weights of the functional concerns.
[0100] In this step, functional concerns refer to the core needs dimensions reflected in the customer's car purchase preference record. The weight of functional concerns refers to the numerical value that quantifies the intensity of the customer's needs.
[0101] In this embodiment, a text analysis model based on an attention mechanism is used to process customer car purchase preference records. First, three types of functional concerns are identified: power performance, intelligent driving, and energy efficiency. Second, the frequency of keyword occurrences is divided by the total number of words to obtain the weight of each functional concern. Then, the semantic strength coefficient is combined with the adjustment to generate a preference feature vector.
[0102] Step 303: Match the functional focus points with the placeholders in the target script template to generate the initial script text.
[0103] In this step, the initial text refers to the intermediate text that has completed placeholder matching but has not yet inserted specific keywords.
[0104] In this embodiment, pattern matching is performed between the functional concerns in the preference feature vector and the placeholders in the target script template to generate initial script text; for example, the power performance concerns are matched with the power parameter placeholders in the template. Initial script text without specific keywords is generated through string replacement operations.
[0105] Step 304: Based on the weight of the functional focus points, select vehicle function keywords from the preset car sales terminology library, replace the keyword placeholders in the initial sales script with the vehicle function keywords, and obtain a car sales script containing vehicle function keywords.
[0106] In this step, the pre-defined automotive sales terminology database refers to a structured database of technical parameters. Vehicle function keywords refer to specific technical descriptions selected from the terminology database, which must simultaneously meet the requirements of matching functional focus and prioritizing by weight. Keyword placeholders are reserved markers in the target sales script template, used to locate the keyword insertion position.
[0107] In this embodiment, vehicle function keywords are sorted according to the weights in the preference feature vector; the vehicle function keywords with the highest weights are selected from the car sales terminology library. For example, a power performance weight of 0.9 corresponds to a dual-motor four-wheel drive system. The keyword placeholders in the initial sales text are replaced with vehicle function keywords to generate a car sales text containing vehicle function keywords.
[0108] This application embodiment uses a weight-driven keyword selection mechanism to ensure that high-profile vehicle features are accurately inserted into the script template.
[0109] This application provides a specific embodiment. Step 104 involves processing the car sales script text to obtain acoustic physical parameters characterizing the prosodic features of the speech. This specifically includes the following steps:
[0110] Step 401: Based on the vehicle function keywords, divide the car sales script text into multiple functional unit fragments.
[0111] In this step, the functional unit fragment refers to the text unit divided by vehicle function keywords.
[0112] In this embodiment of the application, the text of car sales script is processed by using vehicle function keywords as boundary identifiers to divide the text into independent functional unit segments. Each segment contains a core vehicle function description and related modifier statements.
[0113] Step 402: Mark the functional unit fragments containing the vehicle function keywords as strong rhythmic fragments, and mark the functional unit fragments that do not contain the vehicle function keywords as weak rhythmic fragments.
[0114] In this step, strong prosodic segments refer to text units containing keywords related to vehicle functions, and the perceptibility of technical parameters needs to be enhanced by increasing the fundamental frequency amplification value. Weak prosodic segments refer to descriptive text units that do not contain keywords, and non-core information needs to be weakened by reducing the speech rate coefficient.
[0115] In this embodiment of the application, each functional unit segment is scanned. If the functional unit segment contains at least one vehicle function keyword, it is marked as a strong prosodic segment; if the functional unit segment does not contain any keyword, it is marked as a weak prosodic segment.
[0116] Step 403: Using a preset car sales rhythm rule library, determine the high-intensity rhythm parameters corresponding to the strong rhythm segment and the low-intensity rhythm parameters corresponding to the weak rhythm segment.
[0117] In this step, the preset automotive sales rhythm rule base refers to a database that stores rhythm configurations specific to the automotive field, including a set of high-intensity parameters for technical parameters and a set of low-intensity parameters for comfort. High-intensity rhythm parameters refer to acoustic control values used for strong rhythmic segments. Low-intensity rhythm parameters refer to acoustic control values used for weak rhythmic segments.
[0118] In this embodiment, a preset automobile sales prosody rule library is invoked. Based on the strong prosody segment markers, the technical parameter class configuration in the preset automobile sales prosody rule library is queried, and the fundamental frequency amplification coefficient within a certain range is extracted as a high-intensity prosody parameter. Based on the weak prosody segment markers, the descriptive class configuration is queried, and the speech rate attenuation coefficient within a certain range is extracted as a low-intensity prosody parameter.
[0119] Step 404: Combine all high-intensity prosodic parameters and low-intensity prosodic parameters to obtain acoustic physical parameters that characterize the prosodic properties of speech.
[0120] In this embodiment, the high-intensity prosodic parameters of a strong prosodic segment and the low-intensity prosodic parameters of a weak prosodic segment are linearly concatenated according to the original temporal order of the functional unit segments. Each parameter is appended with a millisecond-level timestamp, ultimately generating a temporally continuous sequence of acoustic physical parameters.
[0121] This application embodiment achieves acoustic layering enhancement of technical parameters and descriptive content by differentiating the parameters of strong and weak prosodic segments.
[0122] This application provides a specific embodiment. Step 404 involves combining all high-intensity prosodic parameters and low-intensity prosodic parameters to obtain acoustic physical parameters characterizing the prosodic properties of speech. This specifically includes the following steps:
[0123] Step 411: Assign a position weight coefficient to each functional unit segment according to the temporal order of the functional unit segments in the car sales script text.
[0124] In this step, the position weight coefficient refers to the attenuation factor calculated based on the order in which the segments appear, reflecting the pattern of customer attention decay.
[0125] In this embodiment, position numbers are generated based on the order in which functional unit segments appear in the car sales script text, and position weight coefficients are calculated. These weight coefficients reflect the attenuation law of importance of segment positions in time sequence. For example, if the coefficient for the first segment is 1.5, then the position weight coefficient for the third segment is 1.3.
[0126] Step 412: Based on the preset car sales weight allocation rules, assign a first proportional coefficient to the strong rhythmic segment and a second proportional coefficient to the weak rhythmic segment.
[0127] In this step, the preset car sales weight allocation rule refers to a database that stores the mapping relationship between segment types and proportional coefficients. The first proportional coefficient refers to an amplification factor specifically for strong prosodic segments, used to enhance the pronunciation intensity of technical parameters. The second proportional coefficient refers to a compression factor specifically for weak prosodic segments, used to weaken non-core descriptions.
[0128] In this embodiment of the application, a preset automobile sales weight allocation rule is invoked. When a strong rhythmic segment is detected, the first proportional coefficient of the technical parameter class configuration in the preset automobile sales weight allocation rule is extracted; when a weak rhythmic segment is detected, the second proportional coefficient of the comfort description class configuration is extracted.
[0129] Step 413: Multiply each high-intensity prosodic parameter by its corresponding position weight coefficient and the first proportional coefficient to obtain multiple enhanced prosodic parameters.
[0130] In this step, the enhanced prosodic parameter refers to the acoustic modulation value after being amplified by position weights and scaling factors.
[0131] In this embodiment of the application, a triple operation is performed on each strong prosodic segment: the high-intensity prosodic parameter is multiplied by the position weight coefficient, and then multiplied by the first proportional coefficient to obtain the enhanced prosodic parameter, that is, enhanced prosodic parameter = original high-intensity prosodic parameter × position weight coefficient × second proportional coefficient.
[0132] Step 414: Multiply each low-intensity prosodic parameter by its corresponding position weight coefficient and the second proportional coefficient to obtain multiple weakened prosodic parameters.
[0133] In this step, the weakened prosodic parameter refers to the acoustic modulation value after compression by position weights and scaling factors.
[0134] In this embodiment of the application, a triple operation is performed on each weak prosodic segment: the low-intensity prosodic parameter is multiplied by the position weight coefficient, and then multiplied by the second proportional coefficient to obtain the weakened prosodic parameter, that is, the weakened prosodic parameter = the original low-intensity prosodic parameter × the position weight coefficient × the second proportional coefficient.
[0135] Step 415: According to the original temporal order of the functional unit segments, combine all enhancement parameters and all weakening parameters to obtain acoustic physical parameters characterizing the prosodic properties of speech.
[0136] In this step, the original temporal sequence refers to the original arrangement order of functional unit fragments in the car sales script text, which determines the timeline of the parameter sequence.
[0137] In this embodiment, the enhanced prosodic parameter values and the weakened prosodic parameter values are arranged sequentially according to the order in which the functional unit segments appear in the car sales script text, to obtain the acoustic physical parameters characterizing the prosodic characteristics of speech.
[0138] This application embodiment solves the problem of intensity attenuation of key selling points in the voice stream by coupling the calculation of temporal position weight and type proportional coefficient.
[0139] This application provides a specific embodiment. Step 105 involves generating synthesized speech associated with the vehicle function keywords using the acoustic physical parameters, specifically including the following steps:
[0140] Step 501: Convert the acoustic physical parameters into speech synthesis control parameters, which include pitch adjustment values and speech rate adjustment values.
[0141] In this step, the speech synthesis control parameters refer to the set of control instructions that drive speech synthesis. The pitch adjustment value refers to the fundamental frequency amplification coefficient in the acoustic physical parameters. The speech rate adjustment value refers to the duration proportion coefficient in the prosodic parameters.
[0142] In this embodiment, acoustic parameter conversion technology is used to process the time-series data of acoustic physical parameters. The pitch adjustment value is obtained by extracting the fundamental frequency trajectory value from the time-series data and dividing it by a preset reference fundamental frequency value. The speech rate adjustment value is obtained by extracting the syllable duration coefficient from the time-series data and multiplying it by a preset reference speech rate value. Finally, speech synthesis control parameters containing the pitch adjustment value and speech rate adjustment value at each time point are obtained.
[0143] Step 502: Generate corresponding initial speech waveforms based on the text content of all functional unit segments, and adjust all initial speech waveforms using the pitch adjustment value and speech rate adjustment value to generate intermediate speech waveforms corresponding to each functional unit segment, and use the physical duration of each intermediate speech waveform as the actual duration.
[0144] In this step, the initial speech waveform refers to the raw, unparalleled synthesized speech signal. The intermediate speech waveform refers to the speech signal after pitch and rate adjustments. The physical duration refers to the objective length of the intermediate speech waveform. The actual duration refers to the true playback time of the intermediate speech waveform in its unscaled state.
[0145] In this embodiment, the text content of all functional unit segments is processed to generate an initial speech waveform. The pitch adjustment value in the speech synthesis control parameters is applied to modify the fundamental frequency envelope of the initial speech waveform. The speech rate adjustment value is applied to scale the syllable duration of the initial speech waveform. An intermediate speech waveform is output, and the physical duration of each intermediate speech waveform is measured as the actual duration.
[0146] Step 503: Calculate the target duration of each functional unit segment based on its text length and prosody type.
[0147] In this step, text length refers to the number of valid characters contained in the functional unit segment; Chinese characters and English words are counted as single characters. Prosody type refers to the acoustic classification attribute of the functional unit segment; strong prosody type corresponds to technical parameter segment, and weak prosody type corresponds to descriptive segment.
[0148] In this embodiment, the base duration is obtained by multiplying the text length of the functional unit segment by a preset baseline syllable duration. When the prosodic type is strong, an enhancement coefficient that increases the duration is applied; when the prosodic type is weak, a weakening coefficient that decreases the duration is applied. Based on the enhancement and weakening coefficients, the base duration is updated to generate the target duration of each functional unit segment.
[0149] Step 504: Based on the actual duration and the target duration, perform time-domain scaling on all intermediate speech waveforms to generate the corresponding aligned speech waveforms.
[0150] In this step, the target duration refers to the ideal playback duration optimized for the sales scenario. The aligned speech waveform refers to a standardized speech signal segment whose actual duration precisely matches the target duration after time-domain scaling.
[0151] In this embodiment, the ratio difference between the actual duration and the target duration is calculated. A temporal scaling algorithm is used to process the intermediate speech waveform; high-fidelity scaling is used for strong prosodic segments, while efficient compression is used for weak prosodic segments, resulting in an aligned speech waveform.
[0152] Step 505: According to the original temporal sequence of the functional unit segments, all aligned speech waveforms are spliced together to generate synthesized speech.
[0153] In this embodiment, the aligned speech waveforms are joined end to end according to the original temporal order of the functional unit segments in the original speech text, and fade intervals are added between adjacent functional unit segments to eliminate pops, ultimately generating synthesized speech.
[0154] This application embodiment ensures millisecond-level synchronization between vehicle function explanation and 3D display through waveform scaling processing guided by target duration.
[0155] This application provides a specific embodiment. Step 504 involves performing time-domain scaling on all intermediate speech waveforms based on the actual duration and the target duration to generate the corresponding aligned speech waveform. This specifically includes the following steps:
[0156] Step 511: Based on the vehicle function keyword type of each functional unit segment, query the preset automotive terminology pronunciation benchmark table to determine the syllable density weight coefficient.
[0157] In this step, "vehicle function keyword type" refers to the functional classification of technical terms, including categories reflecting power performance such as motor power, categories reflecting intelligent driving assistance systems, and categories characterizing energy efficiency such as driving range and energy consumption. The "preset automotive terminology pronunciation benchmark table" refers to a database storing the pronunciation difficulty of automotive-related terms.
[0158] In this embodiment, the vehicle function keyword type of the functional unit segment is first identified. By accessing a preset automotive terminology pronunciation benchmark table, the corresponding syllable density weight coefficient is extracted according to the vehicle function keyword type. For example, the coefficient for the power performance keyword "dual motor four-wheel drive" is 1.8, and the coefficient for the comfort keyword "leather seats" is 1.0.
[0159] Step 512: Using the ratio of the actual duration to the target duration as the base scaling factor, multiply the syllable density weight coefficient and the base scaling factor to obtain the dynamic scaling parameters corresponding to each functional unit segment.
[0160] In this step, the syllable density weighting coefficient refers to the adjustment factor that quantifies the pronunciation complexity of terms; the coefficient for dynamic performance terms is generally higher than that for comfort terms. The dynamic scaling parameter refers to the composite parameter that ultimately controls the scaling of the waveform.
[0161] In this embodiment, the ratio of the actual duration to the target duration is used as the base scaling factor. Then, the syllable density weighting coefficient is multiplied by the base scaling factor to generate the dynamic scaling parameter. For example, dividing the actual duration of 5 seconds by the target duration of 6 seconds yields a base scaling factor of 0.83. Multiplying the base scaling factor of 0.83 by the syllable density weighting coefficient of 1.8 yields the dynamic scaling parameter of 1.49.
[0162] Step 513: Based on all dynamic scaling parameters, segment and scale the intermediate speech waveform corresponding to each functional unit segment to obtain the corresponding aligned speech waveform, so that the actual duration of each functional unit segment matches the corresponding target duration.
[0163] In this embodiment, the intermediate speech waveform is differentiated according to dynamic scaling parameters. When the parameter is greater than a preset value, the waveform is stretched; when it is less than the preset value, the waveform is compressed. The scaling boundaries are strictly aligned with the start and end times of the functional unit segments, and the output is a target duration-aligned speech waveform that matches the actual duration.
[0164] The embodiments of this application solve the problem of pronunciation integrity of complex technical terms through a dynamic scaling mechanism that adapts syllable density.
[0165] Figure 2 This application provides a schematic diagram of the structure of an AI voice interaction system, as shown in the embodiment. Figure 2 As shown, the system includes:
[0166] Module 21 is used to obtain customer type information and customer car purchase preference records from the historical sales interaction database of the automobile OEM;
[0167] Classification module 22 is used to classify the customer type information using preset customer classification rules and generate customer type identifiers, including first-time purchase type, replacement type and additional purchase type;
[0168] The parsing module 23 is used to parse the customer's car purchase preference record, generate a preference feature vector, and combine the customer type identifier and the preset car sales script template to generate car sales script text containing vehicle function keywords.
[0169] Processing module 24 is used to process the car sales script text to obtain acoustic physical parameters characterizing the prosodic characteristics of speech.
[0170] The association module 25 is used to generate synthesized speech associated with the vehicle function keywords using the acoustic physical parameters;
[0171] During the sales process, the display module 26 plays the synthesized voice and simultaneously triggers a 3D display of the car's functions corresponding to the car sales script text.
[0172] Figure 2 The AI voice interaction system described above can execute Figure 1 The implementation principle and technical effects of the AI voice interaction method described in the illustrated embodiment will not be repeated here. The specific methods by which each module and unit of the AI voice interaction system in the above embodiments perform operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0173] In one possible design, Figure 2 The AI voice interaction system shown in the embodiment can be implemented as a computing device, such as... Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;
[0174] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 32.
[0175] The processing component 32 is used for the above Figure 1 The embodiment describes an AI voice interaction method.
[0176] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.
[0177] Storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0178] Of course, computing devices may also include other components, such as input / output interfaces, display components, communication components, etc.
[0179] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.
[0180] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.
[0181] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.
[0182] This application also provides a computer storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 1 An AI voice interaction method is shown in the embodiment.
[0183] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0184] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0185] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. An AI voice interaction method, characterized in that, include: Obtain customer type information and customer car purchase preference records from the historical sales interaction database of automobile manufacturers; The customer type information is classified using preset customer classification rules to generate customer type identifiers, which include first-time purchase type, replacement type, and additional purchase type. The customer's car purchase preference record is parsed to generate a preference feature vector. Combined with the customer type identifier and the preset car sales script template, a car sales script text containing vehicle function keywords is generated. The car sales script text is processed to obtain acoustic physical parameters that characterize the prosodic features of the speech. Using the acoustic physical parameters, synthesized speech associated with the vehicle function keywords is generated; During the sales process, the synthesized voice is played, and a 3D display of the car's functions corresponding to the car sales script is triggered simultaneously. The customer's car purchase preference record is parsed to generate a preference feature vector. Combined with the customer type identifier and a preset car sales script template, a car sales script text containing vehicle feature keywords is generated, including: Based on the customer type identifier, select the target sales script template from the preset car sales script templates; The customer's car purchase preference record is parsed to generate a preference feature vector, which includes functional concerns and the weights of the functional concerns. The functional focus points are matched with the placeholders in the target script template to generate the initial script text; Based on the weight of the functional focus points, vehicle function keywords are selected from a preset car sales terminology library, and the keyword placeholders in the initial sales script are replaced with the vehicle function keywords to obtain a car sales script containing vehicle function keywords.
2. The method according to claim 1, characterized in that, The customer type information is classified using preset customer classification rules to generate customer type identifiers. These identifiers include first-time purchasers, replacement purchasers, and add-on purchasers. Extract purchase frequency and consultation text data from the customer type information; Using the main matching rule in the preset customer classification rules, the number of purchases is matched with a preset threshold to generate an initial customer type identifier, which includes first-time purchase type, replacement type and additional purchase type; Using auxiliary rules in the preset customer classification rules, key attention features in the consultation text data are detected; Based on the key features of interest, the initial customer type identifier is modified to generate a new customer type identifier.
3. The method according to claim 1, characterized in that, The text of the car sales pitch is processed to obtain acoustic physical parameters characterizing the prosodic features of the speech, including: Based on the vehicle function keywords, the car sales script text is divided into multiple functional unit fragments; Functional unit segments containing the vehicle function keywords are marked as strong rhythmic segments, and functional unit segments that do not contain the vehicle function keywords are marked as weak rhythmic segments. Using a pre-defined car sales rhythm rule library, determine the high-intensity rhythm parameters corresponding to the strong rhythm segment and the low-intensity rhythm parameters corresponding to the weak rhythm segment; By combining all high-intensity prosodic parameters and low-intensity prosodic parameters, acoustic physical parameters characterizing the prosodic properties of speech are obtained.
4. The method according to claim 3, characterized in that, By combining all high-intensity and low-intensity prosodic parameters, acoustic physical parameters characterizing the prosodic properties of speech are obtained, including: Based on the temporal order of the functional unit segments in the car sales script text, assign a positional weight coefficient to each functional unit segment; Based on the preset car sales weight allocation rules, a first proportional coefficient is assigned to the strong rhythmic segment, and a second proportional coefficient is assigned to the weak rhythmic segment. Each high-intensity prosodic parameter is multiplied by its corresponding position weight coefficient and the first proportional coefficient to obtain multiple enhanced prosodic parameters; Each low-intensity prosodic parameter is multiplied by its corresponding position weight coefficient and the second proportional coefficient to obtain multiple weakened prosodic parameters; According to the original temporal order of the functional unit segments, all enhancement parameters and all weakening parameters are combined to obtain acoustic physical parameters that characterize the prosodic properties of speech.
5. The method according to claim 1, characterized in that, Using the aforementioned acoustic physical parameters, synthesized speech associated with the vehicle function keywords is generated, including: The acoustic physical parameters are converted into speech synthesis control parameters, which include pitch adjustment values and speech rate adjustment values. Based on the text content of all functional unit segments, generate corresponding initial speech waveforms, and adjust all initial speech waveforms using the pitch adjustment value and speech rate adjustment value to generate intermediate speech waveforms corresponding to each functional unit segment. The physical duration of each intermediate speech waveform is used as the actual duration. Calculate the target duration of each functional unit segment based on its text length and prosody type; Based on the actual duration and the target duration, all intermediate speech waveforms are subjected to time-domain scaling to generate corresponding aligned speech waveforms. According to the original temporal sequence of the functional unit segments, all aligned speech waveforms are spliced together to generate synthesized speech.
6. The method according to claim 5, characterized in that, Based on the actual duration and the target duration, all intermediate speech waveforms are subjected to temporal scaling to generate corresponding aligned speech waveforms, including: Based on the vehicle function keyword type of each functional unit segment, a preset automotive terminology pronunciation benchmark table is queried to determine the syllable density weight coefficient. The ratio of the actual duration to the target duration is used as the basic scaling factor. The syllable density weight coefficient is multiplied by the basic scaling factor to obtain the dynamic scaling parameters corresponding to each functional unit segment. Based on all dynamic scaling parameters, the intermediate speech waveform corresponding to each functional unit segment is segmented and scaled to obtain the corresponding aligned speech waveform, so that the actual duration of each functional unit segment matches the corresponding target duration.
7. An AI voice interaction system, characterized in that, include: The acquisition module is used to retrieve customer type information and customer car purchase preference records from the historical sales interaction database of the car manufacturer; The classification module is used to classify the customer type information using preset customer classification rules and generate customer type identifiers, which include first-time purchase type, replacement type and additional purchase type. The parsing module is used to parse the customer's car purchase preference record, generate a preference feature vector, and combine it with the customer type identifier and the preset car sales script template to generate car sales script text containing vehicle function keywords. The processing module is used to process the car sales script text to obtain acoustic physical parameters that characterize the prosodic features of speech. The association module is used to generate synthesized speech associated with the vehicle function keywords using the acoustic physical parameters; The display module plays the synthesized voice during the sales process and simultaneously triggers a 3D display of the car's functions corresponding to the car sales script text; The customer's car purchase preference record is parsed to generate a preference feature vector. Combined with the customer type identifier and a preset car sales script template, a car sales script text containing vehicle feature keywords is generated, including: Based on the customer type identifier, select the target sales script template from the preset car sales script templates; The customer's car purchase preference record is parsed to generate a preference feature vector, which includes functional concerns and the weights of the functional concerns. The functional focus points are matched with the placeholders in the target script template to generate the initial script text; Based on the weight of the functional focus points, vehicle function keywords are selected from a preset car sales terminology library, and the keyword placeholders in the initial sales script are replaced with the vehicle function keywords to obtain a car sales script containing vehicle function keywords.
8. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement an AI voice interaction method as described in any one of claims 1 to 6.
9. A computer storage medium, characterized in that, The device contains a computer program that, when executed by a computer, implements an AI voice interaction method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Speech skill generation method, artificial intelligence-based interpretation method and devices
CN110990550A
Automobile sales system, module and method based on big data
CN118350839A