AI voice rate adjusting method and system of AI platform
By analyzing the historical data of user conversations with AI and dynamically adjusting the AI voice rate, the problem of poor voice rate adjustment in the existing technology is solved, and the user's personalized experience is improved.
Patent Information
- Application Number
- CN202510582109.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing AI voice interaction system has shortcomings in voice rate adjustment and cannot achieve a truly personalized experience. Users need to manually adjust the voice rate repeatedly, which is cumbersome and has poor experience.
By obtaining historical data of user conversations with AI, extracting user behavior characteristics, forming standard feature groups, and calculating feature deviation amounts in real-time conversations, dynamically adjusting AI voice rate to match users' interaction habits.
It realizes automatic matching of AI voice rate and user interaction habits, eliminates the tedious process of user manual adjustment, and significantly improves the naturalness and user satisfaction of AI voice interaction.
Smart Images

Figure CN120220684A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing. More specifically, the present invention relates to an AI voice rate adjustment method and system for an AI platform. Background Art
[0002] With the rapid development of artificial intelligence technology, AI voice interaction has become an important way of human-computer interaction and is widely used in many fields such as intelligent assistants, customer service systems, and education platforms. However, there are still obvious deficiencies in the voice rate adjustment of existing AI voice interaction systems, and it is difficult to achieve a truly personalized experience.
[0003] Traditional AI voice systems usually adopt a fixed voice rate setting and cannot adapt to the speech rate preferences of different users. Although some systems allow users to manually adjust the voice rate, this static adjustment method requires users to continuously try different settings, which is a cumbersome process and the experience is not good. Users often need to enter the settings interface multiple times and repeatedly adjust the voice rate until they find a relatively comfortable speed, which greatly reduces the fluency of interaction and user satisfaction. Summary of the Invention
[0004] In order to overcome the problem of low fluency of interaction in the prior art, the present invention proposes an AI voice rate adjustment method and system for an AI platform to solve the above problems.
[0005] The present invention provides the following technical solutions:
[0006] An AI voice rate adjustment method for an AI platform, comprising:
[0007] Obtaining historical data of the conversation between the user and the AI in the AI platform, where the historical data includes user behavior data and user feedback data, and the user feedback data is divided into four levels: excellent, good, medium, and poor;
[0008] Performing feature extraction on the user behavior data to obtain corresponding behavior features; screening out the data subsets with excellent and good user feedback from the historical data, and extracting the behavior features corresponding to the data subsets to form a standard feature group;
[0009] Calculating a standard behavior feature based on the standard feature group;
[0010] During the real-time conversation between the AI and the user, collecting the current user behavior data, and performing feature extraction on the current user behavior data to obtain the corresponding behavior feature denoted as the real-time behavior feature;
[0011] Calculating the feature deviation amount between the real-time behavior feature and the standard behavior feature;
[0012] Obtain the real-time speech rate of the current AI speech, and dynamically adjust the AI speech rate according to the real-time speech rate and the feature deviation amount.
[0013] Preferably, the user behavior data includes the user's speaking rate, the user's pause frequency and duration, the frequency of interrupting the AI, and the frequency of repeated questions. Among them, the user's speaking rate is obtained by calculating the number of characters input by the user per unit time or the number of words after speech recognition; the user's pause frequency and duration are obtained by detecting the silent segments during the user's speech input and counting their occurrence times and durations; the frequency of interrupting the AI is obtained by detecting the number of times of voice input or operation behavior initiated by the user when the AI speech output is not completed; the frequency of repeated questions is obtained by semantic analysis to identify the number of times of expressing the same or similar intentions in the user's continuous input content.
[0014] Preferably, the step of extracting features from the user behavior data to obtain the corresponding behavior features includes:
[0015] Segment the user's speaking rate data by time window, calculate the average rate, the standard deviation of rate change, and the rate change trend within each window to obtain the user's speaking speed feature;
[0016] Conduct statistical analysis on the user's pause data, calculate the average pause duration, pause frequency, and pause distribution pattern to obtain the user's pause feature;
[0017] Conduct time series analysis on the data of the frequency of interrupting the AI, calculate the number of interruptions per unit time and the interruption timing distribution to obtain the user's interruption feature;
[0018] Conduct pattern recognition on the data of the frequency of repeated questions, calculate the time interval of repeated questions to obtain the user's repetition feature;
[0019] Perform feature vectorization processing on the user's speaking speed feature, user's pause feature, user's interruption feature, and user's repetition feature to form the user behavior feature.
[0020] Preferably, the steps of calculating the standard behavior features based on the standard feature group include:
[0021] Normalize each behavior feature in the standard feature group so that each feature value is distributed in the same numerical interval;
[0022] Calculate the mean value of each normalized behavior feature to obtain the preliminary standard behavior feature;
[0023] Perform weighted processing on the preliminary standard behavior features, where the weight coefficient is determined according to the correlation between the corresponding behavior feature and the user satisfaction;
[0024] Take the weighted feature vector as the standard behavior feature.
[0025] Preferably, the steps for obtaining the weight coefficient include:
[0026] Extract data samples containing all user feedback levels from historical data;
[0027] Convert the user feedback levels from excellent, good, medium, and poor from large to small into different numerical scores;
[0028] Calculate the Pearson correlation coefficient between each behavioral feature and the user feedback score to obtain the correlation strength of each feature;
[0029] Standardize the Pearson correlation coefficient to convert it into a value between 0 and 1; calculate the variance ratio of each feature in samples of different feedback levels;
[0030] Perform a weighted average of the standardized correlation coefficient and the variance ratio to obtain the comprehensive importance score of each feature;
[0031] Normalize the comprehensive importance score so that the sum of all weight coefficients is 1 to obtain the final weight coefficient.
[0032] Preferably, the steps for obtaining the feature deviation amount include:
[0033] Normalize the real-time behavioral feature so that it is in the same numerical range as the standard behavioral feature;
[0034] Calculate the Euclidean distance between the normalized real-time behavioral feature and the standard behavioral feature to obtain the feature distance value;
[0035] Calculate the feature direction vector based on the feature distance value to indicate the deviation direction of the real-time behavioral feature relative to the standard behavioral feature;
[0036] According to the feature direction vector and the feature distance value, map the feature deviation amount to the range between -1 and 1, where a positive value indicates that the user behavioral feature is faster than the standard behavioral feature, a negative value indicates that the user behavioral feature is slower than the standard behavioral feature, and the absolute value indicates the degree of deviation.
[0037] Preferably, the steps for dynamically adjusting the AI speech rate according to the real-time speech rate and the feature deviation amount include:
[0038] Obtain the preset maximum speech rate and minimum speech rate;
[0039] Use the formula Calculate the relative position coefficient β of the current speech rate within the range of the maximum and minimum speech rates, where Vd is the real-time speech rate, Vmax is the maximum speech rate, and Vmin is the minimum speech rate;
[0040] Obtain the feature deviation amount θ;
[0041] When θ is positive, the rate adjustment coefficient λ is obtained through the formula λ = (1 - β) × θ;
[0042] When θ is negative, the rate adjustment coefficient λ is obtained through the formula λ = β × θ;
[0043] The adjusted speech rate Vt is calculated through the formula Vt = Vd × (1 + λ);
[0044] The speech rate is smoothly adjusted from Vd to Vt.
[0045] The present invention also provides an AI speech rate adjustment system for an AI platform, which is used to implement an AI speech rate adjustment method for an AI platform, including:
[0046] A historical data acquisition module, which is used to acquire the historical data of the conversation between the user and the AI in the AI platform. The historical data includes user behavior data and user feedback data, and the user feedback data is divided into four levels: excellent, good, medium, and poor;
[0047] A feature processing module, which is used to extract features from the user behavior data to obtain corresponding behavior features; screen out the data subset with excellent and good user feedback from the historical data, and extract the behavior features corresponding to the data subset to form a standard feature group;
[0048] A standard feature calculation module, which is used to calculate the standard behavior features based on the standard feature group;
[0049] A real-time feature acquisition module, which is used to collect the current user behavior data during the real-time conversation between the AI and the user, and extract features from the current user behavior data to obtain the corresponding behavior features, denoted as real-time behavior features;
[0050] A feature deviation analysis module, which is used to calculate the feature deviation amount between the real-time behavior features and the standard behavior features;
[0051] A speech rate adjustment module, which is used to obtain the real-time speech rate of the current AI speech, and dynamically adjust the AI speech rate according to the real-time speech rate and the feature deviation amount.
[0052] The present invention provides an AI speech rate adjustment method and system for an AI platform, which have the following beneficial effects:
[0053] First, by collecting multi-dimensional behavioral data such as the user's speech rate and pause frequency, a comprehensive user interaction behavior model is established, avoiding the limitations of traditional methods that rely only on a single parameter and being able to capture the user's interaction habits and preferences more accurately. Second, by screening the data with excellent and good user feedback to form a standard feature group, this solution establishes a standard behavior feature model oriented to user satisfaction, ensuring that the direction of speech rate adjustment always meets the user's expectations and solving the problem of the disconnection between traditional systems and user needs. Third, this solution can calculate the deviation between the user's current behavior features and the standard features in real time, accurately judge whether the user's behavior is faster or slower than the standard features, and automatically adjust the AI speech rate accordingly, obtaining a voice experience that matches the user's interaction habits without the need for manual settings by the user. Finally, this solution adopts a smooth adjustment mechanism and sets a reasonable rate range to ensure the continuity and naturalness of the speech rate change, avoiding the negative impact of abrupt changes on the user experience.
[0054] In summary, this technical solution realizes the automatic matching of the AI speech rate with the user's interaction habits, eliminates the cumbersome process of the user needing to manually adjust repeatedly, and significantly improves the naturalness of AI speech interaction and user satisfaction. Brief Description of the Drawings
[0055] Figure 1 It is a schematic flowchart of the AI speech rate adjustment method of an AI platform of the present invention;
[0056] Figure 2 It is a schematic block diagram of the AI speech rate adjustment system of an AI platform of the present invention. Detailed Embodiments
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] Embodiment 1
[0059] Please refer to Figure 1 , in this embodiment, an AI speech rate adjustment method for an AI platform includes:
[0060] S1. Obtain the historical data of the user's conversation with the AI in the AI platform, where the historical data includes user behavior data and user feedback data, and the user feedback data is divided into four levels: excellent, good, medium, and poor;
[0061] The user behavior data includes the user's speaking rate, the frequency and duration of user pauses, the frequency of interrupting the AI, and the frequency of repeated questions. Among them, the user's speaking rate is obtained by calculating the number of characters input by the user per unit time or the number of words after speech recognition; the frequency and duration of user pauses are obtained by detecting the silent segments during the user's voice input process and counting their occurrence times and durations; the frequency of interrupting the AI is obtained by detecting the number of voice inputs or operation behaviors initiated by the user when the AI's voice output is not completed; the frequency of repeated questions is obtained by semantic analysis to identify the number of times of expressing the same or similar intentions in the user's consecutive input content.
[0062] In this embodiment, obtaining the historical data of the conversation between the user and the AI in the AI platform can be carried out according to the following steps: First, integrate the data acquisition system with the AI platform, set up an automatically triggered user feedback scoring interface, and divide the scores into four levels: excellent (90 - 100 points), good (75 - 89 points), medium (60 - 74 points), and poor (0 - 59 points).
[0063] Then, the method for collecting user behavior data is as follows: The user's speaking rate is calculated by recording the text input time and the number of characters (for example, 200 characters are input in 10 seconds, and the rate is 20 characters / second) or the voice input duration and the number of recognized words (for example, 5 seconds of voice is recognized as 50 words, and the rate is 10 words / second); the frequency and duration of user pauses are detected by the audio analysis module for silent segments below -40 dB, and a pause is recorded if it lasts for more than 300 milliseconds, and the number of times and the duration are counted; the frequency of interrupting the AI is detected by the interaction monitoring module for new input behaviors of the user when the AI's voice output is not completed, and the number of interruptions per unit time is recorded; the frequency of repeated questions is calculated by the natural language processing module for the semantic similarity of consecutive input content, and if the similarity exceeds 80%, it is determined as a repeated question.
[0064] S2. Extract features from the user behavior data to obtain corresponding behavior features; screen out the data subsets with excellent and good user feedback from the historical data, and extract the behavior features corresponding to this data subset to form a standard feature group;
[0065] The extraction of features from the user behavior data to obtain corresponding behavior features includes:
[0066] Segment the user's speaking rate data by time window, calculate the average rate, the standard deviation of rate change, and the rate change trend within each window to obtain the user's speech rate feature;
[0067] Conduct statistical analysis on the user's pause data, calculate the average pause duration, the pause frequency, and the pause distribution pattern to obtain the user's pause feature;
[0068] Conduct time series analysis on the data of the frequency of interrupting the AI, calculate the number of interruptions per unit time and the interruption timing distribution to obtain the user's interruption feature;
[0069] Perform pattern recognition on the repeated question frequency data, calculate the time interval of repeated questions, and obtain the user's repetition characteristics;
[0070] Vectorize the user speech rate characteristics, user pause characteristics, user interruption characteristics, and user repetition characteristics to form user behavior characteristics.
[0071] In this embodiment, the process of feature extraction from user behavior data is as follows: First, extract all user behavior data from the database, and filter out the data subsets with excellent (90 - 100 points) and good (75 - 89 points) user feedback scores as standard samples.
[0072] Then, extract features from various types of behavior data: For user speech rate data, segment it using a sliding time window with a unit of 30 seconds and a window overlap rate of 50%, and calculate the average rate (e.g., 15 words / second), standard deviation of rate change (e.g., ±3 words / second), and rate change trend (e.g., linear regression slope 0.2, indicating a gradual increase) within each window; For user pause data, calculate the average pause duration (e.g., 0.8 seconds), pause frequency (e.g., 8 times / minute), and pause distribution pattern (e.g., end-of-sentence pauses account for 70%, in-sentence pauses account for 30%); For the frequency of interrupting the AI data, perform time series analysis on a per-conversation basis, and calculate the number of interruptions per minute (e.g., 0.5 times / minute) and interruption timing distribution (e.g., interruptions at 20% of the AI's answer progress account for 40%, at 50% account for 35%, and at 80% account for 25%); For the repeated question frequency data, identify the repeated question pattern and calculate the time interval (e.g., average interval of 45 seconds).
[0073] Finally, vectorize all the extracted features to form a feature vector with a fixed dimension, which serves as the basis for the standard feature group.
[0074] S3. Calculate the standard behavior characteristics based on the standard feature group;
[0075] The steps of calculating the standard behavior characteristics based on the standard feature group include:
[0076] Normalize each behavior feature in the standard feature group so that each feature value is distributed in the same numerical range;
[0077] Calculate the mean of each normalized behavior feature to obtain the preliminary standard behavior characteristics;
[0078] Perform weighted processing on the preliminary standard behavior characteristics, where the weight coefficient is determined according to the correlation between the corresponding behavior feature and user satisfaction;
[0079] Take the weighted feature vector as the standard behavior characteristics.
[0080] The steps for obtaining the weight coefficients include:
[0081] Extract data samples containing all user feedback levels from historical data;
[0082] Convert the user feedback levels from excellent, good, medium, and poor from large to small into different numerical scores;
[0083] Calculate the Pearson correlation coefficient between each behavioral feature and the user feedback score to obtain the correlation strength of each feature;
[0084] Standardize the Pearson correlation coefficient to convert it into a value between 0 and 1; calculate the variance ratio of each feature in samples with different feedback levels;
[0085] Perform a weighted average of the standardized correlation coefficient and the variance ratio to obtain the comprehensive importance score of each feature;
[0086] Normalize the comprehensive importance score so that the sum of all weight coefficients is 1 to obtain the final weight coefficients.
[0087] In this embodiment, the process of calculating the standard behavioral features based on the standard feature group is as follows: First, perform Min-Max normalization on each behavioral feature in the standard feature group to map all feature values to the interval [0, 1]. For example, the average rate in the user speech rate feature is normalized from the original value of 15 words per second to 0.6, and the standard deviation of rate change from ±3 words per second is normalized to 0.4. Then, calculate the mean of each behavioral feature after normalization. For example, the mean of the speech rate feature is [0.6, 0.4, 0.55], and the mean of the pause feature is [0.45, 0.7, 0.65], etc., to form a preliminary standard behavioral feature vector.
[0088] Next, calculate the weight coefficients: Extract 10,000 samples containing all feedback levels from the database, and convert the user feedback levels into numerical scores (excellent = 4 points, good = 3 points, medium = 2 points, poor = 1 point); calculate the Pearson correlation coefficient between each feature and the feedback score. For example, the correlation coefficient of the average speech rate is 0.75, and the correlation coefficient of the pause frequency is 0.68; standardize the correlation coefficient to the interval [0, 1]. For example, the average speech rate is standardized to 0.9, and the pause frequency is standardized to 0.82; calculate the variance ratio of each feature in samples with different feedback levels. For example, the variance ratio of the average speech rate is 0.85, and the variance ratio of the pause frequency is 0.78; perform a weighted average of the standardized correlation coefficient and the variance ratio at a ratio of 6:4 to obtain the comprehensive importance score. For example, the score of the average speech rate is 0.88, and the score of the pause frequency is 0.8;
[0089] Finally, normalize all scores so that the sum is 1 to obtain the final weight coefficients (for example, the weight of the average speech rate is 0.15, and the weight of the pause frequency is 0.12). Multiply the preliminary standard behavior features by the corresponding weight coefficients to obtain the weighted standard behavior feature vector as the final result.
[0090] S4. During the real-time conversation between the AI and the user, collect the current user behavior data, and extract features from the current user behavior data to obtain the corresponding behavior features, denoted as real-time behavior features;
[0091] In this embodiment, the implementation of collecting behavior data during the real-time conversation between the AI and the user is as follows: The system continuously captures the user behavior data in the current conversation through the real-time monitoring module, and all these real-time data are processed through the same feature extraction process as in step S2 to obtain the real-time behavior feature vector of the current conversation.
[0092] S5. Calculate the feature deviation amount between the real-time behavior features and the standard behavior features;
[0093] The steps for obtaining the feature deviation amount include:
[0094] Normalize the real-time behavior features so that they are in the same numerical range as the standard behavior features;
[0095] Calculate the Euclidean distance between the normalized real-time behavior features and the standard behavior features to obtain the feature distance value;
[0096] Calculate the feature direction vector based on the feature distance value to indicate the deviation direction of the real-time behavior features relative to the standard behavior features;
[0097] According to the feature direction vector and the feature distance value, map the feature deviation amount to the range between -1 and 1, where a positive value indicates that the user behavior features are faster than the standard behavior features, a negative value indicates that the user behavior features are slower than the standard behavior features, and the absolute value indicates the degree of deviation.
[0098] In this embodiment, the process of calculating the feature deviation amount between the real-time behavior feature and the standard behavior feature is as follows: First, perform Min-Max normalization on the real-time behavior feature vector to make it in the same [0, 1] numerical interval as the standard behavior feature. Then, calculate the Euclidean distance between the normalized real-time feature vector and the standard behavior feature vector, and obtain a feature distance value of 0.28. Next, calculate the feature direction vector, that is, subtract the standard feature vector from the real-time feature vector to obtain a vector representing the deviation direction. Finally, based on the component of the direction vector and the feature distance value, calculate the feature deviation amount: First, calculate the weighted average value of the component related to the speech rate in the direction vector as 0.09 (a positive value indicates that the user's speech rate is faster than the standard), and the weighted average value of the component related to the pause as -0.02 (a negative value indicates that the user's pause is slightly slower than the standard). Combine all components and consider their weights to finally obtain a feature deviation amount of 0.35.
[0099] S6. Obtain the real-time speech rate of the current AI voice, and dynamically adjust the AI voice rate according to the real-time speech rate and the feature deviation amount.
[0100] The step of dynamically adjusting the AI voice rate according to the real-time speech rate and the feature deviation amount includes:
[0101] Obtain the preset maximum speech rate and minimum speech rate;
[0102] Use the formula Calculate the relative position coefficient β of the current speech rate within the range of the maximum and minimum speech rates, where Vd is the real-time speech rate, Vmax is the maximum speech rate, and Vmin is the minimum speech rate;
[0103] Obtain the feature deviation amount θ;
[0104] When θ is a positive value, obtain the rate adjustment coefficient λ through the formula λ = (1 - β) × θ;
[0105] When θ is a negative value, obtain the rate adjustment coefficient λ through the formula λ = β × θ;
[0106] Calculate the adjusted speech rate Vt through the formula Vt = Vd × (1 + λ);
[0107] Smoothly adjust the speech rate from Vd to Vt.
[0108] In this embodiment, the process of dynamically adjusting the AI voice rate is as follows: First, the system obtains the preset maximum voice rate (such as 300 words per minute) and minimum voice rate (such as 120 words per minute), and reads the real-time voice rate of the current AI voice (such as 220 words per minute). Then, it calculates the relative position coefficient of the current voice rate within the range of the maximum and minimum voice rates. In this example, it is 0.56, indicating that the current rate is slightly higher than the middle value. Next, the system obtains the feature deviation amount of 0.35 (positive value) calculated in the previous step. Since the feature deviation amount is positive, indicating that the overall user behavior characteristics are faster than the standard characteristics, the system calculates a rate adjustment coefficient of 0.154, which means that the current voice rate needs to be increased by about 15.4%. Subsequently, the system calculates the adjusted voice rate as 220×(1 + 0.154) = 253.88 words per minute. To ensure the naturalness of the change in the voice rate, the system uses the linear interpolation method to smoothly transition the voice rate from 220 words per minute to 254 words per minute within the next 5 seconds, increasing the rate by about 6.8 words per minute per second. In this way, the AI voice rate can be dynamically adjusted according to the user behavior characteristics, providing a more user-friendly dialogue experience.
[0109] Embodiment 2
[0110] Please refer to Figure 2 , the present invention provides an AI voice rate adjustment system for an AI platform, which is used to implement an AI voice rate adjustment method for an AI platform, including:
[0111] A historical data acquisition module, configured to acquire historical data of the conversation between the user and the AI in the AI platform, where the historical data includes user behavior data and user feedback data, and the user feedback data is divided into four levels: excellent, good, medium, and poor;
[0112] A feature processing module, configured to extract features from the user behavior data to obtain corresponding behavior features; screen out a data subset of the user feedback as excellent and good from the historical data, and extract the behavior features corresponding to the data subset to form a standard feature group;
[0113] A standard feature calculation module, configured to calculate standard behavior features based on the standard feature group;
[0114] A real-time feature acquisition module, configured to collect the current user behavior data during the real-time conversation between the AI and the user, and extract features from the current user behavior data to obtain corresponding behavior features denoted as real-time behavior features;
[0115] A feature deviation analysis module, configured to calculate the feature deviation amount between the real-time behavior features and the standard behavior features;
[0116] The voice rate adjustment module is used to obtain the real-time voice rate of the current AI voice and dynamically adjust the AI voice rate according to the real-time voice rate and the feature deviation amount.
[0117] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only one type, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0118] As mentioned above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.
[0119] Finally: The above is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for adjusting the AI speech rate of an AI platform, characterized in that: include: Obtain historical data of user-AI conversations in the AI platform, wherein the historical data includes user behavior data and user feedback data, wherein the user feedback data is divided into four levels: excellent, good, medium, and poor; Extracting features from the user behavior data to obtain corresponding behavior features; selecting a data subset with excellent and good user feedback from the historical data, and extracting the behavior features corresponding to the data subset to form a standard feature group; Based on the standard feature group, a standard behavior feature is calculated; During the real-time conversation between AI and users, the current user behavior data is collected and features are extracted from the current user behavior data to obtain corresponding behavior features, which are recorded as real-time behavior features. Calculating a feature deviation between the real-time behavior feature and the standard behavior feature; The real-time speech rate of the current AI speech is obtained, and the AI speech rate is dynamically adjusted according to the real-time speech rate and the characteristic deviation.
2. The AI speech rate adjustment method of an AI platform according to claim 1, characterized in that: The user behavior data includes user speaking rate, user pause frequency and duration, AI interruption frequency and repeated questioning frequency, wherein the user speaking rate is obtained by calculating the number of characters input by the user per unit time or the number of words after voice recognition; the user pause frequency and duration are obtained by detecting the silent segments during the user's voice input process and counting their occurrence times and duration; the AI interruption frequency is obtained by detecting the number of voice inputs or operation behaviors initiated by the user when the AI voice output is not completed; the repeated questioning frequency is obtained by identifying the number of times the user expresses the same or similar intentions in the continuous input content through semantic analysis.
3. The AI speech rate adjustment method of an AI platform according to claim 2, characterized in that: The feature extraction of user behavior data to obtain corresponding behavior features includes: The user's speaking rate data is segmented into time windows, and the average rate, rate change standard deviation and rate change trend in each window are calculated to obtain the user's speaking rate characteristics; Perform statistical analysis on user pause data, calculate the average pause duration, pause frequency and pause distribution pattern, and obtain user pause characteristics; Perform time series analysis on the interruption frequency data of AI, calculate the number of interruptions per unit time and the distribution of interruption timing, and obtain the user interruption characteristics; Perform pattern recognition on the frequency data of repeated questions, calculate the time interval between repeated questions, and obtain the user's repetition characteristics; The user speech speed feature, user pause feature, user interruption feature and user repetition feature are subjected to feature vectorization processing to form user behavior features.
4. The AI speech rate adjustment method of an AI platform according to claim 3, characterized in that: The step of calculating the standard behavior feature based on the standard feature group includes: Normalize each behavior feature in the standard feature group so that each feature value is distributed in the same numerical range; Calculate the mean of each normalized behavior feature to obtain the preliminary standard behavior feature; Performing weighted processing on the preliminary standard behavior characteristics, wherein the weight coefficient is determined according to the correlation between the corresponding behavior characteristics and user satisfaction; The weighted feature vector is used as the standard behavior feature.
5. The AI speech rate adjustment method of an AI platform according to claim 4, characterized in that: The step of obtaining the weight coefficient comprises: Extract data samples containing all user feedback levels from historical data; Convert user feedback levels into different numerical scores from excellent, good, medium and poor; Calculate the Pearson correlation coefficient between each behavioral feature and the user feedback score to obtain the correlation strength of each feature; The Pearson correlation coefficient was standardized and converted to a value between 0 and 1; the variance ratio of each feature in samples with different feedback levels was calculated; The standardized correlation coefficient and variance ratio are weighted averaged to obtain the comprehensive importance score of each feature; The comprehensive importance score is normalized so that the sum of all weight coefficients is 1 to obtain the final weight coefficient.
6. The AI speech rate adjustment method of an AI platform according to claim 5, characterized in that: The step of obtaining the characteristic deviation comprises: Normalize the real-time behavior features so that they are in the same numerical range as the standard behavior features; Calculate the Euclidean distance between the normalized real-time behavior feature and the standard behavior feature to obtain a feature distance value; Calculate the feature direction vector based on the feature distance value, indicating the deviation direction of the real-time behavior feature relative to the standard behavior feature; According to the feature direction vector and the feature distance value, the feature deviation is mapped to a range between -1 and 1, where a positive value indicates that the user behavior feature is faster than the standard behavior feature, a negative value indicates that the user behavior feature is slower than the standard behavior feature, and the absolute value indicates the degree of deviation.
7. The AI speech rate adjustment method of an AI platform according to claim 6, characterized in that: The step of dynamically adjusting the AI speech rate according to the real-time speech rate and the characteristic deviation comprises: Get the preset maximum and minimum speech rates; Using the formula Calculate the relative position coefficient β of the current speech rate within the range of maximum and minimum speech rates, where Vd is the real-time speech rate, Vmax is the maximum speech rate, and Vmin is the minimum speech rate; Obtain feature deviation θ; When θ is a positive value, the rate adjustment coefficient λ is obtained by the formula λ=(1-β)×θ; When θ is a negative value, the rate adjustment coefficient λ is obtained by the formula λ=β×θ; The adjusted speech rate Vt is calculated by the formula Vt=Vd×(1+λ); Smoothly adjust the speech rate from Vd to Vt.
8. An AI voice rate adjustment system for an AI platform, used to implement an AI voice rate adjustment method for an AI platform as described in any one of claims 1 to 7, characterized in that: include: A historical data acquisition module is used to acquire historical data of user-AI dialogues in the AI platform, wherein the historical data includes user behavior data and user feedback data, wherein the user feedback data is divided into four levels: excellent, good, medium, and poor; The feature processing module is used to extract features from the user behavior data to obtain corresponding behavior features; filter out a data subset with excellent and good user feedback from the historical data, and extract the behavior features corresponding to the data subset to form a standard feature group; A standard feature calculation module, used for calculating and obtaining standard behavior features based on the standard feature group; The real-time feature collection module is used to collect the current user behavior data during the real-time dialogue between AI and users, and extract features from the current user behavior data to obtain corresponding behavior features, which are recorded as real-time behavior features. A feature deviation analysis module, used to calculate the feature deviation between the real-time behavior feature and the standard behavior feature; The speech rate adjustment module is used to obtain the real-time speech rate of the current AI speech and dynamically adjust the AI speech rate according to the real-time speech rate and the characteristic deviation.
Citation Information
Cited By
AI tourism personalized explanation system and service system
CN121053717A