AI voice rate adjusting method of AI platform
Through the autoregressive pre-training language model, the speech signal keywords are extracted and the speech speed is analyzed, and the speech speed of the AI speech system is automatically adjusted, which solves the problems of user speech rate differences and environmental adaptation, and improves the interactive experience and user satisfaction.
Patent Information
- Application Number
- CN202510383561.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
The existing AI voice system cannot effectively adapt to the voice rate needs of different users, resulting in poor interactive experience. Especially under different environments and user habits, manual adjustments are cumbersome and not real-time, which may affect security.
The autoregressive pre-trained language model is used to extract speech signal keywords, combine speech speed analysis of speech signals, and automatically adjust the speech speed of AI speech to match the user's speech speed habits and environmental needs.
Real-time speech speed adaptation of AI voice systems in different users and environments has been realized, which improves interaction fluency and user satisfaction, and enhances users' dependence on voice services.
Smart Images

Figure CN120260581A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of AI speech rate, and particularly relates to an AI speech rate adjustment method for an AI platform. Background Art
[0002] With the rapid development of technology, artificial intelligence (AI) has been widely integrated into various fields of people's lives, and the application of AI speech technology is particularly prominent. In many scenarios such as smart home, intelligent vehicle, intelligent customer service, and voice assistants, the interaction between users and AI through voice has become increasingly frequent.
[0003] In a smart home environment, a user may be busy with housework while talking to a smart speaker to query the weather, play music, or control home appliances; in an intelligent vehicle scenario, a driver needs to rely on voice commands to complete operations such as navigation settings and making phone calls during driving; intelligent customer service needs to respond to a large number of customer inquiries in real time and provide accurate and effective services; voice assistants have become a powerful helper for people to obtain information and execute tasks in daily life. However, there are significant differences in voice interaction among different users.
[0004] On the one hand, users' own language habits vary. Some people speak at a relatively fast speed, with quick thinking and concise and efficient expression; while some people speak at a slower speed and tend to elaborate their views more clearly and slowly. On the other hand, the environments where users are located are also complex and diverse. In noisy public places, such as shopping malls and stations, the background noise is relatively large, and it may be necessary for AI speech to appropriately slow down the speed so that users can hear the content clearly; while in a quiet private space, such as at home or in the office, users may hope that the voice can convey information quickly to improve the interaction efficiency. At the same time, for users of different age groups, their preferences for speech rate also vary. The elderly may be more adaptable to a slower speech rate for easier understanding and receiving information, while young people are often more able to accept a faster speech rate in pursuit of high efficiency and convenience.
[0005] Most early AI speech systems adopted a fixed speech rate. This "one-size-fits-all" approach obviously cannot meet the diverse needs of users. Even if some systems allow users to manually adjust the speech rate, this process is often rather cumbersome. Users need to carefully search for relevant parameters in the settings options and make adjustments, which is quite difficult for users who are not familiar with the operations. Moreover, users may need to frequently adjust the rate in different scenarios, and manual operations are difficult to adapt to changes in real time, resulting in a poor interaction experience. In the vehicle scenario, if a driver is distracted to operate and adjust the speech rate while driving, it may also pose a safety hazard. Summary of the Invention
[0006] The present invention provides an AI voice rate adjustment method for an AI platform, which combines an autoregressive pre-trained language model to improve the accuracy and efficiency of keyword extraction from voice signals, and adjusts the speech rate of the AI voice response based on the response text obtained from the keywords and the speech rate of the user's voice, greatly enhancing the user experience. Whether the user is accustomed to expressing quickly or prefers to narrate slowly, or is in a noisy or quiet environment, the AI voice can adapt to their speech rate requirements in real time, making the interaction smooth and natural, and enhancing the user's satisfaction and dependence on the voice service.
[0007] The present invention provides an AI voice rate adjustment method for an AI platform, including:
[0008] Obtain the voice signal input by the user, and perform keyword extraction on the voice signal using a preset autoregressive pre-trained language model;
[0009] Search and determine the response text for the voice content corresponding to the voice signal according to the extracted keywords;
[0010] Perform speech rate analysis on the voice signal to obtain the speech rate of the voice signal as the initial speech rate, and at the same time obtain the speech rate of the AI voice for the response text as the target speech rate;
[0011] Compare the initial speech rate with the target speech rate. When the difference between the initial speech rate and the target speech rate is less than a set threshold, use the target speech rate to perform AI voice broadcast on the response text;
[0012] When the difference between the initial speech rate and the target speech rate is greater than or equal to the set threshold, adjust the target speech rate according to the initial speech rate.
[0013] Further, the step of obtaining the voice signal input by the user and performing keyword extraction on the voice signal using a preset autoregressive pre-trained language model includes:
[0014] Obtain the voice signal input by the user, and construct word vector features for the voice signal to obtain the word vector features of the primary selected keywords;
[0015] Perform duplicate removal processing on the primary selected keywords according to the word vector features and word frequencies of the primary selected keywords to obtain a keyword candidate set;
[0016] Input the keyword candidate set into a preset autoregressive pre-trained language model to extract the target keywords of the voice signal.
[0017] Further, the step of obtaining the voice signal input by the user and constructing word vector features for the voice signal to obtain the word vector features of the primary selected keywords includes:
[0018] An audio processing tool is used to read the user's voice signal and draw the time-domain waveform diagram of the voice signal. At each sampling point of the voice signal, the superimposed amplitude of the voice signal is defined as:
[0019] X0 = s i + n i (i = 1, 2, …, p)
[0020] where s i and n i represent the amplitudes of the voice signal and the noise signal at the sampling point i respectively; p represents the number of sampling points; the voice signal is subjected to multi-modal decomposition to obtain voice components with w scale features, and the calculation formula is:
[0021]
[0022] where C j represents the filtering scale factor of the jth voice component; r0 represents the noise intensity; x(t) represents the voice signal after decomposition processing; furthermore, the non-stationary feature and non-linear feature of the voice signal are obtained, and the expression is as follows:
[0023]
[0024] where h1 represents the time variable; v0 represents the time window width of the low-frequency band of the voice signal; p t represents the Stockwell coefficient; f(t) represents the soft limiting function; based on the voice multi-modal decomposition result, the voice is denoised, and the expression is:
[0025]
[0026] where q1 represents the reconstructed state space dimension of the signal, q2 represents the scalar signal sequence, h2 represents the sampling point embedding dimension, v1 represents the corrected state vector, and α0 represents the sampling delay time; the feature matrix of the lexical method is obtained, and the expression is as follows:
[0027] S f = a(t) × d w
[0028] where d w represents the length of the voice signal; the construction expression of the word vector feature of the voice signal is obtained as:
[0029]
[0030] where V c represents the word vector feature of the primary selected keyword, β0 represents the word vector dimension, and ω t represents the unit length basis vector of the voice signal.
[0031] Further, the step of performing duplicate removal processing on the primary selected keywords according to the word vector features and word frequencies of the primary selected keywords to obtain a keyword candidate set includes:
[0032] According to the word frequency information included in the word vector features of the keywords in the voice signal, a feature dictionary is established, and the text is represented by a bag-of-words model. Then, the main word matrix of each word frequency can be expressed as:
[0033]
[0034] where z l represents the distribution probability of the l-th word frequency, and e m represents the distribution probability of the m-th topic word; a classification threshold is set for the voice sample set, and the word frequencies greater than the classification threshold are determined as topic words. The expression is as follows:
[0035]
[0036] where W r represents the topic vector containing text information, and t u represents the feature information fusion coefficient; the gradient value of the topic words in the voice signal is initialized, that is:
[0037]
[0038] where θ n′ represents the maximum depth parameter of the word vector, and T0 represents the discretized eigenvalue after normalization processing; assuming E u is a function characterizing the frequency change of the voice signal, then the zero-crossing rate of the signal is expressed as:
[0039]
[0040] where x(m) represents the sign function; the spectra of the voice signal are superimposed and calculated to obtain the horizontal structure differential vector of the signal, that is:
[0041]
[0042] where a b represents the linear spectrum of the voice signal, g0 represents the frequency warping factor, J t represents the Mel cepstrum order, and y p represents the number of iterations;
[0043] Add a granularity identifier at the starting position of the voice signal to obtain the semantic vector parameters of the full-segment voice signal. The expression is as follows:
[0044]
[0045] Among them, f0 represents the length of the speech sequence, P0 represents the output vector corresponding to the starting position identifier, and t r represents the sampling time; all the obtained semantic vector parameters are clustered to obtain the coefficient of variation of keywords with similar or synonymous semantics. The calculation formula is as follows:
[0046] L c =B y ×H t ×μ
[0047] Among them, H t represents the coefficient of inner linearization approximation, and μ represents the number of clustering centers; keywords with similar and repeated semantics are excluded from the candidate list, that is:
[0048]
[0049] Among them, Q t represents the candidate set of keywords excluding those with similar or repeated semantics, φ0 represents the dimension of the speech sequence length, and j p represents the feature order of the speech keyword.
[0050] Furthermore, the step of inputting the keyword candidate set into a preset autoregressive pre-trained language model to extract the target keyword of the speech signal includes:
[0051] Using an autoencoder pre-trained language model with bidirectional feature representation to extract keywords of the speech signal. The process is expressed as:
[0052]
[0053] Among them, Q t represents the candidate set of keywords excluding those with similar or repeated semantics, represents the model regularization coefficient, h g represents the model control parameter, R x represents the model training task; the trained model is deeply trained in a large-scale corpus to output more general semantic representation information. The expression is:
[0054]
[0055] Among them, b x represents the semantic attribute of the keyword, h s represents a multivariate non-linear function; the speech signal is input into the trained model by means of context semantic embedding to obtain the most critical text subsequence in the audio, that is:
[0056]
[0057] Among them, δ0 represents minimizing the loss function, and k t represents the activation function; taking the text subsequence of the voice signal as the semantic representation random vector, then this vector can be expressed as:
[0058] r v = N t × ρ d
[0059] Among them, ρ d represents the attribute constraint; using the cosine similarity to calculate the similarity between the candidate keyword and the semantic vector of the voice signal, and the calculation formula is:
[0060]
[0061] Among them, η represents the voice signal in the input model, and g s represents the hidden state parameter of the pre-trained model, and ω s represents the model hidden vector, represents the signal floating-point number in the interval [0, 1], and g b represents the truncation error of the encoder;
[0062] Using the above formula to calculate the similarity between the keywords in all candidate sets and the voice signal sequence, and sorting the calculation results in descending order, and taking the keywords ranked in the top set number as the voice signal keyword extraction result.
[0063] Furthermore, the step of analyzing the speech rate of the voice signal to obtain the speech rate of the voice signal as the initial speech rate, and at the same time calculating the speech rate of the AI voice for the reply text as the target speech rate includes:
[0064] Intercepting the speech segment of each complete sentence in the voice signal, recording the time length of the speech segment, and at the same time counting the number of words corresponding to the speech segment;
[0065] Calculating the speech rate of the speech segment according to the time length and the number of words of the speech segment; the calculation formula is: speech rate = number of words / time length;
[0066] Calculating the average speech rate of multiple speech segments according to the speech rate of each speech segment, and taking the average speech rate as the initial speech rate;
[0067] Obtaining the AI voice style actively selected by the user for the reply text, and taking the speech rate corresponding to the AI voice style as the target speech rate.
[0068] Furthermore, the step of adjusting the target speech rate according to the initial speech rate when the difference between the initial speech rate and the target speech rate is greater than or equal to the set threshold includes:
[0069] When the difference between the initial speech rate and the target speech rate is greater than or equal to a set threshold, calculate the ratio A of the initial speech rate to the target speech rate; where A = initial speech rate / target speech rate;
[0070] Adjust the target speech rate according to the ratio A; where the adjusted target speech rate = target speech rate × A.
[0071] The present invention also provides an AI speech rate adjustment device for an AI platform, including:
[0072] An acquisition module, configured to acquire a voice signal input by a user, and extract keywords from the voice signal by using a preset autoregressive pre-trained language model;
[0073] A determination module, configured to search and determine a reply text for the voice content corresponding to the voice signal according to the extracted keywords;
[0074] An analysis module, configured to perform speech rate analysis on the voice signal to obtain the speech rate of the voice signal as the initial speech rate, and at the same time obtain the speech rate of the AI voice for the reply text as the target speech rate;
[0075] A comparison module, configured to compare the initial speech rate with the target speech rate. When the difference between the initial speech rate and the target speech rate is less than the set threshold, use the target speech rate to perform AI voice broadcast on the reply text;
[0076] An adjustment module, configured to, when the difference between the initial speech rate and the target speech rate is greater than or equal to the set threshold, adjust the target speech rate according to the initial speech rate.
[0077] The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0078] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0079] The beneficial effects of the present invention are:
[0080] The present invention constructs speech signal word vector features, eliminates duplicate meanings of initially selected keywords, and uses an autoregressive language model to extract keywords from speech signals, achieving effective extraction of keywords from long speech signals. Furthermore, reply texts are searched based on the keywords extracted from speech, and then the AI speech selected by the user is used for reply according to the reply texts. Since different AI speeches have different speeds, the speed of the AI speech is adjusted according to the input speech speed of the user, so as to better adapt to the user's speech speed habit, greatly improving the user experience. Whether the user is used to expressing quickly or prefers slow narration, or is in a noisy or quiet environment, the AI speech can adapt to its speech speed requirements in real time, making the interaction smooth and natural, and enhancing the user's satisfaction and dependence on the speech service. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention.
[0082] Figure 2 It is a schematic structural diagram of the device according to an embodiment of the present invention.
[0083] Figure 3 It is a schematic internal structure diagram of a computer device according to an embodiment of the present invention.
[0084] The implementation, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0085] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0086] As Figure 1 shown, the present invention provides an AI speech speed adjustment method for an AI platform, including:
[0087] S1. Obtain the speech signal input by the user, and extract keywords from the speech signal using a preset autoregressive pre-trained language model.
[0088] Specifically, it includes the following steps:
[0089] S101. Obtain the speech signal input by the user, and construct word vector features for the speech signal to obtain the word vector features of the initially selected keywords.
[0090] The audio processing tool Librosa and the visualization processing tool Matplotlib can be used to read, process, analyze, and visualize audio signals, and plot their time-domain waveform diagrams, thereby extracting the parameter features of speech signals. The speech signal may be mixed with white noise signals within the sampling point interval. Therefore, the background noise of the speech signal is regarded as a Gaussian distribution with a mean of 0 and a fixed standard deviation. Then, at each sampling point of the speech signal, the superimposed amplitude of the speech signal can be expressed as:
[0091] X0 = s i + n i (i = 1, 2, …, p)
[0092] where s i and n i respectively represent the amplitudes of the speech signal and the noise signal at the sampling point i; p represents the number of sampling points.
[0093] For the convenience of subsequent extraction of speech signal keywords, the noise signal in the original audio should be removed. According to experience, pure speech signals include voiced sounds and voiceless sounds. Among them, compared with voiceless sounds, the periodicity of voiced sounds is more obvious and mainly reflects the pitch characteristics of speech signals; while voiceless sounds are reflected in the high-frequency part of the speech signal. Based on this principle, the original speech signal is decomposed by multi-modal decomposition to obtain speech components with w scale features. The calculation formula is as follows:
[0094]
[0095] where C j represents the filtering scale factor of the jth speech component; r0 represents the noise intensity; x(t) represents the speech signal after decomposition processing.
[0096] By performing high-pass scale filtering on the voiced and voiceless components of the low-frequency band noisy signal, the non-stationary and non-linear characteristics of the speech signal can be obtained. The expression is as follows:
[0097]
[0098] where h1 represents the time variable; v0 represents the time window width of the low-frequency band of the speech signal; p t represents the Stockwell coefficient; f(t) represents the soft limiting function. Based on the speech multi-modal decomposition results, the speech is denoised. The expression is:
[0099]
[0100] Among them, q1 represents the reconstructed state space dimension of the signal, q2 represents the scalar signal sequence, h2 represents the sampling point embedding dimension, v1 represents the corrected state vector, and α0 represents the sampling delay time; furthermore, the feature matrix of the lexical method is obtained, and the expression is as follows:
[0101] S f = a(t) × d w
[0102] Among them, d w represents the length of the speech signal.
[0103] Thus, the construction expression of the word vector feature of the speech signal can be obtained as follows:
[0104]
[0105] Among them, V c represents the word vector feature of the primary selected keywords, β0 represents the word vector dimension, and ω t represents the unit length basis vector of the speech signal.
[0106] By performing multi-modal decomposition on the speech signal to obtain high-frequency and low-frequency components, combining the non-stationary and non-linear characteristics of the speech signal, removing the noise interference therein, and then constructing the keyword word vector feature, it creates favorable conditions for removing duplicate meanings of the primary selected keywords.
[0107] S102. Perform duplicate removal processing on the primary selected keywords according to the word vector features and word frequencies of the primary selected keywords to obtain a keyword candidate set.
[0108] There are many topic words to be extracted from a piece of speech signal, and for the extracted keywords, different aspects of the topic to which the speech signal belongs need to be represented. Therefore, based on the keyword word vector features, duplicate removal processing should be performed on the meanings of the primary selected keywords, and a keyword candidate set with different meanings should be retained. According to the word frequency information contained in each word vector feature in the speech signal, a feature dictionary is established for it, and the text is represented by a bag-of-words model. Then, the main word matrix of each word frequency can be expressed as:
[0109]
[0110] Among them, z l represents the distribution probability of the l-th word frequency, and e m represents the distribution probability of the m-th topic word.
[0111] Set a classification threshold for the speech sample set, and determine the word frequency greater than the classification threshold as the topic word. The expression is as follows:
[0112]
[0113] Among them, Wr Denote the topic vector containing text information, t u Denote the feature information fusion coefficient.
[0114] Initialize the gradient value of the topic word of the speech signal, i.e.:
[0115]
[0116] where, θ n′ Denote the maximum depth parameter of the word vector, and T0 denote the discretized eigenvalue after normalization processing.
[0117] Assume E u Is a function representing the frequency change of the speech signal, then the zero-crossing rate of the signal is expressed as:
[0118]
[0119] where, x(m) denotes the sign function.
[0120] Perform superposition calculation on the spectrum of the speech signal to obtain the horizontal structure differential vector of the signal, i.e.:
[0121]
[0122] where, a b Denote the linear spectrum of the speech signal, g0 denote the frequency warping factor, J t Denote the Mel cepstrum order, y p Denote the number of iterations.
[0123] Add a granularity identifier at the starting position of the speech signal to obtain the semantic vector parameter of the full-segment speech signal, and the expression is as follows:
[0124]
[0125] where, f0 denote the length of the speech sequence, P0 denote the output vector corresponding to the starting position identifier, t r Denote the sampling time.
[0126] Cluster all the obtained semantic vector parameters to obtain the coefficient of variation of keywords with similar or synonymous semantics, and the calculation formula is as follows:
[0127] L c = B y × H t × μ
[0128] where, H t Denote the inner linearization approximation coefficient, and μ denote the number of cluster centers.
[0129] Exclude keywords with similar semantics and duplicate semantics from the candidate list, that is:
[0130]
[0131] Among them, Q t represents the candidate set of keywords with similar or duplicate semantics excluded, φ0 represents the dimension of the speech sequence length, and j p represents the feature order of the speech keyword.
[0132] Based on the construction result of the keyword word vector feature of the speech signal, identify the differential vector of the speech signal feature, combine the keyword semantic vector parameters, and use the clustering algorithm to remove the candidate keywords with duplicate semantics among them, thereby obtaining the primary keyword candidate set, laying a foundation for finally realizing keyword extraction.
[0133] S103. Input the keyword candidate set into a preset autoregressive pre-trained language model to extract the target keyword of the speech signal.
[0134] After processing, the speech signal and the keyword signal sequence are already at the same dimension level. Therefore, the keyword with a higher similarity to the full speech signal vector can be directly used as the feature keyword of this segment of speech text.
[0135] Considering that the feature of the speech signal keyword is bidirectional and the structure of a single character is non-entity, therefore, this paper uses the autoregressive pre-trained language model with bidirectional feature representation in NLP technology to extract the keyword of the speech signal. In the model, a multi-head autoencoder is used to enable the model to model and obtain the text sequence, and then obtain smaller-grained sub-words as the basic semantic units, and thus perform autoregressive pre-training on the model. This process can be expressed as:
[0136]
[0137] Among them, Q t represents the candidate set of keywords with similar or duplicate semantics excluded, represents the model regularization coefficient, h g represents the model control parameter, R x represents the model training task.
[0138] Train the trained model in a large-scale corpus for in-depth training to output more general semantic representation information. The expression is:
[0139]
[0140] Among them, b x represents the semantic attribute of the keyword, h s represents the multivariate non-linear function.
[0141] The speech signal is input into the trained model in a context semantic embedding manner to obtain the most critical text subsequence in the audio, that is:
[0142]
[0143] where δ0 represents minimizing the loss function, and k t represents the activation function.
[0144] Taking the speech signal text subsequence as a semantic representation random vector, then this vector can be expressed as:
[0145] r v = N t × ρ d
[0146] where ρ d represents the attribute constraint.
[0147] The cosine similarity is used to calculate the similarity between the candidate keywords and the speech signal semantic vector, and the calculation formula is:
[0148]
[0149] where η represents the speech signal input into the model, and g s represents the hidden state parameter of the pre-trained model, and ω s represents the model hidden vector, represents the signal floating-point number in the range of [0, 1], and g b represents the truncation error of the encoder.
[0150] The similarity between the keywords in all candidate sets and the speech signal sequence is calculated using the above formula, and the calculation results are sorted in descending order. The keywords ranked in the top set number are used as the speech signal keyword extraction results to achieve keyword extraction.
[0151] S2. Search and determine the response text for the speech content corresponding to the speech signal according to the extracted keywords.
[0152] By extracting the keywords in the user's speech, the intelligent speech assistant can accurately understand the user's needs and intentions, and thus provide accurate services and answers. For example, when the user says "I want to know the lyrics of the climax part of Jay Chou's 'Qi Li Xiang'", the speech assistant extracts keywords such as "Jay Chou", "'Qi Li Xiang'", "climax", "lyrics", etc., and can accurately output the corresponding song lyrics.
[0153] S3. Analyze the speech rate of the speech signal to obtain the speech rate of the speech signal as the initial speech rate, and at the same time obtain the speech rate of the AI speech for the response text as the target speech rate.
[0154] The specific steps include:
[0155] S301, intercepting each complete sentence of the speech segment in the speech signal, recording the time length of the speech segment, and counting the number of characters corresponding to the speech segment.
[0156] For example, if a complete speech signal includes four sentences, four speech segments corresponding to the four sentences are respectively intercepted, and four time lengths are recorded, and the number of words corresponding to the four speech segments is counted at the same time.
[0157] S302, calculating the speaking speed of the voice segment according to the time length and the number of characters of the voice segment; wherein the calculation formula is: speaking speed = number of characters / time length.
[0158] S303: Calculate an average speaking rate of the plurality of speech segments according to the speaking rate of each of the speech segments, and use the average speaking rate as the initial speaking rate.
[0159] For example, if the speaking speeds of four speech segments are calculated to be 130 words / minute, 140 words / minute, 145 words / minute, and 135 words / minute respectively, then the average speaking speed of the four segments is calculated to be (130+140+145+135) / 4=137.5 words / minute, which is rounded to 138 words / minute. This speed is used as the initial speaking speed.
[0160] S304: Obtain the AI voice style actively selected by the user for the reply text, and use the speaking speed corresponding to the AI voice style as the target speaking speed.
[0161] The AI platform system stores multiple AI voice styles, such as regular male voice, regular female voice, loli, uncle, etc. Different AI voice styles correspond to different speaking speeds. When using the AI platform system, users will initially select an AI voice style and use the speaking speed of the AI voice style selected by the user as the target speaking speed.
[0162] S4. Compare the initial speaking speed with the target speaking speed. When the difference between the initial speaking speed and the target speaking speed is less than a set threshold, use the target speaking speed to perform AI voice broadcasting on the reply text.
[0163] The calculated difference is compared with the set threshold. When the difference between the initial speech rate and the target speech rate of the AI voice is within the set threshold, it is considered that the conversation speed between the user and the AI matches and a normal conversation can be carried out. When the difference value is outside the preset threshold, it is considered that the conversation speed does not match and will affect the normal conversation.
[0164] S5. When the difference between the initial speech rate and the target speech rate is greater than or equal to a set threshold, adjust the target speech rate according to the initial speech rate.
[0165] If the difference exceeds the set threshold, it indicates that the difference between the target speech rate and the initial speech rate of the current AI speech is too large, and the speech rate of the AI needs to be adjusted so that the difference between the user's speech rate and the AI speech rate is within the set threshold range to ensure the normal progress of the conversation.
[0166] Specifically, it includes the following steps:
[0167] S501. When the difference between the initial speech rate and the target speech rate is greater than or equal to the set threshold, calculate the ratio A of the initial speech rate to the target speech rate; where A = initial speech rate / target speech rate;
[0168] S502. Adjust the target speech rate according to the ratio A; where the adjusted target speech rate = target speech rate × A.
[0169] For example, the user's speech rate, that is, the initial speech rate is 125 words per minute, and the speech rate of the AI speech, that is, the target speech rate is 135 words per minute, then 125 / 135 ≈ 0.93. Then adjust the target speech rate to make the final target speech rate 0.93 times the initial target speech rate, so that the final target speech rate matches the user's speech rate and makes the interaction smooth and natural.
[0170] The present invention constructs speech signal word vector features, performs duplicate removal on the primary selected keyword meanings, and uses an autoregressive language model to achieve keyword extraction from speech signals, realizing effective extraction of keywords from long speech signals. Furthermore, based on the keywords extracted from the speech, a reply text is searched, and then according to the reply text, the AI speech selected by the user is used for reply. Different AI speeches also have different speeds. Therefore, the AI speech rate is adjusted according to the input speech rate of the user to better adapt to the user's speech rate habit, greatly improving the user experience. Whether the user is used to expressing quickly or prefers slow narration, or is in a noisy or quiet environment, the AI speech can adapt to its speech rate requirements in real time, making the interaction smooth and natural, and enhancing the user's satisfaction and dependence on the speech service.
[0171] As Figure 2 shown, the present invention also provides an AI speech rate adjustment device for an AI platform, including:
[0172] An acquisition module 1, configured to acquire the speech signal input by the user, and perform keyword extraction on the speech signal by using a preset autoregressive pre-trained language model;
[0173] A determination module 2, configured to search and determine a reply text for the speech content corresponding to the speech signal according to the extracted keywords;
[0174] An analysis module 3, configured to perform a speech rate analysis on the speech signal to obtain the speech rate of the speech signal as an initial speech rate, and at the same time obtain the speech rate of the AI speech for the reply text as a target speech rate;
[0175] A comparison module 4, configured to compare the initial speech rate with the target speech rate, and when the difference between the initial speech rate and the target speech rate is less than a set threshold, perform AI speech broadcast on the reply text using the target speech rate;
[0176] An adjustment module 5, configured to, when the difference between the initial speech rate and the target speech rate is greater than or equal to the set threshold, adjust the target speech rate according to the initial speech rate.
[0177] In one embodiment, the acquisition module 1 includes:
[0178] A construction unit, configured to acquire a speech signal input by a user, and perform a word vector feature construction on the speech signal to obtain the word vector feature of a preliminary selected keyword;
[0179] A duplicate removal unit, configured to perform duplicate removal processing on the preliminary selected keyword according to the word vector feature and word frequency of the preliminary selected keyword to obtain a keyword candidate set;
[0180] An extraction unit, configured to input the keyword candidate set into a preset autoregressive pre-trained language model to extract the target keyword of the speech signal.
[0181] In one embodiment, the construction unit includes:
[0182] Read the speech signal of the user by using an audio processing tool, and draw a time-domain waveform diagram of the speech signal. At each sampling point of the speech signal, define the superimposed amplitude of the speech signal as:
[0183] X0 = s i + n i (i = 1, 2, …, p)
[0184] where s i , n i respectively represent the amplitudes of the speech signal and the noise signal at the sampling point i; p represents the number of sampling points; perform multi-modal decomposition on the speech signal to obtain a speech component with w scale features, and the calculation formula is:
[0185]
[0186] where C jThe filtering scale factor representing the j-th voice component; r0 represents the noise intensity; x(t) represents the voice signal after decomposition processing; furthermore, the non-stationary feature and non-linear feature of the voice signal are obtained, and the expression is as follows:
[0187]
[0188] Among them, h1 represents the time variable; v0 represents the time window width of the low-frequency band of the voice signal; p t represents the Stockwell coefficient; f(t) represents the soft limiting function; based on the voice multi-modal decomposition result, noise reduction processing is performed on the voice, and the expression is:
[0189]
[0190] Among them, q1 represents the reconstruction state space dimension of the signal, q2 represents the scalar signal sequence, h2 represents the sampling point embedding dimension, v1 represents the corrected state vector, and α0 represents the sampling delay time; the feature matrix of the lexical method is obtained, and the expression is as follows:
[0191] S f = a(t) × d w
[0192] Among them, d w represents the length of the voice signal; the construction expression for obtaining the word vector feature of the voice signal is:
[0193]
[0194] Among them, V c represents the word vector feature of the primary selected keyword, β0 represents the word vector dimension, and ω t represents the unit length basis vector of the voice signal.
[0195] In one embodiment, the duplicate removal unit includes:
[0196] According to the word frequency information included in the word vector feature of the keyword in the voice signal, a feature dictionary is established, and the text is represented by a bag-of-words model. Then, the main word matrix of each word frequency can be expressed as:
[0197]
[0198] Among them, z l represents the distribution probability of the l-th word frequency, and e m represents the distribution probability of the m-th topic word; a classification threshold is set for the voice sample set, and the word frequency greater than the classification threshold is determined as the topic word, and the expression is as follows:
[0199]
[0200] Among them, W r represents the topic vector containing text information, and t u represents the feature information fusion coefficient; initialize the gradient value of the topic word of the speech signal, that is:
[0201]
[0202] Among them, θ n′ represents the maximum depth parameter of the word vector, and T0 represents the discretized eigenvalue after normalization processing; assume that E u is a function representing the frequency change of the speech signal, then the zero-crossing rate of the signal is expressed as:
[0203]
[0204] Among them, x(m) represents the sign function; perform superposition calculation on the spectrum of the speech signal to obtain the horizontal structure differential vector of the signal, that is:
[0205]
[0206] Among them, a b represents the linear spectrum of the speech signal, g0 represents the frequency warping factor, and J t represents the Mel cepstrum order, and y p represents the number of iterations;
[0207] Add a granularity identifier at the starting position of the speech signal to obtain the semantic vector parameter of the full-segment speech signal, and the expression is as follows:
[0208]
[0209] Among them, f0 represents the length of the speech sequence, P0 represents the output vector corresponding to the starting position identifier, and t r represents the sampling time; perform clustering on all the obtained semantic vector parameters to obtain the coefficient of variation of keywords with similar or synonymous semantics, and the calculation formula is as follows:
[0210] L c = B y × H t × μ
[0211] Among them, H t represents the inner linearization approximation coefficient, and μ represents the number of clustering centers; exclude keywords with similar and repeated semantics from the candidate list, that is:
[0212]
[0213] Among them, Q t represents the candidate set of keywords excluding those with similar or repeated semantics, φ0 represents the dimension of the speech sequence length, and jp Represents the feature order of the voice keyword.
[0214] In one embodiment, the extraction unit includes:
[0215] The keyword of the voice signal is extracted by using an autoencoder pre-trained language model with bidirectional feature representation, and the process is expressed as:
[0216]
[0217] where Q t Represents the candidate keyword set after removing semantically similar or duplicate keywords, Represents the model regularization coefficient, h g Represents the model control parameter, R x Represents the model training task; the trained model is deeply trained in a large-scale corpus to output more general semantic representation information, and the expression is:
[0218]
[0219] where b x Represents the semantic attribute of the keyword, h s Represents a multivariate non-linear function; the voice signal is input into the trained model by using the context semantic embedding method to obtain the most critical text subsequence in the audio, that is:
[0220]
[0221] where δ0 represents minimizing the loss function, k t Represents the activation function; the voice signal text subsequence is used as a semantic representation random vector, and the vector can be expressed as:
[0222] r v = N t × ρ d
[0223] where ρ d Represents the attribute constraint; the cosine similarity is used to calculate the similarity between the candidate keyword and the voice signal semantic vector, and the calculation formula is:
[0224]
[0225] where η represents the voice signal input into the model, g s Represents the hidden state parameter of the pre-trained model, ω s Represents the model hidden vector, Represents the signal floating point number in the range of [0,1], g b Represents the truncation error of the encoder;
[0226] Calculate the similarity between the keywords and the speech signal sequence in all candidate sets using the above formula, and sort the calculation results in descending order. Take the keywords ranked in the top set number as the extraction result of the speech signal keywords.
[0227] In one embodiment, the analysis module 3 includes:
[0228] An interception unit, configured to intercept the speech segment of each whole sentence in the speech signal, record the time length of the speech segment, and at the same time count the number of words corresponding to the speech segment;
[0229] A first calculation unit, configured to calculate the speech rate of the speech segment according to the time length and the number of words of the speech segment; the calculation formula is: speech rate = number of words / time length;
[0230] A second calculation unit, configured to calculate the average speech rate of multiple speech segments according to the speech rate of each speech segment, and use the average speech rate as the initial speech rate;
[0231] A style acquisition unit, configured to acquire the AI speech style actively selected by the user for the reply text, and use the speech rate corresponding to the AI speech style as the target speech rate.
[0232] In one embodiment, the adjustment module 5 includes:
[0233] A ratio calculation unit, configured to calculate the ratio A of the initial speech rate to the target speech rate when the difference between the initial speech rate and the target speech rate is greater than or equal to a set threshold; where A = initial speech rate / target speech rate;
[0234] A speech rate adjustment unit, configured to adjust the target speech rate according to the ratio A; where the adjusted target speech rate = target speech rate × A.
[0235] Each of the above modules and units is used to correspondingly execute each step in the AI speech rate adjustment method of the above AI platform, and its specific implementation manner refers to the above method embodiment and will not be elaborated here.
[0236] As Figure 3 shown, the present invention also provides a computer device, which may be a server, and its internal structure may be as Figure 3As shown. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store all the data required for the process of the AI voice rate adjustment method of the AI platform. The network interface of the computer device is used to communicate with an external terminal via a network connection. The computer program, when executed by the processor, implements the AI voice rate adjustment method of the AI platform.
[0237] Those skilled in the art can understand that Figure 3 the structure shown in
[0238] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied.
[0239] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database, or other media provided in this application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0240] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, apparatus, article or method comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, apparatus, article or method. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, apparatus, article or method comprising such element.
[0241] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. An AI voice rate adjustment method for an AI platform, characterized in that, Including: Obtain the voice signal input by the user, and extract keywords from the voice signal by using a preset autoregressive pre-trained language model; Search and determine a reply text for the voice content corresponding to the voice signal according to the extracted keywords; Analyze the speech rate of the voice signal to obtain the speech rate of the voice signal as the initial speech rate, and at the same time obtain the speech rate of the AI voice for the reply text as the target speech rate; Compare the initial speech rate with the target speech rate. When the difference between the initial speech rate and the target speech rate is less than a set threshold, use the target speech rate to perform AI voice broadcast on the reply text; When the difference between the initial speech rate and the target speech rate is greater than or equal to the set threshold, adjust the target speech rate according to the initial speech rate.
2. The AI voice rate adjustment method of the AI platform according to claim 1, wherein, The step of obtaining the voice signal input by the user and extracting keywords from the voice signal by using a preset autoregressive pre-trained language model includes: Obtain the voice signal input by the user, and construct a word vector feature for the voice signal to obtain the word vector feature of the primary selected keywords; Deduplicate the primary selected keywords according to the word vector features and word frequencies of the primary selected keywords to obtain a keyword candidate set; Input the keyword candidate set into a preset autoregressive pre-trained language model to extract the target keywords of the voice signal.
3. The AI voice rate adjustment method of the AI platform according to claim 2, wherein, The step of obtaining the voice signal input by the user and constructing a word vector feature for the voice signal to obtain the word vector feature of the primary selected keywords includes: Use an audio processing tool to read the voice signal of the user, and draw the time-domain waveform diagram of the voice signal. At each sampling point of the voice signal, define the superimposed amplitude of the voice signal as: X0 = s i + n i (i = 1, 2, …, p) where s i and n i represent the amplitudes of the speech signal and the noise signal at sampling point i, respectively; p represents the number of sampling points; Perform multi-modal decomposition on the voice signal to obtain voice components with w scale features. The calculation formula is: Among them, C j represents the filtering scale factor of the j-th voice component; r0 represents the noise intensity; x(t) represents the voice signal after decomposition processing; furthermore, the non-stationary feature and non-linear feature of the voice signal are obtained, and the expression is as follows: where h1 represents a time variable; v0 represents the time window width of the low-frequency band of the voice signal; p t represents the Stockwell coefficient; f(t) represents a soft-limiting function; based on the voice multi-modal decomposition result, voice denoising processing is performed, and the expression is: Among them, q1 represents the reconstructed state space dimension of the signal, q2 represents the scalar signal sequence, h2 represents the sampling point embedding dimension, v1 represents the corrected state vector, and α0 represents the sampling delay time; obtain the feature matrix of morphology, and the expression is as follows: S f = a(t) × d w where d w represents the length of the speech signal; the construction expression for obtaining the word vector features of the speech signal is: Among them, V c represents the word vector feature of the initially selected keywords, β0 represents the dimension of the word vector, and ω t represents the unit-length basis vector of the voice signal.
4. The AI voice rate adjustment method of the AI platform according to claim 3, characterized in that, The step of deduplicating the primary selected keywords according to the word vector features and word frequencies of the primary selected keywords to obtain a keyword candidate set includes: Establish a feature dictionary according to the word frequency information included in the word vector features of the keywords in the voice signal, and represent the text by using a bag-of-words model. Then the main word matrix of each word frequency can be represented as: Among them, z l represents the distribution probability of the l-th word frequency, and e m represents the distribution probability of the m-th topic word; a classification threshold is set for the speech sample set, and the word frequency greater than the classification threshold is determined as the topic word. The expression is as follows: Among them, W r represents the topic vector containing text information, and t u represents the feature information fusion coefficient; initialize the gradient value of the topic words of the speech signal, that is: Among them, θ n′ represents the maximum depth parameter of the word vector, and T0 represents the discretized eigenvalue after normalization; assume E u is a function characterizing the frequency change of the speech signal, then the zero-crossing rate of the signal is expressed as: Among them, x(m) represents the sign function; superimpose and calculate the frequency spectrum of the voice signal to obtain the horizontal structure differential vector of the signal, that is: where a b represents the linear spectrum of the speech signal, g0 represents the frequency warping factor, and J t represents the Mel cepstrum order, and y p represents the number of iterations; Add a granularity identifier at the starting position of the voice signal to obtain the semantic vector parameter of the full-segment voice signal. The expression is as follows: Among them, f0 represents the length of the speech sequence, P0 represents the output vector corresponding to the start position identifier, and t r represents the sampling time; all the obtained semantic vector parameters are clustered to obtain the coefficient of variation of keywords with similar or synonymous semantics, and the calculation formula is as follows: L c = B y × H t × μ Among them, H t represents the inner linearization approximation coefficient, and μ represents the number of cluster centers; exclude the keywords with similar semantics and repeated semantics from the candidate list, that is: Among them, Q t represents the candidate set of keywords with similar or repeated semantics removed, φ0 represents the dimension of the speech sequence length, and j p represents the feature order of the speech keyword.
5. The AI voice rate adjustment method of the AI platform according to claim 4, characterized in that The step of inputting the keyword candidate set into a preset autoregressive pre-trained language model to extract the target keywords of the voice signal includes: Use an autoencoder pre-trained language model with bidirectional feature representation to extract keywords of the voice signal. The process is represented as: Among them, Q t represents the candidate set of keywords with similar or repeated semantics removed, represents the model regularization coefficient, h g represents the model control parameter, R x represents the model training task; the trained model is deeply trained in a large-scale corpus to output more general semantic representation information, and the expression is: where b x represents the semantic attribute of the keyword, and h s represents a multivariable non-linear function; the speech signal is input into the trained model by using the context semantic embedding method to obtain the most critical text subsequence in the audio, that is: Among them, δ0 represents minimizing the loss function, and k t represents the activation function; taking the text subsequence of the speech signal as the semantic representation random vector, then this vector can be expressed as: r v = N t × ρ d where ρ d represents an attribute constraint; the cosine similarity is used to calculate the similarity between the candidate keyword and the semantic vector of the speech signal, and the calculation formula is: Among them, η represents the speech signal in the input model, and g s represents the hidden state parameter of the pre-trained model, ω s represents the model hidden vector, ζ0 represents the signal floating-point number in the interval [0, 1], and g b represents the truncation error of the encoder; Calculate the similarity between the keywords and the speech signal sequence in all candidate sets using the above formula, and sort the calculation results in descending order. Take the keywords ranked among the top set number as the extraction result of the speech signal keywords.
6. The AI voice rate adjustment method of the AI platform according to claim 1, characterized in that The step of analyzing the speech rate of the speech signal to obtain the speech rate of the speech signal as the initial speech rate, and at the same time calculating the speech rate of the AI speech for the reply text as the target speech rate includes: Intercept the speech segments of each complete sentence in the speech signal, record the time length of the speech segments, and at the same time count the number of words corresponding to the speech segments; Calculate the speech rate of the speech segment according to the time length and the number of words of the speech segment; the calculation formula is: speech rate = number of words / time length; Calculate the average speech rate of multiple speech segments according to the speech rate of each speech segment, and take the average speech rate as the initial speech rate; Obtain the AI speech style actively selected by the user for the reply text, and take the speech rate corresponding to the AI speech style as the target speech rate.
7. The AI voice rate adjustment method of the AI platform according to claim 1, characterized in that The step of adjusting the target speech rate according to the initial speech rate when the difference between the initial speech rate and the target speech rate is greater than or equal to a set threshold includes: When the difference between the initial speech rate and the target speech rate is greater than or equal to a set threshold, calculate the ratio A of the initial speech rate to the target speech rate; where A = initial speech rate / target speech rate; Adjust the target speech rate according to the ratio A; where the adjusted target speech rate = target speech rate × A.
8. An AI voice rate adjustment device for an AI platform, characterized in that, Includes: An acquisition module for acquiring the speech signal input by the user and extracting keywords from the speech signal using a preset autoregressive pre-trained language model; A determination module for searching and determining a reply text for the speech content corresponding to the speech signal according to the extracted keywords; An analysis module for analyzing the speech rate of the speech signal to obtain the speech rate of the speech signal as the initial speech rate, and at the same time obtaining the speech rate of the AI speech for the reply text as the target speech rate; A comparison module for comparing the initial speech rate with the target speech rate. When the difference between the initial speech rate and the target speech rate is less than the set threshold, use the target speech rate to perform AI speech broadcast on the reply text; An adjustment module for adjusting the target speech rate according to the initial speech rate when the difference between the initial speech rate and the target speech rate is greater than or equal to the set threshold.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.