A user interaction method based on a smart TV
Through the voice authentication module of smart TV, users' voice characteristics are collected and analyzed, and the TV chaos caused by multiple people is solved, achieving more accurate user interaction.
Patent Information
- Application Number
- CN202411134473.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-19
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-08-19
AI Technical Summary
Smart TVs are easy to recognize chaos when multiple people control at the same time, resulting in chaos in control and affecting the user experience.
Through the voice authentication module, users' voice characteristics are collected, volume, speech speed and tone matching degree are calculated, and user authentication is performed to ensure that only specific users can perform voice control.
It effectively avoids TV chaos caused by multiple people's control and improves the accuracy and experience of user interaction.
Smart Images

Figure CN119172600B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of smart TVs, and specifically to a user interaction method based on smart TVs. Background Art
[0002] Smart TVs are new products formed under the impact of the Internet wave, aiming to bring users a more convenient experience. Currently, they have become the trend of TVs. Smart TVs break the shackles of traditional TVs by remote controls, realizing the four major functions of taking and watching, classifying and watching, multi-screen watching, and watching at any time, pushing the development of smart TVs to a new height. With the rapid development of smart TVs, the number of smart TV users is increasing. Different from traditional digital TVs, smart TVs have added various new functions such as web browsing, email sending and receiving, TV shopping, distance learning, telemedicine, stock trading, and information consultation. Under these new functions, more and more user interactions are required. Traditional digital TVs rely on the interaction of infrared remote controls. With the increasing number of TV channels, the disadvantages of traditional TV remote control methods for turning on / off the TV and changing channels are becoming more and more obvious. Smart TVs with voice control, gesture control, and even face recognition control have entered ordinary people's homes one after another, and many brands at home and abroad already have smart voice-controlled TVs.
[0003] However, there are still certain deficiencies in the current smart TVs in terms of the voice control channel selection function. For example, when multiple people are present at the same time, the TV will recognize the voices of multiple people and be controlled by multiple people, which is likely to cause chaos in TV control, making it difficult for consumers to watch normally and affecting the viewing experience, and unable to create a good interaction experience for users. Summary of the Invention
[0004] (I) Technical Problems to be Solved
[0005] In view of the deficiencies of the prior art, the present invention provides a user interaction method based on smart TVs, which has the advantages of authenticating the voices of specific users through a voice authentication module, thus avoiding being controlled by multiple people during the voice control process and causing chaos in TV control, and solves the above problems.
[0006] (II) Technical Solutions
[0007] To achieve the above object, the present invention provides the following technical solution: A user interaction method based on smart TVs, including the following steps:
[0008] S1. The data acquisition module collects the user's voice through a built-in microphone. The data acquisition module extracts the real-time audio signal Xh and the real-time syllable count Yjs from the collected user voice. The data acquisition module sends the collected data to the data processing module through the network;
[0009] S2. The data processing module calculates the volume Yl based on the real-time audio signal Xh, calculates the speech rate Ys based on the real-time syllable count Yjs, has a preset autocorrelation function Zxhs in the data processing module, calculates the pitch Yd based on the real-time audio signal Xh and the autocorrelation function Zxhs, preprocesses the real-time audio signal Xh to obtain the feature vector Tzxl, calculates the volume matching degree Ylpp, the speech rate matching degree Yspp, and the pitch matching degree Ydpp respectively based on the volume Yl, the speech rate Ys, and the pitch Yd, has a preset target feature vector Tzmb in the data processing module, and calculates the matching degree score Ppdf based on the feature vector Tzxl and the target feature vector Tzmb. The data processing module sends these data to the data analysis module through the network;
[0010] S3. The data analysis module has a preset volume matching degree threshold Ylpp yz , a speech rate matching degree threshold Yspp yz , a pitch matching degree threshold Ydpp yz , and a matching degree score threshold Ppdf yz . The data analysis module compares the volume matching degree Ylpp, the speech rate matching degree Yspp, the pitch matching degree Ydpp, and the matching degree score Ppdf with the volume matching degree threshold Ylpp yz , the speech rate matching degree threshold Yspp yz , the pitch matching degree threshold Ydpp yz , and the matching degree score threshold Ppdf yz respectively, and sends the comparison results to the authentication module through the network;
[0011] S4. When any two of the volume matching degree Ylpp, the speech rate matching degree Yspp, and the pitch matching degree Ydpp are greater than the corresponding thresholds and the matching degree score Ppdf is greater than the matching degree score threshold Ppdf yz , the authentication module determines that the user authentication is successful. When any two of the volume matching degree Ylpp, the speech rate matching degree Yspp, and the pitch matching degree Ydpp are less than the corresponding thresholds or the matching degree score Ppdf is less than the matching degree score threshold Ppdf yz , the authentication module determines that the user authentication is unsuccessful;
[0012] S5. The authentication module sends the authentication result to the control module through the network. When the user authentication is successful, the user can perform corresponding voice commands on the TV through the control module. When the user authentication is unsuccessful, the user cannot perform corresponding voice commands on the TV through the control module.
[0013] Preferably, the expression of the feature vector Tzxl is {Tzxl1, Tzxl2,..., Tzxln}, and the expression of the target feature vector Tzmb is {Tzmb1, Tzmb2,..., Tzmb n}.
[0014] Preferably, the expression of the volume Yl algorithm is as follows:
[0015]
[0016] In the formula, Yl represents the volume, Xh i represents the audio signal value of the i-th sample, and N represents the total number of signal samples.
[0017] Preferably, the expression of the speech rate Ys algorithm is as follows:
[0018]
[0019] In the formula, Ys represents the speech rate, Yjs represents the real-time syllable number, and Cxsj represents the duration.
[0020] Preferably, the formula of the autocorrelation function Zxhs(τ) is as follows:
[0021]
[0022] In the formula, Zxhs(τ) represents the autocorrelation value of the signal at a lag of τ, Xh i represents the audio signal value of the i-th sample, Xh (i+τ) represents the audio signal value of the (i + τ)-th sample, N represents the total number of signal samples, τ represents the delay, which is the independent variable of the autocorrelation function and represents the time offset of the signal. Find the peak according to the autocorrelation function and denote it as Zxhs max , and the lag value τ corresponding to the peak Zxhs max is the fundamental period T.
[0023] Preferably, the expression of the pitch Yd algorithm is as follows:
[0024]
[0025] In the formula, Yd represents the pitch and T represents the fundamental period.
[0026] Preferably, the expression of the volume matching degree Ylpp algorithm is as follows:
[0027]
[0028] In the formula, Ylpp represents the volume matching degree, Yl represents the volume, Ylck represents the volume reference value, and Ylcy max represents the maximum value of the volume change;
[0029] The expression of the speech rate matching degree Yspp algorithm is as follows:
[0030]
[0031] In the formula, Ylpp represents the speech rate matching degree, Ys represents the speech rate, Ysck represents the speech rate reference value, and Ysbh max represents the maximum value of the speech rate change;
[0032] The expression of the pitch matching degree Ydpp algorithm is as follows:
[0033]
[0034] In the formula, Ydpp represents the pitch matching degree, Yd represents the pitch, Ydck represents the pitch reference value, and Ydbh max represents the maximum value of the pitch change.
[0035] Preferably, the expression of the matching degree score Ppdf algorithm is as follows:
[0036]
[0037] In the formula, Ppdf represents the matching degree score, Tzxl i represents the i-th feature vector, Tzmb i represents the i-th target feature vector, where i = 1 means starting from Tzxl1 * Tzmb1, n means calculating up to the nth data, that is, calculating up to Tzxl n * Tzmb n up to, n means there are n data in this dataset, and the calculated result is the dot product of the feature vector and the target feature vector, means adding up the squares of all feature vectors, and the calculated result is the square root of the sum, means adding up the squares of all target feature vectors, and the calculated result is the square root of the sum.
[0038] Preferably, the data analysis module compares the volume matching degree Ylpp, the speech rate matching degree Yspp, the pitch matching degree Ydpp, and the matching degree score Ppdf with the volume matching degree threshold Ylpp yz , the speech rate matching degree threshold Yspp yz , the pitch matching degree threshold Ydpp yz , and the matching degree score threshold Ppdf yzFor comparison, the data analysis module sends the comparison result to the authentication module via the network. When any two of the volume matching degree Ylpp, speech rate matching degree Yspp, and pitch matching degree Ydpp are greater than their corresponding thresholds and the matching score Ppdf is greater than the matching score threshold Ppdf yz the authentication module determines that the user authentication is successful;
[0039] When any two of the volume matching degree Ylpp, speech rate matching degree Yspp, and pitch matching degree Ydpp are less than their corresponding thresholds or the matching score Ppdf is less than the matching score threshold Ppdf yz the authentication module determines that the user authentication is unsuccessful.
[0040] Preferably, the authentication module sends the authentication result to the control module via the network. When the user authentication is successful, the user can issue corresponding voice commands to the TV through the control module. When the user authentication is unsuccessful, the user cannot issue corresponding voice commands to the TV through the control module.
[0041] Compared with the prior art, the present invention provides a user interaction method based on a smart TV, having the following beneficial effects:
[0042] The present invention calculates the volume matching degree, speech rate matching degree, pitch matching degree, and matching score, and then compares the volume matching degree, speech rate matching degree, pitch matching degree, and matching score with their corresponding thresholds respectively. When any two of the volume matching degree, speech rate matching degree, and pitch matching degree are greater than their corresponding thresholds and the matching score is greater than the matching score threshold, the authentication module determines that the user authentication is successful. When any two of the volume matching degree, speech rate matching degree, and pitch matching degree are less than their corresponding thresholds or the matching score is less than the matching score threshold, the authentication module determines that the user authentication is unsuccessful. Thus, voice recognition authentication can be performed on a specific user, avoiding the chaos of TV control caused by multiple people's voice control of the TV. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic diagram of the method steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Please refer to Figure 1, A user interaction method based on a smart TV, comprising the following steps:
[0046] S1. The data acquisition module collects the user's voice through a built-in microphone. The data acquisition module extracts the real-time audio signal Xh and the real-time syllable count Yjs from the collected user voice, and the data acquisition module sends the collected data to the data processing module through the network;
[0047] S2. The data processing module calculates the volume Yl based on the real-time audio signal Xh. The algorithm expression of the volume Yl is as follows:
[0048]
[0049] In the formula, Yl represents the volume, Xh i represents the audio signal value of the i-th sample, N represents the total number of samples of the signal. The volume Yl is also called loudness or intensity, which is a basic attribute of sound and corresponds to the auditory perception of the loudness of sound by people. The data processing module calculates the volume matching degree Ylpp according to the volume Yl. The algorithm expression of the volume matching degree Ylpp is as follows:
[0050]
[0051] In the formula, Ylpp represents the volume matching degree, Yl represents the volume, Ylck represents the volume reference value, and Ylcy max represents the maximum value of the volume change. Here, the volume reference value Ylck and Ylcy max the maximum value of the volume change are all known numbers;
[0052] The data processing module calculates the speech rate Ys according to the real-time syllable count Yjs. The algorithm expression of the speech rate Ys is as follows:
[0053]
[0054] In the formula, Ys represents the speech rate, Yjs represents the real-time syllable count, and Cxsj represents the duration. The speech rate Ys, that is, the speaking speed, is a factor that cannot be ignored. It not only affects feature extraction and frame synchronization analysis, but also has an important impact on the accuracy of acoustic models and language models. The data processing module calculates the speech rate matching degree Yspp according to the speech rate Ys. The algorithm expression of the speech rate matching degree Yspp is as follows:
[0055]
[0056] In the formula, Ylpp represents the speech rate matching degree, Ys represents the speech rate, Ysck represents the speech rate reference value, and Ysbh max represents the maximum value of the speech rate change. Here, the speech rate reference value Ysck and Ysbh maxThe maximum values of the speech rate changes are all known numbers;
[0057] The data processing module is preset with an autocorrelation function Zxhs. The formula for the autocorrelation function Zxhs(τ) is as follows:
[0058]
[0059] In the formula, Zxhs(τ) represents the autocorrelation value of the signal at a lag of τ, and Xh i represents the audio signal value of the i-th sample, and Xh (i+τ) represents the audio signal value of the (i + τ)-th sample. N represents the total number of samples of the signal, τ represents the delay, which is the independent variable of the autocorrelation function and represents the time offset of the signal. Find the peak according to the autocorrelation function and denote it as Zxhs max The peak Zxhs max The corresponding lag value τ is the fundamental period T. The data processing module calculates the pitch Yd based on the fundamental period T and the real-time audio signal Xh. The algorithm expression for the pitch Yd is as follows:
[0060]
[0061] In the formula, Yd represents the pitch, T represents the fundamental period. The pitch Yd is also called the pitch height and is mainly related to the frequency of the sound wave. The level of the pitch Yd directly determines the auditory perception of the sound. Sound waves with a high frequency sound like a higher pitch, while those with a low frequency sound like a lower pitch. The data processing module calculates the pitch matching degree Ydpp based on the pitch Yd. The algorithm expression for the pitch matching degree Ydpp is as follows:
[0062]
[0063] In the formula, Ydpp represents the pitch matching degree, Yd represents the pitch, Ydck represents the pitch reference value, and Ydbh max represents the maximum value of the pitch change. Here, the pitch reference value Ydck and the maximum value of the pitch change Ydbh max are all known numbers;
[0064] The data processing module is preset with a target feature vector Tzmb. The data processing module calculates the matching degree score Ppdf based on the feature vector Tzxl and the target feature vector Tzmb. Among them, the expressions for the feature vector Tzxl and the target feature vector Tzmb are {Tzxl1, Tzxl2,..., Tzxl n} and {Tzmb1, Tzmb2,..., Tzmb n}, where the feature vector Tzxl and the target feature vector Tzmb are both numerical arrays of a fixed length, used to represent the features of a certain aspect of the data. Each element usually represents a dimension of the feature. The algorithm expression for the matching degree score Ppdf is as follows:
[0065]
[0066] In the formula, Ppdf represents the matching degree score, Tzxl i represents the i-th feature vector, Tzmb i represents the i-th target feature vector, where i = 1 in [[ ]] means starting from Tzxl1 * Tzmb1 for calculation, and n means calculating up to the n-th data, that is, calculating up to Tzxl n * Tzmb n up to, and n means there are n data in this dataset. The calculated result is the dot product of the feature vector and the target feature vector, means adding up the squares of all feature vectors and then taking the square root of the calculated sum, means adding up the squares of all target feature vectors and then taking the square root of the calculated sum. In user recognition, the matching degree score Ppdf is an index to evaluate the similarity between the input speech and the target speaker model. The matching degree score Ppdf is a key index to evaluate the similarity between the input speech and the known template. Its main role is to judge and identify the pronunciation content by quantifying the similarity between the input speech and the reference template;
[0067] In audio analysis and processing tasks, pitch Yd, volume Yl, and speech rate Ys are important features. In a speech recognition system, the matching degrees of pitch Yd, volume Yl, and speech rate Ys have an important impact on the recognition accuracy. Therefore, it is necessary to use pitch Yd, volume Yl, and speech rate Ys to complete the authentication of users;
[0068] The data processing module sends this data to the data analysis module through the network;
[0069] S3. There are preset volume matching degree thresholds Ylpp yz and speech rate matching degree thresholds Yspp yz and pitch matching degree thresholds Ydpp yz and matching degree score thresholds Ppdf yz in the data analysis module. The data analysis module compares the volume matching degree Ylpp, speech rate matching degree Yspp, pitch matching degree Ydpp, and matching degree score Ppdf with the volume matching degree threshold Ylpp yz and speech rate matching degree threshold Yspp yz and pitch matching degree threshold Ydpp yzand the matching score threshold Ppdf yz For comparison, the data analysis module sends the comparison result to the authentication module through the network;
[0070] S4. When any two of the volume matching degree Ylpp, the speech rate matching degree Yspp, and the pitch matching degree Ydpp are greater than the corresponding thresholds and the matching score Ppdf is greater than the matching score threshold Ppdf yz the authentication module determines that the user authentication is successful. When any two of the volume matching degree Ylpp, the speech rate matching degree Yspp, and the pitch matching degree Ydpp are less than the corresponding thresholds or the matching score Ppdf is less than the matching score threshold Ppdf yz the authentication module determines that the user authentication is unsuccessful;
[0071] S5. The authentication module sends the authentication result to the control module through the network. When the user authentication is successful, the user can perform corresponding voice commands on the TV through the control module. When the user authentication is unsuccessful, the user cannot perform corresponding voice commands on the TV through the control module.
[0072] This system calculates the volume matching degree, the speech rate matching degree, the pitch matching degree, and the matching score, and then compares the volume matching degree, the speech rate matching degree, the pitch matching degree, and the matching score with the corresponding thresholds respectively. When any two of the volume matching degree, the speech rate matching degree, and the pitch matching degree are greater than the corresponding thresholds and the matching score is greater than the matching score threshold, the authentication module determines that the user authentication is successful. When any two of the volume matching degree, the speech rate matching degree, and the pitch matching degree are less than the corresponding thresholds or the matching score is less than the matching score threshold, the authentication module determines that the user authentication is unsuccessful. Thus, voice recognition authentication for specific users can be performed to avoid the chaos of TV control caused by multiple people's voice control of the TV.
[0073] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A user interaction method based on a smart TV, characterized in that: Including the following steps: S1. The data acquisition module collects the user's voice through the built-in microphone. From the collected user voice, the data acquisition module extracts the real-time audio signal Xh and the real-time syllable count Yjs, and the data acquisition module sends the collected data to the data processing module through the network; S2. The data processing module calculates the volume Yl according to the real-time audio signal Xh, calculates the speech rate Ys according to the real-time syllable count Yjs. There is a preset autocorrelation function Zxhs in the data processing module. The data processing module calculates the pitch Yd according to the real-time audio signal Xh and the autocorrelation function Zxhs. The data processing module preprocesses the real-time audio signal Xh to obtain the feature vector Tzxl. The data processing module calculates the volume matching degree Ylpp, the speech rate matching degree Yspp and the pitch matching degree Ydpp according to the volume Yl, the speech rate Ys and the pitch Yd respectively. There is a preset target feature vector Tzmb in the data processing module. The data processing module calculates the matching degree score Ppdf according to the feature vector Tzxl and the target feature vector Tzmb. The data processing module sends these data to the data analysis module through the network; S3. There is a preset volume matching degree threshold Ylpp in the data analysis module yz , speech rate matching degree threshold Yspp yz , pitch matching degree threshold Ydpp yz and matching degree score threshold Ppdf yz . The data analysis module compares the volume matching degree Ylpp, speech rate matching degree Yspp, pitch matching degree Ydpp and matching degree score Ppdf with the volume matching degree threshold Ylpp yz , speech rate matching degree threshold Yspp yz , pitch matching degree threshold Ydpp yz and matching degree score threshold Ppdf yz respectively, and sends the comparison results to the authentication module through the network; S4. When any two of the volume matching degree Ylpp, the speech rate matching degree Yspp, and the pitch matching degree Ydpp are greater than the corresponding thresholds and the matching score Ppdf is greater than the matching score threshold Ppdf yz the authentication module determines that the user authentication is successful. When any two of the volume matching degree Ylpp, the speech rate matching degree Yspp, and the pitch matching degree Ydpp are less than the corresponding thresholds or the matching score Ppdf is less than the matching score threshold Ppdf yz the authentication module determines that the user authentication is unsuccessful; S5. The authentication module sends the authentication result to the control module through the network. When the user authentication is successful, the user can issue corresponding voice commands to the TV through the control module. When the user authentication is unsuccessful, the user cannot issue corresponding voice commands to the TV through the control module.
2. The user interaction method based on a smart TV according to claim 1, wherein: The expression of the feature vector Tzxl is {Tzxl1, Tzxl2,..., Tzxl n}, and the expression of the target feature vector Tzmb is {Tzmb1, Tzmb2,..., Tzmb n}.
3. The user interaction method based on a smart TV according to claim 2, wherein: The algorithm expression of the volume Yl is as follows: In the formula, Yl represents the volume, and Xh i represents the audio signal value of the i-th sample, and N represents the total number of samples of the signal.
4. The user interaction method based on a smart TV according to claim 3, wherein: The algorithm expression of the speech rate Ys is as follows: In the formula, Ys represents the speech rate, Yjs represents the real-time syllable count, and Cxsj represents the duration.
5. The user interaction method based on a smart TV according to claim 4, characterized in that: The formula of the autocorrelation function Zxhs(τ) is as follows: In the formula, Zxhs(τ) represents the autocorrelation value of the signal at a lag of τ, and Xh i represents the audio signal value of the i-th sample, and Xh (i+τ) represents the audio signal value of the (i + τ)-th sample. N represents the total number of samples of the signal, τ represents the delay, which is the independent variable of the autocorrelation function and represents the time offset of the signal. Find the peak according to the autocorrelation function and denote it as Zxhs max , and the peak Zxhs max The corresponding lag value τ is the fundamental period T.
6. The user interaction method based on a smart TV according to claim 5, wherein: The algorithm expression of the pitch Yd is as follows: In the formula, Yd represents the pitch and T represents the fundamental period.
7. A user interaction method based on a smart TV according to claim 6, characterized in that: The algorithm expression of the volume matching degree Ylpp is as follows: In the formula, Ylpp represents the volume matching degree, Yl represents the volume, Ylck represents the volume reference value, and Ylbh max represents the maximum value of the volume change; The algorithm expression of the speech rate matching degree Yspp is as follows: In the formula, Yspp represents the speech rate matching degree, Ys represents the speech rate, Ysck represents the speech rate reference value, and Ysbh max represents the maximum value of the speech rate change; The algorithm expression of the pitch matching degree Ydpp is as follows: In the formula, Ydpp represents the pitch matching degree, Yd represents the pitch, Ydck represents the pitch reference value, and Ydbh max represents the maximum value of pitch change.
8. The user interaction method based on a smart TV according to claim 7, wherein: The algorithm expression of the matching degree score Ppdf is as follows: In the formula, Ppdf represents the matching degree score, and Tzxl i represents the i-th feature vector, and Tzmb i represents the i-th target feature vector. where i = 1 indicates starting from Tzxl1 * Tzmb1 for calculation, and n indicates calculating up to the n-th data, that is, calculating up to Tzxl n * Tzmb n up to. n indicates that there are n data in this dataset, and the calculated result is the dot product of the feature vector and the target feature vector. represents the accumulation and summation of the squares of all feature vectors, and the calculated result is the square root of the accumulated sum. represents the accumulation and summation of the squares of all target feature vectors, and the calculated result is the square root of the accumulated sum.
Citation Information
Patent Citations
Television voice remote control system and television voice remote control method
CN106358061A
Authentic identification method based on language processing
CN116863961A