Operation control methods, devices, storage media and terminals
By determining the base frequency of the voice signal in the terminal and determining the target screen range based on the base frequency, the problem of low voice recognition accuracy is solved, resulting in higher operation control accuracy and a better user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING XIAOMI MOBILE SOFTWARE CO LTD
- Filing Date
- 2022-03-17
- Publication Date
- 2026-05-26
Smart Images

Figure CN114464183B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of terminal technology, and in particular to an operation control method, device, storage medium and terminal. Background Technology
[0002] With the widespread adoption of smartphones, tablets, and other mobile devices, applications for these devices have also developed significantly. Users can install various applications, especially games. Currently, most mobile games are controlled via keyboard, mouse, gamepad, or touchscreen, requiring manual operation from the user.
[0003] In related technologies, games can be controlled by voice, for example, by recognizing specific characters (up, down, left, right, etc.) in the voice to control the movement of objects in the game. However, for some users who speak dialects or have non-standard Mandarin, the accuracy of voice recognition is relatively low, resulting in low accuracy of operation control and thus affecting the user's operating experience. Summary of the Invention
[0004] To overcome the problems existing in related technologies, this disclosure provides an operation control method, apparatus, storage medium, and terminal.
[0005] According to a first aspect of the present disclosure, an operation control method is provided, which is applied to a terminal, the method comprising:
[0006] In response to the user's input voice signal, determine the base frequency corresponding to the voice signal;
[0007] Based on the base frequency, the target screen range corresponding to the voice signal is determined from the screen of the terminal;
[0008] Based on the target screen range, control the terminal to perform a preset operation.
[0009] Optionally, determining the fundamental frequency corresponding to the voice signal includes:
[0010] The speech signal is segmented into frames to obtain multiple speech frames corresponding to the speech signal;
[0011] For each of the speech frames, determine the frequency domain energy and frequency corresponding to the speech frame;
[0012] The fundamental frequency corresponding to the speech signal is determined based on the multiple frequency domain energies and the multiple frequencies.
[0013] Optionally, determining the frequency domain energy corresponding to the speech frame includes:
[0014] The speech frame is pre-emphasized to obtain the first speech frame corresponding to the speech frame;
[0015] Filter out the DC component in the first speech frame to obtain the second speech frame corresponding to the speech frame;
[0016] Windowing is applied to the second speech frame to obtain the third speech frame corresponding to the speech frame;
[0017] Perform a Fast Fourier Transform on the third speech frame to obtain the fourth speech frame corresponding to the speech frame;
[0018] Determine the frequency domain energy corresponding to the fourth speech frame.
[0019] Optionally, determining the fundamental frequency corresponding to the speech signal based on the plurality of frequency domain energies and the plurality of frequencies includes:
[0020] The amplitude spectrum corresponding to the speech signal is determined based on the multiple frequency domain energies and the multiple frequencies;
[0021] Obtain the preset frequency range;
[0022] The fundamental frequency corresponding to the speech signal is determined based on the preset frequency range and the amplitude spectrum.
[0023] Optionally, determining the target screen range corresponding to the voice signal from the screen of the terminal based on the baseband includes:
[0024] Obtain the preset maximum baseband, preset minimum baseband, and preset number of operating intervals;
[0025] Based on the base frequency, the preset maximum base frequency, the preset minimum base frequency, and the preset number of operating intervals, the target screen interval corresponding to the voice signal is determined from the screen of the terminal.
[0026] Optionally, determining the target screen interval corresponding to the voice signal from the terminal's screen based on the baseband, the preset maximum baseband, the preset minimum baseband, and the preset number of operating intervals includes:
[0027] Based on the preset maximum base frequency and the preset minimum base frequency, the screen is divided into the preset number of preset screen intervals.
[0028] Determine the base frequency range corresponding to each preset screen interval to obtain the operation interval correlation relationship;
[0029] Based on the base frequency, the target screen range corresponding to the voice signal is determined through the operation range correlation.
[0030] Optionally, dividing the screen into the preset number of preset screen intervals based on the preset maximum base frequency and the preset minimum base frequency includes:
[0031] Determine the maximum Mel-frequency corresponding to the preset maximum fundamental frequency;
[0032] Determine the minimum Mel-frequency corresponding to the preset minimum fundamental frequency;
[0033] Based on the maximum Mel frequency and the minimum Mel frequency, the screen is divided into a preset number of preset screen intervals.
[0034] Optionally, determining the baseband range corresponding to each of the preset screen intervals includes:
[0035] For each preset screen interval, the Mel fundamental frequency range corresponding to the preset screen interval is determined based on the maximum Mel fundamental frequency and the minimum Mel fundamental frequency, the frequency range corresponding to the Mel fundamental frequency range is determined, and the frequency range is used as the fundamental frequency range corresponding to the preset screen interval.
[0036] Optionally, the voice signals include multiple signals, with different voice signals corresponding to different users; determining the target screen range corresponding to the voice signal from the terminal's screen based on the baseband includes:
[0037] Determine the largest fundamental frequency among the fundamental frequencies corresponding to the multiple speech signals;
[0038] Based on the maximum base frequency, the target screen range corresponding to the voice signal is determined from the screen of the terminal.
[0039] According to a second aspect of the present disclosure, an operation control device is provided for a terminal, the device comprising:
[0040] The baseband determination module is configured to determine the baseband corresponding to the voice signal in response to the voice signal input by the user.
[0041] The target screen range determination module is configured to determine the target screen range corresponding to the voice signal from the screen of the terminal based on the base frequency;
[0042] The control module is configured to control the terminal to perform a preset operation based on the target screen range.
[0043] Optionally, the baseband determination module is further configured to:
[0044] The speech signal is segmented into frames to obtain multiple speech frames corresponding to the speech signal;
[0045] For each of the speech frames, determine the frequency domain energy and frequency corresponding to the speech frame;
[0046] The fundamental frequency corresponding to the speech signal is determined based on the multiple frequency domain energies and the multiple frequencies.
[0047] Optionally, the baseband determination module is further configured to:
[0048] The speech frame is pre-emphasized to obtain the first speech frame corresponding to the speech frame;
[0049] Filter out the DC component in the first speech frame to obtain the second speech frame corresponding to the speech frame;
[0050] Windowing is applied to the second speech frame to obtain the third speech frame corresponding to the speech frame;
[0051] Perform a Fast Fourier Transform on the third speech frame to obtain the fourth speech frame corresponding to the speech frame;
[0052] Determine the frequency domain energy corresponding to the fourth speech frame.
[0053] Optionally, the baseband determination module is further configured to:
[0054] The amplitude spectrum corresponding to the speech signal is determined based on the multiple frequency domain energies and the multiple frequencies;
[0055] Obtain the preset frequency range;
[0056] The fundamental frequency corresponding to the speech signal is determined based on the preset frequency range and the amplitude spectrum.
[0057] Optionally, the target screen range determination module is further configured to:
[0058] Obtain the preset maximum baseband, preset minimum baseband, and preset number of operating intervals;
[0059] Based on the base frequency, the preset maximum base frequency, the preset minimum base frequency, and the preset number of operating intervals, the target screen interval corresponding to the voice signal is determined from the screen of the terminal.
[0060] Optionally, the target screen range determination module is further configured to:
[0061] Based on the preset maximum base frequency and the preset minimum base frequency, the screen is divided into the preset number of preset screen intervals.
[0062] Determine the base frequency range corresponding to each preset screen interval to obtain the operation interval correlation relationship;
[0063] Based on the base frequency, the target screen range corresponding to the voice signal is determined through the operation range correlation.
[0064] Optionally, the target screen range determination module is further configured to:
[0065] Determine the maximum Mel-frequency corresponding to the preset maximum fundamental frequency;
[0066] Determine the minimum Mel-frequency corresponding to the preset minimum fundamental frequency;
[0067] Based on the maximum Mel frequency and the minimum Mel frequency, the screen is divided into a preset number of preset screen intervals.
[0068] Optionally, the target screen range determination module is further configured to:
[0069] For each preset screen interval, the Mel fundamental frequency range corresponding to the preset screen interval is determined based on the maximum Mel fundamental frequency and the minimum Mel fundamental frequency, the frequency range corresponding to the Mel fundamental frequency range is determined, and the frequency range is used as the fundamental frequency range corresponding to the preset screen interval.
[0070] Optionally, the voice signals include multiple signals, with different voice signals corresponding to different users; the target screen range determination module is further configured to:
[0071] Determine the largest fundamental frequency among the fundamental frequencies corresponding to the multiple speech signals;
[0072] Based on the maximum base frequency, the target screen range corresponding to the voice signal is determined from the screen of the terminal.
[0073] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the steps of the operation control method provided in the first aspect of the present disclosure.
[0074] According to a fourth aspect of the present disclosure, a terminal is provided, comprising:
[0075] A memory on which computer programs are stored;
[0076] A processor is configured to execute the computer program in the memory to implement the steps of the operation control method provided in the first aspect of this disclosure.
[0077] The technical solutions provided by the embodiments of this disclosure can include the following beneficial effects: determining the base frequency corresponding to the voice signal in response to user input; determining the target screen interval corresponding to the voice signal from the screen of the terminal based on the base frequency; and controlling the terminal to perform a preset operation based on the target screen interval. In other words, this disclosure can determine the target screen interval based on the base frequency corresponding to the voice signal, and control the terminal to perform a preset operation based on the target screen interval. Since the base frequency is not affected by the user's own pronunciation, the preset operation performed based on the base frequency is more accurate, thereby improving the accuracy of operation control and enhancing the user's operating experience.
[0078] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0079] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0080] Figure 1 This is a flowchart illustrating an operation control method according to an exemplary embodiment of the present disclosure;
[0081] Figure 2 This is a flowchart illustrating another operation control method according to an exemplary embodiment of the present disclosure;
[0082] Figure 3 This is a block diagram illustrating an operation control device according to an exemplary embodiment of the present disclosure;
[0083] Figure 4 This is a block diagram illustrating a terminal according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0084] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0085] It should be noted that all actions involving the acquisition of signals, information, or data in this application are carried out in compliance with the relevant data protection laws and policies of the country where the application is located, and with the authorization granted by the owner of the relevant device.
[0086] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0087] Before detailing the specific embodiments of this disclosure, the application scenarios of this disclosure will first be explained. Currently, in the field of voice games, there are many ways to control game operations through volume or voice recognition. However, the accuracy of volume-based game control is relatively low because volume is easily affected by external environmental interference. Similarly, the accuracy of voice recognition-based game control is easily affected by the user's accent, resulting in low overall accuracy.
[0088] To overcome the technical problems existing in the above-mentioned related technologies, this disclosure provides an operation control method, device, storage medium and terminal. The target screen range is determined according to the base frequency corresponding to the voice signal, and the terminal is controlled to perform preset operations according to the target screen range. In this way, since the base frequency is not affected by the user's own voice, the preset operations performed according to the base frequency are more accurate, thereby improving the accuracy of operation control and enhancing the user's operating experience.
[0089] The present disclosure will now be described in conjunction with specific embodiments.
[0090] Figure 1 This is a flowchart illustrating an operation control method according to an exemplary embodiment of the present disclosure, such as... Figure 1 As shown, this method is applied to a terminal, which may include a mobile device, such as a smartphone, smart wearable device, smart speaker, smart tablet, personal computer, etc., and the method may include:
[0091] S101. In response to the voice signal input by the user, determine the base frequency corresponding to the voice signal.
[0092] In this step, the user's voice signal can be received through the terminal's microphone. After receiving the voice signal, the voice signal can be segmented into frames to obtain multiple voice frames corresponding to the voice signal. For each voice frame, the frequency domain energy and frequency corresponding to the voice frame are determined. Based on the multiple frequency domain energies and multiple frequencies, the fundamental frequency corresponding to the voice signal is determined.
[0093] It should be noted that the above method for determining the base frequency of the speech signal is only an example. The base frequency of the speech signal can also be determined by other methods in the prior art. This disclosure does not limit this method.
[0094] S102. Based on the base frequency, determine the target screen range corresponding to the voice signal from the screen of the terminal.
[0095] In this step, after determining the base frequency corresponding to the voice signal, the target screen interval corresponding to the voice signal can be determined based on the base frequency and a pre-created operation interval association relationship. This operation interval association relationship can include the correspondence between different base frequency ranges and preset screen intervals. For example, if the terminal's screen is evenly divided into three preset screen intervals: a preset upper screen interval, a preset middle screen interval, and a preset lower screen interval, with a maximum base frequency of 400 Hz and a minimum base frequency of 100 Hz, then the base frequency range corresponding to the preset upper screen interval can be determined to be 300 Hz to 400 Hz, the base frequency range corresponding to the preset middle screen interval to be 200 Hz to 300 Hz, and the base frequency range corresponding to the preset lower screen interval to be 100 Hz to 200 Hz. If the base frequency corresponding to the voice signal is determined to be 230 Hz, then the target screen interval corresponding to the voice signal can be determined to be the preset middle screen interval.
[0096] The terminal may receive multiple voice signals simultaneously, with different voice signals corresponding to different users. In this case, the maximum base frequency among the multiple voice signals can be determined. Based on this maximum base frequency, the target screen range corresponding to the voice signal can be determined from the terminal's screen. Taking a game scenario as an example, during a multiplayer game, the terminal may receive multiple voice signals simultaneously. For instance, if there are three voice signals, with the first voice signal corresponding to a base frequency of 120 Hz, the second voice signal corresponding to a base frequency of 180 Hz, and the third voice signal corresponding to a base frequency of 300 Hz, then the base frequency corresponding to the third voice signal can be determined as the maximum base frequency. Based on the base frequency corresponding to the third voice signal, the target screen range corresponding to the voice signal can be determined.
[0097] S103. Based on the target screen range, control the terminal to perform a preset operation.
[0098] The preset operation can be determined according to the application scenario of the terminal. Different application scenarios can correspond to different preset operations. For example, if the application scenario is a game scenario, the preset operation can be to control the object in the game to move to the target screen area; if the application scenario is a text display scenario, the preset operation can be to select the text content of the target screen area.
[0099] In this step, after determining the target screen range corresponding to the voice signal, the current application scenario of the terminal can be obtained. For example, the application currently running on the terminal can be determined, and the current application scenario of the terminal can be determined based on the application. Then, through a pre-created preset operation association, the preset operation corresponding to the application scenario is determined. This preset operation association can include the correspondence between different application scenarios and preset operations.
[0100] Using the above method, the target screen range can be determined based on the base frequency corresponding to the voice signal, and the terminal can be controlled to perform preset operations based on the target screen range. In this way, since the base frequency is not affected by the user's own pronunciation, the preset operations performed based on the base frequency are more accurate, thereby improving the accuracy of operation control and enhancing the user's operating experience.
[0101] Figure 2 This is a flowchart illustrating another operation control method according to an exemplary embodiment of the present disclosure, such as... Figure 2 As shown, the method may include:
[0102] S201. In response to the user-input voice signal, perform frame segmentation on the voice signal to obtain multiple voice frames corresponding to the voice signal.
[0103] In this step, after receiving the user's input voice signal, the voice signal can be segmented into frames using existing techniques to facilitate analysis. To avoid excessive variation between adjacent frames, an overlapping region can exist between two adjacent voice frames. The number of sampling points in this overlapping region can be 1 / 3 to 2 / 3 of the total number of sampling points in the voice frame. For example, the number of sampling points in the overlapping region can be 3 / 5 of the total number of sampling points in the voice frame. If the total number of sampling points in the voice frame is 400, then the number of sampling points in the overlapping region can be 240. Typically, the sampling frequency of the voice signal is 8kHz or 16kHz. Taking a sampling frequency of 16kHz as an example, if the number of sampling points in the voice frame is 400, then the corresponding time length of the voice frame can be determined to be 25ms.
[0104] S202. For each speech frame, determine the frequency domain energy and frequency corresponding to that speech frame.
[0105] In this step, after obtaining multiple speech frames corresponding to the speech signal, for each speech frame, pre-emphasis processing can be performed to obtain the first speech frame corresponding to the speech frame; the DC component in the first speech frame is filtered out to obtain the second speech frame corresponding to the speech frame; windowing processing is performed on the second speech frame to obtain the third speech frame corresponding to the speech frame; fast Fourier transform is performed on the third speech frame to obtain the fourth speech frame corresponding to the speech frame; and the frequency domain energy corresponding to the fourth speech frame is determined.
[0106] For example, the speech frame can first be passed through a high-pass filter, the system function of which can be:
[0107] H(z) = 1 - a*z -1 , 0.9 < a < 1.0 (1)
[0108] Where H(z) is the speech frame after high-pass filtering, z is the speech frame, and a is 0.97.
[0109] After the speech frame is filtered by the high-pass filter, the filtered speech frame is obtained. Then, the filtered speech frame can be pre-emphasized using formula (2) to obtain the first speech frame corresponding to the speech frame:
[0110] S(n)=p(n)-a*p(n-1) (2)
[0111] Wherein, S(n) is the first speech frame, p(n) is the filtered speech frame corresponding to the speech frame, p(n-1) is the filtered speech frame corresponding to the speech frame adjacent to and preceding the speech frame, and n represents the sequence number of the speech frame.
[0112] The pre-emphasis processing described above enhances the high-frequency components of the speech frame, resulting in a flatter spectrum. This allows for consistent signal-to-noise ratio (SNR) across the entire frequency band from low to high frequencies. Furthermore, it eliminates the effects of the vocal cords and lips during vocalization, compensating for the suppression of high-frequency components by the vocal system. This highlights high-frequency formants and improves the accuracy of determining the fundamental frequency corresponding to the speech frame.
[0113] After obtaining the first speech frame corresponding to the speech frame, the DC component in the first speech frame can be filtered out using formula (3) to obtain the second speech frame corresponding to the speech frame:
[0114] V(n)=S(n)-S(n-1)+k1*V(n-1) (3)
[0115] Wherein, V(n) is the second speech frame, S(n) is the first speech frame corresponding to the speech frame, S(n-1) is the first speech frame corresponding to the speech frame adjacent to and preceding the speech frame, V(n-1) is the second speech frame corresponding to the speech frame adjacent to and preceding the speech frame, and k1 is 0.997.
[0116] The second speech frame filters out the governance components in the first speech frame, which can avoid the impact on the low spectrum and improve the accuracy of determining the fundamental frequency corresponding to the speech frame.
[0117] After obtaining the second speech frame corresponding to the first speech frame, the second speech frame can be windowed using formula (4) to obtain the third speech frame corresponding to the first speech frame:
[0118] V′(n)=V(n)*W(n) (4)
[0119] Where V′(n) is the third speech frame, V(n) is the second speech frame, and W(n) is the window function.
[0120] For example, the window function can be a Hamming window, and the expression for the Hamming window can be:
[0121]
[0122] Where W(n) is the window function of the Hamming window, t is 0.46, and N is the number of speech frames contained in the speech signal.
[0123] By windowing each second speech frame, the signal discontinuity that may be caused at the two ends of each frame can be eliminated, thus improving the accuracy of the speech frame.
[0124] After obtaining the third speech frame corresponding to the speech frame, the fourth speech frame corresponding to the speech frame can be obtained by performing a fast Fourier transform on the third speech frame using formula (6):
[0125]
[0126] Where X(k, l) represents the spectral amplitude value of the k-th frequency band in the l-th frame, Q is the change length of the fast Fourier transform, j is a complex number, x(d, l) represents the value of the d-th sampling point in the l-th frame, and Q is 512.
[0127] Furthermore, after obtaining the fourth speech frame corresponding to the speech frame, the frequency domain energy corresponding to the fourth speech frame can be calculated using formula (7):
[0128]
[0129] Where E(l) is the frequency domain energy corresponding to the fourth speech frame, xc,l,r Let x represent the real part of the c-th frequency band in the l-th frame. c,l,i R represents the imaginary part of the c-th frequency band in the l-th frame, where R is 50% of Q.
[0130] Furthermore, the frequency corresponding to this voice frame can be determined using existing technologies, which will not be elaborated here.
[0131] S203. Determine the fundamental frequency corresponding to the speech signal based on multiple frequency domain energies and multiple frequencies.
[0132] In this step, after determining the frequency domain energy and frequency corresponding to each speech frame, the amplitude spectrum corresponding to the speech signal can be determined based on multiple frequency domain energies and multiple frequencies to obtain a preset frequency range. Then, based on the preset frequency range and the amplitude spectrum, the fundamental frequency corresponding to the speech signal is determined. The preset frequency range can be the frequency range of human voice, for example, 100Hz to 500Hz. For example, for each frequency, the frequency domain energy corresponding to that frequency is determined to obtain the amplitude spectrum.
[0133] Furthermore, after determining the amplitude spectrum and preset frequency range corresponding to the speech signal, the fundamental frequency corresponding to the speech signal can be calculated using formula (8):
[0134]
[0135] Where F(f) is the fundamental frequency corresponding to the speech signal, h is the preset frequency range, X(gf) is the amplitude spectrum, f is the horizontal axis (frequency) of the curve corresponding to the amplitude spectrum, and g is the vertical axis (signal amplitude) of the curve corresponding to the amplitude spectrum.
[0136] S204. Obtain the preset maximum base frequency, preset minimum base frequency, and preset number of operating intervals.
[0137] The preset maximum fundamental frequency can be the maximum fundamental frequency of human voice under normal circumstances, and the preset minimum fundamental frequency can be the minimum fundamental frequency of human voice under normal circumstances. For example, the preset maximum fundamental frequency can be 400 Hz, and the preset minimum fundamental frequency can be 100 Hz. The number of preset operating intervals can be predetermined according to the application scenario. For example, the number of preset operating intervals can be 3.
[0138] S205. Based on the base frequency, the preset maximum base frequency, the preset minimum base frequency, and the preset number of operating intervals, determine the target screen interval corresponding to the voice signal from the screen of the terminal.
[0139] In one possible implementation, the terminal's screen can be divided evenly according to the preset number of operation intervals. For example, if the preset number of operation intervals is 3, the terminal's screen can be divided into 3 preset screen intervals. Based on this screen interval division method, after determining the baseband, the preset maximum baseband, the preset minimum baseband, and the preset number of operation intervals, the target screen interval corresponding to the voice signal can be determined using the following formula:
[0140]
[0141] Wherein, P(u) is the target screen range, u is the base frequency, MIN is the preset minimum base frequency, MAX is the preset maximum base frequency, and T is the preset number of operating ranges.
[0142] In another possible implementation, after determining the base frequency, the preset maximum base frequency, the preset minimum base frequency, and the preset number of operating intervals, the screen can be divided into the preset number of preset screen intervals based on the preset maximum base frequency and the preset minimum base frequency; the base frequency range corresponding to each preset screen interval is determined to obtain the operating interval correlation; and the target screen interval corresponding to the voice signal is determined based on the base frequency and the operating interval correlation.
[0143] First, the maximum Mel-frequency corresponding to the preset maximum base frequency can be determined, and the minimum Mel-frequency corresponding to the preset minimum base frequency can be determined. Then, based on the maximum and minimum Mel-frequencys, the screen is divided into a preset number of preset screen intervals for the preset operation interval. For example, the maximum and minimum Mel-frequencys can be calculated using the following formulas:
[0144] F mel (q)=1125ln(1+q / 700) (10)
[0145] Where, if F mel (q) is the maximum Mel frequency, then q is the preset maximum fundamental frequency, if F mel (q) is the minimum Mel frequency, then q is the preset minimum fundamental frequency.
[0146] After determining the maximum Mel-frequency corresponding to the preset maximum fundamental frequency and the minimum Mel-frequency corresponding to the preset minimum fundamental frequency, for each preset screen interval, the Mel-frequency range corresponding to the preset screen interval can be determined based on the maximum Mel-frequency and the minimum Mel-frequency, and the frequency range corresponding to the Mel-frequency range can be determined, and this frequency range can be used as the fundamental frequency range corresponding to the preset screen interval. For example, based on the preset number of operating intervals, the total Mel-frequency range corresponding to the maximum and minimum Mel-frequency can be divided into the preset number of Mel-frequency ranges. For instance, if the preset number of operating intervals is 3, the total Mel-frequency range can be divided into 3 Mel-frequency ranges. Then, for each preset screen interval, the Mel-frequency range corresponding to the preset screen interval is determined, and the actual frequency corresponding to the two boundary points of each Mel-frequency range is calculated using the following formula:
[0147]
[0148] Where f′ is the actual frequency corresponding to the boundary point of the Mel fundamental frequency range, and m is the Mel frequency corresponding to the boundary point of the Mel fundamental frequency range.
[0149] Finally, for each Mel fundamental frequency range, the frequency range corresponding to the Mel fundamental frequency range is determined based on the actual frequencies corresponding to the two boundary points of the Mel fundamental frequency range, and this frequency range is used as the fundamental frequency range corresponding to the preset screen interval.
[0150] Furthermore, after obtaining the baseband range corresponding to each preset screen interval, the operational interval correlation can be obtained. Then, based on the baseband, the target screen interval corresponding to the voice signal can be determined through the operational interval correlation. For example, the target baseband range to which the baseband belongs can be determined from multiple baseband ranges. For instance, if the baseband is greater than the lower boundary of the y-th baseband range and less than or equal to the lower boundary of the (y+1)-th baseband range, then the baseband belongs to the y-th baseband range. Then, through the operational interval correlation, the preset screen interval corresponding to the target baseband range is determined. The preset screen interval corresponding to the target baseband range is the target screen interval corresponding to the voice signal.
[0151] It should be noted that the correlation of the operation interval can also be determined in advance. After determining the base frequency corresponding to the voice signal, the correlation of the operation interval can be obtained directly. This can improve the efficiency of determining the target screen interval, thereby improving the efficiency of operation control.
[0152] S206. Based on the target screen range, control the terminal to perform a preset operation.
[0153] The preset operation can be determined according to the application scenario of the terminal. Different application scenarios can correspond to different preset operations. For example, if the application scenario is a game scenario, the preset operation can be to control the object in the game to move to the target screen area; if the application scenario is a text display scenario, the preset operation can be to select the text content of the target screen area.
[0154] Using the above method, the target screen interval can be determined based on the base frequency corresponding to the voice signal. Based on this target screen interval, the terminal can be controlled to execute preset operations. Since the base frequency is not affected by the user's own pronunciation, the preset operations executed based on this base frequency are more accurate, thereby improving the accuracy of operation control and enhancing the user experience. Furthermore, before determining the target screen interval based on the base frequency of the voice signal, the voice signal is pre-emphasized, DC components are filtered out, windowed, and subjected to a Fast Fourier Transform, making the determined base frequency of the voice signal even more accurate, thus further improving the accuracy of operation control.
[0155] Figure 3 This is a block diagram illustrating an operation control device according to an exemplary embodiment of the present disclosure, such as... Figure 3 As shown, the device is applied to a terminal and includes:
[0156] The baseband determination module 301 is configured to determine the baseband corresponding to the voice signal input by the user in response to the voice signal input by the user.
[0157] The target screen range determination module 302 is configured to determine the target screen range corresponding to the voice signal from the screen of the terminal based on the base frequency.
[0158] The control module 303 is configured to control the terminal to perform a preset operation based on the target screen range.
[0159] Optionally, the baseband determination module 301 is further configured to:
[0160] The speech signal is segmented into frames to obtain multiple speech frames corresponding to the speech signal.
[0161] For each speech frame, determine the corresponding frequency domain energy and frequency.
[0162] The fundamental frequency corresponding to the speech signal is determined based on multiple frequencies and multiple frequencies in that frequency domain.
[0163] Optionally, the baseband determination module 301 is further configured to:
[0164] The speech frame is pre-emphasized to obtain the first speech frame corresponding to the speech frame.
[0165] The DC component in the first speech frame is filtered out to obtain the second speech frame corresponding to the speech frame;
[0166] Windowing is applied to the second speech frame to obtain the corresponding third speech frame;
[0167] Perform a Fast Fourier Transform on the third speech frame to obtain the corresponding fourth speech frame.
[0168] Determine the frequency domain energy corresponding to the fourth speech frame.
[0169] Optionally, the baseband determination module 301 is further configured to:
[0170] The amplitude spectrum corresponding to the speech signal is determined based on multiple frequencies and multiple frequencies in that frequency domain.
[0171] Obtain the preset frequency range;
[0172] Based on the preset frequency range and the amplitude spectrum, the fundamental frequency corresponding to the speech signal is determined.
[0173] Optionally, the target screen range determination module 302 is further configured to:
[0174] Obtain the preset maximum baseband, preset minimum baseband, and preset number of operating intervals;
[0175] Based on the base frequency, the preset maximum base frequency, the preset minimum base frequency, and the preset number of operating intervals, the target screen interval corresponding to the voice signal is determined from the screen of the terminal.
[0176] Optionally, the target screen range determination module 302 is further configured to:
[0177] Based on the preset maximum base frequency and the preset minimum base frequency, the screen is divided into a preset number of preset screen intervals for the preset operation interval;
[0178] Determine the base frequency range corresponding to each preset screen interval to obtain the operation interval correlation;
[0179] Based on the base frequency, the target screen range corresponding to the voice signal is determined through the correlation of the operating range.
[0180] Optionally, the target screen range determination module 302 is further configured to:
[0181] Determine the maximum Mel-frequency corresponding to the preset maximum fundamental frequency;
[0182] Determine the minimum mezzanine frequency corresponding to the preset minimum fundamental frequency;
[0183] Based on the maximum Mel frequency and the minimum Mel frequency, the screen is divided into a preset number of preset screen intervals for the preset operation interval.
[0184] Optionally, the target screen range determination module is further configured to:
[0185] For each preset screen interval, the Mel fundamental frequency range corresponding to the preset screen interval is determined based on the maximum Mel fundamental frequency and the minimum Mel fundamental frequency, the frequency range corresponding to the Mel fundamental frequency range is determined, and the frequency range is used as the fundamental frequency range corresponding to the preset screen interval.
[0186] Optionally, the voice signals include multiple signals, with different voice signals corresponding to different users; the target screen range determination module 302 is further configured to:
[0187] Determine the largest fundamental frequency among the multiple fundamental frequencies corresponding to this speech signal;
[0188] Based on the maximum base frequency, the target screen range corresponding to the voice signal is determined from the screen of the terminal.
[0189] The aforementioned device can determine the target screen range based on the base frequency corresponding to the voice signal, and control the terminal to execute preset operations based on the target screen range. Since the base frequency is not affected by the user's own pronunciation, the preset operations executed based on the base frequency are more accurate, thereby improving the accuracy of operation control and enhancing the user's operating experience.
[0190] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0191] This disclosure also provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the operation control method provided in this disclosure.
[0192] Figure 4 This is a block diagram illustrating a terminal 400 according to an exemplary embodiment of the present disclosure. For example, terminal 400 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.
[0193] Reference Figure 4 Terminal 400 may include one or more of the following components: processing component 402, memory 404, power component 406, multimedia component 408, audio component 410, input / output (I / O) interface 412, sensor component 414, and communication component 416.
[0194] Processing component 402 typically controls the overall operation of terminal 400, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 402 may include one or more processors 420 to execute instructions to complete all or part of the steps of the operation control method described above. Furthermore, processing component 402 may include one or more modules to facilitate interaction between processing component 402 and other components. For example, processing component 402 may include a multimedia module to facilitate interaction between multimedia component 408 and processing component 402.
[0195] Memory 404 is configured to store various types of data to support operation on terminal 400. Examples of this data include instructions for any application or method operating on terminal 400, contact data, phonebook data, messages, pictures, videos, etc. Memory 404 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0196] Power component 406 provides power to various components of terminal 400. Power component 406 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to terminal 400.
[0197] Multimedia component 408 includes a screen that provides an output interface between the terminal 400 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 408 includes a front-facing camera and / or a rear-facing camera. When the terminal 400 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0198] Audio component 410 is configured to output and / or input audio signals. For example, audio component 410 includes a microphone (MIC) configured to receive external audio signals when terminal 400 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 404 or transmitted via communication component 416. In some embodiments, audio component 410 also includes a speaker for outputting audio signals.
[0199] I / O interface 412 provides an interface between processing component 402 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0200] Sensor assembly 414 includes one or more sensors for providing state assessments of various aspects of terminal 400. For example, sensor assembly 414 may detect the on / off state of terminal 400, the relative positioning of components such as the display and keypad of terminal 400, changes in the position of terminal 400 or a component of terminal 400, the presence or absence of user contact with terminal 400, the orientation or acceleration / deceleration of terminal 400, and temperature changes of terminal 400. Sensor assembly 414 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 414 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 414 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0201] Communication component 416 is configured to facilitate wired or wireless communication between terminal 400 and other devices. Terminal 400 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 416 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 416 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0202] In an exemplary embodiment, terminal 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described operation control method.
[0203] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 404 including instructions, which can be executed by a processor 420 of a terminal 400 to complete the above-described operation control method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0204] In another exemplary embodiment, a computer program product is also provided, the computer program product comprising a computer program executable by a programmable device, the computer program having a code portion for performing the above-described operation control method when executed by the programmable device.
[0205] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of this disclosure. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0206] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An operation control method, characterized in that, Applied to a terminal, the method includes: In response to the user's input voice signal, determine the base frequency corresponding to the voice signal; Based on the base frequency, the target screen range corresponding to the voice signal is determined from the screen of the terminal; Based on the target screen range, the terminal is controlled to perform a preset operation. The preset operation is determined according to the application scenario of the terminal, and different application scenarios correspond to different preset operations.
2. The method according to claim 1, characterized in that, Determining the fundamental frequency corresponding to the voice signal includes: The speech signal is segmented into frames to obtain multiple speech frames corresponding to the speech signal; For each of the speech frames, determine the frequency domain energy and frequency corresponding to the speech frame; The fundamental frequency corresponding to the speech signal is determined based on the multiple frequency domain energies and the multiple frequencies.
3. The method according to claim 2, characterized in that, Determining the frequency domain energy corresponding to the speech frame includes: The speech frame is pre-emphasized to obtain the first speech frame corresponding to the speech frame; Filter out the DC component in the first speech frame to obtain the second speech frame corresponding to the speech frame; Windowing is applied to the second speech frame to obtain the third speech frame corresponding to the speech frame; Perform a Fast Fourier Transform on the third speech frame to obtain the fourth speech frame corresponding to the speech frame; Determine the frequency domain energy corresponding to the fourth speech frame.
4. The method according to claim 2, characterized in that, The step of determining the fundamental frequency corresponding to the speech signal based on the plurality of frequency domain energies and the plurality of frequencies includes: The amplitude spectrum corresponding to the speech signal is determined based on the multiple frequency domain energies and the multiple frequencies; Obtain the preset frequency range; The fundamental frequency corresponding to the speech signal is determined based on the preset frequency range and the amplitude spectrum.
5. The method according to claim 1, characterized in that, Determining the target screen range corresponding to the voice signal from the screen of the terminal based on the base frequency includes: Obtain the preset maximum baseband, preset minimum baseband, and preset number of operating intervals; Based on the base frequency, the preset maximum base frequency, the preset minimum base frequency, and the preset number of operating intervals, the target screen interval corresponding to the voice signal is determined from the screen of the terminal.
6. The method according to claim 5, characterized in that, The step of determining the target screen interval corresponding to the voice signal from the screen of the terminal based on the base frequency, the preset maximum base frequency, the preset minimum base frequency, and the preset number of operating intervals includes: Based on the preset maximum base frequency and the preset minimum base frequency, the screen is divided into the preset number of preset screen intervals. Determine the base frequency range corresponding to each preset screen interval to obtain the operation interval correlation relationship; Based on the base frequency, the target screen range corresponding to the voice signal is determined through the operation range correlation.
7. The method according to claim 6, characterized in that, The step of dividing the screen into a preset number of preset screen intervals based on the preset maximum base frequency and the preset minimum base frequency includes: Determine the maximum Mel-frequency corresponding to the preset maximum fundamental frequency; Determine the minimum Mel-frequency corresponding to the preset minimum fundamental frequency; Based on the maximum Mel frequency and the minimum Mel frequency, the screen is divided into a preset number of preset screen intervals.
8. The method according to claim 7, characterized in that, Determining the base frequency range corresponding to each preset screen interval includes: For each preset screen interval, the Mel fundamental frequency range corresponding to the preset screen interval is determined based on the maximum Mel fundamental frequency and the minimum Mel fundamental frequency, the frequency range corresponding to the Mel fundamental frequency range is determined, and the frequency range is used as the fundamental frequency range corresponding to the preset screen interval.
9. The method according to any one of claims 1-8, characterized in that, The voice signals include multiple signals, and different voice signals correspond to different users; Determining the target screen range corresponding to the voice signal from the screen of the terminal based on the base frequency includes: Determine the largest fundamental frequency among the fundamental frequencies corresponding to the multiple speech signals; Based on the maximum base frequency, the target screen range corresponding to the voice signal is determined from the screen of the terminal.
10. An operation control device, characterized in that, Applied to a terminal, the device includes: The baseband determination module is configured to determine the baseband corresponding to the voice signal in response to the voice signal input by the user. The target screen range determination module is configured to determine the target screen range corresponding to the voice signal from the screen of the terminal based on the base frequency; The control module is configured to control the terminal to perform a preset operation based on the target screen range. The preset operation is determined according to the application scenario of the terminal, and different application scenarios correspond to different preset operations.
11. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method described in any one of claims 1-9.
12. A terminal, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-9.