Processing method, intelligent terminal and storage medium
By receiving and identifying voice information in the voice wake-up system and performing conversion processing to fix abnormal wake-up words, the problem of decreasing recognition accuracy when user vocal cords or voice-making abnormalities is solved, the efficiency and accuracy of device wake-up are improved, and the user experience is enhanced.
Patent Information
- Application Number
- CN202510169155.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-10
AI Technical Summary
The existing voice wake-up system has reduced recognition accuracy in the case of organic lesions of the user's vocal cords or dysarthria, resulting in poor user experience.
By receiving voice information, preliminary recognition and conversion processing are performed, abnormal wake-up words are repaired as normal wake-up words, and the efficiency and accuracy of device wake-up are improved.
Improves the efficiency and accuracy of device wake-up and enhances the user experience, especially when the user's vocal cords or voice accompaniment abnormalities.
Smart Images

Figure CN120126470A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of intelligent terminals, and particularly relates to a processing method, an intelligent terminal, and a storage medium. Background Art
[0002] With the continuous development of Internet technology, most intelligent terminals on the market will integrate an AI assistant, and this kind of AI assistant often supports voice wake-up of users. Users can wake up the intelligent terminal in the standby state through voice, so that the intelligent terminal enters the running state from the standby state. Voice wake-up means that the intelligent terminal detects specific keywords from a continuous voice stream, emits a signal when the specific keywords are detected, and then wakes up the intelligent terminal. Among them, the specific keywords are wake-up words. Users can customize wake-up words in the intelligent terminal and wake up the corresponding intelligent terminal through the voice carrying the wake-up words.
[0003] The applicant found that in the actual use process, when the organic lesions of the user's vocal cords reach a critical level, it often induces abnormal states of vocal cord vibration; and / or, dysarthria or cold symptoms, such as hoarseness of voice, increased nasal resonance, and coughing, will all affect the voice characteristics. Since the voice wake-up system mainly relies on these voice features to identify wake-up words, any atypical change in the voice may interfere with the vocalization model established during the system learning process, thereby reducing the recognition accuracy. Although the current mobile phone voice wake-up technology has achieved relatively satisfactory wake-up effects under normal pronunciation conditions, it still fails to fully address the recognition difficulties caused by abnormal voices, and the user experience is poor.
[0004] The foregoing description is for providing general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] In view of the above technical problems, this application provides a processing method, an intelligent terminal, and a storage medium, which can automatically identify the user's voice state, repair abnormal wake-up words into normal wake-up words, and improve the efficiency and / or accuracy of device wake-up.
[0006] To solve the above technical problems, this application provides a processing method, optionally applied to an intelligent terminal, including the steps of:
[0007] S11: Receive voice information and recognize the voice information;
[0008] S12: Output a voice wake-up interface according to the recognition result.
[0009] Optionally, the S11 step includes at least one of the following:
[0010] Receive voice information and perform a preliminary recognition on the voice information;
[0011] Perform conversion processing on the speech information according to the preliminary recognition result;
[0012] Perform secondary recognition on the converted speech information.
[0013] Optionally, the preliminary recognition of the speech information includes at least one of the following:
[0014] Extract the voiceprint feature of the speech information and determine whether it contains a preset voiceprint feature;
[0015] Extract the spectral feature of the speech information and determine whether it contains a preset spectral feature;
[0016] Extract the harmonic distortion value of the speech information and determine whether it exceeds a preset harmonic distortion value.
[0017] Optionally, the processing method further includes:
[0018] Perform fuzzy detection of the wake-up word on the speech information;
[0019] If the fuzzy detection passes, extract the voiceprint feature, or spectral feature, or harmonic distortion value of the speech information according to the preset wake-up mode.
[0020] Optionally, the conditions for passing the fuzzy detection include at least one of the following:
[0021] The pitch feature of the speech information and the wake-up word exceeds a first preset similarity;
[0022] The timbre feature of the speech information and the wake-up word exceeds a second preset similarity;
[0023] The similarity between the speech information and the wake-up word output by the preset neural network exceeds a third preset similarity.
[0024] Optionally, the conversion processing of the speech information includes at least one of the following:
[0025] Extract the feature information in the speech information and adjust the feature parameters of the speech information according to the feature information;
[0026] Extract the semantic information in the speech information and regenerate the speech information according to the semantic information and the user speech library;
[0027] Obtain the visual feature and / or tactile feature when the user is speaking, and perform conversion processing on the speech information in combination with the visual feature and / or tactile feature.
[0028] Optionally, the step S12 includes:
[0029] If the recognition result includes a first result indicating failure twice, a preset wake-up mode is enabled;
[0030] Perform a third recognition on the voice information according to the preset wake-up mode;
[0031] If the third recognition result includes a second result indicating success, output a voice wake-up interface.
[0032] Optionally, the method further includes:
[0033] If the third recognition result includes a first result indicating failure, output a manual wake-up interface including preset controls.
[0034] This application also provides an intelligent terminal. The intelligent terminal includes a memory and a processor. A processing program is stored on the memory. When the processing program is executed by the processor, the steps of any one of the above-mentioned processing methods are implemented.
[0035] This application also provides a storage medium. The storage medium stores a processing program. When the processing program is executed by the processor, the steps of any one of the above-mentioned processing methods are implemented.
[0036] As described above, the processing method of this application can receive voice information, recognize the voice information, and output a voice wake-up interface according to the recognition result. Through the technical solution of this application, the voice state of the user can be automatically recognized, and the abnormal wake-up word can be repaired into a normal wake-up word, improving the efficiency and / or accuracy of device wake-up. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 Schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of this application;
[0039] Figure 2 Schematic diagram of a communication network system architecture provided by an embodiment of this application;
[0040] Figure 3 First flowchart of the processing method provided by an embodiment of this application;
[0041] Figure 4 First scenario diagram of the processing method provided by an embodiment of this application;
[0042] Figure 5 It is a schematic diagram of the second scenario of the processing method provided by the embodiment of the present application;
[0043] Figure 6 It is a schematic diagram of the setting interface provided by the embodiment of the present application;
[0044] Figure 7 It is a schematic diagram of the wake-up mode selection interface provided by the embodiment of the present application;
[0045] Figure 8 It is a schematic diagram of the wake-up interface provided by the embodiment of the present application;
[0046] Figure 9 It is a schematic diagram of the second process of the processing method provided by the embodiment of the present application;
[0047] Figure 10 It is a schematic diagram of the structure of the processing device provided by the embodiment of the present application;
[0048] Figure 11 It is a schematic diagram of the structure of another processing device provided by the embodiment of the present application.
[0049] The realization of the purpose of the present application, functional features and advantages will be further described in conjunction with the embodiments with reference to the accompanying drawings. Through the above-mentioned accompanying drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0050] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0051] It should be noted that in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element. In addition, components, features, and elements with the same name in different embodiments of this application may have the same meaning or different meanings, and their specific meanings need to be determined based on their explanations in the specific embodiments or further in combination with the context in the specific embodiments.
[0052] It should be understood that although the terms first, second, third, etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this document, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining". Furthermore, as used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context indicates otherwise. It should be further understood that the terms "comprise", "include" indicate the presence of the stated features, steps, operations, elements, components, items, kinds, and / or groups, but do not exclude the presence, occurrence or addition of one or more other features, steps, operations, elements, components, items, kinds, and / or groups. The terms "or", "and / or", "include at least one of the following" and the like used in this application can be interpreted inclusively, or mean any one or any combination. For example, "include at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C", and again, "A, B or C" or "A, B and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C". An exception to this definition only occurs when the combination of elements, functions, steps or operations is inherently mutually exclusive in some way.
[0053] It should be understood that although the steps in the flowchart in the embodiments of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and they can be executed in other orders. Moreover, at least a part of the steps in the figure may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0054] Depending on the context, as used herein, the words "if", "when" can be interpreted as "when...", "when...", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined" or "if it is detected (stated condition or event)" can be interpreted as "when it is determined", "in response to determining", "when it is detected (stated condition or event)", or "in response to detecting (stated condition or event)".
[0055] It should be noted that in this article, step codes such as S11, S12, etc. are used. The purpose is to more clearly and briefly express the corresponding content and do not constitute a substantial limitation in order. Those skilled in the art may execute S12 first and then S11 during specific implementation, etc., but these should all be within the protection scope of the present application.
[0056] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] In the following description, the suffixes such as "module", "component" or "unit" used to represent elements are only for the convenience of the description of the present application, and they have no specific meaning in themselves. Therefore, "module", "component" or "unit" can be used interchangeably.
[0058] The intelligent terminal can be implemented in various forms. For example, the intelligent terminal described in the present application may include mobile terminals such as mobile phones, tablet computers, laptop computers, palm computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc. In other embodiments, it may also include fixed terminals such as digital TVs, desktop computers, etc.
[0059] In the following description, a mobile terminal will be taken as an example for illustration. Those skilled in the art will understand that, except for the components specifically for mobile purposes, the structure according to the embodiments of the present application can also be applied to fixed-type terminals.
[0060] Please refer to Figure 1 , which is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application. The mobile terminal 100 may include components such as an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (audio / video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111. Those skilled in the art can understand that Figure 1 the mobile terminal structure shown in
[0061] does not limit the mobile terminal. The mobile terminal may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. Figure 1 The following will optionally introduce each component of the mobile terminal:
[0062] The radio frequency unit 101 can be used for receiving and sending information or signals during communication. Optionally, after receiving the downlink information of the base station, it is sent to the processor 110 for processing; in addition, the uplink data is sent to the base station. Generally, the radio frequency unit 101 includes but is not limited to antennas, at least one amplifier, transceivers, couplers, low-noise amplifiers, duplexers, etc. In addition, the radio frequency unit 101 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), the fifth-generation (5G) mobile communication system, 6G, etc.
[0063] WiFi belongs to short-distance wireless transmission technology. The mobile terminal can help users send and receive emails, browse the web, and access streaming media through the WiFi module 102, which provides users with wireless broadband Internet access. Although Figure 1 the WiFi module 102 is shown, it can be understood that it is not an essential component of the mobile terminal and can be omitted entirely within the scope of not changing the essence of the invention according to needs.
[0064] The audio output unit 103 can convert the audio data received by the RF unit 101 or the WiFi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in modes such as a call signal reception mode, a call mode, a recording mode, a voice recognition mode, a broadcast reception mode, etc. Moreover, the audio output unit 103 can also provide an audio output related to a specific function executed by the mobile terminal 100 (e.g., a call signal reception sound, a message reception sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0065] The A / V input unit 104 is used to receive an audio or video signal. The A / V input unit 104 may include a Graphics Processing Unit (GPU) 1041 and a microphone 1042. The graphics processor 1041 processes the image data of a still picture or a video obtained by an image capturing device (such as a camera) in a video capture mode or an image capture mode. The processed image frame can be displayed on the display unit 106. The processed image frame can be stored in the memory 109 (or other storage media) or transmitted via the RF unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) via the microphone 1042 in operation modes such as a phone call mode, a recording mode, a voice recognition mode, etc., and can process such sound into audio data. The processed audio (voice) data can be output in a format that can be transmitted to a mobile communication base station via the RF unit 101 in the case of a phone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to cancel (or suppress) the noise or interference generated during the reception and transmission of the audio signal.
[0066] The mobile terminal 100 further includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as a pedometer, a tap), etc.; as for other sensors that the mobile phone can also be configured with, such as a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., they will not be elaborated here.
[0067] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, and the display panel 1061 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.
[0068] The user input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function controls of the mobile terminal. Optionally, the user input unit 107 may include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 1071), and drive corresponding connection devices according to a preset program. The touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 110, and can receive commands sent by the processor 110 and execute them. In addition, the touch panel 1071 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may further include other input devices 1072. Optionally, the other input devices 1072 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc., and specific details are not limited here.
[0069] Optionally, the touch panel 1071 may cover the display panel 1061. After the touch panel 1071 detects a touch operation on or near it, it is transmitted to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides a corresponding visual output on the display panel 1061 according to the type of touch event. Although in Figure 1 the touch panel 1071 and the display panel 1061 are implemented as two independent components to realize the input and output functions of the mobile terminal, in some embodiments, the touch panel 1071 and the display panel 1061 may be integrated to realize the input and output functions of the mobile terminal, and specific details are not limited here.
[0070] The interface unit 108 serves as an interface through which at least one external device can be connected to the mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. The interface unit 108 can be used to receive inputs from an external device (such as data information, power, etc.) and transmit the received inputs to one or more components within the mobile terminal 100 or can be used to transfer data between the mobile terminal 100 and the external device.
[0071] The memory 109 can be used to store software programs and various data. The memory 109 mainly includes a program storage area and a data storage area. Optionally, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 109 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0072] The processor 110 is the control center of the mobile terminal, connecting various parts of the entire mobile terminal using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by invoking data stored in the memory 109, it executes various functions of the mobile terminal and processes data, thereby monitoring the mobile terminal as a whole. The processor 110 can include one or more processing units; preferably, the processor 110 can integrate an application processor and a modem processor. Optionally, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 110.
[0073] The mobile terminal 100 can also include a power supply 111 (such as a battery) for powering each component. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system.
[0074] Although Figure 1 not shown, the mobile terminal 100 can also include a Bluetooth module, etc., which will not be elaborated here.
[0075] To facilitate understanding of the embodiments of the present application, the communication network system on which the mobile terminal of the present application is based will be described below.
[0076] Please refer toFigure 2 , Figure 2 This is an architecture diagram of a communication network system provided by an embodiment of the present application. The communication network system is an LTE system of the Universal Mobile Telecommunications Technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and an operator's IP service 204 that are communicatively connected in sequence.
[0077] Optionally, the UE 201 may be the above-mentioned mobile terminal 100, which will not be elaborated here.
[0078] The E-UTRAN 202 includes an eNodeB 2021 and other eNodeBs 2022, etc. Optionally, the eNodeB 2021 may be connected to other eNodeBs 2022 through a backhaul (such as an X2 interface). The eNodeB 2021 is connected to the EPC 203, and the eNodeB 2021 may provide access for the UE 201 to the EPC 203.
[0079] The EPC 203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gate Way) 2034, a PGW (PDN Gate Way) 2035, and a PCRF (Policy and Charging Rules Function) 2036, etc. Optionally, the MME 2031 is a control node that processes the signaling between the UE 201 and the EPC 203 and provides bearer and connection management. The HSS 2032 is used to provide some registers to manage functions such as a home location register (not shown in the figure) and stores some user-specific information such as service characteristics and data rates. All user data can be sent through the SGW 2034. The PGW 2035 may provide IP address allocation for the UE 201 and other functions. The PCRF 2036 is a policy and charging control policy decision point for service data flows and IP bearer resources, and it selects and provides available policy and charging control decisions for a policy and charging enforcement functional unit (not shown in the figure).
[0080] The IP service 204 may include the Internet, an intranet, IMS (IP Multimedia Subsystem), or other IP services, etc.
[0081] Although the above has been described by taking the LTE system as an example, those skilled in the art should understand that this application is not only applicable to the LTE system, but also applicable to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, and future new network systems (such as 6G), etc., which are not limited herein.
[0082] Based on the above mobile terminal hardware structure and communication network system, each embodiment of this application is proposed.
[0083] Please refer to Figure 3 , Figure 3 is the first process schematic diagram of the processing method provided by the embodiment of this application. The processing method includes the steps:
[0084] S11. Receive voice information and recognize the voice information.
[0085] The execution subject of the embodiment of this application may be an intelligent terminal, or a processing device provided in the intelligent terminal. Optionally, the processing device may be implemented by software, or by a combination of software and hardware. The above intelligent terminal may be an intelligent phone, a tablet computer, a notebook computer, an Ultra-mobile Personal Computer (UMPC), a netbook, a Personal Digital Assistant (PDA), etc., and is not limited thereto.
[0086] Optionally, the above voice information may be obtained through a microphone in the intelligent terminal for subsequent voice wake-up process. Among them, voice wake-up refers to the way of activating the voice assistant and starting the subsequent interaction process by saying a preset wake-up word. As Figure 4 shown, it is the application scenario schematic diagram of the embodiment of this application. The application scenario diagram may include an intelligent terminal 210 and a server 220. The intelligent terminal 210 may be, for example, any device such as an intelligent speaker, a mobile phone, a computer, etc. that involves the need for voice wake-up word detection. Based on the above intelligent terminal 210, the user can interact with the intelligent terminal 210 through voice commands. And, in some embodiments, when the intelligent terminal 210 is in the standby state, it can receive the voice information input by the user and recognize the voice information.
[0087] Optionally, the intelligent terminal 210 may be installed with a voice wake-up system, which has a voice wake-up word detection function or a function of initiating a voice wake-up word detection request. The voice wake-up system involved in the embodiments of the present application may be a software client, or a client such as a web page or a mini-program, and the server 220 is a server corresponding to the software or the web page, mini-program, etc., without limiting the specific type of the client. The server 220 may be, for example, an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or may also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, i.e., Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited thereto. It should be noted that the voice wake-up word detection method in the embodiments of the present application may be executed independently by the intelligent terminal 210 or the server 220, or may be jointly executed by the server 220 and the intelligent terminal 210.
[0088] Optionally, the intelligent terminal 210 and the server 220 may be directly or indirectly communicatively connected through one or more networks 230. The network 230 may be a wired network or a wireless network. For example, the wireless network may be a mobile cellular network or a Wireless-Fidelity (WIFI) network. Of course, it may also be other possible networks, and the embodiments of the present application do not limit this.
[0089] When a voice wake-up word has been configured in the intelligent terminal 210, the user can wake up the intelligent terminal 210 by the voice wake-up word. In some embodiments, the voice wake-up word may be a fixed keyword configured for the intelligent terminal 210, or may be a keyword customized by the user in the intelligent terminal 210. Optionally, before using the voice wake-up function in this embodiment, the user may pre-enter the wake-up word. As Figure 5 shown, the user says "Hello, Little X" to the intelligent terminal, and the intelligent terminal receives the voice of "Hello, Little X" and stores it as the wake-up word.
[0090] Optionally, after obtaining the wake-up word input by the user, the intelligent terminal system may also extract the voiceprint feature of the user in the normal pronunciation state through audio processing and machine learning algorithms. The system will preprocess the collected voiceprint data, including operations such as removing noise and enhancing signals, and then convert it into a digital feature vector through a feature extraction algorithm (such as Mel Frequency Cepstral Coefficients - MFCC, etc.) and store it in the local database or the cloud server.
[0091] Optionally, before recognizing the voice information, this embodiment can also pre-select the wake-up mode of the smart terminal. Specifically, it can include the normal wake-up mode and the repair wake-up mode. For example Figure 6 As shown, the user can enter the secondary menu of voice settings through the "Voice Assistant" setting in the system settings. This secondary menu can include "Personal Information", "Voiceprint Management", "Tone Setting", "Volume Setting", "Button Wake-up", "Wake-up Word Setting", and "Wake-up Mode". When the user clicks on "Wake-up Mode", they can jump to Figure 7 the interface and select the normal wake-up mode or the repair wake-up mode in the pop-up window.
[0092] Optionally, for users who may be at risk of abnormal voice, such as those suffering from chronic pharyngitis, often having colds, or engaged in occupations that frequently use the voice, it is recommended to enable the repair wake-up mode. Specifically, the system can identify users at risk of abnormal voice through some predefined rules or the user's self-declaration. For example, provide an option in the settings interface for the user to choose whether they have a history of voice diseases, whether they have symptoms of upper respiratory tract infection recently, etc. Or, the system can also automatically recommend enabling the repair wake-up mode based on the voice quality analysis results when the user first enters the wake-up word, such as detecting that indicators such as the clarity and pitch stability of the voice are lower than a certain threshold.
[0093] Optionally, in the normal wake-up mode, the smart terminal system will only perform wake-up word recognition and voiceprint verification based on the normal voice samples and voiceprint characteristics initially entered by the user. When the user says the wake-up word in a normal pronunciation state, the system will process it according to the preset recognition process. This mode is suitable for the situation where the user's voice state is normal and can provide a relatively fast and simple wake-up experience. In the repair wake-up mode, the smart terminal system can further use voice conversion technology to repair possible abnormal voices to improve the wake-up success rate. For example, the system will load the corresponding voice repair model and algorithm in the background. When detecting abnormal voice, it will automatically start the repair process, convert the abnormal voice into a voice close to the normal state, and then perform recognition and voiceprint verification.
[0094] Optionally, in the repair wake-up mode, the system can continuously monitor the ambient sound in the primary wake-up state. When a potential wake-up word is captured, the system can perform a preliminary recognition on it to evaluate whether the captured voice information conforms to the normal voiceprint characteristics. If abnormal voice is detected (such as hoarseness or nasal congestion), the system can use voice conversion technology to process the abnormal voice to convert the abnormal wake-up word into a voice close to the normal state for secondary recognition. Then, the repaired voice is transmitted to the secondary wake-up engine for in-depth recognition. If the voice matches the preset wake-up word and the voiceprint verification passes, it can be determined that the wake-up is successful and the subsequent steps can be jumped to.
[0095] Optionally, after detecting a voice anomaly, the system can first perform in-depth feature extraction on the captured voice. For example, algorithms such as Mel Frequency Cepstral Coefficients (MFCC) can be used to analyze the voice signal from multiple dimensions in the time domain and frequency domain to obtain key acoustic features including fundamental frequency, formants, and harmonic structure. For hoarse voices, features such as harmonic noise ratio and fundamental frequency jitter are analyzed; for nasal voices, attention is paid to the characteristics of enhanced low-frequency energy and nasal resonance frequency bands. At the same time, a comparison is made with the stored normal voiceprint features to determine the specific type of anomaly. Then, according to the type of voice anomaly, the system selects a suitable model from the pre-trained voice conversion model library. Such as recurrent neural network (RNN) or convolutional neural network (CNN) models based on deep learning, which are trained on a large number of normal and abnormal voice pair data to learn the conversion rules of voice features. For hoarse voices, a model specifically for repairing vocal cord vibration anomalies may be selected; for nasal voices, a model for adjusting nasal resonance and low-frequency energy is chosen. Then, the system further adjusts the conversion parameters of the model according to the degree of voice anomaly. For example, in hoarse voice conversion, if the fundamental frequency jitter is severe, the weight for fundamental frequency smoothing is increased; if the low-frequency energy of nasal voice is excessively enhanced, the low-frequency gain parameter is appropriately reduced. By dynamically adjusting the parameters, it is ensured that the converted voice is closer to normal voice in terms of acoustic features while retaining the individual characteristics of the user's voice, avoiding voice distortion caused by excessive conversion. Finally, the voice conversion operation is performed using the model with adjusted parameters to generate the converted voice.
[0096] Optionally, post-processing can also be performed on the converted voice. Specifically, signal enhancement techniques can be used to improve the clarity and intelligibility of the voice and remove possible conversion noise. For example, the voice signal is enhanced through methods such as Wiener filtering, and after multiple iterations of optimization, the wake-up word voice close to the normal state is finally obtained for subsequent secondary wake-up recognition and voiceprint verification processes.
[0097] S12: Output a voice wake-up interface according to the recognition result.
[0098] When it is determined that the above wake-up word matches and the voiceprint verification passes, the voice assistant interface can be directly opened, and subsequent voice commands from the user can be obtained, such as Figure 8 shown. The voice assistant interface can also include a text input box. When receiving subsequent voice commands from the user, they can be automatically converted into text and displayed in the text input box for the user to check whether the current text conversion is accurate.
[0099] Optionally, if the smart terminal is currently in the normal wake-up mode, when it is determined that the received voice information fails to be recognized twice in a row, the repair wake-up mode can be automatically enabled, and the voice information can be converted and then recognized for the third time. If the recognition is successful, the voice wake-up interface is output. That is, if the recognition result contains the first result indicating failure twice, the preset wake-up mode is enabled, and the voice information is recognized for the third time according to the preset wake-up mode. If the third recognition result contains the second result indicating success, the voice wake-up interface is output.
[0100] Specifically, in the normal wake-up mode, when the first recognition result shows that it does not match the preset wake-up word, the system does not immediately switch the mode but waits for the second recognition result. If both of these two recognition results are failures, the start condition of the repair wake-up mode is triggered. After entering the repair wake-up mode, the voice information can be repaired and then recognized for the third time. If the repaired voice information matches the preset wake-up word and the voiceprint verification passes, it can be determined that the wake-up is successful, and the voice wake-up interface is output.
[0101] Optionally, if the third recognition result still contains the first result indicating failure, a manual wake-up interface including preset controls is output. The manual wake-up interface can include a password unlocking control, a fingerprint unlocking control, a pattern unlocking control, etc. For those users who cannot successfully use voice wake-up due to voice problems in the short term, they can quickly enter the device by means of familiar manual unlocking methods, avoiding being unable to use due to voice function disorders, and ensuring the continuity and convenience of device use.
[0102] As can be seen from the above, the processing method of the embodiment of the present application can receive voice information, recognize the voice information, and output a voice wake-up interface according to the recognition result. Through the technical solution of the present application, the voice state of the user can be automatically recognized, and the abnormal wake-up word can be repaired into a normal wake-up word, improving the efficiency and / or accuracy of device wake-up.
[0103] The embodiment of the present application also provides a processing method. Please refer to Figure 9 , Figure 9 which is the second process schematic diagram of the processing method provided by the embodiment of the present application. The method includes the steps:
[0104] S21. Receive voice information and perform preliminary recognition on the voice information.
[0105] Optionally, the smart terminal in this embodiment defaults to enabling the repair and wake-up mode. In this mode, if it is determined that the initially recognized voice information is abnormal voice, the voice information can be directly repaired and recognized again. Among them, the steps of initially recognizing the voice information may include at least one of the following: extracting the voiceprint characteristics of the voice information and determining whether it contains preset voiceprint characteristics; extracting the spectral characteristics of the voice information and determining whether it contains preset spectral characteristics; extracting the harmonic distortion value of the voice information and determining whether it exceeds the preset harmonic distortion value.
[0106] Optionally, the terminal system can extract the voiceprint characteristics in the voice information based on voiceprint recognition algorithms such as Gaussian Mixture Model (GMM) or Deep Neural Network (DNN). These algorithms can capture unique acoustic characteristics from the voice signal, such as the voiceprint patterns reflected by the vocal tract shape, pronunciation habits, etc. During the judgment process, the extracted voiceprint characteristics are compared with the preset voiceprint characteristics stored in advance. Among them, the preset voiceprint characteristics are stored in the system after multiple collections of the user's normal voice samples and analysis and processing when the user first sets the voice wake-up function. If the extracted voiceprint characteristics do not match the preset voiceprint characteristics, it can be determined that the voice information is abnormal.
[0107] Optionally, using spectrum analysis techniques such as Fast Fourier Transform (FFT), the terminal system can convert the voice signal from the time domain to the frequency domain to obtain the spectral characteristics of the voice, which may include information such as frequency distribution and energy concentration region. For normal voice, its spectral characteristics have certain regularity. For example, there are relatively stable energy peaks in certain frequency bands, corresponding to the pronunciation characteristics of vowels and consonants. When the voice is abnormal, such as due to vocal cord lesions or nasal congestion, the spectral characteristics will change significantly. The system compares with the preset spectral characteristic template, and if obvious differences are found, it will determine that the voice information is abnormal.
[0108] Optionally, the harmonic structure in the voice signal reflects the characteristics of vocal cord vibration. Under normal circumstances, there is a relatively stable proportional relationship between harmonics. When abnormal situations such as hoarseness occur, the vocal cord vibration is irregular, resulting in an increase in harmonic distortion. The system calculates the harmonic distortion value of the voice information and compares it with the preset harmonic distortion value. If it exceeds the preset value, it indicates that the voice information is abnormal.
[0109] Optionally, before extracting the voiceprint feature, spectral feature, or harmonic distortion value of the voice information, the voice information can also be subjected to a wake-word fuzzy detection. If the fuzzy detection passes, the subsequent steps are continued. Among them, the conditions for passing the fuzzy detection can include at least one of the following: the pitch feature of the voice information exceeds the first preset similarity with the wake word; the timbre feature of the voice information exceeds the second preset similarity with the wake word; the similarity between the voice information and the wake word output by the preset neural network exceeds the third preset similarity.
[0110] Optionally, the pitch feature reflects the fundamental frequency change of the voice. The terminal system can analyze the fundamental frequency curve of the voice information and compare it with the standard fundamental frequency curve of the wake word. For example, the dynamic time warping (DTW) algorithm is used to calculate the similarity of the two fundamental frequency curves. If the pitch feature of the voice information exceeds the first preset similarity with the wake word, it indicates that the voice has a high possibility of being related to the wake word in terms of pitch.
[0111] Optionally, the timbre feature depends on factors such as the harmonic structure and formants of the voice. The terminal system can use techniques such as linear predictive coding (LPC) to extract the timbre feature parameters of the voice information and the wake word, and then judge by calculating the distance or similarity measure (such as Euclidean distance, cosine similarity, etc.) between these parameters. When the timbre feature of the voice information exceeds the second preset similarity with the wake word, the possibility that the voice is the wake word is further increased.
[0112] Optionally, when using the preset neural network to judge the similarity, this neural network can be an architecture such as a convolutional neural network (CNN) or a recurrent neural network (RNN) based on deep learning. In the training stage, the neural network is trained using a large number of wake-word samples and related voice variant samples to learn the complex mapping relationship between different voices and the wake word. During detection, the voice information is input into the neural network, and the network outputs the similarity between the voice information and the wake word. If this similarity exceeds the third preset similarity, it also meets the condition for passing the fuzzy detection.
[0113] S22. Perform conversion processing on the voice information according to the preliminary recognition result.
[0114] Optionally, the conversion processing of the voice information can include at least one of the following: extracting the feature information in the voice information and adjusting the feature parameters of the voice information according to the feature information; extracting the semantic information in the voice information and regenerating the voice information according to the semantic information and the user voice library; obtaining the visual feature and / or tactile feature when the user is speaking, and performing conversion processing on the voice information in combination with the visual feature and / or tactile feature.
[0115] For example, detailed movement data of the vocal cords when the user is speaking can be obtained. In the case of hoarse voice, a virtual pathological model of the vocal cords is constructed based on this data. Then, combined with the acoustic physical model of healthy vocal cords, a conversion model is trained through machine learning algorithms (such as deep generative adversarial network - GAN). This model can generate speech that is physically close to the normal vocal cord pronunciation according to the movement parameters and acoustic features of the pathological vocal cords. For example, it simulates parameters such as the vibration frequency and amplitude of healthy vocal cords to adjust acoustic elements such as the fundamental frequency and harmonics of abnormal speech, thus achieving speech conversion.
[0116] For another example, a speech database containing a large number of normal speech samples of the user is established in advance. These speech samples cover different speech styles, speaking speeds, intonations, etc. For abnormal speech, first use a style transfer network in deep learning (such as a style transfer model based on the Transformer architecture) to extract the semantic content of the abnormal speech, and at the same time learn the style features in the normal speech database. During the style transfer process, a special pathological speech feature elimination module is designed. This module can identify and weaken abnormal speech features such as hoarseness or nasal congestion. For example, for the hoarseness feature, the noise component can be specifically reduced by analyzing the harmonic - to - noise ratio in the speech signal; for the nasal congestion feature, the abnormal low - frequency resonance caused by nasal obstruction is identified and adjusted. Then the semantic content and the adjusted style features are recombined to generate speech close to the normal speech style.
[0117] For yet another example, a camera can be used to obtain the facial expressions and lip movement information of the speaker. In the case of abnormal voice, visual information is combined to assist in speech conversion. For example, when the voice is hoarse, by precisely analyzing the lip movement (a lip - reading model in deep learning can be used to improve the analysis accuracy), the expected normal movement state of the vocal cords is inferred. Based on clues such as the lip opening and closing speed and tongue position in the visual information, combined with acoustic features such as the fundamental frequency and amplitude of the speech signal, a more normal - sounding speech is generated through a multi - modal fusion neural network (such as a multi - modal Transformer).
[0118] S23. Perform wake - word matching on the converted speech information.
[0119] Optionally, the terminal system can adopt an efficient speech recognition algorithm, such as a hidden Markov model (HMM) - based or deep - learning - based end - to - end speech recognition model, to accurately compare the converted speech information with a preset wake - word. During this process, information such as the acoustic features and phoneme sequence of the speech is analyzed to determine its similarity to the wake - word. This embodiment does not further limit this.
[0120] S24. If the wake - word matching is successful, perform speaker verification on the converted speech information.
[0121] Optionally, if the wake-up word match is successful, the voiceprint matching can be continued. In the voiceprint matching stage, advanced voiceprint recognition technologies can be used, such as the Gaussian Mixture Model-Universal Background Model (GMM-UBM) or the deep neural network voiceprint recognition system. By extracting the voiceprint features of the converted speech and comparing them with the stored user-specific voiceprint template, combined with the unique frequency, rhythm and other characteristics of the voiceprint, the legitimacy of the voice source is ensured.
[0122] S25. Output a voice wake-up interface according to the matching result.
[0123] Optionally, if both the wake-up word and the voiceprint match successfully, the terminal system will activate the voice wake-up interface to provide functions such as subsequent intelligent assistant services for the user; if the match fails, corresponding prompts or retry operations can be performed according to the system settings.
[0124] As can be seen from the above, the embodiments of the present application can receive voice information, perform preliminary recognition on the voice information, perform conversion processing on the voice information according to the preliminary recognition result, perform wake-up word matching on the converted voice information. If the wake-up word match is successful, perform voiceprint matching on the converted voice information, and output a voice wake-up interface according to the matching result. Through the technical solution of the present application, the voice state of the user can be automatically recognized, the abnormal wake-up word can be repaired into a normal wake-up word, and the efficiency and / or accuracy of device wake-up is improved.
[0125] Figure 10 It is a schematic structural diagram of a processing device provided by an embodiment of the present application. The processing device can be set in an intelligent terminal or is the intelligent terminal itself. Please refer to Figure 10 , the processing device 30 includes:
[0126] An identification module 301, configured to receive voice information and perform identification on the voice information;
[0127] An output module 302, configured to output a voice wake-up interface according to the identification result.
[0128] Optionally, continue to refer to Figure 11 , the above identification module 301 may specifically include:
[0129] A first identification sub-module 3011, configured to receive voice information and perform preliminary identification on the voice information;
[0130] A conversion sub-module 3012, configured to perform conversion processing on the voice information according to the preliminary identification result;
[0131] A second identification sub-module 3013, configured to perform secondary identification on the converted voice information.
[0132] Optionally, the above conversion sub-module 3012 is specifically configured to extract the feature information in the voice information and adjust the feature parameters of the voice information according to the feature information; or extract the semantic information in the voice information and regenerate the voice information according to the semantic information and the user voice library; or obtain the visual feature and / or tactile feature when the user is speaking, and perform conversion processing on the voice information in combination with the visual feature and / or tactile feature.
[0133] Optionally, the output module 302 may specifically include:
[0134] The activation sub-module 3021 is configured to activate a preset wake-up mode when the recognition result includes a first result indicating two failures in representation;
[0135] The third recognition sub-module 3022 is configured to perform a third recognition on the voice information according to the preset wake-up mode;
[0136] The output sub-module 3023 is configured to output a voice wake-up interface when the third recognition result includes a second result indicating successful representation.
[0137] Optionally, the processing device 30 may further include:
[0138] The manual wake-up module 303 is configured to output a manual wake-up interface including preset controls when the third recognition result includes a first result indicating failure in representation.
[0139] The processing device provided in the embodiment of the present application can receive voice information, perform recognition on the voice information, and output a voice wake-up interface according to the recognition result. Through the technical solution of the present application, the voice state of the user can be automatically recognized, and the abnormal wake-up word can be repaired into a normal wake-up word, improving the efficiency and / or accuracy of device wake-up.
[0140] The embodiment of the present application further provides an intelligent terminal, which includes a memory and a processor. A processing program is stored on the memory, and when the processing program is executed by the processor, the steps of the processing method in any of the above embodiments are implemented.
[0141] The embodiment of the present application further provides a storage medium, on which a processing program is stored, and when the processing program is executed by the processor, the steps of the processing method in any of the above embodiments are implemented.
[0142] In the embodiments of the intelligent terminal and the storage medium provided in the present application, all the technical features of any of the above processing method embodiments may be included. The extended and explanatory content of the specification is basically the same as that of the above embodiments of the method, and will not be repeated here.
[0143] The embodiments of the present application also provide a computer program product. The computer program product includes computer program code. When the computer program code runs on a computer, the computer is caused to execute the methods in the above various possible embodiments.
[0144] The embodiments of the present application also provide a chip, including a memory and a processor. The memory is used for storing a computer program, and the processor is used for calling and running the computer program from the memory, so that a device installed with the chip executes the methods in the above various possible embodiments.
[0145] It can be understood that the above scenarios are only examples and do not constitute a limitation on the application scenarios of the technical solutions provided by the embodiments of the present application. The technical solutions of the present application can also be applied to other scenarios. For example, as is known to those of ordinary skill in the art, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0146] The serial numbers of the embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.
[0147] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs.
[0148] The modules or units in the devices of the embodiments of the present application can be combined, divided, and deleted according to actual needs.
[0149] In the present application, for the description of the same or similar term concepts, technical solutions, and / or application scenarios, generally only a detailed description is made when it first appears. When it appears repeatedly later, for the sake of brevity, it is generally not described again. When understanding the technical solutions and other contents of the present application, for the same or similar term concepts, technical solutions, and / or application scenarios that are not described in detail later, reference can be made to their relevant detailed descriptions before.
[0150] In the present application, the descriptions of the various embodiments have their own emphases. For the parts not described in detail or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0151] The technical features of the technical solutions of the present application can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should all be considered as the scope recorded in the present application.
[0152] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of the present application.
[0153] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (for example, floppy disk, storage disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)), etc.
[0154] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.
Claims
1. A processing method, characterized in that: Includes steps: S11: receiving voice information and recognizing the voice information; S12: Outputting a voice wake-up interface according to the recognition result.
2. The processing method according to claim 1, characterized in that: Step S11 includes at least one of the following: Receiving voice information and performing preliminary recognition on the voice information; Performing conversion processing on the voice information according to the preliminary recognition result; Perform secondary recognition on the converted voice information.
3. The processing method according to claim 2, characterized in that: The preliminary recognition of the voice information includes at least one of the following: Extracting voiceprint features of the voice information, and determining whether the voiceprint features are preset; Extracting the frequency spectrum features of the speech information, and determining whether the speech information contains preset frequency spectrum features; The harmonic distortion value of the voice information is extracted, and it is determined whether it exceeds a preset harmonic distortion value.
4. The processing method according to any one of claims 1 to 3, characterized in that: Also includes: Performing wake-up word fuzzy detection on the voice information; If the fuzzy detection passes, the voiceprint features, or the spectrum features, or the harmonic distortion values of the voice information are extracted according to the preset wake-up mode.
5. The processing method according to claim 4, characterized in that: The conditions for passing the fuzzy detection include at least one of the following: The tone feature of the voice information and the wake-up word exceeds a first preset similarity; The timbre characteristics of the voice information and the wake-up word exceed a second preset similarity; The similarity between the voice information output by the preset neural network and the wake-up word exceeds a third preset similarity.
6. The processing method according to claim 2, characterized in that: The converting and processing the voice information includes at least one of the following: Extracting feature information from the voice information, and adjusting feature parameters of the voice information according to the feature information; Extracting semantic information from the voice information, and regenerating the voice information according to the semantic information and a user voice library; The visual features and / or tactile features of the user when speaking are obtained, and the voice information is converted and processed in combination with the visual features and / or tactile features.
7. The processing method according to any one of claims 1 to 3, characterized in that: Step S12 includes: If the recognition result includes the first result of two characterization failures, starting the preset wake-up mode; Recognize the voice information for a third time according to the preset wake-up mode; If the third recognition result includes the second result representing success, the voice wake-up interface is output.
8. The processing method according to claim 7, characterized in that: Also includes: If the third recognition result includes the first result indicating failure, a manual wake-up interface including preset controls is output.
9. An intelligent terminal, characterized in that: include: A memory and a processor, wherein a processing program is stored in the memory, and when the processing program is executed by the processor, the steps of the processing method according to any one of claims 1 to 8 are implemented.
10. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the processing method according to any one of claims 1 to 8.