Intelligent dialogue wake-up control method, dialogue system, intelligent device and storage medium
By combining microphones and touch sensors, the probability of user interaction is calculated, and the interaction target between the user and the smart device is identified. This solves the problem of adding operation steps to the wake words of smart devices and improves the user experience.
Patent Information
- Application Number
- CN202511150972.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-18
AI Technical Summary
In existing technologies, the voice wake-up method for smart devices requires a wake word, which increases the operation steps and leads to a reduced user experience.
By collecting voice data through a microphone and detecting touch data from buttons through a touch sensor, the probability of user interaction is calculated. Based on the interaction probability and data, the interaction target between the user and the smart device is determined, and the smart device is controlled.
It can identify user interaction intent without the need for a wake word, reducing the difficulty of operation and improving the user experience.
Smart Images

Figure CN120748400B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart devices, and more particularly to a smart dialogue wake-up control method, dialogue system, smart device, and storage medium. Background Technology
[0002] With the development of smart devices, setting up AI dialogue modules in these devices can effectively enable interaction with users and receive voice commands through AI, thereby significantly improving the user experience. Currently, AI dialogue usually requires voice wake-up before it can begin. During the voice data collection process, wake words are identified, and the corresponding AI dialogue function is activated after the wake word is identified. This wake word-based wake-up method increases the number of steps required for dialogue, which in turn leads to a decrease in user experience.
[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main objective of this invention is to provide an intelligent dialogue wake-up control method, a dialogue system, an intelligent device, and a storage medium, aiming to improve the user's experience in waking up intelligent devices. To achieve the above objective, this invention provides an intelligent dialogue wake-up control method, which includes the following steps:
[0005] Control the microphone to collect voice data and receive touch detection data from the touch sensor to detect buttons;
[0006] Calculate the probability of user interaction based on the voice data and the touch detection data;
[0007] When the probability of user interaction is higher than the preset probability, the interaction target between the user and the smart device is determined based on the voice data and the touch detection data.
[0008] Control the smart device according to the interaction target.
[0009] Optionally, the step of determining the user's interaction target with the smart device based on the voice data and the touch detection data includes:
[0010] When the voice data does not include semantic data, and the touch detection data is light touch data, the interaction target is determined to be waiting for user input instructions, and the light touch data is when the user touches the button for less than a first preset time;
[0011] When the touch detection data is long-term touch data, the interaction target is determined to be accurate command input. The accurate command input is determined by the interception time corresponding to the voice data based on the touch detection data, and the control of the smart device is determined based on the interception time and the voice data. The light touch data is the time when the user touches the button for a period of time greater than or equal to a first preset time.
[0012] When the touch detection data is swipe touch data, the interaction target is determined to be to execute a preset control command.
[0013] Optionally, the step of calculating the probability of user interaction based on the voice data and the touch detection data includes:
[0014] Extract target feature data from the speech data, the target feature data including: semantic feature data, acoustic environment feature data, personnel information feature data, and noise feature data;
[0015] A first score is generated based on the target feature data, and a second score is generated based on the touch detection data.
[0016] The user interaction probability is calculated based on the first rating and the second rating.
[0017] Optionally, the step of generating a first score corresponding to the speech data based on the target feature data includes:
[0018] Generate corresponding feature vector data based on the semantic feature data, the acoustic environment feature data, the personnel information feature data, and the noise feature data;
[0019] The feature vector data is input into a preset scoring and recognition model to obtain the output result of the preset scoring and recognition model;
[0020] The first score is determined based on the output results.
[0021] Optionally, the step of generating a first score corresponding to the speech data based on the target feature data includes:
[0022] When a preset operation instruction exists in the semantic feature data, a first type of rating label corresponding to the voice data is generated;
[0023] When the acoustic environment feature data is a preset scenario, a second type of scoring label corresponding to the speech data is generated;
[0024] A third type of rating label is generated based on the personnel information feature data;
[0025] When the number of noise features is greater than the preset noise data, a fourth type of scoring label corresponding to the speech data is generated;
[0026] The first score is determined based on the scoring markers corresponding to the voice data.
[0027] Optionally, the step of calculating the user interaction probability based on the first rating and the second rating includes:
[0028] Calculate the total score based on the first score and the second score;
[0029] The user interaction probability is determined based on the total score data and the preset mapping relationship.
[0030] Optionally, before the step of controlling the microphone to acquire voice data, the method further includes:
[0031] Control the real-time orientation data of the gravity sensor detection device;
[0032] When the real-time orientation data is the preset orientation data, the step of controlling the microphone to collect voice data is executed.
[0033] Furthermore, to achieve the above objectives, the present invention also provides a dialogue system, the dialogue system comprising:
[0034] The acquisition module is used to control the microphone to acquire voice data and receive touch detection data from the touch sensor to detect buttons;
[0035] The calculation module is used to calculate the probability of user interaction based on the voice data and the touch detection data;
[0036] The recognition module is used to determine the interaction target between the user and the smart device based on the voice data and the touch detection data when the user interaction probability is higher than a preset probability.
[0037] An execution module is used to control the smart device according to the interaction target.
[0038] Furthermore, to achieve the above objectives, the present invention also provides an intelligent device, the intelligent device comprising: a memory, a processor, and an intelligent dialogue wake-up control program stored in the memory and executable on the processor, the intelligent dialogue wake-up control program being configured to implement the steps of the intelligent dialogue wake-up control method described in any of the above claims.
[0039] In addition, to achieve the above objectives, the present invention also provides a storage medium storing an intelligent dialogue wake-up control program, wherein the intelligent dialogue wake-up control program, when executed by a processor, implements the steps of the intelligent dialogue wake-up control method described in any of the above claims.
[0040] This invention proposes an intelligent dialogue wake-up control method. This method collects voice data via a microphone and receives touch detection data from a touch sensor. It calculates the user interaction probability based on the voice data and the touch detection data. When the user interaction probability is higher than a preset probability, it determines the user's interaction target with the smart device based on the voice data and the touch detection data. The method then controls the smart device according to the interaction target. Compared to methods that require a wake word to activate the smart device, this method can identify whether the user needs to interact through multiple methods without requiring an accurate wake word. This effectively reduces the user's operational difficulty and effectively identifies the user's interaction target, thereby significantly improving the user experience. Attached Figure Description
[0041] Figure 1 This is a schematic diagram of the structure of an intelligent device in the hardware operating environment involved in the embodiments of the present invention;
[0042] Figure 2 This is a flowchart illustrating the first embodiment of the intelligent dialogue wake-up control method of the present invention;
[0043] Figure 3 This is a flowchart illustrating the second embodiment of the intelligent dialogue wake-up control method of the present invention;
[0044] Figure 4 This is a flowchart illustrating the third embodiment of the intelligent dialogue wake-up control method of the present invention;
[0045] Figure 5 A front view of the hardware for an AI-powered chatbot;
[0046] Figure 6 A top-down view of the hardware for an AI-powered chatbot;
[0047] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0048] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0049] Reference Figure 1 , Figure 1 This is a schematic diagram of the intelligent device structure of the hardware operating environment involved in the embodiments of the present invention.
[0050] like Figure 1As shown, the smart device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, an interaction device 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The interaction device 1003 may include a display screen and an input unit such as a keyboard. Optionally, the interaction device 1003 may also be connected to the communication bus via a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed random access memory (RAM) or a stable non-volatile memory (NVM), such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0051] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on smart devices and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0052] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a data storage module, a network communication module, a user interface module, and an intelligent dialogue wake-up control program.
[0053] exist Figure 1 In the smart device shown, the network interface 1004 is mainly used for data communication with other devices; the interactive device 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the smart device of the present invention can be set in the smart device, and the smart device calls the smart dialogue wake-up control program stored in the memory 1005 through the processor 1001 and executes the smart dialogue wake-up control method provided in the embodiment of the present invention.
[0054] Optionally, the interactive device may include a microphone and a touch sensor. In addition, the processor 1001 may also connect to the AI agent via an application protocol.
[0055] The specific composition of the module is not limited; optionally, it may include the following components:
[0056] Outer shell: Protects internal components and is usually made of materials such as plastic, metal or wood, with a certain degree of aesthetics and durability.
[0057] Speaker: Converting electrical signals into sound, it is one of the core components of conversational robot devices. The quality and performance of the speaker directly affect the sound quality.
[0058] Microphone: Used to receive user voice commands. Multiple microphones are usually arranged in an array to improve the accuracy of voice recognition.
[0059] Motherboard: It integrates the core components of the conversational robot device, such as the control circuit, audio processing chip, and wireless communication module, and is the control center of the conversational robot device.
[0060] Power supply: Provides power to the conversational robot device, usually through a built-in battery.
[0061] Button and indicator light assembly: This may also include indicator lights, buttons, sensors, and other auxiliary components to enable various functions of the conversational robot device.
[0062] Touch sensor and touch processing module: Capacitive touch sensors can be used to receive user touch interaction, which is then processed by the touch processing module to wake up the device to receive user voice input.
[0063] Application protocols and AI agents: The device interfaces with a large model or AI agent to process semantic understanding and generate responses, which are then converted into voice output.
[0064] This invention provides an intelligent dialogue wake-up control method, referring to... Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of an intelligent dialogue wake-up control method according to the present invention.
[0065] In this embodiment, the intelligent dialogue wake-up control method includes:
[0066] Step S1: Control the microphone to collect voice data and receive touch detection data from the touch sensor to detect buttons;
[0067] The intelligent dialogue wake-up control method described here can be applied to smart devices. Preferably, in this embodiment, the smart device is a human-computer dialogue robot. The microphone can collect voice data in real time. It should be noted that the voice data here refers to the sound data collected by the microphone. In fact, since users do not interact with the smart device via voice at every moment, the voice data here may not include the user's voice information. Specifically, the order of the steps of controlling the microphone to collect voice data and receiving touch detection data from the touch sensor to detect buttons is not limited. Specifically, the above steps can be performed simultaneously.
[0068] Step S2: Calculate the user interaction probability based on the voice data and the touch detection data;
[0069] In this embodiment, the voice data is analyzed to extract semantic data. Words in the semantic data correspond to different user interaction probabilities. Preferably, the higher the relevance of the words to the functions of the smart device, the higher the corresponding user interaction probability. In this embodiment, in addition to the semantic data, the voice data is analyzed. During the analysis, the higher the number of users speaking in the voice data, the lower the corresponding user interaction probability.
[0070] It should be noted that, for touch detection data, when a user touches a smart device, the probability of user interaction is determined to be 100%. If the user interaction probability output by the voice data and the user interaction probability corresponding to the touch detection data are not the same, the data with the highest user interaction probability shall prevail.
[0071] Step S3: When the user interaction probability is higher than the preset probability, determine the interaction target between the user and the smart device based on the voice data and the touch detection data;
[0072] The preset probability here can be set by the user. When the user interaction probability is higher than the preset probability, it is determined that the user needs to interact with the smart device. Therefore, the interaction target between the user and the smart device needs to be determined through the voice data and the touch detection data. The interaction target here can be issuing commands, asking questions, etc. In addition, the interaction target here also corresponds to different interaction methods. Optionally, the interaction methods here generally include a voice interaction mode and a command execution mode. In the voice interaction mode, the smart device can issue corresponding prompt sounds. In the command execution mode, the smart device does not issue corresponding sounds and only executes the corresponding commands from the user.
[0073] Step S4: Control the smart device according to the interaction target.
[0074] In this embodiment, the various operating modules of the smart device can be activated according to the interaction target. For example, after activation, a large language model on a cloud server is invoked via API to send the voice data to the operating server, where a corresponding response is generated in the large language model. Alternatively, local computing resources can be used to generate the corresponding response. Optionally, adjustment commands can be sent via network interfaces, serial communication, etc., to adjust the operation of the appliances connected to the smart device.
[0075] In this embodiment, voice data is collected by controlling the microphone, and touch detection data of the touch sensor is received from the touch sensor. The user interaction probability is calculated based on the voice data and the touch detection data. When the user interaction probability is higher than a preset probability, the interaction target between the user and the smart device is determined based on the voice data and the touch detection data. The smart device is then controlled according to the interaction target. Compared to requiring a wake word to wake up the smart device, this method can identify whether the user needs to interact in multiple ways without using an accurate wake word. This effectively reduces the user's operational difficulty and effectively identifies the user's interaction target, thereby effectively improving the user experience.
[0076] Furthermore, based on the first embodiment, a second embodiment of the intelligent dialogue wake-up control method of the present invention is proposed. In this embodiment, reference is made to... Figure 3 The step of determining the user's interaction target with the smart device based on the voice data and the touch detection data includes:
[0077] Step S31: When the voice data does not include semantic data and the touch detection data is light touch data, the interaction target is determined to be waiting for user input instructions, and the light touch data is that the time for the user to touch the button is less than a first preset time.
[0078] In this embodiment, it is determined that the user interaction probability is higher than a preset probability, thus indicating that the user needs to interact. Specifically, no semantic data is extracted from the voice data, meaning the semantic data is not null. In reality, the voice data often includes environmental data, noise, etc. The touch data here refers to the user waking up the smart device by touching a button. Touch data corresponds to the user's action of touching a button, thereby determining that the interaction target is to wait for user input. Specifically, the microphone records the voice data input within a short period after the user completes the button touch action. Commonly, touch actions lasting less than 2 seconds are not considered touches; however, the first preset time can also be 0.5 seconds, 1 second, etc.
[0079] Step S32: When the touch detection data is long-term touch data, the interaction target is determined to be accurate command input. The accurate command input is to determine the interception time corresponding to the voice data based on the touch detection data, and to determine the control of the smart device based on the interception time and the voice data. The light touch data is the time when the user touches the button for a period of time greater than or equal to a first preset time.
[0080] Specifically, accurate command input refers to the interactive target of voice data input within a precise time range. The interception time corresponding to the voice data is determined based on touch detection data. The interception time includes a start time and a stop time. Based on the start and stop times, the required target voice data is intercepted from the continuous voice data. Semantic analysis is performed to extract the target voice data, and the smart device is controlled based on the semantic data extracted from the target voice data. Preferably, in this embodiment, the interactive target of accurate command input needs to extract voiceprint data to determine the semantic data of users with registered voiceprints. Preferably, when the intercepted target voice data includes semantic data of users with registered voiceprints and semantic data of users with unregistered voiceprints, the smart device is controlled primarily based on the semantic data of users with registered voiceprints. When the intercepted target voice data includes semantic data of multiple users with registered voiceprints, the volume corresponding to the semantic data of the multiple registered users is compared, and the semantic data of the user with the higher volume is executed first. That is, in this embodiment, accurate command input is for accurately controlling the smart device according to a user's command.
[0081] Step S33: When the touch detection data is sliding touch data, determine that the interaction target is to execute a preset control command.
[0082] In this embodiment, preferably, different sliding data can correspond to different preset control commands, which can be set by the user. For example, sliding downwards can adjust the ambient lighting and temperature in the room. The rotational sliding mode corresponds to the command execution mode. In the command execution mode, the smart device does not emit corresponding sounds, but only executes the user's corresponding commands.
[0083] In this embodiment, after recognizing the need to wake up the smart device, there are multiple different operating states, which can effectively correspond to different actual user scenarios and thus effectively improve the user experience.
[0084] Furthermore, based on the first or second embodiment, a third embodiment of the intelligent dialogue wake-up control method of the present invention is proposed. In this embodiment, reference is made to... Figure 4 The step of calculating the probability of user interaction based on the voice data and the touch detection data includes:
[0085] Step S21: Extract target feature data from the speech data. The target feature data includes: semantic feature data, acoustic environment feature data, personnel information feature data, and noise feature data.
[0086] In this embodiment, target feature data is extracted from the speech data through Fourier transform, filtering, and other methods. Commonly, semantic feature data can be determined through automatic speech recognition. Acoustic environment feature data is determined through spectral statistical analysis, and the number of people in the personnel information is determined through voiceprint recognition and emotion recognition. Furthermore, noise feature data is determined by calculating the signal-to-noise ratio.
[0087] Step S22: Generate a first score corresponding to the voice data based on the target feature data, and generate a second score based on the touch detection data;
[0088] In this embodiment, a first score corresponding to the speech data needs to be generated by integrating various target feature data. Different semantic feature data, acoustic environment feature data, personnel information feature data, and noise feature data can each correspond to different scores. Preferably, the second score may only include two different scores.
[0089] Step S23: Calculate the user interaction probability based on the first score and the second score.
[0090] In this embodiment, the highest data value is determined between the first rating and the second rating, and the user interaction probability is calculated based on the highest data value. A common approach is to set a mapping formula to obtain the user interaction probability.
[0091] In this embodiment, target feature data is extracted from the voice data. The target feature data includes semantic feature data, acoustic environment feature data, personnel information feature data, and noise feature data. A first score corresponding to the voice data is generated based on the target feature data, and a second score is generated based on the touch detection data. The user interaction probability is calculated based on the first score and the second score, thereby enabling the identification of whether the user needs to interact with the smart device through multiple dimensions, and thus improving the accuracy of wake-up recognition without setting a wake word.
[0092] Furthermore, the step of generating a first score corresponding to the speech data based on the target feature data includes:
[0093] Generate corresponding feature vector data based on the semantic feature data, the acoustic environment feature data, the personnel information feature data, and the noise feature data;
[0094] The feature vector data is input into a preset scoring and recognition model to obtain the output result of the preset scoring and recognition model;
[0095] The first score is determined based on the output results.
[0096] Optionally, feature data with different dimensions and ranges are uniformly scaled to a similar numerical range to eliminate the impact of dimensional differences on the model and ensure the stability and effectiveness of model training. The semantic feature data, acoustic environment feature data, personnel information feature data, and noise feature data are then concatenated in a preset order. It should be noted that the preset rating recognition model is pre-trained. During training, it primarily learns from historical rating data, using a deep learning model to learn the inherent relationships and decision boundaries between ratings. After the feature vector data is input into the preset rating recognition model, the model obtains the output result. In some embodiments, the output result is directly used to determine the first rating.
[0097] Furthermore, the step of generating a first score corresponding to the speech data based on the target feature data includes:
[0098] When a preset operation instruction exists in the semantic feature data, a first type of rating label corresponding to the voice data is generated;
[0099] When the acoustic environment feature data is a preset scenario, a second type of scoring label corresponding to the speech data is generated;
[0100] A third type of rating label is generated based on the personnel information feature data;
[0101] When the number of noise features is greater than the preset noise data, a fourth type of scoring label corresponding to the speech data is generated;
[0102] The first score is determined based on the scoring markers corresponding to the voice data.
[0103] Optionally, in this embodiment, corresponding labels can be generated for each target feature data, and the labels can be used to determine the corresponding scores through a label-rating mapping table. Different types of labels can correspond to different scores. Specifically, the rating labels can include: a first type of rating label, a second type of rating label, a third type of rating label, and a fourth type of rating label. All rating labels corresponding to the target feature data are counted, and the total score is calculated based on all the rating labels and the label-rating mapping table as the first score.
[0104] In this embodiment, a corresponding rating tag is generated based on the target feature data, and the rating contribution of each feature data is calculated based on the rating tag, thereby avoiding the need for a single wake word recognition and improving the accuracy of determining the first rating.
[0105] Furthermore, based on any of the above embodiments, a third embodiment of the intelligent dialogue wake-up control method of the present invention is proposed, wherein the step of calculating the user interaction probability based on the first score and the second score includes:
[0106] Calculate the total score based on the first score and the second score;
[0107] The user interaction probability is determined based on the total score data and the preset mapping relationship.
[0108] In this embodiment, optionally, the first score and the second score are added together to obtain the total score. Optionally, each point in the total score data corresponds to a 1% probability of user interaction, and when the total score data is higher than 100, the probability of user interaction is determined to be 100%. Of course, this is only one mapping relationship, and other mapping relationships can also be selected in this embodiment.
[0109] Furthermore, prior to the step of controlling the microphone to collect voice data, the method further includes:
[0110] Control the real-time orientation data of the gravity sensor detection device;
[0111] When the real-time orientation data is the preset orientation data, the step of controlling the microphone to collect voice data is executed.
[0112] In this embodiment, to help users determine whether they can directly interact with the smart device, a convenient interaction method can be used to determine whether the smart device has enabled its convenient interaction by real-time orientation.
[0113] Furthermore, embodiments of the present invention also propose a dialogue system, the dialogue system comprising:
[0114] The acquisition module is used to control the microphone to acquire voice data and receive touch detection data from the touch sensor to detect buttons;
[0115] The calculation module is used to calculate the probability of user interaction based on the voice data and the touch detection data;
[0116] The recognition module is used to determine the interaction target between the user and the smart device based on the voice data and the touch detection data when the user interaction probability is higher than a preset probability.
[0117] An execution module is used to control the smart device according to the interaction target.
[0118] Furthermore, this invention also proposes an intelligent device, comprising: a memory, a processor, and an intelligent dialogue wake-up control program stored in the memory and executable on the processor. The intelligent dialogue wake-up control program is configured to implement the steps of the intelligent dialogue wake-up control method described above. Preferably, the intelligent device can be an artificial intelligence dialogue robot, the hardware of which can refer to... Figure 5 and Figure 6 . Reference Figure 5 and Figure 6 , Figure 5 and Figure 6 These are the front and top views of the AI conversational robot hardware.
[0119] Furthermore, embodiments of the present invention also propose a storage medium storing an intelligent dialogue wake-up control program, wherein the intelligent dialogue wake-up control program, when executed by a processor, implements the steps of the intelligent dialogue wake-up control method described above.
[0120] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0121] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0122] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0123] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for intelligent dialogue wake-up control, characterized in that, The intelligent dialogue wake-up control method includes the following steps: Control the microphone to collect voice data and receive touch detection data from the touch sensor to detect buttons; Calculate the probability of user interaction based on the voice data and the touch detection data; When the probability of user interaction is higher than the preset probability, the interaction target between the user and the smart device is determined based on the voice data and the touch detection data. Control the smart device according to the interaction target; The step of determining the user's interaction target with the smart device based on the voice data and the touch detection data includes: When the voice data does not include semantic data, and the touch detection data is light touch data, the interaction target is determined to be waiting for user input instructions, and the light touch data is when the user touches the button for less than a first preset time; When the touch detection data is long-term touch data, the interaction target is determined to be accurate command input. The accurate command input is determined by the interception time corresponding to the voice data based on the touch detection data, and the control of the smart device is determined based on the interception time and the voice data. The light touch data is the time when the user touches the button for a period of time greater than or equal to a first preset time. When the touch detection data is swipe touch data, the interaction target is determined to be to execute a preset control command.
2. The intelligent dialogue wake-up control method as described in claim 1, characterized in that, The step of calculating the probability of user interaction based on the voice data and the touch detection data includes: Extract target feature data from the speech data, the target feature data including: semantic feature data, acoustic environment feature data, personnel information feature data, and noise feature data; A first score is generated based on the target feature data, and a second score is generated based on the touch detection data. The user interaction probability is calculated based on the first rating and the second rating.
3. The intelligent dialogue wake-up control method as described in claim 2, characterized in that, The step of generating a first score corresponding to the speech data based on the target feature data includes: Generate corresponding feature vector data based on the semantic feature data, the acoustic environment feature data, the personnel information feature data, and the noise feature data; The feature vector data is input into a preset scoring and recognition model to obtain the output result of the preset scoring and recognition model; The first score is determined based on the output results.
4. The intelligent dialogue wake-up control method as described in claim 2, characterized in that, The step of generating a first score corresponding to the speech data based on the target feature data includes: When a preset operation instruction exists in the semantic feature data, a first type of rating label corresponding to the voice data is generated; When the acoustic environment feature data is a preset scenario, a second type of scoring label corresponding to the speech data is generated; A third type of rating label is generated based on the personnel information feature data; When the number of noise features is greater than the preset noise data, a fourth type of scoring label corresponding to the speech data is generated; The first score is determined based on the scoring markers corresponding to the voice data.
5. The intelligent dialogue wake-up control method as described in claim 2, characterized in that, The step of calculating the user interaction probability based on the first score and the second score includes: Calculate the total score based on the first score and the second score; The user interaction probability is determined based on the total score data and the preset mapping relationship.
6. The intelligent dialogue wake-up control method as described in any one of claims 1 to 5, characterized in that, Before the step of controlling the microphone to collect voice data, the method further includes: Control the real-time orientation data of the gravity sensor detection device; When the real-time orientation data is the preset orientation data, the step of controlling the microphone to collect voice data is executed.
7. A dialogue system, characterized in that, The dialogue system includes: The acquisition module is used to control the microphone to acquire voice data and receive touch detection data from the touch sensor to detect buttons; The calculation module is used to calculate the probability of user interaction based on the voice data and the touch detection data; The recognition module is configured to determine the interaction target between the user and the smart device based on the voice data and the touch detection data when the user interaction probability is higher than a preset probability; the step of determining the interaction target between the user and the smart device based on the voice data and the touch detection data includes: When the voice data does not include semantic data, and the touch detection data is light touch data, the interaction target is determined to be waiting for user input instructions, and the light touch data is when the user touches the button for less than a first preset time; When the touch detection data is long-term touch data, the interaction target is determined to be accurate command input. The accurate command input is determined by the interception time corresponding to the voice data based on the touch detection data, and the control of the smart device is determined based on the interception time and the voice data. The light touch data is the time when the user touches the button for a period of time greater than or equal to a first preset time. When the touch detection data is swipe touch data, the interaction target is determined to be to execute a preset control command; An execution module is used to control the smart device according to the interaction target.
8. A smart device, characterized in that, The smart device includes: a memory, a processor, and a smart dialogue wake-up control program stored in the memory and executable on the processor, the smart dialogue wake-up control program being configured to implement the steps of the smart dialogue wake-up control method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium stores an intelligent dialogue wake-up control program, which, when executed by a processor, implements the steps of the intelligent dialogue wake-up control method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Predictive pre-recording of audio for voice input
US20110238191A1