A method, apparatus, storage medium and electronic device of a voice wake-up device
By collecting, separating, and analyzing ambient sound from terminal devices, identifying wake words, and generating response commands, the problem of false wake-up of voice wake-up devices is solved, improving the accuracy of wake word recognition and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HAIER YOUJIA INTELLIGENT TECH (BEIJING) CO LTD
- Filing Date
- 2023-03-30
- Publication Date
- 2026-05-19
AI Technical Summary
Existing voice wake-up devices are prone to false wake-ups, which affects user experience.
By collecting ambient sound from the terminal device, separating and identifying valid sounds, analyzing wake words, and using voice processing strategies to generate response commands, the target device can be controlled.
Reduce the possibility of false wake-ups, improve the accuracy of wake word recognition, and enhance the user experience.
Smart Images

Figure CN116564285B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and in particular to a method, apparatus, storage medium and electronic device for a voice wake-up device. Background Technology
[0002] Currently, with the rapid development of artificial intelligence and Internet of Things (IoT) technologies, voice interaction technology has been widely applied in information retrieval, business processing, entertainment, and other scenarios, making people's lives more convenient. For example, smart air conditioners, smart speakers, smart TVs, smart cars, and so on can gradually be woken up by artificial intelligence. The main wake-up method now is still through wake words. For example, if the wake word "Hai Hai" is set for a smart air conditioner, the air conditioner can listen to external sounds in real time. If it recognizes the voice input of "Hai Hai," it will wake up the air conditioner; that is, the device is woken up through a wake word.
[0003] However, waking up the device using a wake word can easily lead to false wake-ups. That is, if the user says something else and does not say "Haihai", but the recognition system misidentifies it as "Haihai", it will result in a false wake-up and seriously affect the user experience.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] In view of this, this application provides a method, apparatus, storage medium, and electronic device for voice wake-up devices, which can reduce the possibility of false wake-up, improve the accuracy of wake-up word recognition, and thus enhance the user's intelligent experience.
[0006] In a first aspect, this application provides a method for voice-activated device, comprising:
[0007] Collect ambient sound from the environment where the terminal device is located;
[0008] The ambient sound is processed to determine the valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake word, and the wake word includes at least: a quick wake word, a full wake word, and an invalid wake word;
[0009] The wake word is analyzed and processed;
[0010] When the wake word matches the quick wake word, the effective sound is processed using a preset voice processing strategy to generate target semantic information, and a response command is generated based on the target semantic information.
[0011] Based on the response instruction, the target device corresponding to the response instruction is made to perform the operation of the response instruction, wherein the target device includes at least the terminal device.
[0012] Preferably, according to the method for a voice wake-up device provided in this application, after the step of analyzing and processing the wake-up word, the method further includes:
[0013] If the wake word is a quick wake word, the time of the quick wake word will be determined as the wake word time;
[0014] Obtain the current time and determine the waiting period based on the current time and the wake word time.
[0015] Preferably, according to the method for a voice wake-up device provided in this application, after the step of determining a waiting period based on the current time and the wake-up word time, the method further includes:
[0016] If the waiting time period is less than a preset time period, determine whether to generate full wake-up event information;
[0017] In the case of generating the full wake-up event information, the quick wake-up word is marked to obtain quick wake-up mark information.
[0018] Preferably, according to the method for a voice wake-up device provided in this application, after the step of processing the quick wake-up word marker to obtain quick wake-up marker information, the method further includes:
[0019] The full wake-up event information, the quick wake-up flag information, and the target device parameter information are concatenated to generate target request information; and
[0020] The first audio corresponding to the full wake-up event and the second audio corresponding to the valid sound are spliced together to generate the target audio.
[0021] Preferably, according to the method for a voice wake-up device provided in this application,
[0022] The step of generating the response instruction based on the target semantic information includes:
[0023] The target semantic information is extracted and processed to obtain the target function corresponding to the target semantic information;
[0024] The target function is matched with multiple execution functions in a preset function implementation library;
[0025] Based on a first matching result where the target function matches at least one of the execution functions, a response instruction to implement the target function is generated.
[0026] Preferably, according to the method for a voice wake-up device provided in this application, after the step of analyzing and processing the wake-up word, the method further includes:
[0027] When the wake word matches the full set of wake words, the valid sound is processed using a preset response strategy to generate the response command, so that the target device executes the operation of the response command; and
[0028] If the wake word is matched with at least one of the quick wake word and / or the full wake word and no match is found, the wake word is determined to be an invalid wake word. If the wake word is determined to be an invalid wake word, the ambient sound of the environment in which the terminal device is located is re-acquired.
[0029] Preferably, according to the method for a voice wake-up device provided in this application, the step of processing the ambient sound to determine the valid sound includes:
[0030] The ambient sound is subjected to sound separation processing to obtain human voice sound;
[0031] The human voice is input into a voice recognition model for processing, and the effective voice is output. The voice recognition model is obtained by training on human voice samples.
[0032] Preferably, according to the method for a voice wake-up device provided in this application,
[0033] The operation of causing the target device corresponding to the response instruction to execute the response instruction based on the response instruction includes:
[0034] The terminal information corresponding to the terminal device and the target information of the target device are compared and processed.
[0035] Based on a first comparison result showing that the terminal information and the target information are identical, the terminal device is instructed to execute the response instruction.
[0036] Preferably, according to the method for a voice wake-up device provided in this application,
[0037] After the step of comparing the terminal information corresponding to the terminal device and the target information of the target device, the method further includes:
[0038] Based on a second comparison result showing that the terminal information and the target information are different, the response instruction is sent from the terminal device to the target device, so that the target device executes the operation of the response instruction according to the received response instruction.
[0039] Secondly, this application also provides a device for a voice wake-up device, comprising:
[0040] The acquisition module is used to collect ambient sound from the environment where the terminal device is located.
[0041] The determination module is used to process the ambient sound and determine the valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake word, and the wake word includes at least: a quick wake word, a full wake word, and an invalid wake word;
[0042] The analysis module is used to analyze and process the wake word;
[0043] The generation module is used to process the effective sound using a preset voice processing strategy when the wake word matches the quick wake word, generate target semantic information, and generate a response instruction based on the target semantic information.
[0044] A response module is configured to, based on the response instruction, cause a target device corresponding to the response instruction to perform the operation of the response instruction, wherein the target device includes at least the terminal device.
[0045] Thirdly, this application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute, through the computer program, a method for implementing any of the above-described voice wake-up devices.
[0046] Fourthly, this application also provides a computer-readable storage medium comprising a stored program, wherein the program, when executed, performs a method for implementing any of the voice wake-up devices described above.
[0047] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the above-described voice wake-up devices.
[0048] This application provides a method, apparatus, storage medium, and electronic device for voice wake-up of a device. The method involves: collecting ambient sound from the environment where the terminal device is located; processing the ambient sound to determine valid sounds, wherein the valid text information corresponding to the valid sounds includes at least a wake-up word; analyzing and processing the wake-up word; if the wake-up word matches a quick wake-up word, processing the valid sound using a preset voice processing strategy to generate target semantic information, and generating a response command based on the target semantic information; and, based on the response command, causing a target device corresponding to the response command to execute the operation of the response command, wherein the target device includes at least the terminal device. This reduces the possibility of false wake-ups, improves the accuracy of wake-up word recognition, and thus enhances the user's intelligent experience. Attached Figure Description
[0049] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0051] Figure 1 This is a schematic diagram of the hardware environment for a method of voice wake-up device provided in this application;
[0052] Figure 2 This is one of the flowcharts illustrating a method for a voice wake-up device provided in this application;
[0053] Figure 3 This is a second flowchart illustrating a method for a voice wake-up device provided in this application;
[0054] Figure 4 This is a schematic diagram of the structure of a voice wake-up device provided in this application;
[0055] Figure 5 This is a schematic diagram of the electronic device provided in this application. Detailed Implementation
[0056] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0057] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0058] First, the technical terms used in the embodiments of this application will be explained:
[0059] Voice technology, including speech recognition and voice interaction, is an important area of artificial intelligence.
[0060] Speech recognition is a technology that enables machines to convert speech signals into corresponding text or commands through recognition and understanding processes. It mainly includes three aspects: feature extraction technology, pattern matching criteria, and model training technology.
[0061] Voice interaction is a technology that allows machines and users to interact, communicate, and exchange information using voice as the information carrier. Compared to traditional human-computer interaction, it has the advantages of being convenient, fast, and providing a high level of user comfort.
[0062] Natural Language Processing (NLP) is a science that studies computer systems, especially their software systems, that can effectively achieve natural language communication. It is an important direction in the fields of computer science and artificial intelligence.
[0063] Deep learning (DL) is a new research direction in the field of machine learning (ML). It is a science that learns the inherent patterns and representation levels of sample data, enabling machines to have analytical and learning capabilities like humans, and to recognize data such as text, images, and sound. It is widely used in speech and image recognition.
[0064] The following is combined with Figures 1-5 This application describes a method, apparatus, storage medium, and electronic device for a voice wake-up device. The embodiments provided in this application can reduce the possibility of false wake-ups, improve the accuracy of wake-up word recognition, and thus enhance the user's intelligent experience.
[0065] According to one aspect of the embodiments of this application, a method for voice-activated device wake-up is provided. This method is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-described method for voice-activated device wake-up can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0066] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0067] like Figure 2As shown, it is one of the implementation flow diagrams of a method for waking up a device with voice provided in the embodiments of this application. A method for waking up a device with voice may include, but is not limited to, steps S100 to S500.
[0068] S100 collects ambient sound from the environment where the terminal device is located;
[0069] S200, process the ambient sound to determine the valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake word, and the wake word includes at least: a quick wake word, a full wake word, and an invalid wake word;
[0070] S300, the wake word is analyzed and processed;
[0071] S400, when the wake word matches the quick wake word, the effective sound is processed using a preset voice processing strategy to generate target semantic information, and a response command is generated based on the target semantic information;
[0072] S500, based on the response instruction, to cause the target device corresponding to the response instruction to perform the operation of the response instruction, wherein the target device includes at least the terminal device.
[0073] In step S100 of some embodiments, ambient sound of the environment in which the terminal device is located is collected.
[0074] It should be noted that the execution subject of the voice wake-up device method in the embodiments of this application can be a hardware device with data information processing capabilities and / or the necessary software to drive the hardware device to work.
[0075] Optionally, the executing entity may include, but is not limited to, workstations, servers, computers, user terminals, and other intelligent devices. User terminals include, but are not limited to, mobile phones, computers, intelligent voice interaction devices, smart home appliances, and in-vehicle terminals.
[0076] Optionally, the terminal device has a voice acquisition device, which can be a microphone, a microphone array, etc. The microphone array of the voice acquisition device is a circular 6-microphone array, and the voice acquisition device with a circular 6-microphone array is used to acquire the ambient sound of the environment in which the terminal device is located.
[0077] It should be noted that the terminal device can be a smart speaker, smart tablet, smartphone, etc.
[0078] The circular 6-microphone array achieves a far-field pickup distance of up to 5 meters using far-field recognition and noise reduction technology. Within 5 meters, the recognition rate is comparable to near-field performance, exceeding 90%. The circular 6-microphone array effectively adapts to the characteristics of far-field pickup, achieving 360° voice signal acquisition and accurately determining the speaker's direction through sound source localization. Compared to the previous circular 5-microphone array, the 6+0 microphone array forms six pickup beams, each corresponding to a 60° range, which is finer than the 90° range beam of the 5-microphone array, resulting in more accurate sound source angle localization.
[0079] The circular 6-microphone array configuration is widely used. The circular 6-microphone array has a prototype layout, with six microphones evenly distributed around the circumference, and one microphone placed equidistantly in the same shape, with a radius of 35mm. Compared to the 4+1 five-microphone array, the circular 6-microphone array represents a significant breakthrough in configuration. The empty space in the middle can be utilized in later product designs, making it more widely applicable.
[0080] The circular 6-microphone array supports voice wake-up and continuous wake-up with a success rate exceeding 90%. The voice wake-up function allows users to change the interaction state by speaking keywords, such as from sleep to waiting, or from interactive to waiting. It also supports continuous wake-up, allowing for multiple wake-ups at any time from any angle, outputting a ring beam for audio recognition, with a wake-up success rate exceeding 90%.
[0081] A circular 6-microphone array echo cancellation system can eliminate the echo caused by the microphone picking up sound from the speaker during simultaneous playback and recording, thus affecting recording quality. By using a reference signal and filtering the recording signal, this echo cancellation technology can block out the speaker sound, ensuring accurate sound recognition even when playback and recording are happening simultaneously.
[0082] The circular 6-microphone array supports voice interruption; the device can still be woken up while it is broadcasting, thus achieving the interruption effect.
[0083] The circular 6-microphone array has efficient suppression of steady-state and dynamic noise, and can easily cope with noisy environments. The algorithm of the 6-microphone circular array will enhance the sound within the beam range and weaken the sound outside the beam, so that the enhancement effect of the sound source angle is better and the noise suppression effect is better, thereby improving the recognition accuracy.
[0084] In step S200 of some embodiments, the ambient sound is processed to determine a valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake word.
[0085] It should be noted that ambient sound includes at least human voices, scene sounds, and environmental noise.
[0086] The specific execution steps can be as follows: performing sound separation processing on the environmental sound to obtain human voice sound, inputting the human voice sound into a sound recognition model for processing, and outputting the effective sound, wherein the sound recognition model is obtained by training on human voice sound samples.
[0087] Human voices, scene sounds, and environmental noise correspond to different audio frequencies. The audio frequency information of these sounds is pre-stored in an audio database. By matching the audio frequencies of the collected environmental sounds with those in the audio database, human voices, scene sounds, and environmental noise can be distinguished. Then, a preset sound separation algorithm is used to separate the human voices, scene sounds, and environmental noise, finally filtering out the human voices.
[0088] The human voice is then input into the voice recognition model for processing, and the effective voice is output.
[0089] It should be noted that a large number of human voice samples are input into a pre-set neural network for training to obtain a well-trained voice recognition model.
[0090] The sound recognition model can accurately output valid sounds, thus avoiding interference from invalid sounds, such as human voices played on a television or other smart speakers or smart devices.
[0091] It should be noted that the valid text information corresponding to the valid sound includes at least a wake word.
[0092] The valid sounds are preprocessed to obtain valid text information corresponding to the valid sounds. The valid text information includes at least the wake word.
[0093] It should be noted that wake words include at least quick wake words, full wake words, and invalid wake words.
[0094] The quick wake word in this embodiment is a two-syllable wake word, not a conventional four-syllable wake word.
[0095] A full set of wake words includes at least: a quick wake word and a wake command word.
[0096] For example, the quick wake-up word is "Haihai" or "Erer". The characters in the word can be replaced as needed, but the number of words is fixed at two syllables.
[0097] The full wake-up phrases are "Haihai, open XX", "Haihai, close XX", etc.
[0098] An invalid wake-up word is determined when the wake-up word is not matched against at least one of the quick wake-up words and / or the full set of wake-up words. If the wake-up word is determined to be invalid, the ambient sound of the terminal device's environment is re-acquired.
[0099] In step S300 of some embodiments, the wake word is analyzed and processed.
[0100] It should be noted that by analyzing the wake word, whether it is a quick wake word, a full wake word, or an invalid wake word, we can obtain the first analysis result if the wake word is a quick wake word, the second analysis result if the wake word is a full wake word, and the third analysis result if the wake word is an invalid wake word.
[0101] In step S400 of some embodiments, when the wake word matches the quick wake word, the effective sound is processed using a preset voice processing strategy to generate target semantic information, and a response instruction is generated based on the target semantic information.
[0102] Understandably, after analyzing and processing the wake word, if the wake word matches the quick wake word, the time of the quick wake word will be determined as the wake word time, and the current time will be obtained. Finally, a waiting period will be determined based on the current time and the wake word time.
[0103] If the waiting time period is less than a preset time period, determine whether to generate full wake-up event information. If the full wake-up event information is generated, process the quick wake-up word to obtain quick wake-up mark information.
[0104] The full wake-up event information, the quick wake-up marker information, and the device parameter information of the target device are concatenated to generate target request information. The first audio corresponding to the full wake-up event and the second audio corresponding to the valid sound are concatenated to generate target audio.
[0105] It should be noted that the speech processing strategy includes at least: a speech recognition strategy and a semantic understanding strategy;
[0106] The target audio is processed using the speech recognition strategy to generate target text information corresponding to the target audio. The target request information and the target text information are then processed using the semantic understanding strategy to generate target semantic information. Finally, the response instruction is generated based on the target semantic information.
[0107] It should be further noted that the speech recognition strategy is a strategy that utilizes the ASR algorithm to process the target audio and automatically generate target text information.
[0108] ASR stands for Automatic Speech Recognition, a technology that converts human speech into text.
[0109] Semantic understanding strategies utilize natural language understanding (NLU), which is a collective term for all methods, models, or tasks that support machine understanding of text content. NLU plays a crucial role in text information processing systems and is an essential module for recommendation, question answering, and search systems. Using natural semantic understanding strategies, target semantic information can be directly determined from the target request information and the target text information.
[0110] More specifically, the target request information and the target text information are processed by machine translation, text mining, and information extraction to generate target semantic information.
[0111] Machine translation processing: Automatically translates the input target audio into text in a different language.
[0112] Text mining processing includes text clustering, classification, information extraction, summarization, sentiment analysis, and visualization and interactive interface representation of the mined information and knowledge.
[0113] Information extraction processing: Extracting important information from a given text.
[0114] After determining the target semantic information, the target semantic information is extracted and processed to obtain the target function corresponding to the target semantic information. The target function is then matched with multiple execution functions in a preset function implementation library. Based on the first matching result of the target function matching with at least one of the execution functions, the response instruction that implements the target function is generated.
[0115] For example, if the target function corresponding to the target semantic information is "turn on the air conditioner", then the target function of "turn on the air conditioner" will be matched with multiple execution functions in the preset function implementation library to obtain a first matching result that matches the target function with at least one of the execution functions. Based on the first matching result that is successfully matched, the response instruction that implements the target function will be automatically generated.
[0116] The program matches the target function "turn on the air conditioner" with multiple execution functions in a preset function implementation library. If a second matching result is obtained where the target function does not match any of the execution functions, a no-response instruction is generated based on the second matching result, and the program terminates. No response operation is performed based on the no-response instruction.
[0117] If the target function is A, the multiple execution functions in the function implementation library are A, B, C, etc.
[0118] The target text information corresponding to target function A is "turn on the air conditioner".
[0119] The corresponding text information in the function implementation library might be multiple text messages such as "start the air conditioner" or "turn on the air conditioner". In this case, it is also considered a successful match. The messages do not necessarily have to be exactly the same, as long as the functions they perform are the same.
[0120] For example, if the target text information corresponding to the target function P is "start the car", but the text information pre-stored in the function implementation library does not have the execution function of "start the car", then a second matching result of mismatch is obtained.
[0121] In step S500 of some embodiments, the target device corresponding to the response instruction is caused to perform the operation of the response instruction based on the response instruction.
[0122] First, it should be noted that the target device includes at least the terminal device.
[0123] The target device is a device that needs to be controlled based on human voice commands, such as smart air conditioners, smart range hoods, smart refrigerators, smart ovens, smart stoves, smart washing machines, smart water heaters, etc.
[0124] The terminal devices are smart devices with a circular 6-microphone array, such as smart speakers and smart tablets.
[0125] This is because not every smart device has a circular 6-microphone array, such as a smart door lock. The target device can be a smart device with a circular 6-microphone array or a smart device without a circular 6-microphone array.
[0126] First, obtain the terminal information of the terminal device and the target information of the target device. Then, compare the terminal information of the terminal device and the target information of the target device. Specifically, this can involve comparing the ID of the terminal device and the ID of the target device.
[0127] Based on a first comparison result showing that the terminal information and the target information are identical, the terminal device is instructed to execute the response instruction.
[0128] For example, at this time, the terminal device and the target device are the same smart device, which is a smart speaker.
[0129] Based on a second comparison result showing that the terminal information and the target information are different, the response instruction is sent from the terminal device to the target device, so that the target device executes the operation of the response instruction according to the received response instruction.
[0130] For example, in this case, the terminal device is a smart speaker, and the target device is a smart water heater.
[0131] In some embodiments of this application, after the step of analyzing the wake word, the method includes:
[0132] When the wake word matches the full set of wake words, the valid sound is processed using a preset response strategy to generate the response command, so that the target device executes the operation of the response command.
[0133] For example, if the full wake-up word is "Haihai, turn on the air conditioner", the server directly calls the preset response strategy to perform natural language processing on the valid voice and parses out the corresponding control text information. Based on the control text information, it automatically generates a response command to turn on the air conditioner, thereby causing the target device to execute the operation of turning on the air conditioner.
[0134] In some embodiments of this application, after the step of analyzing and processing the wake word, if the wake word is matched with at least one of the quick wake word and / or the full set of wake words and no match is found, the wake word is determined to be the invalid wake word.
[0135] If the wake word is determined to be invalid, the ambient sound of the environment in which the terminal device is located is re-collected, thus avoiding waking up and controlling the device due to invalid sound and improving the accuracy of voice wake-up.
[0136] Figure 3 This is a second flowchart illustrating a method for a voice wake-up device provided in this application. It involves collecting ambient sound from the environment where the terminal device is located, determining whether a full wake-up event should be generated based on the ambient sound, and storing the corresponding audio information. If a full wake-up event is generated, a response command is directly generated according to a preset response strategy so that the target device can execute the operation of the response command.
[0137] If a quick wake-up event is generated, the audio information of the quick wake-up event is first stored. Then, it is determined whether a full wake-up event was captured within time T. If so, a quick wake-up flag is added to generate quick wake-up flag information. A concatenation process is then performed, specifically concatenating the full wake-up event information, the quick wake-up flag information, and the device parameter information of the target device to generate target request information. The first audio corresponding to the full wake-up event and the second audio corresponding to the valid sound are then concatenated to generate the target audio.
[0138] The target audio is processed using ASR technology to generate target text information. Then, NLU technology is used to process the target text information and target request information to generate target semantic information. The target semantic information is then extracted and processed to obtain the target function corresponding to it. This target function is matched against multiple execution functions in a preset function implementation library. Based on a first matching result between the target function and at least one of the execution functions, a response instruction to implement the target function is generated. Based on the response instruction, the target device corresponding to the response instruction executes the operation of the response instruction.
[0139] This application provides a method, apparatus, storage medium, and electronic device for voice wake-up of a device. The method involves: collecting ambient sound from the environment where the terminal device is located; processing the ambient sound to determine valid sounds, wherein the valid text information corresponding to the valid sounds includes at least a wake-up word; analyzing and processing the wake-up word; if the wake-up word matches a quick wake-up word, processing the valid sound using a preset voice processing strategy to generate target semantic information, and generating a response command based on the target semantic information; and, based on the response command, causing a target device corresponding to the response command to execute the operation of the response command, wherein the target device includes at least the terminal device. This reduces the possibility of false wake-ups, improves the accuracy of wake-up word recognition, and thus enhances the user's intelligent experience.
[0140] The following describes an apparatus for a voice wake-up device provided in this application. The apparatus for a voice wake-up device described below and the method for a voice wake-up device described above can be referred to in correspondence.
[0141] like Figure 4 The diagram shown is a structural schematic of a voice wake-up device provided in this application. The voice wake-up device includes:
[0142] The acquisition module 410 is used to acquire ambient sound from the environment where the terminal device is located.
[0143] The determination module 420 is used to process the ambient sound and determine the valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake word, and the wake word includes at least: a quick wake word, a full wake word, and an invalid wake word;
[0144] Analysis module 430 is used to analyze and process the wake word;
[0145] The generation module 440 is used to process the effective sound using a preset voice processing strategy when the wake word matches the quick wake word, generate target semantic information, and generate a response instruction based on the target semantic information.
[0146] The response module 450 is configured to, based on the response instruction, cause a target device corresponding to the response instruction to perform the operation of the response instruction, wherein the target device includes at least the terminal device.
[0147] Preferably, the apparatus for a voice wake-up device according to this application is further configured to, when the wake-up word is a quick wake-up word, determine the time of determining the quick wake-up word as the wake-up word time;
[0148] Obtain the current time and determine the waiting period based on the current time and the wake word time.
[0149] Preferably, the apparatus for a voice wake-up device according to this application is further configured to determine whether to generate full wake-up event information when the waiting time period is less than a preset time period;
[0150] In the case of generating the full wake-up event information, the quick wake-up word is marked to obtain quick wake-up mark information.
[0151] Preferably, the apparatus for a voice wake-up device according to this application is further configured to concatenate the full wake-up event information, the quick wake-up marker information, and the device parameter information of the target device to generate target request information; and
[0152] The first audio corresponding to the full wake-up event and the second audio corresponding to the valid sound are spliced together to generate the target audio.
[0153] Preferably, in the device for a voice wake-up device provided in this application, the generation module 440 is used to extract and process the target semantic information to obtain a target function corresponding to the target semantic information;
[0154] The target function is matched with multiple execution functions in a preset function implementation library;
[0155] Based on a first matching result where the target function matches at least one of the execution functions, a response instruction to implement the target function is generated.
[0156] Preferably, the apparatus for a voice wake-up device according to this application is further configured to, when the wake-up word matches the full set of wake-up words, process the valid sound using a preset response strategy to generate the response command, so that the target device executes the operation of the response command; and
[0157] If the wake word is matched with at least one of the quick wake word and / or the full wake word and no match is found, the wake word is determined to be an invalid wake word. If the wake word is determined to be an invalid wake word, the ambient sound of the environment in which the terminal device is located is re-acquired.
[0158] Preferably, in the device for a voice wake-up device provided in this application, the determining module 420 is used to perform sound separation processing on the ambient sound to obtain a human voice.
[0159] The human voice is input into a voice recognition model for processing, and the effective voice is output. The voice recognition model is obtained by training on human voice samples.
[0160] Preferably, in the device for a voice wake-up device provided in this application, the response module 450 is used to compare and process the terminal information corresponding to the terminal device and the target information of the target device;
[0161] Based on a first comparison result showing that the terminal information and the target information are identical, the terminal device is instructed to execute the response instruction.
[0162] Preferably, the apparatus for a voice wake-up device provided in this application is further configured to send the response instruction from the terminal device to the target device based on a second comparison result showing that the terminal information and the target information are different, so that the target device executes the operation of the response instruction according to the received response instruction.
[0163] This application provides a method, apparatus, storage medium, and electronic device for voice wake-up of a device. The method involves: collecting ambient sound from the environment where the terminal device is located; processing the ambient sound to determine valid sounds, wherein the valid text information corresponding to the valid sounds includes at least a wake-up word; analyzing and processing the wake-up word; if the wake-up word matches a quick wake-up word, processing the valid sound using a preset voice processing strategy to generate target semantic information, and generating a response command based on the target semantic information; and, based on the response command, causing a target device corresponding to the response command to execute the operation of the response command, wherein the target device includes at least the terminal device. This reduces the possibility of false wake-ups, improves the accuracy of wake-up word recognition, and thus enhances the user's intelligent experience.
[0164] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a method for a voice wake-up device. This method includes: acquiring ambient sound from the environment where the terminal device is located; processing the ambient sound to determine valid sounds, wherein the valid text information corresponding to the valid sounds includes at least a wake-up word; analyzing and processing the wake-up word; if the wake-up word matches a shortcut wake-up word, processing the valid sound using a preset voice processing strategy to generate target semantic information, and generating a response instruction based on the target semantic information; and, based on the response instruction, causing a target device corresponding to the response instruction to execute the operation of the response instruction, wherein the target device includes at least the terminal device.
[0165] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0166] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a method for a voice wake-up device provided by the methods described above. The method includes: collecting ambient sound of the environment in which the terminal device is located; processing the ambient sound to determine valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake-up word; analyzing and processing the wake-up word; when the wake-up word matches the quick wake-up word, processing the valid sound using a preset voice processing strategy to generate target semantic information, and generating a response instruction based on the target semantic information; and, based on the response instruction, causing a target device corresponding to the response instruction to execute the operation of the response instruction, wherein the target device includes at least the terminal device.
[0167] In another aspect, this application also provides a computer-readable storage medium, the computer-readable storage medium including a stored program, wherein the program, when running, executes a method for a voice wake-up device provided by the methods described above, the method comprising: acquiring ambient sound of the environment in which the terminal device is located; processing the ambient sound to determine valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake-up word; analyzing and processing the wake-up word; when the wake-up word matches the quick wake-up word, processing the valid sound using a preset voice processing strategy to generate target semantic information, and generating a response instruction based on the target semantic information; and, based on the response instruction, causing a target device corresponding to the response instruction to execute the operation of the response instruction, wherein the target device includes at least the terminal device.
[0168] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0169] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for waking up a device by voice, characterized in that, include: Collect ambient sound from the environment where the terminal device is located; The ambient sound is processed to determine the valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake word, and the wake word includes at least: a quick wake word, a full wake word, and an invalid wake word; The wake word is analyzed and processed, and if the wake word is a quick wake word, the time of the quick wake word is determined as the wake word time; the current time is obtained, and the waiting time is determined based on the current time and the wake word time; If the waiting time is less than a preset time, determine whether to generate full wake-up event information; if the full wake-up event information is generated, process the quick wake-up word to obtain quick wake-up mark information; The full wake-up event information, the quick wake-up marker information, and the device parameter information of the target device are concatenated to generate target request information; and the first audio corresponding to the full wake-up event and the second audio corresponding to the valid sound are concatenated to generate target audio. When the wake word matches the quick wake word, the effective sound is processed using a preset voice processing strategy to generate target semantic information, and a response command is generated based on the target semantic information. Based on the response instruction, the target device corresponding to the response instruction is made to perform the operation of the response instruction, wherein the target device includes at least the terminal device.
2. The method for voice wake-up device according to claim 1, characterized in that, The step of generating the response instruction based on the target semantic information includes: The target semantic information is extracted and processed to obtain the target function corresponding to the target semantic information; The target function is matched with multiple execution functions in a preset function implementation library; Based on a first matching result where the target function matches at least one of the execution functions, a response instruction to implement the target function is generated.
3. The method for voice wake-up device according to claim 1, characterized in that, After the step of analyzing and processing the wake word, the method further includes: When the wake word matches the full set of wake words, the valid sound is processed using a preset response strategy to generate the response command, so that the target device executes the operation of the response command; and If the wake word is matched with at least one of the quick wake word and / or the full wake word and no match is found, the wake word is determined to be an invalid wake word. If the wake word is determined to be an invalid wake word, the ambient sound of the environment in which the terminal device is located is re-acquired.
4. The method for voice wake-up device according to claim 1, characterized in that, The process of processing the ambient sound to determine valid sounds includes: The ambient sound is subjected to sound separation processing to obtain human voice sound; The human voice is input into a voice recognition model for processing, and the effective voice is output. The voice recognition model is obtained by training on human voice samples.
5. The method for voice wake-up device according to claim 1, characterized in that, The operation of causing the target device corresponding to the response instruction to execute the response instruction based on the response instruction includes: The terminal information corresponding to the terminal device and the target information of the target device are compared and processed. Based on a first comparison result showing that the terminal information and the target information are identical, the terminal device is instructed to execute the response instruction.
6. The method for voice wake-up device according to claim 5, characterized in that, After the step of comparing the terminal information corresponding to the terminal device and the target information of the target device, the method further includes: Based on a second comparison result showing that the terminal information and the target information are different, the response instruction is sent from the terminal device to the target device, so that the target device executes the operation of the response instruction according to the received response instruction.
7. A device for a voice wake-up device, characterized in that, include: The acquisition module is used to collect ambient sound from the environment where the terminal device is located. The determination module is used to process the ambient sound and determine the valid sound, wherein the valid text information corresponding to the valid sound includes at least a wake word, and the wake word includes at least: a quick wake word, a full wake word, and an invalid wake word; The analysis module is used to analyze and process the wake word, and when the wake word is a quick wake word, determine the time of the quick wake word as the wake word time; obtain the current time, and determine the waiting time based on the current time and the wake word time; if the waiting time is less than a preset time, determine whether to generate full wake-up event information; if the full wake-up event information is generated, mark the quick wake word to obtain quick wake-up mark information; concatenate the full wake-up event information, the quick wake-up mark information, and the device parameter information of the target device to generate target request information; and concatenate the first audio corresponding to the full wake-up event and the second audio corresponding to the valid sound to generate target audio. The generation module is used to process the effective sound using a preset voice processing strategy when the wake word matches the quick wake word, generate target semantic information, and generate a response command based on the target semantic information. A response module is configured to, based on the response instruction, cause a target device corresponding to the response instruction to perform the operation of the response instruction, wherein the target device includes at least the terminal device.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 6.
9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 6 through the computer program.