Voice control method and device, electronic equipment and storage medium
By directly matching the voice data input by the user based on the preset reference information, voice control is realized directly without the need for wake-up words, solving the problem of low efficiency of the voice wake-up process in the prior art, and improving interaction efficiency and flexibility.
Patent Information
- Application Number
- CN202510459082.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing voice wake-up technology requires the voice assistant to wake up through the wake-up word before receiving subsequent instructions, resulting in low processing efficiency.
By receiving voice data input by the user and matching it with the voice data based on the preset reference information, the target voice data containing the reference information is determined, thereby directly performing voice control, and the wake-up process is avoided.
It shortens the response time of instructions, improves the efficiency and flexibility of human-computer interaction, reduces the processing of irrelevant voice data, and improves data processing efficiency.
Smart Images

Figure CN120220682A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice control technology, and in particular, to a voice control method, apparatus, electronic device, and storage medium. Background Art
[0002] Voice wake-up refers to the technology of waking up a display device or an application by recognizing a specific wake-up word, so that it enters a working state or performs a specific operation. This technology can be applied to fields such as smart homes, smart phones, in-vehicle systems, etc., providing a better interaction experience for users. However, currently, voice wake-up must go through the wake-up process to receive subsequent instructions for human-computer interaction, and the processing efficiency is relatively low. Summary of the Invention
[0003] This application aims to solve at least one of the technical problems in the related art to some extent.
[0004] To this end, this application proposes a method, apparatus, electronic device, and storage medium.
[0005] An embodiment of one aspect of this application proposes a voice control method, including:
[0006] Receiving voice data input by a user;
[0007] Matching the preset reference information with the voice data to determine target voice data containing the reference information;
[0008] Performing voice control according to the target voice data; wherein, the reference information is information for human-computer interaction, and the reference information includes at least one of instruction information or a preset wake-up word.
[0009] Optionally, the matching the preset reference information with the voice data to determine target voice data containing the reference information includes:
[0010] Converting the voice data into corresponding text data;
[0011] Retrieving the reference information in the text data, determining the text data containing the reference information as target text data, and determining the voice data corresponding to the target text data as the target voice data.
[0012] Optionally, the retrieving the reference information in the text data and determining the text data containing the reference information as target text data includes at least one of the following:
[0013] Responding to the target position in the text data containing the preset wake-up word, and determining the text data as the target text data;
[0014] In response to any one of the instruction information being included in the text data, determine that the text data is the target text data.
[0015] Optionally, the determining that the text data is the target text data in response to any one of the instruction information being included in the text data includes:
[0016] Determine the text similarity between the text data and the instruction information, and when the text similarity is greater than a preset text similarity threshold, determine that the text data is the target text data.
[0017] Optionally, the performing voice control according to the target voice data includes:
[0018] Determine a corresponding target control instruction according to the target text data;
[0019] Execute the target control instruction according to the target text data, the target voice data, and the target control instruction.
[0020] Optionally, the determining a corresponding target control instruction according to the target text data includes:
[0021] Extract a first feature corresponding to the target text data;
[0022] Determine the similarity between the first feature and a preset second feature, where the second feature is a feature corresponding to a preset control instruction;
[0023] Determine the control instruction corresponding to the second feature with the highest similarity as the target control instruction determined by the target text data.
[0024] Optionally, the executing the target control instruction according to the target text data, the target voice data, and the target control instruction includes:
[0025] Generate a discrimination result according to the target text data, the target voice data, the target control instruction, and the features of historical interaction data, where the discrimination result is used to represent the probability of the target voice data being a human-computer interaction;
[0026] When the probability is greater than a preset probability threshold, determine to execute the target control instruction.
[0027] Optionally, the method further includes:
[0028] Generate a prompt voice according to the target control instruction and play the prompt voice.
[0029] Optionally, the method further includes:
[0030] Obtain the second prompt voice corresponding to the second target control instruction. In response to the first prompt voice corresponding to the first target control instruction being in a playing state, interrupt the playing of the first prompt voice and play the second prompt voice, where the generation time of the first prompt voice is earlier than that of the second prompt voice.
[0031] Optionally, the voice data is voice data of multiple users, and the receiving the voice data input by the user includes:
[0032] Determine the first time point when the collection of the voice data is completed, and pause the collection of the voice data at the first time point;
[0033] Start collecting voice data at a second time point, where the second time point is related to the first time point and a preset pause duration.
[0034] Another embodiment of the present application provides a voice control device, including:
[0035] A receiving module, configured to receive voice data input by a user;
[0036] A processing module, configured to match the voice data with preset reference information to determine target voice data including the reference information;
[0037] A control module, configured to perform voice control according to the target voice data; where the reference information is information for human-computer interaction, and the reference information includes at least one of instruction information or a preset wake-up word. The above voice control device is used to execute the method described in the foregoing aspect.
[0038] Another embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the foregoing aspect is implemented.
[0039] Another embodiment of the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the foregoing aspect is implemented.
[0040] Another embodiment of the present application provides a computer program product, on which a computer program is stored. When the program is executed by a processor, the method described in the foregoing aspect is implemented.
[0041] The voice control method, device, electronic device, chip, and storage medium proposed in this application. Since the reference information can cover common basic instructions or pre-wake word commands of users, there is no need to first wake up the voice assistant to enter the actual instruction conversation, which shortens the response time of instructions and makes the interaction method more flexible. In addition, since voice control is performed by screening voice data containing target information, the processing of irrelevant voice data is avoided, and the data processing efficiency in the voice interaction scenario is also improved. Especially in the intelligent cockpit scenario, a more flexible and efficient conversation method can be provided.
[0042] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Brief Description of the Drawings
[0043] The above-mentioned and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0044] Figure 1 is a schematic flowchart of a voice control method provided by an embodiment of the present application;
[0045] Figure 2 is a schematic structural diagram of a voice control device provided by an embodiment of the present application;
[0046] Figure 3 is a schematic flowchart of a voice control method provided by an embodiment of the present application;
[0047] Figure 4 is a schematic structural diagram of a voice control device provided by an embodiment of the present application;
[0048] Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0049] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0050] Voice wake-up refers to a technology that can only wake up a display device or an application by recognizing a specific wake-up word and then enable it to enter the working state or perform a specific operation. This technology can be applied to fields such as smart home, smart phone, and vehicle-mounted system to provide a better interaction experience for users. However, currently, voice wake-up must go through the wake-up process to receive subsequent instructions for human-computer interaction, and the processing efficiency is relatively low.
[0051] In the related art, before interacting with a voice assistant, the user needs to wake up the voice assistant with a wake-up word first, and then can speak out their own instructions to the voice assistant. In the usage scenario of starting a vehicle, the user has to say "XX classmate" first to wake up the in-vehicle voice assistant, and then speak out specific instructions for subsequent interaction to prompt the voice assistant to execute specific instructions.
[0052] Next, the voice control method, device, electronic device, chip, and storage medium according to the embodiments of the present application will be described with reference to the accompanying drawings.
[0053] Figure 1 It is a schematic flowchart of a voice control provided by an embodiment of the present application.
[0054] As an implementation manner, the voice control method of the embodiments of the present application can be configured in a voice control device, and the voice control device can be applied to any electronic device so that the electronic device can perform the voice control function.
[0055] Among them, the electronic device can be any device with computing power. For example, it can be a mobile terminal, and the mobile terminal can be a hardware device such as a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a vehicle, etc. with various operating systems, touch screens, and / or display screens.
[0056] As Figure 1 shown, the method may include the following steps:
[0057] Step 101, receiving voice data input by the user;
[0058] Step 102, matching the voice data with preset reference information to determine target voice data containing the reference information;
[0059] Step 103, performing voice control according to the target voice data; where the reference information is information for human-computer interaction, and the reference information includes at least one of instruction information or a preset wake-up word.
[0060] In this embodiment, the vehicle-mounted computer collects the voice of the user in the vehicle in real time through a built-in sound sensor in the vehicle, that is, the voice data. After collecting the voice data, it is matched with the preset reference information. If the voice data matches the reference information, the voice data is determined to be the target voice information.
[0061] The voice data matching the reference information means that the voice data contains the reference information, or the meaning expressed by the voice information is similar to the reference information. The reference information is information for human-computer interaction. If the voice data matches the reference information, it means that the user speaks the voice to interact with the vehicle and obtain the feedback of the vehicle.
[0062] To determine whether the voice data spoken by the user contains the intention of human-machine interaction with the vehicle, it can be determined from two aspects. One aspect is to determine the instructions for the vehicle. The instructions for the vehicle are clearly distinguishable from the instructions for other objects. The interaction with the vehicle includes commanding the in-vehicle computer to control one or more components in the vehicle (such as windows, air conditioners, etc.), or obtaining specific information from the in-vehicle computer (such as vehicle position information, temperature information, etc.). Another aspect is to determine the object the user is speaking to, that is, the addressing word in the voice data. The object addressed by the user during the speech reflects the object the user wants to interact with. During the interaction with the vehicle, the user will address the name of the voice assistant and add instructions for the vehicle before or after the name.
[0063] Therefore, the reference information set in the embodiment includes at least one of instruction information or a preset wake-up word. The appearance of any one of these two items in the voice data indicates that the voice data has the intention of human-machine interaction and is the target voice data.
[0064] After determining the target voice data, there is no need to wake up the voice assistant. The target voice data is directly input into the voice assistant so that the voice assistant performs voice control according to the target voice data. Compared with the related technology where the wake-up word is spoken to wake up the voice assistant and then the instruction for the in-vehicle computer is spoken. In the solution of this embodiment, the in-vehicle computer responds faster to the user's instructions, has higher efficiency in processing the user's instructions, and the user experience during the interaction with the in-vehicle computer is more natural.
[0065] Optionally, the step 102 of matching the preset reference information with the voice data to determine the target voice data containing the reference information includes:
[0066] Convert the voice data into corresponding text data;
[0067] Retrieve the reference information in the text data, determine the text data containing the reference information as the target text data, and determine the voice data corresponding to the target text data as the target voice data.
[0068] In this embodiment, after collecting the speech data, in order to analyze the semantics of the speech data, the speech data is converted into editable and searchable text data through Automatic Speech Recognition (ASR). ASR is a technology that converts human speech signals into text, and its goal is to enable a computer to understand human language and accurately convert speech information into text output. For the collected speech data, through processing steps such as preprocessing, feature extraction, acoustic modeling, language modeling, and decoding, the optimal text output is obtained. One piece of speech data contains one or more sentences of text. Then, the reference information is retrieved from the text data to determine whether the text data contains the reference information.
[0069] In a possible embodiment, determining whether the text data contains the reference information includes: setting a whitelist that contains multiple pieces of reference information, comparing the text data with the reference information in the whitelist, and if the text data contains the reference information, determining that the text data contains the reference information. Or, calculating the text similarity between the text data and each piece of reference information in the whitelist, and if the text similarity is higher than a certain similarity threshold, determining that the text data contains the reference information.
[0070] In a possible embodiment, the text data and the reference information are respectively tokenized to form a set of words, a word frequency vector is constructed based on the set of words, and the cosine similarity formula is used to calculate the text similarity in combination with the word frequency vector. In a possible embodiment, during the process of determining whether the text data contains the reference information, for a preset wake-up word in the reference information, as long as the preset wake-up word appears in the text data, it can be determined that the text data is the target text data.
[0071] In a possible embodiment, the text data is "Open half of the window", and one of the instruction messages in the reference information is "Open the window". After analyzing the similarity between the text data and the reference information, it is determined that the text similarity between the text data and the reference information is higher than the threshold, and "Open half of the window" is determined as the target text data.
[0072] In a possible embodiment, the text data is "Little X, please fully open the sunroof of the vehicle", and the reference information does not contain instruction information related to the vehicle sunroof. However, "Little X" is a preset wake-up word in the reference information, so it also indicates that the text data has the intention of human-computer interaction, and "Little X, please fully open the sunroof of the vehicle" is determined as the target text data.
[0073] Optionally, retrieving the reference information from the text data and determining the text data containing the reference information as the target text data includes at least one of the following:
[0074] Determine that the text data is the target text data in response to the target position in the text data containing the preset wake-up word;
[0075] Determine that the text data is the target text data in response to the text data containing any one of the instruction information.
[0076] In this embodiment, the reference information may be instruction information or a preset wake-up word. Optionally, the preset wake-up word is a nickname for a voice assistant, such as "XX classmate" or "Little X Little X". If the text data corresponding to the voice data contains the preset wake-up word, it indicates that the sentence spoken by the user may contain information about human-computer interaction, such as an instruction issued to the voice assistant. Then determine that the text data is the target text data and perform further analysis on the target text data.
[0077] The instruction information is an instruction issued for the vehicle. The instruction information contains the operation intention of a person. If the text data corresponding to the voice data contains the instruction information, it indicates that the sentence spoken by the user may contain information about human-computer interaction. Then determine that the text data is the target text data and perform further analysis on the target text data.
[0078] Optionally, the instruction information is information configured after statistical analysis of the user's historical requests. In a possible embodiment, the matching of the instruction information is a whitelist matching mechanism. By statistically analyzing the high-frequency requests in the user's historical requests (for example, the number of requests is higher than a certain threshold), these high-frequency requests are added to the whitelist. Only when the text data contains the instruction information or the similarity between the text data and the instruction information is higher than a certain threshold can it be determined that the matching is successful.
[0079] Two retrieval methods for reference information in text data (preset wake-up word or instruction information) are clarified, which can quickly identify the user's intention and improve the flexibility of voice control. Through the combination of the two retrieval methods, it is ensured that the system can adapt to the expression methods of different users and enhance the robustness of the system.
[0080] Optionally, the determining that the text data is the target text data in response to the text data containing any one of the instruction information includes:
[0081] Determine the text similarity between the text data and the instruction information. When the text similarity is greater than the preset text similarity threshold, determine that the text data is the target text data.
[0082] In this embodiment, it is determined whether the text data contains instruction information through text similarity. By performing a detailed analysis of the text data and the instruction information through text similarity, it is possible to more accurately determine whether the text data contains instruction information, improve the accuracy of the vehicle's feedback to the user's instructions in human-computer interaction, and enhance the user experience.
[0083] In a possible embodiment, the text similarity calculation is a similarity analysis based on morphology. The specific process is as follows: Word segmentation, where the two sentences are segmented separately to extract words or phrases; Stop word removal, removing meaningless words (such as "de", "shi", "he", etc.); Calculating the word overlap degree, and counting the number of words that appear in both sentences as the similarity. For example, if the text data is "Navigate to the parking lot" and one of the instruction messages is "Navigate to the parking lot", there are 5 overlapping characters, so the text overlap degree is 5 / 6.
[0084] In a possible embodiment, the text similarity calculation is a similarity analysis based on word vectors. The specific process is as follows: Word vector representation: Using pre-trained word vectors to represent each word as a vector. Sentence vectorization: Averaging the vectors of all words in the sentence to obtain a vector representation of the sentence. Calculating similarity: Using cosine similarity or other methods to calculate the similarity between the sentence vectors corresponding to the text data and the instruction information as the text similarity. By extracting vectors for similarity calculation, the semantic features in the text can be better utilized, and the text similarity analysis can be more accurate. Optionally, step 103 performs voice control according to the target voice data, including:
[0085] Determining a corresponding target control instruction according to the target text data;
[0086] Executing the target control instruction according to the target text data, the target voice data, and the target control instruction.
[0087] In this embodiment, the target text data is analyzed by Natural Language Processing (NLP) technology to determine the corresponding target control instruction. The NLP technology analyzes the target text data through processes such as text preprocessing, morphology analysis, syntax analysis, semantic understanding, word segmentation, text classification, text similarity processing, sentiment analysis, text generation, etc. to determine the target control instruction that matches the target text data.
[0088] Then, a comprehensive analysis is performed on the target text data, the target voice data, and the target control instruction to analyze whether the target control instruction is for human-computer interaction, so as to further determine whether to execute the target control instruction.
[0089] The combination of the target text data, the target voice data, and the target instruction code ensures the accuracy and consistency of the instruction execution.
[0090] The corresponding target control instruction is extracted from the target voice data. The target control instruction is a code that the machine can understand, and the system can perform corresponding operations according to the target control instruction, such as: opening the window, turning on the air conditioner, adjusting the air conditioner temperature, etc.
[0091] By generating the control instruction and performing the corresponding operation, the voice control becomes more intelligent, can directly complete tasks according to the user's voice instruction, and improves the user experience.
[0092] Optionally, determining the corresponding target control instruction according to the target text data includes:
[0093] Extracting the first feature corresponding to the target text data;
[0094] Determining the similarity between the first feature and the preset second feature, where the second feature is the feature corresponding to the preset control instruction;
[0095] Determining the control instruction corresponding to the second feature with the highest similarity as the target control instruction determined by the target text data.
[0096] In this embodiment, the first feature corresponding to the target text data is extracted through the operation of feature extraction, and the similarity between the first feature and the preset second feature is used to characterize the similarity between the target text data and the control instruction. The higher the similarity, the more similar the target text data and the control instruction are. The text data most similar to the target text data is determined as the target text data.
[0097] Through feature extraction and similarity calculation, the target instruction code can be more accurately matched, avoiding misjudgment of the instruction code. It improves the efficiency and accuracy of the instruction code matching, and ensures the stability and reliability of the voice control system.
[0098] Optionally, executing the target control instruction according to the target text data, the target voice data, and the target control instruction includes:
[0099] Generating a discrimination result according to the features of the target text data, the target voice data, the target control instruction, and the historical interaction data, where the discrimination result is used to characterize the probability of the target voice data being a human-computer interaction;
[0100] When the probability is greater than the preset probability threshold, it is determined to execute the target control instruction.
[0101] In this embodiment, a semantic rejection recognition model is used to determine whether a target control instruction is a human-computer interaction. Through multi-dimensional analysis of these data, such as target text data, target voice data, target control instructions, and historical interaction data, it is possible to better determine whether the target voice data contains the intention of human-computer interaction. The probability is used to characterize the likelihood of the target voice data being a human-computer interaction. The higher the probability, the higher the likelihood of the target voice data being a human-computer interaction.
[0102] The semantic rejection recognition model can filter out invalid target voice data. For example, if the target voice data is the content of communication between people or the communication content with unclear intentions, the corresponding probability will be lower than the preset threshold.
[0103] If the target voice data is human-computer interaction content with obvious intentions, then the probability will be greater than the preset threshold, and the target control instruction needs to be executed subsequently.
[0104] The generation of discriminant results and threshold judgment are introduced, which can more intelligently judge whether the voice data is a valid instruction for human-computer interaction. Through the feature analysis of historical interaction data, the self-learning ability of the system is enhanced, and the accuracy and security of voice control are improved.
[0105] Optionally, the method further includes:
[0106] Generate a prompt voice according to the target control instruction and play the prompt voice. In this embodiment, after determining to execute the target control instruction, feedback needs to be given to the user to let the user know that his instruction has been executed. Therefore, it is necessary to generate a corresponding prompt voice according to the target control instruction and play it.
[0107] The generation of the prompt voice uses text-to-speech (TTS) technology. Through text processing, application of acoustic models, and speech synthesis, playable audio data is generated. Analyze and process the text, including text regularization, word segmentation, part-of-speech tagging, prosody prediction, etc. This step ensures that the text can be correctly understood and pronounced. Convert the processed text into acoustic features, such as linear spectrograms, mel spectrograms, etc. The acoustic model needs to solve the mapping problem between unequal-length sequences to generate natural and fluent speech. Finally, the acoustic features are converted into playable speech waveforms through a vocoder.
[0108] In a possible embodiment, the target control instruction is used to control a certain module inside the vehicle to perform a certain operation. For example, control the in-vehicle air conditioner inside the vehicle to turn on, adjust the temperature, turn off, etc. The corresponding prompt voice is used to inform the user that the operation has been completed. For example, the prompt voice generated after the execution of the target control instruction "turn on the air conditioner" is "The air conditioner has been turned on".
[0109] In a possible embodiment, the target control instruction is used to control the vehicle to obtain specific information and feedback it to the user. For example, it controls the in-vehicle computer to obtain data such as the vehicle speed, the temperature inside the vehicle, and the tire pressure. Or, it controls the in-vehicle computer to obtain data such as the temperature today and the vehicle's location through the network. The corresponding prompt voice is used to feedback the obtained information to the user. For example, if the target control instruction "What's the weather like today?" obtains weather data of sunny and 20 degrees after execution, the generated prompt voice will be "It's sunny today, and the temperature is 20 degrees Celsius."
[0110] By generating and playing the prompt voice, it provides instant feedback to the user and improves the user experience. The prompt voice can help the user confirm whether the system correctly understands the instruction, enhancing the interactivity and friendliness of voice control.
[0111] The method further includes:
[0112] Obtain the second prompt voice corresponding to the second target control instruction. In response to the first prompt voice corresponding to the first target control instruction being in the playing state, interrupt the playing of the first prompt voice and play the second prompt voice. The generation time of the first prompt voice is earlier than that of the second prompt voice.
[0113] In this embodiment, in the intelligent cockpit, in order to accurately collect the voices of each user, corresponding sound zones are divided for each user, and corresponding voice collection modules and sound playback modules are set in the sound zones corresponding to each user. When the user inputs voice data, the corresponding target control instruction is determined according to the voice data. After the in-vehicle computer executes the target control instruction, the corresponding prompt voice is generated, and the execution result of the target control instruction is included in the prompt voice. In order to more accurately feedback the execution result of the target control instruction to the user, after the prompt voice is generated, determine the sound zone where the voice collection module that collected the voice data is located, and play the prompt voice in the sound playback module of this sound zone, so that the user can accurately receive the prompt voice and improve the user's usage experience.
[0114] During the playback of the first prompt voice corresponding to the first target control instruction corresponding to the voice input by the user, if the second target control instruction generated by the voice data input by other users has been executed and the corresponding second prompt voice has been generated, if the first prompt voice and the second prompt voice are played simultaneously at this time, it will cause different voices to be played simultaneously in different sound zones inside the vehicle, and the sounds inside the vehicle will be mixed, reducing the user experience. In this embodiment, by interrupting the first prompt voice and playing the second prompt voice, it avoids the chaos caused by the simultaneous broadcast of prompt voices in multiple sound zones inside the vehicle and improves the user experience.
[0115] In a possible embodiment, the target control instruction is used to control a certain module inside the vehicle to perform a certain operation. For example, it controls the on-board air conditioner inside the vehicle to turn on, adjust the temperature, turn off, etc. The corresponding prompt voice is used to inform the user that the operation has been completed. For example, the prompt voice generated after the execution of the target control instruction "turn on the air conditioner" is "The air conditioner has been turned on".
[0116] In a possible embodiment, the driver inputs "What is the temperature inside the vehicle", and the vehicle head unit generates the first prompt voice "The temperature inside the vehicle is 20 degrees Celsius". During the playback of the first prompt voice, the passenger inputs "What's the weather today", and the vehicle head unit generates the second prompt voice "It's light rain today". At the moment when the second prompt voice "It's light rain today" is generated, if the first prompt voice has not finished playing, the playback of the first prompt voice is interrupted and "It's light rain today" is played.
[0117] Optionally, the voice data is the voice data of multiple users. The receiving of the voice data input by the user includes:
[0118] Determine the first time point when the voice data collection is completed, and pause the collection of the voice data at the first time point;
[0119] Start collecting voice data at the second time point, and the second time point is related to the first time point and a preset pause duration.
[0120] In this embodiment, in the intelligent cockpit, in order to accurately collect the voices of each user, corresponding sound zones are divided for each user, and corresponding voice collection modules and sound playback modules are set in the sound zones corresponding to each user. Since the space inside the vehicle is small, the voice collection module in one sound zone may collect the sound in another sound zone. To prevent this situation, after collecting a segment of voice data, control all the voice collection modules inside the vehicle to pause the sound collection, and then control all the voice collection modules inside the vehicle to start collecting voice data again after the pause duration. This embodiment improves the operating efficiency of the system by preventing the same voice from being repeatedly collected by different sound zones for ineffective voice recognition through the voice collection interruption mechanism.
[0121] In a possible embodiment, the pause duration is at least 300 ms, that is, after collecting a segment of voice data, pause the voice data collection for 300 ms, and then start collecting voice data again.
[0122] In a possible embodiment, a voice control device is provided. Figure 2 This is a schematic structural diagram of a voice control device provided by an embodiment of the present application. As Figure 2 shown, the steps executed by each module are:
[0123] 1) Wake up the APP: A voice detection module integrated in the in-vehicle infotainment system (IVI) is used to determine whether there is sound in the current sound zone of the cockpit.
[0124] 2) Wake-free ASR: A module for converting audio to text integrated in the IVI, which converts the audio recognized by the wake-up APP into text.
[0125] 3) Wake-up rough screening: Based on the recognition result of the wake-free ASR, it is determined whether the current sound is a wake-free request sent by the user to Xiaoai. This module is mainly used to ensure the recall of wake-free requests and is equivalent to the first verification of wake-free requests. If the wake-up rough screening determines that it is a wake-free request, the audio of the current request is transmitted to the voice assistant client.
[0126] 4) Voice assistant client: Integrates various capabilities of Xiaoai on the terminal side and is also responsible for interacting with the cloud. After receiving the audio from the wake-up rough screening, it will send the audio to speech in the form of a stream.
[0127] 5) Speech: The cloud scheduling module of Xiaoai is responsible for scheduling data and business flows to enable close cooperation among various cloud modules.
[0128] 6) ASR: A cloud module for converting audio to text, which receives the audio from speech and converts it into text.
[0129] 7) NLP: A cloud text processing module that receives text input, understands the user's request, and generates execution instructions.
[0130] 8) Wake-free judgment: Based on the understanding result of NLP and the original audio information, it is determined whether the current request is a human-computer interaction. This module is equivalent to the secondary verification of wake-free requests, ensuring the accuracy of wake-free requests for users, suppressing non-human-computer interaction requests, and reducing interference to users. If the wake-free judgment passes, speech will request the TTS module with the instruction set generated by the NLP module.
[0131] 9) TTS: A cloud text processing module used to convert text into audio. This is the last step in cloud execution. The generated audio file and the execution instructions generated by the NLP module will be scheduled by the speech module to the voice assistant client for final execution and audio playback.
[0132] Figure 3 It is a schematic flow diagram of a voice control method provided by an embodiment of this application. As Figure 3 shown, the method includes:
[0133] After receiving a wake-free request, generate target text data corresponding to the audio according to ASR recognition;
[0134] Understand the target control instruction corresponding to the target text data through NLP;
[0135] Then, according to the rejection recognition model, perform a hands-free wake-up judgment to determine whether the target text data is text for human-computer interaction;
[0136] If so, through the judgment, issue the target control instruction, and perform TTS synthesis on the prompt voice according to the target control instruction, and send the prompt voice to the client for playback;
[0137] If not, the judgment fails, issue an end session instruction, and the client will no longer execute the instruction after receiving it.
[0138] To implement the above embodiments, an embodiment of the present application also proposes a voice control device.
[0139] Figure 4 It is a schematic structural diagram of a voice control device provided by an embodiment of the present application.
[0140] As Figure 4 shown, the device may include:
[0141] A receiving module 410, configured to receive voice data input by a user;
[0142] A processing module 420, configured to match the voice data with preset reference information to determine target voice data including the reference information;
[0143] A control module 430, configured to perform voice control according to the target voice data; wherein, the reference information is information for human-computer interaction, and the reference information includes at least one of instruction information or a preset wake-up word.
[0144] Optionally, the processing module 420 includes:
[0145] A conversion module, configured to convert the voice data into corresponding text data;
[0146] A target text determination module, configured to retrieve the reference information in the text data, determine the text data including the reference information as the target text data, and determine the voice data corresponding to the target text data as the target voice data.
[0147] Optionally, the target text determination module includes at least one of the following:
[0148] A first determination module, configured to determine the text data as the target text data in response to the target position in the text data including the preset wake-up word;
[0149] A second determination module, configured to determine the text data as the target text data in response to any one of the instruction information being included in the text data.
[0150] Optionally, the first determination module includes:
[0151] A similarity determination module, configured to determine the text similarity between the text data and the instruction information, and determine the text data as the target text data when the text similarity is greater than a preset text similarity threshold.
[0152] Optionally, the control module 430 includes:
[0153] An instruction determination module, configured to determine a corresponding target control instruction according to the target text data;
[0154] An execution module, configured to execute the target control instruction according to the target text data, the target voice data, and the target control instruction.
[0155] Optionally, the instruction determination module includes:
[0156] A first feature extraction module, configured to extract a first feature corresponding to the target text data;
[0157] A similarity determination module, configured to determine the similarity between the first feature and a preset second feature, where the second feature is a feature corresponding to a preset control instruction;
[0158] A control instruction determination module, configured to determine the control instruction corresponding to the second feature with the highest similarity as the target control instruction corresponding to the target text data.
[0159] Optionally, the execution module includes:
[0160] A probability determination module, configured to generate a discrimination result according to the features of the target text data, the target voice data, the target control instruction, and historical interaction data, where the discrimination result is used to represent the probability that the target voice data is for human-computer interaction;
[0161] An instruction execution module, configured to determine to execute the target control instruction when the probability is greater than a preset probability threshold.
[0162] Optionally, the device further includes:
[0163] A prompt module, configured to generate a prompt voice according to the target control instruction and play the prompt voice.
[0164] Optionally, the device further includes:
[0165] An interruption module, configured to obtain a second prompt voice corresponding to a second target control instruction, and in response to the first prompt voice corresponding to the first target control instruction being in a playing state, interrupt the playing of the first prompt voice and play the second prompt voice, wherein the generation time of the first prompt voice is earlier than that of the second prompt voice.
[0166] Optionally, the voice data is voice data of multiple users, and the receiving module includes:
[0167] A pause acquisition module, configured to determine a first time point when the acquisition of the voice data is completed, and pause the acquisition of the voice data at the first time point;
[0168] A resume acquisition module, configured to start acquiring voice data at a second time point, where the second time point is related to the first time point and a preset pause duration. It should be noted that the foregoing explanations of the method embodiments also apply to the apparatus of this embodiment, and will not be elaborated here.
[0169] To implement the above embodiments, the present application also proposes a non-transitory computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the foregoing method embodiments is implemented.
[0170] To implement the above embodiments, the present application also proposes a computer program product, on which a computer program is stored. When the computer program is executed by a processor, the method described in the foregoing method embodiments is implemented.
[0171] To implement the above embodiments, the present application also proposes an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the foregoing method embodiments is implemented.
[0172] Figure 5 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device 800 may be a mobile terminal, a smart speaker, a vehicle, a smart cockpit, etc.
[0173] Refer to Figure 5 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0174] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above - mentioned methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0175] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non - volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read - only memory (EEPROM), erasable programmable read - only memory (EPROM), programmable read - only memory (PROM), read - only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0176] The power component 806 provides power to various components of the electronic device 800. The power component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0177] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front - facing camera and / or a rear - facing camera. When the electronic device 800 is in an operation mode, such as a shooting mode or a video mode, the front - facing camera and / or the rear - facing camera can receive external multimedia data. Each front - facing camera and rear - facing camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0178] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0179] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module may be a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0180] The sensor component 814 includes one or more sensors for providing an assessment of various aspects of the status of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor component 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 may further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0181] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0182] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0183] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0184] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0185] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0186] Any process or method description in the flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0187] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0188] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0189] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0190] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0191] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A voice control method, characterized in that: The following steps are involved: Receiving voice data input by a user; Matching the voice data with the preset reference information to determine target voice data containing the reference information; Voice control is performed according to the target voice data; wherein the reference information is information used for human-computer interaction, and the reference information includes at least one of instruction information or a preset wake-up word.
2. The method according to claim 1, characterized in that: The matching of the preset reference information with the voice data to determine the target voice data containing the reference information includes: Converting the speech data into corresponding text data; The reference information is retrieved from the text data, the text data containing the reference information is determined as target text data, and the voice data corresponding to the target text data is determined as the target voice data.
3. The method according to claim 2, characterized in that The step of retrieving the reference information from the text data and determining the text data containing the reference information as the target text data includes at least one of the following: In response to a target position in the text data containing the preset wake-up word, determining that the text data is the target text data; In response to the text data containing any one of the instruction information, the text data is determined to be the target text data.
4. The method according to claim 3, characterized in that In response to the text data containing any one of the instruction information, determining that the text data is the target text data includes: The text similarity between the text data and the instruction information is determined, and when the text similarity is greater than a preset text similarity threshold, the text data is determined to be the target text data.
5. The method according to claim 4, characterized in that The performing voice control according to the target voice data comprises: Determine a corresponding target control instruction according to the target text data; The target control instruction is executed according to the target text data, the target voice data and the target control instruction.
6. The method according to claim 5, characterized in that The step of determining the corresponding target control instruction according to the target text data comprises: Extracting a first feature corresponding to the target text data; Determine a similarity between the first feature and a preset second feature, wherein the second feature is a feature corresponding to a preset control instruction; The control instruction corresponding to the second feature with the highest similarity is determined as the target control instruction corresponding to the target text data.
7. The method according to claim 6, characterized in that The executing the target control instruction according to the target text data, the target voice data and the target control instruction comprises: Generate a discrimination result according to the characteristics of the target text data, the target voice data, the target control instruction and the historical interaction data, wherein the discrimination result is used to characterize the probability that the target voice data is human-computer interaction; When the probability is greater than a preset probability threshold, it is determined to execute the target control instruction.
8. The method according to claim 7, characterized in that The method further comprises: Generate a prompt voice according to the target control instruction, and play the prompt voice.
9. The method according to claim 8, characterized in that The method further comprises: A second prompt voice corresponding to a second target control instruction is obtained, and in response to the first prompt voice corresponding to the first target control instruction being in a playing state, the playing of the first prompt voice is interrupted, and the second prompt voice is played, wherein the generation time of the first prompt voice is earlier than that of the second prompt voice.
10. The method according to claim 1, characterized in that The voice data is voice data of multiple users, and the receiving of voice data input by the user includes: Determine a first time point at which the voice data collection is completed, and suspend the collection of the voice data at the first time point; The voice data is collected starting at a second time point, where the second time point is related to the first time point and a preset pause duration.
11. A voice control device, characterized in that: The device is used to perform the method according to any one of claims 1 to 10, comprising: A receiving module, used for receiving voice data input by a user; A processing module, used for matching the voice data with preset reference information to determine target voice data containing the reference information; A control module is used to perform voice control according to the target voice data; wherein the reference information is information used for human-computer interaction, and the reference information includes at least one of instruction information or a preset wake-up word.
12. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 10 is implemented.
13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.
14. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.