A voice recognition wake-up method, system, terminal device and storage medium
By analyzing voice features and location information to select the appropriate wake-up device, the security and device conflict issues of voice recognition wake-up in existing technologies are resolved, achieving safer and more accurate device wake-up.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-04
- Publication Date
- 2026-03-24
AI Technical Summary
Existing voice recognition wake-up technology lacks specific restrictions on personnel, which may lead to security risks and device conflicts. It cannot effectively distinguish between legitimate and unauthorized access, and conflicts exist between multiple devices.
By parsing sound information to obtain speech features, determining whether they meet preset standards, converting them into target text, and combining location information and wake-up semantic information, a suitable target wake-up device is selected to reduce device conflicts and improve security.
It effectively eliminates unauthorized control, reduces device conflicts, improves the security and accuracy of voice recognition wake-up, and provides a comfortable environment for output.
Smart Images

Figure CN116682438B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic equipment technology, and in particular to a voice recognition wake-up method, system, terminal device and storage medium. Background Technology
[0002] Voice recognition wake-up technology is a device function wake-up technology that allows users to wake up smart devices by speaking specific words. By using voice recognition wake-up technology, the device can respond to the user's voice immediately without pressing any buttons or other manual operations.
[0003] In practical applications, waking up smart devices without specific restrictions on personnel may pose security risks, such as unauthorized access, malicious attacks, and data breaches. Furthermore, it could render the system insecure and unable to withstand attacks. Similarly, conflicts between multiple devices can also cause voice recognition wake-up technology to malfunction. Summary of the Invention
[0004] To improve the voice recognition wake-up effect of the device, this application provides a voice recognition wake-up method, system, terminal device and storage medium.
[0005] In a first aspect, this application provides a voice recognition wake-up method, comprising the following steps:
[0006] Acquire sound information;
[0007] Analyze the sound information to obtain the corresponding speech features;
[0008] If the speech features meet the preset speech feature standards, then the speech signal corresponding to the speech features is converted into target text according to the preset speech recognition rules;
[0009] If the target text contains basic wake-up semantic information corresponding to a preset wake-up device, then the number of the preset wake-up devices is obtained;
[0010] If there are multiple preset wake-up devices, then determine whether the target text contains target wake-up semantic information corresponding to the preset wake-up device;
[0011] If the target text contains the target wake-up semantic information corresponding to the preset wake-up device, then the corresponding target wake-up device is obtained based on the target wake-up semantic information;
[0012] If there are multiple target wake-up devices, then determine whether the specified keyword exists in the target wake-up semantic information;
[0013] If the specified keyword is not present in the target wake-up semantic information, then the location information of the target person is obtained;
[0014] According to the preset wake-up rules corresponding to the location information, the target wake-up device is woken up based on the target wake-up semantic information.
[0015] By employing the above technical solution, the collected sound information in the current application scenario can be analyzed to obtain the speech features corresponding to the current speaker. By judging whether the speech features meet the preset speech feature standards, the interference of non-target personnel's voices and illegal control can be effectively eliminated. Furthermore, according to the preset speech recognition rules, the speech signal corresponding to the speech features of the target personnel is converted into target text. If the target text contains basic wake-up semantic information corresponding to the preset wake-up device, in order to reduce the conflict between multiple preset wake-up devices, it is further judged whether the target text contains target wake-up semantic information corresponding to the specific preset wake-up device type. If it exists, in order to reduce the conflict between target wake-up devices of the same type, it is further judged whether the target wake-up semantic information contains a specified keyword indicating a specific target wake-up device. If it does not exist, the location information of the current target personnel is obtained, and the corresponding target wake-up device is woken up by combining the preset wake-up rules corresponding to the location information and the target wake-up semantic information. By comprehensively analyzing the speech features of the target personnel and related wake-up speech information, and then selecting the appropriate target wake-up device to wake up based on the analysis, the speech recognition wake-up effect of the device is improved.
[0016] Optionally, after acquiring the sound information, the following steps are also included:
[0017] Analyze the sound information to obtain the corresponding noise characteristics;
[0018] If the noise feature does not meet the preset noise feature recognition standard, then the noise feature is added to the preset noise feature recognition standard as the updated preset noise recognition feature standard.
[0019] By adopting the above technical solution, while acquiring and recognizing user voice data, the noise characteristics in the current scene are learned and classified, simplifying the noise analysis process and thus improving the noise filtering effect.
[0020] Optionally, after parsing the sound information and obtaining the corresponding speech features, the method further includes the following steps:
[0021] If the speech features do not conform to the preset speech feature standard, then the corresponding speech features to be processed are obtained.
[0022] Based on the voice features to be processed, generate corresponding personnel identity verification items;
[0023] If the verification result corresponding to the personnel identity verification item is passed, then the voice feature to be processed is added to the preset voice feature standard as the updated voice feature standard.
[0024] By adopting the above technical solution, the personnel identity verification items generated based on the voice features to be processed are used to improve the security of voice wake-up of the corresponding devices and to facilitate users in confirming and adding newly added voice features.
[0025] Optionally, after generating the corresponding personnel identity verification item based on the voice information to be processed, the following steps are further included:
[0026] If the verification result corresponding to the personnel identity verification item is pending verification, then it is determined whether the target voice feature corresponding to the voice feature to be processed exists in the preset voice feature collection library;
[0027] If the target speech feature corresponding to the speech feature to be processed exists in the preset speech feature collection library, then it is determined whether there are multiple target speech features;
[0028] If there are multiple target speech features, then by combining the feature similarity between each target speech feature and the speech feature to be processed, a speech feature verification reference table corresponding to the personnel identity verification item is generated.
[0029] By adopting the above technical solution, for voice features that users are unsure about, a voice feature verification reference table is generated based on the feature similarity between the target voice features and the voice features to be processed in the preset voice feature collection library, which facilitates the user's analysis and confirmation of the person's identity.
[0030] Optionally, after generating the voice feature verification reference table corresponding to the personnel identity verification item by combining the similarity between each target voice feature and the voice feature to be processed if there are multiple target voice features, the following steps are further included:
[0031] The speech features are identified and the reference table is verified to obtain the speech data corresponding to the target speech features.
[0032] Based on the feature similarity corresponding to the target speech features, the playback priority corresponding to the speech data is set;
[0033] The audio data is played according to the playback priority.
[0034] By adopting the above technical solution, the corresponding voice data can be played according to the playback priority, which can reconfirm the identity of the non-user and thus improve the security of the voice recognition wake-up function authorization.
[0035] Optionally, the step of waking up the corresponding target wake-up device based on the target wake-up semantic information according to the preset wake-up rule corresponding to the location information includes the following steps:
[0036] Based on the location information, the movement direction corresponding to the target person is obtained;
[0037] If the direction of movement matches the detection direction corresponding to the target wake-up device, then it is determined whether the target distance between the target person and the target wake-up device is within the preset wake-up range;
[0038] If the target distance between the target person and the target wake-up device is within the preset wake-up range, then a corresponding wake-up command is generated based on the target wake-up semantic information identified according to the preset wake-up rules.
[0039] The target wake-up device is woken up according to the wake-up command.
[0040] By adopting the above technical solution, the movement direction and distance of the target personnel are analyzed and confirmed simultaneously, thereby improving the accuracy of waking up the target device.
[0041] Optionally, waking up the target's wake-up device according to the wake-up command includes the following steps:
[0042] Identify the wake-up command and obtain the corresponding target keyword;
[0043] If no quantitative keyword is found among the target keywords, the target wake-up device is activated based on the standard setting parameters corresponding to the target person.
[0044] By adopting the above technical solution and combining it with the conventional settings parameters corresponding to the target personnel to wake up the target wake-up device, it is helpful for the target wake-up device to output a relatively comfortable environment for the target personnel when waking up.
[0045] Secondly, this application provides a voice recognition wake-up system, comprising:
[0046] The first acquisition module is used to acquire sound information;
[0047] The parsing module is used to parse the sound information and obtain the corresponding speech features;
[0048] The conversion module, if the speech features conform to a preset speech feature standard, is used to convert the speech signal corresponding to the speech features into target text according to a preset speech recognition rule;
[0049] The second acquisition module, if the target text contains basic wake-up semantic information corresponding to a preset wake-up device, is used to acquire the number of the preset wake-up devices;
[0050] The first judgment module, if there are multiple preset wake-up devices, is used to determine whether there is target wake-up semantic information corresponding to the preset wake-up device in the target text;
[0051] If the target text contains the target wake-up semantic information corresponding to the preset wake-up device, the third acquisition module is used to acquire the corresponding target wake-up device based on the target wake-up semantic information.
[0052] If there are multiple target wake-up devices, the second judgment module is used to determine whether a specified keyword exists in the target wake-up semantic information.
[0053] The fourth acquisition module, if the specified keyword is not present in the target wake-up semantic information, is used to acquire the location information of the target person;
[0054] The wake-up module is used to wake up the corresponding target wake-up device based on the target wake-up semantic information according to the preset wake-up rules corresponding to the location information.
[0055] By adopting the above technical solution, the parsing module analyzes the sound information collected in the current application scenario to obtain the speech features of the current speaker. Then, it determines whether the speech features meet the preset speech feature standards, which can effectively eliminate the interference of non-target personnel's voices and illegal control. Further, the conversion module converts the speech signal of the target person's corresponding speech features into target text according to the preset speech recognition rules. If the target text contains basic wake-up semantic information corresponding to the preset wake-up device, in order to reduce the conflict between multiple preset wake-up devices, the first judgment module further judges whether there is target wake-up semantic information corresponding to the specific preset wake-up device type in the target text. If it exists, in order to reduce the conflict between target wake-up devices of the same type, the second judgment module further judges whether there is a specified keyword indicating a specific target wake-up device in the target wake-up semantic information. If it does not exist, the location information of the current target person is obtained, and the corresponding target wake-up device is woken up by the wake-up module in combination with the preset wake-up rules corresponding to the location information and the target wake-up semantic information. By comprehensively analyzing the speech features of the target person and related wake-up speech information, and then selecting the appropriate target wake-up device to wake up, the speech recognition wake-up effect of the device is improved.
[0056] Thirdly, this application provides a terminal device, which adopts the following technical solution:
[0057] A terminal device includes a memory and a processor. The memory stores computer instructions that can run on the processor. When the processor loads and executes the computer instructions, it employs the aforementioned voice recognition wake-up method.
[0058] By adopting the above technical solution, a computer instruction is generated by the above-mentioned voice recognition wake-up method and stored in the memory, so that it can be loaded and executed by the processor. Thus, a terminal device can be made based on the memory and the processor, which is convenient to use.
[0059] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution:
[0060] A computer-readable storage medium storing computer instructions, wherein when the computer instructions are loaded and executed by a processor, the above-described voice recognition wake-up method is employed.
[0061] By adopting the above technical solution, a computer instruction is generated by the above-mentioned voice recognition wake-up method and stored in a computer-readable storage medium for loading and execution by the processor. The computer-readable storage medium facilitates the reading and storage of the computer instruction.
[0062] In summary, this application includes at least one of the following beneficial technical effects: by analyzing the sound information collected in the current application scenario, the speech features corresponding to the current speaker can be obtained. By judging whether the speech features meet the preset speech feature standards, the interference of non-target personnel's voices can be effectively eliminated. Furthermore, according to the preset speech recognition rules, the speech signal of the speech features corresponding to the target personnel is converted into target text. If the target text contains basic wake-up semantic information corresponding to the preset wake-up device, in order to reduce the conflict between multiple preset wake-up devices, it is further judged whether the target text contains target wake-up semantic information corresponding to the specific preset wake-up device type. If it exists, in order to reduce the conflict between target wake-up devices of the same type, it is further judged whether the target wake-up semantic information contains a specified keyword indicating a specific target wake-up device. If it does not exist, the location information of the current target personnel is obtained, and the corresponding target wake-up device is woken up by combining the preset wake-up rules corresponding to the location information and the target wake-up semantic information. By comprehensively analyzing the speech features of the target personnel and related wake-up speech information, and then selecting the appropriate target wake-up device to wake up based on the analysis, the speech recognition wake-up effect of the device is improved. Attached Figure Description
[0063] Figure 1 This is a flowchart illustrating steps S101 to S109 of a voice recognition wake-up method according to this application.
[0064] Figure 2 This is a flowchart illustrating steps S201 to S202 in a voice recognition wake-up method according to this application.
[0065] Figure 3 This is a flowchart illustrating steps S301 to S303 of a voice recognition wake-up method according to this application.
[0066] Figure 4 This is a flowchart illustrating steps S401 to S403 of a voice recognition wake-up method according to this application.
[0067] Figure 5 This is a flowchart illustrating steps S501 to S503 of a voice recognition wake-up method according to this application.
[0068] Figure 6 This is a flowchart illustrating steps S601 to S604 of a voice recognition wake-up method according to this application.
[0069] Figure 7 This is a flowchart illustrating steps S701 to S702 in a voice recognition wake-up method according to this application.
[0070] Figure 8 This is a schematic diagram of a voice recognition wake-up system according to this application.
[0071] Explanation of reference numerals in the attached figures:
[0072] 1. First acquisition module; 2. Parsing module; 3. Conversion module; 4. Second acquisition module; 5. First judgment module; 6. Third acquisition module; 7. Second judgment module; 8. Fourth acquisition module; 9. Wake-up module. Detailed Implementation
[0073] The following is in conjunction with the appendix Figure 1-8 This application will be described in further detail.
[0074] This application discloses a voice recognition wake-up method, such as... Figure 1 As shown, it includes the following steps:
[0075] S101. Obtain sound information;
[0076] S102. Analyze the sound information and obtain the corresponding speech features;
[0077] S103. If the speech features meet the preset speech feature standards, then convert the speech signal corresponding to the speech features into target text according to the preset speech recognition rules;
[0078] S104. If the target text contains basic wake-up semantic information corresponding to the preset wake-up device, then obtain the number of preset wake-up devices;
[0079] S105. If there are multiple preset wake-up devices, determine whether there is target wake-up semantic information corresponding to the preset wake-up devices in the target text;
[0080] S106. If the target text contains target wake-up semantic information corresponding to a preset wake-up device, then obtain the corresponding target wake-up device based on the target wake-up semantic information;
[0081] S107. If there are multiple target wake-up devices, determine whether the specified keyword exists in the target wake-up semantic information;
[0082] S108. If the specified keyword is not present in the target wake-up semantic information, then obtain the location information of the target person;
[0083] S109. Based on the preset wake-up rules corresponding to the location information, wake up the corresponding target wake-up device based on the target wake-up semantic information.
[0084] In step S101, the sound information refers to the sound information collected by the sound collection device in the current application scenario, which can be indoors or outdoors.
[0085] In step S102, by parsing the sound information, sound data such as volume, frequency, and waveform can be obtained. Furthermore, based on the frequency, volume, and waveform of the sound, human voices and non-human voices in the sound information can be distinguished. Non-human voices refer to other sounds besides human voices.
[0086] Furthermore, speech features corresponding to human voices can be obtained through speech recognition technology. Specifically, speech recognition technology can detect key information in the sound, such as the frequency, volume, waveform, and delay, to obtain specific speech features.
[0087] In step S103, the preset voice feature standard refers to the voice recognition authentication standard set by the user in advance. Only when the voice feature meets the preset voice feature standard will the system further identify and analyze the semantic information corresponding to the voice feature. If the voice feature meets the preset voice feature standard, it means that the speaker corresponding to the current voice feature is the user or a person authorized by the user in advance, thereby improving the wake-up security of the smart device.
[0088] Furthermore, the speech signal corresponding to the speech features is converted into target text according to the preset speech recognition rules. The preset speech recognition rules refer to the speech keyword recognition rules set by the user in advance. These preset speech recognition rules can be set into a set of corresponding speech keywords according to the user's personal preferences. After the speech features meet the preset speech feature standards, the corresponding speech keywords are obtained by analyzing the speech features, namely frequency, volume, waveform and time delay. Then, it is determined whether the speech keywords match the keywords recorded in the set of speech keywords set by the user in advance. If they exist, the corresponding speech signal is converted into target text information to control the target wake-up device to achieve the corresponding function.
[0089] For example, if the system determines that the current voice features meet the preset voice feature standards, it will further identify the voice features and obtain the corresponding semantic content as "turn on ambient light ballroom mode". If the set of known voice keywords contains "ambient light ballroom mode", it can be determined that the semantic content, i.e., the voice keyword, matches the keyword recorded in the user's preset set of voice keywords. Then, the corresponding voice signal is converted into the target text information for controlling the ambient light to turn on the ballroom mode.
[0090] For example, if the system determines that the current voice characteristics do not meet the preset voice characteristic standards, meaning that the person issuing the wake-up voice is not authorized by the user, the system will then record their voice characteristics.
[0091] In step S104, if the target text contains basic wake-up semantic information corresponding to a preset wake-up device, the preset wake-up device refers to a device with wake-up function that is preset by the user. The basic wake-up semantic information refers to the basic semantic information of the wake-up device, such as the basic wake-up semantic information that does not refer to the device, such as turning on, turning off, raising the temperature, and lowering the temperature.
[0092] In practical applications, if the target text contains basic wake-up semantic information corresponding to the preset wake-up device, the wake-up device recognition may fail because the basic wake-up semantic information does not contain specific device wake-up indication semantics, i.e., wake-up of a specific type or a certain number of wake-up smart devices. Therefore, it is necessary to further obtain and analyze the number of preset wake-up devices.
[0093] For example, if the target text contains the semantic information of "on", it can be determined that there is basic wake-up semantic information corresponding to the preset wake-up device in the target text. Since it is unclear what specific type and number of wake-up smart devices need to be woken up, the number of currently woken-up smart devices, i.e. the preset wake-up devices, can be further obtained.
[0094] For example, if the target text contains the semantic information "vacation", it can be determined that the target text does not contain the basic wake-up semantic information corresponding to the preset wake-up device. In this case, the wake-up smart device will not read the above semantic information, and the retrieval system in the wake-up smart device will continue to identify and retrieve the semantic information in the target text.
[0095] In step S105, if there are multiple preset wake-up devices, it means that there are multiple wake-up smart devices currently existing. These can be multiple devices of the same type or multiple devices of different types. In order to reduce conflicts between multiple wake-up smart devices, it is determined whether there is target wake-up semantic information corresponding to the preset wake-up device in the target text. Target wake-up semantic information refers to information that has specific indicative semantics for wake-up smart devices.
[0096] For example, if there are three preset wake-up devices, including two air conditioners and one ambient light, and the target text contains the semantic information "turn on the air conditioner", then it can be determined that the target text contains the target wake-up semantic information corresponding to the preset wake-up device. If the target text contains the semantic information "turn on", then it can be determined that the target text does not contain the target wake-up semantic information corresponding to the preset wake-up device.
[0097] For example, if there is only one preset wake-up device, namely an air conditioner, then the target text only needs to contain basic wake-up semantic information. For example, if the basic wake-up semantic information is "turn on", then the air conditioner will start directly according to this instruction.
[0098] In steps S106 to S107, if the target text contains target wake-up semantic information corresponding to a preset wake-up device, there may be multiple wake-up smart devices of the same type. In order to reduce the conflict between multiple wake-up smart devices of the same type, the corresponding target wake-up device is obtained according to the target wake-up semantic information.
[0099] Furthermore, if multiple target wake-up devices are identified, it is determined whether a specified keyword exists in the target wake-up semantic information. The specified keyword refers to the semantic keyword that distinguishes similar wake-up smart devices. For example, if there are two target wake-up devices, namely two air conditioners, one of which is located in the living room and the other in the bedroom, then the specified keywords could be "living room" and "bedroom".
[0100] Specifically, by combining the basic wake-up semantic information, target wake-up semantic information, and specified keywords, a specific wake-up smart device can be selected. The selected wake-up smart device then implements the corresponding wake-up function based on the above information.
[0101] In steps S108 to S109, if the specified keyword is not present in the target wake-up semantic information, then in order to reduce conflicts between similar wake-up smart devices, the location information of the target person is obtained. The location information refers to the location information of the target person who is currently speaking.
[0102] Furthermore, based on the aforementioned location information and its corresponding preset wake-up rules, the corresponding target wake-up device is woken up based on the target wake-up semantic information. The preset wake-up rules refer to the rule strategy for waking up the smart device based on the current location information of the target person. Then, the target wake-up device is woken up by combining the target wake-up semantic information in the target text that implements the specific wake-up function.
[0103] The voice recognition wake-up method provided in this embodiment analyzes the sound information collected in the current application scenario to obtain the voice features corresponding to the current speaker. By judging whether the voice features meet the preset voice feature standards, it can effectively eliminate voice interference from non-target personnel and illegal control. Furthermore, according to the preset voice recognition rules, the voice signal of the voice features corresponding to the target personnel is converted into target text. If the target text contains basic wake-up semantic information corresponding to the preset wake-up device, in order to reduce the conflict between multiple preset wake-up devices, it is further judged whether the target text contains target wake-up semantic information corresponding to the specific preset wake-up device type. If it exists, in order to reduce the conflict between target wake-up devices of the same type, it is further judged whether the target wake-up semantic information contains a specified keyword indicating a specific target wake-up device. If it does not exist, the location information of the current target personnel is obtained, and the corresponding target wake-up device is woken up by combining the preset wake-up rules corresponding to the location information and the target wake-up semantic information. By comprehensively analyzing the voice features of the target personnel and related wake-up voice information, and then selecting the appropriate target wake-up device to wake up based on the analysis, the voice recognition wake-up effect of the device is improved.
[0104] In one embodiment of this example, such as Figure 2 As shown, after obtaining the sound information in step S101, the following steps are also included:
[0105] S201. Analyze the sound information and obtain the corresponding noise characteristics;
[0106] S202. If the noise features do not meet the preset noise feature recognition standard, the noise features are added to the preset noise feature recognition standard as the updated preset noise recognition feature standard.
[0107] In step S201, the corresponding noise characteristics can be obtained by parsing the sound information, including frequency, time delay, phase, spectrum, etc., and the source of noise can be analyzed through these characteristics.
[0108] In step S202, the preset noise feature recognition standard refers to a pre-set noise feature database, which contains various noise feature standards. By comparing the currently acquired noise features with the noise features in the noise feature database, the type of noise feature can be quickly identified. If the noise feature does not conform to the preset noise feature recognition standard, it means that the current noise feature is the first noise type to appear. In order to better filter noise during the voice recognition wake-up process, the first noise feature type can be learned and recorded, and the noise feature can be added to the current preset noise feature recognition standard as the updated preset noise recognition feature standard. By continuously training the preset noise recognition feature standard, the ability to recognize and identify various noises can be improved, thereby improving the efficiency of noise filtering.
[0109] The voice recognition wake-up method provided in this embodiment learns and classifies the noise features in the current scene while acquiring and recognizing user voice data, which simplifies the noise analysis process and thus improves the noise filtering effect.
[0110] In one embodiment of this example, such as Figure 3 As shown, after parsing the sound information and obtaining the corresponding speech features in step S102, the following steps are also included:
[0111] S301. If the speech features do not conform to the preset speech feature standard, then obtain the corresponding speech features to be processed;
[0112] S302. Generate corresponding personnel identity verification items based on the voice features to be processed;
[0113] S303. If the verification result corresponding to the personnel identity verification item is passed, the voice features to be processed are added to the preset voice feature standard as the updated voice feature standard.
[0114] In step S301, the preset voice feature standard refers to the pre-set voice feature recognition standard. This preset voice feature standard can be a voice feature database pre-set by the user. As long as the currently acquired voice features match the voice features in the voice feature database, the permission to wake up the corresponding device can be obtained.
[0115] If the current voice characteristics do not conform to the preset voice characteristic standard, it indicates that the person making the voice has not obtained the corresponding wake-up permission. Therefore, the voice characteristics of the unauthorized person are further obtained for user confirmation and analysis. It should be noted that this implementation is designed for scenarios where the user is not physically present; the user can choose to remotely authorize others to control the wake-up of the smart device.
[0116] In step S302, based on the acquired voice features to be processed, a corresponding personnel identity verification item is generated. This personnel identity verification item can be set as a prompt box and sent to the user via email. The personnel identity verification item displays the voice features to be processed of the person corresponding to the current voice features. The voice features to be processed can be confirmed by playing voice messages for the user to confirm.
[0117] In step S303, if the verification result corresponding to the personnel identity verification item is passed, it indicates that the user has completed the confirmation of the above-mentioned voice features to be processed and authorized the current person's wake-up permission. In order to reduce the tediousness of repeatedly authorizing the same person, the currently passed voice features to be processed can be added to the preset voice feature standard as the updated voice feature standard.
[0118] The voice recognition wake-up method provided in this embodiment improves the security of voice wake-up of corresponding devices by generating personnel identity verification items based on the voice features to be processed, and makes it easier for users to confirm and add newly added voice features.
[0119] In one embodiment of this example, such as Figure 4 As shown, step S302, which generates the corresponding personnel identity verification item based on the voice features to be processed, includes the following steps:
[0120] S401. If the verification result corresponding to the personnel identity verification item is pending verification, then determine whether there is a target speech feature corresponding to the speech feature to be processed in the preset speech feature collection library;
[0121] S402. If the target speech feature corresponding to the speech feature to be processed exists in the preset speech feature collection library, then determine whether there are multiple target speech features;
[0122] S403. If there are multiple target speech features, then combine the feature similarity between each target speech feature and the speech feature to be processed to generate a speech feature verification reference table corresponding to the personnel identity verification item.
[0123] In step S401, if the verification result corresponding to the personnel identity verification item is pending verification, it means that the current user is uncertain about the identity of the person corresponding to the current voice feature. In order to facilitate the user to further confirm the identity of the current person, it is determined whether there is a target voice feature corresponding to the voice feature to be processed in the preset voice feature collection library.
[0124] The preset voice feature collection library refers to a pre-set voice feature collection database. This preset voice feature collection library collects and records voice segments of daily communication among various people in its application scenarios, and then identifies the voice segment to generate corresponding voice features. The target voice feature refers to the voice feature in the preset voice feature collection library that matches the voice feature corresponding to the personnel identity verification item.
[0125] In step S402, if the target voice feature corresponding to the voice feature to be processed exists in the preset voice feature collection library, it means that the person whose voice feature corresponds to the personnel identity verification item has appeared in its application scenario in the past. In order to further analyze and confirm its voice feature, it is determined whether there are multiple target voice features matched in the preset voice feature collection library.
[0126] In step S403, if there are multiple target speech features, in order to facilitate the user's confirmation of the corresponding person's identity in the personnel identity verification, a speech feature verification reference table corresponding to the personnel identity verification item is generated by combining the feature similarity between each target speech feature and the speech feature to be processed. While the user confirms the speech features by hearing, the speech feature verification reference table can provide the user with a more intuitive way to analyze the person's identity.
[0127] The voice recognition wake-up method provided in this embodiment further generates a voice feature verification reference table based on the feature similarity between the target voice feature and the voice feature to be processed in the preset voice feature collection library for users who are uncertain about the voice features to be processed. This facilitates the user's analysis and confirmation of the person's identity.
[0128] In one implementation method provided in this embodiment, such as Figure 5 As shown, after step S403, which states that if there are multiple target speech features, the similarity between each target speech feature and the speech feature to be processed is combined to generate a speech feature verification reference table corresponding to the personnel identity verification item, the following steps are also included:
[0129] S501. Verify the speech feature recognition reference table and obtain the speech data corresponding to the target speech feature;
[0130] S502. Set the playback priority of the speech data according to the feature similarity corresponding to the target speech features;
[0131] S503. Play voice data according to playback priority.
[0132] In steps S501 to S502, the user can actively obtain the corresponding voice features in the personnel identity verification item, or set the playback priority of the voice data according to the feature similarity of the target voice features. In the implementation, the feature similarity of the target voice features and its playback priority can be set to a direct proportional relationship, that is, the higher the feature similarity of the target voice features, the higher its playback priority. The purpose is to enable the user to quickly confirm the identity of the target personnel.
[0133] For example, if the feature similarity of target speech feature A is 98%, the feature similarity of target speech feature B is 97%, and the feature similarity of target speech feature C is 95%, then after system recognition, target speech feature A corresponds to level 1 playback priority, target speech feature B corresponds to level 2 playback priority, and target speech feature C corresponds to level 3 playback priority. Level 1 playback priority is higher than level 2, and level 2 playback priority is higher than level 3. 501
[0134] In step S503, based on the playback priority obtained above, the playback order corresponding to each target speech feature is set, and then the speech information, i.e., speech data, corresponding to each target speech feature is played according to the playback order.
[0135] The voice recognition wake-up method provided in this embodiment plays corresponding voice data according to the playback priority, which can reconfirm the identity of the current non-user personnel, thereby improving the security of voice recognition wake-up function authorization.
[0136] In one embodiment of this example, such as Figure 6 As shown, step S109, which is to wake up the corresponding target wake-up device based on the target wake-up semantic information according to the preset wake-up rule corresponding to the location information, includes the following steps:
[0137] S601. Based on the location information, obtain the movement direction corresponding to the target person;
[0138] S602. If the direction of movement matches the detection direction corresponding to the target wake-up device, determine whether the target distance between the target person and the target wake-up device is within the preset wake-up range;
[0139] S603. If the target distance between the target person and the target wake-up device is within the preset wake-up range, then the target wake-up semantic information is identified according to the preset wake-up rules to generate the corresponding wake-up command;
[0140] S604. Wake up the target wake-up device according to the wake-up command.
[0141] In step S601, the orientation information refers to the geographical location information of the current target person. Further, the corresponding movement direction is obtained based on the current orientation information of the target person. The orientation information and corresponding movement direction of the current target person can be obtained through a camera.
[0142] In step S602, if the direction of movement matches the detection direction corresponding to the target wake-up device, it means that the current direction of movement of the target person is towards the target wake-up device. In order to further determine the actual activity direction or area of the target person, it is determined whether the target distance between the target person and the target wake-up device is within the preset wake-up range. The preset wake-up range refers to the perception range that is preset for the target wake-up device.
[0143] In steps S603 to S604, if the target distance between the target person and the target wake-up device is within the preset wake-up range, it indicates that the target person intends to wake up the target wake-up device in the current area. The preset wake-up rule refers to the wake-up rule set in advance for the target wake-up device under the premise that there is no specified keyword in the target wake-up semantic information. The preset wake-up rule can be the aforementioned judgment rule, namely whether the movement direction and the target distance between the target person and the target wake-up device are within the preset wake-up range. If both are met, the target device further identifies the target wake-up semantic information and generates the corresponding wake-up command, and then wakes up the target wake-up device according to the wake-up command.
[0144] The voice recognition wake-up method provided in this embodiment performs dual analysis and confirmation of the target person's movement direction and target distance, thereby improving the accuracy of waking up the target device.
[0145] In one embodiment of this example, such as Figure 7 As shown, step S604, which involves waking up the target wake-up device according to the wake-up command, includes the following steps:
[0146] S701. Recognize the wake-up command and obtain the corresponding target keyword;
[0147] S702. If there are no quantitative keywords in the target keywords, then obtain and wake up the target wake-up device according to the regular setting parameters corresponding to the target personnel.
[0148] In step S701, the corresponding target keywords are obtained by recognizing the wake-up command generated by the recognition system. The target keywords refer to the keywords for waking up the control device in the target wake-up semantic information. For example, the target keywords for "turn on the air conditioner" are "turn on" and "air conditioner".
[0149] In step S702, if there are no quantified keywords in the target keywords, in order to facilitate the target wake-up device to output a more comfortable environment for the user or to meet the user's needs to the greatest extent, the target wake-up device is obtained and woken up according to the conventional setting parameters corresponding to the target personnel.
[0150] In this context, quantified keywords refer to the specific scope of the target wake-up device's functionality. For example, the target wake-up semantic information might be "adjust the air conditioner to 18 degrees Celsius," where "18 degrees Celsius" is the quantified keyword. Regular setting parameters refer to the adjustment parameters that the target user typically uses on the target wake-up device. Therefore, even when the target wake-up semantic information lacks specific quantified keywords, waking the target wake-up device based on the user's corresponding regular setting parameters can improve the user experience.
[0151] The voice recognition wake-up method provided in this embodiment wakes up the target wake-up device by combining the conventional setting parameters corresponding to the target person, thereby helping the target wake-up device to output a relatively comfortable environment for the target person when waking up.
[0152] This application discloses a voice recognition wake-up system, such as... Figure 8 As shown, it includes:
[0153] The first acquisition module 1 is used to acquire sound information;
[0154] Parsing module 2 is used to parse sound information and obtain corresponding speech features;
[0155] If the speech features conform to the preset speech feature standard, the conversion module 3 is used to convert the speech signal corresponding to the speech features into target text according to the preset speech recognition rules.
[0156] If the target text contains basic wake-up semantic information corresponding to a preset wake-up device, the second acquisition module 4 is used to acquire the number of preset wake-up devices.
[0157] The first judgment module 5, if there are multiple preset wake-up devices, is used to determine whether there is target wake-up semantic information corresponding to the preset wake-up device in the target text;
[0158] If the target text contains target wake-up semantic information corresponding to a preset wake-up device, the third acquisition module 6 is used to acquire the corresponding target wake-up device based on the target wake-up semantic information.
[0159] If there are multiple target wake-up devices, the second judgment module 7 is used to determine whether a specified keyword exists in the target wake-up semantic information.
[0160] If the specified keyword is not present in the target wake-up semantic information, the fourth acquisition module 8 is used to acquire the location information of the target person.
[0161] The wake-up module 9 is used to wake up the corresponding target wake-up device based on the target wake-up semantic information according to the preset wake-up rules corresponding to the location information.
[0162] The voice recognition wake-up system provided in this embodiment analyzes the sound information collected in the current application scenario by the parsing module 2 to obtain the voice features corresponding to the current speaker. Then, it determines whether the voice features meet the preset voice feature standards, which can effectively eliminate voice interference from non-target personnel and illegal control. Further, the conversion module 3 converts the voice signal of the voice features corresponding to the target personnel into target text according to the preset voice recognition rules. If the target text contains basic wake-up semantic information corresponding to the preset wake-up device, in order to reduce the conflict between multiple preset wake-up devices, the first judgment module 5 further determines whether there is target wake-up semantic information corresponding to the specific preset wake-up device type in the target text. If it exists, in order to reduce the conflict between target wake-up devices of the same type, the second judgment module 7 further determines whether there is a specified keyword indicating a specific target wake-up device in the target wake-up semantic information. If it does not exist, the location information of the current target personnel is obtained, and the corresponding target wake-up device is woken up by the wake-up module 9 in combination with the preset wake-up rules corresponding to the location information and the target wake-up semantic information. By comprehensively analyzing the voice features of the target personnel and related wake-up voice information, and then selecting the appropriate target wake-up device to wake up based on the analysis, the voice recognition wake-up effect of the device is improved.
[0163] It should be noted that the voice recognition wake-up system provided in this application embodiment also includes each module and / or corresponding sub-module corresponding to the logical function or logical step of any of the above-mentioned voice recognition wake-up methods, to achieve the same effect as each logical function or logical step, which will not be elaborated here.
[0164] This application also discloses a terminal device, including a memory, a processor, and computer instructions stored in the memory and capable of running on the processor, wherein the processor executes the computer instructions using any of the voice recognition wake-up methods described in the above embodiments.
[0165] The terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server. The terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may also include input / output devices, network access devices, and buses.
[0166] The processor can be a central processing unit (CPU). Of course, depending on the actual use, it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc., and this application does not limit it.
[0167] The memory can be an internal storage unit of the terminal device, such as a hard disk or RAM of the terminal device, or an external storage device of the terminal device, such as a plug-in hard disk, smart memory card (SMC), secure digital card (SD), or flash memory card (FC) equipped on the terminal device. Furthermore, the memory can be a combination of internal storage units and external storage devices of the terminal device. The memory is used to store computer instructions and other instructions and data required by the terminal device. The memory can also be used to temporarily store data that has been output or will be output. This application does not limit this.
[0168] In this terminal device, any one of the voice recognition wake-up methods in the above embodiments is stored in the terminal device's memory and loaded and executed on the terminal device's processor for convenient use.
[0169] This application also discloses a computer-readable storage medium, which stores computer instructions, wherein when the computer instructions are executed by a processor, any of the voice recognition wake-up methods described in the above embodiments are employed.
[0170] The computer instructions can be stored in a computer-readable medium. The computer instructions include computer instruction code, which can be in the form of source code, object code, executable file, or certain middleware. The computer-readable medium includes any entity or device capable of carrying computer instruction code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the computer-readable medium includes, but is not limited to, the above-mentioned components.
[0171] In this computer-readable storage medium, any one of the voice recognition wake-up methods in the above embodiments is stored in the computer-readable storage medium and loaded and executed on the processor to facilitate the storage and application of the above methods.
[0172] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A voice recognition wake-up method, characterized in that, Includes the following steps: Acquire sound information; Analyze the sound information to obtain the corresponding speech features; If the speech features meet the preset speech feature standards, then the speech signal corresponding to the speech features is converted into target text according to the preset speech recognition rules; If the target text contains basic wake-up semantic information corresponding to a preset wake-up device, then the number of the preset wake-up devices is obtained; If there are multiple preset wake-up devices, then determine whether the target text contains target wake-up semantic information corresponding to the preset wake-up device; If the target text contains the target wake-up semantic information corresponding to the preset wake-up device, then the corresponding target wake-up device is obtained based on the target wake-up semantic information; If there are multiple target wake-up devices, then determine whether the specified keyword exists in the target wake-up semantic information; If the specified keyword is not present in the target wake-up semantic information, then the location information of the target person is obtained; According to the preset wake-up rule corresponding to the location information, the target wake-up device is woken up based on the target wake-up semantic information. After parsing the sound information and obtaining the corresponding speech features, the following steps are also included: If the speech features do not conform to the preset speech feature standard, then the corresponding speech features to be processed are obtained. Based on the voice features to be processed, generate corresponding personnel identity verification items; If the verification result corresponding to the personnel identity verification item is passed, then the voice feature to be processed is added to the preset voice feature standard as the updated voice feature standard; After generating the corresponding personnel identity verification item based on the voice features to be processed, the following steps are also included: If the verification result corresponding to the personnel identity verification item is pending verification, then it is determined whether the target voice feature corresponding to the voice feature to be processed exists in the preset voice feature collection library; If the target speech feature corresponding to the speech feature to be processed exists in the preset speech feature collection library, then it is determined whether there are multiple target speech features; If there are multiple target speech features, then by combining the feature similarity between each target speech feature and the speech feature to be processed, a speech feature verification reference table corresponding to the personnel identity verification item is generated.
2. The voice recognition wake-up method according to claim 1, characterized in that, After acquiring the sound information, the following steps are also included: Analyze the sound information to obtain the corresponding noise characteristics; If the noise feature does not meet the preset noise feature recognition standard, then the noise feature is added to the preset noise feature recognition standard as the updated preset noise recognition feature standard.
3. The voice recognition wake-up method according to claim 1, characterized in that, If there are multiple target speech features, then after generating the speech feature verification reference table corresponding to the personnel identity verification item by combining the similarity between each target speech feature and the speech feature to be processed, the following steps are also included: The speech features are identified and the reference table is verified to obtain the speech data corresponding to the target speech features. Based on the feature similarity corresponding to the target speech features, the playback priority corresponding to the speech data is set; The audio data is played according to the playback priority.
4. The voice recognition wake-up method according to claim 1, characterized in that, The step of waking up the target wake-up device based on the target wake-up semantic information according to the preset wake-up rule corresponding to the location information includes the following steps: Based on the location information, the movement direction corresponding to the target person is obtained; If the direction of movement matches the detection direction corresponding to the target wake-up device, then it is determined whether the target distance between the target person and the target wake-up device is within the preset wake-up range; If the target distance between the target person and the target wake-up device is within the preset wake-up range, then a corresponding wake-up command is generated based on the target wake-up semantic information identified according to the preset wake-up rules. The target wake-up device is woken up according to the wake-up command.
5. The voice recognition wake-up method according to claim 4, characterized in that, The step of waking up the target's wake-up device according to the wake-up command includes the following steps: Identify the wake-up command and obtain the corresponding target keyword; If no quantitative keyword is found among the target keywords, the target wake-up device is activated based on the standard setting parameters corresponding to the target person.
6. A voice recognition wake-up system, characterized in that, For implementing a voice recognition wake-up method as described in any one of claims 1 to 5, the voice recognition wake-up system comprises: The first acquisition module (1) is used to acquire sound information; The parsing module (2) is used to parse the sound information and obtain the corresponding speech features; If the speech feature conforms to the preset speech feature standard, the conversion module (3) is used to convert the speech signal corresponding to the speech feature into target text according to the preset speech recognition rules; If the target text contains basic wake-up semantic information corresponding to a preset wake-up device, the second acquisition module (4) is used to acquire the number of preset wake-up devices; First judgment module (5): If there are multiple preset wake-up devices, the judgment module is used to determine whether there is target wake-up semantic information corresponding to the preset wake-up device in the target text; If the target text contains the target wake-up semantic information corresponding to the preset wake-up device, the third acquisition module (6) is used to acquire the corresponding target wake-up device according to the target wake-up semantic information; If there are multiple target wake-up devices, the second judgment module (7) is used to determine whether there is a specified keyword in the target wake-up semantic information; The fourth acquisition module (8) is used to acquire the location information of the target person if the specified keyword is not present in the target wake-up semantic information. The wake-up module (9) is used to wake up the target wake-up device corresponding to the target wake-up device based on the target wake-up semantic information according to the preset wake-up rules corresponding to the location information.
7. A terminal device, comprising a memory and a processor, characterized in that, The memory stores computer instructions that can run on the processor. When the processor loads and executes the computer instructions, it employs a voice recognition wake-up method as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are loaded and executed by the processor, a voice recognition wake-up method as described in any one of claims 1 to 5 is employed.
Citation Information
Patent Citations
Terminal awakening method and device
CN116013280A