Voice Control Method, Device, Electronic Device and Storage Medium

By analyzing and judging the to-process voice, the wake-up control of the target type voice is achieved, the problem of false wake-up of voice devices is solved, and the wake-up accuracy and user experience are improved.

CN114023335BActive Publication Date: 2025-08-01APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202111314250.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-08
Publication Date
2025-08-01
Estimated Expiration
2041-11-08

AI Technical Summary

Technical Problem

Existing voice devices are susceptible to interfering voice during wake-up, resulting in low wake-up accuracy.

Method used

By obtaining the to-processed voice, performing feature analysis to obtain the speech feature vector, and determining whether the to-processed voice is the target type voice based on the speech feature vector, thereby performing wake-up control to avoid false wake-up.

Benefits of technology

It improves the accuracy of voice device wake-up, reduces the occurrence of false wake-ups, and improves the user's interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114023335B_ABST
    Figure CN114023335B_ABST
Patent Text Reader

Abstract

The present disclosure provides a voice control method, apparatus, electronic device, and storage medium, which relate to the field of computer technologies, and specifically to artificial intelligence technologies such as vehicle networking and intelligent cockpits. The method includes: obtaining a voice to be processed, performing feature analysis on the voice to be processed to obtain a voice feature vector, then determining whether the voice to be processed is a target type voice according to the voice feature vector, and when the voice to be processed is a target type voice, performing wake-up control on a target device according to the voice to be processed. Since the wake-up control of the target device is performed according to the type of the voice to be processed, it is possible to effectively avoid false wake-up caused by other types of voices, effectively improve the accuracy of device wake-up, and effectively improve the effect of voice wake-up control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular to artificial intelligence technologies such as vehicle networking and intelligent cockpits. Specifically, it relates to a voice control method, device, electronic device, and storage medium. Background Art

[0002] Artificial intelligence is a discipline that studies enabling computers to simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning, deep learning, big data processing technology, and knowledge graph technology.

[0003] In related technologies, in scenarios where a voice device (the voice device can be, for example, a smart speaker, a smart watch, etc.) is awakened, there are a large number of interfering type voices, which may cause the voice device to be misawakened, resulting in a low wake-up accuracy rate of the voice device. Summary of the Invention

[0004] The present disclosure provides a voice control method, device, electronic device, storage medium, and computer program product.

[0005] According to a first aspect of the present disclosure, there is provided a voice control method, including: obtaining a voice to be processed; performing feature parsing on the voice to be processed to obtain a voice feature vector; judging whether the voice to be processed is a target type voice according to the voice feature vector; if the voice to be processed is the target type voice, performing wake-up control on a target device according to the voice to be processed.

[0006] According to a second aspect of the present disclosure, there is provided a voice control device, including: an obtaining module, configured to obtain a voice to be processed; an analyzing module, configured to perform feature parsing on the voice to be processed to obtain a voice feature vector; a judging module, configured to judge whether the voice to be processed is a target type voice according to the voice feature vector; a wake-up module, configured to perform wake-up control on a target device according to the voice to be processed when the voice to be processed is the target type voice.

[0007] According to a third aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the voice control method as in the first aspect of the present disclosure.

[0008] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the voice control method as in the first aspect of the present disclosure.

[0009] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the steps of the voice control method as in the first aspect of the present disclosure.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0013] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure;

[0016] Figure 5 shows a schematic block diagram of an exemplary electronic device for implementing the voice control method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0018] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure.

[0019] Among them, it should be noted that the execution subject of the voice control method in this embodiment is a voice control device, which can be implemented in a software and / or hardware manner, and the device can be configured in an electronic device, and the electronic device can include but is not limited to a terminal, a server, etc.

[0020] The embodiments of the present disclosure relate to artificial intelligence technology fields such as vehicle networking and intelligent cockpits

[0021] Among them, Artificial Intelligence (AI) is a new technical science that studies, develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence.

[0022] The concept of vehicle networking originates from the Internet of Things, that is, the Internet of Vehicles. It takes the vehicles in motion as the information perception objects and, with the help of the new generation of information and communication technologies, realizes the network connection between vehicle and vehicle, vehicle and person, vehicle and road, vehicle and service platform, improves the overall intelligent driving level of vehicles, provides users with safe, comfortable, intelligent, and efficient driving experiences and traffic services, and at the same time improves traffic operation efficiency and the intelligent level of social traffic services.

[0023] The intelligent cockpit, by carrying intelligent / networked in-vehicle devices or services, such as digital instrument clusters, central control large screens, streaming media rearview mirrors, head-up displays, intelligent air conditioners, intelligent ambient lights, voice, visual interactions, etc., makes the interaction content between "human-vehicle-road-cloud" richer and the information of each system fully integrated; it can achieve personalized definition and provide a better experience for drivers and passengers.

[0024] It should be noted that the voice control method described in the embodiments of the present disclosure can be applied to the interaction scenario between users and voice devices, or can be applied to any other possible scenarios where a certain type of voice is used to interact with voice devices, and there is no limitation on this.

[0025] Among them, a voice device refers to a device that can respond to the voice of a user and perform some basic operations and actions. The voice device can specifically be, for example, a smart speaker, a smart voice assistant, etc., and there is no limitation on this.

[0026] For example, a user can perform voice wake-up control on a voice device through voice to achieve the interaction between the user and the voice device. Or, the type of the certain type of voice can be adaptively set. For example, the type of noise generated by a television device, or the voice type of the voice control command generated by an air conditioner device. When the type of the certain type of voice is the type of noise generated by a television device, the voice device can identify the type of the noise to call the corresponding noise reduction algorithm to send a noise reduction command to the television device. When the type of the certain type of voice is the voice type of the voice control command generated by an air conditioner device, the voice device can identify the type of the voice control command to call the corresponding air conditioner control algorithm to send a corresponding control command to the air conditioner device, and there is no limitation on this.

[0027] In the embodiments of the present disclosure, the voice control method is exemplified in the interaction scenario between the user and the voice device, and no limitation is imposed thereon.

[0028] As Figure 1 shown, the voice control method includes:

[0029] S101: Obtain the voice to be processed.

[0030] Among them, when in the above-mentioned interaction scenario between the user and the intelligent device, the voice detected in the scenario can be referred to as the voice to be processed. The voice to be processed can be a single voice segment collected by an electronic device with a recording function such as a mobile phone or a microphone, or can also be a partial voice segment in the collected voice. For example, the collected voice can be cut to obtain multiple voice segments, and all or part of the multiple voice segments can be used as the voice to be processed, and no limitation is imposed thereon.

[0031] In the embodiments of the present disclosure, to obtain the voice to be processed, a corresponding voice collection module (such as a microphone) can be pre-configured for the voice control device. Then, the microphone can collect multiple voices to be processed in its environment (the multiple voices to be processed can be human voices in its environment, or can also be other voices in the environment, such as the voice broadcast by an electronic device, and no limitation is imposed thereon). Then, corresponding processing can be performed on the voice to be processed to implement the voice control method described in the embodiments of the present disclosure according to the voice to be processed, and no limitation is imposed thereon.

[0032] It should be noted that in the embodiments of the present disclosure, regarding the acquisition, processing, storage, and use of the voice to be processed, the process complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0033] S102: Perform feature analysis on the voice to be processed to obtain a voice feature vector.

[0034] After obtaining the voice to be processed, feature analysis can be performed on the voice to be processed to obtain a voice feature vector, which can be used to describe the acoustic features of the voice to be processed.

[0035] In some embodiments, performing feature analysis on the voice to be processed can be inputting the voice to be processed into a pre-trained convolutional neural network to obtain multiple voice feature vectors output by the convolutional neural network, and no limitation is imposed thereon.

[0036] In other embodiments, feature analysis can also be performed on the voice to be processed to obtain voice features, and then the voice features are mapped to a vector space for vectorization processing to obtain a vector representation in the vector space that can represent the voice features as the voice feature vector, and no limitation is imposed thereon.

[0037] Of course, other arbitrary possible methods can also be used to perform feature parsing on the speech to be processed to obtain a speech feature vector.

[0038] S103: Determine whether the speech to be processed is a target type of speech according to the speech feature vector.

[0039] In the embodiments of the present disclosure, the target type of speech can be, for example, a human type of speech, or can also be a speech of an adaptive configuration type according to the actual speech control scenario requirements. For example, it can be a speech of a noise type generated by a television device, or a speech of a voice control instruction type generated by an air conditioner device. There is no limitation in this regard.

[0040] That is to say, in the embodiments of the present disclosure, it is supported to pre-configure the types that can be used to wake up and interact with the target device. This type can be, for example, a noise type generated by a television device, or a speech of a voice control instruction type generated by an air conditioner device. Then, for the corresponding type, configure one or more segments of speech used to trigger wake-up and interaction control as the target type of speech. This target type of speech can be used to match the collected speech to be processed. When the match passes, the corresponding voice control logic can be triggered. There is no limitation in this regard.

[0041] Correspondingly, in the embodiments of the present disclosure, it can be determined whether the speech to be processed is a target type of speech according to the speech feature vector.

[0042] For example, it can be determined whether the speech to be processed is a human type of speech according to the speech feature vector, or it can be determined whether the speech to be processed is a speech of a noise type generated by a television device according to the speech feature vector, or it can be determined whether the speech to be processed is a speech of a voice control instruction type generated by an air conditioner device according to the speech feature vector. There is no limitation in this regard.

[0043] In some embodiments, to determine whether the speech to be processed is a target type of speech according to the speech feature vector, it can be to determine the similarity between the speech to be processed and the target type of speech, and compare the calculated similarity with a pre-set similarity threshold (the similarity threshold can be adaptively configured according to the actual speech recognition and control scenario). If the similarity is greater than or equal to the similarity threshold, it can be determined that the speech to be processed is a target type of speech.

[0044] In some other embodiments, other arbitrary possible methods can also be used to determine whether the speech to be processed is a target type of speech according to the speech feature vector. There is no limitation in this regard.

[0045] For example, the frequencies corresponding to the voice to be processed and the human voice type can be analyzed. If the frequencies of the voice to be processed and the human voice type are the same, it can be determined that the voice to be processed is a human voice type. Additionally, the timbres of the voice to be processed and the human voice type can be compared. If the timbres of the voice to be processed and the human voice type are the same, it can be determined that the voice to be processed is a human voice type. There is no limitation in this regard.

[0046] Optionally, in some other embodiments, to determine whether the voice type of the voice to be processed is the target voice type based on the voice feature vector, the voice feature vector can be input into a feature matching model to obtain the output result of the feature matching model. This output result can be used to describe the determination situation of whether the voice to be processed is a human voice type. Since the feature matching model is used to determine whether the voice to be processed is the target voice, the determination processing logic can be effectively simplified, the determination efficiency can be effectively improved, and other subjective judgment factors can be avoided, effectively enhancing the objectivity and accuracy of the determination result of the voice to be processed.

[0047] Among them, the feature matching model can be used to determine whether the voice type of the voice to be processed is the target voice type. The feature matching model can be pre-trained based on the voice features of the human voice type. The feature matching model can be an artificial intelligence model, specifically, for example, a neural network model or a machine learning model. Of course, other arbitrary possible artificial intelligence models capable of performing feature matching tasks can also be used. There is no limitation in this regard.

[0048] That is to say, after performing feature analysis on the voice to be processed to obtain the voice feature vector, the voice feature vector can be input into the pre-trained feature matching model. Then, it can be determined whether the voice feature vector is within the feature interval corresponding to the feature matching model. If the voice feature vector is within the feature interval corresponding to the feature matching model, it can be determined that the voice to be processed corresponding to the feature voice is the target voice type. There is no limitation in this regard.

[0049] S104: If the voice to be processed is the target voice type, the target device is awakened and controlled according to the voice to be processed.

[0050] Among them, the device to be awakened and controlled currently, that is, can be referred to as the target device.

[0051] When it is determined that the voice to be processed is the target voice type as described above, the target device can be awakened and controlled according to the voice to be processed. Thus, voice recognition between the user and the target device can be achieved.

[0052] In this embodiment, by obtaining the voice to be processed, parsing the features of the voice to be processed to obtain a voice feature vector, and then judging whether the voice to be processed is a target type voice according to the voice feature vector, and when the voice to be processed is a target type voice, performing wake-up control on the target device according to the voice to be processed. Since the wake-up control of the target device is performed according to the type of the voice to be processed, it can effectively avoid false wake-up caused by other types of voices, effectively improve the accuracy of device wake-up, and effectively improve the effect of voice wake-up control.

[0053] Figure 2 It is a schematic diagram according to the second embodiment of the present disclosure.

[0054] As Figure 2 shown, the voice control method includes:

[0055] S201: Obtain the voice to be processed.

[0056] S202: Parse the features of the voice to be processed to obtain a voice feature vector.

[0057] S203: Judge whether the voice to be processed is a target type voice according to the voice feature vector.

[0058] For the description of S201 - S203, specific reference can be made to the above embodiments, which will not be elaborated here.

[0059] S204: Determine the target control sensitivity corresponding to the target device.

[0060] Among them, the target control sensitivity can be used to describe the sensitivity of the wake-up control of the target device. The higher the target control sensitivity, the more sensitive the target device is to the voice it receives for recognition control. On the contrary, the lower the target control sensitivity, the more delayed the target device is in recognizing and controlling the voice it receives.

[0061] S205: Perform wake-up control on the target device according to the target control sensitivity combined with the voice to be processed.

[0062] In the embodiments of the present disclosure, the target device can be wake-up controlled in combination with the target control sensitivity to improve the user's wake-up control experience. That is, in an application scenario where the target device needs to be frequently woken up (for example, when the user has a high demand for interacting with the target device), the target device can be controlled to maintain a high target control sensitivity, so that the user can quickly and agilely wake up the target device, thereby effectively reducing the wake-up time and effectively improving the wake-up efficiency. Correspondingly, in an application scenario where the target device does not need to be frequently woken up (for example, when the user has no need to interact with the target device), the target device can be controlled to maintain a low target control sensitivity. At this time, the target device is not sensitive to the voice it receives, thereby effectively avoiding false wake-up and effectively improving the user's interaction experience.

[0063] In some embodiments, wake-up controlling the target device according to the target control sensitivity in combination with the voice to be processed may be wake-up controlling the target device in combination with a preset sensitivity threshold (the sensitivity threshold can be adaptively configured according to the user's wake-up control requirements) in combination with the voice to be processed.

[0064] For example, when the target control sensitivity is greater than or equal to the sensitivity threshold, the target device can be woken up according to the voice to be processed, or a corresponding control coefficient can be generated according to the target control sensitivity, and then the target device can be wake-up controlled in combination with the control coefficient and the voice to be processed. There is no limitation in this regard.

[0065] Optionally, in some other embodiments, wake-up controlling the target device according to the target control sensitivity in combination with the voice to be processed may be performing voice segmentation on the voice to be processed to obtain a plurality of voice segments, parsing a plurality of feature sub-vectors corresponding to the plurality of voice segments from the voice feature vector, determining a plurality of voice scores corresponding to the plurality of feature sub-vectors according to the target control sensitivity, and wake-up controlling the target device according to the plurality of voice scores. Thus, voice segmentation of the voice to be processed can be first realized, and then the target device can be wake-up controlled according to the voice scores of the feature sub-vectors of the voice segments obtained from the voice segmentation. Since the voice score can be used to express the similarity between the features of the voice segment and the voice features that can trigger wake-up control of the target device, when wake-up controlling the target device based on the voice score, the voice to be processed with a higher feature similarity can be focused on, thereby effectively ensuring the accuracy of wake-up control of the target device.

[0066] Among them, the feature vectors parsed from the voice feature vector and corresponding to each voice segment can be referred to as feature sub-vectors, that is, the plurality of feature sub-vectors together constitute the voice feature vector.

[0067] That is to say, after obtaining the voice to be processed, the voice to be processed can be segmented into multiple voice segments, and then multiple feature sub-vectors corresponding to the multiple voice segments can be parsed from the voice feature vectors. Then, according to the target control sensitivity, the multiple feature sub-vectors are scored to obtain multiple scoring results corresponding to the multiple feature sub-vectors respectively. This scoring result can be referred to as voice scoring, which can be used to describe the similarity between the features of the voice segment and the voice features that can trigger the wake-up control of the target device. The higher the voice scoring, the easier it is to trigger the wake-up control of the target device. On the contrary, the lower the voice scoring, the less likely it is to trigger the wake-up control of the target device.

[0068] Optionally, in some embodiments, feature parsing can be performed on multiple voice segments to obtain the voice features corresponding to the multiple voice segments respectively. Then, the voice features corresponding to the multiple voice segments can be mapped to the vector space corresponding to the voice feature vectors for vectorization processing to obtain the vector representations in the vector space that can represent the voice features corresponding to the multiple voice segments respectively as the feature sub-vectors.

[0069] In some other embodiments, the voice feature vectors and multiple voice segments can also be input into a pre-trained feature parsing model to obtain multiple feature sub-vectors corresponding to the multiple voice segments output by the feature parsing model, and there is no limitation on this.

[0070] Optionally, in some embodiments, to determine the multiple voice scores corresponding to the multiple feature sub-vectors according to the target control sensitivity, the target control sensitivity and the multiple feature sub-vectors can be input into a pre-trained voice scoring model to obtain the multiple voice scores corresponding to the multiple voice feature vectors output by the pre-trained voice scoring model. Or, according to the target control sensitivity, the multiple similarity values between the multiple feature sub-vectors and the voice features that can trigger the wake-up control of the target device can be determined respectively, and then the determined multiple similarity values are used as the multiple voice scores corresponding to the multiple feature sub-vectors respectively to implement the determination of the multiple voice scores corresponding to the multiple feature sub-vectors according to the target control sensitivity, and there is no limitation on this.

[0071] After determining the multiple voice scores corresponding to the multiple feature sub-vectors according to the target control sensitivity as described above, the target device can be wake-up controlled according to the multiple voice scores.

[0072] Optionally, in some embodiments, waking up and controlling the target device according to the multiple voice scores may be generating a control score according to the multiple voice scores, triggering waking up and controlling the target device when the control score is greater than or equal to the score threshold, and not triggering waking up and controlling the target device when the control score is less than the score threshold. Since the corresponding control score is generated according to the voice to be processed, and not triggering waking up and controlling the target device when the control score is less than the score threshold, false wake-up can be effectively avoided. Moreover, it also supports adjusting and configuring the score threshold to meet the personalized wake-up control requirements of different voice control scenarios.

[0073] Among them, the score used to control the target device can be referred to as the control score.

[0074] In some embodiments, generating a control score according to the multiple voice scores may be using the multiple voice scores as the input of a pre-trained neural network model to obtain the control score output by the pre-trained neural network model.

[0075] Alternatively, it can also be dividing the multiple voice scores into multiple voice score intervals and determining the control scores corresponding to the multiple voice score intervals. Correspondingly, generating a control score according to the multiple voice scores may be determining the score interval corresponding to the voice score and using the control score corresponding to this voice score interval as the control score corresponding to this voice score, and there is no limitation on this.

[0076] For example, the multiple voice scores can be divided into two voice score intervals of (0 - 5) and (5 - 10) according to the size of the voice scores, and it is determined that the control score of the (0 - 5) voice score interval is score A, and the control score of the (5 - 10) voice score interval is score B. Then, generating a control score according to the voice score may be determining that the voice score interval corresponding to the voice score 4 is (0 - 5), and then the control score A of the (0 - 5) voice score interval can be used as the control score corresponding to the voice score 4.

[0077] Among them, the critical value of the pre-set control score can be referred to as the score threshold, and this score threshold can be used to assist in waking up and controlling the target device, that is, triggering waking up and controlling the target device when the control score is greater than or equal to the score threshold, and not triggering waking up and controlling the target device when the control score is less than the score threshold, and there is no limitation on this.

[0078] In some other embodiments, waking up and controlling the target device according to the multiple voice scores may also be generating corresponding wake-up control instructions according to the multiple voice scores, and the target device can determine whether to maintain the wake-up or non-wake-up state in response to the corresponding wake-up control instructions, and there is no limitation on this.

[0079] S206: If the wake-up control of the target device is not triggered, determine the first cumulative count value of the unsuccessful trigger of the wake-up control.

[0080] In the embodiments of the present disclosure, if the wake-up control of the target device is not triggered, the situation where the wake-up control of the target device is not triggered can be cumulatively counted to obtain the corresponding cumulative count value, and this cumulative count value can be referred to as the first cumulative count value, which can be used to describe the number of times the wake-up control of the target device is not triggered.

[0081] S207: If the first cumulative count value is greater than or equal to the first count threshold, adjust the target control sensitivity to the first control sensitivity, and the first control sensitivity is lower than the target control sensitivity.

[0082] Among them, the preset critical value of the number of times the wake-up control of the target device is not triggered can be referred to as the first count threshold, and the first count threshold can be adaptively configured according to the wake-up control requirements of the actual wake-up control scenario, and there is no limitation on this.

[0083] In the embodiments of the present disclosure, after determining the first cumulative count value of the unsuccessful trigger of the wake-up control, the first cumulative count value can be compared with the preset first count threshold. If the first cumulative count value is greater than or equal to the first count threshold, the target control sensitivity can be adjusted, and the adjusted control sensitivity is used as the first control sensitivity, and the first control sensitivity is lower than the target control sensitivity.

[0084] That is to say, when the first cumulative count value is greater than or equal to the first count threshold, it can indicate that the current application scenario does not require frequent wake-up of the target device, that is, the current user's interaction demand for the target device is low. At this time, the target control sensitivity can be adjusted to the first control sensitivity.

[0085] S208: If the first cumulative count value is less than the first count threshold, keep the target control sensitivity.

[0086] That is to say, after determining the first cumulative count value of the unsuccessful trigger of the wake-up control, the first cumulative count value can be compared with the preset first count threshold. If the first cumulative count value is less than the first count threshold, keep the target control sensitivity.

[0087] In the embodiments of the present disclosure, by determining the first cumulative count value of the unsuccessful trigger for wake-up control, when the first cumulative count value is greater than or equal to the first count threshold, the target control sensitivity is adjusted to the first control sensitivity, and when the first cumulative count value is less than the first count threshold, the target control sensitivity is maintained. Thus, when the user's interaction demand for the target device is low, the wake-up sensitivity of the target device can be reduced, so that while effectively ensuring the user's interaction demand for the target device, ineffective wake-up can be effectively avoided, thereby effectively improving the user's interaction experience.

[0088] S209: If a trigger for wake-up control of the target device is detected, determine the second cumulative count value of the successful trigger for wake-up control.

[0089] In the embodiments of the present disclosure, if a successful trigger for wake-up control of the target device is detected, the situations of the successful trigger for wake-up control of the target device can be cumulatively counted to obtain the corresponding cumulative count value, which can be referred to as the second cumulative count value. The second cumulative count value can be used to describe the number of times of the successful trigger for wake-up control of the target device.

[0090] S210: If the second cumulative count value is greater than or equal to the second count threshold, adjust the target control sensitivity to the second control sensitivity, and the second control sensitivity is higher than the target control sensitivity.

[0091] Among them, the preset critical value of the number of successful triggers for wake-up control, which can be referred to as the second cumulative count value, and the second count threshold can be adaptively configured according to the wake-up control requirements of the actual wake-up control scenario, and there is no limitation on this.

[0092] In the embodiments of the present disclosure, after determining the second cumulative count value of the successful trigger for wake-up control, the second cumulative count value can be compared with the preset second count threshold. If the second cumulative count value is greater than or equal to the second count threshold, the target control sensitivity can be adjusted accordingly, and the adjusted control sensitivity is used as the second control sensitivity, and the second control sensitivity is higher than the target control sensitivity.

[0093] That is to say, when the second cumulative count value is greater than or equal to the second count threshold, it can indicate that the current application scenario requires frequent wake-up of the target device, that is, the current user's interaction demand with the target device is high. At this time, the target control sensitivity can be adjusted to the second control sensitivity to meet the user's interaction demand.

[0094] S211: If the second cumulative count value is less than the second count threshold, maintain the target control sensitivity.

[0095] In an embodiment of the present disclosure, after determining the second cumulative count value that successfully triggers wake-up control, the second cumulative count value can be compared with a preset second count threshold. If the second cumulative count value is less than the second count threshold, the target control sensitivity is maintained.

[0096] In an embodiment of the present disclosure, by determining the second cumulative count value that successfully triggers wake-up control, when the second cumulative count value is greater than or equal to the second count threshold, the target control sensitivity is adjusted to the second control sensitivity. When the second cumulative count value is less than the second count threshold, the target control sensitivity is maintained. Thus, when the user's interaction demand for the target device is relatively high, the wake-up sensitivity of the target device can be improved, so that the target device can quickly respond to the user's wake-up control demand, effectively reducing the response time of the target device's wake-up control, and thus enabling efficient wake-up control of the target device, effectively meeting the user's interaction demand, and effectively enhancing the user's experience.

[0097] In this embodiment, by obtaining the voice to be processed, parsing the features of the voice to be processed to obtain a voice feature vector, then judging whether the voice to be processed is a target type voice according to the voice feature vector, and performing wake-up control on the target device according to the target control sensitivity in combination with the voice to be processed to improve the user's wake-up control experience. Then, by determining the first cumulative count value that fails to successfully trigger wake-up control, when the first cumulative count value is greater than or equal to the first count threshold, the target control sensitivity is adjusted to the first control sensitivity. When the first cumulative count value is less than the first count threshold, the target control sensitivity is maintained. Thus, when the user's interaction demand for the target device is relatively low, the wake-up sensitivity of the target device can be reduced, so that while effectively ensuring the user's interaction demand for the target device, ineffective wake-up can be effectively avoided, effectively enhancing the user's interaction experience. And by determining the second cumulative count value that successfully triggers wake-up control, when the second cumulative count value is greater than or equal to the second count threshold, the target control sensitivity is adjusted to the second control sensitivity. When the second cumulative count value is less than the second count threshold, the target control sensitivity is maintained. Thus, when the user's interaction demand for the target device is relatively high, the wake-up sensitivity of the target device can be improved, so that the target device can quickly respond to the user's wake-up control demand, effectively reducing the response time of the target device's wake-up control, and thus enabling efficient wake-up control of the target device, effectively meeting the user's interaction demand, and effectively enhancing the user's experience.

[0098] Figure 3 It is a schematic diagram according to the third embodiment of the present disclosure.

[0099] As Figure 3 shown, the voice control device 30 includes:

[0100] An acquisition module 301, configured to acquire the voice to be processed;

[0101] An analysis module 302, configured to perform feature analysis on the voice to be processed to obtain a voice feature vector;

[0102] A judgment module 303, configured to judge whether the voice to be processed is a target type voice according to the voice feature vector;

[0103] A wake-up module 304, configured to perform wake-up control on the target device according to the voice to be processed when the voice to be processed is a target type voice.

[0104] In some embodiments of the present disclosure, as Figure 4 shown, Figure 4 is a schematic diagram according to the fourth embodiment of the present disclosure. The voice control device 40 includes: an acquisition module 401, an analysis module 402, a judgment module 403, and a wake-up module 404. Among them, the voice control device 40 further includes:

[0105] A first determination module 405, configured to determine a target control sensitivity corresponding to the target device after judging whether the voice to be processed is a target type voice according to the voice feature vector;

[0106] Among them, the wake-up module 404 is specifically configured to:

[0107] Perform wake-up control on the target device according to the target control sensitivity in combination with the voice to be processed.

[0108] In some embodiments of the present disclosure, among them, the wake-up module 404 includes:

[0109] A sub-module 4041 for voice segmentation, configured to perform voice segmentation on the voice to be processed to obtain a plurality of voice segments;

[0110] An analysis sub-module 4042, configured to analyze and obtain a plurality of feature sub-vectors corresponding to the plurality of voice segments respectively from the voice feature vector;

[0111] A determination sub-module 4043, configured to determine a plurality of voice scores corresponding to the plurality of feature sub-vectors respectively according to the target control sensitivity;

[0112] A wake-up sub-module 4044, configured to perform wake-up control on the target device according to the plurality of voice scores.

[0113] In some embodiments of the present disclosure, among them, the wake-up sub-module 4044 is specifically configured to:

[0114] Generate a control score according to the plurality of voice scores;

[0115] If the control score is greater than or equal to the score threshold, trigger wake-up control for the target device;

[0116] If the control score is less than the score threshold, do not trigger wake-up control for the target device.

[0117] In some embodiments of the present disclosure, wherein the target type of voice is a human voice type, and the judgment module 403 is specifically configured to:

[0118] Input the voice feature vector into the feature matching model to obtain the result output by the feature matching model, where the output result describes the judgment of whether the voice to be processed is a human voice type.

[0119] In some embodiments of the present disclosure, the voice control device 40 further includes:

[0120] The second determination module 406 is configured to, after triggering wake-up control for the target device according to the voice to be processed, if the wake-up control for the target device is not triggered, determine the first cumulative number value of the unsuccessful triggering of the wake-up control;

[0121] The first adjustment module 407 is configured to, when the first cumulative number value is greater than or equal to the first number threshold, adjust the target control sensitivity to the first control sensitivity, where the first control sensitivity is lower than the target control sensitivity, and when the first cumulative number value is less than the first number threshold, maintain the target control sensitivity.

[0122] In some embodiments of the present disclosure, the voice control device 40 further includes:

[0123] The third determination module 408 is configured to, after triggering wake-up control for the target device according to the voice to be processed, if the wake-up control for the target device is triggered, determine the second cumulative number value of the successful triggering of the wake-up control;

[0124] The second adjustment module 409 is configured to, when the second cumulative number value is greater than or equal to the second number threshold, adjust the target control sensitivity to the second control sensitivity, where the second control sensitivity is higher than the target control sensitivity, and when the second cumulative number value is less than the second number threshold, maintain the target control sensitivity.

[0125] It can be understood that the voice control device 40 in this embodiment Figure 4 and the voice control device 40 in the above embodiment, the acquisition module 401 and the acquisition module 301 in the above embodiment, the parsing module 402 and the parsing module 302 in the above embodiment, the judgment module 403 and the judgment module 303 in the above embodiment, and the wake-up module 404 and the wake-up module 304 in the above embodiment may have the same functions and structures.

[0126] It should be noted that the foregoing explanation of the voice control method also applies to the voice control device of this embodiment.

[0127] In this embodiment, by obtaining the voice to be processed and performing feature analysis on the voice to be processed to obtain a voice feature vector, and then judging whether the voice to be processed is a target type voice according to the voice feature vector, and when the voice to be processed is a target type voice, performing wake-up control on the target device according to the voice to be processed. Since the wake-up control of the target device is performed according to the type of the voice to be processed, it can effectively avoid false wake-up caused by other types of voices, effectively improve the accuracy of device wake-up, and effectively improve the effect of voice wake-up control.

[0128] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0129] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device for implementing the voice control method of the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0130] As Figure 5 shown, the device 500 includes a computing unit 501, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 505 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0131] Multiple components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0132] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the voice control method. For example, in some embodiments, the voice control method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the voice control method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the voice control method by any other suitable means (e.g., by means of firmware).

[0133] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0134] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0135] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0136] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0137] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0138] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs that run on the respective computers and have a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server may also be a server of a distributed system, or a server combined with a blockchain.

[0139] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this is not limited herein.

[0140] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A voice control method, comprising: Obtaining the voice to be processed; Performing feature analysis on the voice to be processed to obtain a voice feature vector; Judging whether the voice to be processed is a target type voice according to the voice feature vector; If the voice to be processed is the target type voice, determining a target control sensitivity corresponding to the target device, where the target control sensitivity is used to describe the sensitivity of the wake-up control of the target device. The higher the target control sensitivity, the more sensitive the target device is to the voice it receives for recognition control, and the lower the target control sensitivity, the more delayed the target device is in recognizing and controlling the voice it receives; Performing voice segmentation on the voice to be processed to obtain a plurality of voice segments; Parsing from the voice feature vector to obtain a plurality of feature sub-vectors respectively corresponding to the plurality of voice segments; Determining a plurality of voice scores respectively corresponding to the plurality of feature sub-vectors according to the target control sensitivity; Generating a control score according to the plurality of voice scores; If the control score is greater than or equal to a score threshold, triggering wake-up control for the target device; If the control score is less than the score threshold, not triggering wake-up control for the target device.

2. The method according to claim 1, wherein, The target type voice is a human voice type. Among them, judging whether the voice to be processed is the target type voice according to the voice feature vector includes: Inputting the voice feature vector into a feature matching model to obtain the result output by the feature matching model, where the output result describes the judgment situation of whether the voice to be processed is the human voice type.

3. The method according to claim 1, after performing wake-up control on the target device according to the voice to be processed, further comprising: If the wake-up control for the target device is not triggered, determining a first cumulative count value of the unsuccessful trigger of the wake-up control; If the first cumulative count value is greater than or equal to a first count threshold, adjusting the target control sensitivity to a first control sensitivity, where the first control sensitivity is lower than the target control sensitivity; If the first cumulative count value is less than the first count threshold, maintaining the target control sensitivity.

4. The method according to claim 1, after performing wake-up control on the target device according to the voice to be processed, further comprising: If the wake-up control for the target device is triggered, determining a second cumulative count value of the successful trigger of the wake-up control; If the second cumulative count value is greater than or equal to a second count threshold, adjusting the target control sensitivity to a second control sensitivity, where the second control sensitivity is higher than the target control sensitivity; If the second cumulative count value is less than the second count threshold, maintaining the target control sensitivity.

5. A voice control device, comprising: An acquisition module, configured to acquire the voice to be processed; An analysis module, configured to perform feature analysis on the voice to be processed to obtain a voice feature vector; A judgment module, configured to judge whether the to-be-processed voice is a target type voice according to the voice feature vector; A determination module, configured to, when the to-be-processed voice is the target type voice, determine a target control sensitivity corresponding to a target device, where the target control sensitivity is used to describe the sensitivity degree of the wake-up control of the target device. The higher the target control sensitivity is, the more sensitive the target device is to the voice recognition control it receives. The lower the target control sensitivity is, the more delayed the target device is in the voice recognition control it receives; A wake-up module, including: A sub-module for voice segmentation, configured to perform voice segmentation on the to-be-processed voice to obtain a plurality of voice segments; An analysis sub-module, configured to analyze, from the voice feature vectors, a plurality of feature sub-vectors respectively corresponding to the plurality of voice segments; A determination sub-module, configured to determine a plurality of voice scores respectively corresponding to the plurality of feature sub-vectors according to the target control sensitivity; A wake-up sub-module, configured to generate a control score according to the plurality of voice scores; if the control score is greater than or equal to a score threshold, trigger wake-up control of the target device; if the control score is less than the score threshold, do not trigger wake-up control of the target device.

6. The device according to claim 5, wherein The target type voice is a human voice type. Wherein, the judgment module is specifically configured to: Input the voice feature vector into a feature matching model to obtain a result output by the feature matching model, where the output result describes the judgment situation of whether the to-be-processed voice is the human voice type.

7. The device according to claim 5, further including: A first determination module, configured to, after performing wake-up control on the target device according to the to-be-processed voice, if the wake-up control of the target device is not triggered, determine a first cumulative number value of the failed wake-up control trigger; A first adjustment module, configured to, when the first cumulative number value is greater than or equal to a first number threshold, adjust the target control sensitivity to a first control sensitivity, where the first control sensitivity is lower than the target control sensitivity, and when the first cumulative number value is less than the first number threshold, maintain the target control sensitivity.

8. The device according to claim 5, further including: A second determination module, configured to, after performing wake-up control on the target device according to the to-be-processed voice, if the wake-up control of the target device is triggered, determine a second cumulative number value of the successful wake-up control trigger; A second adjustment module, configured to, when the second cumulative number value is greater than or equal to a second number threshold, adjust the target control sensitivity to a second control sensitivity, where the second control sensitivity is higher than the target control sensitivity, and when the second cumulative number value is less than the second number threshold, maintain the target control sensitivity.

9. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.

11. A computer program product, comprising a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Awakening sensitivity adjustment method and device, and terminal

    CN109672775A

  • Method for voice wake-up, electronic equipment, storage medium, and program

    CN114038457A

  • Dynamic thresholds for always listening speech trigger

    US20160077794A1

  • Speech classification of audio for wake on voice

    US20190043529A1

  • System and method for updating an adaptive speech recognition model

    US9697822B1