Speech Recognition Method, Apparatus, Device, System and Computer Readable Storage Medium
By acquiring and analyzing the channel temperature characteristics and voiceprint characteristics in the reflected voice signal, combined with the neural network model, the problem that the voiceprint unlocking method is easily attacked by forged voiceprints is solved, and the security and recognition accuracy of the PC are improved.
Patent Information
- Application Number
- CN202011300780.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-03
- Filing Date
- 2020-11-19
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2040-11-19
AI Technical Summary
The existing voiceprint unlocking methods are easily attacked by forged voiceprint data, resulting in reduced PC security.
By obtaining the reflected voice signal, using the channel temperature characteristics and voiceprint characteristics in the reflected voice signal, combining the neural network model to determine whether the sound source is the target user, and improve recognition accuracy.
It effectively avoids attacks of forged voiceprint data, and improves the security and recognition accuracy of voiceprint unlocking PCs.
Smart Images

Figure CN114067802B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technologies, and in particular, to a voice recognition method, apparatus, device, system, and computer-readable storage medium. Background Art
[0002] In recent years, with the rapid development of Internet technologies, electronic devices such as personal computers (PCs) have become an indispensable part of people's lives. To improve the security of PCs and ensure user privacy, researchers have begun to attempt to use voiceprint recognition technology to unlock PCs, with the aim of enabling only the target user to unlock the corresponding PC and operate it. Thus, voiceprint unlocking has emerged.
[0003] The principle of voiceprint unlocking is mainly as follows: voiceprint data is used to verify the legitimacy of a user, that is, by comparing the voiceprint data input by the user with the voiceprint data of the target user to determine whether the user is the target user. However, some illegal users can forge the voiceprint data of the target user through machines. For example, they can forge the voiceprint data of the target user through voice synthesis technology or record the voiceprint data of the target user using a recording device, and then use the forged voiceprint data to unlock the corresponding PC, which will pose a great security risk to the PCs using the above voiceprint unlocking method. Summary of the Invention
[0004] This application discloses a voice recognition method, apparatus, device, system, and computer-readable storage medium, which can improve the accuracy of voice recognition and further enhance the security of unlocking a PC with a voiceprint.
[0005] In a first aspect, this application provides a voice recognition method, which includes the following steps:
[0006] Obtain a reflected voice signal, where the reflected voice signal is a sound signal obtained after the original voice signal emitted by a sound source is reflected by a sound reflection device; and determine whether the sound source comes from a target user according to the reflected voice signal.
[0007] By implementing the method described in the first aspect, by obtaining the reflected voice signal obtained after one or more reflections of the original sound signal and determining whether the sound source comes from the target user according to the reflected voice signal, compared with the practice of determining whether the sound source comes from the target user according to the original voice signal emitted by the sound source, the former can avoid the loopholes in the existing methods for detecting the original voice signal, thereby improving the accuracy of voice recognition.
[0008] In a possible implementation, determining whether the sound source comes from the target user based on the reflected voice signal includes: obtaining the characteristics of the transmission medium (referred to as the first channel) between the sound source and the microphone according to the reflected voice signal; and determining whether the sound source comes from the target user according to the characteristics of the first channel.
[0009] According to the above approach, when determining whether the sound source comes from the target user, it will be determined by the characteristics of the transmission medium between the sound source and the microphone, thereby eliminating the security vulnerability of using forged voiceprint data to unlock the device in current voice recognition and improving the security of voice recognition.
[0010] In a possible implementation, the characteristics of the first channel include the temperature characteristics of the first channel. Obtaining the characteristics of the first channel according to the reflected voice signal includes: determining the transmission speed of the reflected voice signal in the first channel; and obtaining the temperature characteristics of the first channel according to the transmission speed of the reflected voice signal in the first channel.
[0011] It can be understood that when the sound source is a real person, the temperature of the first channel is affected by two factors, namely the oral temperature of the real person and the current air temperature, so it will be between the current air temperature and the oral temperature of the real person; while when the sound source is not a real person (for example, a tape recorder), the temperature of the first channel is the current air temperature. Therefore, it is possible to determine whether the sound source is a real person according to the temperature of the first channel. In addition, since the oral temperatures of different people are different, it is also possible to determine whether the sound source comes from the target user according to the different temperatures of the first channel corresponding to different users.
[0012] In a possible implementation, determining the transmission speed of the sound in the first channel according to the reflected voice signal includes: determining the time period during which the reflected voice signal is transmitted between the microphone and the sound source; and determining the transmission speed of the reflected voice signal in the first channel according to the time period during which the reflected voice signal is transmitted between the microphone and the sound source.
[0013] In the above implementation, according to the time period during which the reflected voice signal is transmitted between the microphone and the sound source, the transmission speed of the reflected voice signal in the first channel is determined, thereby determining the temperature characteristics of the first channel. Then, combined with the fact that the temperature of the first channel is different when the sound source comes from the target user and when the sound source does not come from the target user, it is determined whether the sound source comes from the target user. It can be seen that the above implementation can avoid the situation where a non-target user forges the voiceprint information of the target user through an electronic device (for example, a tape recorder) to unlock the PC, thereby improving the accuracy of voice recognition and the security of the PC unlocked by using this method.
[0014] In a possible implementation, obtaining the reflected speech signal includes: obtaining a speech signal; filtering the speech signal to obtain an original speech signal and a reflected speech signal. It can be understood that since the intensity of the sound weakens after reflection, the signal intensity of the reflected speech signal is weaker than that of the original speech signal. Then, this property can be used to separate the original speech signal and the reflected speech signal from the speech signal.
[0015] In a possible implementation, determining the speed at which the reflected speech signal travels in the first channel includes: determining a first time when the first sound signal is received and a second time when the second sound signal is received, where the reflected speech signal includes the first sound signal and the second sound signal, the first sound signal is the signal obtained by reflecting the original speech signal by a sound reflection device, and the second sound signal is the signal obtained by reflecting the first sound signal by a sound source; determining the speed at which the reflected speech signal travels in the first channel according to the first time and the second time.
[0016] In a possible implementation, determining the first time when the first sound signal is received and the second time when the second sound signal is received includes: extracting the features of the original speech signal, where the features of the original speech signal include a first feature, and the first feature is used to indicate the reflected speech signal corresponding to the original speech signal after one or more reflections; determining the first sound signal and the second sound signal with the first feature according to the features of the original speech signal, so as to determine the first time and the second time.
[0017] In a possible implementation, the features of the original speech signal further include a second feature, and the second feature is used to reflect the signal intensity of the original speech signal. Determining the first sound signal and the second sound signal with the first feature according to the features of the original speech signal, so as to determine the first time and the second time, includes: obtaining a first signal intensity and a second signal intensity according to the second feature, where the first signal intensity is used to determine the signal intensity of the first sound signal, and the second signal intensity is used to determine the signal intensity of the second sound signal; determining the first sound signal and the second sound signal according to the first feature, the first signal intensity, and the second signal intensity; determining the first time and the second time according to the first sound signal and the second sound signal.
[0018] It can be understood that since the reflected speech signal is obtained by reflecting the original speech signal one or more times, the reflected speech signal and the original speech signal have the same first features (including phoneme features, voiceprint features, etc.), and the signal intensity of the reflected speech signal is less than that of the original speech signal, and the more the number of reflections, the smaller the signal intensity. Therefore, the first sound signal and the second sound signal can be extracted according to the first feature and the second feature, so as to determine the first time and the second time, and further determine the temperature feature of the first channel.
[0019] In a possible implementation, the above method further includes: extracting a voiceprint feature from the original voice signal and / or the reflected voice signal; judging whether it is from the target user according to the feature of the first channel, including: judging whether the sound source is from the target user according to the feature of the first channel and the voiceprint feature.
[0020] It can be understood that since the voiceprint feature of the sound source is included in the voice signal, the original voice signal, and the reflected voice signal, the voiceprint feature can be extracted from the voice signal (i.e., the original voice signal and the reflected voice signal), or from the original voice signal, or from the reflected voice signal. It can also be understood that on the basis of the feature of the first channel, combining the voiceprint feature to judge whether the sound source is from the target user further improves the accuracy of voice recognition and the security of the PC unlocked by using this method.
[0021] In a possible implementation, judging whether the sound source is from the target user according to the feature of the first channel and the voiceprint feature includes: comparing the temperature feature of the first channel with the standard temperature to obtain a first matching degree; comparing the voiceprint feature with the standard voiceprint feature to obtain a second matching degree; judging whether the sound source is from the target user according to the first matching degree and the second matching degree.
[0022] In the above approach, when judging whether the sound source is from the target user, the temperature feature of the first channel and the voiceprint feature are simultaneously considered, thereby further improving the accuracy of the recognition result.
[0023] In a possible implementation, the standard temperature is related to the current air temperature.
[0024] Since the temperature of the first channel is affected by the oral temperature of a real person and the current air temperature, making the standard temperature related to the current air temperature is beneficial to improving the accuracy of the recognition result.
[0025] In a possible implementation, the standard voiceprint feature is set by the user.
[0026] By allowing the user to set the standard voiceprint feature, it is beneficial to improve the flexibility of the technical solution of the present application.
[0027] In a possible implementation, the above method further includes: in response to a voiceprint setting request input by a target user, prompting the target user to input voiceprint setting information, where the voiceprint setting information includes the voiceprint information of the target user, the current air temperature, and the distance of the first channel; receiving the voiceprint setting information input by the target user, and obtaining a standard temperature and a standard voiceprint feature according to the voiceprint setting information. It can be understood that when the target user wants to use the voiceprint to unlock the PC, it is necessary to first perform relevant settings for the voiceprint lock in the PC to obtain the standard temperature and the standard voiceprint feature, and then judge whether the sound source comes from the target user by calculating the first matching degree and the second matching degree.
[0028] In a possible implementation, the first feature library includes at least one set of corresponding temperature features of the first channel and the air temperature, and at least one set of corresponding temperature features of the first channel and the air temperature includes the standard temperature and the current air temperature. Comparing the temperature feature of the first channel with the standard temperature includes: determining the corresponding standard temperature according to the current air temperature; comparing the temperature feature of the first channel with the standard temperature.
[0029] In a possible implementation, the second feature library includes one or more voiceprint features, and each voiceprint feature corresponds to a target user.
[0030] In a possible implementation, judging whether the sound source comes from the target user according to the first matching degree and the second matching degree includes: when the first matching degree is greater than the first threshold and the second matching degree is greater than the second threshold, determining that the sound source comes from the target user, where the first threshold or the second threshold is set by the user.
[0031] In a possible implementation, the first threshold is different from the second threshold, and the above method further includes: obtaining the recognition requirement input by the target user, and updating the first threshold, or the second threshold, or the first threshold and the second threshold according to the recognition requirement.
[0032] It can be seen that the user can adjust the first threshold and the second threshold according to their own needs, making the unlocking process more in line with the user's usage. For example, when the user hopes that the process of unlocking the PC is more secure, the first threshold and the second threshold can be appropriately increased. Another example is that when the user hopes to easily unlock, the first threshold and the second threshold can be appropriately decreased.
[0033] In a possible implementation, the feature of the first channel is obtained according to the reflected voice signal and the channel feature extraction model, where the channel feature extraction model is obtained by training a neural network with multiple training sets, and each training set includes a reflected voice signal sample and its corresponding temperature feature sample of the first channel.
[0034] In a possible implementation, the voiceprint feature of the sound source is obtained based on the original voice signal and the voiceprint feature extraction model, where the voiceprint feature extraction model is obtained by training a neural network with multiple original voice signal samples and their corresponding voiceprint feature samples.
[0035] In the above implementation, the model obtained by training with a neural network can more efficiently and accurately extract the features of the first channel and the voiceprint feature of the sound source.
[0036] In a possible implementation, determining whether the sound source is from the target user based on the reflected voice signal includes: obtaining the features of the second channel according to the reflected voice signal, where the second channel is the transmission medium between the microphone and the sound reflection device; determining whether the sound source is from the target user according to the features of the second channel.
[0037] It can be understood that similar to the temperature feature of the first channel, when the sound source is a real person, the temperature of the second channel is also affected by two factors, the oral temperature of the real person and the current air temperature, so it will be between the current air temperature and the oral temperature of the real person; when the sound source is a non-real person (for example, a tape recorder), the temperature of the second channel is the current air temperature, so it is possible to determine whether the sound source is a real person according to the temperature of the second channel. In addition, since the oral temperatures of different people are different, it is also possible to determine whether the sound source is from the target user according to the different temperatures of the second channel corresponding to different users. Thus, the security vulnerability of using forged voiceprint data to unlock the device in current voice recognition is eliminated, and the security of voice recognition is improved.
[0038] In a possible implementation, the material of the sound reflection device includes aluminum or aluminum alloy or glass.
[0039] It can be seen that without adding sensors, the reflected voice signal is obtained through the sound reflection device, so that the temperature feature of the first channel can be obtained. This implementation can improve the correctness of voice recognition, and thus further improve the security of the PC unlocked using this method.
[0040] In a second aspect, the present application provides a voice recognition device, which includes an acquisition unit and a recognition unit.
[0041] The acquisition unit is used to acquire the reflected voice signal, which is the voice signal obtained after the original voice signal emitted by the sound source is reflected by the sound reflection device; the recognition unit is used to determine whether the sound source is from the target user according to the reflected voice signal.
[0042] In a possible implementation, the recognition unit is specifically configured to: obtain the characteristics of the transmission medium (referred to as the first channel) between the sound source and the microphone according to the reflected speech signal; and determine whether the sound source comes from the target user according to the characteristics of the first channel.
[0043] In a possible implementation, the characteristics of the first channel include the temperature characteristics of the first channel. The recognition unit is specifically configured to: determine the transmission speed of the reflected speech signal in the first channel; and obtain the temperature characteristics of the first channel according to the transmission speed of the reflected speech signal in the first channel.
[0044] In a possible implementation, the recognition unit is specifically configured to: determine the time period during which the reflected speech signal is transmitted between the microphone and the sound source; and determine the transmission speed of the reflected speech signal in the first channel according to the time period.
[0045] In a possible implementation, the acquisition unit is specifically configured to: acquire a speech signal; and filter the speech signal to obtain an original speech signal and a reflected speech signal.
[0046] In a possible implementation, the speech recognition device further includes a voiceprint feature extraction unit, which is configured to extract voiceprint features from the original speech signal and / or the reflected speech signal. The recognition unit is specifically configured to: determine whether the sound source comes from the target user according to the characteristics of the first channel and the voiceprint features.
[0047] In a possible implementation, the recognition unit is configured to: compare the temperature characteristics of the first channel with a standard temperature to obtain a first matching degree; compare the voiceprint features with standard voiceprint features to obtain a second matching degree; and determine whether the sound source comes from the target user according to the obtained first matching degree and second matching degree.
[0048] In a possible implementation, the standard temperature is related to the current air temperature.
[0049] In a possible implementation, the standard voiceprint features are set by the user.
[0050] In a possible implementation, the recognition unit is specifically configured to: when the first matching degree is greater than a first threshold and the second matching degree is greater than a second threshold, determine that the sound source comes from the target user, where the first threshold or the second threshold is set by the user.
[0051] In a possible implementation, the characteristics of the first channel are obtained according to the reflected speech signal and a channel feature extraction model, where the channel feature extraction model is obtained by training a neural network with multiple training sets.
[0052] In a possible implementation, the recognition unit is further configured to: obtain the characteristics of the second channel according to the reflected voice signal, where the second channel is the transmission medium between the microphone and the sound reflection device; and determine whether the sound source comes from the target user according to the characteristics of the second channel.
[0053] In a possible implementation, the material of the sound reflection device includes aluminum or aluminum alloy or glass.
[0054] In a third aspect, the present application provides an electronic device, which includes a processor and a memory, and the processor executes the code in the memory to implement the method described in the first aspect or any optional manner of the first aspect.
[0055] In a fourth aspect, the present application provides a computer-readable storage medium, storing computer instructions, and the computer instructions are used to implement the method described in the first aspect or any optional manner of the first aspect.
[0056] In a fifth aspect, the present application provides an identification system, which includes a microphone, a sound reflection device and a voice recognition device, and the voice recognition device is used to execute the method described in the first aspect or any optional manner of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 is a scenario of voiceprint unlocking a PC provided by the present application;
[0059] Figure 2A is a possible voiceprint unlocking method provided by the present application;
[0060] Figure 2B is another possible voiceprint unlocking method provided by the present application;
[0061] Figure 3 is a schematic structural diagram of an identification system provided by the present application;
[0062] Figure 4A is a schematic diagram of a possible deployment manner of an identification system provided by the present application;
[0063] Figure 4B is a schematic diagram of another possible deployment manner of an identification system provided by the present application;
[0064] Figure 4CIt is a schematic diagram of another possible deployment method of the recognition system provided by this application;
[0065] Figure 5 It is a schematic flowchart of a voice recognition method provided by this application;
[0066] Figure 6 It is a specific flowchart of a voice recognition device 230 provided by this application for determining the speed of sound transmission in the first channel according to the reflected voice signal;
[0067] Figure 7A It is a schematic diagram of a possible deployment method of a sound source, a sound reflection device 210, and a microphone 220 provided by this application;
[0068] Figure 7B It is a schematic diagram of another possible deployment method of a sound source, a sound reflection device 210, and a microphone 220 provided by this application;
[0069] Figure 8 It is a specific flowchart of a voice recognition device 230 provided by this application for determining whether the sound source comes from the target user according to the temperature characteristics of the first channel;
[0070] Figure 9 It is a schematic structural diagram of a channel feature extraction model provided by this application;
[0071] Figure 10 It is a schematic structural diagram of a voiceprint feature extraction model provided by this application;
[0072] Figure 11 It is a schematic structural diagram of a voice recognition device provided by this application;
[0073] Figure 12 It is a schematic structural diagram of an electronic device provided by this application. Detailed implementation manners
[0074] Next, the technical solutions provided by this application will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the embodiments described in this application are only a part of the embodiments, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this invention.
[0075] To facilitate understanding of the technical solutions provided by this application, the terms involved in this application are first explained:
[0076] Voice recognition, also known as speaker recognition, is a technology that converts the voice signal emitted by a speaker into an electrical signal and then uses a computer to recognize the electrical signal to determine the identity of the speaker.
[0077] Target user: refers to a person who has the permission to use the target intelligent terminal, and this target user can legally unlock, set, etc. the target intelligent terminal.
[0078] Introduce the main application scenarios involved in this application: scenarios of unlocking electronic devices using voice recognition technology. For example, using voice to unlock desktop computers, laptops, tablets, game consoles, smart speakers, etc. Subsequently, for the convenience of introducing the solution, the example of unlocking a personal computer (PC) including desktop computers and laptops will be used for illustration, but the types of electronic devices protected by this application are not limited.
[0079] In recent years, with the rapid development of Internet technology, intelligent terminals have gradually changed people's lifestyles and living habits and have become an irreplaceable part of people's lives, especially PCs (such as desktop computers, laptops, etc.). The emergence of PCs has brought great convenience to people's lives and work. For example, people can work, socialize, shop, read, etc. on laptops anytime and anywhere. Therefore, more and more personal data is stored on users' PCs. To improve the security of PCs and ensure users' privacy, researchers have begun to try to use voiceprint recognition technology to unlock PCs, aiming to achieve the goal that only the target user can unlock the corresponding PC and operate on this PC.
[0080] Please refer to Figure 1 , Figure 1 which shows a scenario of unlocking a PC with a voiceprint. Figure 1 The PC 110 in is a personal device of a user who has the permission to use the PC. For example, it is the personal device of user A, and user A has pre-set a voiceprint lock on the PC 110. When the target user 120 wants to use this PC 110, they need to first unlock the PC 110 through the voiceprint lock on the PC 110. The following gives two common unlocking methods:
[0081] Method 1: As Figure 2AAs shown, a microphone is provided on the PC 110, and a voice acquisition chip is provided in the microphone. The voice acquisition chip is used to pick up voice data. After the target user 120 sends an unlock request to the PC 110, the processor of the PC 110 generates text data of a dynamic password and displays the text data of the dynamic password (for example, "Open Sesame") on the screen. The target user 120 reads out the dynamic password according to the text data displayed on the screen. Then, the microphone picks up the voice data of the dynamic password read out by the target user 120 and sends it to the processor of the PC 110. The processor of the PC 110 performs voiceprint extraction processing on the voice data of the target user 120 to obtain the voiceprint data of the target user 120. Then, the voiceprint data of the target user 120 is matched with the preset voiceprint data. If the voiceprint data of the target user 120 matches the preset voiceprint data, it is determined that the target user 120 is user A, and at this time, the PC is unlocked.
[0082] It should be understood that during the process of user A setting a voiceprint lock on the PC 110, the PC 110 may record multiple voice data of user A (for example, user A reads out the text data of multiple passwords displayed on the screen of the PC 110 in sequence), and then analyzes the voiceprint characteristics of these voice data to obtain the voiceprint data of user A, that is, the preset voiceprint data.
[0083] Method 2: On the basis of Method 1, an airflow sensor is additionally provided in the microphone in Method 2. As Figure 2B shown, when the target user 120 reads out the password, the voice acquisition chip in the microphone picks up the voice data of the target user 120, and the airflow sensor in the microphone collects the airflow data when the target user 120 reads out the password. Then, the microphone sends the voice data of the target user 120 and the airflow data of the target user to the processor of the PC 110. The processor of the PC 110 matches the voice data of the target user with the preset voice data and matches the airflow data of the target user (for example, the magnitude of the airflow) with the preset airflow data. When the voice data of the target user matches the preset voice data and the airflow data of the target user matches the preset airflow data, it is determined that the target user is user A, and at this time, the PC 110 is unlocked.
[0084] It can be seen that the voiceprint unlocking method described in Method 1 determines whether the target user 120 is User A by matching the voiceprint data of the target user 120 and the preset voiceprint data. However, the target user 120 can also unlock the PC 110 by playing forged voice data. For example, the target user 120 can use an audio converter to convert the voiceprint data of the target user into the voiceprint data of User A; the target user 120 can also use a voice synthesizer to synthesize the voiceprint data of User A; the target user 120 can also use a recorder to replay the voice data of User A to the PC 110, etc. In these cases, Method 1 will not be able to accurately determine the identity of the target user 120, thus posing a security risk to the PC 110.
[0085] The voiceprint unlocking method described in Method 2 determines whether the target user 120 is User A through dual recognition. Specifically, compared with Method 1, the airflow data of the target user 120 and the preset airflow data are also matched in Method 2. Since there are significant differences between the airflow when a real person speaks and the airflow when an electronic device speaks, when the airflow data of the sound source does not match the preset airflow data, the sound source may be an electronic device, thus determining that the target user 120 is not User A. It can be seen that Method 2 makes up for the inability of Method 1 to identify forged voice data and greatly improves the security of voiceprint unlocking. However, it can be understood that the distance between the sound source and the airflow sensor has a great influence on the airflow characteristics. For example, when the sound source is close to the microphone, more airflow is collected by the airflow sensor, and forged voice data can be more accurately identified through the airflow data; while when the sound source is far from the microphone, less airflow is collected by the airflow sensor, and even if the sound source is User A, it may cause the airflow data of the sound source to not match the preset airflow data, thus affecting the accuracy rate of voiceprint recognition.
[0086] To solve the above problems, the present application provides an identification system, which can further improve the accuracy of voice recognition, thereby improving the security of unlocking the PC with voiceprint. As Figure 3 shown, the identification system includes a sound reflection device 210, a microphone 220, and a voice recognition device 230.
[0087] The sound reflection device 210 is used to reflect a part of the sound signal emitted by the sound source to the microphone 220, that is, to send the reflected sound signal to the microphone 220. Among them, the sound reflection device 210 can be a physical material with good sound reflection effect. Good sound reflection effect means that when the sound absorption coefficient of the physical material for sound waves of 4000 hertz (Hz) is less than 0.05, this application considers that the physical material has a good sound reflection effect. For example, materials such as aluminum, aluminum alloy, and glass. Therefore, the sound reflection device 210 can specifically be an aluminum foil sheet, a thin sheet made of aluminum alloy, a glass plate, etc. It should be noted that when actually selecting the material of the sound reflection device 210, not only the effect of reflecting sound but also the specific application scenario needs to be considered. For example, for a PC with a low price, a glass plate with a lower cost can be selected as the sound reflection device 210. Another example is that for a PC with high anti-drop performance, an aluminum foil sheet can be selected as the sound reflection device, and so on.
[0088] In a specific embodiment, the sound reflection device 210 can be deployed on the PC to be unlocked. Specifically, it can be deployed on the display of the PC to be unlocked. For example, the frame of the display of the PC to be unlocked is made of aluminum alloy; the sound reflection device 210 can also be deployed on the keyboard of the PC to be unlocked. For example, the keyboard housing of the PC to be unlocked is made of aluminum alloy material; the sound reflection device 210 can also be deployed on the host of the PC to be unlocked. For example, the host housing of the PC to be unlocked is made of aluminum alloy material, etc. No specific limitation is made here.
[0089] The microphone 220 is used to pick up another part of the sound signal emitted by the sound source, as well as the reflected sound signal emitted by the sound reflection device 210, and send these picked-up voice signals to the voice recognition device 230. The microphone 220 can be deployed on the PC to be unlocked. For example, the microphone 220 is built into the cover plate around the keyboard of the PC to be unlocked (as Figure 2A or Figure 2B the shown microphone). The microphone 220 can also be deployed at a position relatively close to the PC. For example, a conference microphone, wireless headphones, a head-mounted microphone, etc. This application does not limit this.
[0090] The voice recognition device 230 is used to perform voice recognition on the voice signals picked up by the microphone 220, so as to determine whether the user corresponding to the sound source is the target user, that is, the user who has the permission to use this PC. The deployment of the voice recognition device 230 is flexible. It can be deployed on a local computing node (for example, a PC), as Figure 4A shown; it can also be deployed on a remote computing node, as Figure 4B shown; it can also be partly deployed on a local computing node and partly deployed on a remote computing node, as Figure 4CAs shown in the figure. Among them, the remote computing node can be a cloud computing node provided by a distributed network or cloud data center composed of several servers.
[0091] It should be noted that when deploying the sound reflection device 210 and the microphone 220 in this application, it should be ensured that: for a frame of sound signal emitted by the sound source, the microphone 220 picks up this frame of sound signal first, and after a period of time, picks up the reflected sound signal corresponding to this frame of sound signal. For the specific reason for such deployment, please refer to the description of the part of extracting channel features in the following text.
[0092] The technical solution provided by this application can be specifically implemented based on Figure 3 the recognition system 100 with the structure shown in the example or its variant structure. Next, how the speech recognition device 230 in the recognition system 100 performs speech recognition will be introduced in detail.
[0093] As Figure 5 shown, Figure 5 shows a schematic flowchart of a speech recognition method provided by this application. The speech recognition method involved in this application includes but is not limited to the following steps:
[0094] S101: The speech recognition device 230 acquires the reflected speech signal.
[0095] In a specific embodiment, the reflected speech signal is the sound signal of the original speech signal emitted by the sound source passing through the sound reflection device 210. For example, after the original speech signal emitted by the sound source is transmitted to the sound reflection device 210, the sound signal obtained after the sound reflection device 210 reflects it, or for another example, after the original speech signal emitted by the sound source is transmitted to the sound reflection device 210, the sound reflection device 210 reflects it to the sound source, and then the sound signal obtained after the sound source reflects it.
[0096] In a specific embodiment, the speech recognition device 230 acquires the reflected speech signal, including: the speech recognition device 230 acquires the speech signal; filters the speech signal to obtain the original speech signal and the reflected speech signal. Among them, the speech signal is picked up by the microphone 220, and when the sound source makes a sound, the microphone 220 will pick up the original speech signal emitted by the sound source and the reflected speech signal passing through the sound reflection device 210. Therefore, the speech signal picked up by the microphone 220 is a superimposed signal of the original speech signal and the reflected speech signal.
[0097] Optionally, after the voice recognition device 230 obtains the voice signal, it first samples the voice signal according to a preset sampling frequency, and then filters the voice signal. It should be understood that since the frequency of human voice is usually between 0.3 - 4 kHz, according to the Nyquist sampling theorem: when the sampling frequency is greater than twice the highest frequency in the signal, the sampled signal can completely retain the information of the original signal. Therefore, the preset sampling frequency should be higher than 8 kHz. In a possible implementation, in order to ensure that the sampled signal completely retains the information of the original signal, the preset sampling frequency in the embodiments of the present application can be set to a value far exceeding 8 kHz, such as 48 kHz. It should also be understood that since the signal directly picked up by the microphone 220 is an analog signal, the voice recognition device 230 needs to sample the voice signal first before processing it. However, when the microphone 220 has a sampling function (for example, the microphone 220 is built-in with a voice processing chip), the voice recognition device 230 does not need to sample the voice signal anymore.
[0098] It can be understood that since the intensity of the sound will decrease after reflection, and in the present application, when deploying the sound reflection device 210 and the microphone 220, the microphone 220 picks up the original voice signal first and then the reflected voice signal. Therefore, the signal intensity of the reflected voice signal in the present application is weaker than that of the original voice signal. Then, the difference between the signal intensity of the reflected voice signal and the signal intensity of the original voice signal can be used to separate the original voice signal and the reflected voice signal from the voice signal.
[0099] In a specific embodiment, the difference between the signal intensity of the above-mentioned reflected voice signal and the signal intensity of the original voice signal can be reflected by a preset ratio. In a possible implementation, the preset ratio is the ratio of the signal intensity of the reflected voice signal to the signal intensity of the voice signal. In another possible implementation, the preset ratio is the ratio of the signal intensity of the original voice signal to the signal intensity of the voice signal. Optionally, the preset ratio can also be the ratio of the signal intensity of the original voice signal to the signal intensity of the reflected voice signal, etc., which is not specifically limited here. The specific setting method of the preset ratio will be described in detail in combination with the signal separation network 310 in the following text ( Figure 9 ).
[0100] Taking the ratio of the signal intensity of the reflected speech signal to the signal intensity of the speech signal as the preset ratio as an example, the specific process of the speech recognition device 230 filtering the speech signal according to the preset ratio to obtain the original speech signal and the reflected speech signal is as follows: First, calculate the product of the signal intensity of the speech signal and the preset ratio to obtain the signal intensity of the reflected speech signal, thereby obtaining the reflected speech signal. Then, calculate the difference between the signal intensity of the speech signal and the signal intensity of the reflected speech signal to obtain the signal intensity of the original speech signal, thereby obtaining the original speech signal. Similarly, when the preset ratio is the ratio of the signal intensity of the original speech signal to the signal intensity of the speech signal, or the preset ratio is the ratio of the signal intensity of the original speech signal to the signal intensity of the reflected speech signal, the specific process of the speech recognition device 230 separating the original speech signal and the reflected speech signal from the speech signal is similar to the above process and will not be elaborated here for simplicity. It should be noted that the signal intensity of the speech signal, the signal intensity of the original speech signal, and the signal intensity of the reflected speech signal mentioned above can all be the sound pressure level or the sound intensity, etc. in this application, and no specific limitation is made here.
[0101] S102: The speech recognition device 230 determines whether the sound source comes from the target user according to the reflected speech signal.
[0102] In a specific embodiment, the speech recognition device 230 determines whether the sound source comes from the target user according to the reflected speech signal, which specifically includes the following two implementation manners, namely, manner one and manner two, and will be described separately below.
[0103] Manner one, the speech recognition device 230 determines whether the sound source comes from the target user according to the reflected speech signal, including: The speech recognition device 230 obtains the characteristics of the first channel according to the reflected speech signal; determines whether the sound source comes from the target user according to the characteristics of the first channel. Wherein, the first channel refers to the sound channel from the sound source to the microphone 220, and the sound channel refers to the transmission medium of the sound signal, so the first channel can be understood as the medium for transmitting the sound signal between the sound source and the microphone 220.
[0104] In a specific embodiment, the characteristics of the first channel include the temperature characteristics of the first channel. The speech recognition device 230 obtains the characteristics of the first channel according to the reflected speech signal, including: The speech recognition device 230 determines the transmission speed of the reflected speech signal in the first channel; obtains the temperature characteristics of the first channel according to the transmission speed of the reflected speech signal in the first channel.
[0105] In a specific embodiment, since the first channel is the transmission medium between the sound source and the microphone 220, the temperature characteristics of the first channel include the temperature characteristics between the sound source and the microphone 220. When the sound source is a real person, the sound signal emitted by the sound source is emitted through the oral cavity of the real person. Therefore, the first channel can be understood to include the oral cavity of the real person and the air between the real person's mouth and the microphone 220. The temperature characteristics of the first channel are affected by two factors: the temperature characteristics of the oral cavity of the real person and the temperature characteristics of the air between the real person's mouth and the microphone 220. When the sound source is a non-real person (e.g., a tape recorder), the sound signal emitted by the sound source is directly transmitted to the microphone through the air. At this time, the first channel only includes the air between the sound source and the microphone 220, and the temperature characteristics of the first channel are the temperature characteristics of the air between the sound source and the microphone 220.
[0106] In a specific embodiment, the speech recognition device 230 determines the speed of sound transmission in the first channel according to the reflected speech signal, including: determining the time period during which the reflected speech signal is transmitted between the microphone 220 and the sound source; and determining the speed of the reflected speech signal transmitted in the first channel according to this time period. The specific content of this step will be specifically described in steps S1021 - S1022 later.
[0107] It should be noted that the reason why the speech recognition device 230 can determine whether the sound source comes from the target user according to the characteristics of the first channel is as follows: According to the foregoing content, when the sound source is a real person, the temperature of the first channel can be understood as the weighted average of the oral temperature of the real person and the temperature of the air between the oral cavity and the microphone 220. When the sound source is a non-real person, the temperature of the first channel is the current air temperature. Therefore, the temperature characteristics of the first channel corresponding to a real person as the sound source are different from the temperature characteristics of the first channel corresponding to a non-real person as the sound source. That is to say, it is possible to distinguish whether the sound source is a real person according to the temperature characteristics of the first channel. In addition, since the oral temperatures of different people may be different, when different users speak, the corresponding temperature of the first channel will also be different. That is to say, it is also possible to distinguish different users according to the temperature characteristics of the first channel. In summary, it is possible to use the temperature characteristics of the first channel to determine whether the sound source comes from the target user.
[0108] It can be understood that in recent years, with the continuous improvement of people's pursuit of quality of life, air conditioners are installed in both office areas, home environments, and other places, enabling people to be in the most comfortable environment at any time. Generally, the temperature that makes the human body feel comfortable is 18°C to 23°C. Therefore, this application believes that the temperature of the indoor air is 18°C to 23°C. In addition, since the normal oral temperature of the human body is 36.3°C to 37.2°C, the oral temperature of a real person is higher than the air temperature. Then, the temperature of the first channel corresponding to a real person as the sound source is less than the oral temperature of the real person and greater than the current air temperature (i.e., the temperature of the first channel corresponding to a non-real person as the sound source). That is to say, when the temperature of the first channel is greater than the current air temperature, it can be determined that the sound source is a real person. Furthermore, when the temperature of the first channel is consistent with the temperature of the first channel corresponding to the sound source from the target user, it can also be determined that the sound source comes from the target user. It should be noted that in practical applications, in order to improve the accuracy of speech recognition, the voiceprint characteristics of the sound source can also be combined to further determine whether the sound source comes from the target user. The specific content of this step will be specifically described in steps S1023 - S1024 below.
[0109] Method 2: The speech recognition device 230 determines whether the sound source comes from the target user according to the reflected speech signal, including: The speech recognition device 230 obtains the characteristics of the second channel according to the reflected speech signal, where the second channel is the transmission medium between the sound reflection device 210 and the microphone 220; determines whether the sound source comes from the target user according to the characteristics of the second channel.
[0110] In a specific embodiment, the characteristics of the second channel include the temperature characteristics of the second channel. It can be understood that when the distance of the second channel is small, the speech recognition device 230 can also determine whether the sound source is a real person according to the temperature characteristics of the second channel. This is because: when the sound source is a real person, the sound signal emitted by the sound source carries the oral temperature of the real person. If the distance between the sound source and the sound reflection device 210 is relatively close, the original sound signal emitted by the sound source reaches the sound reflection device 210, and the process of being reflected by the sound reflection device 210 to the microphone 220 is still affected by the oral temperature of the real person. Then, similar to the first channel, the temperature of the second channel is also determined by the oral temperature of the real person and the current air temperature. The temperature of the second channel is greater than the current air temperature and less than the oral temperature of the real person. When the sound source is not a real person, the temperature of the second channel is the current air temperature. Therefore, when the distance of the second channel is small, the speech recognition device 230 can also determine whether the sound source is a real person according to the second channel.
[0111] Similar to the temperature characteristic of the first channel, the voice recognition device 230 can obtain the temperature characteristic of the second channel in the following manner: The voice recognition device 230 determines the moment of receiving the third sound signal according to the first characteristic and the second characteristic, where the third sound signal is the sound signal picked up by the microphone 220 when the original voice signal is reflected by the sound reflection device 210 to the microphone 220, that is, the microphone 220 picks up for the first time the sound signal with the first characteristic reflected by the sound reflection device 210. Then, the voice recognition device 230 calculates half of the difference between the moment of receiving the original voice signal and the moment of receiving the third sound signal, thereby obtaining the time for the sound to travel in the second channel. Then, according to the distance of the second channel, the voice recognition device 230 determines the speed of the sound traveling in the second channel, thereby obtaining the temperature characteristic of the second channel.
[0112] Optionally, when the distance of the second channel is small, the voice recognition device 230 can also determine whether the sound source is a real person according to the time for the sound to travel in the second channel or the speed of the sound traveling in the second channel. Specifically, when the sound source is a real person, since the temperature of the second channel is higher than the current air temperature, the speed of the sound traveling in the second channel is greater than the speed of the sound traveling in the air v2, and the time for the sound to travel in the second channel is less than Δt2'. When the sound source is not a real person, since the temperature of the second channel is the current air temperature, the speed of the sound traveling in the second channel is equal to the speed of the sound traveling in the air v2, and the time for the sound to travel in the second channel is equal to Δt2'. Among them, according to the positive correlation relationship between the speed of sound and the temperature of the channel mentioned in step S1022, the speed of the sound traveling in the air v2 = k1 + k2×T2 can be obtained, where T2 represents the current air temperature. For example, when k1 = 331 and k2 = 0.6, if the current air temperature is 0°C (i.e., T2 = 0), then v2 is 331 meters per second (m / s); if the current air temperature is 25°C (i.e., T2 = 25), then v2 is 346 m / s. According to the speed of the sound traveling in the air v2, the corresponding propagation time Δt2' = D2 / v2 can be obtained, where D2 represents the distance of the second channel. Since the distance of the second channel in this application corresponds to the preset ratio mentioned in S101, the distance of the second channel can be determined according to the preset ratio. In addition, the manner in which the voice recognition device 230 obtains the time and speed for the sound to travel in the second channel can refer to the specific process of the voice recognition device 230 obtaining the temperature characteristic of the second channel, which will not be elaborated here.
[0113] The following combines Figure 6 with steps S1021 - S1022 to explain in detail the specific process of the voice recognition device 230 in step S102 for determining the speed of the sound traveling in the first channel according to the reflected voice signal.
[0114] S1021: The voice recognition device 230 determines the time period during which the reflected voice signal is transmitted between the microphone 220 and the sound source.
[0115] In a specific embodiment, the voice recognition device 230 determines the time period during which the reflected voice signal is transmitted between the microphone 220 and the sound source, including: the voice recognition device 230 determines a first moment of receiving a first sound signal and a second moment of receiving a second sound signal, where the reflected voice signal includes the first sound signal and the second sound signal, the first sound signal is a signal that the original voice signal is reflected by the sound reflection device 210 and transmitted to the sound source, and the second sound signal is a signal that the first sound signal is reflected by the sound source and transmitted to the sound reflection device 210; according to the first moment and the second moment, the time period during which the reflected voice signal is transmitted between the microphone 220 and the sound source is determined.
[0116] In a specific embodiment, the voice recognition device 230 determines the first moment of receiving the first sound signal and the second moment of receiving the second sound signal, including: the voice recognition device 230 extracts the features of the original voice signal, where the features of the original voice signal include a first feature, and the first feature is used to indicate the reflected voice signal corresponding to the original voice signal after one or more reflections. Then, the voice recognition device 230 determines the first sound signal and the second sound signal with the first feature according to the features of the original voice signal, so as to determine the first moment and the second moment.
[0117] In a specific embodiment, the features of the original voice signal further include a second feature, and the second feature is used to reflect the signal strength of the original voice signal. Then, the specific process of the voice recognition device 230 determining the first sound signal and the second sound signal with the first feature according to the features of the original voice signal is: first, the voice recognition device 230 obtains a first signal strength and a second signal strength according to the second feature, where the first signal strength is used to determine the signal strength of the first sound signal, the second signal strength is used to determine the signal strength of the second sound signal, the first signal strength can be a value or an interval, and the second signal strength can be a value or an interval. Then, the voice recognition device 230 determines the first sound signal and the second sound signal according to the first feature, the first signal strength, and the second signal strength.
[0118] Optionally, the voice recognition device 230 may first determine a plurality of sound signals with the first feature according to the first feature, where the plurality of sound signals include the first sound signal and the second sound signal. Then, according to the second feature, the first signal strength and the second signal strength are obtained, and then the first sound signal is screened out from the plurality of sound signals according to the first signal strength, and the second sound signal is screened out from the plurality of sound signals according to the second signal strength, so as to further obtain the first moment and the second moment.
[0119] Optionally, the voice recognition device 230 may also first determine multiple voice signals with a first signal strength and multiple voice signals with a second signal strength according to the second feature, and then screen out the first voice signal from the multiple voice signals with the first signal strength and the second voice signal from the multiple voice signals with the second signal strength according to the first feature, so as to further obtain the first moment and the second moment.
[0120] In a specific embodiment, the voice recognition device 230 may determine which voice signals have the first feature in the following manner: the voice recognition device 230 respectively extracts features from the original voice signal and the reflected voice signal, so as to obtain the first feature and features h1, h2,... h m , where, since the reflected voice signal includes voice signals reflected by the sound reflection device 210 once or multiple times, the features h1, h2,... h m are all features of the reflected voice signal, and m is a positive integer. Then, calculate the similarity degrees d1, d2,..., d m between the first feature and the features h1, h2,... h m . If the similarity degree d n is greater than the preset similarity degree, it means that the first feature matches the feature h n , and the reflected voice signal corresponding to the feature h n is a voice signal with the first feature, where 1 ≤ n ≤ m and n is a positive integer.
[0121] It can be understood that since the signal strength of the original voice signal in this application is greater than that of the reflected voice signal, and the signal strength of the reflected voice signal will decrease after being reflected again, and the attenuation of the signal is positively correlated with the distance of the channel. Therefore, according to the characteristics of the signal attenuation and combined with the first feature, the first voice signal and the second voice signal can be determined. For example, the first voice signal is the signal that the original voice signal is reflected by the sound reflection device 210 and transmitted to the sound source, and the second voice signal is the signal that the first voice signal is reflected by the sound source and transmitted to the sound reflection device 210. Then, the signal strength of the original voice signal is greater than the signal strength of the first voice signal, and the signal strength of the first voice signal is greater than the signal strength of the second voice signal. Then, according to the second feature and the signal attenuation feature, the first signal strength and the second signal strength can be determined. Among them, the signal attenuation feature is determined according to the transmission path of the signal, as well as the distance from the sound source to the microphone 220, the distance from the sound reflection device 210 to the microphone 220, and the distance from the sound source to the sound reflection device 210. The first signal strength can specifically be the signal strength of the first voice signal, or it can be an interval that is less than the signal strength of the original voice signal and greater than the signal strength of the second voice signal. Then, the first signal strength can filter out the signal strengths greater than the signal strength of the first voice signal (for example, the signal strength of the original voice signal) and less than the signal strength of the first voice signal (for example, the signal strengths of the second voice signal and some signals of diffuse reflection), so as to obtain one or more voice signals whose signal strengths are close to the signal strength of the first voice signal. Similarly, the second signal strength can also be determined, so as to obtain one or more voice signals whose signal strengths are close to the signal strength of the second voice signal. Combining with the multiple voice signals obtained according to the first feature, the first voice signal and the second voice signal can finally be determined.
[0122] In a specific embodiment, the first feature may be the phoneme feature of the original voice signal, or the voiceprint feature of the original voice signal, or the emotion feature of the original voice signal, etc., which is not specifically limited here. The voice recognition device 230 can extract the phoneme feature of the original voice signal through a neural network (for example, a deep neural networks-hidden markov model (DNN-HMM)), can extract the voiceprint feature of the original voice signal through a neural network (for example, a convolutional neural network (CNN)), and can extract the emotion feature of the original voice signal through a neural network (for example, a recurrent neural network (RNN)). It should be understood that the embodiments of this application do not specifically limit the extraction method of the first feature.
[0123] In a specific embodiment, the second feature may be the sound intensity feature of the original voice signal, or the sound pressure feature of the original voice signal, etc., which is not specifically limited herein. The voice recognition device 230 may determine the sound intensity feature of the original voice signal by detecting the average sound energy passing through a unit area perpendicular to the sound wave propagation direction per unit time of the original voice signal, and may determine the sound pressure feature by detecting the sound pressure value of the original voice signal. It should be understood that the extraction method of the second feature in the embodiments of the present application is not specifically limited.
[0124] In a specific embodiment, the difference between the second moment and the first moment is the time period during which the reflected voice signal is transmitted between the microphone 220 and the sound source, or the difference between the second moment and the first moment includes the time period during which the reflected voice signal is transmitted between the microphone 220 and the sound source and the time period during which the reflected voice signal is transmitted between the sound reflection device 210 and the microphone 220.
[0125] In a specific embodiment, the first moment may be the moment when the microphone 220 picks up the first sound signal, or the moment when the voice recognition device 230 receives the first sound signal. Similarly, the second moment may be the moment when the microphone 220 picks up the second sound signal, or the moment when the voice recognition device 230 receives the second sound signal, which is not specifically limited herein.
[0126] S1022: The voice recognition device 230 determines the transmission speed of the reflected voice signal in the first channel according to the time period during which the reflected voice signal is transmitted between the microphone 220 and the sound source.
[0127] In a specific embodiment, the voice recognition device 230 determines the transmission speed of the reflected voice signal in the first channel according to the distance of the first channel and the time period during which the reflected voice signal is transmitted between the microphone 220 and the sound source. Among them, when the sound source is a real person, the distance of the first channel is the sum of the oral cavity distance of the real person and the distance from the microphone 220 to the real person's mouth. When the sound source is not a real person, the distance of the first channel is the distance between the microphone 220 and the sound generation position of the non-real person. It should be noted that the distance of the first channel in the present application is known, and the distance of the first channel can be obtained by recording the unlocking habit of the target user. The unlocking habit of the target user refers to the distance that the target user is accustomed to maintaining from the microphone 220 when unlocking the PC.
[0128] It should be noted that when the sound source is a real person, the distance of the first channel when the real person's mouth is open is the sum of the oral cavity distance and the distance from the microphone 220 to the real person's mouth, and the distance of the first channel when the real person's mouth is closed is the distance from the microphone 220 to the real person's mouth. Then, there may be a situation where the distance for the first sound signal to be transmitted from the microphone 220 to the sound source is different from the distance for the second sound signal to be transmitted from the sound source to the microphone 220. For example, when the first sound signal is transmitted from the microphone 220 to the sound source and the real person's mouth is open, the distance for the first sound signal to be transmitted is the sum of the real person's oral cavity distance and the distance from the microphone 220 to the real person's mouth. When the sound source reflects the first sound signal, if the real person's mouth is closed, the distance for the second sound signal to be transmitted from the sound source to the microphone 220 is only the distance from the real person's mouth to the microphone 220, which will interfere with determining the transmission speed of the reflected voice signal in the first channel. To avoid this situation, the microphone 220 can pick up the sound signal only when the real person's mouth is open, so that the distance of the corresponding first channel when the sound source is a real person is the sum of the real person's oral cavity distance and the distance from the microphone 220 to the real person's mouth. In a specific implementation, considering that when a person is speaking, the opening and closing of the oral cavity have a fixed frequency, the microphone 220 can determine the moment when the real person's mouth is open according to this fixed frequency, so as to pick up only the sound signal reflected when the real person's mouth is open. It is also possible to pick up all the sound signals first and then screen out the sound signals that are obviously reflected at the moment when the real person's mouth is closed. The present application does not limit the specific implementation manner.
[0129] In addition, when the sound source is a real person, a sound will be emitted when the real person's mouth is open. At this time, since the sound signal emitted by the real person carries the oral cavity temperature of the real person, the temperature of the first channel for transmitting this sound signal is jointly determined by the oral cavity temperature of the real person and the air temperature; when the real person's mouth is closed, no sound is emitted, so that the temperature of the first channel where no sound signal needs to be transmitted is the air temperature. Since the microphone 220 picks up the voice signal by continuously sampling multiple times within a period of time, and the real person's mouth emits a sound signal only when it is open, the microphone 220 may not pick up the sound signal emitted by the real person at multiple moments. Then, at these moments, the temperature of the first channel is the air temperature. To avoid this situation, it is also necessary to make the microphone 220 pick up the sound signal only when the real person's mouth is open, so that the temperature of the corresponding first channel when the sound source is a real person is jointly determined by the oral cavity temperature of the real person and the air temperature.
[0130] To understand steps S1021 - S1022 more clearly, the following will combine Figure 7A to introduce in detail why the speech recognition device 230 can obtain the temperature characteristics of the first channel according to the reflected voice signal.
[0131] Figure 7A shows a schematic diagram of the deployment of a possible sound source, a sound reflection device 210, and a microphone 220. In Figure 7A , the sound source, the sound reflection device 210, and the microphone 220 are on the same horizontal line geographically. Assume that the sound source emits a frame of sound signal (denoted as S1) at time t0. Then, the signal S1 propagates in the first channel, and the microphone 220 receives the signal S1 at time t1. Next, the signal S1 propagates in the second channel. When the signal S1 reaches the sound reflection device 210, the sound reflection device 210 reflects the signal S1. At this time, the reflected signal (denoted as F1) will continue to propagate in the second channel, and the microphone 220 receives the signal F1 at time t2. Next, the signal F1 will propagate in the first channel. When the signal F1 reaches the sound source, the sound source reflects the signal F1 again. At this time, the reflected signal (denoted as F2) will pass through the first channel again and propagate to the microphone 220, and the microphone 220 receives the signal F2 at time t3.
[0132] Since the signal F1 propagates in the first channel, when the distance of the first channel is fixed, the time taken for the signal F1 to propagate in the first channel depends on the propagation speed of the signal F1 in the first channel, and the propagation speed of the signal F1 in the first channel depends on the temperature of the first channel. That is to say, the time taken for the signal F1 to propagate in the first channel depends on the temperature of the first channel. Similarly, the time taken for the signal F2 to propagate in the first channel also depends on the temperature of the first channel. It can be understood that when the sound source is a real person, if the time t0 is small enough (for example, the time t0 is the moment when the sound source starts to emit sound), then when the signal F1 or the signal F2 propagates in the first channel, the sound source is still emitting sound signals, and at this time, the temperature of the first channel is still jointly determined by the oral temperature of the real person and the current air temperature. That is to say, the temperature of the first channel corresponding to the propagation of the signal F1 in the first channel is equal to the temperature of the first channel corresponding to the propagation of the signal F2 in the first channel. When the sound source is not a real person, the temperature of the first channel is the current air temperature. Then, the temperature of the first channel corresponding to the propagation of the signal F1 in the first channel is equal to the temperature of the first channel corresponding to the propagation of the signal F2 in the first channel. In summary, this application believes that the time taken for the signal F1 to propagate in the first channel (from the microphone 220 → the sound source) is equal to the time taken for the signal F2 to propagate in the first channel (from the sound source → the microphone 220). Then, the time Δt1 for the sound to propagate in the first channel can be calculated by the formula Δt1 = (t3 - t2) / 2. When the time for the sound to propagate in the first channel is determined, the propagation speed v1 of the sound in the first channel can be calculated by the formula v1 = D1 / Δt1, where D1 represents the distance of the first channel. A large number of studies have shown that the speed of sound has a positive correlation with the temperature of the channel, which can be specifically expressed by the formula V = k1 + k2×T, where V represents the speed of sound, T represents the temperature of the channel, and k1 and k2 are variable parameters. In practical applications, k1 can take the value of 331 and k2 can take the value of 0.6. Therefore, when the propagation speed v1 of the sound in the first channel is determined, the temperature T1 of the first channel can be calculated by the formula T1 = (v1 - k1) / k2.
[0133] In a specific embodiment, the original voice signal, the first sound signal, the second sound signal, the first moment, and the second moment in step S1021 can correspond to the signal S1, the signal F1, the signal F2, the moment t2, and the moment t3 here in sequence. Then, in this case, the voice recognition device 230 can obtain the temperature T1 of the first channel according to the following formula.
[0134]
[0135] It can be understood that after the signal F2 reaches the microphone 220, it will continue to propagate and then be transmitted from the microphone 220 to the sound reflection device 210 via the second channel again. By analogy with the propagation of the reference signal S1, the signal F1, and the signal F2, the signal will then propagate repeatedly along the path of the sound reflection device 210 → microphone 220 → sound source → microphone 220 → sound reflection device 210 until it attenuates to 0. Therefore, the first sound signal can also be the sound signal that the sound reflection device 210 reflects the signal F2 (the reflected signal is denoted as the signal F3) and then transmits to the sound source. The second sound signal can also be the sound signal that the sound source reflects the signal F3 (the signal after re - reflection is denoted as the signal F4) and then transmits to the sound reflection device 210, and so on. That is to say, in a possible implementation manner of the present application, the first sound signal can include the signal that the original voice signal is reflected by the sound reflection device 210 and transmitted to the sound source, or the signal that the second voice signal is reflected by the sound reflection device 210 for the Nth time and transmitted to the sound source, such as Figure 7A the signal F1 and the signal F3 in Figure 7A ; and the second sound signal can include the signal that the first sound signal is reflected by the sound source for the Nth time and transmitted to the sound reflection device 210, for example
[0136] the signal F2 and the signal F4 in
[0137] Figure 7B Figure 7B shown, where N is a positive integer. Figure 9 In a specific embodiment, the time for the sound to propagate in the second channel can be determined as follows: The speech recognition device 230 determines the third moment when the microphone 220 receives the first sound signal obtained by the sound reflection device 210 reflecting the original voice signal. Then, the speech recognition device 230 calculates 1 / 2 of the difference between the moment when the original voice signal is received and the third moment, and then according to the length of the second channel, the time Δt2 for the sound to propagate in the second channel can be obtained.
[0138] It should be noted that in practical applications, the sound source, the sound reflection device 210, and the microphone 220 may not be on the same horizontal line geographically, as Figure 8And steps S1023 - S1024 are used to explain in detail the specific process in step S102 where the voice recognition device 230 determines whether the sound source comes from the target user according to the temperature characteristics of the first channel.
[0139] S1023: The voice recognition device 230 extracts the voiceprint features of the sound source.
[0140] In a specific embodiment, the voice recognition device 230 extracts voiceprint features from the original voice signal, or the reflected voice signal, or the voice signal. It can be understood that since the voice signal, the original voice signal, and the reflected voice signal all include the voiceprint features of the sound source, the voice recognition device 230 can extract voiceprint features from the voice signal, or from the original voice signal, or from the reflected voice signal. However, compared with the voice signal and the reflected voice signal, the voiceprint features extracted from the original voice signal have higher accuracy and can better represent the characteristics of the sound emitted by the sound source. This is because: the signal intensity of the reflected voice signal is weak, and the reflected voice signal is the original voice signal emitted by the sound source reflected by the sound reflection device 210 to the microphone 220. In this process, the reflected voice signal may include more noise signals. Similarly, since the voice signal is a superimposed signal of the original voice signal and the reflected voice signal, the voice signal also includes more noise signals compared with the original voice signal. That is to say, more accurate voiceprint features can be extracted from the original voice signal.
[0141] In the embodiments of the present application, the voiceprint features include linear predictive coding coefficients (LPC), linear predictive cepstrum coefficients (LPCC), mel-frequency cepstrum coefficients (MFCC), etc. Among them, since MFCC can make full use of the auditory characteristics of the human ear, the present application takes MFCC as an example to illustrate the specific process of the voice recognition device 230 extracting the voiceprint features of the sound source.
[0142] In a possible implementation, the speech recognition device 230 filters the original speech signal using a Mel filter bank to obtain MFCC. Specifically, first, the speech recognition device 230 preprocesses the original speech signal (including pre-emphasis, framing, windowing, etc.), and then performs a Fourier transform on each frame of the preprocessed signal to obtain the spectrum corresponding to each frame of the signal, and uses the Mel filter bank to perform band-pass filtering on each frame of the spectrum, and sums the energies within each filter band to obtain the spectral energy; finally, takes the logarithm of the spectral energy output by each Mel filter in the Mel filter bank, and then performs a discrete cosine transform to obtain MFCC.
[0143] In another possible implementation, the speech recognition device 230 inputs the original speech signal into a voiceprint feature extraction model to obtain voiceprint features. Among them, the voiceprint feature extraction model is obtained by repeatedly training multiple original speech signal samples and multiple voiceprint feature samples corresponding to the multiple original speech signal samples. The specific content of this step will be described in detail in conjunction with the voiceprint feature extraction model 400 in the following text ( Figure 10 ).
[0144] It should be understood that the specific process of the speech recognition device 230 extracting voiceprint features from the reflected speech signal or from the speech signal is similar to the specific process of the speech recognition device 230 extracting voiceprint features from the original speech signal, and for the sake of simplicity, it will not be elaborated in this application.
[0145] S1024: The speech recognition device 230 determines whether the sound source comes from the target user according to the characteristics of the first channel and the voiceprint features of the sound source.
[0146] Among them, the characteristics of the first channel are used to determine whether the sound source is a real person, and the voiceprint features of the sound source are used to determine whether the sound signal emitted by the sound source matches the voiceprint features of the target user.
[0147] Optionally, the speech recognition device 230 can first determine whether the voiceprint features of the sound source match the voiceprint features of the target user. If they do not match, it is determined that the sound source does not come from the target user; if they match, it continues to determine whether the sound source is a real person. If the sound source is a real person, it is determined that the sound source comes from the target user. If the sound source is not a real person, it is determined that the sound source does not come from the target user.
[0148] Optionally, the speech recognition device 230 can also first determine whether the sound source is a real person. If the sound source is not a real person, it is determined that the sound source does not come from the target user; if the sound source is a real person, it continues to determine whether the voiceprint features of the sound source match the voiceprint features of the target user. If they match, it is determined that the sound source comes from the target user. If they do not match, it is determined that the sound source does not come from the target user.
[0149] Optionally, the voice recognition device 230 can also determine whether the sound source is a real person and whether the voiceprint feature of the sound source matches the voiceprint feature of the target user at the same time. When the sound source is a real person and the voiceprint feature of the sound source matches the voiceprint feature of the target user, it is determined that the sound source comes from the target user.
[0150] In a specific embodiment, the voice recognition device 230 determines whether the sound source comes from the target user according to the feature of the first channel and the voiceprint feature of the sound source, including: the voice recognition device 230 compares the temperature feature of the first channel with the standard temperature to obtain a first matching degree; compares the voiceprint feature of the sound source with the standard voiceprint feature to obtain a second matching degree; and determines whether the sound source comes from the target user according to the first matching degree and the second matching degree.
[0151] In a specific implementation, the voice recognition device 230 obtains the first matching degree by calculating the similarity between the temperature feature of the first channel and the standard temperature, and obtains the second matching degree by calculating the similarity between the voiceprint feature of the sound source and the standard voiceprint feature. Optionally, the voice recognition device 230 can obtain the similarity between the temperature feature of the first channel and the standard temperature by calculating the cosine distance, Euclidean distance, Manhattan distance, etc. between the temperature feature of the first channel and the standard temperature. Similarly, the similarity between the voiceprint feature of the sound source and the standard voiceprint feature can also be obtained in the above manner.
[0152] In a specific embodiment, the standard temperature is related to the current air temperature, and the standard temperature refers to the temperature of the first channel when the sound source comes from the target user.
[0153] In a specific embodiment, the standard voiceprint feature is set by the user, and the standard voiceprint feature refers to the voiceprint feature of the target user.
[0154] In a specific embodiment, when the target user wants to use the voiceprint to unlock the PC, the relevant settings of the voiceprint lock need to be made in the PC first. Therefore, the standard temperature and the standard voiceprint feature can be obtained in the following manner: in response to the voiceprint setting request input by the target user, the voice recognition device 230 prompts the target user to input voiceprint setting information, where the voiceprint setting information includes the voiceprint information of the target user, the current air temperature, and the distance of the first channel. Then, the voiceprint setting information input by the target user is received, and the standard temperature and the standard voiceprint feature are obtained according to the voiceprint setting information.
[0155] In a specific implementation, the target user first turns on the PC (the PC includes a voice recognition device 230) by entering a password, and then can input a voiceprint setting request to the PC through devices such as a mouse and a keyboard. When the PC receives the voiceprint setting request, it triggers the authentication of the target user. The target user can authenticate through methods such as password and face recognition. When the authentication is successful, the PC displays the relevant information required for voiceprint setting on the display screen to prompt the target user. Among them, the relevant information required for voiceprint setting includes password information, the current air temperature, and the distance between the target user and the microphone 220 when the target user reads out the password. Then, the target user adjusts the microphone 220 to a comfortable position and inputs the current air temperature and the current distance between the target user and the microphone 220 according to the information displayed on the PC. When the target user needs to input password information, options of "Start recording password" and "End recording password" will pop up on the PC display screen. When the target user clicks the option of "Start recording password", the target user reads out the password. When the target user finishes reading the password, clicks the option of "End recording password". Next, the PC obtains the time for the sound to propagate in the first channel according to the moment when the target user starts to read the password and the moment when the microphone 220 starts to receive the password, and then calculates the propagation speed of the sound in the first channel according to the distance between the target user and the microphone 220 when the target user reads out the password, so as to obtain the standard temperature. The PC obtains the standard voiceprint feature according to the password read out by the target user.
[0156] In a specific embodiment, a first feature library is set in the voice recognition device 230. The first feature library includes at least one set of corresponding temperature features and air temperature of the first channel. At least one set of corresponding temperature features and air temperature of the first channel includes the standard temperature and the current air temperature. Comparing the temperature feature of the first channel with the standard temperature includes: determining the corresponding standard temperature according to the current air temperature; comparing the temperature feature of the first channel with the standard temperature. Optionally, each temperature feature in the first feature library also corresponds to a distance of the first channel. This is because: when the distance of the first channel increases, the oral temperature carried by the original sound signal emitted by the sound source will become lower as the propagation distance increases. Therefore, the temperature feature of the first channel is not only related to the current air temperature but also related to the distance of the first channel.
[0157] In a specific embodiment, a second feature library is also set in the voice recognition device 230. The second feature library includes one or more voiceprint features, and each voiceprint feature corresponds to a target user. One or more voiceprint features include the standard voiceprint feature. As described above, the target user refers to a user who has the permission to unlock and use the above-mentioned PC. Then, when a PC can be used by multiple target users, a PC can be unlocked by multiple target users. Therefore, the second feature library can include the voiceprint features of multiple target users.
[0158] It is understandable that the specific setting methods of the first feature library and the second feature library are similar to the process by which the voice recognition device 230 obtains the standard temperature and the standard voiceprint features. For the sake of simplicity, it will not be elaborated here.
[0159] In a specific embodiment, considering that the air temperature, the unlocking habit of the target user (the distance of the first channel), etc. may change each time unlocking is performed, or the oral temperature of the target user may change within a certain period of time (for example, having a cold), or the voice of the target user may change within a certain period of time (for example, having a sore throat), etc., the voice recognition device 230 can update the first feature library according to the current air temperature and the distance of the first channel regularly or when the target user requests to update the first feature library or the second feature library, and update the second feature library according to the current voiceprint features of the target user.
[0160] In a specific embodiment, the specific process by which the voice recognition device 230 determines whether the voice is that of the target user according to the first matching degree and the second matching degree is as follows: when the first matching degree is greater than the first threshold and the second matching degree is greater than the second threshold, the voice recognition device 230 determines that the sound source is the target user. Among them, the first threshold or the second threshold is set by the user. The first threshold and the second threshold may be the same or different. The first threshold and the second threshold may be fixed values or values adjusted accordingly according to the needs of the target user. The present application does not make specific limitations on this.
[0161] In a specific embodiment, the voice recognition device 230 adjusts the first threshold and the second threshold in the following manner: the voice recognition device 230 obtains the recognition request input by the target user and adjusts the first threshold, or the second threshold, or the first threshold and the second threshold according to the recognition request.
[0162] Optionally, the recognition request includes the target user's requirement for the accuracy of the speech recognition result. For example, the recognition request may specifically include a security recognition request and a convenience recognition request. The requirement for the accuracy of the speech recognition result in the security recognition request is higher than that in the convenience recognition request. Then, the first threshold corresponding to the security recognition request is greater than the first threshold corresponding to the convenience recognition request, and the second threshold corresponding to the security recognition request is greater than the second threshold corresponding to the convenience recognition request. For example, during the process of the target user setting up the PC voiceprint lock, the target user inputs the corresponding recognition request according to their own consideration of the security of the PC. When the target user inputs a security recognition request, the speech recognition device 230 sets the first threshold to 90% and the second threshold to 95%. When the target user inputs a convenience recognition request, the speech recognition device 230 sets the first threshold to 80% and the second threshold to 90%. It can be understood that the higher the first matching degree, the higher the probability that the sound source is a real person. The higher the second matching degree, the more similar the voiceprint feature of the sound source is to the voiceprint feature of the target user, and the higher the probability that the sound source is the target user. Then, the higher the first threshold and the second threshold are set, the higher the first matching degree and the second matching degree required to unlock the PC. That is to say, only when the sound source is the target user can the PC be unlocked. Therefore, compared with the convenience recognition request, the unlocking requirement corresponding to the security recognition request is more stringent and has a higher security factor. That is to say, when the target user pays more attention to the security of unlocking, the target user can choose to input a security recognition request. When the target user pays more attention to the convenience of unlocking, the target user can choose to input a convenience recognition request.
[0163] In addition, optionally, the security recognition request further includes different security levels, such as level 1 security and level 2 security. The second thresholds corresponding to level 1 security and level 2 security are the same, and the third threshold corresponding to level 1 security is less than the third threshold corresponding to level 2 security. Optionally, the convenience recognition request further includes different convenience levels, such as level 1 convenience and level 2 convenience. The second thresholds corresponding to level 1 convenience and level 2 convenience are the same, and the third threshold corresponding to level 1 convenience is less than the third threshold corresponding to level 2 convenience. It should be understood that the recognition request further includes many other forms, which are not limited herein.
[0164] The foregoing content relates to how to extract the temperature feature of the first channel (step S102) and the voiceprint feature of the sound source (step S1023). In the specific implementation process, in order to improve the efficiency and accuracy of voiceprint unlocking, the embodiments of the present application may use an artificial intelligence (AI) model (for example, a neural network model, a support vector machine, etc.) to extract the temperature feature of the first channel and the voiceprint feature of the sound source. Specifically, the embodiments of the present application use a channel feature extraction model to extract the temperature feature of the first channel from the voice signal, and use a voiceprint feature extraction model to extract the voiceprint feature of the sound source from the voice signal.
[0165] In a possible implementation manner, the present application adopts the open-source deep learning framework TensorFlow, and implements the training of the channel feature extraction model and the voiceprint feature extraction model on the TensorFlow platform, and uses the trained channel feature extraction model to extract the temperature feature of the first channel, and uses the trained voiceprint feature extraction model to extract the voiceprint feature. Thus, the channel feature extraction model and the voiceprint feature extraction model tend to be lightweight, so that the hardware resources required when using the channel feature extraction model and the voiceprint feature extraction model are within the available capabilities of the PC, thereby realizing the unlocking task.
[0166] Optionally, the training process of the channel feature extraction model and the voiceprint feature extraction model may be completed in the cloud. The PC regularly downloads the trained channel feature extraction model and the voiceprint feature extraction model from the cloud for use in realizing the unlocking task of the PC. Optionally, the training process of the channel feature extraction model and the voiceprint feature extraction model, and using the trained channel feature extraction model to extract the temperature feature of the first channel, and using the trained voiceprint feature extraction model to extract the voiceprint feature are all completed in the cloud. After the cloud extracts the features of the first channel and the voiceprint feature, it sends them to the PC for use in realizing the unlocking task.
[0167] The following will combine Figure 9 and Figure 10 to introduce the channel feature extraction model and the voiceprint feature extraction model in detail respectively.
[0168] (I) Channel Feature Extraction Model
[0169] As Figure 9 shown, Figure 9 shows a schematic structural diagram of a channel feature extraction model. Figure 9In the example, channel feature extraction model 300 includes a signal separation network 310 and a channel feature inference network 320. Signal separation network 310 is used to separate the original speech signal and the reflected speech signal from the speech signal, and input the original speech signal and the reflected speech signal into channel feature inference network 320. Channel feature inference network 320 is used to infer the temperature characteristics of the first channel based on the original speech signal and the reflected speech signal. The following describes the various components of channel feature extraction model 300.
[0170] (1) Signal separation network 310
[0171] The signal separation network 310 may specifically include a signal strength detection unit 311 and a filtering unit 312. The signal strength detection unit 311 is configured to detect the input voice signal at preset times t1, t2, ..., t m The signal strength of d1, d2, ..., d m The filter unit 312 is used to call the preset ratio in the database (the preset ratio here is the ratio of the signal strength of the reflected voice signal to the signal strength of the voice signal), and respectively convert the signal strengths d1, d2, ..., d m Multiplying by the preset ratio gives the signal strength d 11 ,d 12 ,...,d 1m , thus obtaining the reflected voice signal, and then the signal strength d1, d2, ..., d m With signal strength d 11 ,d 12 ,...,d 1m Subtract and get the signal strength d 21 ,d 22 ,...,d 2m , thus obtaining the original speech signal.
[0172] In a specific embodiment, the database includes a plurality of preset ratios, and each preset ratio corresponds to a distance of a first channel and a distance of a second channel. As can be seen from step S101, the preset ratio can be the ratio of the signal intensity of the reflected voice signal to the signal intensity of the voice signal, or the ratio of the signal intensity of the reflected voice signal to the signal intensity of the original voice signal, etc. Taking the preset ratio as the ratio of the signal intensity of the reflected voice signal to the signal intensity of the voice signal as an example, the preset ratio can be obtained in the following way: First, detect the distance of the first channel and the distance of the second channel. Then, when the sound source (including a real person, a tape recorder, etc.) emits sound, respectively detect the signal intensity of the voice signal picked up by the microphone 220 and the signal intensity of the reflected voice signal. Then, calculate the ratio of the signal intensity of the reflected voice signal sample to the signal intensity of the voice signal sample as the preset ratio. After calculating the preset ratio, associate and store the distance of the first channel, the distance of the second channel, and the preset ratio in the database.
[0173] It should be noted that: Since the signal intensity of the original voice signal is inversely proportional to the distance between the sound source and the microphone 220 (i.e., the distance of the first channel), and the signal intensity of the reflected voice signal is inversely proportional to the distance between the sound reflection device 210 and the microphone 220 (i.e., the distance of the second channel). In practical applications, for the same user, the distance of the first channel depends on the user's unlocking habit. Each time the user unlocks, the user will continuously adjust the position of the microphone 220 and / or their own position according to their habit to make the position of the microphone 220 meet the user's habit requirements. Therefore, the distance of the first channel remains basically unchanged each time the user unlocks, so the signal intensity of the original voice signal is hardly affected by the distance of the first channel. The distance of the second channel depends on the positions of the sound reflection device 210 and the microphone 220. In practical applications, each time the user unlocks, the sound reflection device 210 or the microphone 220 may not be fixed at a certain position. For example, when the microphone 220 is a wireless earphone or a head-mounted microphone, the position of the microphone 220 changes with the position of the wearer. In this case, each time the user unlocks, the distance of the second channel may change randomly, so the signal intensity of the reflected voice signal will be affected by the distance of the second channel. Therefore, a plurality of preset ratios can be stored in the database, and the plurality of preset ratios should include: a plurality of preset ratios obtained by changing the distance of the second channel while the distance of the first channel remains unchanged. Optionally, the plurality of preset ratios can also include: a plurality of preset ratios obtained by changing the distance of the first channel while the distance of the second channel remains unchanged.
[0174] Generally speaking, before and after unlocking, if the distances of the first channel and the second channel remain unchanged, the database may include a preset ratio. However, considering that during actual use, the distance of the first channel or the second channel may change, the database may include multiple preset ratios to adapt to various possible situations.
[0175] (2) Channel Feature Inference Network 320
[0176] The channel feature inference network 320 may specifically include a time inference unit 321, a speed calculation unit 322, and a temperature calculation unit 323. Among them, the time inference unit 321 is used to infer the time for sound to propagate in the first channel based on the original voice signal and the reflected voice signal. For the specific inference process, please refer to the relevant description in step S1021. The speed calculation unit 322 is used to calculate the speed of sound propagation in the first channel based on the distance of the first channel and the time for sound to propagate in the first channel obtained by the time inference unit 321. For details, please refer to the relevant description in step S1022. The temperature calculation unit 323 is used to calculate the temperature of the first channel based on the speed of sound propagation in the first channel.
[0177] Combined with Figure 9 As shown in the channel feature extraction model 300, after inputting the voice signal into the channel feature extraction model 300, the temperature feature of the first channel can be output. However, before using the channel feature extraction model to extract the features of the first channel from the voice signal, the channel feature extraction model needs to be trained first so that the channel feature extraction model has the function of extracting the features of the first channel. The training process of the channel feature extraction model may specifically be:
[0178] First, obtain multiple training sets. Each training set includes a voice signal sample, a temperature feature sample of the first channel, the distance of the first channel, and the current air temperature. In a specific implementation, during the process of the target user setting up the voiceprint lock for the PC, the target user inputs their own voice information into the voice recognition device 230, thereby obtaining multiple training sets. Taking one training set as an example, when the target user sets the voiceprint lock of the PC, the text information of the preset password will be displayed on the display screen of the PC. The target user reads out the preset password according to the text information prompted on the PC display screen. At this time, the sound signal picked up by the microphone 220 is the voice signal sample. During this process, it is also necessary to detect the current air temperature, the distance between the sound source of the target user and the microphone 220 (i.e., the distance of the first channel), and the time when the target user starts to speak (denoted as time A), and the time when the microphone 220 starts to pick up the sound (time B). Then, the time it takes for the sound to travel from the sound source to the microphone 220 (i.e., the time the sound travels in the first channel) is the difference between time A and time B. By calculating the distance of the first channel and the time the sound travels in the first channel, the speed of the sound in the first channel can be obtained, and further the temperature of the first channel can be obtained as the temperature feature sample of the first channel.
[0179] Then, input the multiple training sets into the channel feature extraction model 300 for repeated training until the channel feature extraction model 300 can extract accurate first channel features, thereby obtaining the trained channel feature extraction model 300. In a specific implementation, input the multiple voice signal samples into the channel feature extraction model 300 for training to obtain multiple predicted temperature features of the first channel. Compare the multiple predicted temperature features of the first channel with the multiple temperature feature samples of the first channel, calculate the loss function, and use the calculated loss function to further adjust the parameters in the channel feature extraction model 300. Iteratively execute the above process until the value of the loss function converges and the value of the loss function is less than the preset threshold, then stop the iteration. At this time, the channel feature extraction model 300 is trained and has the function of extracting the temperature features of the first channel from the voice signal.
[0180] After the channel feature extraction model 300 is trained, use the trained channel feature extraction model 300 to extract the temperature features of the first channel from the voice signal. The specific process is as follows: When the target user needs to unlock the PC, the target user will read out the unlock password. At this time, the microphone 220 will pick up the voice signal and send the picked-up voice signal to the voice recognition device 230. After the voice recognition device 230 obtains the voice signal, it inputs the voice signal into the signal intensity detection unit 311. The signal intensity detection unit 311 detects the voice signal at preset times t1, t2,..., t mThe signal strength. Then, the voice signal is at preset times t1, t2,..., t m The signal strength is input into the filtering unit 312. The filtering unit 312 separates the reflected voice signal and the original voice signal from the voice signal according to a preset ratio. Then, the original voice signal and the reflected voice signal are input into the time inference unit 321. The time inference unit 321 infers the time for sound to propagate in the first channel based on the original voice signal and the reflected voice signal. Then, the time for sound to propagate in the first channel is input into the speed calculation unit 322. The speed calculation unit 322 calculates the speed of sound propagation in the first channel according to the distance of the first channel and the time for sound to propagate in the first channel. Then, the speed of sound propagation in the first channel is input into the temperature calculation unit 323, so as to calculate the temperature characteristic of the first channel.
[0181] It should be noted that in the process of training the channel feature extraction model 300 described above, not only multiple speech signal samples and multiple temperature feature samples of the first channel need to be input into the channel feature extraction model 300, but also the current air temperature samples and the distance samples of the first channel in multiple training sets need to be input into the channel feature extraction model 300. The purpose is as follows: As can be seen from step S102, when the sound source is a real person, the temperature of the first channel should be greater than the current air temperature. When the sound source is not a real person, the temperature of the first channel should be equal to the current air temperature. Therefore, it is possible to determine whether the sound source is a real person based on the current air temperature. From the relevant description of the signal separation network 310, the distance of the first channel can be used to determine a preset ratio corresponding to the distance of the first channel in the database, so as to obtain the original speech signal and the reflected speech signal. Also, since the target user has specific unlocking habits, and generally, the air temperature during training is the same as the air temperature during unlocking. Therefore, when the target user unlocks the PC, the distance of the first channel and the preset ratio required by the channel feature extraction model 300 are known. Thus, only by inputting the speech signal into the channel feature extraction model 300 can the temperature feature of the first channel be obtained. In practical applications, there may be a situation where the air temperature during unlocking is different from the air temperature during training. For example, the air conditioner is working normally during training, but the air conditioner fails during unlocking. At this time, it is necessary to retrain the channel feature extraction model 300 according to the current air temperature to update the network parameters (including the time Δt2) in the channel feature extraction model 300, so that the subsequent use of the channel feature extraction model 300 can extract accurate temperature features of the first channel. There may also be a situation where the unlocking habit of the target user changes, resulting in a difference between the distance of the first channel detected during training and the distance of the first channel during unlocking. At this time, it is necessary to retrain the channel feature extraction model 300 according to the new unlocking habit of the target user (i.e., the distance of the first channel during unlocking) to update the network parameters (including the distance of the first channel) in the channel feature extraction model 300, so that the subsequent use of the channel feature extraction model 300 can extract accurate temperature features of the first channel. It can be understood that in order to improve the real-time performance of the channel feature extraction model 300, the channel feature extraction model 300 can also be updated according to the update instruction input by the target user or regularly, so that the channel feature extraction model 300 can extract accurate temperature features of the first channel in subsequent unlockings.
[0182] It should be noted that the channel feature model 300 described in the above embodiments can directly obtain the temperature feature of the first channel based on the voice signal picked up by the microphone 220. It can be understood that the channel feature extraction model can also be various variants based on the channel feature extraction model 300. For example, the channel feature extraction model only includes the channel feature inference network 320. In this case, it is necessary to first separate the reflected voice signal from the voice signal and then input the reflected voice signal into the channel feature extraction model to obtain the temperature feature of the first channel. It can also be understood that when the channel feature extraction model is various variant models based on the channel feature extraction model 300, the training process and the using process of the channel feature extraction model are similar to those of the above channel feature extraction model 300. For the sake of simplicity, no further elaboration will be made.
[0183] (2) Voiceprint feature extraction model
[0184] The voiceprint feature extraction model in the embodiments of the present application can adopt any one of the neural network models in the industry that have better effects in extracting voiceprint features. For example, convolutional neural network (CNN), deep neural networks (DNN), etc. The present application does not make specific limitations.
[0185] Taking CNN as an example below, the process of extracting voiceprint features from the original voice signal in step S1023 will be elaborated.
[0186] As Figure 10 shown, Figure 10 shows a schematic structural diagram of a voiceprint feature extraction model. Figure 10 The voiceprint feature extraction model 400 in
[0187] Before using the voiceprint feature extraction model 400 to extract the voiceprint features of the sound source from the original voice signal, it is necessary to train the voiceprint feature extraction model 400 first, so that the voiceprint feature extraction model 400 has the function of extracting voiceprint features. During training, first obtain a plurality of original voice signal samples and a plurality of voiceprint feature samples. Among them, the plurality of voiceprint feature samples are extracted from the plurality of original voice signal samples. Optionally, the plurality of voiceprint feature samples can be obtained by filtering the plurality of original voice signal samples using a Mel filter bank, and the plurality of voiceprint feature samples can also be extracted by other means. This application does not limit this. Then, input the plurality of original voice signal samples into the voiceprint feature extraction model 400 for training to obtain a plurality of predicted voiceprint features, compare the plurality of predicted voiceprint features with the plurality of voiceprint feature samples, calculate the loss function, and use the calculated loss function to further adjust the parameters in the voiceprint feature extraction model 400. Iteratively execute the above training process until the value of the loss function converges and the value of the loss function is less than the preset threshold, then stop the iteration. At this time, the voiceprint feature extraction model 400 is trained and has the function of extracting voiceprint features from the original voice signal.
[0188] After the voiceprint feature extraction model 400 is trained, use the trained voiceprint feature extraction model 400 to extract the voiceprint features from the original voice signal. The specific process is as follows: First, input the original voice signal output by the signal separation network 310 into the preprocessing unit 420, and obtain the spectrogram corresponding to the original voice signal after being processed by the preprocessing unit 420. Then, input the spectrogram corresponding to the original voice signal into the convolutional layer 410, and generate a voiceprint feature image after convolution processing in one or more convolutional layers 410, so as to obtain the voiceprint features of the sound source.
[0189] It should be understood that Figure 8 The shown voiceprint feature extraction model is only an example. In practical applications, the voiceprint feature extraction model can also be in other forms. For example, the voiceprint feature extraction model can also include one or more pooling layers, one or more fully connected layers, etc., which are not specifically limited here.
[0190] In practical applications, since the voiceprint features of the target user may change. For example, the target user has a sore throat. In this case, the voiceprint features extracted by using the previous voiceprint feature extraction model 400 may not be accurate. Therefore, in order to improve the real-time performance of the voiceprint feature extraction model 400, the network parameters in the voiceprint feature extraction model 400 can be updated according to the update instruction input by the target user or regularly according to the training process of the above voiceprint feature extraction model 400, so that the voiceprint feature extraction model 400 can extract accurate voiceprint features in subsequent unlocking.
[0191] The foregoing content elaborates in detail on the speech recognition method provided by the embodiments of the present application. To better implement the method provided by the embodiments of the present application, the related devices provided by the embodiments of the present application for cooperating to implement the above method will be introduced next.
[0192] As Figure 11 shown, Figure 11 A schematic structural diagram of a speech recognition device provided by an embodiment of the present application is shown. The speech recognition device 230 includes an acquisition unit 510 and an identification unit 520.
[0193] The acquisition unit 510 is used to acquire a reflected speech signal, where the reflected speech signal is a sound signal of the original speech signal emitted by a sound source passing through the sound reflection device 210;
[0194] The identification unit 520 is used to determine whether the sound source comes from a target user according to the reflected speech signal.
[0195] In a specific embodiment, the identification unit 520 is specifically used to: obtain the characteristics of the first channel according to the reflected speech signal, where the first channel is the transmission medium between the sound source and the microphone 220; determine whether the sound source comes from a target user according to the characteristics of the first channel.
[0196] In a specific embodiment, the characteristics of the first channel include the temperature characteristics of the first channel. The identification unit 520 is specifically used to: determine the transmission speed of the reflected speech signal in the first channel; obtain the temperature characteristics of the first channel according to the transmission speed of the reflected speech signal in the first channel.
[0197] In a specific embodiment, the identification unit 520 is specifically used to: determine the time period during which the reflected speech signal is transmitted between the microphone and the sound source; determine the transmission speed of the reflected speech signal in the first channel according to the time period.
[0198] In a specific embodiment, the acquisition unit 510 is specifically used to: acquire a speech signal; filter the speech signal to obtain the original speech signal and the reflected speech signal.
[0199] In a specific embodiment, the speech recognition device 230 further includes a voiceprint feature extraction unit 530. The voiceprint feature extraction unit 530 is used to extract voiceprint features from the original speech signal and / or the reflected speech signal; the identification unit 520 is specifically used to: determine whether the sound source comes from a target user according to the characteristics of the first channel and the voiceprint features.
[0200] In a specific embodiment, the recognition unit 520 is configured to: compare the temperature feature of the first channel with a standard temperature to obtain a first matching degree; compare the voiceprint feature with a standard voiceprint feature to obtain a second matching degree; and determine whether the sound source comes from the target user according to the first matching degree and the second matching degree.
[0201] In a specific embodiment, the standard temperature is related to the current air temperature.
[0202] In a specific embodiment, the standard voiceprint feature is set by the user.
[0203] In a specific embodiment, the recognition unit 520 is specifically configured to: when the first matching degree is greater than a first threshold and the second matching degree is greater than a second threshold, determine that the sound source comes from the target user, where the first threshold or the second threshold is set by the user.
[0204] In a specific embodiment, the feature of the first channel is obtained according to the reflected voice signal and a channel feature extraction model, where the channel feature extraction model is obtained by training a neural network with multiple training sets.
[0205] In a specific embodiment, the recognition unit 520 is further configured to: obtain the feature of a second channel according to the reflected voice signal, where the second channel is the transmission medium between the microphone and the sound reflection device 210; and determine whether the sound source comes from the target user according to the feature of the second channel.
[0206] In a specific embodiment, the material of the sound reflection device 210 includes aluminum or aluminum alloy or glass.
[0207] The voice recognition device 230 in the embodiments of the present application is only illustrated by dividing the above functional modules. In actual applications, the above functions may be allocated to different functional modules according to needs, that is, the internal structure of the voice recognition device 230 is divided into different functional modules to complete all or part of the functions described above. In addition, the voice recognition device 230 provided in the above embodiments and the Figure 3 、 Figure 4A 、 Figure 4B and Figure 4C shown voice recognition device 230 belong to the same concept. The specific implementation process is detailed in the above method embodiments. Among them, the acquisition unit 510 corresponds to execute Figure 5 S101 in Figure 5 , the recognition unit 520 corresponds to execute Figure 6 S102 in Figure 8 , S1021 and S1022 in Figure 8S1023 in it will not be elaborated here for simplicity.
[0208] As Figure 12 shown, Figure 12 FIG. shows a schematic structural diagram of an electronic device provided by the present application. The electronic device 600 includes at least one processor 610, a communication bus 620, a memory 630, and at least one communication interface 640.
[0209] The processor 610 may be a general-purpose central processing unit (CPU), a microprocessor, or may be one or more integrated circuits for implementing the solution of the present application. For example, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0210] The communication bus 620 is used to transmit information between the above components. The communication bus 620 may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0211] The memory 630 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, or it can be a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 630 can exist independently and be connected to the processor 610 through the communication bus 620. The memory 630 can also be integrated with the processor 610.
[0212] The communication interface 640 uses any device such as a transceiver for communicating with other devices or communication networks. The communication interface 640 includes a wired communication interface and can also include a wireless communication interface. Among them, the wired communication interface can be, for example, an Ethernet interface. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface can be a wireless local area networks (WLAN) interface, a cellular network communication interface, or a combination thereof, etc.
[0213] In a specific implementation, as an embodiment, the processor 610 can include a CPU and a DSP, such as Figure 12 CPU0 and DSP1 shown in
[0214] In a specific implementation, as an embodiment, the electronic device 600 can include multiple processors, such as Figure 12 the processor 610 and the processor 650 shown in
[0215] In some embodiments, the memory 630 is used to store the program code 660 for executing the solution of this application, and the processor 610 can execute the program code 660 stored in the memory 630. That is, the electronic device 600 can implement the method provided in the above method embodiments through the processor 610 and the program code 660 in the memory 630. The software components loaded on the memory 630 can be generally summarized functionally or logically. For example, Figure 11 the acquisition unit 510, the recognition unit 520, and the voiceprint feature extraction unit 530 shown.
[0216] The electronic device 600 in the embodiments of this application includes the above voice recognition device 230, that is, the electronic device 600 can be either the above voice recognition device 230 or an intelligent terminal such as a PC or a smart speaker including the voice recognition device. Moreover, the processor 610, the communication interface 640, etc. in the electronic device 600 can implement the functions and / or various steps and methods implemented by the voice recognition device 230 in the above method embodiments. For the sake of brevity, details are not described herein again.
[0217] This application also provides another computer-readable storage medium. The computer-readable storage medium stores computer instructions. When the computer instructions run on a computing device (for example, Figure 11 the voice recognition device 230 shown or Figure 12 the electronic device 600 shown), the computing device is caused to execute the method executed by the voice recognition device 230 in the above method embodiments.
[0218] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The above computer program product includes one or more computer instructions. When the above computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The above computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The above computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium can be a magnetic medium (such as a floppy disk, storage disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD), etc. In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0219] In several embodiments provided by the present application, it should be understood that the disclosed device can also be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the above units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the indirect coupling, direct coupling, or communication connection shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0220] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present application.
[0221] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0222] When the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium may include, for example, various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.
[0223] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A speech recognition method, characterized in that, The method includes: Obtaining a reflected speech signal, where the reflected speech signal is a sound signal of an original speech signal emitted by a sound source after passing through a sound reflection device; Obtaining characteristics of a sound channel according to the reflected speech signal, where the sound channel is a transmission medium for transmitting the original speech signal and the reflected speech signal; Judging whether the sound source comes from a target user according to the characteristics of the sound channel and the voiceprint characteristics of the sound source.
2. The method according to claim 1, wherein The obtaining of the characteristics of the sound channel according to the reflected speech signal includes: Obtaining characteristics of a first channel according to the reflected speech signal, where the first channel is a transmission medium between the sound source and a microphone.
3. The method according to claim 2, wherein The characteristics of the first channel include the temperature characteristics of the first channel. The obtaining of the characteristics of the first channel according to the reflected speech signal includes: Determining the transmission speed of the reflected speech signal in the first channel; Obtaining the temperature characteristics of the first channel according to the transmission speed of the reflected speech signal in the first channel.
4. The method according to claim 3, wherein The determining of the transmission speed of the reflected speech signal in the first channel includes: Determining the time period during which the reflected speech signal is transmitted between the microphone and the sound source; Determining the transmission speed of the reflected speech signal in the first channel according to the time period.
5. The method according to claim 4, characterized in that The obtaining of the reflected speech signal includes: Obtaining a speech signal; Filtering the speech signal to obtain the original speech signal and the reflected speech signal.
6. The method according to any one of claims 2-5, wherein The method further includes: Extracting the voiceprint characteristics from the original speech signal and / or the reflected speech signal; The judging of whether it comes from the target user according to the characteristics of the sound channel and the voiceprint characteristics of the sound source includes: Judging whether the sound source comes from the target user according to the characteristics of the first channel and the voiceprint characteristics.
7. The method according to claim 6, wherein Judging whether the sound source comes from the target user according to the characteristics of the first channel and the voiceprint characteristics includes: Comparing the temperature characteristics of the first channel with a standard temperature to obtain a first matching degree; Comparing the voiceprint characteristics with standard voiceprint characteristics to obtain a second matching degree; Judging whether the sound source comes from the target user according to the first matching degree and the second matching degree.
8. The method according to claim 7, wherein The standard temperature is related to the current air temperature.
9. The method according to claim 7, wherein The standard voiceprint characteristics are set by a user.
10. The method according to claim 7, wherein The judging of whether the sound source comes from the target user according to the first matching degree and the second matching degree includes: When the first matching degree is greater than a first threshold and the second matching degree is greater than a second threshold, determining that the sound source comes from the target user, where the first threshold or the second threshold is set by the user.
11. The method according to claim 2, characterized in that The characteristics of the first channel are obtained according to the reflected speech signal and a channel feature extraction model, where the channel feature extraction model is obtained by training a neural network with a plurality of training sets.
12. The method according to claim 1, wherein obtaining the characteristics of the sound channel based on the reflected voice signal includes: obtaining the characteristics of a second channel based on the reflected voice signal, where the second channel is the transmission medium between the microphone and the sound reflection device.
13. The method according to claim 1 or 12, characterized in that, The material of the sound reflection device includes aluminum, aluminum alloy or glass.
14. A voice recognition device, characterized in that, The voice recognition device includes: an acquisition unit configured to acquire a reflected voice signal, where the reflected voice signal is the voice signal of the original voice signal emitted by a sound source after passing through the sound reflection device; an identification unit configured to: obtain the characteristics of the sound channel based on the reflected voice signal, where the sound channel is the transmission medium for transmitting the original voice signal and the reflected voice signal; and determine whether the sound source comes from a target user based on the characteristics of the sound channel and the voiceprint characteristics of the sound source.
15. The voice recognition device according to claim 14, wherein The identification unit is configured to: obtain the characteristics of a first channel based on the reflected voice signal, where the first channel is the transmission medium between the sound source and the microphone.
16. The voice recognition device according to claim 15, characterized in that, The characteristics of the first channel include the temperature characteristics of the first channel, and the identification unit is configured to: determine the transmission speed of the reflected voice signal in the first channel; obtain the temperature characteristics of the first channel based on the transmission speed of the reflected voice signal in the first channel.
17. The voice recognition device according to claim 16, characterized in that, The identification unit is configured to: determine the time period during which the reflected voice signal is transmitted between the microphone and the sound source; determine the transmission speed of the reflected voice signal in the first channel based on the time period.
18. The voice recognition device according to claim 17, characterized in that, The acquisition unit is configured to: acquire a voice signal; filter the voice signal to obtain the original voice signal and the reflected voice signal.
19. The voice recognition device according to any one of claims 15-18, characterized in that, The voice recognition device further includes a voiceprint feature extraction unit, where the voiceprint feature extraction unit is configured to extract the voiceprint characteristics from the original voice signal and / or the reflected voice signal; and the identification unit is configured to determine whether the sound source comes from the target user based on the characteristics of the first channel and the voiceprint characteristics.
20. The voice recognition device according to claim 19, characterized in that, The identification unit is configured to: compare the temperature characteristics of the first channel with a standard temperature to obtain a first matching degree; compare the voiceprint characteristics with standard voiceprint characteristics to obtain a second matching degree; and determine whether the sound source comes from the target user based on the first matching degree and the second matching degree.
21. The voice recognition device according to claim 20, characterized in that, The standard temperature is related to the current air temperature.
22. The voice recognition device according to claim 20, characterized in that, The standard voiceprint characteristics are set by the user.
23. The voice recognition device according to claim 20, characterized in that, The identification unit is configured to: when the first matching degree is greater than a first threshold and the second matching degree is greater than a second threshold, determine that the sound source comes from the target user, where the first threshold or the second threshold is set by the user.
24. The voice recognition device according to claim 15, wherein The characteristics of the first channel are obtained based on the reflected voice signal and a channel feature extraction model, where the channel feature extraction model is obtained by training a neural network with multiple training sets.
25. The voice recognition device according to claim 14, characterized in that, The identification unit is configured to: Based on the reflected voice signal, characteristics of a second channel are obtained, where the second channel is a transmission medium between a microphone and the sound reflection device.
26. The voice recognition device according to claim 14 or 25, characterized in that, The material of the sound reflection device includes aluminum or aluminum alloy or glass.
27. An electronic device, characterized in that, The electronic device includes a processor and a memory, and the processor executes the code in the memory to implement the method according to any one of claims 1-13.
28. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for implementing the method according to any one of claims 1-13.
29. An identification system, characterized in that, The recognition system includes a microphone, a sound reflection device, and a voice recognition device, and the voice recognition device is used to execute the method according to any one of claims 1-13.
Citation Information
Patent Citations
Voice waking-up method, device thereof, terminal and storage medium
CN108735218A