Speech recognition method and device based on online speech model, and storage medium
By upgrading the online voice model when the user is not in the environment, the problem of reduced recognition accuracy of the online voice model due to environmental changes is solved, and more accurate voice command recognition is achieved.
Patent Information
- Application Number
- CN202311679268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-07
- Publication Date
- 2025-06-10
AI Technical Summary
During use, the online speech model's recognition accuracy decreases due to changes in the speech recognition environment, and the model parameters cannot be updated in time to recognize voice commands.
By obtaining environmental information and sound information of the speech recognition environment, it is determined whether there is a user, and when the user is not present, it is determined whether the sound information meets the preset conditions. If not, the online speech model is upgraded to obtain the upgraded model to improve recognition accuracy.
When the user is not in the environment, the upgraded online voice model can identify sound information that does not meet the conditions, thereby more accurately identifying voice commands and improving the accuracy of voice recognition.
Smart Images

Figure CN120126465A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular, to a speech recognition method, device, and storage medium based on an online speech model. Background Art
[0002] Currently, the development of speech AI technology is advancing by leaps and bounds. For example, in speech AI technologies used in smart home appliances and smart homes, people only need to speak relevant commands, and after the speech device recognizes them, it can perform corresponding operations according to the speech commands. This provides an extremely simple interaction method between users and home appliances and smart homes, unleashes the potential of smart home appliances and smart homes, optimizes the controllability of smart home appliances and smart homes, and improves the comfort of users' life in the home environment.
[0003] An online speech device refers to a speech device that communicates with a host computer or at least one other device. When applying speech AI technology to various online speech devices, the following problems are faced: Online speech models are usually set according to general scenarios at the time of factory production. During the use of online speech devices, the corresponding online speech models only need to work according to the factory settings. However, in actual applications, the speech recognition environments of online speech devices vary greatly, and the speech recognition environment of the same online speech device also changes over time. The change in the speech recognition environment means that the recognition accuracy of the online speech model trained based on prior data decreases. Specifically, it is manifested that the model parameters of the online speech model cannot be updated in a timely manner with the change of the speech recognition environment, resulting in the online speech model being unable to more accurately recognize speech commands from environmental sounds.
[0004] Therefore, proposing a new speech recognition method based on an online speech model to improve the accuracy of speech recognition is an urgent problem to be solved during the use of online speech devices. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a speech recognition method, device, and storage medium based on an online speech model, in which the model parameters of the online speech model can be updated in a timely manner with the change of the speech recognition environment, so as to more accurately recognize speech commands from environmental sounds, thereby improving the accuracy of speech recognition.
[0006] To solve the above technical problem, in the first aspect of the present invention, a speech recognition method based on an online speech model is disclosed, and the method includes:
[0007] Obtain the environmental information of the speech recognition environment and the sound information in the speech recognition environment, and determine whether there is a user in the speech recognition environment according to the environmental information and the sound information;
[0008] When it is determined that the user does not exist in the voice recognition environment, determine whether the voice information meets a preset first voice recognition condition;
[0009] When it is determined that the voice information does not meet the preset first voice recognition condition, perform an upgrade operation on the online voice model according to the voice information to obtain the upgraded online voice model;
[0010] Perform a voice recognition operation on the voice recognition environment through the upgraded online voice model to obtain a voice recognition result.
[0011] As an alternative implementation manner, in the first aspect of the present invention, the method further includes:
[0012] When it is determined that the user exists in the voice recognition environment, determine whether the voice information meets a preset second voice recognition condition. When it is determined that the voice information meets the preset second voice recognition condition, analyze the voice information to obtain a voice analysis result, determine a voice control instruction matching the voice analysis result, and control at least one target device in the voice recognition environment to perform an operation matching the voice control instruction, where the voice analysis result includes voice instruction information issued by the user;
[0013] When it is determined that the voice information does not meet the preset second voice recognition condition, determine whether the voice information exists in the voice database corresponding to the online voice model. When it is determined that the voice information does not exist in the voice database, store the voice information in the voice database.
[0014] As an alternative implementation manner, in the first aspect of the present invention, before performing the upgrade operation on the online voice model according to the voice information to obtain the upgraded online voice model, the method further includes:
[0015] According to all the voice information, determine whether the online voice model meets a preset model self-check condition;
[0016] When it is determined that the online voice model meets the preset model self-check condition, determine whether the online voice model needs to be updated according to all the voice information stored in the online voice model;
[0017] When it is determined that the online voice model needs to be updated, trigger the operation of performing the upgrade operation on the online voice model according to the voice information to obtain the upgraded online voice model.
[0018] As an alternative implementation, in the first aspect of the present invention, determining whether the online speech model meets the preset model self-checking conditions based on all the voice information includes:
[0019] Based on all the voice information, determining the information quantity of all the voice information stored in the online speech model and the update time when the online speech model was last updated, and based on the update time when the online speech model was last updated, determining the update interval duration of the online speech model;
[0020] Judging whether the information quantity is greater than or equal to a preset storage quantity threshold and whether the update interval duration of the online speech model is greater than or equal to a preset self-checking duration threshold;
[0021] When it is judged that the information quantity is greater than or equal to the preset storage quantity threshold and / or the update interval duration of the online speech model is greater than or equal to the preset self-checking duration threshold, determining that the online speech model meets the preset model self-checking conditions;
[0022] When it is judged that the information quantity is less than the preset storage quantity threshold and the update interval duration of the online speech model is less than the preset self-checking duration threshold, determining that the online speech model does not meet the preset model self-checking conditions.
[0023] As an alternative implementation, in the first aspect of the present invention, determining whether the online speech model needs to be updated based on all the voice information stored in the online speech model includes:
[0024] Generating a test speech set according to all the voice information and a plurality of predetermined voice test instructions;
[0025] Performing a test operation on the online speech model according to the test speech set to obtain a model test result, where the model test result includes the speech test result corresponding to each voice test instruction;
[0026] According to the model test result, determining the target quantity of target test results that meet the preset voice test conditions, and based on the target quantity, judging whether the target quantity is greater than or equal to a preset test quantity threshold;
[0027] When it is judged that the target quantity is greater than or equal to the preset test quantity threshold, determining that the online speech model does not need to be updated;
[0028] When it is judged that the target quantity is less than the preset test quantity threshold, determining that the online speech model needs to be updated.
[0029] As an alternative implementation, in the first aspect of the present invention, the step of determining whether there is a user in the speech recognition environment according to the environmental information and the voice information includes:
[0030] Extract voice feature information from the voice information according to the voice information, and generate a voice feature set based on all the voice feature information;
[0031] Judge whether there is target user information related to the user in the environmental information and whether there is a pre-determined background voice feature in the voice information feature set;
[0032] When it is judged that the target user information exists in the environmental information and / or the background voice feature exists in the voice information feature set, it is determined that there is a user in the speech recognition environment;
[0033] When it is judged that the target user information does not exist in the environmental information and the background voice feature does not exist in the voice information feature set, it is determined that there is no user in the speech recognition environment;
[0034] Wherein, the pre-determined background voice feature is determined by the following method:
[0035] Obtain environmental voice information within a preset voice acquisition duration corresponding to the speech recognition environment, determine the environmental voice information as background voice information, and determine the background voice feature according to the background voice information;
[0036] Wherein, the preset voice acquisition duration includes a first time period before obtaining the user voice command and / or a second time period after obtaining the user voice command.
[0037] As an alternative implementation, in the first aspect of the present invention, before the step of performing an upgrade operation on the online voice model according to the voice information to obtain the upgraded online voice model, the method further includes:
[0038] For each environment to be recognized, respectively set a model parameter set corresponding to the environment to be recognized, and the model parameter set includes at least one parameter in the online voice model trained to convergence;
[0039] For each environment to be recognized, generate an online voice model corresponding to the environment to be recognized according to the model parameter set corresponding to the environment to be recognized;
[0040] And, the step of performing an upgrade operation on the online voice model according to the voice information to obtain the upgraded online voice model includes:
[0041] Determine a target online speech model corresponding to the speech recognition environment from all the online speech models corresponding to the to-be-recognized environments according to the voice information;
[0042] Perform an upgrade operation on the target online speech model according to the voice information to obtain the upgraded target online speech model.
[0043] A second aspect of the present invention discloses a speech recognition device based on an online speech model, and the device includes:
[0044] An acquisition module, configured to acquire the environmental information of the speech recognition environment and the voice information in the speech recognition environment;
[0045] A first judgment module, configured to judge whether there is a user in the speech recognition environment according to the environmental information and the voice information;
[0046] A second judgment module, configured to judge whether the voice information meets a preset first speech recognition condition when the first judgment module judges that there is no such user in the speech recognition environment;
[0047] An upgrade module, configured to perform an upgrade operation on the online speech model according to the voice information to obtain the upgraded online speech model when the second judgment module judges that the voice information does not meet the preset first speech recognition condition;
[0048] A speech recognition module, configured to perform a speech recognition operation on the speech recognition environment through the upgraded online speech model to obtain a speech recognition result.
[0049] As an optional implementation manner, in the second aspect of the present invention, the second judgment module is further configured to judge whether the voice information meets a preset second speech recognition condition when the first judgment module judges that there is such user in the speech recognition environment;
[0050] The device further includes:
[0051] A control module, configured to analyze the voice information to obtain a voice analysis result, determine a voice control instruction matching the voice analysis result, and control at least one target device in the speech recognition environment to perform an operation matching the voice control instruction when the second judgment module judges that the voice information meets the preset second speech recognition condition, where the voice analysis result includes the voice instruction information issued by the user;
[0052] A third judgment module, configured to, when the second judgment module determines that the voice information does not meet the preset second voice recognition condition, determine whether the voice information exists in the voice database corresponding to the online voice model, and when it is determined that the voice information does not exist in the voice database, store the voice information in the voice database.
[0053] As an optional implementation manner, in the second aspect of the present invention, the device further includes:
[0054] A model self-check module, configured to, before the upgrade module performs an upgrade operation on the online voice model according to the voice information to obtain the upgraded online voice model, determine whether the online voice model meets a preset model self-check condition according to all the voice information;
[0055] The second judgment module is further configured to, when the model self-check module determines that the online voice model meets the preset model self-check condition, determine whether the online voice model needs to be updated according to all the voice information stored in the online voice model;
[0056] The model self-check module is further configured to, when it is determined that the online voice model needs to be updated, trigger the upgrade module to perform the operation of performing an upgrade operation on the online voice model according to the voice information to obtain the upgraded online voice model.
[0057] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the model self-check module determines whether the online voice model meets a preset model self-check condition according to all the voice information includes:
[0058] According to all the voice information, determine the information quantity of all the voice information stored in the online voice model and the update time when the online voice model was last updated, and based on the update time when the online voice model was last updated, determine the update interval duration of the online voice model;
[0059] Judge whether the information quantity is greater than or equal to a preset storage quantity threshold and whether the update interval duration of the online voice model is greater than or equal to a preset self-check duration threshold;
[0060] When it is determined that the information quantity is greater than or equal to the preset storage quantity threshold and / or the update interval duration of the online voice model is greater than or equal to the preset self-check duration threshold, determine that the online voice model meets the preset model self-check condition;
[0061] When it is determined that the number of the information is less than the preset storage quantity threshold and the update interval duration of the online voice model is less than the preset self-check duration threshold, it is determined that the online voice model does not meet the preset model self-check condition.
[0062] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the second determination module determines whether the online voice model needs to be updated according to all the voice information stored in the online voice model includes:
[0063] According to all the voice
[0064] information and a plurality of pre-determined voice test instructions, a test voice set is generated;
[0065] A test operation is performed on the online voice model according to the test voice set to obtain a model test result, where the model test result includes a voice test result corresponding to each voice test instruction;
[0066] According to the model test result, the target quantity of the target test results that meet the preset voice test conditions is determined, and based on the target quantity, it is determined whether the target quantity is greater than or equal to a preset test quantity threshold;
[0067] When it is determined that the target quantity is greater than or equal to the preset test quantity threshold, it is determined that the online voice model does not need to be updated;
[0068] When it is determined that the target quantity is less than the preset test quantity threshold, it is determined that the online voice model needs to be updated.
[0069] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the first determination module determines whether there is a user in the voice recognition environment according to the environment information and the voice information includes:
[0070] According to the voice information, voice feature information in the voice information is extracted, and a voice feature set is generated based on all the voice feature information;
[0071] It is determined whether there is target user information related to the user in the environment information and whether there is a pre-determined background voice feature in the voice information feature set;
[0072] When it is determined that the target user information exists in the environment information and / or the background voice feature exists in the voice information feature set, it is determined that there is a user in the voice recognition environment;
[0073] When it is determined that the target user information does not exist in the environmental information and the background sound feature does not exist in the sound information feature set, it is determined that there is no user in the voice recognition environment;
[0074] Among them, the pre-determined background sound feature is determined by the following method:
[0075] Obtain the environmental sound information within the preset voice acquisition duration corresponding to the voice recognition environment, determine the environmental sound information as the background sound information, and determine the background sound feature according to the background sound information;
[0076] Among them, the preset voice acquisition duration includes a first time period before obtaining the user voice command and / or a second time period after obtaining the user voice command.
[0077] As an optional implementation manner, in the second aspect of the present invention, the device further includes:
[0078] A parameter configuration module, configured to, before the upgrade module performs an upgrade operation on the online voice model according to the sound information to obtain the upgraded online voice model, for each environment to be recognized, respectively set a model parameter set corresponding to the environment to be recognized, where the model parameter set includes at least one parameter in the online voice model trained to convergence; for each environment to be recognized, generate an online voice model corresponding to the environment to be recognized according to the model parameter set corresponding to the environment to be recognized;
[0079] And, the specific manner in which the upgrade module performs an upgrade operation on the online voice model according to the sound information to obtain the upgraded online voice model includes:
[0080] According to the sound information, determine a target online voice model corresponding to the voice recognition environment from all the online voice models corresponding to the environments to be recognized;
[0081] Perform an upgrade operation on the target online voice model according to the sound information to obtain the upgraded target online voice model.
[0082] The third aspect of the present invention discloses another voice recognition device based on an online voice model, and the device includes:
[0083] A memory storing executable program code;
[0084] A processor coupled to the memory;
[0085] The processor calls the executable program code stored in the memory and executes the voice recognition method based on the online voice model disclosed in the first aspect of the present invention.
[0086] The fourth aspect of the present invention discloses a computer-readable storage medium. The computer-readable storage medium stores computer instructions, which when called, are used to execute the speech recognition method based on an online speech model disclosed in the first aspect of the present invention.
[0087] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0088] In the embodiments of the present invention, environmental information of the speech recognition environment and sound information in the speech recognition environment are obtained. According to the environmental information and the sound information, it is determined whether there is a user in the speech recognition environment; when it is determined that there is no user in the speech recognition environment, it is determined whether the sound information meets a preset first speech recognition condition; when it is determined that the sound information does not meet the preset first speech recognition condition, an upgrade operation is performed on the online speech model according to the sound information to obtain an upgraded online speech model; a speech recognition operation is performed on the speech recognition environment through the upgraded online speech model to obtain a speech recognition result. It can be seen that implementing the present invention can determine whether it is necessary to upgrade the online speech model according to the first speech recognition condition based on the sound information in the environment, so as to upgrade the online speech model when the user is not in the recognition environment. The upgraded online speech model can recognize the sound information that does not meet the first speech recognition condition, and thus more accurately recognize the speech command from the environmental sound, improving the accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0090] Figure 1 is a schematic diagram of a speech recognition scenario based on an online speech model disclosed in the embodiments of the present invention;
[0091] Figure 2 is a schematic flowchart of a speech recognition method based on an online speech model disclosed in the embodiments of the present invention;
[0092] Figure 3 is a schematic flowchart of another speech recognition method based on an online speech model disclosed in the embodiments of the present invention;
[0093] Figure 4 is a schematic structural diagram of a speech recognition device based on an online speech model disclosed in the embodiments of the present invention;
[0094] Figure 5 It is a schematic structural diagram of another voice recognition device based on an online voice model disclosed in an embodiment of the present invention;
[0095] Figure 6 It is a schematic structural diagram of yet another voice recognition device based on an online voice model disclosed in an embodiment of the present invention. Detailed implementation manners
[0096] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.
[0097] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal including a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.
[0098] Referring to "embodiments" herein means that the specific features, structures or characteristics described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0099] The present invention discloses a voice recognition method, device and storage medium based on an online voice model. The model parameters of the online voice model can be updated in a timely manner as the voice recognition environment changes, so as to more accurately recognize voice commands from environmental sounds, thereby improving the accuracy of voice recognition. The following will be described in detail respectively.
[0100] To better understand the voice recognition method, device and storage medium based on an online voice model disclosed in the present invention, the scenario of voice recognition based on an online voice model will be described first. Specifically, the scenario of voice recognition based on an online voice model can be as Figure 1 shown Figure 1 It is a schematic diagram of the scenario of the voice recognition method based on an online voice model disclosed in an embodiment of the present invention. AsFigure 1 As shown, the voice recognition environment can be various environments such as a home environment, a company environment, a shopping mall environment, a carriage environment, etc., and the embodiments of the present invention do not make limitations. The voice recognition environment may include a user, an online voice device, an online perception device, and a noise source. Among them, the online voice device can obtain the environmental information and sound information in the voice recognition environment. Among them, the sound information includes the voice control sound information issued by the user and noise. The noise can be the sound in the voice recognition environment other than the voice control sound information issued by the user, or some typical noises in the voice recognition environment. For example, it can be the sound emitted by a washing machine, the sound of a hair dryer, the sound of a car engine, the sound of a TV set, etc. When the above various noises are mixed with the voice control commands issued by the user, the voice recognition accuracy of the online voice model will be reduced.
[0101] Optionally, the online voice model can communicate with the online voice device and the online perception device through a cloud server, or can be integrated in the online voice device and communicate with the online voice device and the online perception device. The online voice model is preliminarily trained when leaving the factory, used to analyze the sound information, obtain a sound analysis result, determine a voice control command matching the sound analysis result, and control at least one target device in the voice recognition environment to perform an operation matching the voice control command. The target device can be devices such as an air conditioner, a light, an electric curtain, a TV set, etc., and the embodiments of the present invention do not make limitations.
[0102] Optionally, the environmental information includes target user information related to the user. The target user information can be recognized by an online perception device set in the voice recognition environment. It can be infrared information related to the target user detected by an infrared detector, radar information related to the target user detected by a radar detector, or signal information related to the wearable device carried by the target user detected by a signal detector. The embodiments of the present invention do not make limitations. According to whether there is target user information related to the user in the environmental information, it can be judged whether the user exists in the voice recognition environment. Optionally, the first voice recognition condition can be whether the sound information can be recognized by the online voice model, or whether the characteristics of the sound information have been pre-stored in a storage database that the online voice model can read.
[0103] Optionally, for example, when the user leaves the speech recognition environment, the environmental information obtained by the online sensing device no longer contains environmental information related to the user. By analyzing this environmental information, information indicating that the user is not present in the speech recognition environment is obtained. At this time, it is determined whether the sound information emitted by the noise source satisfies a preset first speech recognition condition. When the sound information emitted by the noise source does not satisfy the preset first speech recognition condition, it indicates that the current sound information can be used as the sound information for training the online language model. Then, an upgrade operation is performed on the online speech model according to this sound information to obtain an upgraded online speech model. A speech recognition operation is performed on the speech recognition environment through the upgraded online speech model to obtain a speech recognition result. The upgraded online speech model can recognize the latest noise information in the speech environment, thereby distinguishing the noise information from the user's voice command information in the sound information, and thus more accurately recognizing the voice command from the environmental sound, thereby improving the accuracy of speech recognition.
[0104] It should be noted that Figure 1 The scene architecture shown is only for representing the scene used in the speech recognition method based on the online speech model. The users, online sensing devices, online speech devices, and noise sources involved are only shown schematically. The specific structure / size / shape / location / installed method, etc. can be adaptively adjusted according to the actual scene. Figure 1 The scene architecture shown does not limit this.
[0105] The application scenarios used in the speech recognition method based on the online speech model are described above. Next, the speech recognition method and device, and storage medium based on the online speech model will be described in detail.
[0106] Embodiment 1
[0107] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a speech recognition method based on an online speech model disclosed in an embodiment of the present invention. Among them, Figure 2 The described speech recognition method based on the online speech model can be applied to a speech recognition device based on the online speech model. The speech recognition device based on the online speech model can be integrated in a cloud server or a local server, which is not limited in the embodiments of the present invention. As Figure 2 shown, the speech recognition method based on the online speech model may include the following operations:
[0108] 201. Obtain the environmental information of the speech recognition environment and the sound information in the speech recognition environment, and determine whether there is a user in the speech recognition environment according to the environmental information and the sound information.
[0109] In an embodiment of the present invention, optionally, the voice recognition environment may be an environment in a home, an environment in an office, or an environment in a shopping mall. The embodiment of the present invention does not make any limitation thereto.
[0110] In an embodiment of the present invention, optionally, the environmental information includes target user information related to the user. The target user information can be recognized by an online sensing device disposed in the voice recognition environment, which may be infrared information related to the target user detected by an infrared detector, radar information related to the target user detected by a radar detector, or signal information related to a wearable device carried by the target user detected by a signal detector. The embodiment of the present invention does not make any limitation thereto. The sound information includes all sound information in the voice recognition environment, which may include voice command information issued by the user and noise information other than the voice command information. The sound information can be obtained by an independent sound sensor disposed in the voice recognition environment, or by a voice acquisition device built in an online voice device in the voice recognition environment, such as a voice acquisition device built in an air conditioner that can be voice-controlled. The embodiment of the present invention does not make any limitation thereto.
[0111] In an embodiment of the present invention, based on the environmental information and the sound information, it is determined whether a user exists in the voice recognition environment, which can comprehensively determine whether a user exists in the voice recognition environment by combining the user information detected by the detection device and the sound information in the environment, thereby avoiding an inaccurate judgment result caused when the detection device fails to accurately detect the information indicating the existence of the user. For example, when the online sensing device fails to sense the information indicating the existence of the user, but the sound of a hair dryer is analyzed from the sound information, and according to the living habit, the sound of the hair dryer only appears when the user exists in the voice recognition environment. Therefore, it should be determined at this time that the user exists in the voice recognition environment. It can be seen that the embodiment of the present invention can more accurately determine whether a user exists in the voice recognition environment, and avoids the problem of inaccurate judgment results caused when the online sensing device fails to sense the information indicating the existence of the user or the sensing result is incorrect.
[0112] In an optional embodiment, based on the environmental information and the sound information, determining whether a user exists in the voice recognition environment may include:
[0113] According to the sound information, extract the sound feature information in the sound information, and generate a sound feature set based on all the sound feature information;
[0114] Determine whether there is target user information related to the user in the environmental information and whether there is a pre-determined background sound feature in the sound information feature set;
[0115] When it is determined that there is target user information in the environmental information and / or there is background sound feature in the sound information feature set, it is determined that there is a user in the speech recognition environment; when it is determined that there is no target user information in the environmental information and there is no background sound feature in the sound information feature set, it is determined that there is no user in the speech recognition environment.
[0116] In this optional embodiment, the background sound information is the sound that only exists when there is a user in the speech recognition environment, such as the sound of a hair dryer, the sound of watching TV, the sound of a range hood, the sound of opening a door, etc. When the above sound information exists in the sound information, it indicates the presence of a user. In this optional embodiment, the background sound information is obtained by acquiring the sound information in a specific time period before and after the user issues a voice command.
[0117] Specifically, the pre-determined background sound feature can be determined in the following way: acquire the environmental sound information within the pre-set speech acquisition duration corresponding to the speech recognition environment, determine the environmental sound information as the background sound information, and determine the background sound feature according to the background sound information;
[0118] Among them, the pre-set speech acquisition duration includes a first time period before acquiring the user's voice command and / or a second time period after acquiring the user's voice command.
[0119] In this optional embodiment, the sound feature information in the sound information is extracted to generate a sound feature set, and it is judged whether there is a pre-determined background sound feature in the sound information feature set. Only when it is determined that there is no target user information in the environmental information and there is no background sound feature in the sound information feature set, it is determined that there is no user in the speech recognition environment, so as to more accurately judge whether there is a user in the speech recognition environment.
[0120] 202. When it is determined that there is no user in the speech recognition environment, judge whether the sound information meets the pre-set first speech recognition condition.
[0121] In an embodiment of the present invention, optionally, the first voice recognition condition is used to measure whether the voice information can be used to upgrade the online voice model. For example, it can be whether the online voice model can recognize the voice information. For voice information that cannot be recognized, the online voice model is trained with this voice information so that the online voice model can recognize this voice information. It is also possible to combine the voice information with a plurality of predetermined voice commands to generate a first voice recognition test set, perform a test operation on the online voice model according to the first voice recognition test set, and determine whether the voice information meets the preset first voice recognition condition based on the number of voice commands that meet the test results. For example: Combine the voice information with voice command information such as "turn on the air conditioner", "raise the air conditioner temperature by one degree", "turn off the air conditioner", "raise the air conditioner temperature by one degree", "turn on the bedroom light", "turn off the bedroom light", "open the curtain", "close the curtain" and other predetermined voice commands to generate a first voice recognition test set, perform a test operation on the online voice model according to the first voice recognition test set, and determine that the voice information meets the preset first voice recognition condition when the online voice model can correctly recognize all the above voice command information, and determine that the voice information does not meet the preset first voice recognition condition when the online voice model cannot correctly recognize all the above voice command information.
[0122] It can be seen that in an embodiment of the present invention, by determining whether the voice information meets the preset first voice recognition condition, it is determined whether to use the voice information to perform an upgrade operation on the online voice model, so as to screen out useful voice information to perform an upgrade operation on the online voice model, improving the efficiency of upgrading the online voice model.
[0123] In an embodiment of the present invention, the method may further include:
[0124] When it is determined that there is a user in the voice recognition environment, determine whether the voice information meets the preset second voice recognition condition. When it is determined that the voice information meets the preset second voice recognition condition, analyze the voice information to obtain a voice analysis result, determine a voice control command that matches the voice analysis result, and control at least one target device in the voice recognition environment to perform an operation that matches the voice control command, where the voice analysis result includes the voice command information issued by the user;
[0125] When it is determined that the voice information does not meet the preset second voice recognition condition, determine whether the voice information exists in the voice database corresponding to the online voice model. When it is determined that the voice information does not exist in the voice database, store the voice information in the voice database.
[0126] In an embodiment of the present invention, optionally, the second voice recognition condition is used to measure whether the current voice information is voice control instruction information issued by the user. For example, it may be whether the voice information corresponds to a pre-stored voice instruction, or whether there is a preset wake-up voice information in the voice information. For example, whether there are preset wake-up voice information such as "Hello", "Please note", "Start recognition", "Hello". For example, when the voice information is "Hello, please turn on the air conditioner", when it is determined that "Hello" exists in the voice information, it is determined that the voice information meets the preset second voice recognition condition. Analyze the voice information, and the voice analysis result is "Please turn on the air conditioner". Determine the voice control instruction matching the voice analysis result as "Turn on the air conditioner", and control the air conditioner in the voice recognition environment to turn on.
[0127] In an embodiment of the present invention, optionally, the voice database corresponding to the online voice model may be a voice feature information database that saves the voice information after converting it into voice feature information, or a storage database that directly stores the recorded voice information. All voice information is saved in the voice database. Further optionally, when performing an upgrade operation on the online voice model, all the voice information saved in the voice database and the currently obtained voice information can be used together as a training set to perform the upgrade operation on the online voice model. Still further optionally, when the online voice model performs a voice recognition operation on the voice recognition environment, all the voice information saved in the voice database can be combined to assist in identifying the noise in the voice information, so as to perform voice recognition more accurately.
[0128] It can be seen that implementing the voice recognition method based on the online voice model in this embodiment can only perform the upgrade operation of the online voice model when it is determined that there is no user in the voice recognition environment. When it is determined that there is a user in the voice recognition environment, for the voice information including the user's voice instruction, corresponding operations are performed according to the voice instruction in the voice information. For the voice information that does not include the user's voice instruction, it is determined whether the voice information exists in the voice database corresponding to the online voice model. When it is determined that the voice information does not exist in the voice database, the voice information is stored in the voice database, and the database can be further used for voice recognition assistance or online voice model upgrade. Upgrading only when there is no user in the voice recognition environment avoids the situation where when there is a user in the voice recognition environment and voice information including the user's voice instruction is obtained, the voice information cannot be recognized to obtain the voice instruction.
[0129] 203. When it is determined that the voice information does not meet the preset first voice recognition condition, perform an upgrade operation on the online voice model according to the voice information to obtain an upgraded online voice model.
[0130] In an embodiment of the present invention, optionally, performing an upgrade operation on the online speech model according to the voice information may include: training the online speech model used by the online speech device according to the voice information until convergence, and then updating at least one parameter in the online speech model to complete the upgrade of the online speech model.
[0131] It can be seen that when implementing the speech recognition method based on the online speech model in this embodiment, when it is determined that the voice information does not meet the preset first speech recognition condition, an upgrade operation is performed on the online speech model according to the voice information to obtain an upgraded online speech model, so that the upgraded online speech model can recognize the voice information that does not meet the first speech recognition condition, and thus more accurately recognize the voice command from the ambient sound, improving the accuracy of speech recognition.
[0132] 204. Perform a speech recognition operation on the speech recognition environment through the upgraded online speech model to obtain a speech recognition result.
[0133] In an embodiment of the present invention, optionally, performing a speech recognition operation on the speech recognition environment through the upgraded online speech model to obtain a speech recognition result may include: analyzing the voice information through the upgraded online speech model to obtain a set of voice features, extracting the voice information in the voice information that the online speech model can recognize, and obtaining a subset of the set of voice features, where each subset of the set of voice features represents a type of voice information that the online speech model can recognize; real-time judging whether the subset of the set of voice features includes the voice command information issued by the user, and when it is determined that a certain subset includes the voice command information issued by the user, determining a voice control command that matches the voice information corresponding to the subset, and controlling at least one target device in the speech recognition environment to perform an operation that matches the voice control command. For example: in the current speech recognition environment, there are hair dryer sounds, TV sounds, and voice command sounds issued by the user. Analyze the voice information through the upgraded online speech model to obtain a set of voice features, extract the voice information in the voice information that the online speech model can recognize, including hair dryer sounds, TV sounds, and voice command sounds issued by the user, and obtain a subset of the set of voice features, including a hair dryer sound feature subset, a TV sound feature subset, and a voice command subset issued by the user. Real-time judge whether the subset of the set of voice features includes the voice command information issued by the user, and when it is determined that the voice command subset issued by the user includes the voice command information issued by the user, determine a voice control command that matches the voice information corresponding to the subset, and control at least one target device in the speech recognition environment to perform an operation that matches the voice control command.
[0134] In an embodiment of the present invention, optionally, the speech recognition operation may include a speech pre-recognition operation and a speech command acquisition operation. Among them, the speech pre-recognition operation includes: acquiring sound information in the speech recognition environment, and determining whether there is user sound information and / or preset wake-up sound information in the sound information. If it is determined that there is user sound information and / or preset wake-up sound information in the sound information, the speech command acquisition operation is triggered. The speech command acquisition operation includes: analyzing the sound information through the upgraded online speech model to obtain a sound analysis result, and determining in real time whether the sound analysis result includes the speech command information issued by the user. When it is determined that the sound analysis result includes the speech command information issued by the user, a speech control command matching the sound analysis result is determined; if it is determined that there is no user sound information and preset wake-up sound information in the sound information, the speech pre-recognition operation is continued.
[0135] In an alternative embodiment, before performing an upgrade operation on the online speech model according to the sound information to obtain an upgraded online speech model, the method further includes: determining whether the online speech model needs to be updated. When it is determined that the online speech model needs to be updated, an operation of performing an upgrade operation on the online speech model according to the sound information to obtain an upgraded online speech model is triggered.
[0136] In this alternative embodiment, optionally, determining whether the online speech model needs to be updated may include: determining whether the online speech model meets a preset model self-check condition according to all the sound information; all the sound information may be all the sound information stored in the speech database in step 202, or may be the sound information obtained in the speech recognition environment within a set time period. When it is determined that the online speech model meets the preset model self-check condition, it is determined whether the online speech model needs to be updated according to all the sound information stored in the online speech model; when it is determined that the online speech model needs to be updated, an operation of performing an upgrade operation on the online speech model according to the sound information to obtain an upgraded online speech model is triggered.
[0137] In this alternative embodiment, optionally, determining whether the online speech model meets a preset model self-check condition according to all the sound information may include:
[0138] According to all the sound information, determine the information quantity of all the sound information stored in the online speech model and the update time of the last update of the online speech model. Based on the update time of the last update of the online speech model, determine the update interval duration of the online speech model;
[0139] Determine whether the information quantity is greater than or equal to a preset storage quantity threshold and whether the update interval duration of the online speech model is greater than or equal to a preset self-check duration threshold;
[0140] When it is determined that the number of messages is greater than or equal to a preset storage quantity threshold and / or the update interval duration of the online speech model is greater than or equal to a preset self-check duration threshold, it is determined that the online speech model meets the preset model self-check conditions; when it is determined that the number of messages is less than the preset storage quantity threshold and the update interval duration of the online speech model is less than the preset self-check duration threshold, it is determined that the online speech model does not meet the preset model self-check conditions.
[0141] In this optional embodiment, the situation of meeting the model self-check conditions is that all the voice messages stored in the online speech model reach a certain quantity and / or the duration since the last update of the online speech model reaches a certain self-check duration threshold; among them, all the voice messages stored in the online speech model can be all the voice messages stored in the speech database corresponding to the online speech model, or can be the voice messages in the speech recognition environment within a set time obtained by the online speech model.
[0142] It can be seen that when implementing this optional embodiment, when the quantity of voice messages obtained by the online speech model is small (that is, the training data set for upgrading and training the online speech model is small) and the online speech model has just been updated not long ago, it does not meet the model self-check conditions and will not trigger the execution of the upgrade operation of the online speech model, which can prevent the online speech model from being upgraded frequently and improve the upgrade efficiency of the online speech model.
[0143] In this optional embodiment, further optionally, judging whether the online speech model needs to be updated according to all the voice messages stored in the online speech model may include:
[0144] Generating a test voice set according to all the voice messages and a plurality of pre-determined voice test instructions;
[0145] Performing a test operation on the online speech model according to the test voice set to obtain a model test result, where the model test result includes the voice test results corresponding to each voice test instruction;
[0146] According to the model test result, determining the target quantity of the target test results that meet the preset voice test conditions, and based on this target quantity, judging whether the target quantity is greater than or equal to a preset test quantity threshold;
[0147] When it is determined that the target quantity is greater than or equal to the preset test quantity threshold, it is determined that the online speech model does not need to be updated; when it is determined that the target quantity is less than the preset test quantity threshold, it is determined that the online speech model needs to be updated.
[0148] In this optional embodiment, the preset voice test condition may be that the voice test result corresponding to the voice test instruction is consistent with the instruction represented by the pre-determined voice test instruction. For example, if the pre-determined voice test instruction is "turn on the air conditioner" and the voice test result corresponding to the voice test instruction is also "turn on the air conditioner", then the test result meets the preset voice test condition; if the pre-determined voice test instruction is "turn on the air conditioner" and the voice test result corresponding to the voice test instruction is "turn off the air conditioner", then the test result does not meet the preset voice test condition.
[0149] For example, this optional embodiment may include: combining all voice information with the voice information of pre-determined voice test instructions such as "turn on the air conditioner", "increase the air conditioner temperature by one degree", "turn off the air conditioner", "increase the air conditioner temperature by one degree", "turn on the bedroom light", "turn off the bedroom light", "open the curtain", "close the curtain", etc. to generate a test voice set, performing a test operation on the online voice model according to the test voice set to obtain a model test result, where the model test result includes the voice test result corresponding to each voice test instruction; according to the model test result, determining the target number of target test results where the voice test result corresponding to the voice test instruction is consistent with the instruction represented by the pre-determined voice test instruction, and based on this target number, determining whether the target number is greater than or equal to a preset test number threshold; when it is determined that the target number is greater than or equal to the preset test number threshold, determining that the online voice model does not need to be updated; when it is determined that the target number is less than the preset test number threshold, determining that the online voice model needs to be updated. For example, the preset test number threshold is 7. If the target number of target test results where the voice test result corresponding to the voice test instruction is consistent with the instruction represented by the pre-determined voice test instruction is 6, then it is determined that the target number 6 is less than the preset test number threshold 7, and it is determined that the online voice model needs to be updated; if the target number of target test results where the voice test result corresponding to the voice test instruction is consistent with the instruction represented by the pre-determined voice test instruction is 8, then it is determined that the target number 8 is greater than or equal to the preset test number threshold 7, and it is determined that the online voice model does not need to be updated.
[0150] It can be seen that implementing this optional embodiment can determine whether the online voice model needs to be updated based on all the voice information stored in the online voice model, thereby realizing a self-check operation on the online voice model and evaluating whether the performance of the online voice model meets the requirements of voice recognition. When the online voice model can meet the requirements of current voice recognition, the voice information used for training model upgrade can be stored first and then used for upgrade, improving the efficiency of online voice model upgrade and thus improving the accuracy of voice recognition.
[0151] It can be seen that implementing Figure 2The described voice recognition method based on an online voice model can obtain the environmental information of the voice recognition environment and the sound information in the voice recognition environment, and determine whether there is a user in the voice recognition environment according to the environmental information and the sound information; when it is determined that there is no user in the voice recognition environment, it is determined whether the sound information meets a preset first voice recognition condition; when it is determined that the sound information does not meet the preset first voice recognition condition, an upgrade operation is performed on the online voice model according to the sound information to obtain an upgraded online voice model; a voice recognition operation is performed on the voice recognition environment through the upgraded online voice model to obtain a voice recognition result. In the embodiments of the present invention, it is determined whether to upgrade the online voice model according to the first voice recognition condition, so that the online voice model is upgraded when the user is not in the recognition environment. The upgraded online voice model can recognize the sound information that does not meet the first voice recognition condition, and thus more accurately recognize the voice command from the environmental sound, improving the accuracy of voice recognition.
[0152] Embodiment 2
[0153] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of another voice recognition method based on an online voice model disclosed in the embodiments of the present invention. Among them, Figure 3 the described voice recognition method based on an online voice model can be applied to a voice recognition device based on an online voice model. The voice recognition device based on an online voice model can be integrated in a cloud server or a local server, which is not limited in the embodiments of the present invention. As Figure 3 shown, the voice recognition method based on an online voice model may include the following operations:
[0154] 301. For each environment to be recognized, a model parameter set corresponding to the environment to be recognized is respectively set, and an online voice model corresponding to the environment to be recognized is generated according to the model parameter set corresponding to the environment to be recognized.
[0155] In the embodiments of the present invention, the model parameter set includes at least one parameter in the online voice model trained to convergence; for different environments to be recognized, different online voice models are used for voice recognition. For example, for a home environment, a company environment, a shopping mall environment, and a carriage environment, the online voice models trained to convergence in the above environments are respectively obtained. When the voice recognition environment of the online voice model switches from the company environment to the home environment, there is no need to upgrade the online voice model again according to the home environment. It only needs to switch the online voice model from the company environment voice model to the home environment voice model, improving the efficiency of upgrading the online voice model.
[0156] In an embodiment of the present invention, optionally, for each environment to be recognized, a set of model parameters corresponding to the environment to be recognized is set respectively. The set of model parameters for a certain environment to be recognized can be obtained by training an online speech model using the voice information in the environment to be recognized and training until convergence. When the online speech model switches from another speech recognition environment to the current speech recognition environment, an online speech model corresponding to the current speech recognition environment is generated according to the set of model parameters corresponding to the current speech recognition environment.
[0157] It can be seen that implementing the speech recognition method based on an online speech model in this embodiment can perform speech recognition using different online speech models for different environments to be recognized. When the speech recognition environment of the online speech model switches from another environment to the current environment, there is no need to upgrade the online speech model according to the voice information in the current environment. It only needs to generate an online speech model corresponding to the current speech recognition environment according to the set of model parameters corresponding to the current environment, thereby improving the efficiency of upgrading the online speech model.
[0158] 302. Obtain the environmental information of the speech recognition environment and the voice information in the speech recognition environment, and determine whether there is a user in the speech recognition environment according to the environmental information and the voice information.
[0159] 303. When it is determined that there is no user in the speech recognition environment, determine whether the voice information meets a preset first speech recognition condition.
[0160] 304. When it is determined that the voice information does not meet the preset first speech recognition condition, determine a target online speech model corresponding to the speech recognition environment from all the online speech models corresponding to the environments to be recognized according to the voice information; perform an upgrade operation on the target online speech model according to the voice information to obtain an upgraded target online speech model.
[0161] In an embodiment of the present invention, optionally, determining a target online speech model corresponding to the speech recognition environment from all the online speech models corresponding to the environments to be recognized according to the voice information may include: obtaining a set of model parameters of the corresponding speech recognition environment according to the voice information of the current speech recognition environment, generating a target online speech model corresponding to the speech recognition environment according to the set of model parameters; performing an upgrade operation on the target online speech model according to the voice information to obtain an upgraded target online speech model.
[0162] 305. Perform a speech recognition operation on the speech recognition environment through the upgraded target online speech model to obtain a speech recognition result.
[0163] In the embodiments of the present invention, for other descriptions of steps 302-303 and step 305, please refer to the detailed descriptions of steps 201-202 and step 204 in Embodiment 1, and the embodiments of the present invention will not be repeated here.
[0164] In an optional embodiment, when it is determined that there is a user in the voice recognition environment, it is determined whether the voice information meets a preset second voice recognition condition. When it is determined that the voice information does not meet the preset second voice recognition condition, it is determined whether there is voice information in the voice database corresponding to the online voice model. When it is determined that there is no voice information in the voice database, the voice information is stored in the voice database.
[0165] In this optional embodiment, the second voice recognition condition is used to measure whether the current voice information is voice control instruction information issued by the user. For example, it can be whether the voice information corresponds to a pre-stored voice instruction, or whether there is a preset wake-up voice information in the voice information, such as whether there are preset wake-up voice information such as "Hello", "Please note", "Start recognition", "Hello". In the embodiments of the present invention, optionally, the voice database corresponding to the online voice model can be a voice feature information database that saves the voice information after converting it into voice feature information, or a storage database that directly stores the recorded voice information. All voice information is saved in this voice database. Further optionally, when performing an upgrade operation on the online voice model, all the voice information saved in this voice database and the currently obtained voice information can be used together as a training set to perform the upgrade operation on the online voice model. Still further optionally, when the online voice model performs a voice recognition operation on the voice recognition environment, all the voice information saved in this voice database can be combined to assist in identifying the noise in the voice information, so as to perform voice recognition more accurately.
[0166] In the embodiments of the present invention, for the voice information that does not meet the preset second voice recognition condition saved in the voice database, optionally, a data cleaning operation can be performed on the database to prevent the storage space of the voice database from being too large. Specifically, the online voice model can be used to test the voice information in the voice database regularly, and the voice information that meets the preset first voice recognition condition is screened out and deleted, so as to avoid reusing the voice information that meets the preset first voice recognition condition in the voice database later, and reduce the calculation amount of the upgrade operation of the online voice model.
[0167] It can be seen that by implementing the speech recognition method based on an online speech model according to this embodiment, different online speech models can be used for speech recognition in different environments to be recognized. When the speech recognition environment of the online speech model switches from another environment to the current environment, there is no need to upgrade the online speech model according to the sound information in the current environment. Instead, only an online speech model corresponding to this speech recognition environment needs to be generated according to the set of model parameters corresponding to the current environment, which improves the efficiency of upgrading the online speech model. The upgraded online speech model can recognize sound information that does not meet the first speech recognition condition, and thus more accurately recognize speech commands from the environmental sounds, improving the accuracy of speech recognition.
[0168] Embodiment III
[0169] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a speech recognition device based on an online speech model disclosed in an embodiment of the present invention. As Figure 4 shown, the speech recognition device based on an online speech model may include:
[0170] An acquisition module 401, configured to acquire environmental information of a speech recognition environment and sound information in the speech recognition environment;
[0171] A first determination module 402, configured to determine whether there is a user in the speech recognition environment according to the environmental information and the sound information;
[0172] A second determination module 403, configured to determine whether the sound information meets a preset first speech recognition condition when the first determination module 402 determines that there is no user in the speech recognition environment;
[0173] An upgrade module 404, configured to perform an upgrade operation on the online speech model according to the sound information to obtain an upgraded online speech model when the second determination module 403 determines that the sound information does not meet the preset first speech recognition condition;
[0174] A speech recognition module 405, configured to perform a speech recognition operation on the speech recognition environment through the upgraded online speech model to obtain a speech recognition result.
[0175] It can be seen that by implementing Figure 4The described device can obtain the environmental information of the speech recognition environment and the sound information in the speech recognition environment, and determine whether there is a user in the speech recognition environment according to the environmental information and the sound information; when it is determined that there is no user in the speech recognition environment, it determines whether the sound information meets a preset first speech recognition condition; when it is determined that the sound information does not meet the preset first speech recognition condition, it performs an upgrade operation on the online speech model according to the sound information to obtain an upgraded online speech model; and performs a speech recognition operation on the speech recognition environment through the upgraded online speech model to obtain a speech recognition result. In the embodiment of the present invention, it is determined whether the online speech model needs to be upgraded according to the first speech recognition condition, so as to upgrade the online speech model when the user is not in the recognition environment. The upgraded online speech model can recognize the sound information that does not meet the first speech recognition condition, and thus more accurately recognize the speech command from the environmental sound, improving the accuracy of speech recognition.
[0176] In an alternative embodiment, as Figure 5 shown, the second determination module 403 of the device is further configured to, when the first determination module 402 determines that there is a user in the speech recognition environment, determine whether the sound information meets a preset second speech recognition condition;
[0177] The device may further include:
[0178] A control module 406, configured to, when the second determination module 403 determines that the sound information meets the preset second speech recognition condition, analyze the sound information to obtain a sound analysis result, determine a speech control command matching the sound analysis result, and control at least one target device in the speech recognition environment to perform an operation matching the speech control command, where the sound analysis result includes the speech command information issued by the user;
[0179] A third determination module 407, configured to, when the second determination module 403 determines that the sound information does not meet the preset second speech recognition condition, determine whether the sound information exists in the speech database corresponding to the online speech model, and when it is determined that the sound information does not exist in the speech database, store the sound information in the speech database.
[0180] It can be seen that implementing Figure 5The described device can perform an upgrade operation on the online speech model only when it is determined that there is no user in the speech recognition environment. When it is determined that there is a user in the speech recognition environment, for the sound information including the user's speech command, corresponding operations are performed according to the speech command in the sound information. For the sound information that does not include the user's speech command, it is determined whether the sound information exists in the speech database corresponding to the online speech model. When it is determined that the sound information does not exist in the speech database, the sound information is stored in the speech database, and this database can be further used for the assistance of speech recognition or the upgrade of the online speech model. Performing the upgrade only when there is no user in the speech recognition environment avoids the situation where, when there is a user in the speech recognition environment and sound information including the user's speech command is obtained, the sound information cannot be recognized to obtain the speech command, thereby improving the accuracy of speech recognition.
[0181] In another alternative embodiment, as Figure 5 shown, the device may further include:
[0182] A model self-checking module 408, configured to, before the upgrade module 404 performs an upgrade operation on the online speech model according to the sound information to obtain an upgraded online speech model, determine whether the online speech model meets a preset model self-checking condition according to all the sound information;
[0183] The second determination module 403 is further configured to, when the model self-checking module 408 determines that the online speech model meets the preset model self-checking condition, determine whether the online speech model needs to be updated according to all the sound information stored in the online speech model;
[0184] The model self-checking module 408 is further configured to, when it is determined that the online speech model needs to be updated, trigger the upgrade module 404 to perform an operation of performing an upgrade operation on the online speech model according to the sound information to obtain an upgraded online speech model.
[0185] It can be seen that implementing Figure 5 the described device can determine whether the online speech model needs to be updated according to all the sound information stored in the online speech model, thereby implementing a self-checking operation on the online speech model and evaluating whether the performance of the online speech model meets the requirements of speech recognition. When the online speech model can meet the current requirements of speech recognition, the sound information used for training model upgrade can be stored first and then upgraded later, improving the efficiency of online speech model upgrade and thus improving the accuracy of speech recognition.
[0186] In yet another alternative embodiment, the specific manner in which the model self-checking module 408 determines whether the online speech model meets the preset model self-checking condition according to all the sound information includes:
[0187] Determine the quantity of all the voice information stored in the online voice model and the update time of the last update of the online voice model according to all the voice information. Based on the update time of the last update of the online voice model, determine the update interval duration of the online voice model;
[0188] Judge whether the quantity of information is greater than or equal to a preset storage quantity threshold and whether the update interval duration of the online voice model is greater than or equal to a preset self-check duration threshold;
[0189] When it is judged that the quantity of information is greater than or equal to the preset storage quantity threshold and / or the update interval duration of the online voice model is greater than or equal to the preset self-check duration threshold, determine that the online voice model meets the preset model self-check condition;
[0190] When it is judged that the quantity of information is less than the preset storage quantity threshold and the update interval duration of the online voice model is less than the preset self-check duration threshold, determine that the online voice model does not meet the preset model self-check condition.
[0191] It can be seen that implementing Figure 5 the described device can, when the quantity of voice information obtained by the online voice model is small (that is, the training data set for upgrading and training the online voice model is small) and the online voice model has just been updated not long ago, not meet the model self-check condition and will not trigger the execution of the upgrade operation of the online voice model, which can prevent the online voice model from being upgraded frequently, improve the upgrade efficiency of the online voice model, and thus improve the accuracy of speech recognition.
[0192] In another optional embodiment, the specific manner in which the second judgment module 403 judges whether the online voice model needs to be updated according to all the voice information stored in the online voice model includes:
[0193] Generate a test voice set according to all the voice information and a plurality of predetermined voice test instructions; perform a test operation on the online voice model according to the test voice set to obtain a model test result, where the model test result includes the voice test result corresponding to each voice test instruction; according to the model test result, determine the target quantity of the target test results that meet the preset voice test conditions, and based on the target quantity, judge whether the target quantity is greater than or equal to a preset test quantity threshold; when it is judged that the target quantity is greater than or equal to the preset test quantity threshold, determine that the online voice model does not need to be updated; when it is judged that the target quantity is less than the preset test quantity threshold, determine that the online voice model needs to be updated.
[0194] It can be seen that implementing Figure 5The described device can determine whether the online voice model needs to be updated based on all the voice information stored in the online voice model, so as to implement a self-check operation on the online voice model and evaluate whether the performance of the online voice model meets the requirements of speech recognition. When the online voice model can meet the requirements of the current speech recognition, the voice information used for training the model upgrade can be stored first and then upgraded later, which improves the efficiency of upgrading the online voice model and further improves the accuracy of speech recognition.
[0195] In another optional embodiment, the specific manner in which the first determination module 402 determines whether there is a user in the speech recognition environment according to the environmental information and the voice information includes:
[0196] According to the voice information, extract the voice feature information in the voice information, and generate a voice feature set based on all the voice feature information; determine whether there is target user information related to the user in the environmental information and whether there are predetermined background voice features in the voice information feature set.
[0197] When it is determined that there is target user information in the environmental information and / or there are background voice features in the voice information feature set, it is determined that there is a user in the speech recognition environment; when it is determined that there is no target user information in the environmental information and there are no background voice features in the voice information feature set, it is determined that there is no user in the speech recognition environment.
[0198] Among them, the predetermined background voice features are determined by the following method:
[0199] Obtain the environmental voice information within the preset voice acquisition duration corresponding to the speech recognition environment, determine the environmental voice information as the background voice information, and determine the background voice features according to the background voice information; wherein, the preset voice acquisition duration includes a first time period before obtaining the user voice command and / or a second time period after obtaining the user voice command.
[0200] It can be seen that the implementation Figure 5 The described device can extract the voice feature information in the voice information, generate a voice feature set, and determine whether there are predetermined background voice features in the voice information feature set. Only when it is determined that there is no target user information in the environmental information and there are no background voice features in the voice information feature set, it is determined that there is no user in the speech recognition environment, so as to more accurately determine whether there is a user in the speech recognition environment, and further improve the accuracy of speech recognition.
[0201] In another optional embodiment, the device further includes:
[0202] A parameter configuration module 409, configured to, before the upgrade module 404 performs an upgrade operation on the online speech model according to the voice information to obtain an upgraded online speech model, for each environment to be recognized, respectively set a set of model parameters corresponding to the environment to be recognized, where the set of model parameters includes at least one parameter in the online speech model trained to convergence; for each environment to be recognized, generate an online speech model corresponding to the environment to be recognized according to the set of model parameters corresponding to the environment to be recognized.
[0203] Moreover, the specific manner in which the upgrade module 404 performs an upgrade operation on the online speech model according to the voice information to obtain an upgraded online speech model includes:
[0204] According to the voice information, determine a target online speech model corresponding to the speech recognition environment from the online speech models corresponding to all environments to be recognized; perform an upgrade operation on the target online speech model according to the voice information to obtain an upgraded target online speech model.
[0205] It can be seen that the Figure 5 described device can perform speech recognition using different online speech models for different environments to be recognized. When the speech recognition environment of the online speech model switches from another environment to the current environment, there is no need to upgrade the online speech model according to the voice information in the current environment. It only needs to generate an online speech model corresponding to the speech recognition environment according to the set of model parameters corresponding to the current environment, which improves the efficiency of upgrading the online speech model and thus improves the accuracy of speech recognition.
[0206] Embodiment 4
[0207] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of another speech recognition device based on an online speech model disclosed in an embodiment of the present invention. As Figure 6 shown, the speech recognition device based on an online speech model may include:
[0208] A memory 501 storing executable program code;
[0209] A processor 502 coupled to the memory 501;
[0210] The processor 502 invokes the executable program code stored in the memory 501 and executes the steps in the speech recognition method based on an online speech model described in Embodiment 1 or Embodiment 2 of the present invention.
[0211] Embodiment 5
[0212] An embodiment of the present invention discloses a computer-readable storage medium. The computer-readable storage medium stores computer instructions, which when called, are used to execute the steps in the speech recognition method based on an online speech model described in Embodiment 1 or Embodiment 2 of the present invention.
[0213] Embodiment 6
[0214] An embodiment of the present invention discloses a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the speech recognition method based on an online speech model described in Embodiment 1 or Embodiment 2.
[0215] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0216] Through the above specific descriptions of the embodiments, those skilled in the art can clearly understand that each implementation can be realized by means of software plus a necessary general hardware platform, and of course, it can also be realized by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0217] Finally, it should be noted that: The voice recognition method, device, and storage medium based on an online voice model disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention. They are only used to illustrate the technical solutions of the present invention, rather than to limit it; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A speech recognition method based on an online speech model, characterized in that, the method includes: obtaining environmental information of the speech recognition environment and sound information in the speech recognition environment, and judging whether there is a user in the speech recognition environment according to the environmental information and the sound information; when it is judged that there is no such user in the speech recognition environment, judging whether the sound information meets a preset first speech recognition condition; when it is judged that the sound information does not meet the preset first speech recognition condition, performing an upgrade operation on the online speech model according to the sound information to obtain the upgraded online speech model; performing a speech recognition operation on the speech recognition environment through the upgraded online speech model to obtain a speech recognition result.
2. The speech recognition method based on an online speech model according to claim 1, characterized in that, the method further includes: when it is judged that there is such user in the speech recognition environment, judging whether the sound information meets a preset second speech recognition condition, and when it is judged that the sound information meets the preset second speech recognition condition, analyzing the sound information to obtain a sound analysis result, determining a speech control instruction matching the sound analysis result, and controlling at least one target device in the speech recognition environment to perform an operation matching the speech control instruction, wherein the sound analysis result includes speech instruction information issued by the user; when it is judged that the sound information does not meet the preset second speech recognition condition, judging whether the sound information exists in the speech database corresponding to the online speech model, and when it is judged that the sound information does not exist in the speech database, storing the sound information into the speech database.
3. The speech recognition method based on an online speech model according to claim 2, characterized in that, before performing the upgrade operation on the online speech model according to the sound information to obtain the upgraded online speech model, the method further includes: judging whether the online speech model meets a preset model self-checking condition according to all the sound information; when it is judged that the online speech model meets the preset model self-checking condition, judging whether the online speech model needs to be updated according to all the sound information stored in the online speech model; when it is judged that the online speech model needs to be updated, triggering the execution of the operation of performing the upgrade operation on the online speech model according to the sound information to obtain the upgraded online speech model.
4. The speech recognition method based on an online speech model according to claim 3, characterized in that, judging whether the online speech model meets a preset model self-checking condition according to all the sound information includes: determining the information quantity of all the sound information stored in the online speech model and the update time of the last update of the online speech model according to all the sound information, and determining the update interval duration of the online speech model based on the update time of the last update of the online speech model; Determine whether the quantity of the information is greater than or equal to a preset storage quantity threshold and whether the update interval duration of the online speech model is greater than or equal to a preset self-check duration threshold; When it is determined that the quantity of the information is greater than or equal to the preset storage quantity threshold and / or the update interval duration of the online speech model is greater than or equal to the preset self-check duration threshold, determine that the online speech model meets the preset model self-check conditions; When it is determined that the quantity of the information is less than the preset storage quantity threshold and the update interval duration of the online speech model is less than the preset self-check duration threshold, determine that the online speech model does not meet the preset model self-check conditions.
5. The speech recognition method based on an online speech model according to claim 3, wherein, the determining whether the online speech model needs to be updated according to all the voice information stored in the online speech model includes: generating a test speech set according to all the voice information and a plurality of pre-determined voice test instructions; performing a test operation on the online speech model according to the test speech set to obtain a model test result, wherein the model test result includes a voice test result corresponding to each voice test instruction; determining the target quantity of target test results that meet the preset voice test conditions according to the model test result, and based on the target quantity, determining whether the target quantity is greater than or equal to a preset test quantity threshold; when it is determined that the target quantity is greater than or equal to the preset test quantity threshold, determine that the online speech model does not need to be updated; when it is determined that the target quantity is less than the preset test quantity threshold, determine that the online speech model needs to be updated.
6. The speech recognition method based on an online speech model according to any one of claims 1-5, wherein, the determining whether there is a user in the speech recognition environment according to the environment information and the voice information includes: extracting voice feature information from the voice information according to the voice information, and generating a voice feature set based on all the voice feature information; determining whether there is target user information related to the user in the environment information and whether there is a pre-determined background voice feature in the voice information feature set; when it is determined that there is the target user information in the environment information and / or there is the background voice feature in the voice information feature set, determine that there is a user in the speech recognition environment; when it is determined that there is no target user information in the environment information and there is no background voice feature in the voice information feature set, determine that there is no user in the speech recognition environment; wherein, the pre-determined background voice feature is determined by the following method: acquiring environment voice information within a preset voice acquisition duration corresponding to the speech recognition environment, determining the environment voice information as background voice information, and determining the background voice feature according to the background voice information; Among them, the preset voice acquisition duration includes a first time period before the user voice command is acquired and / or a second time period after the user voice command is acquired.
7. The voice recognition method based on an online voice model according to any one of claims 1-5, characterized in that, before performing the upgrade operation on the online voice model according to the sound information to obtain the upgraded online voice model, the method further includes: For each environment to be recognized, respectively set a model parameter set corresponding to the environment to be recognized, where the model parameter set includes at least one parameter in the online voice model trained to convergence; For each of the environments to be recognized, generate an online voice model corresponding to the environment to be recognized according to the model parameter set corresponding to the environment to be recognized; And, performing the upgrade operation on the online voice model according to the sound information to obtain the upgraded online voice model includes: According to the sound information, determine a target online voice model corresponding to the voice recognition environment from all the online voice models corresponding to the environments to be recognized; Perform an upgrade operation on the target online voice model according to the sound information to obtain the upgraded target online voice model.
8. A voice recognition device based on an online voice model, characterized in that, the device includes: An acquisition module, configured to acquire environment information of a voice recognition environment and sound information in the voice recognition environment; A first judgment module, configured to judge whether there is a user in the voice recognition environment according to the environment information and the sound information; A second judgment module, configured to judge whether the sound information meets a preset first voice recognition condition when the first judgment module judges that there is no user in the voice recognition environment; An upgrade module, configured to perform an upgrade operation on the online voice model according to the sound information to obtain the upgraded online voice model when the second judgment module judges that the sound information does not meet the preset first voice recognition condition; A voice recognition module, configured to perform a voice recognition operation on the voice recognition environment through the upgraded online voice model to obtain a voice recognition result.
9. A voice recognition device based on an online voice model, characterized in that, the device includes: a memory storing executable program code; a processor coupled to the memory; the processor calls the executable program code stored in the memory and executes the voice recognition method based on an online voice model according to any one of claims 1-7.
10. A computer storage medium, characterized in that, the computer storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the voice recognition method based on an online voice model according to any one of claims 1-7.