Updating method and device, local host, server, system and storage medium
By updating semantic template configuration information in a smart home environment and using the server to obtain semantic recognition results, the problem of offline semantic recognition model is solved, and the flexibility of user experience and voice control is improved.
Patent Information
- Application Number
- CN202510348498.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the iteration of offline semantic recognition models is not timely, resulting in poor user experience and being unable to quickly respond to users' voice control needs in smart home environments.
An update method for offline speech recognition module is provided. By maintaining semantic template configuration information and offline semantic recognition model on the local host, the recognition failure record of semantic request information is used for online updates, and the online semantic recognition model on the server is used to obtain and update semantic recognition results.
It realizes that without frequently updating the offline semantic recognition model, quickly responding to new instructions by updating the semantic template configuration information, improving the user experience and adapting to the language habits and personalized needs of different users.
Smart Images

Figure CN119993127A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of speech recognition technology, and in particular to an updating method, device, local host, server, system and storage medium. Background Art
[0002] In today's era of booming smart home, the offline voice recognition technology of local hosts has made significant progress. At present, the offline voice recognition accuracy of local hosts has generally reached more than 95%, which has brought users an extremely convenient user experience. In the smart home scenario, users can control various home devices through simple voice commands with the help of smart speakers, smart home appliances and other devices. The entire offline voice control interaction process covers several key links: First, the device performs offline voice-to-text processing on the voice commands issued by the user, and converts the voice information into semantic request information in text form; then, the offline semantic recognition model is used to analyze the converted semantic request information to obtain the user's intention; then, the corresponding control instructions are generated according to the recognized intention and sent to the corresponding device; finally, the device feedbacks the execution of the instructions to the user in the form of audio broadcast.
[0003] However, despite the excellent performance of offline speech recognition technology in terms of accuracy, it faces severe challenges in the iteration of offline semantic recognition models. In the smart home environment, the update iteration of offline semantic recognition models cannot respond quickly. Usually, the model needs to undergo a complex training process in the cloud. After the training is completed, the latest model can be deployed to the user's local host through the model OTA (Over-The-Air) technology or waiting for a major version of the system to be updated. This process is often time-consuming, resulting in some entries that users are accustomed to using in daily life. Even if they have high rationality and frequency of use in actual applications, the offline semantic recognition model cannot include them in the recognition scope for optimization and iteration for a long time. This situation seriously affects the user experience, causing users to often encounter problems with voice commands that cannot be accurately recognized or understood when using the offline voice control function, hindering the further development of smart home systems in a more intelligent and humanized direction. Summary of the invention
[0004] In view of this, in order to solve the technical problem in the prior art that the model iteration of the offline semantic recognition model is not timely resulting in a poor user experience, the present disclosure provides an updating method, device, local host, server, system and storage medium for an offline speech recognition module.
[0005] According to a first aspect of an embodiment of the present disclosure, a method for updating an offline speech recognition module is provided, which is applied to a local host, wherein the offline speech recognition module includes semantic template configuration information and an offline semantic recognition model, and the updating method includes:
[0006] Performing semantic recognition on the semantic request information to be recognized based on the semantic template configuration information of the offline speech recognition module and the offline semantic recognition model;
[0007] If the offline speech recognition module fails in semantic recognition of the semantic request information, recognition failure record information corresponding to the semantic request information is generated; wherein the recognition failure record information at least includes the semantic request information;
[0008] Transmitting a set of record information consisting of a plurality of the recognition failure record information to a server, so that an online semantic recognition model on the server performs semantic recognition on the semantic request information in the recognition failure record information to obtain a semantic recognition result;
[0009] Based on the semantic recognition result received from the server that meets the set conditions and the semantic request information corresponding to the semantic recognition result, the semantic template configuration information of the offline speech recognition module is updated so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the semantic recognition result corresponding to it.
[0010] In an alternative embodiment,
[0011] Before transmitting the record information set consisting of the plurality of recognition failure record information to the server, the updating method comprises:
[0012] When the local host and the server are in a communication connection state, the record information set is transmitted to the server at intervals of a first time length.
[0013] In an alternative embodiment,
[0014] Before transmitting the record information set consisting of the plurality of recognition failure record information to the server, the updating method comprises:
[0015] When the local host and the server are in a non-communication connection state, outputting a prompt message at intervals of a second time length; wherein the prompt message is used to remind the local host to establish a communication connection with the server; and / or,
[0016] When the local host and the server are in a non-communication connection state, if the number of recognition failure record information in the record information set reaches a first set number threshold, the prompt information is output.
[0017] In an alternative embodiment,
[0018] The recognition failure record information at least includes the request frequency of the semantic request information for speech recognition failure, and the updating method includes:
[0019] When the local host and the server are in a non-communication connection state, if it is determined that the duration of the non-communication connection state reaches a third time interval, and / or if it is determined that the number of identification failure record information in the record information set reaches a second set number threshold, the record information set is updated based on a least recently used algorithm.
[0020] In an alternative embodiment,
[0021] The updating method comprises:
[0022] If update information for updating the offline semantic recognition model is received, the offline semantic recognition model is updated based on the update information, and newly recognizable semantic request information of the updated offline semantic recognition model is determined based on the update information;
[0023] The speech recognition request information that is the same as the recognizable semantic request information and the semantic recognition result corresponding to the semantic recognition request information are deleted in the semantic template configuration information.
[0024] In an alternative embodiment,
[0025] The performing semantic recognition on the semantic request information to be recognized based on the semantic template configuration information of the offline speech recognition module and the offline semantic recognition model includes:
[0026] Performing semantic recognition on the semantic request information based on the semantic template configuration information;
[0027] If the target semantic recognition result corresponding to the semantic request information is not determined from the semantic template configuration information, semantic recognition is performed on the semantic request information based on the offline semantic recognition model in the offline speech recognition module.
[0028] In an alternative embodiment,
[0029] The semantic recognition result that meets the set condition refers to a control type and / or query type instruction, and the object of the instruction is the device controlled by the local host.
[0030] According to a second aspect of an embodiment of the present disclosure, there is provided a method for updating an offline speech recognition module, which is applied to a server in the updating method according to any one of claims 1 to 7, wherein the offline speech recognition module includes semantic template configuration information and an offline semantic recognition model, and the updating method includes:
[0031] Receive a record information set; wherein the record information set includes a plurality of speech recognition failure record information;
[0032] Performing semantic recognition on the semantic request information in the recognition failure record information based on an online semantic recognition model to obtain a semantic recognition result;
[0033] Determining the semantic recognition result that meets the set condition from the semantic recognition results corresponding to the semantic request information of the recognition failure record information;
[0034] The semantic recognition result that meets the set conditions and the semantic request information corresponding to the semantic recognition result are transmitted to the local host to update the semantic template configuration information of the offline speech recognition module, so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the corresponding semantic recognition result.
[0035] In an alternative embodiment,
[0036] The updating method comprises:
[0037] If it is determined that the offline semantic recognition model has been updated, the update information of the offline semantic recognition model is transmitted to the local host, so that the local host updates the offline semantic recognition model configured on the local host based on the update information, and determines the newly added recognizable semantic request information of the updated offline semantic recognition model based on the update information, so as to delete the speech recognition request information that is identical to the recognizable semantic request information and the semantic recognition result corresponding to the semantic recognition request information in the semantic template configuration information configured on the local host.
[0038] In an alternative embodiment,
[0039] The semantic recognition result that meets the set condition refers to a control type and / or query type instruction, and the object of the instruction is the device controlled by the local host.
[0040] According to a third aspect of an embodiment of the present disclosure, there is provided an updating device for an offline speech recognition module, which is applied to a local host. The updating device comprises an offline speech recognition module, a first information transmission module and an updating module, wherein:
[0041] The offline speech recognition module is used to perform semantic recognition on the semantic request information to be recognized;
[0042] The offline speech recognition module is further configured to generate recognition failure record information corresponding to the semantic request information if the semantic recognition performed on the semantic request information fails; wherein the recognition failure record information at least includes the semantic request information;
[0043] The first information transmission module is used to transmit a record information set consisting of a plurality of recognition failure record information to a server, so that an online semantic recognition model on the server performs semantic recognition on the semantic request information in the recognition failure record information to obtain a semantic recognition result;
[0044] The update module is used to update the semantic template configuration information of the offline speech recognition module based on the semantic recognition result that meets the set conditions and is received from the server, and the semantic request information corresponding to the semantic recognition result, so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the semantic recognition result corresponding to it.
[0045] According to a fourth aspect of an embodiment of the present disclosure, there is provided an updating device for an offline speech recognition module, which is applied to the server, and the updating device comprises a second information transmission module, a semantic recognition module and a determination module, wherein:
[0046] The second information transmission module is used to receive a record information set; wherein the record information set includes a plurality of speech recognition failure record information;
[0047] The semantic recognition module is used to perform semantic recognition on the semantic request information in the recognition failure record information based on the online semantic recognition model to obtain a semantic recognition result;
[0048] The determination module is used to determine the semantic recognition result that meets the set conditions from the semantic recognition results corresponding to the semantic request information of the recognition failure record information;
[0049] The second information transmission module is also used to transmit the semantic recognition result that meets the set conditions and the semantic request information corresponding to the semantic recognition result to the local host, so as to update the semantic template configuration information of the offline speech recognition module, so that the semantic template configuration information increases the correspondence between the semantic request information and the corresponding semantic recognition result.
[0050] According to a fifth aspect of an embodiment of the present disclosure, a local host is provided, the local host comprising:
[0051] a first processor;
[0052] a first memory for storing instructions executable by the first processor;
[0053] Among them, the first processor is configured to execute the updating method of the offline speech recognition module as described in any one of the first aspects.
[0054] According to a fifth aspect of an embodiment of the present disclosure, a server is provided, the server comprising:
[0055] A second processor;
[0056] a second memory for storing instructions executable by the second processor;
[0057] Among them, the second processor is configured to execute the updating method of the offline speech recognition module as described in any one of the second aspects.
[0058] According to a sixth aspect of an embodiment of the present disclosure, a system for updating an offline speech recognition module is provided, the updating system comprising the local host as described in the fourth aspect, and the server as described in the fifth aspect.
[0059] According to a seventh aspect of the embodiments of the present disclosure, a non-transitory computer-readable storage medium is provided.
[0060] When the instructions in the storage medium are executed by the first processor of the local host, the local host is enabled to execute the method for updating the offline speech recognition module as described in any one of the first aspects;
[0061] and / or,
[0062] When the instructions in the storage medium are executed by the second processor of the server, the server is enabled to execute the method for updating the offline speech recognition module as described in any one of the second aspects.
[0063] The technical solution provided by the embodiment of the present disclosure may include the following beneficial effects: the offline speech recognition module in the present disclosure not only includes an offline semantic recognition model, but also adds semantic template configuration information. When the offline speech recognition module fails to perform semantic recognition on the semantic request information to be recognized, recognition failure record information corresponding to the above semantic request information can be formed. With the failure of semantic recognition of different semantic request information, a record information set consisting of multiple semantic recognition failure information can be obtained. The local host can transmit the above record information set to the server, and then the online semantic recognition model of the server performs semantic recognition on the semantic request information in the record information set to obtain the corresponding semantic recognition result. Then, the semantic recognition result that meets the set conditions and the corresponding semantic request information are transmitted to the local host. After the local host receives the above information, the above information can be used to update the semantic template configuration information, that is, the semantic request information received by the local host and the corresponding semantic recognition result are added to the semantic template configuration information. When the above semantic request information is encountered again in the future, the corresponding semantic recognition result can be determined based on the semantic template configuration information, thereby realizing the operation required by the user. Without the need to frequently update the offline semantic recognition model, the present invention can achieve accurate recognition and response to new commands in a short time by updating the semantic template configuration information, thereby ensuring that the offline speech recognition module can continuously adapt to the diverse language habits and personalized needs of different users, and can also well meet the user's voice control needs, thereby improving the user experience of the offline voice control function and enhancing practicality and flexibility.
[0064] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0066] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0067] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0068] Figure 1 The figure is a flowchart of a method for updating an offline speech recognition module according to an exemplary embodiment (applied to a local host).
[0069] Figure 2 is a flowchart of a method for updating an offline speech recognition module according to another exemplary embodiment (applied to a server).
[0070] Figure 3 is a flowchart of a method for updating an offline speech recognition module according to another exemplary embodiment (application updating system).
[0071] Figure 4 is a block diagram of an updating device for an offline speech recognition module according to an exemplary embodiment (applied to a local host).
[0072] Figure 5 is a block diagram of a device for updating an offline speech recognition module according to another exemplary embodiment (applied to a server).
[0073] Figure 6 is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0074] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0075] The disclosure below provides many different embodiments or examples to implement different schemes of the present invention. In order to simplify the disclosure of the present invention, the components and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present invention. In addition, the present invention can repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0076] For ease of description, spatial relative terms may be used herein to describe the relative positional relationship or movement of one element or feature relative to another element or feature as shown in the figure, such as "inside", "outside", "inner side", "outer side", "below", "below", "above", "above", "front", "back", etc. Such spatial relative terms are intended to include different orientations of the device in use or operation in addition to the orientation depicted in the figure. For example, if the device in the figure undergoes a position flip or a posture change or a motion state change, then these directional indications also change accordingly, for example: an element described as "below other elements or features" or "below other elements or features" will subsequently be oriented as "above other elements or features" or "above other elements or features". Therefore, the example term "below..." may include both upper and lower orientations. The device may be otherwise oriented (rotated 90 degrees or in other directions) and the spatial relative descriptors used herein are interpreted accordingly.
[0077] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.
[0078] The following will describe the implementation methods of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, not for limiting the scope of protection of the present application.
[0079] In order to solve the technical problem in the prior art that untimely model iteration of an offline semantic recognition model leads to poor user experience, the present disclosure provides an updating method, device, local host, server, system and storage medium for an offline speech recognition module.
[0080] Among them, the offline speech recognition module in the present disclosure not only includes an offline semantic recognition model, but also adds semantic template configuration information. When the offline speech recognition module fails to perform semantic recognition on the semantic request information to be recognized, the recognition failure record information corresponding to the above semantic request information can be formed. With the failure of semantic recognition of different semantic request information, a record information set consisting of multiple semantic recognition failure information can be obtained. The local host can transmit the above record information set to the server, and then the online semantic recognition model of the server performs semantic recognition on the semantic request information in the record information set to obtain the corresponding semantic recognition result. Then, the semantic recognition result that meets the set conditions and its corresponding semantic request information are transmitted to the local host. After the local host receives the above information, it can use the above information to update the semantic template configuration information, that is, the semantic request information received by the local host and the corresponding semantic recognition result are added to the semantic template configuration information. When the above semantic request information is encountered again in the future, the corresponding semantic recognition result can be determined based on the semantic template configuration information, thereby realizing the operation required by the user. Without the need to frequently update the offline semantic recognition model, the present invention can achieve accurate recognition and response to new commands in a short time by updating the semantic template configuration information, thereby ensuring that the offline speech recognition module can continuously adapt to the diverse language habits and personalized needs of different users, and can also well meet the user's voice control needs, thereby improving the user experience of the offline voice control function and enhancing practicality and flexibility.
[0081] In an exemplary embodiment, a method for updating an offline speech recognition module is provided, which is applied to a local host. The local host may be configured with an offline speech recognition module, which may include semantic template configuration information and an offline semantic recognition model. Figure 1 and Figure 3 As shown, the updating method may include:
[0082] S110, performing semantic recognition on the semantic request information to be recognized based on the semantic template configuration information of the offline speech recognition module and the offline semantic recognition model;
[0083] S120: if the offline speech recognition module fails in semantic recognition of the semantic request information, then generating recognition failure record information corresponding to the semantic request information; wherein the recognition failure record information at least includes the semantic request information;
[0084] S130, transmitting a set of record information consisting of a plurality of recognition failure record information to a server, so that an online semantic recognition model on the server performs semantic recognition on the semantic request information in the recognition failure record information to obtain a semantic recognition result;
[0085] S140. Based on the semantic recognition result received from the server that meets the set conditions and the semantic request information corresponding to the semantic recognition result, update the semantic template configuration information of the offline speech recognition module so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the semantic recognition result corresponding to it.
[0086] In step S110, the semantic template configuration information in the offline speech recognition module is a pre-built rule set, which can store the mapping relationship between semantic request information and its corresponding semantic recognition results, for example, "turn on the living room light" corresponds to the instruction to control the living room light to turn on. The offline semantic recognition model is a model obtained by training a large amount of semantic data.
[0087] When the semantic request information to be recognized is input into the offline speech recognition module, the semantic request information can be first semantically recognized based on the semantic template configuration information. If a semantic recognition result is obtained, subsequent control operations can be performed directly based on the obtained semantic recognition result. If the target semantic recognition result corresponding to the semantic request information is not determined from the semantic template configuration information, the semantic recognition of the semantic request information can continue to be performed based on the offline semantic recognition model. If a semantic recognition result is obtained, subsequent control operations can be performed based on the obtained semantic recognition result. If no semantic recognition result is still obtained, it means that the semantic recognition of the semantic request information by the offline speech recognition module has failed.
[0088] It should be noted that, in some implementations, after receiving the voice information input by the user, the local host may first convert the voice information into semantic request information in text form, and then use the offline voice recognition module to perform semantic recognition on the semantic request information.
[0089] Among them, the semantic recognition step based on the semantic template configuration information may include: trying to match the semantic request information to be recognized with the rules in the semantic template configuration information. For example, the semantic request information may be subjected to lexical analysis, syntactic analysis and other processing to extract key information, and then the semantic template configuration information may be searched for a command pattern that matches it. For example, if a user issues an instruction to "turn off the bedroom fan", keywords such as "turn off", "bedroom" and "fan" may be extracted, and rules containing these keywords and matching logical relationships may be searched in the semantic template configuration information. If a completely matched or highly similar rule is found, the offline speech recognition module can obtain the semantic recognition result corresponding to the above-mentioned semantic request information, that is, the instruction to control the bedroom fan to turn off, and the recognition process ends.
[0090] Among them, the process of semantic recognition based on the offline semantic recognition model may include: if the semantic recognition based on the semantic template configuration information fails, the offline speech recognition module can pass the semantic request information to the offline semantic recognition model. The model can extract features from the information, convert the semantic request information into a digital feature vector, and then perform complex operations and reasoning through the neural network structure inside the model. Based on the semantic request patterns and semantic relationships learned during the training process, the model attempts to predict the intention expressed by the semantic request information and outputs a semantic recognition result. For example, for a relatively novel or complex instruction issued by the user, "switch the air purifier on the balcony to sleep mode", it may not be matched based on the semantic template configuration information, but the offline semantic recognition model can output the semantic recognition result of "control the air purifier on the balcony to switch to sleep mode" by analyzing various features in the instruction.
[0091] It should be noted that, in addition to the above-mentioned methods for semantic recognition, other methods may also be used, which are not limited to this.
[0092] In step S120, the failure of semantic recognition of the semantic request information by the offline speech recognition module means that the semantic recognition based on the semantic template configuration information fails, and the semantic recognition based on the offline semantic recognition model also fails. For example, when semantic recognition is first performed based on the semantic template configuration information, when the semantic recognition based on the semantic template configuration information fails, and then semantic recognition is performed based on the offline semantic recognition model, if the semantic recognition based on the offline semantic recognition model fails, it can be explained that the semantic recognition of the semantic request information by the offline speech recognition module fails.
[0093] In this step, when the offline speech recognition module fails to perform semantic recognition on the semantic request information, recognition failure record information corresponding to the semantic request information can be generated. That is, the semantic request information for which the semantic recognition fails is recorded, thereby obtaining recognition failure record information.
[0094] It should be noted that, in addition to the semantic request information of the recognition failure, the recognition failure record information may also record other information according to the needs, and there is no limitation on this. For example, when the semantic recognition fails, the request frequency corresponding to the semantic request information may also be recorded synchronously, that is, the recognition failure record information may include the semantic request information, and may also include the request frequency corresponding to the above semantic request information.
[0095] In step S130, as the offline speech recognition module fails to perform semantic recognition on multiple semantic request information, more and more recognition failure record information may be recorded on the local host. When certain conditions are met, the local host may transmit a record information set consisting of multiple recognition failure record information to the server. Then, the server may use the online semantic recognition model to perform online semantic recognition on the semantic request information in the record information set, thereby obtaining the corresponding semantic recognition result.
[0096] In some embodiments,
[0097] When the local host is in a communication connection state with the server, for example, when the local host is in a networked state, the record information set can be transmitted to the server at intervals of a first time length. After the local host transmits the record information set to the server, the locally stored record information set can be deleted, and then the record of the identification failure record information can be re-recorded. In this implementation manner, the certain condition refers to that the local host is in a communication connection state with the server and the first time length interval is separated.
[0098] It should be noted that the first time interval can be set according to actual needs, and its specific value may not be limited. For example, it can be set to transmit a record information set once every 24 hours, and the time of transmission can be set to 2 a.m. Of course, it can also be set to other time intervals according to needs, and there is no limitation on this. In addition, the user can also modify the already set first time interval according to actual needs to better meet the needs of different users. In addition, the above-mentioned certain conditions can also be set to other conditions, and there is no limitation on this. For example, certain conditions may include: the local host and the server are in a communication connection state, and the number of recognition failure records in the record information set reaches half of the maximum number that can be stored.
[0099] In addition, in this embodiment, if the local host and the server are in a non-communication connection state, the local host cannot transmit the record information set to the server. In order to better ensure the transmission of the record information set, this embodiment can add a prompt mechanism to remind the local host to establish a communication connection with the server.
[0100] In some embodiments,
[0101] When the local host and the server are in a non-communication connection state, for example, when the local host is not in an online state, a prompt message is output at every second time interval. The prompt message is used to remind the user to establish a communication connection between the local host and the server. That is to say, when the local host and the server cannot communicate, the local host can output a prompt message at every second time interval to remind the user to establish a communication connection between the local host and the server to ensure that the local host can transmit the record information set to the server.
[0102] It should be noted that the second time interval can be set according to actual needs, and its specific value may not be limited. For example, it can be set to output the prompt information once every week or month. Of course, it can also be set to other time intervals according to needs, and there is no limitation on this. In addition, the user can also modify the second time interval that has been set according to actual needs to better meet different user needs.
[0103] In some embodiments,
[0104] When the local host and the server are in a non-communication connection state, for example, when the local host is not connected to the Internet, if the number of recognition failure record information in the record information set reaches a first set number threshold, a prompt message is output to prompt the user to establish a communication connection between the local host and the server to ensure that the local host can transmit the record information set to the server.
[0105] It should be noted that the first set quantity threshold can be set according to actual needs, and its specific value may not be limited. For example, the first set quantity threshold may be the maximum number of recognition failure record information that can be stored in the record information set, or it may be 80% of the above maximum number, and there is no limitation on this. In addition, the user may also modify the already set first set quantity threshold according to actual needs to better meet different user needs.
[0106] In some embodiments,
[0107] When the local host and the server are in a non-communication connection state, if the number of recognition failure record information in the record information set reaches the first set number threshold, a prompt message is output at intervals of a second time length. This prompts the user to establish a communication connection between the local host and the server to ensure that the local host can transmit the record information set to the server. This implementation can reduce the output frequency of the prompt message and better avoid disturbing the user.
[0108] It should be noted that, in addition to reminding the user to establish a communication connection between the local host and the server, other methods may also be used to remind the user, and this is not limited. In addition, the prompt information may include text information, voice information, or image information, and this is not limited.
[0109] In addition, in this embodiment, when the local host and the server are in a non-communication connection state, if it is determined that the duration of the non-communication connection state reaches a third time interval, and / or if it is determined that the number of recognition failure record information in the record information set reaches a second set number threshold. The record information set can be updated based on the least recently used algorithm (LRU (Least Recently Used) algorithm). In other words, the earlier the semantic request information occurs and the less frequently it is used, the more it is deleted from the record information set, so that the limited storage space is used to store more recent and more frequent semantic request information.
[0110] Among them, after the server receives the record information set transmitted by the local host, it can use the online semantic recognition model to perform semantic recognition on the semantic request information in the record information that failed to be recognized, so as to obtain the semantic recognition result. For example, after the local host transmits the record information set to the server, the local host can wait for the result sent back by the server through a message queue or a callback function. After receiving the record information set, the server can perform semantic recognition on the semantic request information in the record information set in turn through an online semantic recognition model (such as an online semantic large model or a large language model), so as to obtain the corresponding semantic recognition result. It should be noted that compared with the offline semantic recognition model, the online semantic recognition model can recognize more and wider semantic request information. Therefore, the online semantic model can generally obtain the semantic recognition result corresponding to the semantic request information.
[0111] In step S140, after the online semantic recognition model of the server recognizes the semantic recognition result, the semantic recognition result can be screened to obtain the semantic recognition result that meets the set conditions. The set conditions may include: first, having a clear intention, that is, an instruction belonging to the control class and / or query class; second, the device corresponding to the intent domain belongs to a device controllable by the local host, that is, the object of the instruction is a device controlled by the local host.
[0112] That is to say, after obtaining the semantic recognition results, first determine the semantic recognition results with clear intent. Then, for the semantic recognition results with clear intent, if the parsed intent domain has a corresponding device in the user's home (that is, the target device corresponding to the intent domain belongs to a device controllable by the local host), for example, the target device corresponding to the intent domain is "air conditioner" and the user's home also has an air conditioner, then the semantic recognition result is determined as a qualified semantic recognition result, and then the semantic request information and semantic recognition result can be added to this update list.
[0113] It should be noted that because the semantic request information not recognized by the offline speech recognition module may be chat, music, weather and other words that the offline speech recognition module itself does not support, even if the semantic recognition results are obtained through the online semantic recognition model analysis, there is no need to add them to the offline template configuration information, because even if the offline speech recognition module recognizes the semantic recognition results, it does not have the ability to perform this function. This part of the semantic recognition results is determined to be semantic recognition results with clear layout intentions.
[0114] In this step, after screening the semantic recognition results to be updated and their corresponding semantic request information, the server can return the semantic recognition results to be added and their corresponding semantic request information to the local host through message queues and other methods. After the local host receives the semantic recognition results that meet the set conditions and the semantic request information corresponding to the semantic recognition results, it can update the semantic template configuration information in the offline speech recognition module based on the above information, so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the corresponding semantic recognition results. In addition, the local host can simultaneously clear the recognition failure record information in the record information set.
[0115] In addition, in this embodiment, if the local host receives update information for updating the offline semantic recognition model, the local host may also update the locally deployed offline semantic recognition model based on the above update information, thereby realizing update iteration of the offline semantic recognition model.
[0116] The update information may include newly added recognizable semantic request information. That is, the local host may determine the newly added recognizable semantic request information of the updated offline semantic recognition model based on the update information, and then query the semantic request information identical to the recognizable semantic request information from the semantic template configuration information, and delete the queried semantic request information and the semantic recognition result corresponding to the above semantic request information to release the storage space of the local host.
[0117] In addition, it should be noted that the local host used in the updating method of the offline voice recognition module of this embodiment may include global voice entrances such as home host and central control, which can be equipped with sufficient memory and processor, can run voice-to-text model, semantic recognition model and voice synthesis module, etc., and can be configured with a complete voice processing link. However, the offline voice of general household appliances is limited by its own limited computing power, and its offline voice processing related model is generally for the "voiceprint: control command" recognition of its own control, which is not within the scope of this embodiment.
[0118] In this embodiment, without the need to frequently update the offline semantic recognition model, accurate recognition and response to new instructions can be achieved in a short time by updating the semantic template configuration information, thereby ensuring that the offline speech recognition module can continuously adapt to the diverse language habits and personalized needs of different users, and can also well meet the user's voice control needs, thereby improving the user experience of using the offline voice control function and enhancing practicality and flexibility.
[0119] In an exemplary embodiment, a method for updating an offline speech recognition module is provided, which is applied to a server. Figure 2 and Figure 3 As shown, the updating method may include:
[0120] S210, receiving a record information set; wherein the record information set includes a plurality of speech recognition failure record information;
[0121] S220, performing semantic recognition on the semantic request information in the recognition failure record information based on the online semantic recognition model to obtain a semantic recognition result;
[0122] S230, determining a semantic recognition result that meets a set condition from the semantic recognition results corresponding to the semantic request information of the recognition failure record information;
[0123] S240. The semantic recognition results that meet the set conditions and the semantic request information corresponding to the semantic recognition results are transmitted to the local host to update the semantic template configuration information of the offline speech recognition module, so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the corresponding semantic recognition results.
[0124] In step S210, the offline speech recognition module of the local host can perform semantic recognition on the semantic request information to be recognized, and if the semantic recognition fails, recognition failure record information corresponding to the semantic request information can be generated. That is, the semantic request information for which the semantic recognition fails is recorded, thereby obtaining recognition failure record information.
[0125] It should be noted that, in addition to the semantic request information of the recognition failure, the recognition failure record information may also record other information according to the needs, and there is no limitation on this. For example, when the semantic recognition fails, the request frequency corresponding to the semantic request information may also be recorded synchronously, that is, the recognition failure record information may include the semantic request information, and may also include the request frequency corresponding to the above semantic request information.
[0126] As the offline speech recognition module fails in semantic recognition of multiple semantic request information, more and more recognition failure record information can be recorded on the local host. When certain conditions are met, the local host can transmit a record information set consisting of multiple recognition failure record information to the server, and the server can receive the record information set.
[0127] It should be noted that the specific contents of the above-mentioned certain conditions can be referred to the description in other embodiments, which will not be elaborated here.
[0128] In step S220, after receiving the record information set, the server may perform semantic recognition on the semantic request information in the record information set based on the online semantic recognition model of the server, thereby obtaining a semantic recognition result.
[0129] For example, after receiving the record information set, the server can sequentially perform semantic recognition on the semantic request information in the record information set through an online semantic recognition model (such as an online semantic large model or a large language model), thereby obtaining corresponding semantic recognition results. It should be noted that compared with the offline semantic recognition model, the online semantic recognition model can recognize more and wider semantic request information. Therefore, the online semantic model can generally obtain the semantic recognition results corresponding to the semantic request information.
[0130] In step S230, after the online semantic recognition model of the server recognizes the semantic recognition result, the semantic recognition result can be screened to obtain the semantic recognition result that meets the set conditions. The set conditions may include: first, having a clear intention, that is, belonging to the control and / or query class instructions; second, the device corresponding to the intent domain belongs to the device controllable by the local host, that is, the object of the instruction is the device controlled by the local host.
[0131] In some embodiments,
[0132] After obtaining the semantic recognition results, first determine the semantic recognition results with clear intent. Then, for the semantic recognition results with clear intent, if the parsed intent domain has a corresponding device in the user's home (that is, the target device corresponding to the intent domain belongs to a device that can be controlled by the local host), for example, the target device corresponding to the intent domain is "air conditioner" and the user's home also has an air conditioner, then the semantic recognition result is determined as a qualified semantic recognition result, and then the semantic request information and semantic recognition result can be added to this update list.
[0133] It should be noted that, in addition to obtaining a semantic recognition result that meets the conditions in the above manner, it can also be determined in other ways, which are not limited to this.
[0134] In step S240, after the server obtains the semantic recognition result that meets the conditions and the corresponding semantic request information, the above information can be transmitted to the local host. After the local host receives the above semantic recognition result that meets the set conditions and the semantic request information corresponding to the above semantic recognition result, it can update the semantic template configuration information in the offline speech recognition module based on the above information, so that the semantic template configuration information adds the corresponding relationship between the above semantic request information and the corresponding semantic recognition result. In addition, the local host can synchronously clear the recognition failure record information in the record information set.
[0135] In addition, in this embodiment, the server can screen the semantic request information (i.e., the semantic request information in the offline template configuration information) that cannot be recognized by the offline speech recognition module according to the feedback from the local host, but can be recognized by the online semantic recognition model, and iterate the semantic request information that is used frequently and is more common to the offline semantic recognition model for training. And every longer period, such as one quarter or half a year, the update information of the offline semantic recognition model is provided to the local host for OTA (over-the-air download technology) upgrade, and the newly added semantic request information is noted in the log of the update information. After the local host receives the above update information, it can update the local offline semantic recognition model based on the above update information, and delete the newly added semantic request information of the offline semantic recognition model and the semantic recognition results corresponding to the above semantic request information from the semantic template configuration information, thereby freeing up the storage space of the local host.
[0136] It should be noted that the server may automatically screen the semantic request information or introduce relevant staff to perform manual screening, and there is no limitation on this.
[0137] In this embodiment, without the need to frequently update the offline semantic recognition model, accurate recognition and response to new instructions can be achieved in a short time by updating the semantic template configuration information, thereby ensuring that the offline speech recognition module can continuously adapt to the diverse language habits and personalized needs of different users, and can also well meet the user's voice control needs, thereby improving the user experience of using the offline voice control function and enhancing practicality and flexibility.
[0138] In an exemplary embodiment, a system for updating an offline speech recognition module is provided. Figure 3 As shown, the update system may include the server and local host in the above embodiments.
[0139] Among them, the local host generally has a certain amount of computing power, which can run a certain amount of speech-to-text models and lightweight home appliance control-related semantic recognition models, but it is generally impossible to place a large semantic understanding model or even a large voice model locally, so it can meet some commonly used expression understandings, such as "turn on the air conditioner", "turn on the air conditioner", "turn on the air conditioner", etc., but for more colloquial or user-specific expressions, it may not be possible to accurately complete semantic recognition, such as "turn on the air conditioner". In this case, if it needs to be added to the offline semantic recognition model, it is necessary to train online and then remotely upgrade the model to the local, and frequent updates cannot be achieved. However, on the offline side (that is, the local host), if a semantic template configuration information is added for matching before the semantic request information is sent to the offline semantic recognition model, a certain number of semantic request information that is not supported by the offline semantic recognition model can be processed, and the addition of semantic template configuration information is faster and more convenient, and it takes effect as soon as it is added, which is more suitable for frequent iterations.
[0140] In this embodiment, after receiving the voice request, the local host first obtains the semantic request information in text form through the speech-to-text model. Then, the target semantic request information matching the above semantic request information is searched from the semantic template configuration information. If the matching target semantic request information is found, the semantic recognition result corresponding to the queried target semantic request information is obtained and sent for execution. If the target semantic request information matching the above semantic request information is not found, the semantic recognition of the above semantic request information is performed through the offline semantic recognition model. If the offline semantic recognition model recognizes the semantic recognition result, the corresponding semantic recognition result is executed and sent. If the offline semantic recognition model fails to recognize, the semantic request information that failed to recognize the semantics and the corresponding request frequency are recorded. The above semantic request information and request frequency constitute the recognition failure record information. The recognition failure record information is stored in the record information set.
[0141] Among them, when the local host can use the network, the recorded record information set is sent to the server (such as a voice server) at every first time interval, such as 2 a.m. every day, and then the server's result is waited for through a message queue or a callback function. After receiving the record information set, the server processes the semantic request information in turn through an online semantic recognition model or a large language model. For semantic request information with clear intent resolution, if the intent domain has a corresponding device in the user's home, such as the intent domain is "air conditioning", and the user's home also has an air conditioner, then the semantic request information and its corresponding semantic recognition result are added to this update list. After screening the semantic request information to be updated and its corresponding semantic recognition results, the server can return the semantic request information to be added and its corresponding semantic recognition results to the local host through methods such as message queues. The local host can then add it to the semantic template configuration information and clear the recognition failure record information in the record information set of the local host.
[0142] In the case where the local host is not connected to the Internet, whenever the time reaches the second set time interval, such as 1 week, or the number of recognition failure record information stored in the record information set reaches the upper limit (i.e., the first set number threshold), a prompt message is pushed to the user through the local screen device or app (application). The prompt message can be, for example, a message reminding the user to connect to the Internet to update the voice function. If the user is slow to connect to the Internet for updating, the local host can update the record information set according to the LRU (least recently used) algorithm, and keep the storage at the upper limit of the record information set until the local host completes the update of the semantic template configuration information.
[0143] In addition, the maintainer of the offline semantic recognition model can also manually screen the semantic request information that is not supported locally but supported in the cloud according to user feedback. For the semantic request information that is used frequently and is more common, it can be iterated to the offline semantic recognition model for training, and a new version of the offline semantic recognition model can be released every longer period, such as one quarter or half a year, for users to upgrade via OTA (over-the-air download technology), and the newly added semantic request information can be noted in the update information, so that the local host can delete the corresponding semantic request information and its semantic recognition results in the semantic template configuration information while updating the offline semantic recognition model. In the subsequent use of the user, the offline speech recognition module will not be able to recognize the above semantic request information based on the semantic template configuration information, but can obtain the semantic recognition results corresponding to the above semantic request information through the offline semantic recognition model.
[0144] It should be noted that, in this embodiment, the semantic template configuration information initially configured in the local host may be an empty set that does not include any semantic request information and semantic recognition results, or may be a non-empty set that pre-configures the correspondence between the initial semantic request information and the semantic recognition results, and there is no limitation on this. In some embodiments, the semantic template configuration information initially configured in the local host is an empty set, and as the user uses it, the correspondence between the semantic request information and the semantic recognition results is gradually added to the semantic template configuration information.
[0145] In this embodiment, without the need to frequently update the offline semantic recognition model, accurate recognition and response to new instructions can be achieved in a short time by updating the semantic template configuration information, thereby ensuring that the offline speech recognition module can continuously adapt to the diverse language habits and personalized needs of different users, and can also well meet the user's voice control needs, thereby improving the user experience of using the offline voice control function and enhancing practicality and flexibility.
[0146] In an exemplary embodiment, an updating device for an offline speech recognition module is provided, which is applied to a local host. The updating device is used to implement the above-mentioned updating method applied to the local host. Figure 4 As shown, the device may include an offline speech recognition module 10 , a first information transmission module 20 and an updating module 30 .
[0147] An offline speech recognition module 10 is used to perform semantic recognition on the semantic request information to be recognized;
[0148] The offline speech recognition module 10 is further configured to generate recognition failure record information corresponding to the semantic request information if the semantic recognition of the semantic request information fails; wherein the recognition failure record information at least includes the semantic request information;
[0149] The first information transmission module 20 is used to transmit a set of record information consisting of a plurality of the recognition failure record information to a server, so that an online semantic recognition model on the server performs semantic recognition on the semantic request information in the recognition failure record information to obtain a semantic recognition result;
[0150] The update module 30 is used to update the semantic template configuration information of the offline speech recognition module based on the semantic recognition result that meets the set conditions and is received from the server, and the semantic request information corresponding to the semantic recognition result, so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the semantic recognition result corresponding to it.
[0151] In an exemplary embodiment, a device for updating an offline speech recognition module is provided, which is applied to a local host. Figure 4 As shown, in the updating device, the first information transmission module 20 can be used to:
[0152] Before transmitting a record information set consisting of a plurality of the identification failure record information to the server, the local host and the server are in a communication connection state, and the record information set is transmitted to the server at intervals of a first time length.
[0153] In an exemplary embodiment, a device for updating an offline speech recognition module is provided, which is applied to a local host. Figure 4 As shown, in the updating device, the first information transmission module 20 can be used to:
[0154] Before transmitting a record information set consisting of a plurality of the identification failure record information to the server, when the local host and the server are in a non-communication connection state, a prompt message is output at every second time interval; wherein the prompt message is used to remind the local host to establish a communication connection with the server; and / or, when the local host and the server are in a non-communication connection state, if the number of the identification failure record information in the record information set reaches a first set number threshold, the prompt message is output.
[0155] In an exemplary embodiment, a device for updating an offline speech recognition module is provided, which is applied to a local host. Figure 4 As shown, in the updating device, the updating module 30 can be used to:
[0156] When the local host and the server are in a non-communication connection state, if it is determined that the duration of the non-communication connection state reaches a third time interval, and / or if it is determined that the number of identification failure record information in the record information set reaches a second set number threshold, the record information set is updated based on a least recently used algorithm.
[0157] In an exemplary embodiment, a device for updating an offline speech recognition module is provided, which is applied to a local host. Figure 4 As shown, in the updating device, the updating module 30 can be used to:
[0158] If update information for updating the offline semantic recognition model is received, the offline semantic recognition model is updated based on the update information, and newly recognizable semantic request information of the updated offline semantic recognition model is determined based on the update information;
[0159] The speech recognition request information that is the same as the recognizable semantic request information and the semantic recognition result corresponding to the semantic recognition request information are deleted in the semantic template configuration information.
[0160] In an exemplary embodiment, a device for updating an offline speech recognition module is provided, which is applied to a local host. Figure 4 As shown, in the updating device, the offline speech recognition module 10 can be used to:
[0161] Performing semantic recognition on the semantic request information based on the semantic template configuration information;
[0162] If the target semantic recognition result corresponding to the semantic request information is not determined from the semantic template configuration information, semantic recognition is performed on the semantic request information based on the offline semantic recognition model in the offline speech recognition module.
[0163] In an exemplary embodiment, an offline speech recognition module updating device is provided, which is applied to a local host. In the updating device, the semantic recognition result that meets the set condition refers to a control type and / or query type instruction, and the object of the instruction is a device controlled by the local host.
[0164] In an exemplary embodiment, an updating device for an offline speech recognition module is provided, which is applied to a server. The updating device can be used to implement the updating method applied to the server in the above embodiment. Figure 5 As shown, the device may include a second information transmission module 40 , a semantic recognition module 50 and a determination module 60 .
[0165] The second information transmission module 40 is used to receive a set of record information; wherein the set of record information includes a plurality of speech recognition failure record information;
[0166] The semantic recognition module 50 is used to perform semantic recognition on the semantic request information in the recognition failure record information based on an online semantic recognition model to obtain a semantic recognition result;
[0167] The determination module 60 is used to determine the semantic recognition result that meets the set conditions from the semantic recognition results corresponding to the semantic request information of the recognition failure record information;
[0168] The second information transmission module 40 is also used to transmit the semantic recognition result that meets the set conditions and the semantic request information corresponding to the semantic recognition result to the local host, so as to update the semantic template configuration information of the offline speech recognition module, so that the semantic template configuration information increases the correspondence between the semantic request information and the corresponding semantic recognition result.
[0169] In an exemplary embodiment, a device for updating an offline speech recognition module is provided, which is applied to a server. Figure 5 As shown, in the updating device, the second information transmission module 40 can be used to:
[0170] If it is determined that the offline semantic recognition model has been updated, the update information of the offline semantic recognition model is transmitted to the local host, so that the local host updates the offline semantic recognition model configured on the local host based on the update information, and determines the newly added recognizable semantic request information of the updated offline semantic recognition model based on the update information, so as to delete the speech recognition request information that is identical to the recognizable semantic request information and the semantic recognition result corresponding to the semantic recognition request information in the semantic template configuration information configured on the local host.
[0171] In an exemplary embodiment, an offline speech recognition module updating device is provided, which is applied to a server. In the updating device, the semantic recognition result that meets the set condition refers to a control type and / or query type instruction, and the object of the instruction is a device controlled by the local host.
[0172] In an exemplary embodiment, an electronic device is provided. The electronic device may be the local host in the above embodiment, or may be the server in the above embodiment, which is not limited.
[0173] Among them, reference Figure 6 As shown, the electronic device 100 may also include: at least one processor 101, a memory 102, at least one network interface 104 and other user interfaces 103. The various components in the electronic device 100 are coupled together through a bus system 105. It can be understood that the bus system 105 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 105 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, all buses are marked as the bus system 105.
[0174] The user interface 103 may include a display, a keyboard, or a pointing electronic device (eg, a mouse, a trackball, a touch pad, or a touch screen).
[0175] It can be understood that the memory 102 in the embodiment of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM). The memory 102 described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0176] In some implementations, the memory 102 stores the following elements, executable units or data structures, or a subset thereof, or an extended set thereof: an operating system 1021 and application programs 1022 .
[0177] Among them, the operating system 1021 includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application 1022 includes various application programs, such as a media player (Media Player), a browser (Browser), etc., which are used to implement various application services. The program that implements the method of the embodiment of the present application can be included in the application 1022.
[0178] In the embodiment of the present application, by calling the program or instruction stored in the memory 102, specifically, the program or instruction stored in the application 1022, the processor 101 is used to execute the method provided by each method embodiment.
[0179] The method disclosed in the above embodiment of the present application can be applied to the processor 101, or implemented by the processor 101. The processor 101 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 101 or the instruction in the form of software. The above processor 101 can be a general processor, a digital signal processor (Digital Signal Processor, DSP), an application specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software units in the decoding processor can be executed. The software unit can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 102, and the processor 101 reads the information in the memory 102 and completes the above method in combination with its hardware.
[0180] It is understood that the embodiments described herein may be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or at least one application specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing electronic device (DSPDevice, DSPD), programmable logic electronic device (PLD), field programmable gate array (FPGA), general purpose processor, controller, microcontroller, microprocessor, other electronic unit for performing the functions described in the present application, or a combination thereof.
[0181] For software implementation, the technology described herein can be implemented by a unit that performs the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or outside the processor.
[0182] It should be noted that when the electronic device is a local host, its processor may be recorded as a first processor, and the memory may be recorded as a first memory, which is used to store the executable instructions of the first processor. In this case, the first processor is configured to execute the update method of the offline speech recognition module applied to the local host in the above embodiment. When the electronic device is a server, its processor may be recorded as a second processor, and the memory may be recorded as a second memory, which is used to store the executable instructions of the second processor. In this case, the second processor is configured to execute the update method of the offline speech recognition module applied to the server in the above embodiment.
[0183] The embodiment of the present application also provides a storage medium (computer-readable storage medium). The storage medium here stores one or at least one program. The storage medium may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a read-only memory, a flash memory, a hard disk or a solid-state drive; the memory may also include a combination of the above-mentioned types of memory.
[0184] When one or at least one program in the storage medium can be executed by one or at least one processor. When the storage medium is applied to an electronic device, the above-mentioned method of execution in the electronic device can be implemented. The processor is used to execute the control program of the electronic device stored in the memory to implement the above-mentioned method of execution in the electronic device.
[0185] It should be noted that when the instructions in the storage medium are executed by the first processor of the local host, the local host can execute the corresponding update method of the offline speech recognition module applied to the local host. When the instructions in the storage medium are executed by the second processor of the server, the server can execute the update method of the offline speech recognition module applied to the server in the above embodiment.
[0186] The professionals should also be further aware that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0187] It should be noted that the phrases "one implementation", "an embodiment", "an exemplary embodiment", "some embodiments", etc. mentioned in the specification indicate that the described embodiments may include certain features, structures or characteristics, but not every embodiment may include the certain features, structures or characteristics. In addition, such phrases do not necessarily refer to the same embodiment. In addition, when describing certain features, structures or characteristics in conjunction with an embodiment, it is within the knowledge of those skilled in the art to implement such features, structures or characteristics in conjunction with other embodiments, whether explicitly or not explicitly described.
[0188] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or electronic device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or electronic device. In the absence of more limitations, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or electronic device including the elements.
[0189] The above embodiments are only preferred embodiments for fully illustrating the present application, and the protection scope of the present application is not limited thereto. Any equivalent substitution or change made by a person skilled in the art based on the present application is within the protection scope of the present application.
Claims
1. An updating method of offline voice recognition module, applied to a local host, characterized in that The offline speech recognition module includes semantic template configuration information and an offline semantic recognition model, and the updating method includes: Based on the semantic template configuration information of the offline speech recognition module and the offline semantic recognition model perform semantic recognition semantic request information to be recognized; If the offline speech recognition module fails in semantic recognition of the semantic request information, recognition failure record information corresponding to the semantic request information is generated; wherein the recognition failure record information at least includes the semantic request information; Transmitting a set of record information consisting of a plurality of the recognition failure record information to a server, so that an online semantic recognition model on the server performs semantic recognition on the semantic request information in the recognition failure record information to obtain a semantic recognition result; Based on the semantic recognition result received from the server that meets the set conditions and the semantic request information corresponding to the semantic recognition result, the semantic template configuration information of the offline speech recognition module is updated so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the semantic recognition result corresponding to it.
2. The updating method according to claim 1, characterized in that Before transmitting the set of record information composed of the plurality of the identification failure record information to the server, the updating method includes: When the local host is in a communication connection state with the server, the set of recorded information is transmitted to the server at a first interval of time.
3. The updating method according to claim 1, characterized in that Before transmitting the set of record information composed of the plurality of the identification failure record information to the server, the updating method includes: When the local host and the server are in a non-communication connection state, outputting a prompt message at intervals of a second time length; wherein the prompt message is used to remind the local host to establish a communication connection with the server; and / or, When the local host and the server are in a non-communication connection state, if the number of recognition failure record information in the record information set reaches a first set number threshold, the prompt information is output.
4. The updating method according to claim 3, characterized in that The recognition failure record information includes at least the request frequency of the semantic request information for speech recognition failure, and the updating method includes: When the local host and the server are in a non-communication connection state, if it is determined that the duration of the non-communication connection state reaches a third time interval, and / or if it is determined that the number of identification failure record information in the record information set reaches a second set number threshold, the record information set is updated based on a least recently used algorithm.
5. The updating method according to claim 1, characterized in that The updating method comprises: If update information for updating the offline semantic recognition model is received, the offline semantic recognition model is updated based on the update information, and newly recognizable semantic request information of the updated offline semantic recognition model is determined based on the update information; Among the semantic template configuration information, the speech recognition request information, which is the same as the recognizable semantic request information and the semantic recognition result corresponding to the semantic recognition request information are deleted.
6. The updating method according to claim 1, characterized in that The semantic template configuration information based on the offline speech recognition module and the offline semantic recognition model perform semantic recognition semantic request information to be recognized, including: Semantic recognition of the semantic request information is performed based on the semantic template configuration information; If the target semantic recognition result corresponding to the semantic request information is not determined from the semantic template configuration information, semantic recognition is performed on the semantic request information based on the offline semantic recognition model in the offline speech recognition module.
7. The updating method according to any one of claims 1-6, characterized in that The semantic recognition result that meets the setting conditions refers to an instruction of the control type and / or query type, and the object of the instruction is a device controlled by the local host.
8. A method for updating an offline voice recognition module, applied to a server in the updating method as claimed in any one of claims 1-7, characterized in that The offline speech recognition module includes semantic template configuration information and an offline semantic recognition model, and the updating method includes: Receive a set of record information; wherein the set of record information includes a plurality of speech recognition failure record information; Based on the online semantic recognition model, semantic recognition information is performed on the semantic request information in the recognition failure record information to obtain semantic recognition results; From the semantic recognition results corresponding to the semantic request information of the identification failure record information, the semantic recognition result that meets the setting conditions is determined; The semantic recognition result that meets the set conditions and the semantic request information corresponding to the semantic recognition result are transmitted to the local host to update the semantic template configuration information of the offline speech recognition module, so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the corresponding semantic recognition result.
9. The updating method according to claim 8, characterized in that The updating method comprises: If it is determined that the offline semantic recognition model has been updated, the update information of the offline semantic recognition model is transmitted to the local host, so that the local host updates the offline semantic recognition model configured on the local host based on the update information, and determines the newly added recognizable semantic request information of the updated offline semantic recognition model based on the update information, so as to delete the speech recognition request information that is identical to the recognizable semantic request information and the semantic recognition result corresponding to the semantic recognition request information in the semantic template configuration information configured on the local host.
10. The method of updating according to claim 8 or 9, characterized in that The semantic recognition result that meets the setting conditions refers to an instruction of the control type and / or query type, and the object of the instruction is a device controlled by the local host.
11. An updating device for an offline voice recognition module, applied to a local host, characterized in that The updating device includes an offline voice recognition module, a first information transmission module and an updating module, wherein The offline speech recognition module is used to semantic recognition of semantic request information to be recognized; The offline speech recognition module is further configured to generate recognition failure record information corresponding to the semantic request information if the semantic recognition performed on the semantic request information fails; wherein the recognition failure record information at least includes the semantic request information; The first information transmission module is used to transmit a record information set consisting of a plurality of recognition failure record information to a server, so that an online semantic recognition model on the server performs semantic recognition on the semantic request information in the recognition failure record information to obtain a semantic recognition result; The update module is used to update the semantic template configuration information of the offline speech recognition module based on the semantic recognition result that meets the set conditions and is received from the server, and the semantic request information corresponding to the semantic recognition result, so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the semantic recognition result corresponding to it.
12. An updating device for an offline voice recognition module, applied to a server, characterized in that The updating device includes a second information transmission module, a semantic recognition module and a determination module, wherein The second information transmission module is used to receive a set of record information; wherein the set of record information includes a plurality of speech recognition failure record information; The semantic recognition module is used to perform semantic recognition of the semantic request information in the recognition failure record information based on the online semantic recognition model to obtain semantic recognition results; The determination module is configured to determine the semantic recognition result that meets the setting conditions from the semantic recognition result corresponding to the semantic request information of the identification failure record information; The second information transmission module is also used to transmit the semantic recognition result that meets the set conditions and the semantic request information corresponding to the semantic recognition result to the local host, so as to update the semantic template configuration information of the offline speech recognition module, so that the semantic template configuration information adds the corresponding relationship between the semantic request information and the corresponding semantic recognition result.
13. A local host, characterized in that: The local host includes: a first processor; a first memory for storing the first processor executable instructions; wherein the first processor is configured to perform an updating method of the offline voice recognition module as claimed in any one of claims 1-7.
14. A server, characterized in that: The server comprises: A second processor; a second memory for storing the second processor executable instructions; wherein the second processor is configured to perform an updating method of the offline voice recognition module as claimed in any one of claims 8-10.
15. An update system for offline voice recognition module, characterized in that The update system comprises a local host as claimed in claim 13 and a server as claimed in claim 14.
16. A non-temporary computer readable storage medium characterized in that When the instructions in the storage medium are executed by the first processor of the local host, the local host is enabled to perform the updating method of the offline voice recognition module as claimed in any one of claims 1-7; and / or, When the instructions in the storage medium are executed by the second processor of the server, the server is enabled to perform the updating method of the offline voice recognition module as claimed in any of claims 8-10.
Citation Information
Patent Citations
Connection reminding method and mobile terminal
CN102111502A
Voice control error reporting method, electrical appliance and computer readable storage medium
CN110364155A
Language offline recognition method, terminal and readable storage medium
CN110992937A
Updating method and device of off-line voice recognition library and voice recognition method and system
CN114610727A
Communication method and device between client and server and electronic equipment
CN116614485A
Cited By
Control method of intelligent robot with body based on layered hybrid model
CN120748403A
Control method of embodied intelligent robot based on hierarchical hybrid model
CN120748403B
Instruction identification method, apparatus and device, and computer readable medium
CN120977303A