Offline intelligent voice control method and device, electronic equipment and storage medium
Through the improved twin label auxiliary module connection and fine-tuning training, an offline speech recognition model suitable for mobile terminals is generated, which solves the security and portability issues of mobile voice control and improves recognition accuracy and model applicability.
Patent Information
- Application Number
- CN202411739890.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing intelligent voice control devices cannot run on mobile devices, and have problems with poor security and portability. In addition, existing voice recognition models cannot take into account both model parameter size and recognition accuracy.
By connecting the outputs of multiple preset speech recognition models using a twin label auxiliary module improved for speech recognition tasks, performing fine-tuning training, and combining preprocessing and data enhancement of imperative speech-text samples, an offline speech recognition model suitable for mobile terminals is generated.
The speech recognition accuracy is improved, making the model suitable for intelligent voice control tasks, while reducing the number of model parameters, achieving the security and portability of running on portable devices and without the need for an Internet connection.
Smart Images

Figure CN119580729B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of voice control technology, and more specifically, relates to an offline intelligent voice control method, device, electronic device and storage medium. Background Art
[0002] At present, intelligent voice control has made great progress, but most of the current popular voice control devices are based on large models and cannot run on mobile terminals. They need to connect to the server through the international Internet for voice recognition operations, which brings possible privacy leaks and security issues, and their portability is poor. The existing voice recognition models cannot take into account both the model parameter size and recognition accuracy at the same time, and therefore cannot meet the needs of mobile terminals. Summary of the Invention
[0003] In response to the defects of the existing technology, the purpose of this application is to provide an offline intelligent voice control method and device, aiming to solve the problems of poor security and portability caused by the existing technology due to its inability to run on mobile terminals and the need to connect to the Internet.
[0004] To achieve the above objectives, in a first aspect, the present application provides an offline intelligent voice control method, comprising:
[0005] The outputs of multiple preset speech recognition models are connected using a twin label auxiliary module improved for speech recognition tasks, and the connected multiple preset speech recognition models are fine-tuned to obtain a trained speech recognition model;
[0006] receiving a target speech signal, and inputting the target speech signal into the trained speech recognition model to obtain a speech recognition result output by the speech recognition model;
[0007] Based on the voice recognition result, control is initiated to the corresponding device.
[0008] Based on the existing speech recognition model, this application connects the outputs of multiple speech recognition models using a twin label auxiliary module improved for speech recognition tasks during the fine-tuning training stage, and then conducts training so that the model can obtain the outputs of other models as auxiliary information, thereby increasing the recognition accuracy of the model in speech recognition, making the model more suitable for intelligent voice control tasks, and at the same time making the model have a smaller number of model parameters, so that the model can run on smaller portable devices without connecting to the network, thereby increasing the security and portability of voice control.
[0009] According to an offline intelligent voice control method provided by the present invention, fine-tuning and training multiple preset voice recognition models connected using twin label auxiliary modules improved for voice recognition tasks include:
[0010] Obtain imperative speech-text samples;
[0011] The command speech-text samples are used to fine-tune and train multiple preset speech recognition models connected using a twin label auxiliary module improved for speech recognition tasks.
[0012] This application uses the collected imperative speech-text samples for pre-training during the fine-tuning training phase, thereby increasing the accuracy of the model in imperative speech recognition and making the model more suitable for intelligent voice control tasks.
[0013] According to an offline intelligent voice control method provided by the present invention, after obtaining the imperative voice-text sample, the method further includes:
[0014] performing noise reduction, filtering, and gain control on the speech samples in the imperative speech-text samples;
[0015] Performing word2vector processing on the text sample in the imperative speech-text sample;
[0016] A data enhancement method is used to perform data enhancement on the command speech-text sample.
[0017] This application processes voice and text samples separately, and increases the sample size through data enhancement methods to improve the accuracy of model training.
[0018] According to an offline intelligent voice control method provided by the present invention, inputting the target voice signal into the trained voice recognition model to obtain a voice recognition result output by the voice recognition model includes:
[0019] Inputting the target speech signal into the trained speech recognition model to obtain text information output by the speech recognition model;
[0020] The key information in the text information is extracted by a keyword matching algorithm as the speech recognition result.
[0021] According to an offline intelligent voice control method provided by the present invention, the method further includes:
[0022] Establishing a correspondence between preset text and the control interface and parameters of the corresponding device;
[0023] Based on the corresponding relationship, a configuration file is created.
[0024] This application configures the voice control interface by using a configuration file, which is simple and convenient, does not limit the type of controlled devices, and makes the application scenarios more flexible and varied.
[0025] According to an offline intelligent voice control method provided by the present invention, initiating control to a corresponding device based on the voice recognition result includes:
[0026] Based on the speech recognition result, obtaining a configuration file corresponding to the corresponding device;
[0027] Based on the control interface and parameters in the configuration file, control is initiated to the corresponding device.
[0028] In a second aspect, the present application provides an offline intelligent voice control device, comprising:
[0029] A training module is used to connect the outputs of multiple preset speech recognition models using a twin label auxiliary module improved for speech recognition tasks, and fine-tune the connected multiple preset speech recognition models to obtain a trained speech recognition model;
[0030] An acquisition module is used to receive a target speech signal, input the target speech signal into the trained speech recognition model, and obtain a speech recognition result output by the speech recognition model;
[0031] The control module is used to initiate control to the corresponding device based on the voice recognition result.
[0032] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the offline intelligent voice control method described in the first aspect or any possible implementation of the first aspect.
[0033] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the offline intelligent voice control method described in the first aspect or any possible implementation of the first aspect.
[0034] In a fifth aspect, the present application provides a computer program product, which, when running on a processor, enables the processor to execute the offline intelligent voice control method described in the first aspect or any possible implementation of the first aspect.
[0035] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0036] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies:
[0037] (1) Based on the existing speech recognition model, during the fine-tuning training stage, the outputs of multiple speech recognition models are connected using a twin label auxiliary module improved for speech recognition tasks, and then trained so that the model can obtain the outputs of other models as auxiliary information. This increases the recognition accuracy of the model in speech recognition, making the model more suitable for intelligent voice control tasks. At the same time, the model has a smaller number of model parameters, so that the model can run on smaller portable devices without connecting to the network, thereby increasing the security and portability of voice control.
[0038] (2) During the fine-tuning training phase, the collected imperative speech-text samples are used for pre-training, thereby increasing the accuracy of the model in imperative speech recognition and making the model more suitable for intelligent voice control tasks.
[0039] (3) The voice control interface can be configured by using a configuration file, which is simple and convenient. It does not limit the type of controlled device, making the application scenarios more flexible and varied. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0041] Figure 1 This is a flow chart of the offline intelligent voice control method provided by an embodiment of the present application;
[0042] Figure 2 This is a schematic diagram of a speech recognition fine-tuning training model provided in an embodiment of the present application, which adds a twin label auxiliary module improved for speech recognition tasks;
[0043] Figure 3 Schematic diagram of the structure of the offline intelligent voice control device provided in the embodiment of the present application;
[0044] Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0046] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.
[0047] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0048] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0049] First, let’s introduce the following contents:
[0050] Speech recognition is one of the core technologies for intelligent voice control. Thanks to advances in machine learning and deep learning, speech recognition technology has achieved significant breakthroughs. Currently, advanced speech recognition systems are capable of highly accurate speech-to-text conversion, providing a reliable foundation for intelligent voice control.
[0051] Voice signal processing is another key technology in intelligent voice control. After the user's voice signal is collected, it undergoes a series of signal processing steps, including noise reduction, filtering, and gain control, to improve its quality. These processing technologies effectively remove background noise and interference, ensuring more accurate speech recognition and semantic understanding.
[0052] Natural language understanding is a crucial step in identifying user intent in intelligent voice control. Using natural language processing technology, the system converts recognized text or commands into structured semantic representations, thereby understanding the user's intent and specific requirements. This enables intelligent voice control systems to respond more intelligently to user commands and provide more personalized services.
[0053] Automatic control technology mainly uses automatic control technology based on intelligent voice to control the controlled devices through open interfaces, including obtaining and analyzing device information through open interfaces, and sending subsequent control commands.
[0054] Next, combine Figures 1-2 The offline intelligent voice control method provided in the embodiments of the present application is introduced.
[0055] Figure 1 This is a flow chart of the offline intelligent voice control method provided by the embodiment of the present application. Figure 1 As shown, the method includes the following steps:
[0056] Step 100: Connect the outputs of multiple preset speech recognition models, and fine-tune the multiple preset speech recognition models connected using the twin label auxiliary module improved for the speech recognition task to obtain a trained speech recognition model;
[0057] Optionally, the preset speech recognition model can be various common speech recognition models, such as the VOSK open source speech recognition model, or it can be a convolutional neural network with various structures designed according to specific circumstances. This application does not limit this.
[0058] Optionally, the outputs of multiple preset speech recognition models can be connected through a twin label auxiliary module improved for speech recognition tasks. The twin label auxiliary module is used to fine-tune the preset speech recognition model. It can combine multiple independent speech recognition models and improve the overall recognition performance in a specific way. It can improve the speech recognition performance of the model without increasing the model parameters, effectively reduce the required computing resources, and improve the recognition accuracy.
[0059] Figure 2 This is a schematic diagram of a speech recognition fine-tuning training model provided by an embodiment of the present application, which adds a twin label auxiliary module improved for speech recognition tasks, such as Figure 2 As shown in the figure, MODEL1 and MODEL2 are two identical speech recognition models with the output layer removed. Their model parameters are independent of each other. The twin label auxiliary module improved for the speech recognition task is an auxiliary training module for fine-tuning. It connects the outputs of MODEL1 and MODEL2 so that the models can assist each other during the training process.
[0060] Output of MODEL1 model , the output of the MODEL2 model , x is the model input, that is, the speech sample. Unlike the ordinary twin label auxiliary module, since the output dimension of the speech model is N×T and the loss function used in fine-tuning is the CTC loss function, the twin label auxiliary module improved for speech recognition tasks improves the CTC loss. The output of the twin label auxiliary module improved for speech recognition tasks is for:
[0061]
[0062] The loss calculation process after the improved twin label auxiliary module is as follows:
[0063]
[0064]
[0065]
[0066]
[0067]
[0068] in are all the selected dynamic programming paths when calculating CTC loss, All paths One of the paths in is the tth node on the path, is the distribution probability output by the improved twin label auxiliary module at this node. The loss function used in training is sila_ctc_loss.
[0069] After training is completed, the twin label auxiliary module used for fine-tuning training is removed. Based on the actual hardware computing resources and classification accuracy requirements, a suitable speech recognition model is selected and used in the actual speech recognition task. The fine-tuned speech recognition model is obtained, which is the speech recognition model actually used.
[0070] The fine-tuned speech recognition model has smaller parameters and can be run on a central processing unit (CPU). Common mobile device configurations can meet the operating requirements, and the voice control device constructed with this model can be run on mobile portable devices.
[0071] In addition, since the model runs entirely on mobile portable devices, it does not require an Internet connection, thus meeting the offline operation requirements of the voice control system and improving security.
[0072] Step 110: receiving a target speech signal and inputting the target speech signal into a trained speech recognition model to obtain a speech recognition result output by the speech recognition model;
[0073] Optionally, the target voice signal can be received by a voice collection device such as a microphone, and the collected signal can be subjected to noise reduction, filtering, gain control, etc., and the voice quality can be improved through a series of preprocessing.
[0074] The collected speech signal is then input into the speech recognition model for speech recognition to obtain speech recognition results corresponding to the speech, such as text information.
[0075] Step 120: Initiate control to the corresponding device based on the voice recognition result.
[0076] After obtaining the voice recognition result, a control signal can be sent to the corresponding device to complete the device control process. If the corresponding control command is to obtain device status information, the returned device status information is received after sending the control command.
[0077] Optionally, a language broadcast module may be used to broadcast the status of device control information transmission, and for a control command for obtaining device status information, the device status result returned by the command also needs to be broadcast.
[0078] The present application provides an offline intelligent voice control method. Based on the existing voice recognition model, in the fine-tuning training stage, the outputs of multiple voice recognition models are connected using a twin label auxiliary module improved for voice recognition tasks, and then training is performed so that the model can obtain the outputs of other models as auxiliary information, thereby increasing the recognition accuracy of the model in voice recognition, making the model more suitable for intelligent voice control tasks, and at the same time making the model have a smaller number of model parameters, so that the model can run on smaller portable devices without connecting to the network, thereby increasing the security and portability of voice control.
[0079] In some embodiments, step 100 specifically includes:
[0080] Step 1001, obtaining a command voice-text sample;
[0081] Step 1002: Use imperative speech-text samples to fine-tune the connected multiple preset speech recognition models.
[0082] The command voice-text sample may be a voice sample with command meaning and its corresponding text sample, such as a voice sample and a text sample of “turn on device A”.
[0083] This application does not limit the method for obtaining the imperative voice-text samples, and the samples can be obtained through various methods such as the Internet.
[0084] After obtaining the imperative speech-text samples, the output layer of the preset speech recognition model is replaced with the twin label auxiliary module for further training. Since the twin label auxiliary module has the characteristics of improving model generalization and accelerating model convergence, a speech recognition model suitable for fine-tuning is constructed.
[0085] The speech recognition model with the output layer replaced is trained using the gradient backpropagation algorithm on the collected imperative speech-text samples. Training is stopped immediately after the model reaches convergence to prevent the model from overfitting on the samples.
[0086] Thanks to the use of the twin label auxiliary module, the model can converge faster and is less prone to overfitting, resulting in a speech recognition model that is more sensitive and accurate to imperative speech.
[0087] In some embodiments, after step 1001, the method further includes:
[0088] Step 10011, performing noise reduction, filtering, and gain control on the speech sample in the command speech-text sample;
[0089] Step 10012, performing word2vector processing on the text sample in the imperative speech-text sample;
[0090] Step 10013: Use a data enhancement method to perform data enhancement on the imperative speech-text sample.
[0091] Optionally, after obtaining the imperative speech-text samples, the speech can be subjected to noise reduction, filtering, and gain control, and the text can be processed using word2vector. Then, the samples can be enhanced using data augmentation methods to obtain the samples required for model fine-tuning training. By processing the speech and text samples separately and increasing the sample size through data augmentation methods, the accuracy of model training can be improved.
[0092] Optionally, this application does not limit the specific data enhancement method used.
[0093] In some embodiments, step 110 specifically includes:
[0094] Step 1101: Input the target speech signal into the trained speech recognition model to obtain text information output by the speech recognition model;
[0095] Step 1102: extract key information from the text information through a keyword matching algorithm as a speech recognition result.
[0096] After obtaining the trained speech recognition model, the target speech signal emitted by the user can be received and input into the trained speech recognition model to obtain the text information output by the speech recognition model.
[0097] Since text information contains key information and non-key information, the key information in the text information can be extracted through a keyword matching algorithm as the speech recognition result.
[0098] Optionally, the keyword matching algorithm may be a string matching algorithm, a fuzzy matching algorithm, etc., which is not limited in this application.
[0099] In some embodiments, the method further comprises:
[0100] Step 130, establishing a correspondence between the preset text and the control interface and parameters of the corresponding device;
[0101] Step 140: Create a configuration file based on the corresponding relationship.
[0102] In order not to limit the type of controlled device and increase the usage scenarios of voice control, the voice control interface can be configured by generating a configuration file. Therefore, before using the offline intelligent voice control method provided by this application, the configuration file can be configured. First, a correspondence between the preset text and the control interface and parameters of the corresponding device is established, and then the configuration information is saved in the configuration file for subsequent use.
[0103] For example, a correspondence between the text "open device A" and the command interface and parameters of device A may be established in advance, and a corresponding configuration file may be generated.
[0104] In one embodiment of the present application, a configuration file in json format is used to save the corresponding configuration information.
[0105] In some embodiments, step 120 specifically includes:
[0106] Step 1201: Based on the speech recognition result, obtain a configuration file corresponding to the corresponding device;
[0107] Step 1202: Initiate control to the corresponding device based on the control interface and parameters in the configuration file.
[0108] After obtaining the speech recognition result, that is, recognizing the text, you can use the key information retrieval configuration in the language understanding stage to obtain the corresponding configuration file and the device interface and parameters therein, and then use the obtained control interface and parameters to initiate a control signal to the corresponding device to complete the device control process.
[0109] Figure 3 This is a schematic diagram of the structure of the offline intelligent voice control device provided in the embodiment of the present application. Figure 3 As shown, the apparatus includes a training module 310, an acquisition module 210 and a control module 330, wherein:
[0110] A training module 310 is configured to connect the outputs of multiple preset speech recognition models using a twin label auxiliary module improved for speech recognition tasks, and to fine-tune the connected multiple preset speech recognition models to obtain a trained speech recognition model;
[0111] An acquisition module 320 is configured to receive a target speech signal and input the target speech signal into a trained speech recognition model to obtain a speech recognition result output by the speech recognition model;
[0112] The control module 330 is used to initiate control to the corresponding device based on the voice recognition result.
[0113] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method and will not be repeated here.
[0114] Based on the method in the above embodiment, Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, an embodiment of the present application provides an electronic device, which may include: a processor (processor) 410, a communication interface (Communications Interface) 420, a memory (memory) 430 and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 can call the logic instructions in the memory 430 to execute the offline intelligent voice control method in the above embodiment.
[0115] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the offline intelligent voice control method described in each embodiment of the present application.
[0116] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the offline intelligent voice control method in the above embodiment.
[0117] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the offline intelligent voice control method in the above embodiment.
[0118] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0119] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC.
[0120] The above embodiments can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially produce the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).
[0121] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
[0122] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. An offline intelligent voice control method, characterized in that: include: The outputs of multiple preset speech recognition models are connected using a twin label auxiliary module improved for speech recognition tasks, and the connected multiple preset speech recognition models are fine-tuned to obtain a trained speech recognition model; receiving a target speech signal, and inputting the target speech signal into the trained speech recognition model to obtain a speech recognition result output by the speech recognition model; Initiate control of the corresponding device based on the voice recognition result; The twin label auxiliary module improved for speech recognition tasks has improved the CTC loss. The output of the twin label auxiliary module improved for speech recognition tasks is for: The loss calculation process after the improved twin label auxiliary module is as follows: in are all the selected dynamic programming paths when calculating CTC loss, All paths One of the paths in is the tth node on the path, is the distribution probability output by the improved twin label auxiliary module at this node, is the loss function used in training.
2. The offline intelligent voice control method according to claim 1, characterized in that: The fine-tuning training of multiple preset speech recognition models connected by the twin label auxiliary modules improved for the speech recognition task includes: Obtain imperative speech-text samples; The command speech-text samples are used to fine-tune multiple preset speech recognition models connected using a twin label auxiliary module improved for speech recognition tasks.
3. The offline intelligent voice control method according to claim 2, characterized in that: After obtaining the command voice-text sample, the method further includes: performing noise reduction, filtering, and gain control on the speech samples in the imperative speech-text samples; Performing word2vector processing on the text sample in the imperative speech-text sample; A data enhancement method is used to perform data enhancement on the command speech-text sample.
4. The offline intelligent voice control method according to claim 1, characterized in that: Inputting the target speech signal into the trained speech recognition model to obtain a speech recognition result output by the speech recognition model includes: Inputting the target speech signal into the trained speech recognition model to obtain text information output by the speech recognition model; The key information in the text information is extracted by a keyword matching algorithm as the speech recognition result.
5. The offline intelligent voice control method according to claim 1 or 4, characterized in that: The method further comprises: Establishing a correspondence between preset text and the control interface and parameters of the corresponding device; Based on the corresponding relationship, a configuration file is created.
6. The offline intelligent voice control method according to claim 5, characterized in that: The initiating control to the corresponding device based on the voice recognition result includes: Based on the speech recognition result, obtaining a configuration file corresponding to the corresponding device; Based on the control interface and parameters in the configuration file, control is initiated to the corresponding device.
7. An offline intelligent voice control device, characterized in that: include: A training module is used to connect the outputs of multiple preset speech recognition models using a twin label auxiliary module improved for speech recognition tasks, and to fine-tune the connected multiple preset speech recognition models to obtain a trained speech recognition model; An acquisition module is used to receive a target speech signal, input the target speech signal into the trained speech recognition model, and obtain a speech recognition result output by the speech recognition model; A control module, configured to initiate control of a corresponding device based on the speech recognition result; The twin label auxiliary module improved for speech recognition tasks has improved the CTC loss. The output of the twin label auxiliary module improved for speech recognition tasks is for: The loss calculation process after the improved twin label auxiliary module is as follows: in are all the selected dynamic programming paths when calculating CTC loss, All paths One of the paths in is the tth node on the path, is the distribution probability output by the improved twin label auxiliary module at this node, is the loss function used in training.
8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the offline intelligent voice control method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program runs on a processor, the processor executes the offline intelligent voice control method according to any one of claims 1 to 6.
10. A computer program product, characterized in that When the computer program product runs on a processor, the processor executes the offline intelligent voice control method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Voice model training method, apparatus and device, and computer readable storage medium
CN114399995A
Picture classification method and system based on twin label auxiliary module, and medium
CN115115874A