Intelligent voice and video command and dispatch method and device
Through voiceprint identification and improved twin label auxiliary module training of the speech recognition model, the problem of low accuracy of intelligent voice and video command and dispatch is solved, and efficient and secure offline voice recognition and video command and dispatch are achieved.
Patent Information
- Application Number
- CN202411739904.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing intelligent voice and video command and dispatch methods have low accuracy and require connection to the international Internet for voice recognition, resulting in insufficient security and efficiency.
The voiceprint identification model is used for permission authentication, and the speech recognition MIN model trained with the improved twin label auxiliary module is used for speech recognition, enabling offline operation, improving accuracy and enhancing security.
It improves the accuracy of speech recognition, enhances the security and flexibility of the system, is suitable for use on mobile portable devices, and meets offline operation requirements.
Smart Images

Figure CN119580709B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of voice control technology, and more specifically, relates to an intelligent voice and video command and dispatch method and device. Background Art
[0002] With the advancement of technology, video surveillance has been widely used in various fields, including public safety, traffic management, and corporate security. Traditional video surveillance and command and dispatch methods rely primarily on manual operations, which are inefficient and prone to errors. Manual operations using keyboards and mice are even more inefficient when there are many source videos to dispatch and many target screens to display after dispatch.
[0003] Intelligent voice and video command and dispatch utilizes technologies such as speech recognition, natural language understanding, and speech synthesis, enabling users to conduct video command and dispatch through voice commands, thereby improving operational efficiency. However, most related intelligent voice and video command and dispatch methods require an internet connection and a server-side voice recognition function, resulting in limited recognition accuracy. Summary of the Invention
[0004] In view of the defects of the existing technology, the purpose of this application is to provide an intelligent voice and video command and dispatch method and device, aiming to solve the problem of low accuracy of intelligent voice and video command and dispatch.
[0005] To achieve the above objectives, in a first aspect, an embodiment of the present application provides an intelligent voice and video command and dispatch method, comprising:
[0006] receiving user's voice commands for video scheduling;
[0007] Input the voice command into the voiceprint identification model and output the voiceprint identification result of whether the user voiceprint corresponding to the voice command is registered;
[0008] If the user voiceprint corresponding to the voice command has been registered, the voice command is input into the speech recognition MIN model, and the speech recognition text is output; the speech recognition MIN model is trained with the assistance of the speech recognition BIG model and the improved twin label auxiliary module;
[0009] Based on the voice recognition text, video command and dispatch parameters and control interface information are determined, control instructions are generated based on the video dispatch parameters and control interface information, and the control interface is mobilized to send the control instructions to perform video command and dispatch.
[0010] In a second aspect, an embodiment of the present application further provides an intelligent voice and video command and dispatch device, comprising:
[0011] A receiving module, configured to receive a user's voice command for video scheduling;
[0012] The voiceprint identification module is used to input the voice command into the voiceprint identification model and output the voiceprint identification result of whether the user voiceprint corresponding to the voice command is registered;
[0013] The speech recognition module is used to input the speech command into the speech recognition MIN model and output speech recognition text if the user voiceprint corresponding to the speech command has been registered; the speech recognition MIN model is trained with the assistance of the speech recognition BIG model and the improved twin label auxiliary module;
[0014] The control module is used to determine the video command and dispatch parameters and control interface information based on the voice recognition text, generate control instructions based on the video dispatch parameters and control interface information, and mobilize the control interface to send control instructions to perform video command and dispatch.
[0015] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any possible implementation of the first aspect.
[0016] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0017] In a fifth aspect, an embodiment of the present application further provides a computer program product, which, when running on a processor, enables the processor to execute the method described in the first aspect or any possible implementation of the first aspect.
[0018] The intelligent voice and video command and dispatch method and device provided in the embodiments of the present application have the following beneficial effects compared with the prior art:
[0019] (1) By identifying the voiceprint of the user who issues the voice command, the authority of the user who operates the video command and dispatch can be identified, which can improve the security of the intelligent voice and video command and dispatch system.
[0020] (2) With the help of the speech recognition BIG model and the improved twin label auxiliary module for auxiliary training, the speech recognition MIN model has a high speech recognition accuracy. At the same time, the number of parameters of the speech recognition MIN model is small, which is suitable for running on the CPU on mobile portable devices. At the same time, the loss function is improved based on the improved twin label auxiliary module, making the overall model more suitable for voice command recognition of video command and dispatch. In addition, since the model can be fully run on mobile portable devices, it does not need to be connected to the international Internet, thereby meeting the offline operation requirements of the intelligent voice and video command and dispatch system and further improving security.
[0021] (3) The voice control interface can be configured by using a configuration file, which is simple and convenient. It does not limit the type of controlled device, making the application scenarios more flexible and varied. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in this application or related technologies, the following is a brief introduction to the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Figure 1 This is a flow chart of the intelligent voice and video command and dispatch method provided in an embodiment of the present application;
[0024] Figure 2 This is a flow chart of voiceprint identification based on a voiceprint identification model provided in an embodiment of the present application;
[0025] Figure 3 This is a schematic diagram of the structure of the speech recognition fine-tuning training model provided in the embodiment of the present application;
[0026] Figure 4 Schematic diagram of a process for determining video command and dispatch parameters according to an embodiment of the present application;
[0027] Figure 5 This is a schematic diagram of the structure of an intelligent voice and video command and dispatch device provided by an embodiment of the present application;
[0028] Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0030] Figure 1 This is a flow chart of the intelligent voice and video command and dispatch method provided by the embodiment of the present application, with reference to Figure 1 The method includes at least the following steps:
[0031] S101: Receive a user's voice instruction for video scheduling.
[0032] Specifically, the user's voice command for video scheduling is received through a microphone, etc. Optionally, the collected voice signal is pre-processed, including noise reduction, filtering, gain control, etc., to improve voice quality.
[0033] S102: Input the voice command into the voiceprint identification model, and output a voiceprint identification result indicating whether the user voiceprint corresponding to the voice command is registered.
[0034] Specifically, by identifying the voiceprint of the user who issues the voice command, the authority of the user who performs the video command and dispatch operation can be identified, thereby improving the security of the intelligent voice and video command and dispatch system.
[0035] Optionally, the voiceprint identification model includes a voiceprint feature extraction network, a voiceprint feature similarity scoring network, and a target voiceprint database. The voiceprint feature extraction network is used to extract voiceprint features, the target voiceprint database is used to store the voiceprint features of authorized / registered users, and the voiceprint feature similarity scoring network is used to compare the similarity between the voiceprint features of real-time speech and the voiceprint features in the target voiceprint database.
[0036] Figure 2 This is a flow chart of voiceprint identification based on the voiceprint identification model provided in the embodiment of the present application, with reference to Figure 2 Before the actual application of the voiceprint identification model, it also includes: the training process of the voiceprint identification model and the establishment process of the target person's voiceprint database.
[0037] For the training process of the voiceprint identification model, a neural network model is used to extract the feature information of the voice sample. The training goal is to ensure that the voice information sent by the same person can obtain the same feature value after the voiceprint feature extraction network extracts the features, and the scoring results given by the voiceprint feature similarity scoring network are all higher than the pre-set confidence threshold.
[0038] In the process of establishing the target voiceprint database, for users who have completed registration / authorization, the voiceprint features are extracted through the trained voiceprint feature extraction network and saved in the target voiceprint database for user authority authentication, that is:
[0039]
[0040] in, represents the target voiceprint database, User voice indicating that registration is complete, Indicates the voiceprint features extracted from the registered user's voice.
[0041] Optionally, S102 specifically includes:
[0042] Step a: extracting the voiceprint features from the received user's voice command based on the voiceprint feature extraction network. Input the received user's voice command into the voiceprint feature extraction network to extract the voiceprint features from the voice command, that is:
[0043]
[0044] in, Represents the extracted voiceprint features, Indicates the voiceprint feature extraction operation, Indicates the received user voice command.
[0045] Step b: Based on the voiceprint feature similarity scoring network, obtain the scoring result of the voiceprint feature and the target person's voiceprint database. The voiceprint feature extracted by the voiceprint feature extraction network is input into the voiceprint feature similarity scoring network to obtain the similarity between the voiceprint feature and the voiceprint feature in the target person's voiceprint database, that is:
[0046]
[0047] in, Indicates the scoring results. Represents the similarity judgment operation, Represents the voiceprint features extracted by the voiceprint feature extraction network, Represents the voiceprint features of the target person in the voiceprint database.
[0048] Step c: Compare the scoring result with a preset confidence threshold and output the voiceprint identification result. If the scoring result is higher than the preset confidence threshold, the voiceprint is considered registered and the subsequent steps are executed. If the scoring result is lower than the preset confidence threshold, the voiceprint is considered unregistered and the voice command may be discarded or the user may be notified of the unregistration through a broadcast or pop-up window.
[0049] In one possible implementation, the voiceprint feature extraction network uses an open source voiceprint identification network—Extended Context-Aware Parallel Attention TimeDelay Neural Network (ECAPA-TDNN).
[0050] S103. If the user voiceprint corresponding to the voice command has been registered, the voice command is input into the voice recognition MIN model, and the voice recognition text is output; the voice recognition MIN model is obtained through the auxiliary training of the voice recognition BIG model and the improved twin label auxiliary module.
[0051] Specifically, the voice commands identified by the voiceprint identification model are input into the speech recognition MIN model, which then outputs the speech recognition text for subsequent video command and dispatch. Speech recognition is the core of voice control. To improve the accuracy of the speech recognition MIN model, the speech recognition BIG model and the improved twin label auxiliary module are used to assist in pre-training the model.
[0052] Optionally, the speech recognition MIN model is trained based on the following steps:
[0053] Step a: Construct a speech recognition BIG model and a speech recognition MIN model with different parameter amounts and speech recognition accuracy.
[0054] The BIG model and MIN model are respectively a model with a large number of parameters and a model with a small number of parameters. The BIG model has a higher accuracy rate, while the MIN model has a lower accuracy rate.
[0055] Step b: Build an improved twin label auxiliary module to connect the outputs of the speech recognition BIG model and the speech recognition MIN model.
[0056] Figure 3 This is a structural diagram of the speech recognition fine-tuning training model provided in the embodiment of the present application, with reference to Figure 3 ,During the training process of the speech recognition MIN model, the improved twin label auxiliary module is used to connect the outputs of the speech recognition BIG model and the speech recognition MIN model.
[0057] Due to the improved use of the twin label auxiliary module, the speech recognition BIG model is used to assist the training of the speech recognition MIN model during fine-tuning training, which significantly improves the generalization and speech recognition accuracy of the speech recognition MIN model.
[0058] Step c: Based on the pre-acquired video scheduling related voice command-text samples, train the speech recognition BIG model and the speech recognition MIN model.
[0059] Optionally, for the speech recognition task of video scheduling, video scheduling related voice instruction-text samples are obtained in advance, the speech is denoised, filtered, and gain controlled, the text is converted into word2vector vectors, and training samples are obtained through data enhancement. The speech recognition BIG model and the speech recognition MIN model are trained based on the training samples.
[0060] Optionally, the input voice commands and output voice recognition text during the actual use of the speech recognition MIN model can be added to the training samples for the next fine-tuning training. This cyclic fine-tuning training can further improve the speech recognition accuracy of the speech recognition MIN model.
[0061] Step d: By improving the twin label auxiliary module, the output of the speech recognition BIG model is obtained as auxiliary information for the speech recognition MIN model to assist in training until the model converges.
[0062] The improved twin label auxiliary module connects the outputs of the speech recognition BIG model and the speech recognition MIN model. During the training process, the speech recognition BIG model transmits auxiliary information through the improved twin label auxiliary module, thereby assisting the training of the speech recognition MIN model until the model converges, thereby improving the speech recognition accuracy of the speech recognition MIN model.
[0063] Optionally, the loss function in the speech recognition MIN model training process is a CTC loss function combined with an improved twin label auxiliary module loss.
[0064] Specifically, CTC Loss (Connectionist Temporal Classification Loss) is a loss function used in sequence labeling tasks. It primarily addresses the problem of aligning input and output sequences. This is particularly true when processing continuous sequence data, where traditional methods are difficult to apply directly because the data may not be explicitly segmented. CTC Loss allows the output sequence length to be inconsistent with the input sequence length and eliminates the need for pre-alignment, enabling effective labeling of unsegmented sequence data.
[0065] Output of the speech recognition BIG model , the output of the speech recognition MIN model , Represents the model input, which is the speech sample. Unlike the ordinary twin label auxiliary module, since the output dimension of the speech model is N×T and the loss function used in fine-tuning is the CTC loss function, the improved twin label auxiliary module mainly improves the CTC loss and improves the output of the twin label auxiliary module. for:
[0066]
[0067] The loss calculation process after improving the twin label auxiliary module is as follows:
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] .
[0073] in, Represents all selected dynamic programming paths when calculating CTC loss, Indicates all paths One of the paths in Indicates the first t nodes, represents the distribution probability output by the improved twin label auxiliary module on the node, Represents the CTC loss function combined with the improved twin label auxiliary module loss.
[0074] Step e: Delete the improved twin label auxiliary module to obtain the trained speech recognition MIN model. After the model converges, remove the improved twin label auxiliary module. The trained speech recognition MIN model obtained is the speech recognition MIN model actually used in S103.
[0075] In the embodiment of the present application, the parameters of the speech recognition MIN model are relatively small, and it is suitable for running on a central processing unit (CPU) on a mobile portable device; at the same time, with the assistance of the speech recognition BIG model and the improved twin label auxiliary module for assisted training, the speech recognition MIN model has a higher speech recognition accuracy; at the same time, the loss function is improved accordingly based on the improved twin label auxiliary module, making the overall model more suitable for voice command recognition of video command and dispatch; in addition, since the model can be fully run on a mobile portable device, it does not need to be connected to the international Internet, thereby meeting the offline operation requirements of the intelligent voice and video command and dispatch system and further improving security.
[0076] In one possible implementation, the speech recognition BIG model and the speech recognition MIN model use the VOSK-BIG model and the VOSK-MIN model, respectively. The parameter size of VOSK-MIN is 42MB and can be run on a portable mobile terminal. The construction and pre-training process of the VOSK-MIN model is as follows:
[0077] Step a-1: Collect voice command-text samples related to video scheduling, perform noise reduction, filtering, and gain control on the voice, perform word2vector vector conversion on the text, and obtain training samples through data enhancement.
[0078] Step a-2: Download the VOSK-MIN and VOSK-BIG models from the VOSK open source website, remove the output layers of the two models, and use the improved twin label auxiliary module to connect the output layers of the two models together to obtain the network structure for fine-tuning training.
[0079] Step a-3: During fine-tuning training, the VOSK-MIN and VOSK-BIG models are trained using the same input sample x, and the outputs of the VOSK-MIN model are obtained respectively. , the output of the VOSK-BIG model .
[0080] Step a-4: Use the improved twin label auxiliary module to connect the outputs of the VOSK-MIN and VOSK-BIG models to obtain the output of the improved twin label auxiliary module :
[0081]
[0082] Use the CTC loss function combined with the twin label auxiliary loss to construct the loss function used for fine-tuning training ,satisfy:
[0083] ;
[0084] ;
[0085] ;
[0086] ;
[0087] .
[0088] in, Represents all selected dynamic programming paths when calculating CTC loss, Indicates all paths One of the paths in Indicates the first t nodes, represents the distribution probability output by the improved twin label auxiliary module on the node, Represents the CTC loss function combined with the improved twin label auxiliary module loss.
[0089] Step a-5: Based on the training samples obtained in step a-1, the gradient back propagation algorithm is used to fine-tune the model. The fine-tuning training network structure with the improved twin label auxiliary module is used, and the model training stops after convergence.
[0090] Step a-6: Remove the improved twin label auxiliary module used for fine-tuning training to obtain the fine-tuned VOSK-MIN model for subsequent speech recognition. Because the VOSK-BIG model has too many parameters and cannot meet the requirements of running in a mobile CPU environment, the VOSK-BIG model is discarded.
[0091] S104: Determine video command and dispatch parameters and control interface information based on the voice recognition text, generate control instructions based on the video command and dispatch parameters and the control interface information, and mobilize the control interface to send the control instructions to perform video command and dispatch.
[0092] Specifically, after completing speech recognition, natural language understanding and corresponding control interface search are performed, and the speech recognition text output by the speech recognition MIN model is used to determine the video command and dispatch parameters and control interface information, thereby generating control instructions and mobilizing the control interface to send control instructions to realize video command and dispatch.
[0093] Optionally, determining video command and dispatch parameters based on voice recognition text specifically includes:
[0094] The speech recognition text is split into multiple keywords, and the video command and dispatch parameters are determined through keyword matching. The video dispatch parameters include video source, target screen and operation.
[0095] Specifically, the speech recognition text output by the speech recognition MIN model is split into keywords, and the video command and dispatch parameters are determined through keyword matching to obtain the key information of the control command.
[0096] Optionally, the keyword matching algorithm may be a string matching algorithm, a fuzzy matching algorithm, etc., which is not limited in this application.
[0097] Figure 4 This is a flow chart of determining video command and dispatch parameters provided by the embodiment of the present application, with reference to Figure 4 The format of the speech recognition text output by the speech recognition MIN model is "video source -> operation -> screen". Assuming that the received voice command is "project video source 3 to screen 2", the output speech recognition text is "video source 3 -> projection -> screen 2". The video command and dispatch parameters obtained through keyword matching include: video source is "video source 3", target screen is "screen 2", and operation is "projection".
[0098] Optionally, determining the interface control information based on the voice recognition text specifically includes:
[0099] A pre-configured configuration file is retrieved based on the video command and dispatch parameters to determine the control interface information of the video command and dispatch; the content of the configuration file includes the video command and dispatch parameters and the corresponding control interface information.
[0100] Specifically, the configuration file is configured in advance, and the configuration content includes video command and dispatch parameters and corresponding control interface information. In actual application, the obtained video command and dispatch parameters are used to retrieve the configuration file, thereby obtaining the control interface information of the video command and dispatch.
[0101] In a possible implementation, an HTTP interface is used as a control interface for video command and dispatch, and "video source", "screen", and "operation" correspond to the "video_source", "screen", and "operate" parameters of the HTTP interface respectively.
[0102] In the embodiment of the present application, the voice control interface is configured by using a configuration file, which is simple and convenient, does not limit the type of controlled device, and makes the application scenarios more flexible and changeable.
[0103] Optionally, the video command and dispatch results can be reported by voice, and the device status can also be reported by voice.
[0104] Optionally, the control instructions and operation results of the video command and dispatch are saved to a log file, and the voice instructions and video data are saved for subsequent viewing and analysis.
[0105] The intelligent voice and video command and dispatch device provided in this application is described below. The intelligent voice and video command and dispatch device described below and the intelligent voice and video command and dispatch method described above can be referenced to each other.
[0106] Figure 5 This is a schematic diagram of the structure of an intelligent voice and video command and dispatch device provided by an embodiment of the present application. Figure 5 As shown, the device at least includes:
[0107] Receiving module 501, used to receive a user's voice instruction for video scheduling;
[0108] The voiceprint identification module 502 is used to input the voice command into the voiceprint identification model and output the voiceprint identification result indicating whether the user voiceprint corresponding to the voice command is registered;
[0109] The speech recognition module 503 is used to input the speech command into the speech recognition MIN model and output speech recognition text if the user voiceprint corresponding to the speech command has been registered; the speech recognition MIN model is trained with the assistance of the speech recognition BIG model and the improved twin label auxiliary module;
[0110] The control module 504 is used to determine video command and dispatch parameters and control interface information based on the voice recognition text, generate control instructions based on the video dispatch parameters and control interface information, and mobilize the control interface to send the control instructions to perform video command and dispatch.
[0111] Optionally, the speech recognition MIN model is trained based on the following steps:
[0112] Construct speech recognition BIG model and speech recognition MIN model with different parameter amounts and speech recognition accuracy;
[0113] Build an improved twin label auxiliary module to connect the outputs of the speech recognition BIG model and the speech recognition MIN model;
[0114] Based on the pre-acquired video scheduling related voice command-text samples, the speech recognition BIG model and speech recognition MIN model are trained;
[0115] By improving the twin label auxiliary module, the output of the speech recognition BIG model is obtained as auxiliary information for the speech recognition MIN model to assist training until the model converges;
[0116] Delete the improved twin label auxiliary module to obtain the trained speech recognition MIN model.
[0117] Optionally, the loss function in the speech recognition MIN model training process is a CTC loss function combined with an improved twin label auxiliary module loss.
[0118] Optionally, the control module 504 determines the video command and dispatch parameters based on the voice recognition text, including:
[0119] The speech recognition text is split into multiple keywords, and the video command and dispatch parameters are determined through keyword matching. The video dispatch parameters include video source, target screen and operation.
[0120] Optionally, the control module 504 determines the control interface information of the video command and dispatch based on the voice recognition text, including:
[0121] A pre-configured configuration file is retrieved based on the video command and dispatch parameters to determine the control interface information of the video command and dispatch; the content of the configuration file includes the video command and dispatch parameters and the corresponding control interface information.
[0122] Optionally, the voiceprint identification module 502 is specifically configured to:
[0123] Extracting voiceprint features from voice commands based on a voiceprint feature extraction network;
[0124] Based on the voiceprint feature similarity scoring network, the scoring results of the voiceprint feature and the target person's voiceprint database are obtained;
[0125] Compare the scoring result with the pre-set confidence threshold and output the voiceprint identification result;
[0126] Among them, the target voiceprint database is constructed based on the voiceprint features of registered users.
[0127] It should be understood that the above-mentioned device is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program module in the device are similar to those described in the above-mentioned method. The working process of the device can refer to the corresponding process in the above-mentioned method and will not be repeated here.
[0128] Based on the methods described in the above embodiments, embodiments of the present application provide an electronic device. The device may include: at least one memory for storing programs and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is configured to execute the methods described in the above embodiments.
[0129] Figure 6 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 6 As shown, the electronic device may include: a processor 601, a communications interface 602, a memory 603, and a communication bus 604. The processor 601, the communications interface 602, and the memory 603 communicate with each other via the communication bus 604. The processor 601 may call software instructions in the memory 603 to execute the methods described in the above embodiments.
[0130] In addition, the logic instructions in the memory 603 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, or the part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application.
[0131] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.
[0132] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.
[0133] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0134] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC.
[0135] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via such a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).
[0136] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.
[0137] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. An intelligent voice and video command and dispatch method, characterized in that: include: receiving user's voice commands for video scheduling; Input the voice command into the voiceprint identification model, and output a voiceprint identification result indicating whether the user voiceprint corresponding to the voice command is registered; If the user voiceprint corresponding to the voice command has been registered, the voice command is input into the voice recognition MIN model, and the voice recognition text is output; the voice recognition MIN model is trained with the assistance of the voice recognition BIG model and the improved twin label auxiliary module; The loss function of auxiliary training is the CTC loss function combined with the improved twin label auxiliary module loss; The CTC loss function combined with the improved twin label auxiliary module loss is: ; in, ; ; ; ; Specifically, For all selected dynamic programming paths when calculating CTC loss, For all paths One of the paths in The first t nodes, The distribution probability output by the improved twin label auxiliary module on the node; in, ; ; ; Specifically, To improve the output of the twin label auxiliary module; Output of the speech recognition BIG model; is the output of the speech recognition MIN model; Represents the model input, i.e., the speech sample; Video command and dispatch parameters and control interface information are determined based on the voice recognition text, control instructions are generated based on the video dispatch parameters and control interface information, and the control interface is mobilized to send the control instructions to perform video command and dispatch.
2. The intelligent voice and video command and dispatch method according to claim 1, characterized in that: The speech recognition MIN model is trained based on the following steps: Construct speech recognition BIG model and speech recognition MIN model with different parameter amounts and speech recognition accuracy; Construct an improved twin label auxiliary module to connect the outputs of the speech recognition BIG model and the speech recognition MIN model; Based on the pre-acquired video scheduling related voice instruction-text samples, training the speech recognition BIG model and the speech recognition MIN model; The output of the speech recognition BIG model is obtained through the improved twin label auxiliary module as auxiliary information of the speech recognition MIN model to assist in training until the model converges; Delete the improved twin label auxiliary module to obtain the trained speech recognition MIN model.
3. The intelligent voice and video command and dispatch method according to claim 2, characterized in that: The loss function in the speech recognition MIN model training process is the CTC loss function combined with the improved twin label auxiliary module loss.
4. The intelligent voice and video command and dispatch method according to claim 1, characterized in that: Determining video command and dispatch parameters based on the voice recognition text includes: The speech recognition text is split into multiple keywords, and the video command scheduling parameters are determined through keyword matching. The video scheduling parameters include video source, target screen and operation.
5. The intelligent voice and video command and dispatch method according to claim 1, characterized in that: Determining control interface information for video command and dispatch based on the voice recognition text includes: A pre-configured configuration file is retrieved based on the video command and dispatch parameters to determine the control interface information of the video command and dispatch; the content of the configuration file includes the video command and dispatch parameters and the corresponding control interface information.
6. The intelligent voice and video command and dispatch method according to claim 1, characterized in that: The outputting of the voiceprint identification result of whether the user voiceprint corresponding to the voice command is registered includes: Extracting voiceprint features from the voice command based on a voiceprint feature extraction network; Based on the voiceprint feature similarity scoring network, obtaining the scoring result of the voiceprint feature and the target person's voiceprint database; Compare the scoring result with a preset confidence threshold and output the voiceprint identification result; The target voiceprint database is constructed based on the voiceprint features of registered users.
7. An intelligent voice and video command and dispatch device, characterized in that: include: A receiving module, configured to receive a user's voice command for video scheduling; a voiceprint identification module, configured to input the voice command into a voiceprint identification model and output a voiceprint identification result indicating whether the user voiceprint corresponding to the voice command is registered; A speech recognition module is used to input the speech command into a speech recognition MIN model and output speech recognition text if the user voiceprint corresponding to the speech command has been registered; the speech recognition MIN model is trained with the assistance of the speech recognition BIG model and the improved twin label auxiliary module; The loss function of auxiliary training is the CTC loss function combined with the improved twin label auxiliary module loss; The CTC loss function combined with the improved twin label auxiliary module loss is: ; in, ; ; ; ; Specifically, For all selected dynamic programming paths when calculating CTC loss, For all paths One of the paths in The first t nodes, The distribution probability output by the improved twin label auxiliary module on the node; in, ; ; ; Specifically, To improve the output of the twin label auxiliary module; Output of the speech recognition BIG model; is the output of the speech recognition MIN model; Represents the model input, i.e., the speech sample; The control module is used to determine the video command and dispatch parameters and control interface information based on the voice recognition text, generate control instructions based on the video dispatch parameters and control interface information, and mobilize the control interface to send the control instructions to perform video command and dispatch.
8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed on a processor, the processor is caused to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that When the computer program product is run on a processor, the processor is enabled to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Monitoring scheduling method and device, computer device and storage medium
CN110099246A
Picture classification method and system based on twin label auxiliary module, and medium
CN115115874A
Personnel voiceprint recognition and authentication method, system and device for power dispatching system
CN116229988A