Voice control execution methods and devices, electronic devices and storage media

By simplifying the offline instruction word system through a speech recognition model, recognizing the posterior probability of speech frames and decoding execution attribute information, the problem of complex structure of the offline instruction word system is solved, achieving the effects of simplified structure and improved recognition rate.

CN119207377BActive Publication Date: 2026-05-26BEIJING CO WHEELS TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2023-06-27
Publication Date
2026-05-26

Smart Images

  • Figure CN119207377B_ABST
    Figure CN119207377B_ABST
Patent Text Reader

Abstract

The voice control execution method, apparatus, electronic device, and storage medium disclosed herein relate to the field of voice processing technology. The main technical solution includes: recognizing the speech to be recognized based on a preset acoustic model in a speech recognition model to obtain the posterior probability corresponding to each frame of speech in the speech to be recognized, where the posterior probability is the probability of the phoneme information corresponding to each frame of speech in the speech to be recognized occurring; calling a preset entity word extraction model in the speech recognition model to decode the posterior probability corresponding to each frame of speech to determine the execution attribute information corresponding to the speech to be recognized; and controlling the execution of the speech to be recognized based on the recognition result of the execution attribute information. Compared with related technologies, the embodiments of this disclosure combine the posterior probability of the speech to be recognized with a preset entity word extraction model, realizing the functions of a language model, decoder, and word segmentation model, thus simplifying the overall structure of the offline command word system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of speech processing technology, and in particular to a speech control execution method and apparatus, electronic device and storage medium. Background Technology

[0002] With the advancement of technology and the progress of the times, more and more electronic devices can be controlled through voice commands. Currently, there are two types of voice control: online control and offline control. For online control, if the network is interrupted, the device cannot connect to the network, directly causing the electronic device to malfunction. On the other hand, electronic devices with offline control include an offline command word system, which is directly embedded in the electronic device. The offline command word system does not need to be connected to the network and can run locally, unaffected by network issues.

[0003] However, in related technologies, offline command word systems obtain audio recognition results through acoustic models, language models, and decoders, then extract control commands from sentences through word segmentation models, and then execute corresponding voice control functions according to the control commands. The overall structure is relatively complex. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for voice control execution. Its main objective is to address the problem of the complex overall structure of offline command word systems in related technologies.

[0005] According to a first aspect of this disclosure, a method for performing voice control is provided, comprising:

[0006] The speech to be recognized is identified based on the preset acoustic model in the speech recognition model, and the posterior probability corresponding to each frame of speech in the speech to be recognized is obtained. The posterior probability is the probability of the phoneme information corresponding to each frame of speech in the speech to be recognized appearing.

[0007] The preset entity word extraction model in the speech recognition model is invoked to decode the posterior probability corresponding to each frame of speech and determine the execution attribute information corresponding to the speech to be recognized.

[0008] Based on the recognition result of the execution attribute information, the speech to be recognized is controlled to be executed.

[0009] Optionally, the step of recognizing the speech based on a preset acoustic model in the speech recognition model to obtain the posterior probability corresponding to each frame of the speech to be recognized includes:

[0010] The phoneme information corresponding to each frame of the speech to be recognized is obtained based on the preset acoustic model in the speech recognition model.

[0011] The posterior probability corresponding to the phoneme information of each frame of speech is calculated using a preset probability algorithm.

[0012] Optionally, before calling the preset entity word extraction model in the speech recognition model to decode the posterior probability corresponding to each frame of speech, the following steps are included:

[0013] The preset entity word extraction model is trained to obtain the trained preset entity word extraction model;

[0014] The trained preset entity word extraction model is loaded into a preset speech recognition device.

[0015] Optionally, training the preset entity word extraction model includes:

[0016] Acquire a preset amount of training voice data;

[0017] The training speech data is input into the preset entity word extraction model to obtain the training execution attribute information corresponding to the training speech data;

[0018] Based on the training execution attribute information, the loss value of the preset entity word extraction model is calculated through a preset loss function. The loss value is a value that measures the degree of difference between the predicted execution attribute information and the actual execution attribute information of the preset entity word extraction model.

[0019] Based on the loss value, the preset entity word extraction model is optimized using a preset optimization algorithm to obtain the trained preset entity word extraction model.

[0020] Optionally, before recognizing the speech to be recognized based on a preset acoustic model in the speech recognition model and obtaining the posterior probability corresponding to each frame of the speech to be recognized, the following steps are included:

[0021] A preset acoustic model long short-term memory network is loaded into a preset speech recognition device. The preset acoustic model long short-term memory network encodes the speech to be recognized to obtain a phoneme vector for each frame of speech. The phoneme vector corresponds to the posterior probability of the speech.

[0022] Optionally, the execution attribute information includes: expectation information, domain information, and classification information; wherein,

[0023] The expected information is the action and / or control expected to be performed in the speech to be recognized;

[0024] The domain information is the customized category information of the speech to be recognized;

[0025] The classification information is the application scenario classification corresponding to the speech to be recognized.

[0026] According to a second aspect of this disclosure, a voice-controlled execution device is provided, comprising:

[0027] The recognition unit is used to recognize the speech to be recognized based on a preset acoustic model in the speech recognition model, and to obtain the posterior probability corresponding to each frame of speech in the speech to be recognized. The posterior probability is the probability of the phoneme information corresponding to each frame of speech in the speech to be recognized appearing.

[0028] The decoding unit is used to call the preset entity word extraction model in the speech recognition model, decode the posterior probability corresponding to each frame of speech, and determine the execution attribute information corresponding to the speech to be recognized.

[0029] An execution unit is used to control the execution of the speech to be recognized based on the recognition result of the execution attribute information.

[0030] Optionally, the identification unit includes:

[0031] The acquisition module is used to acquire the phoneme information corresponding to each frame of the speech to be recognized based on the preset acoustic model in the speech recognition model;

[0032] The calculation module is used to calculate the posterior probability corresponding to the phoneme information of each frame of speech using a preset probability algorithm.

[0033] Optionally, the device further includes:

[0034] The training unit is used to train the preset entity word extraction model before the determining unit calls the preset entity word extraction algorithm and determines the execution attribute information corresponding to the speech to be recognized based on the posterior probability corresponding to each frame of speech, so as to obtain the trained preset entity word extraction model.

[0035] The loading unit is used to load the trained preset entity word extraction model into a preset speech recognition device.

[0036] Optionally, the training unit includes:

[0037] The acquisition module is used to acquire a preset amount of training voice data;

[0038] The input module is used to input the training speech data into the preset entity word extraction model to obtain the training execution attribute information corresponding to the training speech data;

[0039] The calculation module is used to calculate the loss value of the preset entity word extraction model based on the training execution attribute information and through a preset loss function. The loss value is a value that measures the degree of difference between the predicted execution attribute information and the actual execution attribute information of the preset entity word extraction model.

[0040] The optimization module is used to optimize the preset entity word extraction model based on the loss value using a preset optimization algorithm to obtain the trained preset entity word extraction model.

[0041] Optionally, the loading unit is further configured to load a preset acoustic model long short-term memory network into a preset speech recognition device before the recognition unit recognizes the speech to be recognized. The preset acoustic model long short-term memory network encodes the speech to be recognized to obtain a phoneme vector for each frame of speech, and the phoneme vector corresponds to the posterior probability of the speech.

[0042] According to a third aspect of this disclosure, a vehicle is provided, wherein the vehicle includes a voice control actuator as described in a second aspect of this disclosure.

[0043] According to a fourth aspect of this disclosure, an electronic device is provided, comprising:

[0044] At least one processor; and

[0045] A memory communicatively connected to the at least one processor; wherein,

[0046] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect above.

[0047] According to a fifth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are configured to cause the computer to perform the method described in the first aspect above.

[0048] According to a sixth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in the first aspect above.

[0049] The voice control execution method, apparatus, electronic device, and storage medium disclosed herein recognize the speech to be recognized based on a preset acoustic model in a speech recognition model, obtaining the posterior probability corresponding to each frame of the speech to be recognized. The posterior probability is the probability of the phoneme information corresponding to each frame of the speech to be recognized occurring. A preset entity word extraction model in the speech recognition model is invoked to decode the posterior probability corresponding to each frame of the speech, determining the execution attribute information corresponding to the speech to be recognized. Based on the recognition result of the execution attribute information, the speech to be recognized is controlled and executed. Compared with related technologies, the embodiments of this disclosure combine the posterior probability of the speech to be recognized with a preset entity word extraction model, realizing the functions of a language model, decoder, and word segmentation model, thus simplifying the overall structure of the offline command word system.

[0050] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0051] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0052] Figure 1 A flowchart illustrating a voice control execution method provided in an embodiment of this disclosure;

[0053] Figure 2 A schematic diagram of the structure of a voice-controlled execution device provided in an embodiment of this disclosure;

[0054] Figure 3 A schematic diagram of the structure of another voice control execution device provided in an embodiment of this disclosure;

[0055] Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0056] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0057] The following description, with reference to the accompanying drawings, describes a voice control execution method and apparatus, electronic device, and storage medium according to embodiments of the present disclosure.

[0058] Figure 1 This is a flowchart illustrating a voice control execution method provided in an embodiment of the present disclosure.

[0059] like Figure 1 As shown, the method includes the following steps:

[0060] Step 101: Based on the preset acoustic model in the speech recognition model, the speech to be recognized is recognized to obtain the posterior probability corresponding to each frame of speech in the speech to be recognized. The posterior probability is the probability of the phoneme information corresponding to each frame of speech in the speech to be recognized appearing.

[0061] This disclosure can be applied to various offline scenarios and supports multiple voice customizations, such as: voice customization for functional and proprietary industry functions: integrated stove scenario: "turn on heating" "fan speed three levels", smart toilet scenario: "flush", massage chair scenario: "increase intensity" "decompression mode", vehicle scenario: "turn on air conditioning" "turn on music", etc. At the same time, generalized voice commands can also be executed, such as: "turn it up a bit", "set the temperature to 20 degrees", etc.

[0062] In this embodiment of the disclosure, the speech to be recognized needs to be recognized through the acoustic model in the speech recognition model. The posterior probability is a probability value, represented numerically. For details regarding the posterior probability, please refer to the detailed descriptions in related technologies; therefore, they will not be elaborated upon here.

[0063] Step 102: Call the preset entity word extraction model in the speech recognition model to decode the posterior probability corresponding to each frame of speech and determine the execution attribute information corresponding to the speech to be recognized.

[0064] In this embodiment of the disclosure, the execution attribute information can be in text form or other string form. Specifically, this embodiment of the disclosure does not limit the form of the execution attribute information.

[0065] The execution attribute information includes: expected information, domain information, and classification information. In the execution attribute information of the speech to be recognized, the number of domain information can be zero, while the number of expected information and classification information cannot be zero. At the same time, there can be multiple expected information, domain information, and classification information in the execution attribute information of the speech to be recognized. Specifically, the number of expected information, domain information, and classification information in the speech to be recognized is not limited by the embodiments of this disclosure.

[0066] Step 103: Based on the recognition result of the execution attribute information, control the execution of the speech to be recognized.

[0067] In this embodiment, the execution attribute information needs to be output to the execution module. After the execution module recognizes and processes the execution attribute information of the speech to be recognized, it controls the execution of the speech containing the execution attribute information. For details regarding the execution module's recognition and processing of the execution attribute information of the speech to be recognized, please refer to the detailed descriptions in related technologies; therefore, they will not be elaborated upon here.

[0068] The voice control execution method provided in this disclosure identifies the speech to be identified based on a preset acoustic model in a speech recognition model, obtaining the posterior probability corresponding to each frame of the speech to be identified. The posterior probability is the probability of the phoneme information corresponding to each frame of the speech to be identified appearing. A preset entity word extraction model in the speech recognition model is invoked to decode the posterior probability corresponding to each frame of the speech, determining the execution attribute information corresponding to the speech to be identified. Based on the recognition result of the execution attribute information, the speech to be identified is controlled and executed. Compared with related technologies, the embodiments of this disclosure combine the posterior probability of the speech to be identified with a preset entity word extraction model, realizing the functions of a language model, decoder, and word segmentation model, thus simplifying the overall structure of the offline command word system.

[0069] In one possible implementation of this disclosure, the posterior probability corresponding to each frame of speech in the speech to be recognized is essentially the posterior probability corresponding to the phoneme information of each frame of speech in the speech to be recognized. The acquisition of the posterior probability can be achieved, but is not limited to, the following methods: obtaining the phoneme information corresponding to each frame of speech in the speech to be recognized based on a preset acoustic model in the speech recognition model; and calculating the posterior probability corresponding to the phoneme information of each frame of speech using a preset probability algorithm. Here, phoneme information is the smallest unit of speech divided according to the natural attributes of the speech to be recognized. From an acoustic perspective, phoneme information is the smallest unit of speech divided from the perspective of sound quality. From a physiological perspective, one vocalization action forms one phoneme. For specific details regarding phoneme information, please refer to the detailed descriptions in related technologies; therefore, they will not be elaborated upon here.

[0070] The preset probability algorithm is a custom-selected probability algorithm, such as Bayes' theorem, maximum likelihood estimation, maximum a posteriori estimation, etc. Specifically, the selection of the preset probability algorithm can be determined according to the actual application situation, and this embodiment of the disclosure does not impose any restrictions.

[0071] Related to the above embodiments, the acquisition of the preset acoustic model in the speech recognition model can be achieved, but is not limited to, the following method: A preset acoustic model Long Short-Term Memory (LSTM) network is loaded into a preset speech recognition device. The LSTM network encodes the speech to be recognized to obtain a phoneme vector for each frame of speech, where the phoneme vector corresponds to the posterior probability of the speech. The LSTM network can be used as the preset acoustic model. For details regarding LSTM networks, please refer to the detailed descriptions in related technologies; therefore, they will not be elaborated upon here.

[0072] In one possible implementation of this disclosure, a preset entity word extraction model in the speech recognition model needs to be trained to enable the preset entity word extraction model to perform corresponding functions. Therefore, in order to obtain a preset entity word extraction model that can perform the corresponding functions, the following methods can be used, but are not limited to: training the preset entity word extraction model to obtain the trained preset entity word extraction model; and loading the trained preset entity word extraction model into a preset speech recognition device. The preset speech recognition device is the device applied to the offline scenario in this disclosure embodiment; that is, the preset speech recognition device is the device used in the offline scenario in which this disclosure embodiment is applied.

[0073] Related to the above embodiments, training the preset entity word extraction model can be implemented in, but is not limited to, the following manner: acquiring a preset amount of training speech data; inputting the training speech data into the preset entity word extraction model to obtain training execution attribute information corresponding to the training speech data; based on the training execution attribute information, calculating the loss value of the preset entity word extraction model through a preset loss function, wherein the loss value is a value that measures the degree of difference between the predicted execution attribute information and the actual execution attribute information of the preset entity word extraction model; and optimizing the preset entity word extraction model through a preset optimization algorithm based on the loss value to obtain the trained preset entity word extraction model.

[0074] In this embodiment, the preset quantity refers to a custom-set quantity value. During the training of the preset entity word extraction model, the model performs best when the amount of training speech data is within a certain range. If the amount of training speech data is too large, the training speech data may interfere with each other, leading to a high false recognition rate for the preset entity word extraction model. Conversely, if the amount of training speech data is too small, insufficient data during training will also result in a high false recognition rate. Specifically, the amount of training speech data can be determined based on the actual application; this embodiment does not impose any limitations.

[0075] The preset loss function is a user-selected computational function, such as the mean squared error function, the mean absolute value error function, the quantile loss function, etc. The preset optimization algorithm is a user-selected algorithm, such as gradient descent, Newton's method, quasi-Newton method, improved iterative scaling method, etc. Specifically, the selection of the preset loss function and the preset optimization algorithm can be determined according to the actual application situation, and this disclosure embodiment does not impose any restrictions.

[0076] In one possible implementation of this disclosure, as a refinement of step 102 above, the execution attribute information includes: expected information, domain information, and classification information; wherein, the expected information is the action and / or control expected to be performed in the speech to be recognized; the domain information is the custom category information of the speech to be recognized; and the classification information is the application scenario classification corresponding to the speech to be recognized.

[0077] For example, if the voice to be recognized is "turn on the TV," then in the execution attribute information of the voice to be recognized, the expected information is "turn on," the domain information is "function," and the category information is "TV." If the voice to be recognized is "air conditioner fan speed three," then in the execution attribute information of the voice to be recognized, the expected information is "adjust," the domain information is "speed," and the category information is "air conditioner."

[0078] It should be noted that the domain information of similar speech samples to be recognized is the same. This domain information can be custom-named. When training the preset entity word extraction model, the domain information can be custom-named. For example, the speech samples used for training might be "air conditioner fan speed level 3" or "fan speed level 2." During training, these speech samples will be categorized into the same class, and the domain information can be custom-named "speed level." Then, when the trained preset entity word extraction model determines the execution attribute information corresponding to the speech sample, if the speech sample is "air conditioner fan speed level 3" or "fan speed level 2," the output domain information will be "speed level." Specifically, the types of speech samples to be recognized and the custom naming of the domain information can be determined according to the actual application; this embodiment does not impose any limitations.

[0079] In practical applications, since most speech recognition methods achieve high recognition rates in quiet environments but lower rates in noisy or high-background environments, the following methods can be used, but are not limited to: denoising the original speech to obtain the speech to be recognized. Denoising the original speech can significantly improve the recognition rate of the speech to be recognized.

[0080] In summary, the embodiments disclosed herein can achieve the following effects:

[0081] 1. The pre-defined entity word extraction model is used to implement the functions of language model, decoder and word segmentation model, which simplifies the overall structure of the offline instruction word system, reduces memory and resource consumption, and can be run and used under low resource conditions.

[0082] 2. By performing noise reduction processing on the original speech, the speech to be recognized is obtained, which greatly improves the recognition rate of the speech to be recognized.

[0083] Corresponding to the above-described voice control execution method, the present invention also proposes a voice control execution device. Since the device embodiments of the present invention correspond to the above-described method embodiments, details not disclosed in the device embodiments can be referred to the above-described method embodiments, and will not be repeated here.

[0084] Figure 2 This is a schematic diagram of the structure of a voice-controlled execution device provided in an embodiment of the present disclosure, as shown below. Figure 2 As shown, it includes:

[0085] The recognition unit 21 is used to recognize the speech to be recognized based on the preset acoustic model in the speech recognition model, and to obtain the posterior probability corresponding to each frame of speech in the speech to be recognized. The posterior probability is the probability of the phoneme information corresponding to each frame of speech in the speech to be recognized appearing.

[0086] Decoding unit 22 is used to call the preset entity word extraction model in the speech recognition model, decode the posterior probability corresponding to each frame of speech, and determine the execution attribute information corresponding to the speech to be recognized.

[0087] The execution unit 23 is used to control the execution of the speech to be recognized based on the recognition result of the execution attribute information.

[0088] The voice control execution device provided in this disclosure recognizes the speech to be recognized based on a preset acoustic model in a speech recognition model, obtaining the posterior probability corresponding to each frame of the speech to be recognized. The posterior probability is the probability of the phoneme information corresponding to each frame of the speech to be recognized occurring. A preset entity word extraction model in the speech recognition model is invoked to decode the posterior probability corresponding to each frame of the speech, determining the execution attribute information corresponding to the speech to be recognized. Based on the recognition result of the execution attribute information, the speech to be recognized is controlled and executed. Compared with related technologies, the embodiments of this disclosure combine the posterior probability of the speech to be recognized with a preset entity word extraction model, realizing the functions of a language model, decoder, and word segmentation model, thus simplifying the overall structure of the offline command word system.

[0089] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 3As shown, the identification unit 21 includes:

[0090] The acquisition module 211 is used to acquire the phoneme information corresponding to each frame of the speech to be recognized based on the preset acoustic model in the speech recognition model;

[0091] The calculation module 212 is used to calculate the posterior probability corresponding to the phoneme information of each frame of speech through a preset probability algorithm.

[0092] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 3 As shown, the device further includes:

[0093] Training unit 24 is used to train the preset entity word extraction model before the determining unit 22 calls the preset entity word extraction algorithm and determines the execution attribute information corresponding to the speech to be recognized according to the posterior probability corresponding to each frame of speech, so as to obtain the trained preset entity word extraction model, wherein the preset entity word extraction model includes the preset entity word extraction algorithm.

[0094] The loading unit 25 is used to load the trained preset entity word extraction model into a preset speech recognition device.

[0095] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 3 As shown, the training unit 24 includes:

[0096] The acquisition module 241 is used to acquire a preset number of training voice data;

[0097] The input module 242 is used to input the training speech data into the preset entity word extraction model to obtain the training execution attribute information corresponding to the training speech data;

[0098] The calculation module 243 is used to calculate the loss value of the preset entity word extraction model based on the training execution attribute information and through a preset loss function. The loss value is a value that measures the degree of difference between the predicted execution attribute information and the actual execution attribute information of the preset entity word extraction model.

[0099] The optimization module 244 is used to optimize the preset entity word extraction model based on the loss value using a preset optimization algorithm to obtain the trained preset entity word extraction model.

[0100] Furthermore, in one possible implementation of this embodiment, the loading unit 25 is further configured to load a preset acoustic model long short-term memory network into a preset speech recognition device before the recognition unit 21 recognizes the speech to be recognized. The preset acoustic model long short-term memory network encodes the speech to be recognized to obtain a phoneme vector for each frame of speech, and the phoneme vector corresponds to the posterior probability of the speech.

[0101] In this embodiment of the disclosure, a vehicle is also provided, wherein the vehicle is equipped with a voice control actuator.

[0102] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of the embodiments of this disclosure, and the principle is the same. Therefore, the embodiments of this disclosure are not limited thereto.

[0103] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0104] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0105] like Figure 4 As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 402 or a computer program loaded from storage unit 408 into RAM (Random Access Memory) 403. RAM 403 may also store various programs and data required for the operation of device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. I / O (Input / Output) interface 405 is also connected to bus 404.

[0106] Multiple components in device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of monitors, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0107] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the voice control execution method. For example, in some embodiments, the voice control execution method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed on device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform the aforementioned voice control execution method by any other suitable means (e.g., by means of firmware).

[0108] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0109] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0110] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0111] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0112] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0113] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0114] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0115] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0116] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for executing voice control, characterized in that, include: The speech to be recognized is identified based on the preset acoustic model in the speech recognition model, and the posterior probability corresponding to each frame of speech in the speech to be recognized is obtained. The posterior probability is the probability of the phoneme information corresponding to each frame of speech in the speech to be recognized appearing. The preset entity word extraction model in the speech recognition model is invoked to decode the posterior probability corresponding to each frame of speech and determine the execution attribute information corresponding to the speech to be recognized. Based on the recognition result of the execution attribute information, the speech to be recognized is controlled to be executed; wherein, the execution attribute information includes expectation information, domain information and classification information; The step of controlling the execution of the speech to be recognized based on the recognition result of the execution attribute information includes: The execution attribute information is output to the execution module, so that the execution module can identify and process the execution attribute information and then control the execution of the speech to be recognized, which contains the execution attribute information.

2. The method according to claim 1, characterized in that, The speech recognition model based on the preset acoustic model in the speech recognition model recognizes the speech to be recognized, and the posterior probability corresponding to each frame of the speech to be recognized is as follows: The phoneme information corresponding to each frame of the speech to be recognized is obtained based on the preset acoustic model in the speech recognition model. The posterior probability corresponding to the phoneme information of each frame of speech is calculated using a preset probability algorithm.

3. The method according to claim 1, characterized in that, Before calling the preset entity word extraction model in the speech recognition model to decode the posterior probability corresponding to each frame of speech, the method further includes: The preset entity word extraction model is trained to obtain the trained preset entity word extraction model; The trained preset entity word extraction model is loaded into a preset speech recognition device.

4. The method according to claim 3, characterized in that, The training of the preset entity word extraction model includes: Acquire a preset amount of training voice data; The training speech data is input into the preset entity word extraction model to obtain the training execution attribute information corresponding to the training speech data; Based on the training execution attribute information, the loss value of the preset entity word extraction model is calculated through a preset loss function. The loss value is a value that measures the degree of difference between the predicted execution attribute information and the actual execution attribute information of the preset entity word extraction model. Based on the loss value, the preset entity word extraction model is optimized using a preset optimization algorithm to obtain the trained preset entity word extraction model.

5. The method according to claim 1, characterized in that, Before recognizing the speech to be recognized using a preset acoustic model in the speech recognition model and obtaining the posterior probability corresponding to each frame of the speech to be recognized, the method further includes: A preset acoustic model long short-term memory network is loaded into a preset speech recognition device. The preset acoustic model long short-term memory network encodes the speech to be recognized to obtain a phoneme vector for each frame of speech. The phoneme vector corresponds to the posterior probability of the speech.

6. The method according to any one of claims 1-5, characterized in that, The expected information is the action and / or control expected to be performed in the speech to be recognized; the domain information is the custom category information of the speech to be recognized; and the classification information is the application scenario classification corresponding to the speech to be recognized.

7. A voice-controlled execution device, characterized in that, include: The recognition unit is used to recognize the speech to be recognized based on a preset acoustic model in the speech recognition model, and to obtain the posterior probability corresponding to each frame of speech in the speech to be recognized. The posterior probability is the probability of the phoneme information corresponding to each frame of speech in the speech to be recognized appearing. The decoding unit is used to call the preset entity word extraction model in the speech recognition model, decode the posterior probability corresponding to each frame of speech, and determine the execution attribute information corresponding to the speech to be recognized. An execution unit is configured to control the execution of the speech to be recognized based on the recognition result of the execution attribute information; wherein, the execution attribute information includes expectation information, domain information, and classification information; The step of controlling the execution of the speech to be recognized based on the recognition result of the execution attribute information includes: The execution attribute information is output to the execution module, so that the execution module can identify and process the execution attribute information and then control the execution of the speech to be recognized, which contains the execution attribute information.

8. A vehicle, characterized in that, The vehicle includes the voice control actuator as described in claim 7.

9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.