Speech recognition method
By using a speech recognition model trained based on the original loss and auxiliary loss, the speech information to be recognized is recognized twice to remove interference information, which solves the problem of speech recognition accuracy in noisy environments and achieves higher speech information recognition accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing speech recognition models have poor accuracy in the presence of noise, making them difficult to apply effectively in the real world.
A speech recognition model trained based on original loss and auxiliary loss is adopted. The network layer determined by the auxiliary loss is used for information mapping, and the network layer determined by the original loss is used for noise recognition to remove interference information and recognize the original audio information.
By performing two recognition processes, interference information in the speech information to be recognized is effectively removed, thereby improving the accuracy of speech information recognition.
Smart Images

Figure CN115995229B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech information processing technology, and more specifically, to a speech recognition method. Background Technology
[0002] Currently, speech recognition models are usually trained on clean, noise-free corpora. They perform well in recognizing noise-free speech information. However, when the speech information to be recognized contains noise, the performance of such speech recognition models is often poor, that is, the noise resistance of the model is poor. Since real-world speech usually contains background noise, reverberation and other nonlinear distortions, there are usually technical problems of inaccurate recognition when using such speech recognition models to recognize real-world speech.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This invention provides a speech recognition method to at least solve the technical problem of low accuracy in speech information recognition.
[0005] According to one aspect of the present invention, a speech recognition method is provided. The method may include: acquiring monitored speech information to be recognized, wherein the speech information to be recognized includes at least original audio information to be recognized; invoking a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss being obtained based on a noise recognition result, the noise recognition result being obtained by noise recognition of a first noise information sample in a training sample set, the auxiliary loss being obtained based on an information mapping result, and the information mapping result being obtained by information mapping of at least a second noise information sample in the training sample set; using a network layer in the speech recognition model determined at least by the auxiliary loss to perform information mapping on the speech information to be recognized, obtaining a target information mapping result; using a network layer in the speech recognition model determined at least by the original loss to perform noise recognition on the target information mapping result, identifying interference information, wherein the interference information is information that interferes with the original audio information; and removing the interference information from the speech information to be recognized to identify the original audio information.
[0006] According to another aspect of the present invention, a method for determining a speech recognition model is also provided. This method may include: obtaining a first noise information sample and a second noise information sample from a training sample set; determining the original loss of the initial speech recognition model based on the noise recognition result obtained by performing speech recognition on the first noise information sample using the initial speech recognition model, and determining an auxiliary loss of the initial speech recognition model based at least on the information mapping result obtained by performing information mapping on the second noise information sample using the initial speech recognition model; training the initial speech recognition model based on the original loss and the auxiliary loss to obtain a speech recognition model, wherein at least one network layer in the speech recognition model, determined by the auxiliary loss, is used to perform information mapping on the speech information to be recognized to obtain a target information mapping result, the speech information to be recognized contains at least the original audio information to be recognized, and the network layer in the speech recognition model, determined by the original loss, is used to perform noise recognition on the target information mapping result, identifying interference information, the interference information being information that interferes with the original audio information, and is used to identify the original audio information from the speech information to be recognized.
[0007] According to another aspect of the present invention, another speech recognition method is also provided, comprising: acquiring speech information to be recognized sent to a client, wherein the speech information to be recognized includes at least original audio information to be recognized; using a network layer in a speech recognition model determined by at least an auxiliary loss to perform information mapping on the speech information to be recognized, thereby obtaining a target information mapping result, wherein the speech recognition model is obtained by training an initial speech recognition model based on the original loss and the auxiliary loss, the original loss is obtained based on the noise recognition result, the noise recognition result is obtained by noise recognition of a first noise information sample in the training sample set, the auxiliary loss is obtained based on the information mapping result, and the information mapping result is obtained by information mapping of at least a second noise information sample in the training sample set; using a network layer in the speech recognition model determined by at least the original loss to perform noise recognition on the target information mapping result, thereby identifying interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized, thereby recognizing the original audio information; and activating the client based on the original audio information.
[0008] According to another aspect of the present invention, a speech recognition method is also provided, comprising: inputting monitored speech information to be recognized into a virtual reality (VR) device or an augmented reality (AR) device, wherein the speech information to be recognized includes at least original audio information to be recognized; invoking a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss being obtained based on a noise recognition result, the noise recognition result being obtained by noise recognition of a first noise information sample in a training sample set, the auxiliary loss being obtained based on an information mapping result, and the information mapping result being obtained by information mapping of at least a second noise information sample in a training sample set; using a network layer in the speech recognition model determined at least by the auxiliary loss to perform information mapping on the speech information to be recognized, obtaining a target information mapping result; using a network layer in the speech recognition model determined at least by the original loss to perform noise recognition on the target information mapping result, identifying interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized, and identifying the original audio information; and activating the VR device or AR device using the original audio information.
[0009] According to another aspect of the present invention, a speech recognition method is also provided, comprising: acquiring monitored speech information to be recognized by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter being the speech information to be recognized, the speech information to be recognized including at least the original audio information to be recognized; calling a speech recognition model, wherein the speech recognition model is obtained by mapping the initial speech information based on the original loss and the auxiliary loss; using a network layer in the speech recognition model determined by at least the auxiliary loss to perform information mapping on the speech information to be recognized, obtaining a target information mapping result; using a network layer in the speech recognition model determined by at least the original loss to perform noise recognition on the target information mapping result, identifying interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized, identifying the original audio information; and outputting the original audio information by calling a second interface, wherein the second interface includes a second parameter, the parameter value of the second parameter being the original audio information.
[0010] According to one aspect of the present invention, a speech recognition device is provided, comprising: a collection unit for collecting monitored speech information to be recognized, wherein the speech information to be recognized includes at least original audio information to be recognized; a calling unit for calling a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss being obtained based on a noise recognition result, the noise recognition result being obtained by noise recognition of a first noise information sample in a training sample set, the auxiliary loss being obtained based on an information mapping result, and the information mapping result being obtained by information mapping of at least a second noise information sample in a training sample set; a mapping unit for using a network layer in the speech recognition model determined by at least the auxiliary loss to perform information mapping on the speech information to be recognized, thereby obtaining a target information mapping result; a first recognition unit for using a network layer in the speech recognition model determined by at least the original loss to perform noise recognition on the target information mapping result, thereby identifying interference information, wherein the interference information is information that interferes with the original audio information; and a second recognition unit for removing the interference information from the speech information to be recognized, thereby recognizing the original audio information.
[0011] According to one aspect of the present invention, a speech recognition model determination apparatus is provided, comprising: an acquisition unit, configured to acquire a first noise information sample and a second noise information sample from a training sample set; a recognition unit, configured to determine the original loss of the initial speech recognition model based on a noise recognition result obtained by performing speech recognition on the first noise information sample using an initial speech recognition model, and to determine an auxiliary loss of the initial speech recognition model based at least on an information mapping result obtained by performing information mapping on the second noise information sample using the initial speech recognition model; and a training unit, configured to train the initial speech recognition model based on the original loss and the auxiliary loss to obtain a speech recognition model, wherein a network layer in the speech recognition model, at least determined by the auxiliary loss, is used to perform information mapping on the speech information to be recognized to obtain a target information mapping result, the speech information to be recognized contains at least the original audio information to be recognized, and a network layer in the speech recognition model, at least determined by the original loss, is used to perform noise recognition on the target information mapping result to identify interference information, the interference information being information that interferes with the original audio information, and is used to identify the original audio information from the speech information to be recognized.
[0012] According to one aspect of the present invention, a speech recognition device is provided, comprising: a collection unit for collecting speech information to be recognized sent to a client, wherein the speech information to be recognized includes at least original audio information to be recognized; a mapping unit for mapping the speech information to be recognized using a network layer in a speech recognition model determined by at least an auxiliary loss, to obtain a target information mapping result, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss is obtained based on a noise recognition result, the noise recognition result is obtained by noise recognition of a first noise information sample in a training sample set, the auxiliary loss is obtained based on the information mapping result, and the information mapping result is obtained by information mapping of at least a second noise information sample in a training sample set; a first recognition unit for using a network layer in a speech recognition model determined by at least an original loss to perform noise recognition on the target information mapping result, and to identify interference information, wherein the interference information is information that interferes with the original audio information; a second recognition unit for removing interference information from the speech information to be recognized and identifying the original audio information; and an activation unit for activating the client based on the original audio information.
[0013] According to one aspect of the present invention, a speech recognition device is provided, comprising: an input unit for inputting monitored speech information to be recognized into a virtual reality (VR) device or an augmented reality (AR) device, wherein the speech information to be recognized includes at least original audio information to be recognized; and a calling unit for calling a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss being obtained based on a noise recognition result, the noise recognition result being obtained by noise recognition of a first noise information sample in a training sample set, and the auxiliary loss being obtained based on an information mapping result, the information mapping result being... The system employs a mapping unit to map information from the second noise information sample in the training sample set; a mapping unit to map information from the speech information to be recognized using at least a network layer in the speech recognition model determined by the auxiliary loss, thereby obtaining the target information mapping result; a first recognition unit to identify noise from the target information mapping result using at least a network layer in the speech recognition model determined by the original loss, thereby identifying interference information, wherein the interference information is information that interferes with the original audio information; a second recognition unit to remove interference information from the speech information to be recognized, thereby identifying the original audio information; and an activation unit to activate the VR or AR device using the original audio information.
[0014] According to one aspect of the present invention, a speech recognition device is provided, comprising: a first calling unit, configured to acquire monitored speech information to be recognized by calling a first interface, wherein the first interface includes a first parameter, the parameter value of the first parameter being the speech information to be recognized, and the speech information to be recognized including at least original audio information to be recognized; a second calling unit, configured to call a speech recognition model, wherein the speech recognition model is obtained by mapping the initial speech information based on the original loss and the auxiliary loss; a mapping unit, configured to use a network layer in the speech recognition model, at least determined by the auxiliary loss, to perform information mapping on the speech information to be recognized, thereby obtaining a target information mapping result; a first recognition unit, configured to use a network layer in the speech recognition model, at least determined by the original loss, to perform noise recognition on the target information mapping result, thereby recognizing interference information, wherein the interference information is information that interferes with the original audio information; a second recognition unit, configured to remove the interference information from the speech information to be recognized, thereby recognizing the original audio information; and a third calling unit, configured to call a second interface to output the original audio information, wherein the second interface includes a second parameter, the parameter value of the second parameter being the original audio information.
[0015] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored program, wherein, when the program processor is running, it controls the device where the computer storage medium is located to execute a speech recognition method.
[0016] According to another aspect of the embodiments of this application, a processor is also provided, which is used to run a program, wherein the program is executed by the processor to perform a speech recognition method.
[0017] In this embodiment of the invention, the speech recognition model is trained based on the original loss and the auxiliary loss. Based on this, after obtaining the speech information to be recognized, the speech information to be recognized can be recognized once using the network layer determined by the auxiliary loss included in the speech recognition model. Then, the target information mapping result obtained from the first recognition can be recognized a second time using the network layer determined by the original loss included in the speech recognition model to further identify interference information. Thus, interference information can be removed from the speech information to be recognized, and the original audio information can be recognized. That is, in this application, by performing two recognitions on the speech information to be recognized, the purpose of removing interference information in the speech information to be recognized is achieved, thereby improving the accuracy of speech information recognition and solving the technical problem of low accuracy in speech information recognition. Attached Figure Description
[0018] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0019] Figure 1 This is a schematic diagram of the hardware environment of a virtual reality device according to an embodiment of this application;
[0020] Figure 2 This is a structural block diagram of the computing environment for an optional speech recognition method according to an embodiment of this application;
[0021] Figure 3 This is a flowchart of a speech recognition method according to an embodiment of this application;
[0022] Figure 4 This is a flowchart of another speech recognition method according to an embodiment of this application;
[0023] Figure 5 This is a flowchart of another speech recognition method according to an embodiment of this application;
[0024] Figure 6 This is a schematic diagram of a speech recognition method according to an embodiment of this application;
[0025] Figure 7 This is a flowchart of another speech recognition method according to an embodiment of this application;
[0026] Figure 8 This is a schematic diagram illustrating the training of an initial speech recognition model according to an embodiment of this application;
[0027] Figure 9 This is a schematic diagram of a speech recognition device according to an embodiment of this application;
[0028] Figure 10 This is a schematic diagram of a speech recognition model determination device according to an embodiment of this application;
[0029] Figure 11 This is a schematic diagram of another speech recognition device according to an embodiment of this application;
[0030] Figure 12 This is a schematic diagram of another speech recognition device according to an embodiment of this application;
[0031] Figure 13 This is a schematic diagram of another speech recognition device according to an embodiment of this application;
[0032] Figure 14 This is a block diagram of a computer terminal according to an embodiment of this application. Detailed Implementation
[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0036] The HuBERT speech recognition model is a model obtained by training an initial speech recognition model based on the original loss and auxiliary loss.
[0037] The original loss (HuBERT loss) is the loss of the initial speech recognition model determined based on the noise recognition results obtained from the first noise information sample.
[0038] The auxiliary loss (deHuBERT loss) is the loss of the initial speech recognition model determined based on the information mapping result obtained by mapping the second noise information sample.
[0039] Interference information refers to information that interferes with the original audio information.
[0040] Cross-correlation loss The difference between the correlation between multiple input speech information in the initial speech recognition model and the corresponding target correlation;
[0041] Autocorrelation loss The difference between the correlation between different speech information and the corresponding target correlation in the output noise recognition results of the initial speech recognition model.
[0042] Example 1
[0043] According to an embodiment of this application, a speech recognition method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0044] Figure 1 This is a schematic diagram of the hardware environment of a virtual reality device according to an embodiment of the speech recognition method of this application. Figure 1 As shown, the virtual reality device 104 is connected to the terminal 106, and the terminal 106 is connected to the server 102 via a network. The virtual reality device 104 is not limited to virtual reality headsets, virtual reality glasses, virtual reality all-in-one machines, etc. The terminal 106 is not limited to PCs, mobile phones, tablets, etc. The server 102 can be a server corresponding to a media file operator. The network includes, but is not limited to, wide area networks, metropolitan area networks, or local area networks.
[0045] Optionally, the virtual reality device 104 in this embodiment includes a memory, a processor, and a transmission device. The memory stores an application program that can perform the following actions: acquiring monitored speech information to be recognized, wherein the speech information to be recognized includes at least the original audio information to be recognized; calling a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss being obtained based on noise recognition results, the noise recognition results being obtained by noise recognition of a first noise information sample in the training sample set, the auxiliary loss being obtained based on information mapping results, and the information mapping results being obtained by information mapping of at least a second noise information sample in the training sample set; using a network layer in the speech recognition model determined at least by the auxiliary loss to perform information mapping on the speech information to be recognized, obtaining a target information mapping result; using a network layer in the speech recognition model determined at least by the original loss to perform noise recognition on the target information mapping result, identifying interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized, and recognizing the original audio information, thereby solving the technical problem of low accuracy in speech information recognition and achieving the purpose of removing interference information from the speech information to be recognized.
[0046] The terminal in this embodiment can be used to perform the following actions: inputting monitored speech information to be recognized into a virtual reality (VR) device or an augmented reality (AR) device, wherein the speech information to be recognized includes at least the original audio information to be recognized; invoking a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, wherein the original loss is obtained based on the noise recognition result, the noise recognition result is obtained by noise recognition of a first noise information sample in the training sample set, the auxiliary loss is obtained based on the information mapping result, and the information mapping result is obtained by information mapping of at least a second noise information sample in the training sample set; using a network layer in the speech recognition model determined by at least the auxiliary loss to perform information mapping on the speech information to be recognized to obtain a target information mapping result; using a network layer in the speech recognition model determined by at least the original loss to perform noise recognition on the target information mapping result to identify interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized to identify the original audio information; and activating the VR device or AR device using the original audio information.
[0047] Optionally, the eye-tracking HMD (Head-Mounted Display) and eye-tracking module in the virtual reality device 104 of this embodiment function the same as in the embodiments described above. That is, the screen in the HMD is used to display real-time images, and the eye-tracking module in the HMD is used to acquire the real-time movement path of the user's eyes. The terminal in this embodiment acquires the user's position and movement information in real three-dimensional space through the tracking system, and calculates the three-dimensional coordinates of the user's head in virtual three-dimensional space, as well as the user's field of vision orientation in virtual three-dimensional space.
[0048] Figure 1 The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned AR / VR device (or mobile device), but also as an exemplary block diagram of the aforementioned server. In one optional embodiment, Figure 2 The use of the above is illustrated in a block diagram. Figure 1 The AR / VR device (or mobile device) shown is an example of a computing node in computing environment 201. Figure 2 This is a structural block diagram of the computing environment for a speech recognition method according to an embodiment of this application, such as... Figure 2As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (shown as 210-1, 210-2, ..., in the diagram). Each computing node contains local processing and memory resources, and end user 202 can remotely run applications or store data within computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within computing environment 201, representing services "A", "D", "E", and "H", respectively.
[0049] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).
[0050] The services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.
[0051] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, such as Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). Each Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. The proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be Pods similar to Pods.
[0052] During operation, executing a user request from end user 202 may require calling one or more services in computing environment 201, and executing one or more functions of one service may require calling one or more functions of another service. For example... Figure 2As shown, service "A" 220-1 receives user requests from terminal user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to perform one or more functions.
[0053] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0054] Under the aforementioned operating environment, this application provides the following: Figure 3 The speech recognition method shown is illustrated. It should be noted that the speech recognition method in this embodiment can be derived from... Figure 1 The mobile terminal in the illustrated embodiment is executed. Figure 3 This is a flowchart of the speech recognition method according to Embodiment 1 of this application. Figure 3 As shown, the method may include the following steps:
[0055] Step S301: Collect the monitored speech information to be recognized, wherein the speech information to be recognized contains at least the original audio information to be recognized.
[0056] In step S301 above, after responding to the voice information acquisition command, the monitored voice information to be identified can be acquired. The voice information to be identified contains at least the original audio information to be identified, and the voice information to be identified contains noise. The voice information to be identified can be any sound in the real world environment.
[0057] In this embodiment, the noise contained in the speech information to be identified can be real-world sounds such as whispers, writing sounds, typewriter sounds, telephone sounds, laughter, computer keyboard sounds, and printer sounds.
[0058] Step S302: Invoke the speech recognition model, which is obtained by training an initial speech recognition model based on the original loss and auxiliary loss.
[0059] In step S302 above, after acquiring the speech information, a speech recognition model can be invoked. The speech recognition model is obtained by training an initial speech recognition model based on the original loss and auxiliary loss. The original loss is obtained based on the noise recognition result, which is obtained by identifying noise in the first noise information sample in the training sample set. The auxiliary loss is obtained based on the information mapping result, which is obtained by mapping at least the second noise information sample in the training sample set. The information mapping is based on a pre-set information mapping table, which includes the mapping relationship between noise information and loss parameters. Based on this, noise information that matches the noise information contained in the second noise information sample can be found in the information mapping table based on the second noise information sample. The loss parameter corresponding to the found noise information is then determined as the information mapping result of the second noise information sample, and the auxiliary loss of the initial speech recognition model is determined based on the information mapping result.
[0060] In this embodiment, the training sample set may contain multiple noise information samples, which are speech information containing noise. The noise included in the first noise information sample in the training sample set may be a single type of noise, such as car noise, subway noise, and traffic noise. The noise included in the second noise information sample in the training sample set may be mixed noise, such as airport / station noise, coffee shop / air conditioner noise, etc.
[0061] Optionally, the initial speech recognition model can be used to identify noise in the first noise information sample in the training sample set to obtain the noise recognition result. The initial speech recognition model can be the HuBERT model, which uses the Bidirectional Encoder Representation from Transformers (BERT) model in Natural Language Processing (NLP) to perform self-supervised speech recognition learning, allowing the encoder to discover good high-level latent representations of both acoustic and linguistic information from continuous speech signals.
[0062] Optionally, after obtaining the noise identification result, the original loss can be obtained from the noise identification result. The original loss can be used to indicate the degree of deviation between the HuBERT model's identification result and the true result. That is, the original loss can be used to indicate the HuBERT model's accuracy in identifying a single noise. The original loss can be called the HuBERT loss.
[0063] Optionally, the initial speech recognition model may also include an information mapping table, which may include the correspondence between noise information and loss parameters. Based on this, the initial speech recognition model can be used to perform information mapping on the second noise information samples in the training sample set to obtain the information mapping result. Then, the auxiliary loss can be derived from the information mapping result, wherein the auxiliary loss can be used to indicate the accuracy of the initial speech recognition model in recognizing mixed noise.
[0064] Optionally, after determining the original loss and auxiliary loss, the initial speech recognition model can be trained based on the original loss and auxiliary loss to obtain a speech recognition model. This speech recognition model is obtained by training the initial speech recognition model based on noise information samples. Therefore, this speech recognition model has high accuracy in recognizing speech information containing noise.
[0065] Step S303: Using at least one network layer in the speech recognition model determined by the auxiliary loss, information mapping is performed on the speech information to be recognized to obtain the target information mapping result.
[0066] In step S303 above, the speech recognition model includes at least a network layer determined by the auxiliary loss and a network layer determined by the original loss. Based on this, the speech information to be recognized can be mapped using the network layer determined by the auxiliary loss in the speech recognition model to obtain the target information mapping result.
[0067] In this embodiment, the speech information to be recognized can be input into a network layer in the speech recognition model that is determined by at least the auxiliary loss. The network layer determined by the auxiliary loss is used to perform information mapping on the speech information to be recognized to obtain the target information mapping result.
[0068] Step S304: Using at least one network layer in the speech recognition model determined by the original loss, noise recognition is performed on the target information mapping result to identify interference information, wherein the interference information is information that interferes with the original audio information.
[0069] In step S304 above, since the speech recognition model also includes a network layer determined by the original loss, noise recognition can be further performed on the target mapping result determined in step S303 based on the network layer determined by the original loss to identify interference information, wherein the interference information is information that interferes with the original audio information.
[0070] In this embodiment, the target mapping result determined in step S303 can be input to the network layer determined by the original loss. The network layer can identify noise in the target mapping result and identify interference information, wherein the interference information can be high noise information that interferes with the original audio information.
[0071] Step S305: Remove interference information from the speech information to be recognized and identify the original audio information.
[0072] In step S305 above, after determining the interference information in the original audio information to be identified, the interference information can be removed from the speech information to be identified. Then, the speech information to be identified after removing the interference information is identified to obtain the original audio information.
[0073] In this embodiment, interference information in the speech information to be identified can be extracted from the speech information to be identified, and then the speech information to be identified is identified to obtain the original audio information.
[0074] Based on the scheme disclosed in steps S301 to S305 of the above embodiments, the speech recognition model is obtained by training an initial speech recognition model based on the original loss and auxiliary loss. Based on this, after obtaining the speech information to be recognized, the speech information to be recognized can be recognized once using the network layer determined by the auxiliary loss included in the speech recognition model. Then, the target information mapping result obtained from the first recognition can be recognized a second time using the network layer determined by the original loss included in the speech recognition model to further identify interference information. Thus, interference information can be removed from the speech information to be recognized, and the original audio information can be recognized. That is, in this application, by performing two recognitions on the speech information to be recognized, the purpose of removing interference information in the speech information to be recognized is achieved, thereby improving the accuracy of speech information recognition and solving the technical problem of low accuracy of speech information recognition.
[0075] The method described in this embodiment will be further described below.
[0076] As an optional implementation, the method further includes: performing information mapping on the first noise information sample and the second noise information sample respectively based on the information mapping model in the initial speech recognition model; determining the cross-correlation loss of the initial speech recognition model based on the mapped first noise information sample and the mapped second noise information sample, wherein the auxiliary loss includes the cross-correlation loss, which is used to represent the difference between the degree of correlation between multiple input speech information of the initial speech recognition model and the corresponding target correlation degree.
[0077] In this embodiment, the initial speech recognition model includes an information mapping model, which is used to map noise information samples to obtain information mapping results. Specifically, a first noise information sample can be input into the information mapping model to obtain the information mapping result of the first noise information sample. Similarly, a second noise information sample can be input into the information mapping model to obtain the information mapping result of the second noise information sample.
[0078] Optionally, after obtaining the information mapping results of the first noise information sample and the second noise information sample, the cross-correlation loss of the initial speech recognition model can be determined based on the mapped first noise information sample and the mapped second noise information sample. The cross-correlation loss is used to indicate the difference between the degree of correlation between multiple input speech information of the initial speech recognition model and the corresponding target correlation.
[0079] For example, suppose the first noise information sample can be represented by x, the second noise information sample can be represented by x, the first noise information sample obtained after information mapping can be represented by y, and the second noise information sample obtained after information mapping can be represented by y. Therefore, based on this, the cross-correlation loss matrix of the initial speech recognition model can be determined using the following formula:
[0080]
[0081] Where n represents the number of frames of the noise information sample used, i and j are the spatial positions between frames, and C∈[-1,1].
[0082] As an optional implementation, the cross-correlation loss of the initial speech recognition model is determined based on the mapped first noise information sample and the mapped second noise information sample, including: performing cross-correlation processing on the mapped first noise information sample and the mapped second noise information sample to obtain an initial cross-correlation result, wherein the initial cross-correlation result is used to represent the degree of correlation between the mapped first noise information sample and the mapped second noise information sample; and determining the cross-correlation loss based on the initial cross-correlation result and the target cross-correlation result, wherein the target cross-correlation result is used to represent the target degree of correlation between the mapped first noise information sample and the mapped second noise information sample.
[0083] In this embodiment, a pre-set initial cross-correlation algorithm can be used to perform cross-correlation processing on the mapped first noise information sample and the mapped second noise information sample to obtain an initial cross-correlation result, which can be called Empirical Cross-correlation. Similarly, a target cross-correlation algorithm can be used to perform target cross-correlation processing on the mapped first noise information sample and the second noise information sample to obtain a target cross-correlation result, which can be called Target Cross-correlation.
[0084] Optionally, after obtaining the initial cross-correlation result (Empirical Cross-correlation) and the target cross-correlation result (Target Cross-correlation), the cross-correlation loss can be determined using the following formula, which represents the degree of target correlation between the mapped first noise information sample and the second noise information sample.
[0085]
[0086] Here, λ is a penalty coefficient used to maintain the balance between the first invariance term of the loss and the second disentangling term of the loss.
[0087] Optionally, in order to determine whether the cross-correlation loss can reduce noise to obtain invariant features, the cross-correlation loss can be compared with the infoNCE loss, and the comparison results are shown in the following formula (3).
[0088]
[0089] From the above formula (3), it can be seen that the cross-correlation loss of the initial speech recognition model is very similar to the infoNCE loss. Based on this, the two distorted embedded speech contents can be maximized by making the two-dimensional feature components correlated, so as to reduce noise.
[0090] As an optional implementation, information mapping is performed on the first noise information sample and the second noise information sample based on the information mapping model in the initial speech recognition model, including: linearly mapping the first noise information sample to a first matrix and linearly mapping the second noise information sample to a second matrix based on the information mapping model.
[0091] In this embodiment, a first noise information sample can be input into an information mapping model for information mapping to obtain a first matrix, which can be a d-dimensional square matrix. Similarly, a second noise information sample can be input into an information mapping model for information mapping to obtain a second matrix, which can also be a d-dimensional square matrix.
[0092] As an optional implementation, the cross-correlation loss of the initial speech recognition model is determined based on the mapped first noise information sample and the mapped second noise information sample, including: obtaining the cross-correlation matrix between the first matrix and the second matrix, wherein the cross-correlation matrix is associated with the identity matrix; and determining the cross-correlation loss based on the cross-correlation matrix.
[0093] In this embodiment, after obtaining the first matrix and the second matrix, the first matrix and the second matrix can be cross-correlated to obtain the cross-correlation matrix between the first matrix and the second matrix. Since the first matrix and the second matrix are both d-dimensional square matrices, the cross-correlation matrix obtained after cross-correlation of the first matrix and the second matrix can also be a d-dimensional square matrix.
[0094] Optionally, after obtaining the cross-correlation matrix, the cross-correlation matrix can be converted into an identity matrix, and then the cross-correlation loss can be determined based on the identity matrix. This cross-correlation loss is used to indicate the difference between the degree of correlation between multiple input speech information of the initial speech recognition model and the corresponding target correlation.
[0095] As an optional implementation, the method further includes: enhancing the information of the first noise information sample and enhancing the information of the second noise information sample; and mapping the information of the first noise information sample and the second noise information sample based on the information mapping model, including: mapping the enhanced first noise information sample and the enhanced second noise information sample based on the information mapping model.
[0096] In this embodiment, before mapping the first and second noise information samples, the first and second noise information samples can be input into a feedforward neural network for information enhancement, resulting in enhanced first and second noise information samples. The feedforward neural network can be a convolutional neural network (CNN), and the enhanced first noise information sample can be denoted as... The enhanced second noise information samples are denoted as x1, x2, and x3.
[0097] Optionally, after obtaining the enhanced first noise information sample and the enhanced second noise information sample, the enhanced first noise information sample and the enhanced second noise information sample can be input into the information mapping model for information mapping to obtain the first matrix and the second matrix.
[0098] As an optional implementation, the method further includes: obtaining a noise recognition result obtained by the initial speech recognition model performing noise recognition on a first noise information sample; performing information mapping on the noise recognition result based on the information mapping model in the initial speech recognition model; and determining the autocorrelation loss of the initial speech recognition model based on the mapped noise recognition result, wherein the auxiliary loss includes the autocorrelation loss, which is used to represent the difference between the degree of correlation between different speech information in the output noise recognition result of the initial speech recognition model and the corresponding target correlation.
[0099] In this embodiment, as described above, the first noise information sample is input into the initial speech recognition model. Based on this, the noise recognition result obtained by the initial speech recognition model in performing noise recognition on the first noise information sample can be obtained. The noise recognition result is used to represent the noise information in the first noise information sample.
[0100] Optionally, after obtaining the noise recognition result in the first noise information sample, the noise recognition result can be mapped based on the information mapping model in the initial speech recognition model to obtain the mapped noise recognition result. Then, the autocorrelation loss of the initial speech recognition model can also be calculated using the above formulas (1) and (2). The autocorrelation loss is also an auxiliary loss. The autocorrelation loss is used to represent the difference between the degree of correlation between different speech information in the output noise recognition result of the initial speech recognition model and the degree of correlation between the corresponding target.
[0101] As an optional implementation, determining the autocorrelation loss of the initial speech recognition model based on the mapped noise recognition result includes: performing autocorrelation processing on the mapped noise recognition result to obtain an initial autocorrelation result, wherein the initial autocorrelation result is used to represent the degree of correlation between different noise information in the mapped noise recognition result; and determining the autocorrelation loss of the initial speech recognition model based on the initial autocorrelation result and the target autocorrelation result, wherein the auxiliary loss includes the autocorrelation loss, and the target autocorrelation result is used to represent the target degree of correlation between different noise information in the mapped noise recognition result.
[0102] In this embodiment, a preset autocorrelation algorithm can be used to perform autocorrelation processing on the mapped noise identification results to obtain an initial autocorrelation result. This initial autocorrelation result can be called Empirical Self-correlation, and it is used to represent the degree of correlation between different noise information in the mapped noise identification results.
[0103] Optionally, a preset target autocorrelation algorithm can be used to process the mapped noise recognition results to obtain the target autocorrelation result, which can be called the target self-correlation. Then, based on the initial autocorrelation result and the target autocorrelation result, the autocorrelation loss of the initial speech recognition model can be determined.
[0104] Optionally, after obtaining the autocorrelation loss and cross-correlation loss of the initial speech recognition model, the complete optimization loss of the initial speech recognition model can be further determined by the following formula (4). After determining the complete optimization loss of the initial speech recognition model, the initial speech recognition model can be further optimized based on the complete optimization loss to obtain the speech recognition model. This can enhance the recognition accuracy of the speech recognition model for the speech information to be recognized.
[0105]
[0106] in, It can be used to represent HuberT loss. It can be used to represent cross-correlation loss. For autocorrelation loss, α and β are loss parameters. The values of the loss parameters can be preset. For example, α and β can be preset to 0.5.
[0107] As an optional implementation, the method further includes: performing information enhancement on the first noise information sample; obtaining a noise recognition result obtained by the initial speech recognition model performing noise recognition on the first noise information sample, including: obtaining the noise recognition result obtained by the initial speech recognition model performing noise recognition on the enhanced first noise information sample.
[0108] In this embodiment, as described above, the first noise information sample can be input into the CNN encoder for information enhancement to obtain the enhanced first noise information sample. Then, the initial speech recognition model is used to perform noise recognition on the enhanced first noise information sample to obtain the noise recognition result.
[0109] As an optional implementation, the method further includes: randomly selecting a first noise information sample and a second noise information sample from the training sample set, wherein the training sample set includes multiple noise information samples with a signal-to-noise ratio within a signal-to-noise ratio threshold range, wherein the signal-to-noise ratio threshold range can be 0dB to 25dB, and of course, the signal-to-noise ratio threshold range can also be other numerical ranges, which are not specifically limited here.
[0110] Optionally, multiple noise information samples in the training sample set have sound levels. Based on this, a first noise information sample and a second noise information sample can be obtained from the training sample set according to the sound level, wherein the sound level of the first noise information sample is greater than the sound level of the second noise information sample.
[0111] In the above steps, the speech recognition model is trained based on the original loss and the auxiliary loss. Based on this, after obtaining the speech information to be recognized, the network layer included in the speech recognition model, which is determined by the auxiliary loss, can be used to perform a first recognition of the speech information to be recognized. Then, the network layer included in the speech recognition model, which is determined by the original loss, can be used to perform a second recognition of the target information mapping result obtained from the first recognition, further identifying interference information. Thus, interference information can be removed from the speech information to be recognized, and the original audio information can be recognized. That is, in this application, through two recognitions, the purpose of removing interference information in the speech information to be recognized is achieved, improving the accuracy of speech information recognition and solving the technical problem of low accuracy in speech information recognition.
[0112] The method for determining the speech recognition model in the embodiments of this application will be described below.
[0113] Figure 4 This is a flowchart of a method for determining a speech recognition model according to an embodiment of this application. Figure 4 As shown, the method may include the following steps:
[0114] Step S401: Obtain the first noise information sample and the second noise information sample from the training sample set.
[0115] In step S401 above, the training sample set includes multiple noise information samples, and each noise information sample has a sound level. Based on this, a first noise information sample and a second noise information sample can be randomly selected from the training sample set, or the first noise information sample and the second noise information sample can be selected from the training sample set according to the sound level, wherein the sound level of the first noise information sample is greater than the sound level of the second noise information sample.
[0116] Step S402: Based on the noise recognition result obtained by performing speech recognition on the first noise information sample using the initial speech recognition model, determine the original loss of the initial speech recognition model, and at least based on the information mapping result obtained by performing information mapping on the second noise information sample using the initial speech recognition model, determine the auxiliary loss of the initial speech recognition model.
[0117] In step S402 above, after obtaining the first noise information sample and the second noise information sample, the first noise information sample can be input into the initial speech recognition model for noise recognition to obtain the noise recognition result of the first noise information sample, and the original loss of the initial speech recognition model can be determined based on the noise recognition result; in addition, the second noise information sample can be input into the initial speech recognition model for information mapping to determine the auxiliary loss of the initial speech recognition model.
[0118] Step S403: Train the initial speech recognition model based on the original loss and auxiliary loss to obtain a speech recognition model. The speech recognition model has at least one network layer determined by the auxiliary loss, which is used to perform information mapping on the speech information to be recognized to obtain the target information mapping result. The speech information to be recognized contains at least the original audio information to be recognized. The speech recognition model has at least one network layer determined by the original loss, which is used to identify noise on the target information mapping result and identify interference information. The interference information is information that interferes with the original audio information and is used to identify the original audio information from the speech information to be recognized.
[0119] In this embodiment, after determining the original loss and auxiliary loss of the initial speech recognition model, the initial speech recognition model can be trained based on the original loss and auxiliary loss to obtain a speech recognition model. Then, the speech recognition model is used to recognize the speech information to be recognized to obtain the original audio information.
[0120] In steps S401 to S403 above, noise recognition can be performed on the first noise information sample and the second noise information sample based on the initial speech recognition model to obtain the original loss and auxiliary loss of the initial speech recognition model. Then, the initial speech recognition model is trained based on the original loss and auxiliary loss to obtain the speech recognition model. After obtaining the speech recognition model, the speech information to be recognized can be recognized using the speech recognition model to obtain the original audio information. Since the speech recognition model is obtained by training the initial speech recognition model based on the original loss and auxiliary loss, after obtaining the speech information to be recognized, the network layer included in the speech recognition model, which is determined by the auxiliary loss, can be used to perform a first recognition on the speech information to be recognized. Then, the network layer included in the speech recognition model, which is determined by the original loss, can be used to perform a second recognition on the target information mapping result obtained from the first recognition to further identify the interference information. Thus, the interference information can be removed from the speech information to be recognized to identify the original audio information. That is, in this application, the purpose of removing the interference information in the speech information to be recognized is achieved through two recognitions, thereby improving the technical effect of speech information recognition accuracy and solving the technical problem of low speech information recognition accuracy.
[0121] The following section introduces methods for recognizing voice information using a client-side application.
[0122] Figure 5 This is a flowchart of a speech recognition method according to an embodiment of this application. Figure 5 As shown, the method may include the following steps:
[0123] Step S501: Collect the voice information to be recognized sent to the client, wherein the voice information to be recognized contains at least the original audio information to be recognized.
[0124] In step S501 above, after receiving the voice acquisition command, the voice information to be recognized sent to the client can be acquired. The voice information to be recognized contains at least the original audio information to be recognized and contains noise.
[0125] Step S502: Using at least one network layer in the speech recognition model determined by the auxiliary loss, information mapping is performed on the speech information to be recognized to obtain the target information mapping result.
[0126] In step S502 above, the speech recognition model is obtained by training an initial speech recognition model based on the original loss and the auxiliary loss. The original loss is obtained based on the noise recognition result, which is obtained by noise recognition of the first noise information sample in the training sample set. The auxiliary loss is obtained based on the information mapping result, which is obtained by information mapping of at least the second noise information sample in the training sample set. Based on this, the network layer in the speech recognition model that is determined by at least the auxiliary loss can be used to perform information mapping on the speech information to be recognized, and obtain the target information mapping result.
[0127] Step S503: Using at least one network layer in the speech recognition model determined by the original loss, noise recognition is performed on the target information mapping result to identify interference information.
[0128] In step S503 above, after determining the target information mapping result, the network layer in the speech recognition model, which is at least determined by the original loss, can be used to identify noise in the target information mapping result to obtain interference information, which is noise that interferes with the speech information to be recognized.
[0129] Step S504: Remove interference information from the speech information to be recognized and identify the original audio information.
[0130] In step S504 above, after determining the interference information from the speech information to be recognized, the interference information can be removed from the speech information to be recognized, and the speech information to be recognized after removing the interference information can be used as the original audio information.
[0131] Step S505: Activate the client based on the original audio information.
[0132] In step S505 above, after the original audio information is determined, the client can be activated based on the original audio information.
[0133] In steps S501 to S505 above, the speech information to be recognized sent to the client can be collected, and the speech information to be recognized can be mapped using at least one network layer in the speech recognition model determined by the auxiliary loss to obtain the target information mapping result. Then, the target information mapping result can be noise-identified using at least one network layer in the speech recognition model determined by the original loss to identify interference information. This interference information can then be removed from the speech information to be recognized, allowing the original audio information to be identified, and the client can be activated based on the original audio information. In other words, in this embodiment, after collecting the speech information to be recognized, the speech recognition model can be used to identify interference information in the speech information to be recognized. Then, the identified interference information is removed from the speech information to be recognized, allowing the original audio information to be identified. Since the speech recognition model is trained based on the original loss and the auxiliary loss, using this speech recognition model to recognize the speech information to be recognized can more accurately identify noise in the speech information to be recognized, achieving the purpose of removing interference information from the speech information to be recognized, improving the accuracy of speech information recognition, and thus solving the technical problem of low accuracy in speech information recognition.
[0134] The following describes the method for speech recognition using virtual reality (VR) devices or augmented reality (AR) devices in the embodiments of this application.
[0135] Figure 6 This is a flowchart of a speech recognition method according to an embodiment of this application. Figure 6 As shown, the method may include the following steps:
[0136] Step S601: Input the monitored voice information to be recognized into a virtual reality (VR) device or an augmented reality (AR) device, wherein the voice information to be recognized contains at least the original audio information to be recognized.
[0137] In step S601 above, after the voice information to be recognized is detected, the detected voice information to be recognized can be input to a virtual reality (VR) device or an augmented reality (AR) device. After receiving the voice information to be recognized, the virtual reality (VR) device or the augmented reality (AR) device can recognize the voice information to be recognized.
[0138] Step S602: Call the speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on the original loss and the auxiliary loss. The original loss is obtained based on the noise recognition result, which is obtained by noise recognition of the first noise information sample in the training sample set. The auxiliary loss is obtained based on the information mapping result, which is obtained by information mapping of at least the second noise information sample in the training sample set.
[0139] In step S602 above, after receiving the speech information to be recognized, the virtual reality (VR) device or augmented reality (AR) device can call the speech recognition model to recognize the speech information. The speech recognition model is obtained by training an initial speech recognition model based on the original loss and the auxiliary loss. The original loss is obtained based on the noise recognition result, which is obtained by noise recognition of the first noise information sample in the training sample set. The auxiliary loss is obtained based on the information mapping result, which is obtained by information mapping of at least the second noise information sample in the training sample set. Based on this, the speech recognition model has a good recognition effect on the noise in the speech information to be recognized.
[0140] Step S603: Using at least one network layer in the speech recognition model determined by auxiliary loss, information mapping is performed on the speech information to be recognized to obtain the target information mapping result.
[0141] In step S603 above, the speech recognition model includes a network layer for determining auxiliary loss. Based on this, virtual reality (VR) devices or augmented reality (AR) devices can use the network layer in the speech recognition model, which is at least determined by auxiliary loss, to perform information mapping on the speech information to be recognized, and obtain the target information mapping result.
[0142] Step S604: Remove interference information from the speech information to be recognized and identify the original audio information.
[0143] In step S604 above, after identifying the interference information, the virtual reality (VR) device or augmented reality (AR) device can remove the interference information from the speech information to be identified in order to identify the original audio information in the speech information to be identified.
[0144] Step S605: Activate the VR or AR device using the original audio information.
[0145] In step S605 above, after the original audio information is identified, the original audio information can be used to activate the VR device or AR device.
[0146] In steps S601 to S605 above, the monitored speech information to be recognized can be input into a virtual reality (VR) device or an augmented reality (AR) device, and a speech recognition model can be invoked. Using at least one network layer in the speech recognition model determined by the auxiliary loss, information mapping is performed on the speech information to be recognized to obtain a target information mapping result. Then, using at least one network layer in the speech recognition model determined by the original loss, noise recognition is performed on the target information mapping result to identify interference information. Afterward, the interference information is removed from the speech information to be recognized, and the original audio information is identified. In other words, in this embodiment, the speech recognition model can be used to identify interference information in the speech information to be recognized. Then, the identified interference information is removed from the speech information to be recognized to identify the original audio information. Since the speech recognition model is trained based on the original loss and the auxiliary loss, using this speech recognition model to recognize the speech information to be recognized can more accurately identify the noise in the speech information to be recognized, achieving the purpose of removing interference information from the speech information to be recognized, improving the accuracy of speech information recognition, and thus solving the technical problem of low accuracy in speech information recognition.
[0147] Another speech recognition method provided in the embodiments of this application is described below.
[0148] Figure 7 This is a flowchart of a speech recognition method according to an embodiment of this application. Figure 7 As shown, the method may include the following steps:
[0149] Step S701: Obtain the monitored speech information to be recognized by calling the first interface. The first interface includes a first parameter, the value of which is the speech information to be recognized. The speech information to be recognized contains at least the original audio information to be recognized.
[0150] The first interface can be an interface for data interaction between the server and the user. The user can obtain the monitored speech information to be recognized by calling the first interface. The speech information to be recognized is used as a first parameter of the first interface, thus achieving the purpose of obtaining the speech information to be recognized.
[0151] Step S702: Call the speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on the original loss and the auxiliary loss. The original loss is obtained based on the noise recognition result, which is obtained by noise recognition of the first noise information sample in the training sample set. The auxiliary loss is obtained based on the information mapping result, which is obtained by information mapping of at least the second noise information sample in the training sample set.
[0152] Step S703: Using at least one network layer in the speech recognition model determined by the auxiliary loss, information mapping is performed on the speech information to be recognized to obtain the target information mapping result.
[0153] Step S704: Using at least one network layer in the speech recognition model determined by the original loss, noise recognition is performed on the target information mapping result to identify interference information, wherein the interference information is information that interferes with the original audio information.
[0154] Step S705: Remove interference information from the speech information to be recognized and identify the original audio information.
[0155] Step S706: Output the original audio information by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter is the original audio information.
[0156] The second interface can be an interface for data interaction between the server and the user. The server can send the execution result to the client, so that the client can output the original audio information to the second interface as a parameter of the second interface, thus achieving the purpose of sending the original audio information to the user.
[0157] In steps S701 to S706 above, the monitored speech information to be recognized can be obtained by calling the first interface. Then, the speech recognition model is called, and the network layer in the speech recognition model, which is determined by at least the auxiliary loss, is used to perform information mapping on the speech information to be recognized, obtaining the target information mapping result. Then, the network layer in the speech recognition model, which is determined by at least the original loss, is used to identify noise in the target information mapping result, identifying interference information and removing interference information from the speech information to be recognized, thereby recognizing the original audio information. Finally, the original audio information is output by calling the second interface. That is to say, in this embodiment of the application, the speech recognition model can be used to identify interference information in the speech information to be recognized. Then, the identified interference information is removed from the speech information to be recognized to identify the original audio information. Since the speech recognition model is trained based on the original loss and the auxiliary loss, based on this, the speech recognition model can be used to recognize the speech information to be recognized more accurately, achieving the purpose of removing interference information in the speech information to be recognized, improving the accuracy of speech information recognition, and thus solving the technical problem of low accuracy of speech information recognition.
[0158] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0159] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0160] Example 2
[0161] The preferred embodiments of the method described above in this example will be further described below.
[0162] In related technologies, speech recognition models are usually trained on clean, noise-free corpora. They perform well in recognizing noise-free speech information, but when the speech information to be recognized contains noise, the performance of such speech recognition models is often poor. That is, the noise resistance of the model is poor. Since real-world speech usually contains background noise, reverberation and other nonlinear distortions, there are usually technical problems of inaccurate recognition when using such speech recognition models to recognize real-world speech.
[0163] To address the aforementioned problems, this invention proposes a speech recognition method. The method involves collecting monitored speech information to be recognized, which at least includes the original audio information to be recognized. A speech recognition model is then invoked. This model is obtained by training an initial speech recognition model based on a primary loss and an auxiliary loss. The primary loss is obtained based on noise recognition results, which are obtained by identifying noise in a first noise information sample within the training sample set. The auxiliary loss is obtained based on information mapping results, which are obtained by mapping information to at least a second noise information sample within the training sample set. The speech recognition model includes the auxiliary loss... Based on the defined network layers and the network layers determined by the original loss, the speech recognition model can use at least the network layers determined by the auxiliary loss to perform information mapping on the speech information to be recognized, obtain the target information mapping result, and use the network layers determined by at least the original loss in the speech recognition model to perform noise recognition on the target information mapping result, identify the interference information, and remove the interference information from the speech information to be recognized to obtain the original audio information. In this application, through two recognitions, the purpose of removing interference information in the speech information to be recognized can be achieved, thereby improving the technical effect of speech information recognition accuracy and solving the technical problem of low speech information recognition accuracy.
[0164] The initial speech recognition model of this application embodiment is described below.
[0165] In this embodiment, the initial speech recognition model can be a HuBERT model that follows the w2v2 principle. This model architecture includes a convolutional encoder, a BERT encoder, a projection layer, and a code embedding layer. During pre-training, offline clustering (i.e., using the K-Means algorithm) can be used to generate aligned discrete target labels for computing a BERT-like prediction loss from the masked frames. HuBERT training begins with hidden units in K=100 clusters derived from MFCC features of the raw audio data. In subsequent iterations, the target code is updated based on hidden units in K=500 clusters, which are determined at the end of each iteration using the intermediate latent representation of the sixth layer of the HuBERT transform.
[0166] The speech recognition model of the embodiments of this application will be described below.
[0167] In this embodiment, after training the HuBERT model, the resulting model is a speech recognition model, which can be called a deHuBERT model. This deHuBERT model can use a shared CNN encoder to generate second embeddings with different noise enhancement versions in parallel. For example, two sets of noise can be randomly selected and added to the training data, where the signal-to-noise ratio (SNR) of the two randomly selected noise sets is between 0dB and 25dB. Then, the two sets of noise can be encoded to obtain the corresponding encoded features X and . and encoding feature X and The information is passed to the information mapping module to obtain the mapped encoded features Y and Y, respectively. After obtaining the mapped encoded features, the cross-correlation loss of the speech discrimination model can be further determined based on the mapped encoded features.
[0168] In this embodiment, the training dataset can be divided into single noise and overlapping noise. Single noise can include car noise, subway noise, and traffic noise, while overlapping noise can be divided into multi-path overlapping noise, airport / station noise, and coffee shop / air conditioner noise. The initial HuBERT speech recognition model is trained using the dataset to obtain the deHuBERT speech recognition model. Then, the deHuBERT model can be used to recognize the speech information to be recognized, identifying interference information within it. This identified interference information is then removed from the speech information to obtain the original audio information.
[0169] The process of training the initial speech recognition model to obtain the speech recognition model in the embodiments of this application will be further described below.
[0170] Figure 8 This is a schematic diagram illustrating the training of an initial speech recognition model according to an embodiment of this application. To obtain a non-entangled noise representation using the HuBERT model, a deHuBERT training algorithm is proposed to train the HuBERT model, using a shared CNN encoder to generate second embeddings with different noise enhancement versions in parallel. For example, after inputting the signal into the CNN encoder, two sets of noise can be randomly selected from the training dataset. x1, x2, and x3 are added to the CNN encoder to obtain encoded feature representations of two sets of noise, where the signal-to-noise ratio of the two sets of noise is between 0dB and 25dB. Then, the obtained encoded features can be passed to a shared linear projection module to obtain the mapped noise information Y and . Finally, the mapped noise information Y and Cross-correlation processing is performed to obtain the initial cross-correlation result (Empirical Cross-correlation) of the HuBERT model, and the mapped noise information Y and Target cross-correlation processing is performed to obtain the target cross-correlation result of the HuBERT model.
[0171] During the training process, continuous pre-training can be performed using the Fairseq toolkit. Projection blocks can be constructed for the cross-correlation loss and autocorrelation loss of the model. This projection can correspond to a d-dimensional matrix, which can be 2048*4096 in size.
[0172] During training, checkpoints meeting predetermined conditions from pre-training can be used, and training can be performed for 100 hours, 10 hours, 1 hour, and 10 minutes according to typical baseline settings. Fine-tuning of the initial speech recognition model's Automatic Speech Recognition (ASR) involves only the Hubert component. Furthermore, multi-condition training can be employed, where the signal-to-noise ratio of the training noise can range from 0 dB to 20 dB.
[0173] Table 1 shows a subset of premixed test audio samples with various signal-to-noise ratios (SNR) ranging from 0dB to 20dB, illustrating the ASR performance of the model. Table 1 records the data obtained from training the HuberT and deHuBERT models using audio with added single-class noise and audio with added mixed-class noise. The data in Table 1 shows that, under different training durations, the HuberT model consistently outperforms the deHuBERT model in recognizing clean audio compared to audio with added noise, and the HuberT model consistently outperforms the deHuBERT model in recognizing audio with added noise.
[0174] Table 1. Recognition of premixed test audio with various SNR types between 0dB and 20dB.
[0175]
[0176] Table 2 shows the test results for various out-of-domain noises under three different conditions. The three test conditions are (1) original clean audio, (2) with an additional FreeSound noise test device, adding noise of 0-20dB, and (3) with an additional OOD test device, adding noise of 0-20dB. That is, the audio under the three test conditions is identified by the HuBERT model and the deHuBERT model respectively, so as to determine the noise recognition performance of the two models.
[0177] First, training revealed that the deHuBERT model outperforms the HuBERT model in recognizing original, clean audio. This indicates that the deHuBERT model performs well in recognizing original, clean audio and is unaffected by noise pre-training.
[0178] Second, even with fine-tuning of the noise identification device, the deHuBERT model consistently outperforms the HuberT model in the noisy environments of conditions (2) and (3).
[0179] Table 2 Test results of various extraterritorial noises under different conditions
[0180]
[0181] This application provides a novel pre-training framework that separates noise from autocorrelation and cross-correlation losses for more robust speech recognition. The deHuBERT model demonstrates superior performance in recognizing noisy audio without affecting test performance on clean audio.
[0182] Example 3
[0183] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 3 The speech recognition device of the speech recognition method shown.
[0184] Figure 9 This is a schematic diagram of a speech recognition device according to an embodiment of this application. Figure 9 As shown, the voice recognition device 900 may include: a collection unit 901, a recall unit 902, a mapping unit 903, a first recognition unit 904, and a second recognition unit 905.
[0185] The acquisition unit 901 is used to acquire the monitored speech information to be recognized, wherein the speech information to be recognized contains at least the original audio information to be recognized.
[0186] Calling unit 902 is used to call the speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on the original loss and auxiliary loss. The original loss is obtained based on the noise recognition result, which is obtained by noise recognition of the first noise information sample in the training sample set. The auxiliary loss is obtained based on the information mapping result, which is obtained by information mapping of at least the second noise information sample in the training sample set.
[0187] The mapping unit 903 is used to perform information mapping on the speech information to be recognized using at least the network layer in the speech recognition model determined by the auxiliary loss, so as to obtain the target information mapping result.
[0188] The first recognition unit 904 is used to identify noise in the target information mapping result using at least the network layer determined by the original loss in the speech recognition model, and to identify interference information, wherein the interference information is information that interferes with the original audio information.
[0189] The second recognition unit 905 is used to remove interference information from the speech information to be recognized and to recognize the original audio information.
[0190] It should be noted that the above-mentioned acquisition unit 901, calling unit 902, mapping unit 903, first identification unit 904 and second identification unit 905 correspond to steps S301 to S305 in Embodiment 1. It should be noted that the above-mentioned modules or units may be hardware components or software components stored in memory and processed by one or more processors. The above-mentioned modules may also be part of the device and can run in the AR / VR device provided in Embodiment 1.
[0191] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 4 The method for determining the speech recognition model shown is a speech recognition model determination device.
[0192] Figure 10 This is a schematic diagram of a speech recognition model determination device according to an embodiment of this application. Figure 10 As shown, the speech recognition model determination device 1000 may include: an acquisition unit 1001, a recognition unit 1002, and a training unit 1003.
[0193] The acquisition unit 1001 is used to acquire the first noise information sample and the second noise information sample from the training sample set.
[0194] The recognition unit 1002 is used to determine the original loss of the initial speech recognition model based on the noise recognition result obtained by performing speech recognition on the first noise information sample based on the initial speech recognition model, and at least to determine the auxiliary loss of the initial speech recognition model based on the information mapping result obtained by performing information mapping on the second noise information sample based on the initial speech recognition model.
[0195] Training unit 1003 is used to train an initial speech recognition model based on the original loss and auxiliary loss to obtain a speech recognition model. The speech recognition model has at least a network layer determined by the auxiliary loss, which is used to perform information mapping on the speech information to be recognized to obtain the target information mapping result. The speech information to be recognized contains at least the original audio information to be recognized. The speech recognition model has at least a network layer determined by the original loss, which is used to identify noise on the target information mapping result and identify interference information. The interference information is information that interferes with the original audio information and is used to identify the original audio information from the speech information to be recognized.
[0196] It should be noted that the above-mentioned acquisition unit 1001, identification unit 1002, and training unit 1003 correspond to steps S401 to S403 in Embodiment 1. It should be noted that the above-mentioned modules or units may be hardware components or software components stored in memory and processed by one or more processors. The above-mentioned modules may also be part of the device and can run in the AR / VR device provided in Embodiment 1.
[0197] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 5 The speech recognition device of the speech recognition method shown.
[0198] Figure 11 This is a schematic diagram of a speech recognition device according to an embodiment of this application. Figure 11 As shown, the speech recognition device 1100 may include: a collection unit 1101, a mapping unit 1102, a first recognition unit 1103, a second recognition unit 1104, and an activation unit 1105.
[0199] The acquisition unit 1101 is used to acquire the voice information to be recognized sent to the client, wherein the voice information to be recognized contains at least the original audio information to be recognized.
[0200] The mapping unit 1102 is used to perform information mapping on the speech information to be recognized using at least one network layer in the speech recognition model determined by the auxiliary loss, and to obtain the target information mapping result. The speech recognition model is obtained by training an initial speech recognition model based on the original loss and the auxiliary loss. The original loss is obtained based on the noise recognition result, which is obtained by noise recognition of the first noise information sample in the training sample set. The auxiliary loss is obtained based on the information mapping result, which is obtained by information mapping of at least the second noise information sample in the training sample set.
[0201] The first recognition unit 1103 is used to identify noise in the target information mapping result using at least the network layer determined by the original loss in the speech recognition model, and to identify interference information, wherein the interference information is information that interferes with the original audio information.
[0202] The second recognition unit 1104 is used to remove interference information from the speech information to be recognized and to recognize the original audio information.
[0203] Activation unit 1105 is used to activate the client based on the original audio information.
[0204] It should be noted that the acquisition unit 1101, mapping unit 1102, first identification unit 1103, second identification unit 1104 and activation unit 1105 mentioned above correspond to steps S501 to S505 in Embodiment 1. It should also be noted that the above modules or units may be hardware components or software components stored in memory and processed by one or more processors. The above modules may also be part of the device and can run in the AR / VR device provided in Embodiment 1.
[0205] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 6 The speech recognition device of the speech recognition method shown.
[0206] Figure 12 This is a schematic diagram of a speech recognition device according to an embodiment of this application. Figure 12 As shown, the voice recognition device 1200 may include: an input unit 1201, a calling unit 1202, a mapping unit 1203, a first recognition unit 1204, a second recognition unit 1205, and an activation unit 1206.
[0207] The input unit 1201 is used to input the monitored voice information to be recognized into a virtual reality (VR) device or an augmented reality (AR) device, wherein the voice information to be recognized contains at least the original audio information to be recognized.
[0208] Calling unit 1202 is used to call the speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on the original loss and the auxiliary loss. The original loss is obtained based on the noise recognition result, which is obtained by noise recognition of the first noise information sample in the training sample set. The auxiliary loss is obtained based on the information mapping result, which is obtained by information mapping of at least the second noise information sample in the training sample set.
[0209] The mapping unit 1203 is used to perform information mapping on the speech information to be recognized using at least the network layer in the speech recognition model determined by the auxiliary loss, so as to obtain the target information mapping result.
[0210] The first recognition unit 1204 is used to identify noise in the target information mapping result using at least the network layer determined by the original loss in the speech recognition model, and to identify interference information, wherein the interference information is information that interferes with the original audio information.
[0211] The second recognition unit 1205 is used to remove interference information from the speech information to be recognized and to recognize the original audio information.
[0212] Activation unit 1206 is used to activate VR or AR devices using the original audio information.
[0213] It should be noted that the above-mentioned input unit 1201, calling unit 1202, mapping unit 1203, first identification unit 1204, second identification unit 1205 and activation unit 1206 correspond to steps S601 to S606 in Embodiment 1. It should be noted that the above-mentioned modules or units may be hardware components or software components stored in memory and processed by one or more processors. The above-mentioned modules may also be part of the device and can run in the AR / VR device provided in Embodiment 1.
[0214] According to embodiments of the present invention, a method for implementing the above is also provided. Figure 7 The speech recognition device of the speech recognition method shown.
[0215] Figure 13 This is a schematic diagram of a speech recognition device according to an embodiment of this application. Figure 13 As shown, the voice recognition device 1300 may include: a first calling unit 1301, a second calling unit 1302, a mapping unit 1303, a first recognition unit 1304, a second recognition unit 1305, and a third calling unit 1306.
[0216] The first calling unit is used to obtain the monitored speech information to be recognized by calling the first interface. The first interface includes a first parameter, the value of which is the speech information to be recognized. The speech information to be recognized contains at least the original audio information to be recognized.
[0217] The second calling unit is used to call the speech recognition model, which is obtained by mapping the initial speech recognition information based on the original loss and auxiliary loss.
[0218] The mapping unit is used to map the speech information to be recognized using at least the network layer in the speech recognition model determined by the auxiliary loss, and obtain the target information mapping result.
[0219] The first recognition unit is used to identify noise in the target information mapping result using at least the network layer determined by the original loss in the speech recognition model, and to identify interference information, wherein the interference information is information that interferes with the original audio information.
[0220] The second recognition unit is used to remove interference information from the speech information to be recognized and to identify the original audio information.
[0221] The third calling unit is used to call the second interface to output the original audio information. The second interface includes a second parameter, the value of which is the original audio information.
[0222] It should be noted that the first calling unit 1301, the second calling unit 1302, the mapping unit 1303, the first identification unit 1304, the second identification unit 1305, and the third calling unit 1306 mentioned above correspond to steps S701 to S706 in Embodiment 1. It should be noted that the above modules or units may be hardware components or software components stored in memory and processed by one or more processors. The above modules may also be part of the device and can run in the AR / VR device provided in Embodiment 1.
[0223] Example 4
[0224] Optionally, in this embodiment, a memory is also provided, which is further used to provide the processor with instructions for processing the following steps: acquiring monitored speech information to be recognized, wherein the speech information to be recognized contains at least the original audio information to be recognized; calling a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss is obtained based on the noise recognition result, the noise recognition result is obtained by noise recognition of a first noise information sample in the training sample set, the auxiliary loss is obtained based on the information mapping result, and the information mapping result is obtained by information mapping of at least a second noise information sample in the training sample set; using a network layer in the speech recognition model determined by at least the auxiliary loss to perform information mapping on the speech information to be recognized, to obtain a target information mapping result; using a network layer in the speech recognition model determined by at least the original loss to perform noise recognition on the target information mapping result, to identify interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized, and recognizing the original audio information.
[0225] Optionally, in this embodiment, the memory is further configured to provide the processor with instructions to perform the following steps: mapping the first noise information sample and the second noise information sample based on the information mapping model in the initial speech recognition model; determining the cross-correlation loss of the initial speech recognition model based on the mapped first noise information sample and the mapped second noise information sample, wherein the auxiliary loss includes the cross-correlation loss, which is used to represent the difference between the degree of correlation between multiple input speech information of the initial speech recognition model and the corresponding target correlation.
[0226] Optionally, in this embodiment, the memory is further configured to provide the processor with instructions to process the following steps: performing cross-correlation processing on the mapped first noise information sample and the mapped second noise information sample to obtain an initial cross-correlation result, wherein the initial cross-correlation result is used to represent the degree of correlation between the mapped first noise information sample and the mapped second noise information sample; determining the cross-correlation loss based on the initial cross-correlation result and the target cross-correlation result, wherein the target cross-correlation result is used to represent the target degree of correlation between the mapped first noise information sample and the mapped second noise information sample.
[0227] Optionally, in this embodiment, the memory is also used to provide the processor with instructions to process the following steps: linearly mapping a first noise information sample to a first matrix based on an information mapping model, and linearly mapping a second noise information sample to a second matrix; obtaining the cross-correlation matrix between the first matrix and the second matrix, wherein the cross-correlation matrix is associated with the identity matrix; and determining the cross-correlation loss based on the cross-correlation matrix.
[0228] Optionally, in this embodiment, the memory is also used to provide the processor with instructions to perform the following steps: enhance the first noise information sample and enhance the second noise information sample; perform information mapping on the enhanced first noise information sample and the enhanced second noise information sample respectively based on the information mapping model.
[0229] Optionally, in this embodiment, the memory is further configured to provide the processor with instructions for processing the following steps: obtaining the noise recognition result obtained by the initial speech recognition model performing noise recognition on the first noise information sample; performing information mapping on the noise recognition result based on the information mapping model in the initial speech recognition model; and determining the autocorrelation loss of the initial speech recognition model based on the mapped noise recognition result, wherein the auxiliary loss includes the autocorrelation loss, which is used to represent the difference between the degree of correlation between different speech information in the output noise recognition result of the initial speech recognition model and the corresponding target correlation degree.
[0230] Optionally, in this embodiment, the memory is further configured to provide the processor with instructions to process the following steps: performing autocorrelation processing on the mapped noise recognition result to obtain an initial autocorrelation result, wherein the initial autocorrelation result is used to represent the degree of correlation between different noise information in the mapped noise recognition result; determining the autocorrelation loss of the initial speech recognition model based on the initial autocorrelation result and the target autocorrelation result, wherein the auxiliary loss includes the autocorrelation loss, and the target autocorrelation result is used to represent the degree of target correlation between different noise information in the mapped noise recognition result.
[0231] Optionally, in this embodiment, the memory is also used to provide the processor with instructions to perform the following steps: enhance the first noise information sample; and obtain the noise recognition result obtained by the initial speech recognition model performing noise recognition on the enhanced first noise information sample.
[0232] Optionally, in this embodiment, the memory is also used to provide the processor with instructions to process the following steps: randomly selecting a first noise information sample and a second noise information sample from the training sample set, wherein the training sample set includes multiple noise information samples with a signal-to-noise ratio within the signal-to-noise ratio threshold range.
[0233] Example 5
[0234] Embodiments of this application may provide an AR / VR device, which can be any AR / VR device from a group of AR / VR devices. Optionally, in this embodiment, the aforementioned AR / VR device may also be replaced by a terminal device such as a mobile terminal.
[0235] Optionally, in this embodiment, the AR / VR device described above may be located in at least one of a plurality of network devices in a computer network.
[0236] In this embodiment, the AR / VR device described above can execute the following steps of the speech recognition method: inputting the detected speech information to be recognized into the virtual reality (VR) device or augmented reality (AR) device, wherein the speech information to be recognized contains at least the original audio information to be recognized; invoking the speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on the original loss and the auxiliary loss, the original loss is obtained based on the noise recognition result, the noise recognition result is obtained by noise recognition of the first noise information sample in the training sample set, the auxiliary loss is obtained based on the information mapping result, and the information mapping result is obtained by information mapping of at least the second noise information sample in the training sample set; using the network layer in the speech recognition model determined by at least the auxiliary loss, performing information mapping on the speech information to be recognized to obtain the target information mapping result; using the network layer in the speech recognition model determined by at least the original loss, performing noise recognition on the target information mapping result to identify interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized to identify the original audio information; and activating the VR device or AR device using the original audio information.
[0237] Optionally, Figure 14 This is a structural block diagram of a computer terminal according to an embodiment of this application. As shown in the figure, the computer terminal A may include: one or more (only one is shown in the figure) processors 1402, memory 1404, memory controller, and peripheral interfaces, wherein the peripheral interfaces are connected to a radio frequency module, an audio module, and a display.
[0238] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the speech recognition method and device in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned speech recognition method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to computer terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0239] The processor can invoke information and application programs stored in memory via a transmission device to perform the following steps: acquiring monitored speech information to be recognized, wherein the speech information to be recognized contains at least the original audio information to be recognized; invoking a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on a primary loss and an auxiliary loss, the primary loss being obtained based on noise recognition results, the noise recognition results being obtained by noise recognition of a first noise information sample in the training sample set, the auxiliary loss being obtained based on information mapping results, and the information mapping results being obtained by information mapping of at least a second noise information sample in the training sample set; using a network layer in the speech recognition model determined at least by the auxiliary loss to perform information mapping on the speech information to be recognized, obtaining a target information mapping result; using a network layer in the speech recognition model determined at least by the primary loss to perform noise recognition on the target information mapping result, identifying interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized, and recognizing the original audio information.
[0240] Optionally, the processor may also execute program code that performs the following steps: mapping the first noise information sample and the second noise information sample based on the information mapping model in the initial speech recognition model; determining the cross-correlation loss of the initial speech recognition model based on the mapped first noise information sample and the mapped second noise information sample, wherein the auxiliary loss includes the cross-correlation loss, which is used to represent the difference between the degree of correlation between multiple input speech information of the initial speech recognition model and the corresponding target correlation.
[0241] Optionally, the processor may also execute program code for the following steps: performing cross-correlation processing on the mapped first noise information sample and the mapped second noise information sample to obtain an initial cross-correlation result, wherein the initial cross-correlation result is used to represent the degree of correlation between the mapped first noise information sample and the mapped second noise information sample; determining the cross-correlation loss based on the initial cross-correlation result and the target cross-correlation result, wherein the target cross-correlation result is used to represent the target degree of correlation between the mapped first noise information sample and the mapped second noise information sample.
[0242] Optionally, the processor may also execute program code for the following steps: linearly mapping the first noise information sample to a first matrix and linearly mapping the second noise information sample to a second matrix based on the information mapping model; obtaining the cross-correlation matrix between the first matrix and the second matrix, wherein the cross-correlation matrix is associated with the identity matrix; and determining the cross-correlation loss based on the cross-correlation matrix.
[0243] Optionally, the processor may also execute program code that performs the following steps: enhances the first noise information sample and enhances the second noise information sample; and performs information mapping on the enhanced first noise information sample and the enhanced second noise information sample based on the information mapping model.
[0244] Optionally, the processor may also execute program code for the following steps: obtaining the noise recognition result obtained by the initial speech recognition model in performing noise recognition on the first noise information sample; performing information mapping on the noise recognition result based on the information mapping model in the initial speech recognition model; determining the autocorrelation loss of the initial speech recognition model based on the mapped noise recognition result, wherein the auxiliary loss includes the autocorrelation loss, which is used to represent the difference between the degree of correlation between different speech information in the output noise recognition result of the initial speech recognition model and the degree of correlation of the corresponding target.
[0245] Optionally, the processor may also execute program code for the following steps: performing autocorrelation processing on the mapped noise recognition result to obtain an initial autocorrelation result, wherein the initial autocorrelation result is used to represent the degree of correlation between different noise information in the mapped noise recognition result; determining the autocorrelation loss of the initial speech recognition model based on the initial autocorrelation result and the target autocorrelation result, wherein the auxiliary loss includes the autocorrelation loss, and the target autocorrelation result is used to represent the degree of target correlation between different noise information in the mapped noise recognition result.
[0246] Optionally, the processor may also execute program code for the following steps: performing information enhancement on the first noise information sample; and obtaining the noise recognition result obtained by the initial speech recognition model performing noise recognition on the enhanced first noise information sample.
[0247] Optionally, the processor may also execute program code that performs the following steps: randomly selects a first noise information sample and a second noise information sample from the training sample set, wherein the training sample set includes multiple noise information samples with a signal-to-noise ratio within the signal-to-noise ratio threshold range.
[0248] This application provides a speech recognition method. The method uses a speech recognition model to identify the speech information to be recognized, obtaining the original audio information. Since the speech recognition model is trained based on the original loss and auxiliary loss, after obtaining the speech information to be recognized, the network layer included in the speech recognition model, determined by the auxiliary loss, can be used to perform a first recognition. Then, the network layer included in the speech recognition model, determined by the original loss, can be used to perform a second recognition on the target information mapping result obtained in the first recognition, identifying interference information and removing it from the speech information to be recognized. Through these two recognitions, the purpose of removing interference information from the speech information to be recognized is achieved, improving the accuracy of speech information recognition and thus solving the technical problem of low accuracy in speech information recognition.
[0249] Those skilled in the art will understand that the structure shown in the figure is for illustrative purposes only, and the computer terminal may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a mobile internet device (MID), a PAD, and other terminal devices. Figure 14 This does not limit the structure of the aforementioned computer terminal. For example, computer terminal A may also include components that are more complex than those described above. Figure 14 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 14 The different configurations shown.
[0250] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0251] Example 6
[0252] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the speech recognition method provided in Embodiment 1.
[0253] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in the AR / VR device terminal group in the AR / VR device network, or in any mobile terminal in the mobile terminal group.
[0254] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: acquiring monitored speech information to be recognized, wherein the speech information to be recognized includes at least the original audio information to be recognized; invoking a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss being obtained based on noise recognition results, the noise recognition results being obtained by noise recognition of a first noise information sample in the training sample set, the auxiliary loss being obtained based on information mapping results, and the information mapping results being obtained by information mapping of at least a second noise information sample in the training sample set; using a network layer in the speech recognition model determined at least by the auxiliary loss to perform information mapping on the speech information to be recognized, obtaining a target information mapping result; using a network layer in the speech recognition model determined at least by the original loss to perform noise recognition on the target information mapping result, identifying interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the speech information to be recognized, and recognizing the original audio information.
[0255] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: mapping the first noise information sample and the second noise information sample based on the information mapping model in the initial speech recognition model; determining the cross-correlation loss of the initial speech recognition model based on the mapped first noise information sample and the mapped second noise information sample, wherein the auxiliary loss includes the cross-correlation loss, which is used to represent the difference between the degree of correlation between multiple input speech information of the initial speech recognition model and the corresponding target correlation degree.
[0256] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing cross-correlation processing on the mapped first noise information sample and the mapped second noise information sample to obtain an initial cross-correlation result, wherein the initial cross-correlation result is used to represent the degree of correlation between the mapped first noise information sample and the mapped second noise information sample; determining the cross-correlation loss based on the initial cross-correlation result and the target cross-correlation result, wherein the target cross-correlation result is used to represent the target degree of correlation between the mapped first noise information sample and the mapped second noise information sample.
[0257] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: linearly mapping the first noise information sample to a first matrix and linearly mapping the second noise information sample to a second matrix based on the information mapping model; obtaining the cross-correlation matrix between the first matrix and the second matrix, wherein the cross-correlation matrix is associated with the identity matrix; and determining the cross-correlation loss based on the cross-correlation matrix.
[0258] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: enhancing the first noise information sample and enhancing the second noise information sample; and mapping the enhanced first noise information sample and the enhanced second noise information sample based on an information mapping model.
[0259] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: obtaining noise recognition results obtained by the initial speech recognition model for noise recognition of the first noise information sample; performing information mapping on the noise recognition results based on the information mapping model in the initial speech recognition model; determining the autocorrelation loss of the initial speech recognition model based on the mapped noise recognition results, wherein the auxiliary loss includes the autocorrelation loss, which is used to represent the difference between the degree of correlation between different speech information in the output noise recognition results of the initial speech recognition model and the corresponding target correlation degree.
[0260] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: performing autocorrelation processing on the mapped noise recognition result to obtain an initial autocorrelation result, wherein the initial autocorrelation result is used to represent the degree of correlation between different noise information in the mapped noise recognition result; determining the autocorrelation loss of the initial speech recognition model based on the initial autocorrelation result and the target autocorrelation result, wherein the auxiliary loss includes the autocorrelation loss, and the target autocorrelation result is used to represent the target degree of correlation between different noise information in the mapped noise recognition result.
[0261] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: enhancing the first noise information sample; and obtaining the noise recognition result obtained by the initial speech recognition model performing noise recognition on the enhanced first noise information sample.
[0262] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: randomly selecting a first noise information sample and a second noise information sample from a training sample set, wherein the training sample set includes multiple noise information samples with a signal-to-noise ratio within a signal-to-noise ratio threshold range.
[0263] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0264] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0265] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0266] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0267] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0268] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A voice recognition method, characterized by, The method comprises: collecting monitored voice information to be recognized, wherein the voice information to be recognized at least contains original audio information to be recognized; calling a voice recognition model, wherein the voice recognition model is obtained by training an initial voice recognition model based on an original loss and an auxiliary loss, the original loss is obtained based on a noise recognition result, the noise recognition result is obtained by performing noise recognition on a first noise information sample in a training sample set, and the auxiliary loss is obtained based on an information mapping result, the information mapping result is obtained by at least performing information mapping on a second noise information sample in the training sample set; performing information mapping on the voice information to be recognized using at least a network layer of the voice recognition model determined by the auxiliary loss to obtain a target information mapping result; performing noise recognition on the target information mapping result using at least a network layer of the voice recognition model determined by the original loss to identify interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the voice information to be recognized to identify the original audio information. The auxiliary loss comprises a cross-correlation loss and an autocorrelation loss, the cross-correlation loss is used to represent the difference between the correlation degree of a plurality of input voice information of the initial voice recognition model and the corresponding target correlation degree, and the autocorrelation loss is used to represent the difference between the correlation degree of different voice information in the output noise recognition result of the initial voice recognition model and the corresponding target correlation degree.
2. The method of claim 1, wherein, The method further comprises: performing information mapping on the first noise information sample and the second noise information sample respectively based on an information mapping model in the initial voice recognition model; determining the cross-correlation loss of the initial voice recognition model based on the mapped first noise information sample and the mapped second noise information sample.
3. The method of claim 2, wherein, Determining the cross-correlation loss of the initial voice recognition model based on the mapped first noise information sample and the mapped second noise information sample comprises: performing cross-correlation processing on the mapped first noise information sample and the mapped second noise information sample to obtain an initial cross-correlation result, wherein the initial cross-correlation result is used to represent the correlation degree between the mapped first noise information sample and the mapped second noise information sample; determining the cross-correlation loss based on the initial cross-correlation result and a target cross-correlation result, wherein the target cross-correlation result is used to represent the target correlation degree between the mapped first noise information sample and the mapped second noise information sample.
4. The method of claim 2, wherein, Performing information mapping on the first noise information sample and the second noise information sample respectively based on an information mapping model in the initial voice recognition model comprises: linearly mapping the first noise information sample into a first matrix and linearly mapping the second noise information sample into a second matrix based on the information mapping model; Determining the cross-correlation loss of the initial speech recognition model based on the mapped first noise information sample and the mapped second noise information sample includes: obtaining the cross-correlation matrix between the first matrix and the second matrix, wherein the cross-correlation matrix is associated with the identity matrix; and determining the cross-correlation loss based on the cross-correlation matrix.
5. The method of claim 2, wherein, The method further includes: Information enhancement is performed on the first noise information sample and the second noise information sample. Based on the information mapping model, information mapping is performed on the first noise information sample and the second noise information sample, respectively, including: based on the information mapping model, information mapping is performed on the enhanced first noise information sample and the enhanced second noise information sample, respectively.
6. The method of claim 1, wherein, The method further includes: Obtain the noise recognition result obtained by the initial speech recognition model performing noise recognition on the first noise information sample; Information mapping is performed on the noise recognition results based on the information mapping model in the initial speech recognition model; The autocorrelation loss of the initial speech recognition model is determined based on the noise recognition result after mapping.
7. The method of claim 6, wherein, The autocorrelation loss of the initial speech recognition model is determined based on the mapped noise recognition result, including: The mapped noise identification results are subjected to autocorrelation processing to obtain an initial autocorrelation result, wherein the initial autocorrelation result is used to represent the degree of correlation between different noise information in the mapped noise identification results; Based on the initial autocorrelation result and the target autocorrelation result, the autocorrelation loss of the initial speech recognition model is determined, wherein the target autocorrelation result is used to represent the degree of target correlation between different noise information in the mapped noise recognition result.
8. The method of claim 6, wherein, The method further includes: Information enhancement is performed on the first noise information sample; Obtaining the noise recognition result obtained by the initial speech recognition model performing noise recognition on the first noise information sample includes: obtaining the noise recognition result obtained by the initial speech recognition model performing noise recognition on the enhanced first noise information sample.
9. The method according to any one of claims 1 to 8, characterized in that, The method further includes: From the training sample set, the first noise information sample and the second noise information sample are randomly selected, wherein the training sample set includes multiple noise information samples with a signal-to-noise ratio within the signal-to-noise ratio threshold range.
10. The method according to any one of claims 1 to 9, characterized in that, The sound level of the first noise information sample is greater than the sound level of the second noise information sample. 11.A method for determining a speech recognition model, the method comprising: include: Obtain the first noise information sample and the second noise information sample from the training sample set; Based on the noise recognition result obtained by performing speech recognition on the first noise information sample using the initial speech recognition model, the original loss of the initial speech recognition model is determined, and at least based on the information mapping result obtained by performing information mapping on the second noise information sample using the initial speech recognition model, the auxiliary loss of the initial speech recognition model is determined. train the initial speech recognition model based on the original loss and the auxiliary loss to obtain a speech recognition model, wherein at least a network layer of the speech recognition model determined by the auxiliary loss is configured to perform information mapping on to-be-recognized speech information to obtain a target information mapping result, the to-be-recognized speech information at least contains original audio information to be recognized, and at least a network layer of the speech recognition model determined by the original loss is configured to perform noise recognition on the target information mapping result to identify interference information, the interference information is information that interferes with the original audio information, and the original audio information is identified from the to-be-recognized speech information. The auxiliary loss includes a cross-correlation loss and an autocorrelation loss, the cross-correlation loss is used to represent a difference between a correlation degree of a plurality of input speech information of the initial speech recognition model and a corresponding target correlation degree, and the autocorrelation loss is used to represent a difference between a correlation degree of different speech information in an output noise recognition result of the initial speech recognition model and a corresponding target correlation degree.
12. A voice recognition method, characterized by, It includes: collecting to-be-recognized speech information sent to a client, wherein the to-be-recognized speech information at least contains original audio information to be recognized; performing information mapping on the to-be-recognized speech information by using at least a network layer of a speech recognition model determined by an auxiliary loss to obtain a target information mapping result, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and the auxiliary loss, the original loss is obtained based on a noise recognition result, the noise recognition result is obtained by performing noise recognition on a first noise information sample in a training sample set, and the auxiliary loss is obtained based on an information mapping result, the information mapping result is obtained by at least performing information mapping on a second noise information sample in the training sample set; performing noise recognition on the target information mapping result by using at least a network layer of the speech recognition model determined by the original loss to identify interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the to-be-recognized speech information to identify the original audio information; activating the client based on the original audio information; The auxiliary loss includes a cross-correlation loss and an autocorrelation loss, the cross-correlation loss is used to represent a difference between a correlation degree of a plurality of input speech information of the initial speech recognition model and a corresponding target correlation degree, and the autocorrelation loss is used to represent a difference between a correlation degree of different speech information in an output noise recognition result of the initial speech recognition model and a corresponding target correlation degree.
13. A voice recognition method, characterized by, It includes: inputting monitored to-be-recognized speech information on a virtual reality (VR) device or an augmented reality (AR) device, wherein the to-be-recognized speech information at least contains original audio information to be recognized; calling a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss is obtained based on a noise recognition result, the noise recognition result is obtained by performing noise recognition on a first noise information sample in a training sample set, and the auxiliary loss is obtained based on an information mapping result, the information mapping result is obtained by performing information mapping on at least a second noise information sample in the training sample set; performing information mapping on the to-be-recognized speech information by using at least a network layer of the speech recognition model determined based on the auxiliary loss, to obtain a target information mapping result; performing noise recognition on the target information mapping result by using at least a network layer of the speech recognition model determined based on the original loss, to recognize interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the to-be-recognized speech information, to recognize the original audio information; activating the VR device or the AR device by using the original audio information. The auxiliary loss includes a cross-correlation loss and an autocorrelation loss, the cross-correlation loss is used to represent a difference between a correlation degree between a plurality of input speech information of the initial speech recognition model and a corresponding target correlation degree, and the autocorrelation loss is used to represent a difference between a correlation degree between different speech information in an output noise recognition result of the initial speech recognition model and a corresponding target correlation degree.
14. A voice recognition method, characterized by, comprising: obtaining monitored to-be-recognized speech information by calling a first interface, wherein the first interface includes a first parameter, a parameter value of the first parameter is the to-be-recognized speech information, and the to-be-recognized speech information at least includes original audio information to be recognized; calling a speech recognition model, wherein the speech recognition model is obtained by training an initial speech recognition model based on an original loss and an auxiliary loss, the original loss is obtained based on a noise recognition result, the noise recognition result is obtained by performing noise recognition on a first noise information sample in a training sample set, and the auxiliary loss is obtained based on an information mapping result, the information mapping result is obtained by performing information mapping on at least a second noise information sample in the training sample set; performing information mapping on the to-be-recognized speech information by using at least a network layer of the speech recognition model determined based on the auxiliary loss, to obtain a target information mapping result; performing noise recognition on the target information mapping result by using at least a network layer of the speech recognition model determined based on the original loss, to recognize interference information, wherein the interference information is information that interferes with the original audio information; removing the interference information from the to-be-recognized speech information, to recognize the original audio information; outputting the original audio information by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter is the original audio information. The auxiliary loss includes a cross-correlation loss and an autocorrelation loss, the cross-correlation loss is used to represent a difference between a correlation degree between the multiple input speech information of the initial speech recognition model and a corresponding target correlation degree, and the autocorrelation loss is used to represent a difference between a correlation degree between different speech information in the output noise recognition result of the initial speech recognition model and a corresponding target correlation degree.
Citation Information
Patent Citations
Audio recognition method and device, electronic equipment and readable storage medium
CN114171029A