Method and apparatus for speech recognition, and device and storage medium

By transforming the recognition network of the machine learning model from a single network to multiple parallel networks, the problem of existing models being unable to effectively integrate new information is solved, achieving higher speech recognition accuracy and adaptability.

WO2026025914A1PCT designated stage Publication Date: 2026-02-05JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/081292
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-31
Filing Date
2025-03-07
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing machine learning models are unable to effectively integrate information when faced with large amounts of new data, resulting in limited retraining performance. The number of parameters becomes the upper limit of knowledge learning, affecting the accuracy of speech recognition.

Method used

The recognition network of the machine learning model is transformed from a single network to multiple parallel networks, with at least two recognition networks having different network parameters. During the recognition process, the most suitable network combination is selected for speech segment recognition.

Benefits of technology

It significantly expands the model capacity, improves the accuracy and adaptability of speech recognition, can more effectively absorb new information, and enhances the retraining effect in specific domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025081292_05022026_PF_FP_ABST
    Figure CN2025081292_05022026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to a method and apparatus for speech recognition, and a device and a storage medium. The method provided herein comprises: on the basis of content of target speech, determining, from among a first group of recognition networks of a trained machine learning model, a second group of recognition networks for recognizing a target speech segment in the target speech, wherein in the first group of recognition networks of the machine learning model, at least two recognition networks have different network parameters, and the second group of recognition networks comprises at least one recognition network; and on the basis of a recognition result obtained after recognition is performed on the target speech segment by means of the second group of recognition networks, determining text content with respect to the target speech segment.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for speech recognition

[0001] The present application claims priority to the Chinese patent application No. 202411047590.3, filed on July 31, 2024, entitled “Method, apparatus, device and storage medium for speech recognition”, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The example embodiments of the present disclosure generally relate to the field of computer technology, and more specifically, to a method, apparatus, device and storage medium for speech recognition. BACKGROUND

[0003] Automatic Speech Recognition (ASR) technology is a key technology that accurately converts an input speech signal into corresponding text content. Currently, this technology has played an important role in many fields such as e-commerce, finance, logistics, etc. In recent years, with the accumulation of large-scale training data and the rapid progress of deep learning technology, machine learning models trained based on big data for automatic speech recognition have achieved remarkable results. SUMMARY

[0004] In a first aspect of the present disclosure, a method for speech recognition is provided. The method comprises: determining, based on the content of a target speech, a second set of recognition networks for recognizing a target speech segment in the target speech from a first set of recognition networks of a trained machine learning model, wherein the network parameters of at least two recognition networks are different in the first set of recognition networks of the machine learning model, and wherein the second set of recognition networks comprises at least one recognition network; and determining the text content for the target speech segment based on the recognition result of the target speech segment by the second set of recognition networks.

[0005] In a second aspect of the present disclosure, an apparatus for speech recognition is provided. The apparatus comprises: a recognition network determination module configured to determine, based on the content of a target speech, a second set of recognition networks for recognizing a target speech segment in the target speech from a first set of recognition networks of a trained machine learning model, wherein the network parameters of at least two recognition networks are different in the first set of recognition networks of the machine learning model, and wherein the second set of recognition networks comprises at least one recognition network; and a text content generation module configured to determine the text content for the target speech segment based on the recognition result of the target speech segment by the second set of recognition networks.

[0006] In a third aspect of the disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. The instructions, when executed by the at least one processor, cause the device to perform the method of the first aspect.

[0007] In a fourth aspect of the disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon computer-executable instructions that are executable by a processor to implement the method of the first aspect.

[0008] In a fifth aspect of the disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to the first aspect of the disclosure.

[0009] It is to be understood that the details set forth herein are by way of example and not intended to limit the scope of the embodiments of the present disclosure. Other features and aspects of the present disclosure will become apparent from the following detailed description, from the drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, aspects, and advantages of embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements, wherein:

[0011] FIG. 1 shows a schematic diagram of an example of a machine learning model in a scenario;

[0012] FIG. 2 shows a schematic diagram of an environment according to some embodiments of the present disclosure;

[0013] FIG. 3 shows a flowchart of an example process of a method for speech recognition according to some embodiments of the present disclosure;

[0014] FIG. 4 shows a schematic diagram of an example of recognizing a target speech according to some embodiments of the present disclosure;

[0015] FIG. 5 shows a schematic structural block diagram of an apparatus for speech recognition according to some embodiments of the present disclosure; and

[0016] FIG. 6 shows a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure will be thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0018] It should be noted that the titles of any sections / sub-sections provided herein are not limiting. Various embodiments are described throughout this document and any type of embodiment can be included under any section / sub-section. Furthermore, embodiments described in any section / sub-section can be combined with any other embodiments described in the same section / sub-section and / or different section / sub-section in any manner.

[0019] In the description of embodiments of the present disclosure, the term "includes" and its derivatives, such as "including," should be understood in an open, inclusive sense, that is, "including, but not limited to." The term "based on" should be understood as "based at least in part on." The term "one embodiment" or "an embodiment" should be understood as "at least one embodiment." The term "some embodiments" should be understood as "at least some embodiments." Other explicitly and implicitly recited definitions can also be found below. The terms "first," "second," and the like can refer to different or identical objects. Other explicit and implicit definitions can also be found below.

[0020] Data of users, acquisition and / or use of data, etc. can be involved in embodiments of the present disclosure. These aspects all comply with corresponding laws and regulations and relevant provisions. In embodiments of the present disclosure, all collection, acquisition, processing, processing, forwarding, use, etc. of data are performed on the premise that users are aware of and confirm. Accordingly, when implementing embodiments of the present disclosure, the type of data or information that can be involved, the range of use, the scenario of use, etc. should be informed to users and authorized by users in a proper manner according to relevant laws and regulations. The specific informing and / or authorization manner can vary according to actual situations and application scenarios, and the scope of the present disclosure is not limited in this respect.

[0021] In the present specification and embodiments, if personal information processing is involved, it will be processed on the premise of legality (for example, obtaining the consent of the subject of personal information, or being necessary for the performance of a contract, etc.), and only within the prescribed or agreed range. Users refuse to process personal information other than the necessary information required for basic functions, which will not affect the user's use of basic functions.

[0022] As briefly described above, automatic speech recognition technology is widely used in various fields. Currently, a variety of open-source machine learning models are available online for users to choose from. For these open-source models, users can retrain them using their own domain-specific data, allowing the model to better adapt to the chosen application area. However, the capacity of a machine learning model is limited by the number of its parameters. This may prevent the model from fully learning information from new data; in other words, the number of parameters in a machine learning model constitutes the upper limit of the knowledge it can learn.

[0023] As an example, Figure 1 illustrates a schematic diagram of an example 100 of a machine learning model in one scheme. Referring to Figure 1, the machine learning model may at least include a recognition network 110, which may be, for example, a feed-forward network (FFN). The user may provide the speech to be recognized to the machine learning model, which is built at least on the recognition network 110, thereby using the recognition network 110 to recognize the speech content in order to generate text content corresponding to the speech content.

[0024] In speech recognition based on machine learning models, a single recognition network is typically used. The number of network parameters in the recognition network constitutes the upper limit of the knowledge that the machine learning model can learn. When faced with a large amount of new data, the machine learning model may not be able to effectively integrate this new information, thus limiting the effectiveness of retraining.

[0025] In view of this, the present disclosure provides a scheme for speech recognition. According to this scheme, firstly, based on the content of the target speech, a second set of recognition networks is determined from a first set of recognition networks of a trained machine learning model for recognizing target speech segments in the target speech. In the first set of recognition networks of the machine learning model, at least two recognition networks have different network parameters, and the second set of recognition networks includes at least one recognition network. Then, the recognition result of the second set of recognition networks for the target speech segment is determined. Subsequently, based on the recognition result, the text content for the target speech segment is determined.

[0026] As will be more clearly understood from the following description, the scheme disclosed herein transforms the machine learning model from a single recognition network into an architecture comprising multiple parallel recognition networks (i.e., the first set of recognition networks), with at least two of these networks having different network parameters. Compared to the traditional scheme shown in Figure 1, this scheme not only increases the number of recognition networks but also provides more parameter choices when recognizing target speech segments. This design substantially increases the number of parameters in the machine learning model, thereby significantly expanding the model's capacity.

[0027] When faced with the retraining requirements described above, this solution enables the machine learning model to demonstrate superior integration capabilities, more effectively absorbing new information and thus achieving better results in domain-specific retraining. Furthermore, when recognizing target speech content, this solution selects one or more of the most suitable recognition networks from the first set to form a "second set of recognition networks." These selected recognition networks may have diverse network parameters, working together to achieve accurate recognition of target speech segments. In this way, the machine learning model in this solution can comprehensively consider more dimensions of information during the recognition process, thereby achieving higher recognition accuracy.

[0028] The following will further describe in detail various example implementations of this scheme with reference to the accompanying drawings.

[0029] Figure 2 illustrates a schematic diagram of an environment 200 according to some embodiments of the present disclosure. Referring to Figure 2, the example environment 200 may include an electronic device 220.

[0030] In this example environment 200, the machine learning model may include E recognition networks 241-1, 241-2, ..., 241-E, where E is a positive integer. For clarity, recognition networks 241-1, 241-2, ..., 241-E will be collectively referred to as recognition network 241 below.

[0031] In some embodiments, a user may provide speech to be recognized to an electronic device 220. After receiving the speech to be recognized, the electronic device 220 takes the speech to be recognized as the target speech 210, and selects one or more recognition networks 241 from multiple recognition networks 241 in the machine learning model 240 to process each speech segment in the target speech 210, thereby obtaining text content 230 for each speech segment.

[0032] In environment 100 of Figure 2, a user can interact with electronic device 220 through a terminal device to provide the voice to be recognized. The terminal device can present an interface for interaction, and the content presented on the interface changes according to the user's interaction behavior / preset operation, etc.

[0033] In some embodiments, the terminal device can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device may also support any type of user-facing interface (such as "wearable" circuitry).

[0034] Electronic device 220 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 220 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in cloud environments, etc. Electronic device 220 can provide backend services for applications 130 that support content presentation in terminal devices.

[0035] A communication connection can be established between the electronic device 220 and the terminal device. This communication connection can be established via wired or wireless means. The communication connection may include, but is not limited to, Bluetooth connections, mobile network connections, Universal Serial Bus connections, and Wi-Fi connections; the embodiments of this disclosure are not limited in this respect. In the embodiments of this disclosure, the electronic device 220 and the terminal device can achieve signaling interaction through the communication connection between them.

[0036] It should be understood that the structure and function of the various elements in environment 200 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0037] Figure 3 shows a flowchart of an example process 300 of a method for speech recognition according to some embodiments of the present disclosure. Process 300 can be implemented at electronic device 220. Figure 4 shows a schematic diagram of an example 400 of recognizing target speech 210 according to some embodiments of the present disclosure. The process 300 shown in Figure 3 will be described below with reference to Figure 4.

[0038] In box 310, electronic device 220 determines a second set of recognition networks 2412 from a first set of recognition networks 2411 of a trained machine learning model 240, based on the content of the target speech 210, for recognizing target speech segments in the target speech 210. As an example, the machine learning model 240 may include E recognition networks 241-1, 241-2, ..., 241-E, where E is a positive integer. For clarity, recognition networks 241-1, 241-2, ..., 241-E are collectively referred to as recognition networks 241. In some embodiments, at least two recognition networks 2411 in the first set of recognition networks 2411 of the machine learning model 240 have different network parameters, and the second set of recognition networks 2412 includes at least one recognition network 241.

[0039] In some embodiments, the first set of recognition networks 2411 may be a collection of multiple recognition networks 241, which may be parallel and each network has its own characteristics, specifically in that their network parameters are different. For example, the difference in network parameters means that these recognition networks 241 have different capabilities in capturing and parsing speech features. For example, the first set of recognition networks 2411 may be constructed based on a Mixture of Experts (MoE) network, and the recognition networks 241 may include feedforward neural networks.

[0040] In some embodiments, when target speech 210 is input to electronic device 220, electronic device 220 splits the target speech 210 into one or more speech segments (e.g., a frame of content). Then, for any one of these speech segments (i.e., the target speech segment), electronic device 220 selects one or more recognition networks 241 from a first set of recognition networks 2411 that are most suitable for processing this speech segment. These selected recognition networks 241 are also known as a second set of recognition networks 2412. This selection process can be performed based on one or more factors such as the spectral characteristics, speech rate, pitch, and articulation of the target speech segment.

[0041] In some embodiments, the second set of recognition networks 2412 includes at least one recognition network 241. These selected recognition networks 241 will work together to analyze and recognize the target speech segment. Since these selected recognition networks 241 are chosen from the first set of networks with different parameters, they are able to capture more comprehensive information in the speech, thereby improving the accuracy of recognition.

[0042] It should be noted that the number of the first group of recognition networks 2411 and the second group of recognition networks 2412 shown in Figure 4 is only an example. Depending on actual needs, the first group of recognition networks 2411 and the second group of recognition networks 2412 may include more recognition networks 241. Furthermore, the speech segment is not limited to one frame of content; depending on actual needs, the speech segment may include more frames of content.

[0043] In some embodiments, before training the machine learning model 240, the network parameters of the first recognition network and the second recognition network in the first set of recognition networks 2411 are determined to be the same, and after training the machine learning model 240, the network parameters of the first recognition network and the second recognition network are updated to be different.

[0044] In some embodiments, the first identification network may refer to any one of the identification networks 241 in the first set of identification networks 2411, and the second identification network is an identification network 241 in the first set of identification networks 2411 that is different from the first identification network.

[0045] In some embodiments, at the start of training of the machine learning model 240, the first and second recognition networks in the first set of recognition networks 2411 are initialized with the same network parameters. This means that, initially, the two recognition networks 241 process the input data in a similar way and with similar results. Initializing them with the same network parameters can serve as a starting point to ensure that the various recognition networks 241 have a uniform and fair foundation at the start of training.

[0046] It should be clarified that the training process here can refer to the domain-specific retraining process mentioned earlier. Before this retraining process begins, the relevant parameters of the machine learning model 240 can inherit from a pre-trained open-source machine learning model (such as the machine learning model containing a single recognition network 110 shown in Figure 1). Therefore, when the machine learning model 240 enters the retraining process, the network parameters of both the first and second recognition networks can be initialized to be consistent with the network parameters of the recognition network 110 in the pre-trained open-source machine learning model.

[0047] With this initialization method, the retraining of the machine learning model 240 does not need to start from scratch, but can directly inherit the training results of the open-source machine learning model. This means that at the beginning of the retraining stage, the machine learning model 240 of this disclosure already has a certain recognition ability, thereby significantly shortening the training cycle and improving training efficiency. Therefore, although the embodiments of this disclosure increase the number of network parameters of the machine learning model 240, thanks to the inheritance of parameters from the open-source machine learning model, the training time does not increase significantly.

[0048] As the retraining process progresses, the first and second recognition networks update their network parameters based on the training data, loss function, and optimization algorithm. Thanks to the diversity and complexity of the training data, as well as the randomness of the optimization process, the network parameters of the first and second recognition networks gradually become different. This difference in network parameters allows the two recognition networks 241 to output different recognition results when processing the same input, thus providing richer information for subsequent classification or decision-making.

[0049] In some embodiments, the number of identification networks 241 in the second group of identification networks 2412 is less than the total number of identification networks 241 in the first group of identification networks 2411.

[0050] In some embodiments, the number of recognition networks 241 in the second set of recognition networks 2412 is less than the total number of recognition networks 241 in the first set of recognition networks 2411. This means that not all of the first set of recognition networks 2411 will be used to process the target speech segment. This approach helps to improve the targeting and efficiency of processing, because different speech segments may contain different features, which require different recognition networks 241 to process.

[0051] In some embodiments, the electronic device 220 determines target information for a target speech segment based on the content of the target speech 210. The target information instructs the machine learning model 240 to determine the probability that each recognition network 241 in the first set of recognition networks 2411 is a second set of recognition networks 2412. Then, based on the target information, the electronic device 220 selects the second set of recognition networks 2412 from the first set of recognition networks 2411.

[0052] In some embodiments, the target information may be a set of probability values, each probability value corresponding to one of the recognition networks 241 in the first set of recognition networks 2411, indicating the probability that the recognition network 241 is selected as a member of the second set of recognition networks 2412 for the target speech segment.

[0053] As an example, referring to Figure 4, the first set of recognition networks 2411 is constructed based on E recognition networks 241. To reduce the computational load during training and inference, embodiments of this disclosure employ an efficient selection mechanism. Specifically, for a target speech segment, embodiments of this disclosure do not use all E recognition networks 241 for computation, but only select the E' recognition networks 241 with the highest usage probability. Here, E' is less than E, meaning that the number of recognition networks 241 actually used for computation is less than the total number of recognition networks 241.

[0054] In this way, embodiments of this disclosure can skip the low-probability recognition network 241. This approach keeps computational overhead within a certain range. This is crucial for tasks such as handling large-scale datasets or real-time speech recognition, as it allows the model to maintain high performance while also being computationally efficient.

[0055] In some embodiments, the electronic device 220, based on the content of each speech segment in the target speech 210, uses the linear layer 242 of the machine learning model 240 to determine the speech feature representation of each speech segment in the target speech 210 for each recognition network 241 in the first set of recognition networks 2411, so as to obtain a target feature representation matrix. Then, the electronic device 220 determines the target information for the target speech segment based on the target feature representation matrix.

[0056] In some embodiments, the linear layer 242 is capable of effectively classifying and representing target speech segments. Specifically, upon receiving target speech 210, the electronic device 220 can divide the target speech 210 into multiple speech segments based on the linear layer 242. Then, the electronic device 220 can use the linear layer 242 to classify each speech segment within each speech segment, and the classification result can be represented using feature representations, which will be discussed below.

[0057] As an example, for a recognition network 241, given input k∈R T×d The output feature h∈R can be obtained through the recognition network 241. T×d This process can be represented by Equation 1: h = f(k); Equation 1

[0058] Where k can be the spectral feature representation of the target speech 210 (e.g., the Mel spectral feature matrix), R T×d Let R be a real number field of dimension T×d, where T represents the time frame length of the feature, which can be understood as the number of time steps or the sequence length of the input data, and d represents the feature dimension of each time step.

[0059] As an example, for a set of recognition networks 241, given input k∈R T×d After passing through a set of recognition networks 241, the output feature g∈R can be obtained. T×d This process can be represented by Equation 2: g = MoE(k); Equation 2

[0060] Through the processing of the linear layer 242, the electronic device 220 can obtain a target feature representation matrix. Each row of this matrix corresponds to the feature representation of a speech segment (or a frame of content), while each column corresponds to a recognition network 241.

[0061] As an example, each row in the target feature representation matrix can be a set of real numbers, where each real number represents a score of the recognition network 241 in the first set of recognition networks 2411 selected to recognize a speech segment. These scores can be calculated based on factors such as the performance, historical accuracy, and confidence of the recognition network 241. Determining the scores helps in determining the probability of each recognition network 241 being selected as the second set of recognition networks 2412 in subsequent steps.

[0062] As an example, electronic device 220 can utilize linear layer 242, through a weight matrix W∈R d×E The product of the input k and the target feature representation matrix is ​​constructed. Here, E represents the number of the first set of recognition networks 2411.

[0063] As an example, the target feature representation matrix can be represented by Equation 3: s = kW; Equation 3

[0064] Where s represents the target feature representation matrix.

[0065] In this way, electronic device 220 can effectively extract and represent the features of each speech segment in target speech 210 by utilizing the linear layer 242 of machine learning model 240. This not only helps improve the accuracy of speech recognition, but also provides strong support for the subsequent selection of recognition network 241.

[0066] In some embodiments, the electronic device 220 determines target information for the target speech segment based on the speech feature representations related to the target speech segment in the target feature representation matrix and by utilizing the first normalization layer 2431 in the machine learning model 240.

[0067] In some embodiments, the first normalization layer 2431 can normalize the input data (e.g., the output of the linear layer 242) using a normalization operation (e.g., a softmax operation) to conform to a specific form, thereby facilitating processing by subsequent network layers (i.e., the first set of recognition networks 2411). Through the processing of the first normalization layer 2431, the electronic device 220 can determine the target information for the target speech segment based on the speech feature representations related to the target speech segment in the target feature representation matrix. The target information can indicate the probability that each recognition network 241 in the first set of recognition networks 2411 is selected as a member of the second set of recognition networks 2412 for the target speech segment.

[0068] As described above, each row in the target feature representation matrix corresponds to the feature representation of a speech segment (or a frame of content), while each column corresponds to a recognition network 241. Furthermore, each row in the target feature representation matrix can be a set of real numbers (each real number can be a score of each recognition network 241 for a given speech segment). Through the first normalization layer 2431, this set of real numbers can be mapped to a set of probability values. This process can be represented by Equation 4:

[0069] Where s(t,i) represents the score of the i-th recognition network 241 for the t-th speech segment (or the content of the t-th frame), e s(t,i) This indicates that the score of the i-th recognition network 241 with respect to the t-th speech segment is exponentialized. p represents the exponential sum of the scores of all recognition networks 241 for the t-th speech segment. t (i) represents the probability of selecting the i-th recognition network 241 to recognize the t-th speech segment.

[0070] In this way, the electronic device 220 can dynamically select a suitable set of recognition networks 241 (i.e., a second set of recognition networks 2412) for each speech segment based on the target feature representation matrix. This mechanism can improve the flexibility and adaptability of the machine learning model 240, enabling it to better handle complex speech data.

[0071] In box 320, electronic device 220 determines text content 230 for the target speech segment based on the recognition results of the second set of recognition networks 2412.

[0072] In some embodiments, the electronic device 220 determines candidate recognition results for the target speech segment by each recognition network 241 in the second set of recognition networks 2412 to obtain a set of candidate recognition results. Then, the electronic device 220 determines the weights of each recognition network 241 in the second set of recognition networks 2412 for the target speech segment. Subsequently, based on the set of candidate recognition results and the weights of each recognition network 241 in the second set of recognition networks 2412, the electronic device 220 determines the text content 230 for the target speech segment.

[0073] In some embodiments, when processing a target speech segment, the electronic device 220 first determines candidate recognition results through each recognition network 241 in the second set of recognition networks 2412. Each recognition network 241 in these recognition networks recognizes the target speech segment and generates a candidate recognition result, thereby constituting a set of candidate recognition results. Subsequently, the electronic device 220 determines the weights of each recognition network 241 in the second set of recognition networks 2412 for the target speech segment.

[0074] As an example, electronic device 220 can determine the weights of each recognition network 241 based on the probability (or usage probability) of each recognition network 241 being selected as part of the second set of recognition networks 2412. For instance, electronic device 220 can use a second normalization layer 2432 to normalize these usage probabilities (e.g., a softmax operation) based on the usage probabilities of each recognition network 241, thereby converting the usage probabilities of each recognition network 241 into weights. Furthermore, electronic device 220 can combine a set of candidate recognition results with the weights of each recognition network 241 in the second set of recognition networks 2412, and determine the text content 230 for the target speech segment through a weighted fusion process.

[0075] As an example, the weighted fusion process can be represented by Formula 5:

[0076] Where g(t) represents the calculation result obtained after weighted fusion of the t-th speech segment frame, E' represents the number of recognition networks 241 in the second set of recognition networks 2412 for the t-th speech segment, and f i (k(t)) represents the recognition result of the i-th recognition network 241 on the spectral feature representation k(t) of the t-th speech segment. As an example, this recognition result can be a classification label, a predicted value, a feature vector, etc.

[0077] In this way, the electronic device 220 comprehensively utilizes the knowledge of multiple recognition networks 241 to determine the text content 230 of the target speech segment, thereby improving the accuracy of speech recognition.

[0078] The preparation and training of the machine learning model 240 in the embodiments of this disclosure will be described below.

[0079] As described above, when the machine learning model 240 enters the retraining process, the network parameters of the first recognition network and the second recognition network in the first set of recognition networks 2411 can be initialized to be consistent with the network parameters of the recognition network 110 in the pre-trained machine learning model.

[0080] In some embodiments, the parameters of other structures in the machine learning model 240 besides the recognition network 241 can be initialized using the same parameters as those of a pre-trained open-source machine learning model. Furthermore, each of the recognition networks 241 in the first group of recognition networks 2411 can be initialized using the same parameters as the recognition network 110 at the beginning of the retraining process. At this point, the spectral feature representation k(t) for a given t-th speech segment can be expressed by Equation 6:

[0081] Combining formula 1, we get g(t) = h(t). In this way, the machine learning model 240 of this embodiment exhibits performance completely consistent with that of a pre-trained open-source machine learning model in the initial stage of the retraining process. This provides a better starting point for the subsequent model retraining process.

[0082] In some embodiments, the electronic device 220 trains the machine learning model 240 based at least on a balanced loss function, wherein the balanced loss function is configured to, in at least one training iteration, instruct the machine learning model 240 to adjust the probability that each identification network 241 is currently identified as the second identification network 2412 based on the number of times each identification network 241 in the first identification network 2411 has been identified as the second identification network 2412 before determining the second set of identification networks 2412.

[0083] In some embodiments, the machine learning model 240 of this disclosure can be an end-to-end machine learning model 240. In addition to the various network layers described above, the machine learning model 240 of this disclosure may also include attention-based network layers, etc. As an example, the open-source machine learning model may include any open-source model, such as the Paraformer model and the Whisper model, etc.

[0084] In some embodiments, the speech recognition training dataset is denoted as D = {S1, S2, ..., S...} N}, where N is the number of samples in the labeled dataset. For any labeled data sample S n ={x n,y n |n∈[1,N]}, where x n and y n These represent the audio feature matrix (e.g., the Mel spectrum feature matrix) and the corresponding text annotation for the nth training data sample, respectively.

[0085] As an example, a speech recognition training dataset D can be constructed using 100,000 hours of speech data. Optionally, the audio feature matrix can be an 80-dimensional Vermeer spectral feature matrix, where each frame (or each speech segment) has a duration of 25 ms and a step size of 10 ms. Of course, other parameter settings are also feasible, and the embodiments of this disclosure are not limited in this respect.

[0086] In some embodiments, the machine learning model 240 of this disclosure is trained using a speech recognition training dataset D. During training, the loss function used in each iteration includes at least two parts. The first part is the speech recognition loss function. This function can, for example, be the same as the non-autoregressive loss function used in the pre-training phase of the open-source Paraformer model. Loss Function This can be expressed by Formula 7:

[0087] in, Given input audio features x n The output text y predicted by machine learning model 240 n The cross-entropy loss function, Given input audio features x n The length of the output text predicted by machine learning model 240 |y n | loss function.

[0088] Furthermore, in order to balance the usage probability of each recognition network 241 in the first set of recognition networks 2411, a balanced loss function is introduced. Balanced loss function This can be expressed by formula 8:

[0089] in, It represents the usage frequency of the i-th recognition network 241 with the maximum usage probability, where 1 is the indicator function. Let be the average usage probability of the i-th recognition network 241.

[0090] Finally, the loss function used in each iteration This can be expressed by Formula 9:

[0091] Where α is the weighting coefficient, and for example, α = 0.01. It should be noted that the explanation of the weighting coefficient here is only an example, and the weighting coefficient can be other values ​​according to actual needs.

[0092] In some embodiments, the training process may employ at least one of the following methods: Adaptive Moment Estimation (ADAM) optimization algorithm or a warm-up strategy. As an example, the peak learning rate of the warm-up strategy may be set to 0.0002, and the step size may be set to 10,000 steps. In other embodiments, other training parameters or other training algorithms may be set according to specific application needs, and the embodiments of this disclosure do not impose limitations in this regard. It should be noted that the descriptions of the peak learning rate and the step size of the warm-up strategy here are merely examples; depending on actual needs, the peak learning rate and the step size of the warm-up strategy may also be other values.

[0093] It should be noted that the strategies and parameters used in model training described above are merely illustrative examples. Depending on actual needs, different parameters and strategies can be employed when training the machine learning model 240, and the combinations of these parameters and strategies can be numerous.

[0094] As can be clearly understood from the various embodiments described above, according to the embodiments of this disclosure, the machine learning model 240 is transformed from a single recognition network 241 into an architecture containing multiple parallel recognition networks 241 (i.e., the first group of recognition networks 2411), and at least two of these recognition networks 241 have different network parameters. Compared with the conventional scheme shown in Figure 1, the embodiments of this disclosure not only increase the number of recognition networks 241, but also provide more parameter choices when recognizing target speech segments. This design substantially increases the number of parameters of the machine learning model 240, thereby significantly expanding the capacity of the model.

[0095] When faced with the retraining requirements described above, embodiments of this disclosure enable the machine learning model 240 to exhibit superior integration capabilities, more effectively absorbing new information and thus achieving better results in domain-specific retraining. Furthermore, when recognizing the content of the target speech 210, embodiments of this disclosure select one or more of the most suitable recognition networks 241 from the first set of recognition networks 2411 to form a "second set of recognition networks 2412". These selected recognition networks 241 may have different network parameters, working together to achieve accurate recognition of the target speech segments. In this way, the machine learning model 240 of embodiments of this disclosure can comprehensively consider more dimensions of information during the recognition process, thereby achieving higher recognition accuracy.

[0096] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 5 shows a schematic structural block diagram of an apparatus 500 for speech recognition according to some embodiments of this disclosure. The apparatus 500 may be implemented as or included in an electronic device 220. The various modules / components in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.

[0097] Referring to Figure 5, the apparatus 500 includes a recognition network determination module 510 and a text content generation module 520. The recognition network determination module 510 is configured to determine a second set of recognition networks for recognizing target speech segments from a first set of recognition networks in a trained machine learning model, based on the content of the target speech. In the first set of recognition networks of the machine learning model, at least two recognition networks have different network parameters, and the second set of recognition networks includes at least one recognition network. The text content generation module 520 is configured to determine text content for the target speech segment based on the recognition results of the target speech segment by the second set of recognition networks.

[0098] In some embodiments, the identification network determination module 510 is further configured to: determine target information for a target speech segment based on the content of the target speech, wherein the target information indicates the probability that the machine learning model will identify each identification network in the first group of identification networks as a second group of identification networks; and select a second group of identification networks from the first group of identification networks based on the target information.

[0099] In some embodiments, the recognition network determination module 510 is further configured to: determine the speech feature representation of each speech segment in the target speech for each recognition network in the first set of recognition networks based on the content of each speech segment in the target speech and using the linear layer of the machine learning model, so as to obtain the target feature representation matrix; and determine the target information for the target speech segment based on the target feature representation matrix.

[0100] In some embodiments, the identification network determination module 510 is further configured to: determine target information for the target speech segment based on the speech feature representations related to the target speech segment in the target feature representation matrix and using a first normalization layer in a machine learning model.

[0101] In some embodiments, before training the machine learning model, the network parameters of the first recognition network and the second recognition network in the first set of recognition networks are determined to be the same, and after training the machine learning model, the network parameters of the first recognition network and the second recognition network are updated to be different.

[0102] In some embodiments, the apparatus 500 further includes a training module configured to train a machine learning model at least based on a balanced loss function, wherein the balanced loss function is configured to, in at least one training iteration, instruct the machine learning model to adjust the probability of each recognition network currently being identified as a second recognition network based on the number of times each recognition network in the first recognition network has been identified as a second recognition network before determining the second group of recognition networks.

[0103] In some embodiments, the text content generation module 520 is further configured to: determine the candidate recognition results of each recognition network in the second set of recognition networks for the target speech segment, so as to obtain a set of candidate recognition results; determine the weight of each recognition network in the second set of recognition networks for the target speech segment; and determine the text content for the target speech segment based on the set of candidate recognition results and the weight of each recognition network in the second set of recognition networks.

[0104] In some embodiments, the number of identification networks in the second group of identification networks is less than the total number of identification networks in the first group of identification networks.

[0105] Figure 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. This electronic device 600 may, for example, be used to implement the electronic device 220 shown in Figure 2. It should be understood that the electronic device 600 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein.

[0106] Referring to Figure 6, the electronic device 600 is in the form of a general-purpose electronic device. Components of the electronic device 600 may include, but are not limited to, one or more processors 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 may be a physical or virtual processor and is capable of performing various processes according to the program stored in memory 620. In a multiprocessor system, multiple processors 610 execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.

[0107] Electronic device 600 typically includes multiple computer storage media. Such media can be any available media accessible to electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and / or data and accessible within electronic device 600.

[0108] Electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.

[0109] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 600 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0110] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 600, or with any device that enables electronic device 600 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0111] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0112] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0113] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0114] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0116] Various implementations of this disclosure have been described above. The foregoing description is exemplary and not exhaustive, nor is it limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is determined to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for speech recognition, comprising: determining, based on content of a target speech, a second set of recognition networks for recognizing a target speech segment in the target speech from a first set of recognition networks of a trained machine learning model, wherein network parameters of at least two recognition networks in the first set of recognition networks of the machine learning model are different, and wherein the second set of recognition networks comprises at least one recognition network; and determining, based on a recognition result of the target speech segment by the second set of recognition networks, a text content for the target speech segment.

2. The method of claim 1, wherein determining, based on content of a target speech, a second set of recognition networks for recognizing a target speech segment in the target speech from a first set of recognition networks of a trained machine learning model, comprises: determining, based on content of the target speech, target information for the target speech segment, the target information indicating probabilities that the machine learning model determines respective recognition networks in the first set of recognition networks as the second set of recognition networks; and selecting, based on the target information, the second set of recognition networks from the first set of recognition networks.

3. The method of claim 2, wherein determining target information for the target speech segment, comprises: determining, based on content of respective speech segments in the target speech, speech feature representations of respective speech segments in the target speech for respective recognition networks in the first set of recognition networks using a linear layer of the machine learning model to obtain a target feature representation matrix; and determining, based on the target feature representation matrix, the target information for the target speech segment.

4. The method of claim 3, wherein determining, based on the target feature representation matrix, the target information for the target speech segment, comprises: determining, based on speech feature representations in the target feature representation matrix that are related to the target speech segment, the target information for the target speech segment using a first normalization layer in the machine learning model.

5. The method of claim 1, wherein network parameters of a first recognition network and a second recognition network in the first set of recognition networks are determined to be the same before training of the machine learning model, and the network parameters of the first recognition network and the second recognition network are updated to be different after training of the machine learning model.

6. The method of claim 1, further comprising: training the machine learning model based on at least a balanced loss function, wherein the balanced loss function is configured to, in at least one training iteration, instruct the machine learning model to adjust a probability that a respective recognition network in the first set of recognition networks is currently determined as the second set of recognition networks based on a number of times that the respective recognition network has been determined as the second set of recognition networks before the determination of the second set of recognition networks.

7. The method of claim 1, wherein determining, based on the recognition result, a text content for the target speech segment, comprises: ​ ​ ​ determining candidate recognition results of the target speech segment by each recognition network in the second set of recognition networks to obtain a set of candidate recognition results; determining weights of the target speech segment by each recognition network in the second set of recognition networks; and determining the text content of the target speech segment based on the set of candidate recognition results and the weights of each recognition network in the second set of recognition networks. 8.The method of claim 1, wherein a number of recognition networks in the second set of recognition networks is less than a total number of recognition networks in the first set of recognition networks. 9.An apparatus for speech recognition, comprising: a recognition network determination module configured to determine, based on content of a target speech, a second set of recognition networks for recognizing a target speech segment in the target speech from a first set of recognition networks of a trained machine learning model, wherein network parameters of at least two recognition networks in the first set of recognition networks of the machine learning model are different, and wherein the second set of recognition networks comprises at least one recognition network; and a text content generation module configured to determine a text content of the target speech segment based on recognition results of the target speech segment by the second set of recognition networks. 10.An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1 to 8. 11.A computer-readable storage medium having computer-executable instructions stored thereon, the computer-executable instructions executable by a processor to implement the method according to any one of claims 1 to 8. 12.A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8. ​

Citation Information

Patent Citations

  • Training and / or using a language selection model for automatically determining language for speech recognition of spoken utterance

    CN112673421A

  • Mixed voice processing method and device, computer equipment and storage medium

    CN118038887A

  • Systems and methods for response selection in multi-party conversations with dynamic topic tracking

    US20210375280A1

  • Display apparatus and operating method thereof

    US20230267934A1

  • Speech recognition using topic-specific language models

    US9324323B1