Display device and voice recognition method

By introducing text information with the same or similar pronunciations during the speech recognition model training process, generating target label text and conducting multiple iterative training, the problem of low accuracy of the speech recognition system when processing new words is solved, and the accuracy of speech recognition is improved.

CN120452424AInactive Publication Date: 2025-08-08HISENSE ELECTRONIC TECH (WUHAN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510394148.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing speech recognition system has low recognition accuracy when processing new words, especially when film and television content is updated quickly, new words cannot be overwritten in the language model of the speech recognition system, resulting in a significant decrease in recognition accuracy when processing voice commands containing new words.

Method used

By introducing text information with the same or similar pronunciations during the speech recognition model training process, the target labeled text is generated, and the model is optimized through multiple iterative training, so that the model can better learn the mapping relationship between speech data and labeled text, and reduce the word error rate.

Benefits of technology

In the case of context bias, the accuracy of speech recognition is improved, the word error rate is effectively reduced, and the model's recognition ability in processing new words is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452424A_ABST
    Figure CN120452424A_ABST
Patent Text Reader

Abstract

The invention provides a display device and a voice recognition method, and the method comprises the steps: obtaining an initial annotation text in a training corpus corresponding to voice data, and recognizing characters contained in the initial annotation text; determining whether the characters are replaced according to a preset probability; under the condition that replacement is needed, searching a replacement candidate list of the to-be-replaced character from the replacement table, and randomly selecting a target character from the replacement candidate list to replace the to-be-replaced character; generating a target annotation text according to the initial annotation text, the to-be-replaced character and the target character; and training a voice recognition model according to the target labeled text, and recognizing the to-be-recognized voice according to the trained voice recognition model. According to the method, character information with the same or similar pronunciation is introduced in the speech recognition model training process, so that the model can recognize characters in speech more accurately in the recognition process, the word error rate is effectively reduced under the condition of context bias, and the problem of low recognition accuracy in the current speech recognition process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of display devices, and in particular to a display device and a voice recognition method. Background Art

[0002] Speech recognition is widely used in scenarios such as smart TVs and car audio control. By converting voice commands into text, it enables convenient device operation. For example, on smart TVs, users can control the content played on the TV through voice commands such as "Change the channel" or "Play XiX Ji." However, with the rapid update of film and television content, new words, such as new film and television series titles, are constantly emerging. These new words are often not covered by the language model of the speech recognition system. As a result, the recognition accuracy of the speech recognition system drops significantly when processing voice commands containing new words, affecting the user experience.

[0003] By combining the acoustic model with the language model, the speech recognition system can convert the input speech into a series of candidate text sequences and output the recognition result with the highest probability. In the actual speech recognition process, the acoustic model and the language model can be combined by converting the language model stored in the arpa file into the format of the WFST graph. During the speech decoding process, the WFST graph receives the input phone sequence (or syllable, or word, etc.), and finds a path with the best score in the decoding graph. The text sequence corresponding to the path is the final recognition result. In addition, in order to solve the problem of new word recognition, the speech recognition industry has proposed contextual biasing (also known as hot word enhancement) technology, which inputs new words into the speech recognition engine through the configuration interface to improve its recognition rate. In some embodiments, a shallow fusion (ShallowFusion) technical solution can be used to enhance specific words, such as adding an additional edge representing a hot word to the WFST graph and making its score better.

[0004] While existing technologies can meet the needs of speech recognition to a certain extent, they still have shortcomings when processing new words. Acoustic models exhibit a "peaking effect" when outputting classification probabilities. This means the model is overconfident about the most likely classification result, resulting in extremely low probabilities for similar-sounding candidate words. This makes them difficult to retain during the decoding process, making it difficult to improve recognition rates through hot word enhancement. Consequently, speech recognition currently suffers from low recognition accuracy. Summary of the Invention

[0005] Some embodiments of the present application provide a display device and a speech recognition method to solve the problem of low recognition accuracy in the current speech recognition process and improve the speech recognition rate under context bias.

[0006] In a first aspect, some embodiments of the present application provide a display device, including:

[0007] a display configured to display a user interface;

[0008] A user interface configured to obtain raw voice data;

[0009] The controller is configured as:

[0010] Obtaining initial annotated text in a training corpus corresponding to the original speech data, and identifying characters contained in the initial annotated text;

[0011] Determining whether to perform replacement on the text according to a preset probability;

[0012] When it is determined according to the preset probability that the character needs to be replaced, searching a replacement candidate list of the character to be replaced from a replacement table, and randomly selecting a target character from the replacement candidate list to replace the character to be replaced; the character to be replaced is the character to be replaced in the characters; the replacement candidate list includes a plurality of characters that have the same pronunciation as or a similar pronunciation to the character to be replaced;

[0013] Generate a target marked text according to the initial marked text, the character to be replaced and the target character;

[0014] A speech recognition model is trained according to the target annotated text, and the speech to be recognized is recognized according to the trained speech recognition model.

[0015] The above technical solution has the following advantages or beneficial effects: text information with the same or similar pronunciation is introduced during the training process of the speech recognition model, so that the model can more accurately recognize the text in the speech during the recognition process, effectively reduce the word error rate under the condition of context bias, and solve the problem of low recognition accuracy in the current speech recognition process.

[0016] In some embodiments, before determining whether to replace the text according to a preset probability, the controller is further configured to:

[0017] Setting a validation set, and setting multiple groups of candidate values for screening the preset probabilities;

[0018] Generate training annotation text according to the candidate value;

[0019] Training the speech recognition model according to the training annotated text;

[0020] Comparing word error rates under different candidate values in the validation set;

[0021] Determine that the candidate value corresponding to the lowest word error rate is the preset probability.

[0022] The above technical solution has the following advantages or beneficial effects: By comparing the word error rates for different candidate values on a validation set, the impact of different parameter settings on model performance can be quantitatively evaluated. Through comparison, the candidate value with the lowest word error rate is selected as the final preset probability. In this way, the model performs best on the validation set under this parameter setting. By setting a validation set, trying different candidate values, generating annotated text, training the model, evaluating the word error rate, and selecting the optimal parameters, the performance of the speech recognition model can be systematically optimized, making it more accurate and reliable in practical applications.

[0023] In some embodiments, before determining whether to replace the text according to a preset probability, the controller is further configured to:

[0024] Obtaining first specification information of the training corpus and second specification information of the replacement table;

[0025] Acquiring feature information of the speech recognition model;

[0026] The preset probability is determined according to the first specification information, the second specification information and the feature information.

[0027] The above technical solution has the following advantages or beneficial effects: by comprehensively analyzing the characteristic information of the training corpus, the replacement table and the speech recognition model, the preset probability parameters of the model can be optimized to determine an optimal preset probability value.

[0028] In some embodiments, before the step of searching the replacement table for a candidate replacement list of the word to be replaced, the controller is further configured to:

[0029] Constructing the replacement table; the replacement table includes an original word field and multiple replacement word fields; the original word field is used to store all Chinese characters; the replacement word field is used to store characters with the same pronunciation or a similar pronunciation to the Chinese characters;

[0030] Obtaining a dictionary data set containing all Chinese characters, and writing the characters in the dictionary data set into the original character field;

[0031] A candidate list set of characters with the same pronunciation as or similar pronunciation to the Chinese character is obtained, and the candidate list set is written into the replacement character field to generate the replacement candidate list.

[0032] The above technical solution has the following advantages or beneficial effects: when predicting, the model will not only output the text that the model believes to be correct with the highest probability, but also output text with the same or similar pronunciation, thereby increasing the probability of subsequent hot word recognition taking effect, and ultimately improving the recognition rate of the speech recognition model under context bias.

[0033] In some embodiments, before the step of obtaining a candidate list of characters with the same pronunciation as or similar pronunciation to the Chinese character, the controller is further configured to:

[0034] Obtaining the original word filled in the original word field;

[0035] Parsing original character information of the original character; the original character information at least includes flat and retroflex tongue information, tone information, nasal sound information and easily mixed sound information;

[0036] A candidate list set corresponding to the original word is determined according to the original word information.

[0037] The above technical solution has the following advantages or beneficial effects: by analyzing the voice feature information of the original word, a selection list with similar or easily confused pronunciation to the original word is generated, which can provide a basis for text replacement, thereby improving the accuracy and reliability of speech recognition.

[0038] In some embodiments, the controller generates a target annotated text according to the initial annotated text, the character to be replaced, and the target character, and is specifically configured to:

[0039] Identifying the position of the character to be replaced in the initial annotated text;

[0040] adding a location marker for the location;

[0041] Obtain the target text corresponding to the character to be replaced;

[0042] Based on the position mark, the to-be-replaced character is replaced with the target character to generate the target annotated text.

[0043] The above technical solution has the following advantages or beneficial effects: By generating target annotated text, modified annotated data can be provided for speech recognition model training. This modified annotated data enables the model to learn a wider range of pronunciations and character combinations during training, thereby improving the model's adaptability to different speech inputs and recognition accuracy. Furthermore, the generation of the target annotated text is based on the initial annotated text and replacement operations, ensuring the rationality and validity of the annotated data and providing reliable data support for subsequent model training.

[0044] In some embodiments, the speech recognition model is a CTC-based neural network model, and the controller trains the speech recognition model according to the target annotated text, and is specifically configured to:

[0045] Extracting audio features from the original speech data;

[0046] Inputting the target annotated text and the audio features into the speech recognition model;

[0047] The speech recognition model is trained through multiple rounds of iterations so that the speech recognition model learns acoustic features with the same or similar pronunciations.

[0048] The above technical solution has the following advantages or beneficial effects: During the training process, the target annotated text and its corresponding audio file are input into the speech recognition model. Through multiple rounds of training, the model parameters are continuously optimized, enabling it to better learn the mapping relationship between the speech data and the annotated text. After multiple iterations of training, the trained speech recognition model is output.

[0049] In some embodiments, after determining whether to perform the step of replacing the text according to a preset probability, the controller is further configured to:

[0050] When it is determined according to the preset probability that the text does not need to be replaced, the text is ignored.

[0051] The above technical solution has the following advantages or beneficial effects: the display device can avoid unnecessary replacement operations, avoid semantic distortion caused by forced replacement, and enhance the robustness of speech recognition.

[0052] In some embodiments, the preset probability ranges from 0.1 to 0.2.

[0053] The above technical solution has the following advantages or beneficial effects: the display device can effectively control the scope and degree of text replacement, thereby ensuring that the model can learn enough text information with the same or similar pronunciation while avoiding problems such as convergence difficulties caused by excessive replacement.

[0054] In a second aspect, some embodiments of the present application provide a speech recognition method, which can be applied to the display device of the first aspect, the display device including a display and a controller, the method comprising:

[0055] Obtaining initial annotated text in a training corpus corresponding to the speech data, and identifying characters contained in the initial annotated text;

[0056] Determining whether to perform replacement on the text according to a preset probability;

[0057] When it is determined according to the preset probability that the character needs to be replaced, searching a replacement candidate list of the character to be replaced from a replacement table, and randomly selecting a target character from the replacement candidate list to replace the character to be replaced; the character to be replaced is the character to be replaced in the characters; the replacement candidate list includes a plurality of characters that have the same pronunciation as or a similar pronunciation to the character to be replaced;

[0058] Generate a target marked text according to the initial marked text, the character to be replaced and the target character;

[0059] A speech recognition model is trained according to the target annotated text, and the speech to be recognized is recognized according to the trained speech recognition model.

[0060] The above technical solution has the following advantages or beneficial effects: the method introduces text information with the same or similar pronunciation during the training process of the speech recognition model, so that the model can more accurately recognize the text in the speech during the recognition process, effectively reduce the word error rate under the condition of context bias, and solve the problem of low recognition accuracy in the current speech recognition process.

[0061] As can be seen from the above technical solutions, some embodiments of the present application provide a display device and a speech recognition method, the method comprising: obtaining an initial annotated text in a training corpus corresponding to speech data, identifying the text contained in the initial annotated text; determining whether to perform a replacement of the text based on a preset probability; if replacement is required, searching a replacement candidate list of the character to be replaced from a replacement table, and randomly selecting a target text from the replacement candidate list to replace the character to be replaced; generating a target annotated text based on the initial annotated text, the character to be replaced, and the target text; training a speech recognition model based on the target annotated text, and recognizing the speech to be recognized based on the trained speech recognition model. The method introduces text information with the same or similar pronunciation in the speech recognition model training process, so that the model can more accurately recognize the text in the speech during the recognition process, effectively reduce the word error rate under the condition of context bias, and solve the problem of low recognition accuracy in the current speech recognition process. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate some embodiments of the present application or technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0063] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application;

[0064] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments of the present application;

[0065] Figure 3 A schematic diagram of software configuration of a display device provided in some embodiments of the present application;

[0066] Figure 4A schematic diagram of the acoustic model structure based on initials and finals in the hybrid model method provided in some embodiments of the present application;

[0067] Figure 5 A schematic diagram of the structure of an acoustic model based on Chinese characters in the hybrid model method provided in some embodiments of the present application;

[0068] Figure 6 Schematic diagram of the effect of storing the language model provided in some embodiments of the present application as an arpa file;

[0069] Figure 7 A schematic diagram of the structure of a WFST graph used in conjunction with initials and finals provided in some embodiments of the present application;

[0070] Figure 8 A schematic diagram of the structure of a WFST graph used in conjunction with Chinese characters provided in some embodiments of the present application;

[0071] Figure 9 A schematic diagram showing the effect of adding hot words to a WFST graph provided in some embodiments of the present application;

[0072] Figure 10 This is a schematic diagram of the output effect of an AM model with a spike effect shown in some embodiments of the present application;

[0073] Figure 11 This is a schematic diagram illustrating an ideal output effect of an AM model according to some embodiments of the present application;

[0074] Figure 12 A timing diagram of a display device performing a voice recognition method according to some embodiments of the present application;

[0075] Figure 13 A schematic diagram of a flow chart for setting a preset probability for a display device provided in some embodiments of the present application;

[0076] Figure 14 A schematic diagram of a flow chart of setting a preset probability for a display device provided in other embodiments of the present application;

[0077] Figure 15 A schematic diagram of a flow chart of a display device determining a candidate list set according to some embodiments of the present application. DETAILED DESCRIPTION

[0078] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.

[0079] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.

[0080] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.

[0081] The terms "comprise," "include," and "have," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.

[0082] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.

[0083] In the embodiments of the present application, the display device 200 generally refers to a device capable of displaying images and processing data. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.

[0084] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG, a user can operate the display device 200 through touch operation, the mobile terminal 300, and the control device 100. The control device 100 is used to receive operation instructions input by the user and convert the operation instructions into control instructions that the display device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus pen, a handle, etc.

[0085] The mobile terminal 300 can function as a control device for performing human-computer interaction between a user and the display device 200. The mobile terminal 300 can also function as a communication device for establishing a communication connection with the display device 200 and exchanging data. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, enabling connection and communication via a network communication protocol, enabling one-to-one control operations and data communication. Audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 for synchronized display.

[0086] In some embodiments, the mobile terminal 300 or other electronic devices can also simulate the functions of the control device 100 by running an application program for controlling the display device 200 .

[0087] like Figure 1 As shown in FIG, the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0088] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart TV, Internet Protocol television (IPTV), etc.

[0089] Figure 2 Some embodiments of this application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.

[0090] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.

[0091] In some embodiments, detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 may include a light receiver, such as a sensor for collecting ambient light intensity; or an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or a sound collector, such as a microphone, for receiving external sounds.

[0092] In some embodiments, the display 260 includes a display component for presenting images and a driver component for driving image display. The display 260 is configured to receive image signals output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces.

[0093] In some embodiments, the communication device 220 is a component used to communicate with an external device or server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 depending on the supported communication methods. For example, if the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including WiFi functionality. If the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including Bluetooth functionality.

[0094] The communication device 220 can establish a communication connection between the display device 200 and an external device or server 400 via a wireless or wired connection. A wired connection can connect the display device 200 to an external device via a data cable, an interface, or other components. A wireless connection can connect the display device 200 to an external device via a wireless signal or wireless network. The display device 200 can establish a connection with an external device directly or indirectly through a gateway, router, or connection device.

[0095] In some embodiments, the controller 250 may include at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processor, and a power processor, and first to nth interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in a memory. The controller 250 controls the overall operation of the display device 200.

[0096] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0097] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).

[0098] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may further be provided with an external audio output terminal, through which the audio output device may be connected to the display device 200 to output the sound of the display device 200.

[0099] In some embodiments, the user input interface 280 may be configured to receive instructions from a user.

[0100] To facilitate user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources of the display device 200. The operating system can control the display device to provide a user interface. For example, the operating system can directly control the display device to provide a user interface, or it can provide a user interface by running an application program. The operating system also allows the user to interact with the display device 200.

[0101] It should be noted that the operating system may be a native operating system based on a specific operating platform, or a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for the display device.

[0102] The operating system can be divided into different modules or layers according to the functions implemented, e.g. Figure 3 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely, the application layer (referred to as "application layer"), the application framework layer (referred to as "framework layer"), the system library layer and the kernel layer.

[0103] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer can host at least one application, which can include built-in window programs, system settings programs, clock programs, and the like, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0104] The framework layer provides applications with an application programming interface (API) and programming framework. The application framework layer includes predefined functions. The application framework layer acts as a processing center, determining the actions taken by applications in the application layer. Through the API, applications can access system resources and services during execution.

[0105] like Figure 3As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to the system location service; a package manager for retrieving various information related to the application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0106] In some embodiments, the activity manager is used to manage the lifecycle of each application and common navigation back functions, such as controlling application exit, opening, and back. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling changes in display windows, such as shrinking, shaking, or distorting the display window.

[0107] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.

[0108] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 3 As shown, the kernel layer can be configured with hardware drivers, and the drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0109] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.

[0110] The above embodiment illustrates the hardware / software architecture and functional implementation of a display device 200. Based on this display device 200, not only can various media such as movies, television programs, and images be output, but it can also provide speech recognition capabilities. Speech recognition, also known as automatic speech recognition (ASR), aims to convert the content contained in speech signals into computer-readable input, such as text sequences.

[0111] Speech recognition systems are usually composed of an acoustic model (AM) and a language model. The language model is trained using a large amount of scene-related text, such as commands such as "change the channel" and "play XiXji" in smart TV scenarios. However, with the rapid update of film and television content, new words such as the titles of new film and television dramas continue to appear. These new words are often not covered by the language model, resulting in a significant decrease in the recognition accuracy of the speech recognition system when processing voice commands containing new words. For example, if the user does not know the title of the new film and television drama "YuXyao" in advance, it is difficult to accurately recognize it based on pronunciation alone. To solve this problem, the speech recognition industry has proposed contextual biasing (also known as hot word enhancement) technology, which inputs new words (such as "YuXyao") into the speech recognition engine through a configuration interface to improve its recognition rate.

[0112] In some embodiments, speech recognition methods may include a hybrid speech recognition method based on a hybrid speech recognition model and an end-to-end speech recognition method based on an end-to-end model. Hybrid-based speech recognition solutions support streaming recognition, which means they can output recognition results while the speaker is speaking. Therefore, hybrid model solutions are more widely used in human-computer interaction scenarios (such as voice-activated television and car voice control).

[0113] In some embodiments, the hybrid model consists of two parts: an acoustic model and a language model. The acoustic model is responsible for mapping the speech frame sequence into the probability of the acoustic modeling unit, while the language model generates the probability distribution of the word sequence by counting a large amount of text data. In the smart TV scenario, the language model is usually trained based on the text data of the user's voice control instructions, such as "I want to watch XX channel", "Play the tenth episode of West XX", etc. By combining the acoustic model with the language model, the speech recognition system can convert the input speech into a series of candidate text sequences and output the recognition result with the highest probability. The above method can meet the needs of speech recognition in most scenarios, but it still has shortcomings when processing new words. The following is a detailed explanation in combination with actual usage scenarios.

[0114] In some embodiments, a hybrid model-based speech recognition engine generally consists of two parts: an acoustic model and a language model. The input of the acoustic model is a sequence of raw speech frames segmented into fixed time intervals, typically 10 to 30 milliseconds, and the output is the probability of each speech frame corresponding to an acoustic modeling unit. Commonly used modeling units for Mandarin Chinese recognition include phone, syllable, and character.

[0115] Figure 4 A schematic diagram of the acoustic model structure based on initials and finals in the hybrid model method provided in some embodiments of the present application is shown in FIG. Figure 4 As shown, the acoustic model can include an input layer, a hidden layer, and an output layer. The input layer can segment the input speech signal into fixed time intervals, such as 10-30 milliseconds, and assign it to hidden layers, such as Layer 1 and Layer 2. The output layer outputs the probability that each segmented speech frame corresponds to an acoustic modeling unit contained in the acoustic model. The acoustic model can be implemented using a deep neural network (DNN). The language model outputs linguistic scores for different text sequences. The acoustic model and language model, when combined, can convert the input speech into a series of possible text candidates, assigning a corresponding probability value to each text sequence. This probability incorporates both the acoustic model's and the language model's probabilities.

[0116] Figure 5 Schematic diagram of the acoustic model structure based on Chinese characters in the hybrid model method provided in some embodiments of the present application, Figure 4 The biggest difference is that the AM model, an acoustic model based on Chinese characters, directly determines which character the input speech segment belongs to, rather than which phone (initials and finals). Its modeling granularity is larger, and when training the model, it does not require the text pinyin sequence of training data, making it easier to train. Therefore, its scope of application is wider than that of the phone-based AM model.

[0117] In some embodiments, the function of the language model is to output linguistic scores for different text sequences. The acoustic model and the language model are combined to convert the input speech into a series of possible candidate texts. Each text sequence is given a corresponding probability value, which includes the probability of the acoustic model and the probability of the language model. The format of the language model is generally to record the probability of a single word, two words, three words, etc. together. A typical language model is usually stored in the format of an ARPA file. The ARPA file is a standard format that exists in the language model and can be obtained through training texts such as the language model training tool srilm. In some embodiments, the format of the language model can be to record the probability of a single word, two words, three words, etc. together. A typical language model can be stored in the format of an ARPA file.

[0118] For example, Figure 6 Schematic diagram of the effect of storing the language model provided in some embodiments of the present application as an arpa file, such as Figure 6 As shown, an ARPA file can be generated by recording the probabilities of connecting multiple words, such as "don't want an APP" and "will use an APP." These probabilities can be obtained from statistics on a large amount of text. Text can be obtained by collecting a large amount of text used by users to control the TV with voice commands. For example, "I want to watch TV station A," "change channel," "turn up the volume," "play the tenth episode of drama A," etc. These texts can be segmented using a word segmentation tool, and word frequency can be counted to calculate the probability of different words connecting together.

[0119] In some embodiments, when inputting a speech segment and identifying what is said in the speech segment, both the acoustic model and the language model must be considered. For example, if a person is heard saying, "w o3 sh i4 f u2n an2r en2" (the number represents the tone of the vowel, there are 5 tones, 5 represents light). Although from the perspective of the acoustic model, it sounds more like "I am from Funan", but there are no place names like "Funan" and "Flan" in the language model. Therefore, the final result of the above sentence is "I am from Hunan".

[0120] During speech recognition, both the acoustic model and the language model can be considered using the following encoding scheme. First, the language model stored in the ARPA file is converted into a WFST graph format. During speech decoding (i.e., recognition), the WFST graph receives the input phone sequence (or syllable, word, etc.) and searches for the path in the decoding graph that achieves the best score. The text sequence corresponding to this path is the final recognition result.

[0121] Figure 7 This is a schematic diagram of the structure of a WFST graph used in conjunction with consonants and finals provided in some embodiments of the present application. Figure 8 A schematic diagram of the structure of a WFST diagram used in conjunction with Chinese characters provided in some embodiments of this application, combined with Figure 7 and Figure 8 , Figure 7 You need to enter the phone, and Figure 4 Used in conjunction with the AM model in Figure 8 You need to enter a word, and Figure 5 The AM model in Figure 4 AM combination Figure 7 Take the WFST in as an example to illustrate the process of speech recognition.

[0122] The first step is to input a very simple audio, such as 400 milliseconds, such as the word "hello". The second step is to divide the input voice into 4 frames (this is simplified for easy understanding. In practice, one frame is usually 10ms, and 400ms is divided into at least 40 frames). Each frame is input to Figure 4 The AM model predicts which phone (i.e., the initials and finals in Mandarin) the frame corresponds to; in the third step, the first frame is input into the AM, and the AM outputs the score corresponding to each initial and final (considering 5 tones, the total number of initials and finals is about 200). For example, the score corresponding to "n" is 0.2, the score corresponding to "l" is 0.5, the score corresponding to "ao" is 1.2, and so on. The score is obtained by taking the logarithm of the probability (this is done to prevent a series of very small decimals from being multiplied to produce extremely small values), and then multiplying it by -1. Therefore, the smaller the score, the higher the probability; in the fourth step, each of the remaining frames is input into the AM in turn to obtain the score of each initial and final corresponding to each frame. The fifth step is to list all the phone sequences corresponding to the four frames in the form of permutations and combinations, including the correct sequence "n i3 h ao3", some sequences with similar pronunciations, such as "l i3 h ao3", and some completely irrelevant sequences that do not even exist in Mandarin, such as "nnnn", "i3 i3 i3 i3", etc. AM will give each sequence a score, which is the sum of the scores of each frame corresponding to this phone (because the score comes from the logarithm of the probability, so the sum is equal to the multiplication of the probability); the sixth step is to list each phone sequence in the fifth step. Figure 7 Find the corresponding sequence on the WFST graph in , and the corresponding sequence has a corresponding output sequence (such as Figure 7The arc from state 1 to state 2 is labeled "i3:you / 1.2" to indicate that the input is i3 and the output is you, corresponding to a score of 1.2. In step 7, the score of the phone sequence from step 5 is added to the score on the WFST graph from step 6 to obtain the score of the final output text sequence. The final result is the sequence with the lowest score (meaning the highest probability) among all the sequences.

[0123] It is understandable that if Figure 5 The AM model that outputs Chinese characters uses a similar decoding process as above. The output of the third step can be changed to the probability score corresponding to each Chinese character.

[0124] In the above speech recognition algorithm, the fifth step uses a permutation and combination method to list all phone sequences, taking into account both acoustic and language scores. However, since most speech is affected by accents or noise, it is not possible to directly take the best acoustic score result. In actual use scenarios, due to computing resource limitations, it is difficult to implement permutation and combination to list all possible sequences. For example, based on the phone-based AM model, there are about 200 4 There are about 6000 possibilities for the word-based AM model, 4 speech frames 4 The CPU and memory are unable to support this possibility. Therefore, in the actual decoding process, pruning is required to only consider the sequence with the current best score and the sequence with a score difference within a certain range (adjustable hyperparameters). This requires performing a beam search.

[0125] The speech recognition system using the above solution can better solve the speech recognition needs of most scenarios. However, in some areas, such as smart TV scenarios, new film and television dramas are continuously launched, and their names are not common nouns before they are broadcast, such as "The Legend of Zhen X". The speech recognition system cannot correctly recognize the characters corresponding to their pronunciation. For example, if you hear its corresponding pronunciation, you cannot correctly correspond to the words "The Legend of Zhen X". In order to solve this kind of problem, you can use a recognition enhancement solution based on context bias. This solution sets new words to the recognition engine through the API interface, and uses speech bias-related solutions to quickly enhance the recognition effect of new words. The enhancement process generally takes only a few minutes.

[0126] In some embodiments, a shallow fusion technology solution can be used to enhance specific words, such as adding an additional edge representing the hot word on the WFST graph and making its score better. Figure 9 This is a schematic diagram of the effect of adding hot words to the WFST graph provided in some embodiments of the present application, such as Figure 9As shown, taking the original word "hello" as an example, suppose a new word "li hao" appears. Without any enhancement, most likely when a user pronounces "li hao", it will be recognized as "hello". Based on the shallow fusion hot word enhancement method, an additional edge representing the hot word will be added to the WFST graph, and its score will be made better. For example, Figure 9 The WFST graph showing the addition of the hot word "li hao".

[0127] However, the prerequisite for the success of shallow fusion is that when the AM outputs a frame corresponding to a character with the best probability, characters with the same or similar pronunciation can also be output with a relatively small probability, but if the probability is too small, it will be pruned by the beam search. Currently, the character-based AM model generally uses the CTC loss function for training, and there is a "spike effect", that is, it is very confident in the output classification, and the category with the best classification probability occupies a probability of more than 0.999, and the probability values of other categories are very small, almost approaching 0. This may cause the hot word to be enhanced to be pruned in the beam search process due to too small probability and cannot be recognized.

[0128] Exemplarily, Figure 10 Schematic diagram of the output effect of the AM model with spike effect shown in some embodiments of this application. Taking a certain frame of speech as an example, after Figure 5 the AM model, the probability values corresponding to each character can be as Figure 10 shown. The significant problem with the probability that the AM model classifies a frame of speech into each character is the "spike effect" problem, that is, it will allocate the vast majority of probability values to the character that the model considers to be the most correct, and for other characters, regardless of whether the pronunciation is close, the probability that can be obtained is extremely small. If Figure 5 the AM model in Figure 8 is combined with the WFST in

[0129] In some embodiments, for the hot word enhancement based on shallow fusion, it is hoped that when the AM model outputs the classification probability, it will also allocate some probabilities to characters with the same or similar pronunciation. In this way, subsequently, the probability corresponding to the hot word can be increased on the WFST graph, and finally the path containing the hot word can be output. Figure 11 Schematic diagram of the ideal output effect of the AM model shown in some embodiments of this application. Combining Figure 10 and Figure 11 , Figure 11 For Figure 10 The speech is based on the ideal AM model output of shallow fusion technology, and Figure 10 The main difference is that in addition to outputting the correct "you", the AM model can also output text that is close to the pronunciation of "you" with a lower probability.

[0130] In other words, to improve the recognition rate of hot words and increase the probability of including similar-sounding words in the AM model, the AM model needs to be optimized so that when outputting classification probabilities, it also assigns some probability to words with the same or similar pronunciations. This allows us to subsequently increase the probability of hot words on the WFST graph and output paths containing hot words. However, this AM model struggles to achieve ideal output performance through standard training methods (typically CTC training), resulting in low recognition accuracy in current speech recognition.

[0131] Based on this, some embodiments of the present application provide a display device 200, which includes a display 260, a user interface and a controller 250. The display 260 is configured to display a user interface, and the user interface is configured to obtain original voice data; the controller 250 enables the display device 200 to execute a voice recognition method by running an application. The display device 200 modifies the annotated text of the training corpus and randomly replaces the text to be replaced with another word with the same or similar pronunciation to reduce the CTC spike effect of the AM model based on word modeling. In addition to outputting the correct word with the highest probability, the CTC can also output other words with the same or similar pronunciation with a lower probability, that is, improve the output of homophonic or similar words, improve the voice recognition rate under context bias, and solve the problem of low recognition accuracy in the current voice recognition process.

[0132] In order to facilitate the understanding of the technical solutions in some embodiments of the present application, each step is described in detail below in conjunction with some specific embodiments and drawings. Figure 12 A timing diagram of a display device performing a voice recognition method provided in some embodiments of the present application, such as Figure 12 As shown, in some embodiments, when the display device 200 performs the voice recognition method, the following steps may be included, and the specific contents are as follows:

[0133] Step S1: obtaining the initial annotated text in the training corpus corresponding to the original speech data, and identifying the characters contained in the initial annotated text.

[0134] In some embodiments, during the AM model training process, some training corpus needs to be prepared and input into the model. Therefore, the display device 200 can first obtain the training corpus corresponding to the original speech data from the speech data training corpus. These training corpuses can be stored in the form of a data list (DataList), with each row containing a training corpus. Each training corpus contains at least three parts: a unique ID number, an audio file path, and the corresponding text in the audio file. For example, the content of the training corpus can be as shown in Example 1 below:

[0135] 001 wav_0001.wav "Good morning"

[0136] 002 wav_0002.wav "Call Mom"

[0137] 003 wav_0003.wav "Power on"

[0138] 004 wav_0004.wav "Open China XX Channel"

[0139] 005 wav_0005.wav "I want to watch TV"

[0140] 006 wav_0006.wav "Play Xixi Ji"

[0141] 007 wav_0007.wav "Play Liu Xhua's music"

[0142] 008 wav_0008.wav "I want to watch XX TV Sports Channel"

[0143] 009 wav_0009.wav "Play Zhen X Biography"

[0144] 010 wav_0010.wav "Louder"

[0145] 011 wav_0011.wav "Screen Brighter"

[0146] 012 wav_0012.wav "Open system settings"

[0147] 013 wav_0013.wav "Shutdown"

[0148] 014 wav_0014.wav "Fast forward ten minutes"

[0149] 015 wav_0015.wav "Play Hello XX English"

[0150] The above is an example of Datalist for end-to-end speech recognition model training, which contains audio and Chinese character annotation files. A training corpus may be represented as "001wav_0001.wav "Good morning"". By reading these training corpora, the initial annotated text can be obtained, that is, the corresponding text content in the audio file, such as "Good morning". At the same time, the initial annotated text is analyzed to identify each character contained therein, in preparation for subsequent processing steps. For example, the Chinese characters in the annotated text can be extracted one by one through a text parsing tool to form a text sequence to be processed. In this way, the display device 200 can accurately extract the initial annotated text and the text information it contains from the training corpus, providing basic data for subsequent text replacement operations. After step S1 is executed, the following step S2 may be included.

[0151] Step S2: Determine whether to perform text replacement based on a preset probability.

[0152] In some embodiments, after obtaining the initial annotated text and the characters it contains, each character needs to be judged to determine whether it needs to be replaced. For each epoch of training, the annotated characters can be modified according to the following process.

[0153] After traversing the characters in each annotated text in the training corpus, the display device 200 can traverse each character in an annotated text in turn, and decide whether to modify the character according to a preset probability α. Prior to this, the display device 200 can first set a preset probability α, which can be a set hyperparameter used to control the probability of each character being replaced, affecting the results of model training. If α is set too high, the characters in the annotated text may be replaced too much, and the model may not be able to learn correctly, resulting in convergence difficulties; if α is set too low, the replacement is insufficient, and the spike effect of CTC cannot be effectively reduced, so the training effect is not obvious. Therefore, the set α must not only meet the results of model training, but also reduce the spike effect.

[0154] Figure 13 A flow chart of setting a preset probability for a display device provided in some embodiments of the present application, such as Figure 13 As shown, in some embodiments, before determining whether to perform a text replacement based on a preset probability, the display device 200 can set a preset probability α in the following manner. The display device 200 can first set a validation set and multiple sets of candidate values for screening the preset probability. Then, the display device 200 can generate training annotated text based on the candidate values and train a speech recognition model based on the training annotated text. The word error rates under different candidate values are then compared in the validation set, and the candidate value corresponding to the lowest word error rate is determined to be the preset probability.

[0155] Exemplarily, the validation set can be a portion of data independently divided from the training data, which is used to evaluate the performance of the model during the training process. In an embodiment of the present application, the validation set is used to test the word error rate (WER) of the model under different α values to ensure that the selected α has generalization ability. For setting multiple sets of candidate α values, the impact of different parameter settings on model performance can be analyzed. For example, the candidate values can be (such as 0.05, 0.1, 0.15, 0.2, 0.25). The selection of candidate values can cover the range from low replacement rate to high replacement rate. Different candidate values may cause the model to perform differently in terms of recognition accuracy, recall rate, etc. By trying multiple candidate values, the optimal parameter combination can be found, thereby improving the overall performance of the model. Afterwards, training annotation text is generated according to different candidate values to ensure that the training data matches the parameter settings of the model. The generated annotation text is then used to train the speech recognition model so that the model can learn the mapping relationship between speech signals and text. The word error rate WER is an important indicator for measuring the performance of a speech recognition model, which indicates the proportion of words that the model recognizes incorrectly to the total number of words. By comparing word error rates for different candidate values on a validation set, we can quantitatively assess the impact of different parameter settings on model performance. By comparing the candidate value with the lowest word error rate, we select the final preset probability. This parameter setting ensures the model performs best on the validation set. By setting a validation set, trying different candidate values, generating annotated text, training the model, evaluating word error rates, and selecting the optimal parameters, we can systematically optimize the performance of the speech recognition model, making it more accurate and reliable in real-world applications.

[0156] Figure 14 A flow chart of setting a preset probability for a display device provided in other embodiments of the present application is shown in FIG. Figure 14 As shown, in some embodiments, before determining whether to perform a text replacement based on the preset probability, the display device 200 may further set a preset probability α in the following manner. The display device 200 may first obtain first specification information of a training corpus and second specification information of a replacement table (used to replace text in annotated text; details of the replacement table will be described later), then obtain feature information of a speech recognition model, and then determine the preset probability based on the first specification information, the second specification information, and the feature information.

[0157] Exemplarily, the first specification information may include the size of the corpus (such as the total number of words in the corpus, the number of sentences, etc.), the distribution of the corpus (such as the proportion of texts in different topics or fields), the complexity of the corpus (such as the diversity of vocabulary, the complexity of sentence structure, etc.), etc. The second specification information may include the size of the replacement table (such as the number of replacement rules), the coverage of the replacement rules (such as the types of vocabulary involved), etc. The feature information may be the feature parameters used by the speech recognition model in the training and recognition process, such as acoustic features (such as Mel-frequency cepstral coefficients MFCC), language model features, etc. These feature information determines how the model processes and understands speech signals. By comprehensively analyzing the feature information of the training corpus, replacement table and speech recognition model, the preset probability parameters of the model can be optimized to determine an optimal preset probability value.

[0158] In some embodiments, the preset probability value range can be 0.1 to 0.2. In this way, the scope and degree of text replacement can be effectively controlled, thereby ensuring that the model can learn enough text information with the same or similar pronunciation while avoiding problems such as convergence difficulties due to excessive replacement of the model.

[0159] After the preset probability α is set, for each character in the initial annotation text, it can be randomly judged whether the character needs to be replaced according to the preset probability α. In the process of judgment, it can be achieved through the random generator in the program. For example, a range of 1 / n is set to be replaced in a section of annotation text (n is a positive number), and a random label can be thrown. The value of the random label can be 1, 2, 3...n. If the value of the random label corresponding to the character is 1, then the character needs to be replaced. If the value of the random label corresponding to the character is any other number, then the character does not need to be replaced. Judgment can also be made in other ways, and this application does not make specific restrictions on this.

[0160] In some embodiments, when it is determined according to a preset probability that the text does not need to be replaced, the display device 200 can ignore the text, that is, if it is determined according to the random probability α that the text does not need to be modified, then the text is directly skipped and retained in its original state. In this way, the display device 200 can avoid unnecessary replacement operations, avoid semantic distortion caused by forced replacement, and enhance the robustness of speech recognition. Therefore, by introducing a preset probability α to control the randomness of text replacement, the uncertainty in the real speech environment can be simulated, so that the model can learn more diverse pronunciation situations during the training process, thereby improving the model's adaptability and robustness to different pronunciation situations. After judgment, if the text needs to be modified, the subsequent replacement process is entered, that is, the following step S3 is executed.

[0161] Step S3: When it is determined according to the preset probability that the character needs to be replaced, a replacement candidate list of the character to be replaced is searched from the replacement table, and a target character is randomly selected from the replacement candidate list to replace the character to be replaced.

[0162] In some embodiments, when a character is determined to need to be replaced based on a preset probability α, the display device 200 can search a replacement table, such as β, for a list of replacement candidates for the character to be replaced, search the replacement table for a list of replacement candidates for the character to be replaced, and randomly select a target character from the list of replacement candidates to replace the character to be replaced. The character to be replaced is the character in the text that needs to be replaced; the list of replacement candidates contains multiple characters that have the same or similar pronunciation as the character to be replaced. During the replacement, a character is randomly selected to replace the character to be replaced.

[0163] In some embodiments, before searching for a replacement candidate list for a character to be replaced in the replacement table, the display device 200 may construct a replacement table and a replacement candidate list in the following manner. The display device 200 may first construct a replacement table; the replacement table includes an original character field and multiple replacement character fields; the original character field is used to store all Chinese characters; the replacement character field is used to store characters that have the same pronunciation as or a similar pronunciation to the Chinese characters; a dictionary data set containing all Chinese characters is obtained, and the characters in the dictionary data set are written into the original character field; a candidate list set of characters that have the same pronunciation as or a similar pronunciation to the Chinese characters is obtained, and the candidate list set is written into the replacement character field to generate a replacement candidate list.

[0164] For example, referring to Table 1, when constructing the replacement table, the original word field in Table 1 and multiple replacement word fields such as Replacement 1 and Replacement 2 may be included. The original word can be obtained from a public dictionary dataset, such as a Chinese character dictionary, which may include the pinyin for each Chinese character. When filling the original word field, each character in the Chinese character dictionary may be traversed and then entered into the original word field of the replacement table β. The table consisting of multiple replacement fields outside the original word field can be understood as a replacement candidate list.

[0165] In some embodiments, each row of the replacement table β contains a Chinese character and its corresponding multiple homophonic or near - pronunciation characters. For example, for the Chinese character "你", its replacement candidate list may include "倪", "拟", "逆", "尼", "呢", "您", "里", "立", "例", etc. After finding the replacement candidate list of the to - be - replaced character, a target character is randomly selected from this candidate list to replace the to - be - replaced character. By randomly selecting a target character from the replacement table β to replace the to - be - replaced character, the display device 200 can introduce text information with the same or similar pronunciation during the training process, thereby weakening the peak effect of CTC. In this way, when the model makes a prediction, it will not only output the text that the model considers correct with the highest probability, but also output the text with the same or similar pronunciation, thereby increasing the probability of subsequent hot - word recognition taking effect, and ultimately improving the recognition rate of the speech recognition model in the context of bias.

[0166] Table 1: Replacement table β and replacement candidate list

[0167]

[0168]

[0169] Figure 15 It is a schematic flowchart of the process for the display device provided by some embodiments of this application to determine the candidate list set. As Figure 15 shown, in some embodiments, when the display device 200 determines the candidate list set, it can first obtain the original character filled in the original character field, and then parse the original character information of the original character; the original character information at least includes flat - tongue and retroflex - tongue information, tone information, nasal - sound information, and easily confused sound information, and then determine the candidate list set corresponding to the original character according to the original character information.

[0170] Exemplarily, referring to Table 2, after obtaining the original character, the display device 200 can perform a detailed speech feature analysis on the extracted original character, extract its original character information such as key speech information. Then, by parsing this information, more accurately find the candidate characters that are similar in pronunciation or easily confused with the original character. For example, in combination with Table 2, the display device 200 can combine flat - tongue and retroflex - tongue information such as the flat - tongue sounds "z", "c", "s" and the retroflex - tongue sounds "zh", "ch", "sh", combine nasal - sound information such as the nasal sounds "m", "n" and other phonemes, etc., screen out the easily confused phonemes or syllables, and then generate a list of candidate characters that are similar in pronunciation or easily confused with the original character based on the parsed speech feature information. After that, for each character in the β table, find the characters that are the same as it and those with similar pinyin, and fill them into the replacement character part after the table. In this way, by parsing the speech feature information of the original character and generating a candidate list set that is similar in pronunciation or easily confused with the original character, it can provide a basis for character replacement, thereby improving the accuracy and reliability of speech recognition.

[0171] Table 2: Similar pronunciation table

[0172] Original pronunciation Approximate pronunciation 1 Approximate pronunciation 2 Approximate pronunciation 3 Approximate pronunciation 4 Approximate pronunciation 5 ni3 nin3 ni2 li3 lin3 hao3 hao2 zhong1 zong1 zhong2 guo2 guo1 bei3 bei2 jing1 jin1

[0173] After step S3 is completed, the following step S4 may be executed.

[0174] Step S4: Generate target annotated text according to the initial annotated text, the characters to be replaced and the target characters.

[0175] After completing the text replacement operation, the display device 200 may generate a new target marked text according to the initial marked text, the character to be replaced, and the replaced target text.

[0176] Figure 15 A flow chart of a display device generating target annotation text according to some embodiments of the present application is provided, such as Figure 15 As shown, when generating the target annotated text, the display device 200 can first identify the position of the character to be replaced in the initial annotated text and add a position mark to the position; then, obtain the target text corresponding to the character to be replaced, and then replace the character to be replaced with the target text based on the position mark to generate the target annotated text.

[0177] Exemplarily, the display device 200 first needs to find the specific location of the word to be replaced in the initial annotation text, and mark the location of the word to be replaced in the initial annotation text. The location mark can be used to quickly locate the word to be replaced in subsequent steps to ensure the accuracy and efficiency of the replacement operation. The location mark can be a special symbol, label or index, etc. Afterwards, based on the previously added location mark, the location of the word to be replaced can be accurately found and replaced with the target text, thereby generating the target annotation text. In other words, the display device 200 replaces the word to be replaced in the initial annotation text with the target text, thereby obtaining a modified annotation text.

[0178] For example, the following example 2 is the target annotation text formed after the annotation text of the above example 1 is replaced:

[0179] 001 wav_0001.wav "Morning is still good"

[0180] 002 wav_0002.wav "Call Mom"

[0181] 003 wav_0003.wav "Power on"

[0182] 004 wav_0004.wav "Open XX channel"

[0183] 005 wav_0005.wav "I want to publish TV"

[0184] 006 wav_0006.wav "Interview with Reporter X"

[0185] 007 wav_0007.wav "Play the music of Liu X"

[0186] 008 wav_0008.wav "I especially value XX Thai Sports Channel"

[0187] 009 wav_0009.wav "Play X's biography"

[0188] 010 wav_0010.wav "Increase the volume"

[0189] 011 wav_0011.wav "Brighten the screen"

[0190] 012 wav_0012.wav "Open the system settings"

[0191] 013 wav_0013.wav "Shut down"

[0192] 014 wav_0014.wav "Fast forward ten minutes"

[0193] 015 wav_0015.wav "Play Li Hao XX English"

[0194] As described above, after one round of replacement, the training corpus becomes the content shown in Example 2. Among them, the bold font is the replaced part. For example, assume the initial labeled text is "Good morning", and if "早" is determined to be replaced and "尚" is randomly selected from the replacement table as the replacement target word, then the generated target labeled text is "尚上好". In this way, by generating the target labeled text, modified labeled data can be provided for the training of the speech recognition model. These modified labeled data enable the model to learn more diverse pronunciation situations and word combinations during training, thereby improving the model's adaptability and recognition accuracy for different speech inputs. At the same time, the generation of the target labeled text is based on the initial labeled text and replacement operations, which can ensure the rationality and effectiveness of the labeled data and provide reliable data support for subsequent model training. After step S4 is completed, the following step S5 can be included.

[0195] Step S5: Train the speech recognition model according to the target labeled text, and recognize the speech to be recognized according to the trained speech recognition model.

[0196] After the target labeled text is generated, the display device 200 can train the speech recognition model according to the target labeled text, and recognize the speech to be recognized according to the trained speech recognition model.

[0197] In some embodiments, the speech recognition model can be a neural network model based on Connectionist Temporal Classification (CTC), which allows the model to directly map from an input sequence (audio features) to an output sequence (text annotations) without the need for precise time alignment. When training the speech recognition model, the display device 200 can first extract the audio features from the original speech data, then input the target annotated text and audio features into the speech recognition model, and then iterate the speech recognition model for multiple rounds to enable the speech recognition model to learn acoustic features with the same or similar pronunciations.

[0198] Exemplarily, the display device 200 can first extract audio features from the original voice data, and these features are the input of the speech recognition model. Afterwards, the extracted audio features and the target annotated text can be input into the speech recognition model together. The target annotated text is the replaced text, which allows the model to learn the mapping relationship between audio features and text through supervised learning. Through multiple rounds of iterative training, it can be ensured that the model gradually converges on the training data and learns more accurate acoustic features and language models. In this way, during the training process, the target annotated text and its corresponding audio file are input into the speech recognition model, and the parameters of the model are continuously optimized through multiple rounds of training (each training cycle is called an epoch) so that it can better learn the mapping relationship between voice data and annotated text. After multiple iterative training, the trained speech recognition model is output.

[0199] In some embodiments, after the speech recognition model is trained, the trained speech recognition model can be used to recognize the speech to be recognized. For example, the speech to be recognized can be input into the trained model, and the model will analyze and recognize the input speech based on what it has learned during the training process, and output the corresponding text results. Since text information with the same or similar pronunciation is introduced during the training process, the model can more accurately recognize the text in the speech during the recognition process, especially in the case of context bias, which can effectively reduce the word error rate (WER), improve the accuracy and reliability of speech recognition, and thus solve the problem of low recognition accuracy in the current speech recognition process.

[0200] It's understandable that by optimizing the training process of the acoustic model AM, we can improve the accuracy, robustness, and decoding efficiency of the acoustic model and reduce error propagation. These improvements make the combination of the acoustic model and language model in the shallow fusion stage more effective, thereby increasing the success rate of shallow fusion and improving the overall performance of the speech recognition system.

[0201] As can be seen from the above technical solution, the above embodiment provides a display device 200, which obtains the initial annotated text in the training corpus corresponding to the speech data, and recognizes the text contained in the initial annotated text; determines whether to perform a replacement of the text according to a preset probability; when it is determined that the text needs to be replaced according to the preset probability, searches the replacement candidate list of the to-be-replaced character from the replacement table, and randomly selects a target character from the replacement candidate list to replace the to-be-replaced character; the to-be-replaced character is the character that needs to be replaced in the text; the replacement candidate list contains multiple characters that have the same or similar pronunciation as the to-be-replaced character; generates a target annotated text according to the initial annotated text, the to-be-replaced character and the target character; trains a speech recognition model according to the target annotated text, and recognizes the speech to be recognized according to the trained speech recognition model. The display device 200 can introduce text information with the same or similar pronunciation during the speech recognition model training process, so that the model can more accurately recognize the text in the speech during the recognition process, especially in the case of context bias, and can effectively reduce the word error rate (WER), improve the accuracy and reliability of speech recognition, and thus solve the problem of low recognition accuracy in the current speech recognition process.

[0202] Based on the above display device 200, some embodiments of the present application further provide a speech recognition method, which can be applied to the display device 200 in the above embodiment. In some embodiments, the method may include the following:

[0203] Obtaining initial annotated text in a training corpus corresponding to the speech data, and identifying characters contained in the initial annotated text;

[0204] Determining whether to perform replacement on the text according to a preset probability;

[0205] When it is determined according to the preset probability that the character needs to be replaced, searching a replacement candidate list of the character to be replaced from a replacement table, and randomly selecting a target character from the replacement candidate list to replace the character to be replaced; the character to be replaced is the character to be replaced in the characters; the replacement candidate list includes a plurality of characters that have the same pronunciation as or a similar pronunciation to the character to be replaced;

[0206] Generate a target marked text according to the initial marked text, the character to be replaced and the target character;

[0207] A speech recognition model is trained according to the target annotated text, and the speech to be recognized is recognized according to the trained speech recognition model.

[0208] It can be seen from the above technical solution that the above embodiment provides a speech recognition method, which can introduce text information with the same or similar pronunciation during the training process of the speech recognition model, so that the model can more accurately recognize the text in the speech during the recognition process, especially in the case of context bias, and can effectively reduce the word error rate (WER), improve the accuracy and reliability of speech recognition, and thus solve the problem of low recognition accuracy in the current speech recognition process.

[0209] The same and similar parts between the various embodiments in this specification can be referenced to each other and will not be repeated here.

[0210] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solutions in the embodiments of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention or certain portions of the embodiments.

[0211] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

[0212] For ease of explanation, the above description has been made with reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Based on the above teachings, various modifications and variations are possible. The above embodiments are selected and described to better explain the principles and practical applications, so that those skilled in the art can better utilize the embodiments and various different variations of the embodiments suitable for specific use considerations.

Claims

1. A display device, characterized in that: include: a display configured to display a user interface; A user interface configured to obtain raw voice data; The controller is configured as: Obtaining initial annotated text in a training corpus corresponding to the original speech data, and identifying characters contained in the initial annotated text; Determining whether to perform replacement on the text according to a preset probability; When it is determined according to the preset probability that the character needs to be replaced, searching a replacement candidate list of the character to be replaced from a replacement table, and randomly selecting a target character from the replacement candidate list to replace the character to be replaced; The character to be replaced is a character in the character to be replaced; the candidate replacement list contains a plurality of characters that have the same pronunciation as or a similar pronunciation to the character to be replaced; Generate a target marked text according to the initial marked text, the character to be replaced and the target character; A speech recognition model is trained according to the target annotated text, and the speech to be recognized is recognized according to the trained speech recognition model.

2. The display device according to claim 1, wherein Before the step of determining whether to replace the text according to the preset probability, the controller is further configured to: Setting a validation set, and setting multiple groups of candidate values for screening the preset probabilities; Generate training annotation text according to the candidate value; Training the speech recognition model according to the training annotated text; Comparing word error rates under different candidate values in the validation set; Determine that the candidate value corresponding to the lowest word error rate is the preset probability.

3. The display device according to claim 1, wherein Before the step of determining whether to replace the text according to the preset probability, the controller is further configured to: Obtaining first specification information of the training corpus and second specification information of the replacement table; Acquiring feature information of the speech recognition model; The preset probability is determined according to the first specification information, the second specification information and the feature information.

4. The display device according to claim 2 or 3, characterized in that Before the step of searching the replacement candidate list of the to-be-replaced word from the replacement table, the controller is further configured to: Constructing the replacement table; the replacement table includes an original word field and multiple replacement word fields; the original word field is used to store all Chinese characters; the replacement word field is used to store characters with the same pronunciation or a similar pronunciation to the Chinese characters; Obtaining a dictionary data set containing all Chinese characters, and writing the characters in the dictionary data set into the original character field; A candidate list set of characters with the same pronunciation as or similar pronunciation to the Chinese character is obtained, and the candidate list set is written into the replacement character field to generate the replacement candidate list.

5. The display device according to claim 4, wherein: Before the step of obtaining a candidate list of characters with the same pronunciation as or similar pronunciation to the Chinese character, the controller is further configured to: Obtaining the original word filled in the original word field; Parsing original character information of the original character; the original character information at least includes flat and retroflex tongue information, tone information, nasal sound information and easily mixed sound information; A candidate list set corresponding to the original word is determined according to the original word information.

6. The display device according to claim 1, wherein The controller generates a target annotated text according to the initial annotated text, the character to be replaced, and the target character, and is specifically configured to: Identifying the position of the character to be replaced in the initial annotated text; adding a location marker for the location; Obtain the target text corresponding to the character to be replaced; Based on the position mark, the to-be-replaced character is replaced with the target character to generate the target annotated text.

7. The display device according to claim 1, wherein The speech recognition model is a CTC-based neural network model. The controller trains the speech recognition model according to the target annotated text and is specifically configured as follows: Extracting audio features from the original speech data; Inputting the target annotated text and the audio features into the speech recognition model; The speech recognition model is trained through multiple rounds of iterations so that the speech recognition model learns acoustic features with the same or similar pronunciations.

8. The display device according to claim 1, wherein After the controller determines whether to perform the step of replacing the text according to a preset probability, the controller is further configured to: When it is determined according to the preset probability that the text does not need to be replaced, the text is ignored.

9. The display device according to claim 1, wherein The value range of the preset probability is 0.1 to 0.

2.

10. A speech recognition method, applied to the display device according to any one of claims 1 to 9, wherein the display device comprises a display and a controller, wherein: The method comprises: Obtaining initial annotated text in a training corpus corresponding to the speech data, and identifying characters contained in the initial annotated text; Determining whether to perform replacement on the text according to a preset probability; When it is determined according to the preset probability that the character needs to be replaced, searching a replacement candidate list of the character to be replaced from a replacement table, and randomly selecting a target character from the replacement candidate list to replace the character to be replaced; the character to be replaced is the character to be replaced in the characters; the replacement candidate list includes a plurality of characters that have the same pronunciation as or a similar pronunciation to the character to be replaced; Generate a target marked text according to the initial marked text, the character to be replaced and the target character; A speech recognition model is trained according to the target annotated text, and the speech to be recognized is recognized according to the trained speech recognition model.

Citation Information

Cited By

  • Speech recognition method and device and terminal equipment

    CN121838738A

  • A speech recognition method, device and terminal equipment

    CN121838738B