Speech recognition method and related device
Through periodic fine-tuning and multi-level vocabulary management, the problem of new hot word recognition errors and time-consuming in videos is solved, and efficient and accurate speech recognition and optimized user experience is achieved.
Patent Information
- Application Number
- CN202510694122.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-15
AI Technical Summary
Existing voice recognition technology is prone to errors and takes a long time to recognize new hot words in videos, and has poor user experience.
Through periodic fine-tuning training of the trained speech recognition model, new hot words are added and invalid hot words are eliminated, and the dynamic update mechanism of the hot word thesaurus is used, and the multi-level downgrade vocabulary management of hot words is combined to ensure the efficient utilization of model resources.
It improves the accuracy and speed of speech recognition, ensures timely recognition of new hot words, and optimizes the user experience.
Smart Images

Figure CN120496520A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a speech recognition method and related devices. Background Art
[0002] With the continuous development of deep learning technology, a method for generating video commentary based on speech recognition technology has emerged. However, when using acoustic models or language models for speech recognition, this method often encounters recognition errors for newly added hot words (such as character names in the video, artificial words, etc.). In addition, speech recognition using acoustic models and language models takes a long time, resulting in a poor user experience. Summary of the Invention
[0003] In view of the above problems, this application provides a speech recognition method and related devices to achieve the purpose of quickly and accurately identifying new hot words in speech recognition scenarios. The specific solution is as follows:
[0004] The first aspect of the present application provides a speech recognition method, comprising:
[0005] Obtain the user voice input during the user watching the target video;
[0006] Perform speech recognition on the user's voice using a trained speech recognition model to obtain a text comment corresponding to the user's voice;
[0007] In which, the speech recognition model is periodically fine-tuned and trained according to a preset fine-tuning cycle. The hot word library used in the fine-tuning training adds new hot words at each fine-tuning moment corresponding to the fine-tuning cycle, and eliminates invalid hot words at each demotion moment corresponding to a preset demotion cycle. The new hot words do not exist in the hot word library and their frequency of appearance in the interactive area associated with the target video is greater than a preset frequency threshold. The invalid hot words refer to words in the hot word library that meet the preset demotion conditions. The fine-tuning cycle is inversely correlated with the heat of the target video and / or the fine-tuning cycle is less than a preset cycle threshold.
[0008] In a possible implementation, the method further includes:
[0009] While removing the invalid hot words from the hot word library, the invalid hot words are added to a preset first degraded word library, so that when the target video or an associated video of the target video is played, if the playback volume meets a preset playback volume threshold, the invalid hot words in the first degraded word library are added back to the hot word library, wherein the associated video refers to a video whose content relevance to the target video is greater than a preset relevance threshold;
[0010] Invalid hot words that appear in the first degraded vocabulary for a duration reaching a preset first duration threshold are downgraded to the second degraded vocabulary, so that when the target video or the associated video of the target video is played, if the playback volume meets the preset playback volume threshold, in response to the user's control instruction, the invalid hot words in the second degraded vocabulary are re-added to the hot word vocabulary, wherein the newly added hot words do not exist in both the first degraded vocabulary and the second degraded vocabulary.
[0011] In a possible implementation, the process of acquiring the newly added hot words includes:
[0012] Acquiring interaction data within the interaction area;
[0013] Extracting, from the interaction data, words whose occurrence frequency is greater than the frequency threshold and are not in the hot word library, the first degraded word library, and the second degraded word library as candidate words;
[0014] Words related to the content of the target video are selected from the candidate words as the newly added hot words.
[0015] In a possible implementation, the degradation condition is one or more of the following conditions:
[0016] The number of times the target word appears in the interactive area in the current demotion period is not ranked in the top n%, where n>0. The target word refers to any word in the hot word library, and the current demotion period refers to the time period between the current demotion moment and the previous demotion moment.
[0017] The time between the current degradation moment and the last appearance of the target word in the interactive area is greater than a preset second time threshold;
[0018] The duration between the current degradation moment and the last completion of the target video playback is greater than a preset third duration threshold;
[0019] The weight of the target word is less than a preset weight threshold, wherein the weight of the target word is determined according to the last hit time of the target word and the number of times the target word appears in the interactive area within the current degradation period.
[0020] In a possible implementation, the weight of the target word is determined according to the last hit time of the target word and the number of times the target word appears in the interactive area within the current degradation period, including:
[0021] Calculating a hit rate weight according to the number of times the target word appears in the interactive area during the current degradation period;
[0022] Calculating a time decay weight according to the last hit time of the target word;
[0023] According to the hit rate weight and the time decay weight, a comprehensive weight is calculated as the weight of the target word.
[0024] In a possible implementation, the method further includes:
[0025] When the newly added hot word is added to the hot word database, a protection period is set for the newly added hot word so that the newly added hot word is not eliminated during the protection period, and / or an initial weight greater than or equal to the weight threshold is configured for the newly added hot word, and the initial weight becomes invalid when the time when the newly added hot word is added to the hot word database reaches a preset fourth time threshold, or the initial weight becomes invalid when it is less than the comprehensive weight.
[0026] In one possible implementation, performing speech recognition on the user's speech using a trained speech recognition model includes:
[0027] Obtaining a target availability state corresponding to the target video, wherein the target availability state is one of a first state, a second state, and a third state, wherein the network transmission capability of the first state is greater than a preset first capability threshold, and the speech recognition service quality is greater than a preset first quality threshold, wherein the network transmission capability of the third state is less than a preset second capability threshold, and the speech recognition service quality is less than a preset second quality threshold, wherein the first quality threshold is greater than the second quality threshold, and the first capability threshold is greater than the second capability threshold, and the second state refers to a state other than the first state and the third state;
[0028] When the target available state is the first state, performing speech recognition on the user speech by using the speech recognition model without pausing the playing of the target video;
[0029] When the target available state is the second state, the user voice is recognized by the voice recognition model while the target video is paused.
[0030] In a possible implementation, the method further includes:
[0031] When the target available state is the third state, the user is prompted that voice recognition is unavailable when the target video is paused, and a text input component is called so that the user can use the text input component to input text comments.
[0032] A second aspect of the present application provides a speech recognition device, comprising:
[0033] A voice acquisition module is used to acquire the user's voice input during watching the target video;
[0034] A speech recognition module is used to perform speech recognition on the user's speech using a trained speech recognition model to obtain a text comment corresponding to the user's speech;
[0035] In which, the speech recognition model is periodically fine-tuned and trained according to a preset fine-tuning cycle. The hot word library used in the fine-tuning training adds new hot words at each fine-tuning moment corresponding to the fine-tuning cycle, and eliminates invalid hot words at each demotion moment corresponding to a preset demotion cycle. The new hot words do not exist in the hot word library and their frequency of appearance in the interactive area associated with the target video is greater than a preset frequency threshold. The invalid hot words refer to words in the hot word library that meet the preset demotion conditions. The fine-tuning cycle is inversely correlated with the heat of the target video and / or the fine-tuning cycle is less than a preset cycle threshold.
[0036] The third aspect of the present application provides a computer program product, comprising computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements the speech recognition method of the first aspect or any implementation of the first aspect.
[0037] A fourth aspect of the present application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0038] The memory is used to store computer programs;
[0039] The processor is used to execute the computer program so that the electronic device can implement the speech recognition method of the first aspect or any implementation manner of the first aspect.
[0040] The fifth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can use the speech recognition method of the first aspect or any implementation of the first aspect.
[0041] By means of the above technical solution, the speech recognition method provided by the present application obtains the user voice input by the user while watching the target video, and performs speech recognition on the user voice through the trained speech recognition model to obtain the text barrage corresponding to the user voice. The fine-tuning period used for fine-tuning training of the speech recognition model provided by the present application is inversely correlated with the popularity of the target video. Since the more popular the target video is, the higher the possibility of generating new hot words and the greater the number, the new hot words are added to the hot word library and fine-tuning training is performed at each fine-tuning moment corresponding to the fine-tuning period. This can enable the speech recognition model to recognize the new hot words more timely and accurately, thereby improving the accuracy of speech recognition. In addition, setting the fine-tuning period to a smaller period that is less than the preset period threshold can also timely add new hot words to the hot word library and perform fine-tuning training on the speech recognition model based on the new hot words, thereby improving the accuracy of speech recognition.
[0042] Furthermore, the present application will remove invalid hot words from the hot word library at each degradation moment corresponding to the preset degradation cycle, which can effectively save the model resources that the speech recognition model relies on, thereby improving the speed and accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0044] Figure 1 A flowchart of a speech recognition method provided in this application;
[0045] Figure 2 A schematic diagram of the structure of a speech recognition device provided in this application;
[0046] Figure 3 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0047] The following describes the embodiments of the present application in conjunction with the accompanying drawings. The terms used in the implementation methods of the present application are only used to explain the specific embodiments of the present application and are not intended to limit the present application.
[0048] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0049] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0050] The present application provides a speech recognition method and related devices, which can be applied to a terminal with a speech recognition function, a server with a speech recognition function, and a system composed of the above-mentioned terminals and servers.
[0051] In order to enable those skilled in the art to better understand the present application, the following examples are provided for detailed description.
[0052] Reference Figure 1 , which is a flow chart of a speech recognition method provided in an embodiment of the present application, such as Figure 1 As shown, the speech recognition method may include:
[0053] Step S101: Acquire user voice input by a user while watching a target video.
[0054] Specifically, while watching a target video, a user may want to enter a comment. In this embodiment, a voice button is provided, allowing the user to input a user's voice by triggering the voice button, and then the user's voice is recognized to obtain a text comment corresponding to the user's voice. Here, the target video can be any video on the media playback device.
[0055] Optionally, the voice button may be a voice button on a display interface of the target video, or a voice button on a terminal device.
[0056] Optionally, the voice button can be a virtual button on the touch screen or a real physical button.
[0057] Of course, the above-mentioned voice buttons are only examples and are not intended to limit the present application.
[0058] Step S102: Perform speech recognition on the user's speech through the trained speech recognition model to obtain the text barrage corresponding to the user's speech.
[0059] Among them, the speech recognition model is periodically fine-tuned and trained according to a preset fine-tuning cycle. The hot word library used for fine-tuning training adds new hot words at each fine-tuning moment corresponding to the fine-tuning cycle, and eliminates invalid hot words at each demotion moment corresponding to the preset demotion cycle. The newly added hot words do not exist in the hot word library and their frequency of appearance in the interactive area associated with the target video is greater than a preset frequency threshold. Invalid hot words refer to words in the hot word library that meet the preset demotion conditions. The fine-tuning cycle is inversely correlated with the heat of the target video and / or the fine-tuning cycle is less than the preset cycle threshold.
[0060] Specifically, this embodiment provides a trained speech recognition model, and the user speech obtained above can be input into the speech recognition model to obtain the text barrage output by the model.
[0061] Here, the speech recognition model is periodically fine-tuned according to a preset fine-tuning cycle, and each fine-tuning training uses a hot word library as training data.
[0062] Optionally, the initial benchmark model used for fine-tuning training of the speech recognition model can be a domain large model injected with domain data related to the target video, such as the Qwen-Audio large model injected with domain data or the OpenAI Whisper large model, etc.
[0063] Optionally, the domain data can be the script, historical subtitles, and known character names of the target video.
[0064] It should be noted that the domain data can also be other, and this application does not specifically limit it.
[0065] Optionally, when periodically fine-tuning the domain model using the hot word library as training data, the Low-Rank Adaptation (LoRA) method can be used for training. That is, the parameter weights of the LoRA adapter can be adjusted to achieve the effect of the speech recognition model prioritizing high-weight words.
[0066] This embodiment uses a hot word library to perform lightweight fine-tuning training on a large domain model based on the LoRA method, which can avoid the resource consumption of full parameter training and save training time and resource consumption.
[0067] The hot word library is introduced as follows: In this embodiment, the hot word library at the initial moment is empty. New hot words (such as artificial words such as "ex-husband brother" and "hard control", specific nouns such as "county master", character names, etc.) can be added to the hot word library at each fine-tuning moment corresponding to the fine-tuning cycle. At the same time, invalid hot words can be removed from the hot word library at each demotion moment corresponding to the preset demotion cycle (as the popularity of the target video decreases, each new hot word may be converted into an invalid hot word) to keep the hot word library dynamically updated. Here, the new hot word does not exist in the hot word library and its frequency of appearance in the interactive area associated with the target video is greater than the preset frequency threshold. Invalid hot words refer to words in the hot word library that meet the preset demotion conditions.
[0068] Optionally, the interactive area associated with the target video can be the comment area and / or bullet screen area of the target video, or the comment area and / or bullet screen area of an associated video of the target video. An associated video of the target video refers to a video whose content relevance to the target video exceeds a preset relevance threshold. For example, if the target video is the first episode of a television series, the associated videos can be other episodes of the series, or they can be behind-the-scenes clips, feature film clips, promotional videos, and so on.
[0069] In one possible implementation, the fine-tuning cycle can be inversely correlated with the popularity of the target video, that is, the more popular the target video, the smaller the fine-tuning cycle. This approach can ensure that when the target video is popular, all new hot words within the fine-tuning cycle can be identified more promptly, and the speech recognition model can be fine-tuned and trained in a timely manner, thereby improving the real-time recognition capability of the speech recognition model for new hot words, thereby improving the speech recognition accuracy of the speech recognition model.
[0070] Optionally, the process of determining the popularity of a target video may include: performing a weighted summation based on the target video's basic data indicators, user interaction indicators, content quality indicators, and communication effect indicators to obtain the target video's popularity. Here, basic data indicators include, but are not limited to, the following indicators: number of views, number of likes, number of comments, number of reposts, and number of favorites; user interaction indicators include, but are not limited to, the following indicators: completion rate and interaction rate. The completion rate refers to the proportion of users who watch the entire video, and the interaction rate refers to the proportion of user interaction behaviors such as likes, comments, and reposts to the total number of views; content quality indicators include, but are not limited to, the following indicators: content originality and content quality (including the production level, information value, and viewing quality of the video content); and communication effect indicators include, but are not limited to, the following indicators: dissemination range and dissemination speed.
[0071] Of course, the above method for determining the popularity of the target video is only an example and is not intended to limit the present application.
[0072] In another possible implementation, the fine-tuning period may also be smaller than a preset period threshold. Here, the preset period threshold is smaller than or equal to the update period of the existing hot word database.
[0073] Optionally, the degradation period may be a fixed time, such as 2:00 a.m. every day.
[0074] Optionally, the degradation period may be greater than, less than, or equal to the fine-tuning period, which is not specifically limited in this application.
[0075] The speech recognition method provided in the present application obtains the user speech input by the user while watching the target video, and performs speech recognition on the user speech through the trained speech recognition model to obtain the text barrage corresponding to the user speech. The fine-tuning period used for fine-tuning training of the speech recognition model provided in the present application is inversely correlated with the popularity of the target video. Since the more popular the target video is, the higher the possibility of generating new hot words and the greater the number, new hot words are added to the hot word library and fine-tuning training is performed at each fine-tuning moment corresponding to the fine-tuning period, which can enable the speech recognition model to recognize the new hot words more timely and accurately, thereby improving the accuracy of speech recognition; in addition, setting the fine-tuning period to a smaller period that is less than the preset period threshold can also timely add new hot words to the hot word library and perform fine-tuning training on the speech recognition model based on the new hot words, thereby improving the accuracy of speech recognition.
[0076] Furthermore, the present application will remove invalid hot words from the hot word library at each degradation moment corresponding to the preset degradation cycle, which can effectively save the model resources that the speech recognition model relies on, thereby improving the speed and accuracy of speech recognition.
[0077] It is understandable that the popularity of a target video usually does not strictly increase first and then decrease, but may show unpredictable changes in popularity due to factors such as the replay of the target video and the influence of the actors in the target video. Therefore, after a newly added hot word becomes an invalid hot word, it may become a hot word again as the popularity of the target video increases.
[0078] In order to prevent a newly added hot word (such as hard control) from becoming an invalid hot word and then becoming a hot word again due to the popularity of the target video, this embodiment can remove the invalid hot word from the hot word library while adding the invalid hot word to the preset first downgraded vocabulary (that is, downgrade the invalid hot word from the hot word library to the first downgraded vocabulary), so that when the target video or the associated video of the target video is played, and the playback volume meets the preset playback volume threshold, the invalid hot word in the first downgraded vocabulary will be added back to the hot word library.
[0079] Optionally, this embodiment can also downgrade invalid hot words that appear in the first downgraded vocabulary for a period of time reaching a preset first time threshold (for example, 30 days) to a second downgraded vocabulary, so that when the target video or an associated video of the target video is played, and when the playback volume meets the preset playback volume threshold, the invalid hot words in the second downgraded vocabulary can be re-added to the hot word vocabulary in response to the user's control instructions.
[0080] In this embodiment, the newly added hot words do not exist in the hot word library, the first downgraded word library, and the second downgraded word library. That is, only words that do not exist in the hot word library, the first downgraded word library, and the second downgraded word library and whose occurrence frequency is greater than the preset frequency threshold can be used as new hot words in this embodiment.
[0081] Optionally, the first degradation vocabulary is a solid state drive (SSD), and the second degradation vocabulary is a distributed file system (Hadoop Distributed File System, HDFS).
[0082] For example, the data status of invalid hot words in the first degraded vocabulary is "Invalid hot words eliminated in the last 30 days", and the access method is "Supports quick recall". The data status of invalid hot words in the second degraded vocabulary is "Historical archive (elimination time exceeds 30 days)", and the access method is "Supports batch analysis query only".
[0083] This embodiment uses a multi-level downgraded vocabulary to automatically or manually recall invalid hot words from the downgraded vocabulary after they become hot words again, thereby improving the recall speed of invalid hot words.
[0084] In some embodiments of the present application, the process of obtaining the newly added hot words in the previous text is introduced.
[0085] Optionally, the process of acquiring new hot words includes: acquiring interactive data in the interactive area associated with the target video, extracting words whose occurrence frequency is greater than a frequency threshold and are not in the hot word library, the first degraded word library, and the second degraded word library from the interactive data as candidate words, and screening words related to the content of the target video from the candidate words as new hot words.
[0086] Specifically, this embodiment can extract high-frequency new words from the interactive data in the comment area, bullet screen area, and other interactive areas described above as candidate words. The method for extracting high-frequency new words can be to calculate the word frequency and compare it with the words in the hot word library, the first degraded vocabulary, and the second degraded vocabulary. Of course, other methods for extracting high-frequency new words can also be used, which can be determined according to the actual scenario and will not be detailed here.
[0087] Optionally, the candidate words can be directly used as new hot words.
[0088] Preferably, considering that there may be words in the candidate words that are irrelevant to the content of the target video, for example, the target video is an animal science video, and the candidate word is "quantum computing", in order to avoid adding such candidate words to the hot word library and interfering with the fine-tuning training of the speech recognition model, this embodiment can filter out words related to the content of the target video from the candidate words, and use the filtered words as new hot words.
[0089] Optionally, a natural language processing (NLP) algorithm may be used to screen words related to the content of the target video from candidate words. For example, algorithms such as semantic analysis, context association, and domain relevance evaluation may be used to screen and obtain newly added hot words.
[0090] The newly added hot word extraction method provided in this embodiment can ensure that noise words irrelevant to the target video are removed in a timely manner, thereby ensuring that the remaining newly added hot words are meaningful to the current speech recognition scenario and are suitable for incremental training of the speech recognition model.
[0091] The following embodiment introduces the degradation condition provided above.
[0092] Optionally, taking any word in the hot word library (hereinafter referred to as the target word for ease of introduction) as an example, if the target word meets one or more of the following downgrade conditions, the target word is determined to be an invalid hot word.
[0093] The first demotion condition: the number of times the target word appears in the interactive area during the current demotion cycle is not ranked in the top n%, where n>0. The current demotion cycle refers to the time period between the current moment (i.e., the current demotion moment, i.e., the moment when the target word is judged as an invalid hot word based on the demotion condition) and the previous demotion moment.
[0094] In this embodiment, the number of times each word in the hot word list appears in the interactive area can be counted and ranked during the current demotion period. If the target word is not ranked in the top n%, it is considered that the target word is no longer popular and can be determined as an invalid hot word.
[0095] The second degradation condition: the time between the current degradation moment and the last appearance of the target word in the interactive area is greater than the preset second time threshold.
[0096] In this embodiment, the timestamp of the last appearance of each word in the hot word library in the interactive area can be counted. For the convenience of the following description, this timestamp is defined as the last hit time. If the time between the current demotion moment and the last hit time of the target word is greater than a preset second time threshold, it means that the target word has not been used by the user for a period of time. In this case, the target word is considered to be no longer popular and can be determined as an invalid hot word.
[0097] The third degradation condition: the duration between the current degradation moment and the last completion of the target video playback is greater than the preset third duration threshold.
[0098] As mentioned above, the words in the hotword database are extracted from the interactive data in the interactive zone associated with the target video. Therefore, the target words in the hotword database are relevant to the target video. If the time between the current degradation moment and the last playback of the target video is greater than the preset third time threshold, it indicates that the target word may not be used recently and can be determined as an invalid hotword.
[0099] The fourth degradation condition: the weight of the target word is less than a preset weight threshold, wherein the weight is determined according to the last hit time of the target word and the number of times the target word appears in the interactive area in the current degradation period.
[0100] As described above, this embodiment can count the last hit timestamp of each word in the hot word database and the number of times it appears in the interactive area in the current degradation period.
[0101] For example, the storage structure of each word in the hot word database can be as follows:
[0102] "{
[0103] "Hot word": "hard control",
[0104] "Last hit time": "2025-02-15 20:34:21",
[0105] "Time Decay Factor": 0.95
[0106] }".
[0107] In this embodiment, the weight of the target word can be determined based on the last hit time of the target word and the number of times the target word appears in the interactive area in the current demotion cycle, and then the weight of the target word is compared with a preset weight threshold. If the circle of the target word is smaller than the preset weight threshold, it means that the target word is no longer popular, and the target word can be determined as an invalid hot word.
[0108] In one possible implementation, the process of "determining the weight of the target word based on the last hit time of the target word and the number of times the target word appears in the interactive area within the current degradation cycle" may include: calculating the hit rate weight based on the number of times the target word appears in the interactive area within the current degradation cycle, calculating the time decay weight based on the last hit time of the target word, and calculating the comprehensive weight based on the hit rate weight and the time decay weight as the weight of the target word.
[0109] Optionally, the process of "calculating the hit rate weight according to the number of times the target word appears in the interactive area in the current degradation period" can use the following formula (1).
[0110] Formula (1);
[0111] in, Indicates the hit rate weight; Indicates the number of times the target word appears in the interaction area during the current degradation cycle.
[0112] Optionally, the process of “calculating the time decay weight according to the last hit time of the target word” can use the following formula (2).
[0113] Formula (2);
[0114] in, represents the time decay weight; Indicates the last hit time of the target word, Indicates the current moment of degradation; Represents the time decay coefficient, which is a preset value based on the actual scenario. For example, in some scenarios, the initial time decay coefficient can be set to 0.95; 24 represents 24 hours.
[0115] Optionally, the process of "calculating the comprehensive weight according to the hit rate weight and the time decay weight" may include: multiplying the hit rate weight and the time decay weight, and using the product value as the comprehensive weight.
[0116] It should be noted that the above process of calculating the weight of the target word is only an example and is not intended to limit the present application.
[0117] It should also be noted that the above four demotion conditions are only examples. In addition, the demotion conditions can also be other, which are not specifically limited in this application. In addition, the above four demotion conditions can be freely combined into new demotion conditions to meet the use in different scenarios. For example, the first demotion condition and the fourth demotion condition are combined together. Only when the target word meets both the first and fourth demotion conditions, the target word will be downgraded to the first demotion vocabulary.
[0118] It should also be noted that the various thresholds mentioned in this application, such as n, the second duration threshold, the third duration threshold, and the weight threshold, can be set according to the actual scenario, and this application does not make specific limitations. For example, n can be 30, the weight threshold can be 0.4, and so on.
[0119] The above embodiment can determine that words that are no longer popular in the hot word library are invalid hot words by setting the demotion conditions, and then demotion the invalid hot words, which can effectively save model resources. However, new hot words are periodically added to the hot word library. Since the new hot words have just been added to the hot word library, they may be degraded due to meeting the above demotion conditions. For example, The value is relatively small, which results in the newly added hot words meeting the downgrade conditions and being downgraded.
[0120] Optionally, in order to prevent the newly added hot words from being downgraded to the first downgraded vocabulary, this embodiment can set a protection period for the newly added hot words when adding them to the hot word vocabulary, so that the newly added hot words are not removed during the protection period.
[0121] That is, this embodiment can set a protection period for each newly added hot word. During the protection period, the newly added hot word will not be downgraded even if it does not meet the downgrade conditions. However, after the protection period, the newly added hot word will be downgraded to the first downgraded vocabulary if it does not meet the downgrade conditions.
[0122] Optionally, when the demotion conditions include the above-mentioned fourth demotion condition, this embodiment can also configure an initial weight for the newly added hot word that is greater than or equal to the weight threshold. The initial weight will become invalid when the time for the newly added hot word to be added to the hot word database reaches the preset fourth time threshold, or when the initial weight is less than the comprehensive weight.
[0123] That is, this embodiment can set an initial weight for each newly added hot word within the preset fourth duration threshold, and at the same time, the comprehensive weight of each newly added hot word can be calculated according to the weight calculation process provided above at each demotion moment. If the initial weight is always greater than or equal to the comprehensive weight within the fourth duration threshold, the initial weight is always valid, and the newly added hot word is judged to be downgraded according to the initial weight until the duration of the newly added hot word in the hot word database exceeds the fourth duration threshold, at which point the initial weight becomes invalid, and thereafter the newly added hot word is judged to be downgraded according to the calculated comprehensive weight. If the initial weight is less than the calculated comprehensive weight, the initial weight is invalidated regardless of whether it is within the fourth duration threshold, and the newly added hot word is judged to be downgraded according to the comprehensive weight.
[0124] For example, each newly added hot word may go through the following life cycle: joining stage T0 (i.e. the stage when it is first added to the hot word library), popular stage (i.e. the stage when the popularity continues to increase), decline stage (i.e. the stage when the popularity continues to decline but still in the hot word library), elimination warning stage (i.e. the stage when it is downgraded to the first downgraded word library or the second downgraded word library) and recall inspection stage (i.e. the stage when it is recalled from the hot word library due to the replay of the target video or related videos).
[0125] In this embodiment, a protection period (e.g., 24 hours) can be set for a newly added hotword a during the addition phase (T0), or an initial weight of 0.6 (the weight threshold is 0.4) can be set within the fourth time threshold of 24 hours. Due to the protection period or initial weight, the newly added hotword a can be exempted from demotion. Assuming that the newly added hotword a enters the hot broadcast phase at T0+12 hours, due to the large number of appearances in the interactive area, the comprehensive weight of the newly added hotword a rises to 0.85, 0.85 > 0.6, and the initial weight becomes invalid. Thereafter, the newly added hotword a is degraded according to the calculated comprehensive weight. Assuming that the comprehensive weight of the newly added hotword a decays to 0.45 during the decline phase (0.45 > 0.4), the newly added hotword a is not demoted. Assuming that the comprehensive weight of the newly added hotword a decays to 0.18 during the elimination warning phase (0.18 < 0.4), the newly added hotword a is demoted to the first demotion vocabulary.
[0126] In summary, this embodiment, by setting demotion conditions, can demotion words in the hot word library that are no longer popular, i.e., invalid hot words, to the first demotion word library or the second demotion word library, thus saving model resources. To prevent newly added hot words from being degraded due to hitting demotion conditions when they are just added to the hot word library, this application sets a protection period or initial weight to ensure that newly added hot words are not mistakenly identified as invalid hot words, thereby improving the accuracy of invalid hot word identification and demotion.
[0127] Considering that the method of generating barrage by voice recognition consumes more network resources than the method of directly inputting text into barrage and requires the availability of voice recognition services, in order to better combine the method of generating barrage by voice recognition with network resources and voice recognition services, the following embodiment is provided.
[0128] Optionally, the process of step S102 "performing speech recognition on user speech through a trained speech recognition model" may include: obtaining the target availability state corresponding to the target video, where the target availability state is one of the first state, the second state and the third state, the network transmission capacity of the first state is greater than the preset first capacity threshold, and the speech recognition service quality is greater than the preset first quality threshold, the network transmission capacity of the third state is less than the preset second capacity threshold, and the speech recognition service quality is less than the preset second quality threshold, the first quality threshold is greater than the second quality threshold, and the first capacity threshold is greater than the second capacity threshold, and the second state refers to other states except the first state and the third state; when the target availability state is the first state, speech recognition is performed on the user speech through the speech recognition model without pausing the playback of the target video, and when the target availability state is the second state, speech recognition is performed on the user speech through the speech recognition model while pausing the playback of the target video.
[0129] Optionally, when the target availability state is in the third state, the user is prompted that voice recognition is unavailable when the target video is paused, and a text input component is called so that the user can use the text input component to input text comments.
[0130] Optionally, the text input component may be a keyboard input component.
[0131] Optionally, the network transmission capacity can be determined by network bandwidth and / or network speed, and the speech recognition service quality can be determined by speech recognition service latency. Taking network bandwidth as an example, optionally, the first state refers to a network state in which the speech recognition service latency is less than 100ms and the network bandwidth is greater than 10Mbps; the third state refers to a network state in which the speech recognition service latency is greater than or equal to 500ms and the bandwidth is less than or equal to 2Mbps; and the second state refers to any state other than the first and third states.
[0132] Of course, the first state, second state and third state mentioned above may also be other states, which are not specifically limited in this application.
[0133] That is, in this embodiment, when the target video is in the first state, part of the network resources can be used to continuously play the target video, and the remaining network resources can be used to call the speech recognition model to perform speech recognition on the user's voice to generate text barrage.
[0134] When the target video is in the second state, all network resources can be used to call the speech recognition model to perform speech recognition on the user's voice. However, the target video needs to be paused during this period to avoid the user missing important video plots due to slow network response during the entire speech recognition process. When the speech recognition obtains the text barrage, the target video is resumed and the text barrage is sent out at the same time.
[0135] When the target video is in the third state and the target video is forcibly paused, on the one hand, the user is prompted that voice recognition is unavailable, and on the other hand, a keyboard input interface can pop up to allow the user to enter text barrage through the keyboard input interface. When the text barrage is entered, the target video is resumed and the text barrage is sent out at the same time.
[0136] It can be seen that this embodiment can adjust the product interaction form according to the network conditions and the availability of the voice recognition service, ensuring the user's experience of watching the target video.
[0137] The above describes a speech recognition method provided in an embodiment of the present application. The following describes a device for executing the above speech recognition method.
[0138] See also Figure 2 , Figure 2 This is a structural diagram of a speech recognition device provided in an embodiment of the present application. Figure 2 As shown, the speech recognition device may include:
[0139] The voice acquisition module 201 is used to acquire the user voice input by the user while watching the target video;
[0140] The speech recognition module 202 is used to perform speech recognition on the user's speech using a trained speech recognition model to obtain a text comment corresponding to the user's speech;
[0141] Among them, the speech recognition model is periodically fine-tuned and trained according to a preset fine-tuning cycle. The hot word library used for fine-tuning training adds new hot words at each fine-tuning moment corresponding to the fine-tuning cycle, and eliminates invalid hot words at each demotion moment corresponding to the preset demotion cycle. The newly added hot words do not exist in the hot word library and their frequency of appearance in the interactive area associated with the target video is greater than a preset frequency threshold. Invalid hot words refer to words in the hot word library that meet the preset demotion conditions. The fine-tuning cycle is inversely correlated with the heat of the target video and / or the fine-tuning cycle is less than the preset cycle threshold.
[0142] In a possible implementation, the above-mentioned speech recognition module may also be used for:
[0143] While removing invalid hot words from the hot word library, the invalid hot words are added to a preset first degraded vocabulary library, so that when the target video or an associated video of the target video is played, if the playback volume meets the preset playback volume threshold, the invalid hot words in the first degraded vocabulary library are added back to the hot word library, wherein the associated video refers to a video whose content relevance to the target video is greater than the preset relevance threshold.
[0144] Invalid hot words that appear in the first demotion vocabulary for a time period reaching a preset first time period threshold are downgraded to the second demotion vocabulary, so that when the target video or the associated video of the target video is played, when the playback volume meets the preset playback volume threshold, the invalid hot words in the second demotion vocabulary are re-added to the hot word vocabulary in response to the user's control instruction, wherein the newly added hot words do not exist in either the first demotion vocabulary or the second demotion vocabulary.
[0145] In a possible implementation, when acquiring a new hot word, the speech recognition module may be used to:
[0146] Get interaction data within the interaction area;
[0147] Extracting words whose occurrence frequency is greater than a frequency threshold and are not in the hot word library, the first degraded word library, and the second degraded word library from the interaction data as candidate words;
[0148] Filter words related to the content of the target video from the candidate words and use them as new hot words.
[0149] In a possible implementation, the degradation condition in the speech recognition module may be one or more of the following conditions:
[0150] The target word does not appear in the top n% of the interactive area in the current demotion period, where n>0. The target word refers to any word in the hot word database, and the current demotion period refers to the time period between the current demotion moment and the previous demotion moment.
[0151] The time between the current degradation moment and the last appearance of the target word in the interactive area is greater than the preset second time threshold;
[0152] The duration between the current degradation moment and the last playback of the target video is greater than the preset third duration threshold;
[0153] The weight of the target word is less than a preset weight threshold, wherein the weight of the target word is determined according to the last hit time of the target word and the number of times the target word appears in the interactive area in the current degradation period.
[0154] In a possible implementation, the process of determining the weight of the target word in the speech recognition module according to the last hit time of the target word and the number of times the target word appears in the interactive area in the current degradation period may include:
[0155] The hit rate weight is calculated based on the number of times the target word appears in the interactive area during the current degradation cycle;
[0156] Calculate the time decay weight based on the last hit time of the target word;
[0157] According to the hit rate weight and time decay weight, the comprehensive weight is calculated as the weight of the target word.
[0158] In a possible implementation, the above-mentioned speech recognition module may also be used for:
[0159] When a new hot word is added to the hot word database, a protection period is set for the new hot word so that the new hot word is not eliminated during the protection period, and / or an initial weight greater than or equal to the weight threshold is configured for the new hot word. The initial weight becomes invalid when the time since the new hot word was added to the hot word database reaches the preset fourth time threshold, or the initial weight becomes invalid when it is less than the comprehensive weight.
[0160] In one possible implementation, when the speech recognition module performs speech recognition on the user's speech using the trained speech recognition model, it can be specifically used to:
[0161] Obtaining a target availability state corresponding to a target video, where the target availability state is one of a first state, a second state, and a third state; the network transmission capacity of the first state is greater than a preset first capacity threshold, and the voice recognition service quality is greater than a preset first quality threshold; the network transmission capacity of the third state is less than a preset second capacity threshold, and the voice recognition service quality is less than a preset second quality threshold, the first quality threshold is greater than the second quality threshold, and the first capacity threshold is greater than the second capacity threshold; and the second state refers to any state other than the first state and the third state;
[0162] When the target available state is the first state, performing speech recognition on the user's speech through the speech recognition model without pausing the playing of the target video;
[0163] When the target available state is the second state, the user's voice is recognized through the voice recognition model while the target video is paused.
[0164] In one possible implementation, when the speech recognition module performs speech recognition on the user's speech using the trained speech recognition model, it may further be used to:
[0165] When the target available state is the third state, the user is prompted that voice recognition is unavailable when the target video is paused, and the text input component is called so that the user can use the text input component to input text barrage.
[0166] An embodiment of the present application also provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein: the memory is used to store a computer program, and the processor is used to execute the computer program, so that the electronic device can implement the various steps of the speech recognition method as described above.
[0167] For example, reference Figure 3, which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 3 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0168] like Figure 3 As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 302 or programs loaded from a storage device 308 into a random access memory (RAM) 303. When the electronic device is powered on, the RAM 303 also stores various programs and data required for the operation of the electronic device. The processing device 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0169] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a memory card, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0170] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the speech recognition methods provided in the embodiments of the present application.
[0171] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any speech recognition method provided in the embodiment of the present application.
[0172] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0173] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0174] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0175] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A speech recognition method, characterized in that: include: Obtain the user voice input during the user watching the target video; Perform speech recognition on the user's voice using a trained speech recognition model to obtain a text comment corresponding to the user's voice; In which, the speech recognition model is periodically fine-tuned and trained according to a preset fine-tuning cycle. The hot word library used in the fine-tuning training adds new hot words at each fine-tuning moment corresponding to the fine-tuning cycle, and eliminates invalid hot words at each demotion moment corresponding to a preset demotion cycle. The new hot words do not exist in the hot word library and their frequency of appearance in the interactive area associated with the target video is greater than a preset frequency threshold. The invalid hot words refer to words in the hot word library that meet the preset demotion conditions. The fine-tuning cycle is inversely correlated with the heat of the target video and / or the fine-tuning cycle is less than a preset cycle threshold.
2. The speech recognition method according to claim 1, wherein: Also includes: While removing the invalid hot words from the hot word library, the invalid hot words are added to a preset first degraded word library, so that when the target video or an associated video of the target video is played, if the playback volume meets a preset playback volume threshold, the invalid hot words in the first degraded word library are added back to the hot word library, wherein the associated video refers to a video whose content relevance to the target video is greater than a preset relevance threshold; Invalid hot words that appear in the first degraded vocabulary for a duration reaching a preset first duration threshold are downgraded to the second degraded vocabulary, so that when the target video or the associated video of the target video is played, if the playback volume meets the preset playback volume threshold, in response to the user's control instruction, the invalid hot words in the second degraded vocabulary are re-added to the hot word vocabulary, wherein the newly added hot words do not exist in both the first degraded vocabulary and the second degraded vocabulary.
3. The speech recognition method according to claim 2, wherein: The process of acquiring the newly added hot words includes: Acquiring interaction data within the interaction area; Extracting, from the interaction data, words whose occurrence frequency is greater than the frequency threshold and are not in the hot word library, the first degraded word library, and the second degraded word library as candidate words; Words related to the content of the target video are selected from the candidate words as the newly added hot words.
4. The speech recognition method according to any one of claims 1 to 3, characterized in that: The downgrade conditions are one or more of the following conditions: The number of times the target word appears in the interactive area in the current demotion period is not ranked in the top n%, where n>0. The target word refers to any word in the hot word library, and the current demotion period refers to the time period between the current demotion moment and the previous demotion moment. The time between the current degradation moment and the last appearance of the target word in the interactive area is greater than a preset second time threshold; The duration between the current degradation moment and the last completion of the target video playback is greater than a preset third duration threshold; The weight of the target word is less than a preset weight threshold, wherein the weight of the target word is determined according to the last hit time of the target word and the number of times the target word appears in the interactive area within the current degradation period.
5. The speech recognition method according to claim 4, characterized in that The weight of the target word is determined according to the last hit time of the target word and the number of times the target word appears in the interactive area within the current degradation period, including: Calculating a hit rate weight according to the number of times the target word appears in the interactive area during the current degradation period; Calculating a time decay weight according to the last hit time of the target word; According to the hit rate weight and the time decay weight, a comprehensive weight is calculated as the weight of the target word.
6. The speech recognition method according to claim 5, characterized in that Also includes: When the newly added hot word is added to the hot word database, a protection period is set for the newly added hot word so that the newly added hot word is not eliminated during the protection period, and / or an initial weight greater than or equal to the weight threshold is configured for the newly added hot word, and the initial weight becomes invalid when the time when the newly added hot word is added to the hot word database reaches a preset fourth time threshold, or the initial weight becomes invalid when it is less than the comprehensive weight.
7. The speech recognition method according to claim 1, wherein: The performing speech recognition on the user's speech using the trained speech recognition model includes: Obtaining a target availability state corresponding to the target video, wherein the target availability state is one of a first state, a second state, and a third state, wherein the network transmission capability of the first state is greater than a preset first capability threshold, and the speech recognition service quality is greater than a preset first quality threshold, wherein the network transmission capability of the third state is less than a preset second capability threshold, and the speech recognition service quality is less than a preset second quality threshold, wherein the first quality threshold is greater than the second quality threshold, and the first capability threshold is greater than the second capability threshold, and the second state refers to a state other than the first state and the third state; When the target available state is the first state, performing speech recognition on the user speech by using the speech recognition model without pausing the playing of the target video; When the target available state is the second state, the user voice is recognized by the voice recognition model while the target video is paused.
8. The speech recognition method according to claim 7, characterized in that: Also includes: When the target available state is the third state, the user is prompted that voice recognition is unavailable when the target video is paused, and a text input component is called so that the user can use the text input component to input text comments.
9. A speech recognition device, characterized in that: include: A voice acquisition module is used to acquire the user's voice input during watching the target video; A speech recognition module is used to perform speech recognition on the user's speech using a trained speech recognition model to obtain a text comment corresponding to the user's speech; In which, the speech recognition model is periodically fine-tuned and trained according to a preset fine-tuning cycle. The hot word library used in the fine-tuning training adds new hot words at each fine-tuning moment corresponding to the fine-tuning cycle, and eliminates invalid hot words at each demotion moment corresponding to a preset demotion cycle. The new hot words do not exist in the hot word library and their frequency of appearance in the interactive area associated with the target video is greater than a preset frequency threshold. The invalid hot words refer to words in the hot word library that meet the preset demotion conditions. The fine-tuning cycle is inversely correlated with the heat of the target video and / or the fine-tuning cycle is less than a preset cycle threshold.
10. An electronic device, characterized in that: comprising at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is configured to execute the computer program so that the electronic device can implement the speech recognition method according to any one of claims 1 to 8.