Subtitle recognition method and apparatus, device, and medium
By acquiring subtitle replacement data and generating subtitle recognition prompts, the problem of excessive modification of homophones or words with similar pronunciations in subtitle recognition is solved, which improves the accuracy and recall of subtitle recognition and makes the generated subtitles more in line with user needs.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-08-08
- Publication Date
- 2026-04-23
AI Technical Summary
In the current subtitle recognition process, homophones or words with similar pronunciations are easily over-modified, resulting in low accuracy of subtitle recognition.
By acquiring subtitle replacement data, including multiple target replacement data, each containing words before and after correction, subtitle recognition prompts are generated and input into the subtitle recognition model. The user's modification habits and needs are recorded to avoid excessive modification.
It improves the accuracy and recall of subtitle recognition, making the generated subtitles more in line with users' personalized needs.
Smart Images

Figure CN2025113464_23042026_PF_FP_ABST
Abstract
Description
A method, apparatus, device and medium for subtitle recognition
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411441673.0, filed on October 15, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure relates to a method, apparatus, device, and medium for subtitle recognition. Background Technology
[0004] With the development of speech recognition technology, video subtitles can be automatically generated. However, users sometimes need to manually correct the same subtitle errors repeatedly. Summary of the Invention
[0005] To address the aforementioned technical problems, this disclosure provides a method, apparatus, device, and medium for subtitle recognition.
[0006] This disclosure provides a subtitle recognition method, the method comprising:
[0007] Acquire target media data and acquire subtitle replacement data, wherein the subtitle replacement data includes multiple target replacement data, and each target replacement data includes a word before correction and a word after correction;
[0008] Based on the subtitle replacement data, subtitle recognition prompts are generated that include the multiple target replacement data;
[0009] The target media data and the subtitle recognition prompts are input into the subtitle recognition model, and the target subtitles corresponding to the target media data are output.
[0010] This disclosure also provides a subtitle recognition device, the device comprising:
[0011] The acquisition module is used to acquire target media data and subtitle replacement data, wherein the subtitle replacement data includes multiple target replacement data, and each target replacement data includes a word before correction and a word after correction;
[0012] The generation module is used to generate subtitle recognition prompts that include the multiple target replacement data based on the subtitle replacement data;
[0013] The recognition module is used to input the target media data and the subtitle recognition prompts into the subtitle recognition model and output the target subtitles corresponding to the target media data.
[0014] This disclosure also provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the subtitle recognition method provided in this disclosure.
[0015] This disclosure also provides a computer-readable storage medium storing a computer program for performing the subtitle recognition method provided in this disclosure. Attached Figure Description
[0016] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0017] Figure 1 is a flowchart illustrating a subtitle recognition method provided in an embodiment of this disclosure;
[0018] Figure 2 is a flowchart illustrating another subtitle recognition method provided in an embodiment of this disclosure;
[0019] Figure 3 is a flowchart illustrating another subtitle recognition method provided in this embodiment of the present disclosure;
[0020] Figure 4 is a schematic diagram of a subtitle recognition settings page provided in an embodiment of this disclosure;
[0021] Figure 5 is a schematic diagram of another subtitle recognition settings page provided in an embodiment of this disclosure;
[0022] Figure 6 is a schematic diagram of a subtitle recognition method provided in an embodiment of this disclosure;
[0023] Figure 7 is a schematic diagram of the structure of a subtitle recognition device provided in an embodiment of this disclosure;
[0024] Figure 8 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0031] Smart captions use technologies such as automatic speech recognition to automatically generate subtitles for videos. In personalized use cases of smart captions, users sometimes need to manually and repeatedly correct the same subtitle errors.
[0032] To address the aforementioned issues, a vocabulary of multiple optimized words is set up and input into the model along with the media for subtitle recognition. However, this approach often results in over-modification of words in the media that are homophones or similar in pronunciation to the optimized words, leading to lower accuracy in subtitle recognition.
[0033] For example, subtitle optimization can be achieved through word optimization schemes. Specifically, the correct words corresponding to incorrect subtitles are identified as optimized words, and a vocabulary list including multiple optimized words is set up. This vocabulary list is then input into the model along with the media for subtitle recognition. However, while this method can effectively optimize subtitles that have been repeatedly modified by the user, most words that are homophones or similar in pronunciation to the optimized words are also modified to become optimized words, resulting in over-modification and low subtitle recognition accuracy.
[0034] To address the aforementioned problems, this disclosure provides a subtitle recognition method, which will be described below with reference to specific embodiments.
[0035] Figure 1 is a flowchart illustrating a subtitle recognition method according to an embodiment of this disclosure. This method can be executed by a subtitle recognition device, which can be implemented using software and / or hardware and is generally integrated into an electronic device. As shown in Figure 1, the method includes:
[0036] Step 101: Obtain target media data and subtitle replacement data. The subtitle replacement data includes multiple target replacement data, and each target replacement data includes the word before correction and the word after correction.
[0037] The target media data can be the data to be identified by subtitle recognition, and this embodiment does not limit the type of the target media data. For example, the target media data may include video data or audio data. The subtitle replacement data may be word-level replacement modification data extracted from the user's historical records of manually modified subtitles. Word-level replacement modification may be a modification that replaces one word with another. The subtitle replacement data may be obtained by processing the historical subtitle correction data. This embodiment does not limit the timing of the generation of the subtitle replacement data. For example, the subtitle replacement data may be predetermined and stored in local storage space and / or database, or the subtitle replacement data may be generated in real time based on the historical subtitle correction data.
[0038] In this embodiment, the subtitle replacement data may include multiple target replacement data, each representing a single word replacement. Each target replacement data includes a word before correction and a word after correction. The word before correction refers to the word automatically generated by the subtitle system that the user did not manually modify, while the word after correction refers to the correct word that the user manually corrected after discovering an error in the word before correction. In some embodiments of this disclosure, the word before correction and the word after correction are words with the same initial consonant and final vowel, and both the word before correction and the word after correction are nouns. This restricts the word before correction and the word after correction in both pronunciation and part-of-speech dimensions, further avoiding excessive modification caused by excessive intervention in the subtitle generation process.
[0039] In this embodiment, the timing of acquiring the target media data and subtitle replacement data is not limited. For example, the data can be acquired in response to a corresponding operation. Specifically, a user can perform a subtitle recognition operation on the target media data, and in response to this subtitle recognition operation, the subtitle recognition device can acquire the target media data and the subtitle replacement data. Alternatively, the data can be acquired in advance. Specifically, before the user performs the subtitle recognition operation, the subtitle recognition device can acquire the target media data and the subtitle replacement data in advance, and continue with subsequent steps such as generating subtitle recognition prompts after the user performs the subtitle recognition operation. The subtitle recognition operation can be a gesture control operation such as clicking or double-clicking the subtitle recognition control in a video processing application, used to trigger subtitle recognition of an audio or video data. The subtitle recognition control can be a functional control that instructs the target media data to undergo audio recognition to obtain subtitles.
[0040] Figure 2 is a flowchart illustrating another subtitle recognition method provided in an embodiment of this disclosure. As shown in Figure 2, in some embodiments of this disclosure, obtaining subtitle replacement data includes:
[0041] Step 201: Obtain historical subtitle correction data. The historical subtitle correction data includes multiple correction data, and each correction data includes historically generated subtitles and corresponding historically corrected subtitles.
[0042] The subtitle correction history data records user-manually modified subtitles, including at least one of deleting, adding, or replacing characters. Each user has a unique subtitle correction history. In this embodiment, the subtitle correction history data includes multiple correction data sets, each representing a modification to a single subtitle line. Each correction data set includes historically generated subtitles and historically corrected subtitles. Historically generated subtitles can automatically identify subtitles that the user has not manually modified. Historically corrected subtitles can allow users to manually correct errors found in historically generated subtitles.
[0043] In this embodiment, the subtitle recognition device can acquire historically generated subtitles obtained by the user's past subtitle recognition, and acquire historically corrected subtitles after the user has modified the historically generated subtitles. This pair of historically generated subtitles and historically corrected subtitles is identified as correction data. Multiple correction data are determined based on the user's modifications to multiple historically generated subtitles, and the corresponding historical subtitle correction data for the user is determined based on the user's multiple correction data.
[0044] Step 202: Process the historical data of subtitle correction to obtain subtitle replacement data.
[0045] In this embodiment, the subtitle recognition device can perform data processing such as word extraction and word filtering on the subtitle correction data to obtain subtitle replacement data.
[0046] Figure 3 is a flowchart illustrating another subtitle recognition method provided in this embodiment of the present disclosure, which involves processing historical subtitle correction data to obtain subtitle replacement data, including:
[0047] Step 301: Align multiple corrected data using the minimum edit distance rule, and extract multiple first replacement data based on the alignment results.
[0048] The minimum edit distance can be defined as the minimum number of operations required to convert historically generated subtitles into historically corrected subtitles. The alignment result can be the result of aligning historically generated and historically corrected subtitles at the character level. The first replacement data can be used to record character-level replacements performed by the user on historically generated subtitles. This first replacement data can be determined based on character-level replacements in historically generated subtitles; optionally, deletions or additions of characters at the character level in historically generated subtitles do not generate corresponding first replacement data.
[0049] In this embodiment, for each correction data in the historical subtitle correction data, the subtitle recognition device can use the minimum edit distance as the alignment rule to align the historical generated subtitles and historical corrected subtitles in the correction data to obtain an alignment result. In this alignment result, if a character in the historical generated subtitle corresponds to the same character in the historical corrected subtitle, it means that the user did not modify the character. If a character in the historical generated subtitle has an empty corresponding character in the historical corrected subtitle, it means that the user deleted the character in the historical generated subtitle. If a character in the historical corrected subtitle has an empty corresponding character in the historical generated subtitle, it means that the user added the character in the historical generated subtitle. If a character in the historical generated subtitle does not correspond to the same character in the historical corrected subtitle, it means that the user replaced the character in the historical generated subtitle. The subtitle recognition device can use the replaced character in the historical generated subtitle as the original character and the character that replaced other characters in the historical corrected subtitle as the corrected character, thereby generating first replacement data including the original character and its corresponding corrected character.
[0050] Step 302: Map multiple first replacement data to words to obtain multiple second replacement data.
[0051] The second replacement data can be used to record word-level replacements made by users to historically generated subtitles.
[0052] In this embodiment, the subtitle recognition device can map words in multiple first replacement data in historical generated subtitles and historical corrected subtitles to obtain multiple second replacement data.
[0053] In some embodiments of this disclosure, multiple first replacement data are mapped from characters to words to obtain multiple second replacement data, including: for each first replacement data, obtaining the word segmentation result of the corresponding correction data, and determining the word before correction and word after correction mapped to the word before correction and word after correction in the first replacement data based on the word segmentation result, thereby obtaining the second replacement data.
[0054] The word segmentation result can be multiple words obtained by segmenting the subtitles.
[0055] In this embodiment, the subtitle recognition device can call a word segmentation tool to segment the historically generated subtitles in the correction data. Among the multiple word segments of the historically generated subtitles, the word segment containing the character before correction is determined as the "previous word". Furthermore, the device calls a word segmentation tool to segment the historically corrected subtitles in the correction data, and among the multiple word segments of the historically corrected subtitles, the word segment containing the character after correction is determined as the "post-correction word". A correspondence between the "previous word" and the "post-correction word" is established based on the correspondence between the characters before and after correction, and this correspondent "previous word" and "post-correction word" are determined as the second replacement data.
[0056] For example, if the historical generated subtitle is ABAAC and the historical corrected subtitle is ADAAC, after alignment processing, the minimum edit distance is determined to be 1, and the editing of the historical generated subtitle involves correcting B to D. If the historical generated subtitle is ABAAC and the historical corrected subtitle is ADAABC, after alignment processing, the minimum edit distance is determined to be 2, and the editing of the historical generated subtitle includes correcting B to D and directly inserting B between AC. In both examples above, correcting B in the historical generated subtitle to D in the historical corrected subtitle is included. If the word segment containing B is BA and the word segment containing D is DA, then BA in the second replacement data corresponds to DA.
[0057] Step 303: Extract multiple replacement data that are homophones and belong to nouns from multiple second replacement data and determine them as multiple target replacement data, and determine the multiple target replacement data as subtitle replacement data.
[0058] In this context, homophony can be understood as at least one element being the same in the pinyin of two words. Pinyin can include initials, finals, and tones. Initials can include elements such as retroflex consonants, alveolar consonants, velar consonants, and palatal consonants. Finals can include simple finals, compound finals, and nasal finals including front and back nasal finals. Tones can include four elements: level tone, rising tone, falling tone, and turning tone. In other words, homophony can mean that the pinyin is completely identical or partially identical. Completely identical pinyin means that all elements in the pinyin of two words are the same. Partially identical pinyin can be understood as one or more elements being the same in the pinyin of two words. For example, partially identical pinyin can include pinyin elements other than tone being the same, pinyin elements other than front and back nasal consonants being the same, pinyin elements other than retroflex and alveolar consonants being the same, and compatible accents, etc. A replacement data item belonging to a noun can be understood as both the original and modified words in the replacement data being nouns, or the modified word in the replacement data being a noun while the part of speech of the original word is not restricted.
[0059] In this embodiment, for each second replacement data, the subtitle recognition device can call a Chinese character-to-pinyin conversion tool to convert the original word from Chinese characters to pinyin, obtaining the original pinyin; and to convert the original word from Chinese characters to pinyin, obtaining the original pinyin. Based on the above homophony judgment rules, it is determined whether the original pinyin and the original pinyin are homophones. If so, a part-of-speech analysis tool is called to determine the part of speech of the original word and the original word. Further, the second replacement data in which both the original word and the original word are nouns are determined as target replacement data; or, the second replacement data in which the original word is a noun is determined as target replacement data. When the original word is incorrectly identified, its part of speech may not be the same as the original word. By not restricting the part of speech of the original word, the incorrect non-noun can be replaced with the correct noun in the subsequent subtitle recognition process, improving the accuracy of subtitle recognition.
[0060] Furthermore, multiple target replacement data are aggregated into subtitle replacement data.
[0061] In the above scheme, by aligning and mapping words to words, the scheme achieves efficient and accurate determination of words before correction in historically generated subtitles and words after correction in historically corrected subtitles. Furthermore, it imposes restrictions in the pinyin and part-of-speech dimensions to further avoid excessive modification caused by excessive intervention in the subtitle generation process of the model.
[0062] In some embodiments of this disclosure, obtaining subtitle replacement data includes: obtaining pre-stored subtitle replacement data from local storage space or a database, wherein the subtitle replacement data is determined in response to a subtitle correction operation.
[0063] The local storage space can be the space on the user device used to store data, and the user device can be the device currently being used by the user. The database can be the space on the server corresponding to the user device used to store data. The subtitle correction operation can be an operation to modify the subtitles. Understandably, this subtitle correction operation can include at least one of the following: deletion, addition, or replacement of characters in historically generated subtitles by the user.
[0064] In this embodiment, before performing subtitle recognition on the target media data, the subtitle recognition device can pre-determine subtitle replacement data and store the subtitle replacement data in local storage space and / or a database. Specifically, the subtitle recognition device can acquire subtitle correction history data; process the subtitle correction history data to obtain subtitle replacement data. This data processing process may include: aligning multiple correction data using the minimum edit distance rule, extracting multiple first replacement data based on the alignment result; mapping the multiple first replacement data from characters to words to obtain multiple second replacement data; extracting multiple replacement data that are homophones and belong to nouns from the multiple second replacement data to determine multiple target replacement data, and determining the multiple target replacement data as subtitle replacement data. Further, the subtitle replacement data is stored in local storage space and / or a database. Subsequently, the subtitle recognition device can determine the subtitle replacement data corresponding to the user in local storage space or a database, and obtain the subtitle replacement data corresponding to the user.
[0065] In the above scheme, the subtitle replacement data is predetermined and stored in the database. Subsequent retrieval of the subtitle replacement data from this database improves the speed of subtitle recognition.
[0066] Step 102: Generate subtitle recognition prompts that include multiple target replacement data based on the subtitle replacement data.
[0067] The prompt word is an input text fragment that guides the model to generate specific output content. The prompt word guides the subtitle recognition model to perform subtitle recognition on target media data. This prompt word includes multiple target replacement data, which can include a user-manual subtitle replacement process. During the subtitle recognition process of audio or video data, the prompt word replaces words in the audio or video data that have the same or similar pronunciation to the original words in the target replacement data with the corrected words in the target replacement data.
[0068] In this embodiment, the subtitle recognition device can generate subtitle recognition prompts based on all or part of the target replacement data in the subtitle replacement data.
[0069] In some embodiments of this disclosure, generating subtitle recognition prompts based on subtitle replacement data, including multiple target replacement data, includes: inputting each target replacement data in the subtitle replacement data into a preset position of the prompt template in reverse order of correction time to obtain subtitle recognition prompts.
[0070] The correction time can be the execution time of the subtitle correction operation corresponding to the target replacement data. The prompt word template can be a template that configures the content and position included in the subtitle recognition prompt words, used to generate subtitle recognition prompt words. This prompt word template can include requirement sub-prompt words and replacement sub-prompt words. Requirement sub-prompt words are used to indicate that the subtitle recognition model should perform subtitle recognition; this embodiment does not limit the requirement sub-prompt words. For example, the requirement sub-prompt word could be "This is a piece of audio, please recognize the content." Replacement sub-prompt words are used to indicate the requirement to replace words during the subtitle recognition process; this embodiment does not limit the replacement sub-prompt words. For example, the replacement sub-prompt word could be "The following modifications may be needed: a is changed to A, b is changed to B…". The preset position can be a pre-set position for the text to be filled in. This preset position can be marked by a pre-set placeholder. This preset position can include the preset position of the word before correction and the preset position of the word after correction. Taking the replacement sub-prompt words as an example, lowercase letters a and b can represent the preset position of the word before correction, and uppercase letters A and B can represent the preset position of the word after correction.
[0071] In this embodiment, the subtitle recognition device can input the words before and after correction from the target replacement data into the corresponding preset positions of the replacement sub-prompt words in reverse chronological order (i.e., from most recent to oldest). The input replacement sub-prompt words are then appended to the required sub-prompt words to obtain the subtitle recognition prompt words.
[0072] In the above scheme, based on the user's modification process of historically generated subtitles, the subtitle recognition prompt words for input subtitle recognition model are determined. Based on these subtitle recognition prompt words, the word replacement process is recorded, and the user's personalized modification habits are recorded, providing the model with complete replacement information, so that the final generated subtitles better meet the user's needs.
[0073] Optionally, inputting the prompt word template at preset positions in reverse chronological order of correction time yields subtitle recognition prompt words. This includes: inputting the prompt word template at preset positions in reverse chronological order of correction time to obtain initial prompt words; if the length of the initial prompt word exceeds a length threshold, then truncating and retaining the initial prompt word from beginning to end according to the text order to obtain subtitle recognition prompt words with a length not exceeding the length threshold. This retains the more recently generated target replacement data in the subtitle recognition prompt words, making the generated target subtitles more consistent with the user's recent modification habits.
[0074] Step 103: Input the target media data and subtitle recognition prompts into the subtitle recognition model, and output the target subtitles corresponding to the target media data.
[0075] The subtitle recognition model is a model for recognizing subtitles from media data such as audio or video data. This embodiment does not limit the type of subtitle recognition model. The target subtitle can be a subtitle generated by replacing words with target replacement data in the subtitle prompts during the subtitle recognition process of the target media data.
[0076] In this embodiment of the disclosure, the subtitle recognition device can input target media data and subtitle recognition prompts into the subtitle recognition big model. Under the guidance of the subtitle recognition prompts, the subtitle recognition big model performs subtitle recognition on the target media data and outputs the target subtitles.
[0077] For example, if the object name in the first video is incorrectly identified as "XXX A", and the user changes "XXX A" to "XXX A" before reposting the first video, this change is recorded by the subtitle recognition prompt. Subsequently, when the user performs subtitle recognition on the second video, the object name is directly and correctly identified as "XXX A".
[0078] The subtitle recognition scheme provided in this disclosure acquires target media data and subtitle replacement data. The subtitle replacement data includes multiple target replacement data sets, each containing a word before and a word after correction. Based on the subtitle replacement data, a subtitle recognition prompt word is generated, comprising the multiple target replacement data sets. The target media data and the subtitle recognition prompt word are input into a subtitle recognition model, and the target subtitle corresponding to the target media data is output. Using this technical solution, a subtitle recognition prompt word can be generated based on the subtitle replacement data including the word before and after correction. The subtitle recognition model can then identify the corresponding subtitle for the media data based on this prompt word. Since the subtitle recognition prompt word records the word replacement process, it provides complete replacement information for the model. This allows the model to modify the subtitle only when a word is similar to the word before correction, avoiding excessive modification and effectively improving the accuracy of subtitle recognition.
[0079] In some embodiments of this disclosure, the subtitle recognition method further includes: activating a subtitle recognition model in response to a subtitle recognition optimization activation operation.
[0080] The subtitle recognition optimization activation operation can be an operation that activates subtitle recognition through a subtitle recognition model guided by subtitle recognition prompts. This embodiment does not limit this subtitle recognition optimization activation operation. For example, the subtitle recognition optimization activation operation can be an operation to enable the recognition settings control on the subtitle recognition settings page. The recognition settings control can include a model recognition control and a data upload control. The model recognition control can be used to adjust the subtitle recognition model to an available state, and the data upload control can be used to authorize the upload of corrected data to the server's database.
[0081] In this embodiment, the user can trigger the subtitle recognition settings control on the subtitle settings page. This subtitle settings page can be used to configure the subtitle source, recognition method, style template, whether it is bilingual subtitles, whether to highlight key points, and whether to highlight invalid segments. In response to the triggering of the subtitle recognition settings control, the user is redirected to the subtitle recognition settings page, which can be a second-level page under the subtitle settings page, used for configuring subtitle recognition. This subtitle recognition page includes a model recognition control; if the user enables this control, the subtitle recognition model will be changed from an unavailable state to an available state.
[0082] On the subtitle settings page, if the model recognition control is not enabled, the data upload control is unavailable; if the model recognition control is enabled, the data upload control is available. Users can enable or disable the data upload control when it is available. When a user enables the data upload control for the first time, the subtitle recognition model will display a data upload prompt. If the user confirms the prompt, it means the user has authorized the upload of corrected data, and the subtitle replacement data upload function will be enabled. If the user cancels the prompt, it means the user has not authorized the upload of corrected data. The data upload prompt will not be displayed again subsequently.
[0083] Figure 4 is a schematic diagram of a subtitle recognition settings page provided in an embodiment of this disclosure. As shown in Figure 4, control 401 is a closed model recognition control, and control 402 is a data upload control in a grayed-out (i.e., unavailable) state. Figure 5 is a schematic diagram of another subtitle recognition settings page provided in an embodiment of this disclosure. As shown in Figure 5, control 501 is an open model recognition control, and control 502 is an open data upload control.
[0084] In the above solution, based on the user's trigger operation, it determines in fine granularity whether to start the subtitle recognition model or whether to start the subtitle recognition model and upload subtitle replacement data, thereby realizing personalized subtitle recognition according to the user's needs and improving the user experience.
[0085] The subtitle recognition method in this disclosure will be further illustrated by a specific example. Figure 6 is a schematic diagram of a subtitle recognition method provided in this disclosure. As shown in Figure 6, the subtitle recognition method includes:
[0086] Step 1: Determine the first replacement data based on the subtitle correction history data. Specifically, obtain the user's historical subtitle submissions, determine the historical modified subtitles determined based on subtitle correction operations in the historical subtitle submissions, and the historical generated subtitles corresponding to the historical modified subtitles. Generate correction data including one historical generated subtitle and one corresponding historical modified subtitle. Determine the subtitle correction history data based on multiple correction data, and align the correction data in the subtitle correction history data according to the minimum edit distance rule to obtain the first replacement data.
[0087] Step 2: Align the first replacement data and map it from characters to words to obtain the second replacement data. Through analysis of the first replacement data, convert the character-level first replacement data into word-level second replacement data.
[0088] Specifically, the subtitle recognition device can call a word segmentation tool to perform word segmentation on the historical generated subtitles and historical corrected subtitles in the first replacement data, map the characters and words before correction in the first replacement data to the corresponding word segments, obtain the words before correction and the words after correction, and determine the second replacement data based on the words before correction and the words after correction.
[0089] Step 3: Homophone Extraction. The second replacement data at the word level is filtered, selecting those that are homophones before and after correction. Specifically, the subtitle recognition device can use a Chinese character pinyin conversion tool to convert the words before and after correction into pinyin, retaining only the second replacement data with the same pinyin when ignoring tones.
[0090] Step 4: Extract nouns to obtain subtitle replacement data. Select the second replacement data where both the original and corrected words are nouns. Specifically, use part-of-speech analysis tools to determine the part of speech of the second replacement data, retain the second replacement data where both the original and corrected words are nouns, and determine this second replacement data as the target replacement data. Multiple target replacement data are then determined as subtitle replacement data.
[0091] Step 5: Generate subtitle prompts. Based on the subtitle replacement data, generate subtitle recognition prompts that include multiple target replacement data. Specifically, input the target replacement data into the preset positions of the prompt template in reverse chronological order to obtain the subtitle recognition prompts. If the length of the subtitle recognition prompt exceeds a length threshold, truncate the end of the subtitle recognition prompt, ensuring that the length of the truncated subtitle recognition prompt does not exceed the length threshold.
[0092] Step 6: Input the subtitle recognition model. Input the subtitle recognition prompts and target media data into the subtitle recognition model, and output the target subtitle.
[0093] Of the steps described above, steps 1-4 can be triggered after a user modifies a previously generated subtitle, and the subtitle replacement data is stored in the database. Steps 5-6 can be triggered when a user performs subtitle recognition operations on target media data.
[0094] The subtitle recognition scheme provided in this disclosure records the user's modification process of historically generated subtitles by recording the words before and after correction. The modification process is input into the subtitle recognition model by subtitle recognition prompt words. Based on understanding the subtitle recognition prompt words, the subtitle recognition model performs subtitle recognition on the target media data and modifies the words at appropriate positions, thereby improving the accuracy and recall of words that need to be corrected in the subtitles and thus improving the accuracy of subtitle recognition.
[0095] Figure 7 is a schematic diagram of a subtitle recognition device provided in an embodiment of this disclosure. This device can be implemented by software and / or hardware and is generally integrated into an electronic device. As shown in Figure 7, the subtitle recognition device includes:
[0096] The acquisition module 701 is used to acquire target media data and acquire subtitle replacement data, wherein the subtitle replacement data includes multiple target replacement data, and each target replacement data includes a word before correction and a word after correction;
[0097] The generation module 702 is used to generate subtitle recognition prompts including the plurality of target replacement data based on the subtitle replacement data;
[0098] The recognition module 703 is used to input the target media data and the subtitle recognition prompts into the subtitle recognition model and output the target subtitles corresponding to the target media data.
[0099] In some embodiments of this disclosure, obtaining the subtitle replacement data includes:
[0100] Obtain historical data of subtitle correction, wherein the historical data of subtitle correction includes multiple correction data, and each correction data includes historically generated subtitles and corresponding historically corrected subtitles;
[0101] Subtitle replacement data is obtained by processing historical data on subtitle corrections.
[0102] In some embodiments of this disclosure, the process of processing historical subtitle correction data to obtain subtitle replacement data includes:
[0103] The multiple corrected data are aligned using the minimum edit distance rule, and multiple first replacement data are extracted based on the alignment results;
[0104] The multiple first replacement data are mapped from characters to words to obtain multiple second replacement data;
[0105] Multiple replacement data that are homophones and belong to nouns are extracted from the multiple second replacement data and identified as multiple target replacement data, and the multiple target replacement data are identified as subtitle replacement data.
[0106] In some embodiments of this disclosure, the plurality of first replacement data are mapped from characters to words to obtain a plurality of second replacement data, including:
[0107] For each of the first replacement data, the word segmentation result of the corresponding corrected data is obtained, and based on the word segmentation result, the words before and after correction of the words before and after correction in the first replacement data are determined respectively, so as to obtain the second replacement data.
[0108] In some embodiments of this disclosure, obtaining the subtitle replacement data includes:
[0109] The pre-stored subtitle replacement data is retrieved from local storage space or a database, wherein the subtitle replacement data is determined in response to a subtitle correction operation.
[0110] In some embodiments of this disclosure, the generation module 702 is used for:
[0111] Each of the target replacement data in the subtitle replacement data is input into the preset position of the prompt word template in reverse order of the correction time to obtain the subtitle recognition prompt word.
[0112] In some embodiments of this disclosure, the subtitle recognition method further includes:
[0113] The startup module is used to start the subtitle recognition model in response to the subtitle recognition optimization activation operation.
[0114] The subtitle recognition device provided in this disclosure can execute the subtitle recognition method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0115] This disclosure provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the above-described subtitle recognition method.
[0116] Figure 8 is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0117] Referring specifically to Figure 8, which illustrates a structural schematic diagram suitable for implementing the electronic device 800 in the embodiments of this disclosure, the electronic device 800 in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in Figure 8 is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this disclosure.
[0118] As shown in Figure 8, the electronic device 800 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0119] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although FIG8 shows an electronic device 800 with various devices, it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0120] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by the processing device 801, it performs the functions defined in the subtitle recognition method of embodiments of this disclosure.
[0121] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0122] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0123] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0124] The aforementioned computer-readable medium carries one or more programs. When the aforementioned one or more programs are executed by the electronic device, the electronic device causes the electronic device to: acquire target media data and acquire subtitle replacement data, wherein the subtitle replacement data includes multiple target replacement data, each target replacement data including a word before correction and a word after correction; generate subtitle recognition prompt words including multiple target replacement data based on the subtitle replacement data; input the target media data and the subtitle recognition prompt words into a subtitle recognition model, and output the target subtitle corresponding to the target media data.
[0125] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0127] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0128] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0129] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0130] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0131] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0132] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0133] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A subtitle recognition method, comprising: Acquire target media data and acquire subtitle replacement data, wherein the subtitle replacement data includes multiple target replacement data, and each target replacement data includes a word before correction and a word after correction; Based on the subtitle replacement data, subtitle recognition prompts are generated that include the multiple target replacement data; The target media data and the subtitle recognition prompts are input into the subtitle recognition model, and the target subtitles corresponding to the target media data are output.
2. The method according to claim 1, wherein, The acquisition of subtitle replacement data includes: Obtain historical data of subtitle correction, wherein the historical data of subtitle correction includes multiple correction data, and each correction data includes historically generated subtitles and corresponding historically corrected subtitles; Subtitle replacement data is obtained by processing historical data on subtitle corrections.
3. The method according to claim 2, wherein, The process of processing historical subtitle correction data to obtain subtitle replacement data includes: The multiple corrected data are aligned using the minimum edit distance rule, and multiple first replacement data are extracted based on the alignment results; The multiple first replacement data are mapped from characters to words to obtain multiple second replacement data; Multiple replacement data that are homophones and belong to nouns are extracted from the multiple second replacement data and identified as multiple target replacement data, and the multiple target replacement data are identified as subtitle replacement data.
4. The method according to claim 3, wherein, The multiple first replacement data are mapped from characters to words to obtain multiple second replacement data, including: For each of the first replacement data, the word segmentation result of the corresponding corrected data is obtained, and based on the word segmentation result, the words before and after correction of the words before and after correction in the first replacement data are determined respectively, so as to obtain the second replacement data.
5. The method according to claim 1, wherein, The acquisition of subtitle replacement data includes: The pre-stored subtitle replacement data is retrieved from local storage space or a database, wherein the subtitle replacement data is determined in response to a subtitle correction operation.
6. The method according to any one of claims 1-5, wherein, Based on the subtitle replacement data, subtitle recognition prompts are generated that include the multiple target replacement data, including: Each of the target replacement data in the subtitle replacement data is input into the preset position of the prompt word template in reverse order of the correction time to obtain the subtitle recognition prompt word.
7. The method according to any one of claims 1-6, further comprising: In response to the subtitle recognition optimization activation operation, the subtitle recognition model is started.
8. A subtitle recognition device, comprising: The acquisition module is configured to acquire target media data and acquire subtitle replacement data, wherein the subtitle replacement data includes multiple target replacement data, and each target replacement data includes a word before correction and a word after correction; The generation module is configured to generate subtitle recognition prompts including the plurality of target replacement data based on the subtitle replacement data; The recognition module is configured to input the target media data and the subtitle recognition prompts into the subtitle recognition model, and output the target subtitles corresponding to the target media data.
9. An electronic device, comprising: processor; A memory configured to store processor-executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the subtitle recognition method according to any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, wherein, The computer program is used to execute the subtitle recognition method according to any one of claims 1-7.
Citation Information
Patent Citations
Video subtitle adding method, device and equipment and computer readable storage medium
CN115767129A
Video subtitle identification method and device, electronic equipment and storage medium
CN117315639A
Video feature extraction method and device based on large language model, and storage medium
CN117746297A
Language learning through content translation
US20230386360A1