Subtitle language recognition methods, devices, computer equipment and computer-readable media
By parsing the network abstraction layer data of supplemental enhancement information type in the video stream, the character encoding of closed captions is determined, and the subtitle language is identified by using the mapping relationship. This solves the problem of missing subtitle language information and achieves fast and accurate subtitle language identification.
Patent Information
- Application Number
- CN201911416584.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-31
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2039-12-31
AI Technical Summary
In the prior art, the subtitle language information in the video signal is missing, resulting in an inability to accurately identify the language of the subtitles, especially when the number of subtitles in the video signal is limited.
By acquiring network abstraction layer data of supplemental enhancement information type in the video stream, the character encoding of closed captions is parsed, and the language of the captions is determined by using the preset character encoding and language mapping relationship.
Even in the absence of subtitle language information, it can quickly and accurately identify the language of closed captions and supports subtitle recognition in multiple languages.
Smart Images

Figure CN113127701B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of multimedia technology, specifically to a method, apparatus, computer device, and computer-readable medium for identifying subtitle languages. Background Technology
[0002] Many broadcast and multicast video signals now include text that can be displayed on televisions or other display devices; CC (Closed Caption) text is a typical example of this. CC text is a transcription of the text portion of the audio in a video, sometimes also including descriptive narration of a small amount of background noise in the soundtrack. Initially used to facilitate communication for the hearing impaired, CC text has also been used in environments where high or low ambient noise levels make the audio portion of the signal difficult to hear, such as bars, restaurants, airports, and medical clinics. There are two types of CC text: one conforming to the NTSC (National Television Standards Committee) standard EIA (Electronic Industries Association) 608, and the other conforming to the ATSC (Advanced Television Systems Committee) standard EIA 708. EIA 608 text supports six languages: English, French, Spanish, Danish, German, and Portuguese. EIA 608 supports these six languages through three character sets: standard character set, special character set, and extended character set.
[0003] Video signals can also include other text services, which may include program information, electronic program guides, news, sports, emergency announcements, and many other types of information. These text services can be text-based teletext services such as TeleText, Ceefax, and Oracle. Currently, most text services are encoded into the VBI (vertical blanking interval) of the video signal. A smaller portion of text services may also be carried along with the audio and video components of the signal in digital video signals such as MPEG-2 (Moving Picture Experts Group) and MPEG-4 (Moving Picture Experts Group) encoded signals.
[0004] The amount of text that can be carried in any video signal is limited by the encoding system. Video signals using the VBI encoding system have only a limited capacity for carrying text. Since CC text must be carried entirely on line 21 of the VBI, the number of characters that can be encoded into each frame is limited. The raw bitstream lacks subtitle language information, therefore, the language information of the bitstream data cannot be obtained. Summary of the Invention
[0005] This disclosure addresses the aforementioned deficiencies in the prior art by providing a method, apparatus, computer device, and computer-readable medium for identifying subtitle languages.
[0006] In a first aspect, embodiments of this disclosure provide a method for identifying subtitle languages, the method comprising:
[0007] Retrieve the video stream with a preset encoding format from the bitstream;
[0008] Obtain network abstraction layer data of the supplemental enhancement information type from the video stream;
[0009] If a closed caption is obtained from the network abstraction layer data of the supplementary enhancement information type, then the character encoding of the closed caption is determined;
[0010] The language of the closed caption is determined based on the character encoding and the preset mapping relationship between character encoding and language.
[0011] Furthermore, determining the character encoding of the closed caption includes:
[0012] Convert a preset number of bytes of closed captions into a character encoding in a preset base.
[0013] Furthermore, determining the language of the closed caption based on the character encoding and a preset mapping relationship between character encoding and language includes:
[0014] Select a character encoding;
[0015] If a first determination result including a language is obtained based on the character encoding and the mapping relationship, then the language is determined to be the language of the closed caption.
[0016] Furthermore, determining the language of the closed caption based on the character encoding and the preset mapping relationship between character encoding and language further includes:
[0017] If a second determination result including multiple languages is obtained based on the character encoding and the mapping relationship, then the second determination result and the previous processing result are processed to determine the same language in the second determination result and the previous processing result. If the same language is one, then the same language is determined to be the language of the closed caption.
[0018] Furthermore, determining the language of the closed caption based on the character encoding and the preset mapping relationship between character encoding and language further includes:
[0019] If there are multiple identical languages, then other character encodings are selected, and the language of the closed captions is determined based on the selected character encoding and the mapping relationship.
[0020] Furthermore, after obtaining the closed captions from the network abstraction layer data of the supplementary enhancement information type and before determining the language of the closed captions according to the character encoding and the preset character encoding and language mapping relationship, the method further includes: determining the transmission channel of the closed captions;
[0021] The method further includes:
[0022] If it is determined that there are multiple languages in the second determination result and the previous processing result that are the same as the last character encoding of the closed caption, and the transmission channel of the closed caption is one, then the language of the closed caption is determined to be English.
[0023] In another aspect, embodiments of this disclosure also provide a subtitle language recognition device, including: a first acquisition module, a second acquisition module, a first determination module, and a second determination module;
[0024] The first acquisition module is used to acquire a video stream with a preset encoding format from the bitstream;
[0025] The second acquisition module is used to acquire Network Abstraction Layer (NET) data of the Supplemental Enhancement Information (SIE) type from the video stream;
[0026] The first determining module is used to determine the character encoding of the closed caption if a closed caption is obtained from the network abstraction layer data of the supplementary enhancement information type.
[0027] The second determining module is used to determine the language of the closed captions based on the character encoding and a preset mapping relationship between character encoding and language.
[0028] In another aspect, embodiments of this disclosure also provide a computer device, including: one or more processors and a storage device; wherein, the storage device stores one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the subtitle language recognition method provided in the foregoing embodiments.
[0029] This disclosure also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed, implements the subtitle language recognition method provided in the foregoing embodiments.
[0030] The subtitle language identification method provided in this disclosure involves acquiring a video stream with a preset encoding format from the bitstream, obtaining Network Abstraction Layer (NET) data of the supplementary enhancement information type from the video stream, and if closed subtitles are obtained from the NET data of the supplementary enhancement information type, determining the character encoding of the closed subtitles. The language of the closed subtitles is then determined based on the character encoding and a preset mapping relationship between character encodings and languages. This disclosure allows for rapid and accurate identification of the language of closed subtitles even when subtitle language information is lacking in the original bitstream. Attached Figure Description
[0031] Figure 1 A flowchart illustrating a subtitle language recognition method provided in an embodiment of this disclosure;
[0032] Figure 2 A flowchart for determining the language of closed captions based on character encoding and mapping relationships, provided as another embodiment of this disclosure;
[0033] Figure 3 This is a schematic diagram of the subtitle language recognition device provided in another embodiment of the present disclosure. Detailed Implementation
[0034] Exemplary embodiments will be described more fully below with reference to the accompanying drawings; however, these exemplary embodiments may be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will enable those skilled in the art to fully understand the scope of this disclosure.
[0035] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0036] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the said feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded.
[0037] The embodiments described herein can be described with reference to plan views and / or cross-sectional views using the ideal schematic diagrams of this disclosure. Therefore, the example illustrations can be modified according to manufacturing techniques and / or tolerances. Therefore, the embodiments are not limited to those shown in the drawings, but include modifications to configurations formed based on manufacturing processes. Therefore, the areas illustrated in the drawings are schematic in nature, and the shapes of the areas shown in the figures illustrate specific shapes of areas of an element, but are not intended to be limiting.
[0038] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0039] This disclosure provides a method for identifying the language of subtitles, such as... Figure 1 As shown, the method may include the following steps:
[0040] Step 11: Obtain the video stream with the preset encoding format from the bitstream.
[0041] In this embodiment of the disclosure, the preset encoding format can be AVC / H.264 compression format (Advanced Video Coding / MPEG-4 Part 10 compression format). The bitstream can be a channel bitstream containing encoded text received on any broadcast or multicast channel, such as a UDP (User Datagram Protocol), MPEG (Moving Picture Experts Group), or TS (Transport Stream) type bitstream. A UDP MPEG TS type bitstream can be obtained by adding it to a UDP multicast or by binding it to a UDP unicast address.
[0042] In this step, taking the UDP MPEG TS stream as an example, the subtitle language recognition device can first find the TS packet (a basic unit of the UDP MPEG TS stream) with a PID of 0. A TS packet with a PID (PacketIdentifier) of 0 is generally a PAT (Program Association Table). Based on the PIDs of the various PMT (Program Map Tables) provided in the PAT, the PMT table can be found. After parsing the PMT table, the TS packet in the stream with the same PID as the PMT table can be further found. The subtitle language recognition device can parse these TS packets and determine whether each TS packet belongs to a video stream based on its Stream_type value. If it does, it can further determine whether the video stream uses the AVC / H.264 compression format. If so, the video stream with that preset encoding format is obtained.
[0043] Step 12: Obtain network abstraction layer data of the supplemental enhancement information type from the video stream.
[0044] In this step, the AVC / H.264 compressed video stream can generally include NAL (Network Abstraction Layer) data. The subtitle language recognition device can first obtain the NAL from the video stream, and then determine whether the type of the NAL data is SEI (Supplemental Enhancement Information). If so, the SEI NAL data is obtained.
[0045] Step 13: If the closed captions are obtained from the network abstraction layer data of the supplementary enhancement information type, then determine the character encoding of the closed captions.
[0046] In this step, the subtitle language recognition device first parses the SEI NAL data, locates the data area where the payloadType value (effective payload) is 4 (i.e., the user_data_registered_itu_t_t35 data area), and then parses the user_data_registered_itu_t_t35 data area to determine if four consecutive bytes with the ASCII value GA94 exist. If they do, it can be determined that the SEI NAL data includes EIA (Electronic Industries Association) 608 type CC subtitles, and the subtitle language recognition device can then obtain the CC subtitles from the SEI NAL data. Once the CC subtitles are obtained from the SEI NAL data, the subtitle language recognition device can determine the character encoding of the CC subtitles.
[0047] Step 14: Determine the language of the closed captions based on the character encoding and the preset mapping relationship between character encoding and language.
[0048] The mapping relationship between character encoding and language is shown in Table 1.
[0049] Table 1
[0050] Character encoding Language 1 Language 2 Language 3 Language 4 Language 5 …… …… …… …… …… ……
[0051] The character encoding and language mapping relationship (i.e., Table 1) reflects the correspondence between character encodings and the languages supported by CC subtitle text, indicating which languages the characters corresponding to a character encoding can appear in. In this step, the language of the CC subtitle corresponding to a given character encoding can be determined by querying the mapping relationship using the character encoding as an index.
[0052] In this embodiment of the disclosure, the languages include French, Spanish, Danish, German, and Portuguese.
[0053] It should be noted that Table 1 may also include character sets, closed caption disassembly, display, and Unicode name.
[0054] As can be seen from steps 11-14, the subtitle language identification method provided in this embodiment of the present disclosure obtains a video stream with a preset encoding format from the bitstream, obtains network abstraction layer data of supplementary enhancement information type from the video stream, and if a closed subtitle is obtained from the network abstraction layer data of supplementary enhancement information type, then determines the character encoding of the closed subtitle, and determines the language of the closed subtitle based on the character encoding and the preset mapping relationship between character encoding and language. This embodiment of the present disclosure can quickly and accurately identify the language of closed subtitles when the subtitle language information is missing in the original bitstream.
[0055] In some embodiments, determining the character encoding of the closed captions may include converting a preset number of bytes of closed captions into a character encoding in a preset base. In this step, the caption language recognition device can determine how many bytes of closed captions to convert into one character encoding based on the character encoding method used by the closed captions. For example, if the closed captions use UTF-16 (hexadecimal Unicode), then every two bytes of closed captions can be converted into one UTF-16 character encoding. In this step, the caption language recognition device can convert all closed captions in SEI NAL into character encoding.
[0056] In some embodiments, as Figure 2 As shown, determining the language of the closed captions based on character encoding and a preset mapping relationship between character encoding and language may include the following steps:
[0057] Step 21: Select a character encoding.
[0058] In this step, the subtitle language recognition device selects a character encoding as an index to query the mapping relationship.
[0059] Step 22: Determine the language of the closed captions based on character encoding and mapping relationship. If a first determination result including one language is obtained, proceed to step 23. If a second determination result including multiple languages is obtained, proceed to step 24.
[0060] In this step, if only one language is retrieved, it indicates that the character represented by this character encoding is unique to that language and does not exist in other languages. The result of retrieving only one language is the first confirmed result. If multiple languages are retrieved, it indicates that the character represented by this character encoding is not unique to any one language but exists in multiple languages. The result of retrieving multiple languages is the second confirmed result. The subtitle language recognition device can determine whether the language in the confirmed result is unique. If the language in the result is determined to be unique, step 23 can be executed; if the language in the result is determined to be non-unique, step 24 can be executed.
[0061] The determination result of the language of the closed captions based on character encoding and mapping relationships can be represented by binary characters. The number of bits in the binary character is the number of languages in Table 1. Taking 5 languages as an example, according to the order of each language in Table 1, a 5-bit binary determination result is obtained. For example, when the determination result is 10001, 1 indicates that the character corresponding to the character encoding exists in the corresponding language (e.g., language 1 and language 5), and 0 indicates that the character corresponding to the character encoding does not exist in the corresponding language (e.g., language 2, 3, and 4). The caption language recognition device can determine whether the determination result is a first determination result or a second determination result by judging how many "1"s are present in the determination result. For example, when there is one "1" in the determination result, it means that step 22 has obtained a first determination result including one language; when there are multiple "1"s in the determination result, it means that step 22 has obtained a second determination result including multiple languages.
[0062] Step 23: Determine the language as the language of the closed captions.
[0063] In this step, the subtitle language recognition device can directly identify the unique language in the second determination result as the language of the CC subtitle.
[0064] Step 24: Process the second determined result and the previous processing result to determine the same language in the second determined result and the previous processing result.
[0065] The subtitle language identification device can further perform the current processing based on the second determination result and the previous processing result, that is, determine the parts that are the same in the multiple languages in the second determination result and the multiple languages in the previous processing result. If they exist, the same language is determined, and it is determined whether the same language is unique. If it is unique, the unique and same language can be determined as the language of the CC subtitle.
[0066] The processing operation in this step can be an AND operation. For example, if the second determined result is 10001 and the previous processing result is 11100, an AND operation can be performed between 10001 and 11100 to obtain the current processing result of 10000. The "1" appears in the first position, indicating that the language that is the same in the second determined result and the previous processing result is language 1. If the second determined result is 10001 and the previous processing result is 11001, an AND operation can be performed between 10001 and 11001 to obtain the current processing result of 10001. The "1" appears in the first and fifth positions, indicating that the language that is the same in the second determined result and the previous processing result is language 1 and language 5.
[0067] Step 25: Determine if the same language is used. If yes, proceed to step 26; otherwise, proceed to step 27.
[0068] In this step, the number of "1"s in the current processing result can be used to determine if the same language is used only once. For example, if the current processing result is 10000, and there is only one "1", it means that there is only one language that is used in the second determination result and the previous processing result, and step 26 can be executed. If the current processing result is 10011, and there are three "1"s, it means that there are multiple languages that are used in the second determination result and the previous processing result, and step 27 can be executed.
[0069] Step 26: Identify the languages with the same language as the closed captions.
[0070] In this step, when only one language is the same (i.e., when only one language is the same in the second determination result obtained in step 24 and the previous processing result), the subtitle language identification device can directly identify the same language as the language of the closed caption. In other words, the subtitle language identification device can determine the language corresponding to "1" in the current processing result according to Table 1, and identify that language as the language of the closed caption.
[0071] Step 27: Select another character encoding and determine the language of the closed captions based on the selected character encoding and mapping relationship.
[0072] In this step, when the same language is not unique (i.e., when the second determination result obtained in step 24 and the previous processing result have multiple identical languages), the subtitle language recognition device needs to select other character codes as indexes to query the mapping relationship, obtain the determination result, and continue to determine the language of the CC subtitle based on the determination result, i.e., return to execute step 22.
[0073] It should be noted that when determining the language of CC subtitles for the first time based on a character encoding and mapping relationship, the previous processing result is empty. If a second determination result including multiple languages is obtained at this time, another character encoding is directly selected, and the language of the CC subtitles is determined according to the character encoding and mapping relationship selected this time. When determining the language of CC subtitles for the second time based on a character encoding and mapping relationship, the previous processing result is also empty. If a second determination result including multiple languages is obtained at this time, the second determination result and the second determination result of the previous character encoding are processed to determine the languages that are the same.
[0074] It should be noted that when there are no identical languages in the processing results (i.e., when there are 0 "1"s in the processing results), the subtitle language recognition device can directly select other character encodings and determine the language of the CC subtitles based on the selected character encodings and mapping relationships.
[0075] Since English does not include special characters, and non-special characters in English are also included in other languages, English is not included in the character encoding and language mapping relationship (i.e., Table 1). In this embodiment of the disclosure, it can be determined whether the CC subtitle is in English through the transmission channel of the CC subtitle.
[0076] In some embodiments, after obtaining the closed captions from the network abstraction layer data of the supplementary enhancement information type and before determining the language of the closed captions according to the character encoding and the preset character encoding-language mapping relationship, the caption language identification method provided in this disclosure may further include: determining the transmission channel of the closed captions.
[0077] The subtitle language identification method provided in this embodiment may further include: if multiple languages are found to be the same in the second determination result corresponding to the last character encoding of the closed subtitle and the previous processing result, and the transmission channel of the closed subtitle is one, then the language of the closed subtitle is determined to be English.
[0078] In other words, in the subtitle language identification method provided in this embodiment, after obtaining the CC subtitle in the SEINAL data, the transmission channel of the CC subtitle can also be determined. If the language of the CC subtitle cannot be determined until the last character encoding, it can be determined whether the transmission channel of the CC subtitle is unique. If it is unique, the language of the CC subtitle can be directly determined to be English.
[0079] It should be noted that if the CC subtitle channel is not unique, the language of the CC subtitle cannot be determined. In this case, outputting the channel information is sufficient. Considering that the language within the same channel may change, the subtitle language recognition device can use the latest determined unique language as the final language of the CC subtitle within that channel in real time.
[0080] The following is a brief description of the subtitle language recognition method provided in this disclosure, using a specific embodiment as an example. The subtitle language recognition device acquires a video stream with a preset encoding format from the bitstream, obtains SEINAL data from it, acquires the CC subtitles from the SEINAL data, determines the character encoding of the CC subtitles, and then determines the language of the CC subtitles based on the character encoding and mapping relationship. The mapping relationship between the character encoding of the CC subtitles (i.e., the converted hexadecimal value) and the language is shown in Table 2:
[0081] Table 2
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089] If the selected character encoding is c3a7, querying the mapping relationship will yield a definite result of 10000, where "1" appears in the first position. This indicates that the character corresponding to c3a7 exists only in French. This is the first definite result including French, confirming that French is the language of the CC subtitles. If the character encoding is c3a9, querying the mapping relationship will yield a definite result of 11101, where "1" appears in the first, second, third, and fifth positions. This indicates that the character corresponding to c3a9 exists in French, Spanish, Danish, and Portuguese. This is the second definite result including French, Spanish, Danish, and Portuguese. At this point, the subtitle language recognition device processes the second determination result and the previous processing result. If the previous processing result is 11011, then the second determination result 11101 and the previous processing result 11011 are ANDed, and the resulting result is 11001. In this case, "1" appears in the first, second, and fifth positions respectively. The same languages are French, Spanish, and Portuguese. Since there are multiple languages, the subtitle language recognition device needs to select other character encodings and determine the language of the CC subtitles based on the selected character encoding and mapping relationship.
[0090] Based on the same technical concept, this disclosure also provides a subtitle language recognition device, such as... Figure 3 As shown, the device may include: a first acquisition module 301, a second acquisition module 302, a first determination module 303, and a second determination module 303.
[0091] The first acquisition module 301 is used to acquire a video stream with a preset encoding format from the bitstream.
[0092] The second acquisition module 302 is used to acquire network abstraction layer data of the supplementary enhancement information type from the video stream.
[0093] The first determining module 303 is used to determine the character encoding of the closed caption if the closed caption is obtained from the network abstraction layer data of the supplementary enhancement information type.
[0094] The second determining module 304 is used to determine the language of the closed captions based on the character encoding and the preset mapping relationship between character encoding and language.
[0095] In some embodiments, the first determining module 303 is used to convert a preset number of bytes of closed captions into a character encoding in a preset base.
[0096] In some embodiments, the second determining module 304 is used to select a character encoding; if a first determining result including a language is obtained according to the character encoding and mapping relationship, then the language is determined to be the language of the closed caption.
[0097] In some embodiments, the second determining module 304 is used to process the second determining result and the previous processing result if a second determining result including multiple languages is obtained according to the character encoding and mapping relationship, so as to determine the same language in the second determining result and the previous processing result. If the same language is one, the same language is determined to be the language of the closed caption.
[0098] In some embodiments, the second determining module 304 is used to select other character encodings if there are multiple identical languages, and determine the language of the closed captions based on the selected character encodings and mapping relationships.
[0099] In some embodiments, the first determining module 303 is further configured to determine the transmission channel for the closed captions.
[0100] The second determining module 304 is further configured to determine the language of the closed caption as English if the second determining result corresponding to the last character encoding of the closed caption and the previous processing result are the same for multiple languages, and the transmission channel of the closed caption is one.
[0101] This disclosure also provides a computer device, which includes one or more processors and a storage device; wherein the storage device stores one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the subtitle language recognition method provided in the foregoing embodiments.
[0102] This disclosure also provides a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed, implements the subtitle language recognition method provided in the foregoing embodiments.
[0103] It will be understood by those skilled in the art that all or some of the steps in the methods disclosed above, and the functional modules / units in the apparatus, can be implemented as software, firmware, hardware, and suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0104] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.
Claims
1. A method for identifying the language of subtitles, comprising: Retrieve the video stream with a preset encoding format from the bitstream; Obtain network abstraction layer data of the supplemental enhancement information type from the video stream; If a closed caption is obtained from the network abstraction layer data of the supplementary enhancement information type, then the character encoding of the closed caption is determined; The language of the closed caption is determined based on the character encoding and the preset mapping relationship between character encoding and language. The step of determining the language of the closed captions based on the character encoding and the preset mapping relationship between character encoding and language includes: Select a character encoding, and obtain a determination result based on the character encoding and the mapping relationship; wherein, the determination result is represented by binary characters, the number of bits of the binary characters is the number of languages in the mapping relationship between the character encoding and the language, and a preset character in the binary characters indicates that the character corresponding to the character encoding exists in the corresponding language, and the preset character is 0 or 1; The language of the closed captions is determined based on the determination result; wherein, the language of the closed captions is determined based on the number of preset characters.
2. The method as described in claim 1, wherein, Determining the character encoding of the closed caption includes: Convert a preset number of bytes of closed captions into a character encoding in a preset base.
3. The method as described in claim 1 or 2, wherein, Determining the language of the closed captions based on the determination result includes: If the determination result is a first determination result for a language, then the language is determined to be the language of the closed caption.
4. The method of claim 3, wherein, The step of determining the language of the closed captions based on the determination result further includes: If the determination result is a second determination result for multiple languages, then the second determination result and the previous processing result are processed to determine the same language in the second determination result and the previous processing result. If the same language is one, then the same language is determined to be the language of the closed caption.
5. The method of claim 4, wherein, The step of determining the language of the closed captions based on the determination result further includes: If there are multiple identical languages, then other character encodings are selected, and the language of the closed captions is determined based on the selected character encoding and the mapping relationship.
6. The method of claim 5, wherein, After obtaining the closed captions from the network abstraction layer data of the supplementary enhancement information type, and before determining the language of the closed captions according to the character encoding and the preset character encoding-language mapping relationship, the method further includes: determining the transmission channel of the closed captions; The method further includes: If it is determined that there are multiple languages in the second determination result and the previous processing result that are the same as the last character encoding of the closed caption, and the transmission channel of the closed caption is one, then the language of the closed caption is determined to be English.
7. A subtitle language recognition device, comprising: First acquisition module, second acquisition module, first determination module, and second determination module; The first acquisition module is used to acquire a video stream with a preset encoding format from the bitstream; The second acquisition module is used to acquire Network Abstraction Layer (NET) data of the Supplemental Enhancement Information (SIE) type from the video stream; The first determining module is used to determine the character encoding of the closed caption if a closed caption is obtained from the network abstraction layer data of the supplementary enhancement information type. The second determining module is used to determine the language of the closed captions based on the character encoding and a preset mapping relationship between character encoding and language. The second determining module is specifically used to: select a character encoding, obtain a determining result based on the character encoding and the mapping relationship, and determine the language of the closed caption based on the determining result; wherein, the determining result is represented by binary characters, the number of bits of the binary characters is the number of languages in the mapping relationship between the character encoding and the language, and a preset character in the binary characters indicates that the character corresponding to the character encoding exists in the corresponding language; the preset character is 0 or 1; the language of the closed caption is determined based on the number of preset characters.
8. A computer device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the subtitle language recognition method as described in any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed, it implements the subtitle language recognition method as described in any one of claims 1-6.
Citation Information
Patent Citations
Video extension code setting and video playing method and system
CN106658225A
Method and device for identifying external subtitle language of video file
CN108600856A