Subtitle file synthesis method, device, equipment, medium and program product

CN122802725APending Publication Date: 2026-09-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611290371.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-25
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]然而,这种方式需要存储、维护相同语种的多份字幕文件,存储开销较大

Benefits of technology

本申请实施例提供了一种多版本高亮字幕文件合成与分发方案。将完整字幕文件分为原始字幕文件与不同的高亮元数据分离存储,基于第一账号选择的第一视频和第一词汇等级对应的播放请求,获取原始字幕文件以及相应的高亮元数据并合成高亮字幕文件,以在播放第一视频并显示第一视频的字幕内容的过程中,高亮第一词汇等级对应的至少一部分高亮词汇。一方面,对于同一视频、同一语种,仅存储和维护一份原始字幕文件,在获取到播放请求时基于不同的高亮元数据实时合成高亮字幕文件,实现了同一视频同一语种下的多版本高亮字幕展示,避免为不同词汇等级生成多份视频源,也避免为不同词汇等级存储多份完整字幕文件,减少了存储开销和算力成本;另一方面,在新增词汇等级时或者更新词汇时,仅需变更对应的高亮元数据,无需变更原始字幕文件,显著降低存储和转码成本;再一方面,还无需改造播控系统和播放器,降低了改造范围和上线风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802725A_ABST
    Figure CN122802725A_ABST
Patent Text Reader

Abstract

The application discloses a subtitle file synthesis method and device, equipment, medium and program product, and relates to the technical field of computers. The method comprises the following steps: obtaining a playing request of a first video, wherein the playing request is used for indicating the first video and a first vocabulary level, and the first vocabulary level corresponds to a highlight vocabulary to be highlighted in the subtitle content of the first video; based on the playing request, an original subtitle file and highlight metadata are obtained, wherein the original subtitle file is used for indicating the subtitle content of the first video, and the highlight metadata is used for describing highlight information of the highlight vocabulary corresponding to the first vocabulary level; based on the original subtitle file and the highlight metadata, a highlight subtitle file is synthesized; and the highlight subtitle file is used for indicating that at least part of the highlight vocabulary corresponding to the first vocabulary level is highlighted in the process of playing the first video and displaying the subtitle content of the first video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, medium, and program product for synthesizing subtitle files. Background Technology

[0002] With the continuous development of online video services, the number of multilingual video resources with subtitles is growing rapidly. In a video playback scenario, it is necessary to highlight words at different levels in the subtitles of the same video based on the differences in language cognition levels of users of different ages and grades. This requires the same video to be accompanied by multiple differentiated subtitle display effects.

[0003] In related technologies, multiple subtitle files for the same video are generated in advance based on vocabulary at multiple vocabulary levels, and these subtitle files are stored as independent files. During video playback, the user account selects a vocabulary level, and the player loads the corresponding subtitle file and completes the video rendering.

[0004] However, this method requires storing and maintaining multiple subtitle files in the same language, resulting in significant storage overhead. Summary of the Invention

[0005] This application provides a method, apparatus, device, medium, and program product for synthesizing subtitle files. The technical solution is as follows.

[0006] On the one hand, a method for synthesizing subtitle files is provided, the method comprising: Obtain a playback request for a first video, the playback request being used to indicate the first video and a first vocabulary level, the first vocabulary level corresponding to a highlighted vocabulary to be highlighted in the subtitle content of the first video; Based on the playback request, the original subtitle file and highlight metadata are obtained. The original subtitle file is used to indicate the subtitle content of the first video, and the highlight metadata is used to describe the highlight information of the highlighted words corresponding to the first word level. Based on the original subtitle file and the highlighted metadata, a highlighted subtitle file is synthesized. The highlighted subtitle file is used to indicate that during the playback of the first video and the display of the subtitle content of the first video, at least a portion of the highlighted words corresponding to the first word level are highlighted.

[0007] On the other hand, a video playback method is provided, the method comprising: In response to the selection action, determine the first video and the first vocabulary level; Send a playback request for the first video, the playback request being used to indicate the first video and the first vocabulary level, the first vocabulary level corresponding to the highlighted words to be highlighted in the subtitle content of the first video; Receive the returned highlighted subtitle file, and, based on the highlighted subtitle file, during the process of playing the first video and displaying the subtitle content of the first video, highlight at least a portion of the highlighted words corresponding to the first word level; The highlighted subtitle file is synthesized based on the original subtitle file obtained from the playback request and the highlighted metadata. The original subtitle file is used to indicate the subtitle content of the first video, and the highlighted metadata is used to describe the highlighted information of the highlighted words corresponding to the first vocabulary level.

[0008] On the other hand, a subtitle file synthesis apparatus is provided, the apparatus comprising: The acquisition module is used to acquire a playback request for a first video. The playback request is used to indicate the first video and a first vocabulary level. The first vocabulary level corresponds to a highlighted word to be highlighted in the subtitle content of the first video. The processing module is used to obtain the original subtitle file and highlight metadata based on the playback request. The original subtitle file is used to indicate the subtitle content of the first video, and the highlight metadata is used to describe the highlight information of the highlighted words corresponding to the first word level. A synthesis module is used to synthesize a highlighted subtitle file based on the original subtitle file and the highlighted metadata. The highlighted subtitle file is used to indicate that at least a portion of the highlighted words corresponding to the first word level are highlighted during the playback of the first video and the display of the subtitle content of the first video.

[0009] In some embodiments, the synthesis module is configured to: The original subtitle file is parsed to obtain an array of subtitle entries; Based on the subtitle entry array and the highlight metadata, highlight tags are attached to the highlight words in the subtitle entry array that meet the highlight conditions, resulting in a processed subtitle entry array; The processed subtitle entry array is serialized into the highlighted subtitle file.

[0010] In some embodiments, the synthesis module is configured to: Align the subtitle entry array with the highlighted metadata to determine the matched entries; Locate the highlighted words in each of the hit entries that satisfy the highlighted conditions; The highlighted words that meet the highlighted conditions are attached with the highlighted tags to obtain the processed subtitle entry array.

[0011] In some embodiments, the synthesis module is configured to: For each hit entry containing multiple lines of text, identify the line to be highlighted; The highlighted words that meet the highlighted conditions are located sequentially in the row to be highlighted.

[0012] In some embodiments, the synthesis module is configured to: For the j-th highlighted word in the i-th line to be highlighted, initialize a counter and a scanning cursor. The counter is used to record the number of times the j-th highlighted word has been scanned in the line to be highlighted. The scanning cursor is used to mark the starting position for finding the j-th highlighted word from the text to be scanned in the line to be highlighted. i and j are positive integers. Search for the j-th highlighted word from the starting position of the scanning cursor mark; If the j-th type of highlighted word is not found, the processing of the j-th type of highlighted word ends; If the j-th type of highlighted word is found, obtain the current occurrence count of the j-th type of highlighted word recorded by the counter, and perform processing on the j-th type of highlighted word based on the pre-set sequence number array and the current occurrence count; Update the value of the counter, and return to the step of searching for the j-th type of highlighted word from the starting position of the scanning cursor mark until the value of the counter exceeds the maximum index in the index array, and determine the j-th type of highlighted word that meets the highlighting condition from all the j-th type of highlighted words in the i-th row to be highlighted; The sequence number array is a set of occurrence numbers of highlighted words that meet the highlighting conditions.

[0013] In some embodiments, the synthesis module is configured to: If the current occurrence count falls into the sequence number array, the j-th type of highlighted word with the current occurrence count is determined as the j-th type of highlighted word that satisfies the highlighting condition. Also, the remaining text in the i-th line to be highlighted is determined, the scanning cursor is reset to the starting position of the remaining text, and the process of searching for the j-th type of highlighted word from the starting position marked by the scanning cursor is returned. If the current occurrence count does not fall into the sequence number array, it is determined that the j-th type of highlighted word with the current occurrence count does not meet the highlighting condition. The scanning cursor is then moved to the end of the j-th type of highlighted word with the current occurrence count, and the process of searching for the j-th type of highlighted word from the beginning of the scanning cursor mark is returned.

[0014] In some embodiments, the synthesis module is configured to: Based on the end position of the j-th highlighted word with the current occurrence frequency, the text to be scanned in the i-th line to be highlighted that is located after the end position is determined as the remaining text in the i-th line to be highlighted.

[0015] In some embodiments, the synthesis module is configured to: For each of the multi-line text entries that are hit, the proportion of letters is calculated line by line. The line or lines of text in which the proportion of the letters is greater than the threshold are identified as the lines to be highlighted.

[0016] In some embodiments, the synthesis module is further configured to: Read the color configuration table from the configuration center; Write the color configuration table into the style block of the original subtitle file to generate a global style declaration; The color configuration table is used to indicate the class name and the corresponding color value.

[0017] In some embodiments, the synthesis module is configured to: The highlighted words that meet the highlighted conditions are attached with the highlighted tags corresponding to the class names in the global style declaration to obtain the processed subtitle entry array.

[0018] In some embodiments, the highlighted metadata includes at least one of the following: Single-sentence subtitle highlighting information; Highlighted words and their descriptions; Complete data consisting of multiple highlighted single-sentence subtitles.

[0019] In some embodiments, the single-sentence caption highlighting information includes at least one of the following: Subtitle number; Subtitle start time; The set of highlighted words in the caption.

[0020] In some embodiments, the highlighted word description information includes at least one of the following: Highlighted words to be highlighted; Highlighted labels used to indicate colors; The same highlighted words appear in the same subtitle; Vocabulary identifiers used to indicate highlighted words.

[0021] On the other hand, a video playback device is provided, the device comprising: The determination module is used to determine the first video and the first vocabulary level in response to the selection operation; The sending module is used to send a playback request for the first video. The playback request is used to indicate the first video and the first vocabulary level. The first vocabulary level corresponds to the highlighted vocabulary to be highlighted in the subtitle content of the first video. The highlighting module is used to receive the returned highlighted subtitle file, and, based on the highlighted subtitle file, highlight at least a portion of the highlighted words corresponding to the first word level during the process of playing the first video and displaying the subtitle content of the first video; The highlighted subtitle file is synthesized based on the original subtitle file obtained from the playback request and the highlighted metadata. The original subtitle file is used to indicate the subtitle content of the first video, and the highlighted metadata is used to describe the highlighted information of the highlighted words corresponding to the first vocabulary level.

[0022] On the other hand, a computer device is provided, the computer device comprising: a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the subtitle file synthesis method as described above, and / or to implement the video playback method as described above.

[0023] On the other hand, a computer-readable storage medium is provided, which stores a computer program that is loaded and executed by a processor to implement the subtitle file synthesis method described above, and / or to implement the video playback method described above.

[0024] On the other hand, a computer program product is provided, the computer program product including computer instructions stored in a computer-readable storage medium, a processor obtaining the computer instructions from the computer-readable storage medium, causing the processor to load and execute to implement the subtitle file synthesis method as described above, and / or to implement the video playback method as described above.

[0025] The beneficial effects of the technical solutions provided in this application include at least the following: This application provides a scheme for synthesizing and distributing multi-version highlighted subtitle files. The complete subtitle file is divided into an original subtitle file and different highlighted metadata, which are stored separately. Based on the playback request corresponding to the first video and the first vocabulary level selected by the first account, the original subtitle file and the corresponding highlighted metadata are obtained and synthesized into a highlighted subtitle file. This ensures that during the playback of the first video and the display of its subtitle content, at least a portion of the highlighted vocabulary corresponding to the first vocabulary level is highlighted. On one hand, for the same video and the same language, only one original subtitle file is stored and maintained. When a playback request is received, a highlighted subtitle file is synthesized in real time based on different highlighted metadata, enabling the display of multiple versions of highlighted subtitles for the same video and the same language. This avoids generating multiple video sources for different vocabulary levels and also avoids storing multiple complete subtitle files for different vocabulary levels, reducing storage overhead and computing power costs. On the other hand, when adding a new vocabulary level or updating vocabulary, only the corresponding highlighted metadata needs to be changed, without changing the original subtitle file, significantly reducing storage and transcoding costs. Furthermore, it eliminates the need to modify the broadcast control system and player, reducing the scope of modification and the risk of going live. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a structural block diagram of a computer system provided in an exemplary embodiment of this application; Figure 2 This is a schematic diagram of a subtitle file synthesis method and a video playback method provided in an exemplary embodiment of this application; Figure 3 This is a flowchart of a subtitle file synthesis method provided in an exemplary embodiment of this application; Figure 4 This is a schematic diagram illustrating the correspondence between the original subtitle file and different highlight metadata provided in an exemplary embodiment of this application; Figure 5 This is a schematic diagram illustrating selective highlighting according to the occurrence sequence number provided in an exemplary embodiment of this application; Figure 6 This is a flowchart of a video playback method provided in an exemplary embodiment of this application; Figure 7 This is a schematic diagram of an offline generation link for highlighted metadata provided in an exemplary embodiment of this application; Figure 8This is a schematic diagram of the runtime distribution chain of highlighted subtitle files provided in an exemplary embodiment of this application; Figure 9 This is a schematic diagram of precise multidimensional cache invalidation provided in an exemplary embodiment of this application; Figure 10 This is a schematic diagram of the overall architecture of the offline generation link of highlighted metadata and the runtime distribution link of highlighted subtitle files provided in an exemplary embodiment of this application; Figure 11 This is a schematic diagram illustrating the separate storage of the original subtitle file and the highlighted metadata, provided in an exemplary embodiment of this application. Figure 12 This is a schematic diagram of a highlighted metadata structure provided in an exemplary embodiment of this application; Figure 13 This is a schematic diagram illustrating the stripping of highlighted metadata information provided in an exemplary embodiment of this application; Figure 14 This is a schematic diagram illustrating selective highlighting of highlighted words provided in an exemplary embodiment of this application; Figure 15 This is a block diagram of a subtitle file synthesis apparatus provided in an exemplary embodiment of this application; Figure 16 This is a block diagram of a video playback device provided in an exemplary embodiment of this application; Figure 17 This is a structural block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0030] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0031] It should be understood that although the terms first, second, etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0032] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user and user account data and related operations. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps to collect user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise, if no confirmation is received from the user, the steps to collect user data end, and no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0033] This document briefly introduces the terminology used in the embodiments of this application.

[0034] Subtitle files are files that store the subtitle content and timing information of a video, used to display corresponding dialogue or explanatory text during video playback. Subtitle file formats include at least one of the following: SRT (SubRip Text), ASS (Advanced Sub Station), VTT (Web VTT), and SMI (SAMI).

[0035] In some embodiments, the subtitle file uses the VTT format. VTT is a W3C standard subtitle format that supports the use of style blocks (STYLE), pseudo-elements (::cue), and class name tags (...).<c.classname> The style block defines CSS style rules in the file header, which apply to the entire subtitle file or specific classes; pseudo-elements control the overall format of the subtitles through CSS, such as at least one of font, color, shadow, and background; class name tags apply custom class names inline in the subtitle file, and, in conjunction with class name tags in the style block, achieve local style customization, such as at least one of highlight, variant, and italic.

[0036] Original subtitle file: This is the subtitle file delivered by the production company that does not embed any subtitle style information. It is used to indicate the subtitle content of the video, or in other words, to indicate the subtitle content itself. Subtitle style information includes at least one of the following: font, color, size, position, shadow, background, effects, and alignment.

[0037] Highlighted subtitle file: A subtitle file synthesized from the original subtitle file of the video and highlighting metadata. It is used to indicate that during the playback of the video and the display of the subtitle content of the video, at least a portion of the highlighted words corresponding to the word level selected by the user account are highlighted.

[0038] Highlight metadata: This is a data structure stored independently of the original subtitle file. It describes which words in the subtitle content of a video, language, and vocabulary level need to be highlighted, what color they should be highlighted, how many times they appear in the subtitle of a sentence, and the business information of the highlighted words.

[0039] A tiered vocabulary database is a collection of words organized by the operations team according to certain rules, such as age or grade level. Different vocabulary levels correspond to different vocabulary sets, and each set differs in at least one of the following: the number of words, word length, and word difficulty. In some embodiments, vocabulary levels include: beginner level, intermediate level, and expert level.

[0040] Offline subtitle processing service: This service generates and stores highlighted metadata offline. The inputs to the offline subtitle processing service include the original subtitle file and a hierarchical vocabulary database. The processing includes subtitle parsing, vocabulary matching, position recording, and highlighted metadata generation. The output of the offline subtitle processing service is highlighted metadata stored with the video identifier, language, and vocabulary level of a video.

[0041] Subtitle proxy service: Receives business information transmitted through the broadcast control system, generates a Uniform Resource Locator (URL) to retrieve the highlighted subtitle file, and points it to the Content Delivery Network (CDN). The broadcast control system refers to the playback control system, responsible for signal transmission scheduling and equipment management to achieve automatic and secure video playback.

[0042] Figure 1 This is a structural block diagram of a computer system provided in an exemplary embodiment of this application. The computer system 100 can be a system architecture for implementing a subtitle file synthesis method and / or a video playback method. The computer system 100 includes: a first terminal 120, a server 140, and a second terminal 160.

[0043] The first terminal 120 can be an electronic device such as a mobile phone, tablet computer, vehicle terminal (vehicle system), wearable device, personal computer (PC), unmanned reservation terminal, smart home appliance, smart voice interaction device, unmanned vending terminal, etc. The first terminal 120 installs and runs a client of a target application that supports a player. This target application includes at least one of the following: local player, online video application, streaming media, short video application, live streaming application, game application, instant messaging application, travel application, social application, and lifestyle service application. The first terminal 120 can also install and run other applications that provide playback functionality; this application embodiment does not limit this. Furthermore, this application embodiment does not limit the form of the target application, including but not limited to applications (Apps), mini-programs, etc., installed on the first terminal 120, and it can also be in web page form.

[0044] The first terminal 120 is a terminal used by a first user. In some embodiments, the first user logs into a first account on the client of a target application that supports a player on the first terminal 120. The first account selects a first video to be played and a first vocabulary level. The first vocabulary level corresponds to highlighted words in the subtitle content of the first video. During the playback of the first video and the display of the subtitle content of the first video, the client of the target application that supports the player highlights at least a portion of the highlighted words corresponding to the first vocabulary level.

[0045] The first terminal 120 and the second terminal 160 are connected to the server 140 via a wireless network or a wired network.

[0046] Server 140 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud servers, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Server 140 includes at least one of the following: a single server, multiple servers, a cloud computing platform, and a virtualization center.

[0047] Server 140 includes processor 144 and memory 142. Memory 142 includes receiving module 1421, control module 1422, and sending module 1423. Receiving module 1421 is used to receive requests sent by clients; control module 1422 is used to control screen rendering; sending module 1423 is used to send responses to clients; server 140 is used to provide background services for clients of first terminal 120 and second terminal 160.

[0048] Optionally, server 140 undertakes the main computing work, and first terminal 120 and second terminal 160 undertake secondary computing work; or, server 140 undertakes secondary computing work, and first terminal 120 and second terminal 160 undertake the main computing work; or, server 140, first terminal 120 and second terminal 160 collaborate in a distributed computing architecture.

[0049] The second terminal 160 can be an electronic device such as a mobile phone, tablet computer, in-vehicle terminal (vehicle system), wearable device, personal computer (PC), unmanned reservation terminal, smart home appliance, smart voice interaction device, unmanned vending terminal, etc. The second terminal 160 installs and runs a client of a target application that supports a player. This target application includes at least one of the following: local player, online video application, streaming media, short video application, live streaming application, game application, instant messaging application, travel application, social application, and lifestyle service application. The second terminal 160 can also install and run other applications that provide playback functionality; this application embodiment does not limit this. Furthermore, this application embodiment does not limit the form of the target application, including but not limited to applications (Apps), mini-programs, etc., installed on the second terminal 160, and can also be in web page form.

[0050] The second terminal 160 is a terminal used by a second user. In some embodiments, the second user logs into a second account on the client of a target application that supports a player on the second terminal 160. The second account selects a second video to be played and a second vocabulary level. The second vocabulary level corresponds to highlighted words in the subtitle content of the second video. During the playback of the second video and the display of the subtitle content of the second video, the client of the target application that supports the player highlights at least a portion of the highlighted words corresponding to the second vocabulary level.

[0051] In some embodiments, the first video and the second video, and the first vocabulary level and the second vocabulary level, may be the same or different, and this application embodiment does not limit this. Taking the first video and the second video being the same, and the first vocabulary level and the second vocabulary level being different as an example, then for the first terminal 120 and the second terminal 160, during the process of playing the first video and displaying the subtitle content of the first video, the video frame of the first video is the same, the content of the subtitle content of the first video is the same, and at least a portion of the highlighted words in the subtitle content of the first video are different for each.

[0052] Optionally, the client installed on the first terminal 120 and the second terminal 160 may be the same, or the client installed on the two terminals may be the same type of client from different control system platforms. This application embodiment does not limit the form of the client installed on the first terminal 120 and the second terminal 160. The first terminal 120 can refer to one of multiple terminals, and the second terminal 160 can refer to one of multiple terminals; this embodiment only uses the first terminal 120 and the second terminal 160 as examples. The first terminal 120 and the second terminal 160 may have the same device type, but different device models; this application embodiment does not limit this.

[0053] Those skilled in the art will understand that the number of the first terminal 120 and the second terminal 160 can be more or less. For example, there can be one first terminal 120 and one second terminal 160, or there can be multiple first terminals 120 and multiple second terminals 160. This application does not limit the number or type of the first terminal 120 and the second terminal 160 in its embodiments.

[0054] With the continuous development of online video services, the number of multilingual video resources with subtitles is growing rapidly. In a video playback scenario, it is necessary to highlight words at different levels in the subtitles of the same video based on the differences in language cognition levels of users of different ages and grades. This requires the same video to be accompanied by multiple differentiated subtitle display effects.

[0055] The following are some solutions for achieving this requirement in related technologies.

[0056] Hard captioning solution: Using words at different vocabulary levels, after rendering the caption content, the text and styles are directly forced into the video during transcoding. When highlighting words at different vocabulary levels, multiple video sources need to be obtained through transcoding. This method has high compatibility, but requires storing multiple video sources, significantly increasing transcoding and storage costs.

[0057] Multi-version soft subtitle solution: This method pre-generates multiple subtitle files for the same video based on vocabulary levels. These subtitle files are stored as independent files. During video playback, the user account selects a vocabulary level, and the player loads the corresponding subtitle file and completes the video rendering. While this method avoids multiple video files, it still requires storing and managing multiple subtitle files, resulting in significant storage and file management overhead. Furthermore, it necessitates modifications to core services such as the broadcast control system and storage system, leading to a vast scope of changes and extremely high deployment risks.

[0058] The player-side dynamic processing solution involves downloading the original subtitle file and vocabulary at different levels, and then dynamically deciding which words need to be highlighted. This method increases the logical complexity on the player side. Furthermore, different terminals and player kernels have varying capabilities in subtitle parsing and rendering, leading to high costs for multi-platform adaptation and consistency verification. It also prevents manual intervention in the display of a specific video, subtitle, or individual word.

[0059] The solutions in the related technologies have the following drawbacks.

[0060] Disadvantage 1: Highlighting with multiple word levels leads to increased storage of video source or subtitle files. Using a hard subtitle scheme or a multi-version soft subtitle scheme, assuming there are n word levels, requires n subtitle files or n video sources, multiplying storage, transcoding, and file management overhead.

[0061] Disadvantage 2: Subtitle systems typically do not support multiple versions of subtitles in the same language. Related subtitle systems usually store subtitle files according to video identifier and language, with only one subtitle file corresponding to the same video and language. To achieve subtitle files in the same language with multiple vocabulary levels, extensive modifications to the storage system, index management system, and broadcast control system are required, resulting in high modification costs and significant deployment risks.

[0062] Disadvantage 3: High cost of player-side modification and cross-platform consistency. Implementing highlighting logic on each player requires each player to implement subtitle parsing, vocabulary matching, style insertion, and rendering functions. Hardware and kernel differences between different player platforms can easily lead to inconsistencies.

[0063] Disadvantage 4: Ordinary string replacement struggles to handle repeated words and selective highlighting logic. In the subtitles of a video, the same word may appear multiple times in a single subtitle line. If a global replacement is performed based on the subtitle text, all identical words in the same subtitle line may be highlighted, failing to meet the need to highlight only the word appearing at a specific location, and also not supporting the requirement for manual modification of highlighted information by operations staff.

[0064] This application provides a multi-version highlighted subtitle file synthesis and distribution scheme, which is a combination of storing the original subtitle file and different highlighted metadata separately, synthesizing the highlighted subtitle file at runtime using the Content Delivery Network (CDN), accurately highlighting based on the occurrence position index of highlighted words, and accurately invalidating multi-dimensional cache.

[0065] In some embodiments, during the offline phase, the highlighting information of words corresponding to different vocabulary levels is extracted from the original subtitle file into different highlighting metadata. During the runtime phase, a subtitle proxy service points to a highlighted subtitle CDN, and based on the video and vocabulary level selected by the user account, the original subtitle file and highlighting metadata are dynamically synthesized into a highlighted subtitle file, thereby enabling the display of multiple versions of highlighted subtitles for the same video and the same language. In other words, this application embodiment proposes a multi-version highlighted subtitle synthesis and distribution mechanism that, within the limitations of existing video platform architecture, does not add multiple subtitle files of the same language, does not generate multiple video sources, and does not modify the player.

[0066] The embodiments of this application mainly include the following improvements.

[0067] Improvement 1: Replace multiple subtitle files or video sources with a single original subtitle file and different highlight metadata. Separate the original subtitle file and highlight metadata; for the same video and language, only one original subtitle file is retained. When adding or adjusting vocabulary levels or words, only the highlight metadata needs to be updated; the original subtitle file does not need to be modified.

[0068] Improvement point two: Instead of the broadcast control system directly scheduling multiple versions of subtitles, generate a highlighted subtitle CDN address, or highlighted subtitle CDN routing address. The current broadcast control system returns subtitle addresses based on video and language. To support multiple subtitle versions, multiple subtitle files need to be generated in advance, requiring modifications to the broadcast control scheduling architecture and subtitle storage architecture, which carries significant risk. Generating only a single proxy routing address maintains the same logic as before.

[0069] Improvement 3: Instead of generating multiple subtitle files in advance, the highlighted subtitle CDN synthesizes the subtitle file at runtime. The logic for synthesizing the original subtitle file and highlight metadata is migrated to the highlighted subtitle CDN. The player sends a playback request to the highlighted subtitle CDN. The highlighted subtitle CDN receives the playback request, checks its cache based on the request, and if the cache is not found, it retrieves the original subtitle file and highlight metadata, synthesizes the highlighted subtitle file, and returns the standard VTT format highlighted subtitle file to the player. This method allows highlighted subtitle files to be synthesized according to playback requests and cached at edge nodes. For popular videos, playback requests can directly hit the cache.

[0070] Figure 2 This is a schematic diagram illustrating a subtitle file synthesis method and a video playback method provided in an exemplary embodiment of this application. Both the subtitle file synthesis method and the video playback method are executed by a computer device. The computer device executing the subtitle file synthesis method can be... Figure 1 Server 140, the computer device that executes the video playback method can be Figure 1 The first terminal 120 has a client application that supports the player installed and running on it. The first terminal 120 is logged into by the first user's first account. The overall steps of the subtitle file synthesis method and video playback method are briefly described below.

[0071] Step 1: Reference Figure 2 In (1), the player interface 10 is displayed. The player interface 10 includes: a video display area 11, a subtitle display area 12, and a vocabulary level selection area 13. The video display area 11 is used to display the first video to be played. The subtitle display area 12 is used to display the subtitle content of the first video. The vocabulary level selection area 13 is used to display the candidate vocabulary levels. Each candidate also displays the number of words to be highlighted. For example, the candidate includes: all vocabulary · 28 words, beginner · 5 words, advanced · 17 words, and expert · 6 words. The first account performs a selection operation in the player interface 10, selecting the first video to be played and the first vocabulary level "advanced · 17 words". The player responds to the selection operation and confirms the first video and the first vocabulary level "advanced · 17 words".

[0072] Step 2: Reference Figure 2 In (2), the player sends a playback request for the first video to the highlighted subtitle CDN. The playback request is used to indicate the first video and the first vocabulary level "Advanced 17 words". The first vocabulary level corresponds to the highlighted words to be highlighted in the subtitle content of the first video.

[0073] Step 3: Reference Figure 2 In (2), the CDN obtains the playback request of the first video. The playback request is used to indicate the first video and the first vocabulary level "Advanced 17 words". The first vocabulary level "Advanced 17 words" corresponds to the highlighted words to be highlighted in the subtitle content of the first video. For example, the highlighted words in the Chinese subtitles are "like" and "apple", and the highlighted words in the English subtitles are "like" and "apples".

[0074] Step 4: Reference Figure 2In (2), the CDN obtains the original subtitle file 21 and the highlight metadata 22 based on the playback request. The original subtitle file 21 is used to indicate the subtitle content of the first video, and the highlight metadata 22 is used to describe the highlight information of the highlighted words corresponding to the first vocabulary level "Advanced 17 words", such as color and the position of the subtitle in the current sentence.

[0075] Step 5: Reference Figure 2 In (2), the highlighted subtitle CDN synthesizes a highlighted subtitle file 23 based on the original subtitle file 21 and the highlighted metadata 22. The highlighted subtitle file 23 is used to indicate that at least a portion of the highlighted words corresponding to the first vocabulary level "Advanced 17 words" are highlighted during the playback of the first video and the display of the subtitle content of the first video, and to return the highlighted subtitle file 23 to the player.

[0076] Step 6: Reference Figure 2 In (2), the player receives the returned highlighted subtitle file, and, as referenced... Figure 2 In (1), based on the highlighted subtitle file, during the process of playing the first video and displaying the subtitle content of the first video, at least a portion of the highlighted words corresponding to the first vocabulary level "Advanced 17 words" are highlighted, for example, "like", "apple", "like" and "apples" in the subtitle content of the first video are highlighted.

[0077] Next, the method for synthesizing subtitle files will be explained in detail.

[0078] Figure 3 This is a flowchart illustrating a subtitle file synthesis method provided in an exemplary embodiment of this application. The method is executed by a computer device, which may be... Figure 1 The server 140, a computer device, is also used to implement a Content Delivery Network (CDN). The method includes at least some of steps 210, 220, and 230.

[0079] Step 210: Obtain the playback request of the first video. The playback request is used to indicate the first video and the first vocabulary level. The first vocabulary level corresponds to the highlighted words to be highlighted in the subtitle content of the first video.

[0080] The first video is the video selected by the first account and intended to be played in the player. The first account is the user account used by the first user to log in. The type of the first video varies depending on the application scenario. For example, in the education field, the first video could be a bilingual video (including Chinese and English) for English learning. In the entertainment field, the first video could be an entertainment video or a short video.

[0081] The first video also includes subtitles. The subtitles in the first video are used to assist the first user watching the video in understanding the video's visual and audio content. In some embodiments, the subtitles include at least one of the following: dialogue subtitles, explanatory subtitles, title subtitles, interactive subtitles, and decorative subtitles. Dialogue subtitles are used to present at least one of task dialogue, narration, or monologue content, typically one or more lines of text, and are located at the bottom of the screen. For example, if the first video is a bilingual video including Chinese and English, a single dialogue subtitle includes one line of Chinese subtitles and one line of English subtitles. Explanatory subtitles are used to represent at least one of time, location, character introductions, and scene descriptions. Title subtitles are usually located at the beginning or end of the video and include at least one of the following: title, production team list, cast and crew, and production company. Interactive subtitles include at least one of the following: question subtitles, voting subtitles, and bullet comments. Decorative subtitles include at least one of the following: artistic fonts and dynamic effects.

[0082] Vocabulary grade (grade_level), also known as vocabulary difficulty level, is a pre-defined level that can be set according to actual technical needs. Different vocabulary grades correspond to different levels of vocabulary difficulty, which can be represented by at least one of the following: word frequency, word length, word semantics, word structural complexity, and word pronunciation complexity. For example, vocabulary grades may include at least one of beginner, intermediate, and advanced levels, with the vocabulary difficulty increasing accordingly. For a video, different vocabulary grades correspond to the same, different, or not entirely identical words in the video's subtitles. These words are also called highlighted words (highlight_text) and need to be highlighted during the playback of the first video and the display of its subtitles. For example, the vocabulary levels include beginner and intermediate levels. If a Chinese subtitle in a video is "I like to eat apples", the highlighted word in the subtitle for the beginner level is "like". In this case, "like" should be highlighted while the first video is playing and this Chinese subtitle is displayed. In the intermediate level, the highlighted words in the subtitle are "like" and "apple". In this case, "like" and "apple" should be highlighted simultaneously while the first video is playing and this Chinese subtitle is displayed.

[0083] The playback request for the first video is generated and sent by the player. The playback request is used to indicate the first video and the first vocabulary level. The first vocabulary level is the vocabulary level selected by the first account. The first vocabulary level corresponds to the highlighted words to be highlighted in the subtitle content of the first video.

[0084] For example, a computer device obtains a playback request for a first video. In some embodiments, when a first video is a popular video, it may be selected for playback multiple times by multiple user accounts. If multiple playback requests for the first video are obtained simultaneously, and these multiple playback requests indicate the same first vocabulary level, then using the cache penetration concept, a playback request is arbitrarily selected from the multiple playback requests, or the playback request with the earliest time is selected. Based on the selected playback request, the steps of the following embodiments are performed to synthesize a highlighted subtitle file and cache it. For other playback requests, the highlighted subtitle file can be directly obtained from the cache without re-synthesizing.

[0085] Step 220: Based on the playback request, obtain the original subtitle file and highlight metadata. The original subtitle file is used to indicate the subtitle content of the first video, and the highlight metadata is used to describe the highlight information of the highlighted words corresponding to the first vocabulary level.

[0086] The original subtitle file is a subtitle file delivered by the production company that does not embed any subtitle style information. It is used to indicate the subtitle content of the first video, or, as understood, to indicate the subtitle content itself of the first video. Subtitle style information includes at least one of the following: font, color, size, position, shadow, background, effects, and alignment. In some embodiments, the first video has a corresponding original subtitle file, which is used to indicate the subtitle content of the first video.

[0087] Highlight metadata is a data structure stored independently of the original subtitle file. It describes which words in the subtitle content of a video, in a specific language, and at a specific vocabulary level need to be highlighted, what color they should be highlighted, their position in the sentence, and the business information of the highlighted words. In some embodiments, highlight metadata is stored as multi-dimensional information corresponding to three dimensions: video ID (video_id), language, and vocabulary level (grade_level). The index key of the highlight metadata is set as: {video_id, language, grade_level}. Based on the video ID, language, and first vocabulary level of the first video, the corresponding highlight metadata is determined. This highlight metadata describes the highlighting information of the highlighted words corresponding to the first vocabulary level. It should also be noted that the caches of highlight metadata for different vocabulary levels are isolated from each other. If the highlight metadata corresponding to one vocabulary level is updated, only the highlight metadata corresponding to that vocabulary level is invalidated, and the cache is re-triggered for updating. It has no impact on the highlight metadata corresponding to other vocabulary levels, improving cache utilization and ensuring update consistency.

[0088] Taking vocabulary levels as an example: beginner, intermediate, and advanced. Figure 4This is a schematic diagram illustrating the correspondence between the original subtitle file and different highlight metadata provided in an exemplary embodiment of this application. For the first video, the original subtitle file is used to indicate the subtitle content of the first video, without embedding style information, and may include: subtitle sequence number, subtitle start time, and one or more of English / Chinese subtitles. For the same video and the same language, different highlight metadata is stored based on different vocabulary levels, specifically: beginner level highlight metadata, intermediate level highlight metadata, and expert level highlight metadata. These highlight metadata include one or more of the following: highlighted words, color, appearance sequence number, and word identifier. The differences between multiple versions are carried by the highlight metadata. When adding or adjusting vocabulary levels, only the corresponding highlight metadata is updated; the original subtitle file is not copied, nor are multiple video sources regenerated.

[0089] For example, based on a playback request, the computer device obtains the original subtitle file and highlight metadata. The original subtitle file is used to indicate the subtitle content of the first video, and the highlight metadata is used to describe the highlight information of the highlighted words corresponding to the first word level.

[0090] In some embodiments, the highlight metadata includes at least one of the following: single-line subtitle highlight information; highlighted word description information; and complete data composed of multiple single-line subtitle highlight information. The single-line subtitle highlight information includes at least one of the following: subtitle index, corresponding to the original subtitle file number; subtitle start time, in milliseconds (ms); and the set of highlighted words (high_light_texts) in the current subtitle. The highlighted word description information includes at least one of the following: the highlighted word to be highlighted (high_light_text); a highlight label (color) used to indicate color, mapped to a VTT style when composing the highlighted subtitle file; the position of the same highlighted word in the current subtitle (integer_array), used to selectively highlight one or more of the same highlighted word when it appears multiple times in the current subtitle; and a word identifier (text_id) used to indicate the highlighted word, used to extend other resources such as word translation and pronunciation.

[0091] Step 230: Based on the original subtitle file and the highlight metadata, synthesize a highlighted subtitle file. The highlighted subtitle file is used to indicate that at least a portion of the highlighted words corresponding to the first word level are highlighted during the playback of the first video and the display of the subtitle content of the first video.

[0092] The highlighted subtitle file is a subtitle file synthesized based on the original subtitle file of the first video and highlighting metadata. It is used to indicate that at least a portion of the highlighted words corresponding to the first word level should be highlighted during the playback of the first video and the display of its subtitle content. It is understood that the first word level corresponds to multiple highlighted words, and the same highlighted word may appear multiple times in the subtitle content of the first video; for example, it may appear multiple times in the same subtitle line, or it may appear multiple times in the entire subtitle. This same highlighted word that appears multiple times can be highlighted every time it appears, or it can be selectively highlighted. The specific settings can be configured according to actual technical needs. For example, a highlighted word can be set to be highlighted every time a subtitle line appears, or a highlighted word can be set to be highlighted on the first appearance of a subtitle line and ignored on subsequent appearances.

[0093] For example, the computer device synthesizes a highlighted subtitle file based on the original subtitle file and highlight metadata. The highlighted subtitle file is used to indicate that at least a portion of the highlighted words corresponding to a first word level are highlighted during the playback of the first video and the display of the subtitle content of the first video. In some embodiments, both the original subtitle file and the highlighted subtitle file are in VTT format.

[0094] In summary, this application provides a scheme for synthesizing and distributing multi-version highlighted subtitle files. The complete subtitle file is divided into an original subtitle file and different highlighted metadata, which are stored separately. Based on the playback request corresponding to the first video and the first vocabulary level selected by the first account, the original subtitle file and the corresponding highlighted metadata are obtained and synthesized into a highlighted subtitle file. This ensures that at least a portion of the highlighted vocabulary corresponding to the first vocabulary level is highlighted during the playback of the first video and the display of its subtitle content. On one hand, for the same video and the same language, only one original subtitle file is stored and maintained. When a playback request is received, a highlighted subtitle file is synthesized in real time based on different highlighted metadata, enabling the display of multiple versions of highlighted subtitles for the same video and the same language. This avoids generating multiple video sources for different vocabulary levels and also avoids storing multiple complete subtitle files for different vocabulary levels, reducing storage overhead and computing power costs. On the other hand, when adding a new vocabulary level or updating vocabulary, only the corresponding highlighted metadata needs to be changed, without changing the original subtitle file, significantly reducing storage and transcoding costs. Furthermore, it eliminates the need to modify the broadcast control system and player, reducing the scope of modification and the risk of going live.

[0095] In some embodiments, step 230 is implemented as steps 310, 320, and 330.

[0096] Step 310: Parse the original subtitle file to obtain an array of subtitle entries; Step 320: Based on the subtitle entry array and highlight metadata, attach highlight tags to the highlight words in the subtitle entry array that meet the highlight conditions to obtain the processed subtitle entry array; Step 330: Serialize the processed subtitle entry array into a highlighted subtitle file.

[0097] The original subtitle file is in string format and includes at least one of the following: file header, timeline, subtitle number, and subtitle text. The subtitle item array (items) consists of multiple subtitle items arranged in an ordered sequence, consistent with the subtitle playback order. Each subtitle item represents a segment of subtitle content displayed within an independent time period. Each subtitle item in the subtitle item array includes at least one of the following: subtitle number (index), subtitle start time, subtitle end time, and multi-line text. For example, if the first video is a bilingual video including Chinese and English, the multi-line text would include one line of Chinese subtitle and a corresponding line of English subtitle.

[0098] The subtitle entry array is obtained by parsing the strings in the original subtitle file. In some embodiments, the strings in the original subtitle file are preprocessed to obtain a processed original subtitle file. The preprocessing includes at least one of the following: unifying newline characters, filtering blank lines, and removing whitespace characters. The processed original subtitle file is then split according to blank lines to obtain multiple subtitle blocks. For each subtitle block, the strings are parsed line by line from top to bottom to extract subtitle entry information, such as subtitle number, subtitle start time, subtitle end time, or at least one of multiple lines of text. The subtitle entry information is then encapsulated into corresponding subtitle entries to obtain multiple subtitle entries for multiple subtitle blocks. The multiple subtitle entries are then combined into a subtitle entry array according to the order of the multiple subtitle blocks.

[0099] Highlighted words that meet the highlighting criteria refer to at least a portion of the highlighted words that truly need to be highlighted during the playback of the first video and the display of its subtitles. This can also be understood as at least a portion of the highlighted words that are selectively highlighted. Highlight tags are used to indicate highlighted words that meet the highlighting criteria. Highlight tags are used to indicate at least one of the following: color, and optionally, highlighting effect or highlight size.

[0100] A highlighted subtitle file is obtained by serializing a processed array of subtitle entries. For example, a computer device parses the original subtitle file to obtain an array of subtitle entries; based on the array of subtitle entries and highlighting metadata, highlighting tags are attached to the highlighted words in the array of subtitle entries that meet the highlighting conditions, resulting in a processed array of subtitle entries; the processed array of subtitle entries is then serialized into a highlighted subtitle file.

[0101] In this embodiment, by parsing the original subtitle file, a subtitle entry array is obtained, enabling accurate representation of the string content in the original subtitle file. Based on the subtitle entry array and highlight metadata, the highlighted words in the subtitle entry array that meet the highlighting conditions can be accurately identified. By attaching highlight tags, the processed subtitle entry array is obtained, and the processed subtitle entry array is serialized into a highlighted subtitle file. This achieves accurate and selective highlighting of at least some highlighted words, improves the synthesis efficiency of highlighted subtitle files, and avoids storing multiple subtitle files for the same video and the same language, reducing storage overhead.

[0102] In some embodiments, step 320 is implemented as steps 321, 322, and 323.

[0103] Step 321: Align the subtitle entry array with the highlighted metadata to identify the matched entries; Step 322: Locate the highlighted words that meet the highlighting criteria in each hit entry; Step 323: Add highlight tags to the highlighted words that meet the highlighting conditions to obtain the processed subtitle entry array.

[0104] Each subtitle entry in the subtitle entry array includes a subtitle index. The highlight metadata can also be an array, with each metadata entry also carrying a subtitle index to indicate which subtitle the highlighting rule applies to. Alignment is based on subtitle indexes. When the subtitle index of a subtitle entry in the subtitle entry array matches the subtitle index of a metadata record in the highlight metadata, it is considered a match, and this subtitle entry is called a match entry. For example, if the subtitle index of a subtitle entry in the subtitle entry array is 2, and the subtitle index of a metadata record in the highlight metadata is also 2, then it is considered a match, and this subtitle entry is designated as a match entry. In some embodiments, the subtitle entry array and the highlight metadata are aligned according to the subtitle index to determine the match entry.

[0105] Since each subtitle entry in the subtitle entry array includes multiple lines of text, and these multiple lines of text correspond to multiple highlighted words at the first vocabulary level, and these highlighted words are selectively highlighted, it is necessary to locate the highlighted words that meet the highlighting conditions in each hit entry, attach highlighting tags to these highlighted words that meet the highlighting conditions, and obtain the processed subtitle entry array.

[0106] For example, the computer device aligns the subtitle entry array with the highlight metadata to determine the hit entries; locates the highlighted words that meet the highlighting conditions in each hit entry; and attaches highlight tags to the highlighted words that meet the highlighting conditions to obtain the processed subtitle entry array.

[0107] In this embodiment, by aligning the subtitle entry array with the highlight metadata, the hit entries can be accurately determined. For each hit entry, the highlight words that meet the highlight conditions are located. Highlight tags are attached to the highlight words that meet the highlight conditions to obtain the processed subtitle entry array. This achieves the selective highlighting of a highlight word that appears only once in a subtitle, thus achieving precise highlighting.

[0108] In some embodiments, there may be one or more hit entries, and the following steps are performed for each hit entry. Specifically, step 322 is implemented as steps 410 and 420.

[0109] Step 410: For each matched entry containing multiple lines of text, identify the line to be highlighted; Step 420: Locate the highlighted words that meet the highlighting conditions in the highlighted row.

[0110] The line to be highlighted is one or more lines of text within the multi-line text of a matched entry that require highlighted words. This line can also be called the target line. In some embodiments, taking a bilingual video including Chinese and English as an example, and the line to be highlighted as English lines, the letter percentage is calculated line by line for each matched entry's multi-line text. Lines or lines with a letter percentage greater than a threshold are identified as lines to be highlighted. The letter percentage is calculated by dividing the number of letter characters by the total number of characters. If none of the letters in a matched entry's multi-line text are present, it is considered a pure Chinese line or a pure punctuation line, and the matched entry is skipped. This method accurately identifies lines to be highlighted, avoiding unnecessary processing of lines that do not require highlighting and reducing processing overhead.

[0111] In this embodiment, by identifying the line to be highlighted and sequentially locating the highlighted words that meet the highlighting conditions in the line to be highlighted, the highlighted words that meet the highlighting conditions in each line to be highlighted in the hit entry can be accurately determined, thus achieving selective highlighting of a single occurrence of a highlighted word in a subtitle and achieving precise highlighting.

[0112] In some embodiments, the highlighted words that meet the highlighting conditions are located based on their occurrence sequence number in the current sentence subtitle. For each highlighted word in each line to be highlighted, the following steps are performed respectively, and step 420 is implemented as steps 421, 422, 423, 424 and 425.

[0113] Step 421: For the j-th type of highlighted word in the i-th line to be highlighted, initialize the counter and the scanning cursor. The counter is used to record the number of times the j-th type of highlighted word has been scanned in the line to be highlighted. The scanning cursor is used to mark the starting position of searching for the j-th type of highlighted word from the text to be scanned in the line to be highlighted. i and j are positive integers. Step 422: Search for the j-th highlighted word from the starting position of the scanning cursor mark; Step 423: If no j-th type of highlighted word is found, end the processing of the j-th type of highlighted word; Step 424: If the j-th type of highlighted word is found, obtain the current occurrence count of the j-th type of highlighted word recorded by the counter, and perform processing on the j-th type of highlighted word based on the pre-set sequence number array and the current occurrence count; Step 425: Update the counter value, and return to the step of searching for the j-th type of highlighted word from the starting position of the scanning cursor mark until the counter value exceeds the maximum index in the index array. Determine the j-th type of highlighted word that meets the highlighting condition from all the j-th type of highlighted words in the i-th row to be highlighted; where the index array is the set of occurrence indexes of the highlighted words that meet the highlighting condition.

[0114] For the j-th highlighted word in the i-th line to be highlighted, where i and j are positive integers, a counter (foundCount) and a scanning cursor (start) are maintained. The counter records the number of times the j-th highlighted word has been scanned in the line to be highlighted. The initial value of the counter is 0, and this number of occurrences can also be understood as the occurrence sequence number. The scanning cursor marks the starting position for searching for the j-th highlighted word from the text to be scanned in the line to be highlighted. The starting position is initially located at the starting position of the i-th line to be highlighted.

[0115] Search for the j-th type of highlighted word from the starting position of the scanning cursor mark, that is, find the next occurrence position of the j-th type of highlighted word. Depending on whether the j-th type of highlighted word can be found, the following situations exist.

[0116] Scenario 1: If the j-th type of highlighted word is not found, the processing of the j-th type of highlighted word ends, and the next type of highlighted word can be processed, that is, the (j+1)-th type of highlighted word can be processed.

[0117] Scenario 2: If the j-th type of highlighted word is found, the current occurrence count of the j-th type of highlighted word recorded by the counter is obtained. Based on the pre-set sequence number array and the current occurrence count, processing is performed on the j-th type of highlighted word. The specific processing method is as follows in the following embodiment. Here, the sequence number array is the set of occurrence sequence numbers of highlighted words that meet the highlighting conditions. After processing the found j-th type of highlighted word, the counter value is updated, that is, the counter value is incremented by 1. Steps 422 to 424 are repeated until the counter value exceeds the maximum sequence number in the sequence number array. The j-th type of highlighted word that meets the highlighting conditions is then determined from all j-th type highlighted words in the i-th row to be highlighted.

[0118] In some embodiments, step 424 performs processing on the j-th highlighted word based on a pre-defined sequence number array and the current occurrence count, specifically implemented as step 4241 or step 4242.

[0119] Step 4241: If the current occurrence count falls into the sequence number array, determine the j-th type of highlighted word with the current occurrence count as the j-th type of highlighted word that meets the highlighting condition, and determine the remaining text in the i-th line to be highlighted, reset the scanning cursor to the starting position of the remaining text, and return to execute the step of searching for the j-th type of highlighted word from the starting position marked by the scanning cursor. Step 4242: If the current occurrence count does not fall into the sequence number array, determine that the j-th type of highlighted word with the current occurrence count does not meet the highlighting condition, jump the scanning cursor to the end position of the j-th type of highlighted word with the current occurrence count, and return to execute the step of searching for the j-th type of highlighted word from the starting position marked by the scanning cursor.

[0120] The sequence number array is the set of occurrence numbers of highlighted words that meet the highlighting conditions. For example, occurrence numbers 1 and 3 in the sequence number array mean that the highlighted words appearing for the 1st and 3rd time in the current subtitle are highlighted words that meet the highlighting conditions and need to be highlighted. Based on whether the current occurrence number of the j-th type of highlighted word falls into the sequence number array, the following situations exist.

[0121] Case 1: If the current occurrence count of the j-th type of highlighted word falls into the sequence number array, then the j-th type of highlighted word with the current occurrence count is determined as the j-th type of highlighted word that meets the highlighting condition. Also, determine the remaining text in the i-th line to be highlighted, reset the scanning cursor to the starting position of the remaining text, and return to execute the step of searching for the j-th type of highlighted word from the starting position marked by the scanning cursor.

[0122] In some embodiments, for Case 1, the i-th line to be highlighted is divided into three segments: prefix, target word, and remaining text. The prefix retains its original color style, the target word is highlighted, and the remaining text continues to be scanned. Here, the target word refers to the j-th type of highlighted word with the current occurrence count, the prefix is ​​the scanned text in the i-th line before the start position of the j-th type of highlighted word with the current occurrence count, and the remaining text is the text to be scanned in the i-th line after the end position of the j-th type of highlighted word with the current occurrence count. Therefore, step 4241 determines the remaining text in the i-th line by: based on the end position of the j-th type of highlighted word with the current occurrence count, determining the text to be scanned after the end position in the i-th line as the remaining text in the i-th line. This method ensures that subsequent scanning is performed only on the remaining text, thereby guaranteeing that any target word is ultimately wrapped by only one layer of highlighting tags, without nesting conflicts.

[0123] Scenario 2: If the current occurrence count does not fall into the sequence number array, it is determined that the j-th type of highlighted word with the current occurrence count does not meet the highlighting condition. The scanning cursor is then moved to the end of the j-th type of highlighted word with the current occurrence count, and the step of searching for the j-th type of highlighted word from the beginning of the scanning cursor mark is returned. That is, the j-th type of highlighted word with the current occurrence count is skipped, and the scanning cursor is moved to the end of the j-th type of highlighted word with the current occurrence count to continue scanning.

[0124] The above embodiments provide a processing method for each type of highlighted word in each line to be highlighted, which can accurately identify the highlighted words that meet the highlighting conditions in each line to be highlighted, avoid omissions, and achieve selective highlighting of a single occurrence of a highlighted word in a subtitle, thus achieving accurate highlighting and avoiding accidental highlighting.

[0125] Here is an example illustrating the overall processing steps for the j-th type of highlighted word in the i-th row to be highlighted.

[0126] For example, the highlighted word is "apple", and the sequence number array is [1,3]. The original text of the line to be highlighted is: "apple bananaapple orange apple grape apple". The first time "apple" is found, foundCount=1, it falls into the sequence number array, confirming that "apple" meets the highlighting condition and needs to be highlighted; the remaining text continues scanning. The second time "apple" is found, foundCount=2, it does not fall into the sequence number array, so it is skipped. The third time "apple" is found, foundCount=3, it falls into the sequence number array, confirming that "apple" meets the highlighting condition and needs to be highlighted; the remaining text continues scanning. The fourth time "apple" is found, foundCount=4, exceeding the maximum sequence number 3 in the sequence number array; at this point, the loop terminates and no further processing occurs.

[0127] For example, Figure 5 This is a schematic diagram illustrating selective highlighting based on occurrence number, provided in an exemplary embodiment of this application. The English subtitle in the hit entry is: "I Like Blue And Like Books". The highlighted word description is: the second occurrence of "like" in the subtitle is highlighted, with a color of "class A". The counter then sequentially scans the same highlighted word "like", without re-scanning already scanned text. Only the second occurrence of "like" is highlighted, and a highlighted subtitle file is output. In the highlighted subtitle file, this English subtitle is represented as: "I like blue and..."<c.classA> In the "likebooks" file, class A is bound to the color configuration in the style block.

[0128] In some embodiments, a global style declaration is also injected into the original subtitle file. Following step 310, the method further includes steps 312 and 314.

[0129] Step 312: Read the color configuration table from the configuration center; Step 314: Write the color configuration table into the style block of the original subtitle file to generate a global style declaration; the color configuration table is used to indicate the class name and the corresponding color value.

[0130] The configuration center pre-stores multiple color configuration tables, which can be retrieved according to actual technical needs. The color configuration table indicates the class name (className) and its corresponding color value, represented as: {className → CSS color value}. The original subtitle file is in VTT format, which supports the use of style blocks (STYLE), pseudo-elements (::cue), and class name tags (<c.classname> To express the subtitle style, the color configuration table is written to the style block (STYLE), generating a global style declaration in the form of ::cue(.className){color:...}.

[0131] In this embodiment, the color configuration table read from the configuration center is written into the style block of the original subtitle file to generate a global style declaration. This unifies visual specifications, ensures strict consistency in color semantics of subsequently synthesized highlighted subtitle files, and eliminates visual deviations. Furthermore, if a color scheme needs to be modified, only the color configuration table needs to be adjusted, reducing the maintenance cost of the color scheme.

[0132] In some embodiments, when attaching highlight tags to highlight words that meet the highlighting conditions, tag-style binding is also completed based on global style declarations. Therefore, based on the embodiments of steps 312 and 314, step 323 is implemented as step 3230.

[0133] Step 3230: Attach the highlighted words that meet the highlighting conditions to the highlighting tags corresponding to the class names in the global style declaration to obtain the processed subtitle entry array.

[0134] For example, for highlighted words that meet the highlighting criteria, the highlighted words that meet the highlighting criteria will be attached with the highlighting tag corresponding to the class name in the global style declaration.<c.{Color}> The target word, where {Color} is the class name of the declared style, yields an array of processed subtitle entries.

[0135] In this embodiment, the highlighted words that meet the highlighting conditions are attached to the highlighting tags corresponding to the class names in the global style declaration, resulting in a processed array of subtitle entries. This achieves tag-style binding based on the global style declaration, ensuring the correct display of colors.

[0136] Next, we will explain in detail how to play the video.

[0137] Figure 6 This is a flowchart illustrating a video playback method provided in an exemplary embodiment of this application. The method is executed by a computer device, which may be... Figure 1 The first terminal 120. The method includes at least some of the steps 610, 620, and 630.

[0138] Step 610: In response to the selection operation, determine the first video and the first vocabulary level.

[0139] The first video is the video selected by the first account and intended to be played in the player. The first account is the user account used by the first user to log in. The type of the first video varies depending on the application scenario. For example, in the education field, the first video could be a bilingual video (including Chinese and English) for English learning. In the entertainment field, the first video could be an entertainment video or a short video.

[0140] The first video also includes subtitles. The subtitles in the first video are used to assist the first user watching the video in understanding the video's visual and audio content. In some embodiments, the subtitles include at least one of the following: dialogue subtitles, explanatory subtitles, title subtitles, interactive subtitles, and decorative subtitles. Dialogue subtitles are used to present at least one of task dialogue, narration, or monologue content, typically one or more lines of text, and are located at the bottom of the screen. For example, if the first video is a bilingual video including Chinese and English, a single dialogue subtitle includes one line of Chinese subtitles and one line of English subtitles. Explanatory subtitles are used to represent at least one of time, location, character introductions, and scene descriptions. Title subtitles are usually located at the beginning or end of the video and include at least one of the following: title, production team list, cast and crew, and production company. Interactive subtitles include at least one of the following: question subtitles, voting subtitles, and bullet comments. Decorative subtitles include at least one of the following: artistic fonts and dynamic effects.

[0141] Vocabulary grade (grade_level), also known as vocabulary difficulty level, is a pre-defined level that can be set according to actual technical needs. Different vocabulary grades correspond to different levels of vocabulary difficulty, which can be represented by at least one of the following: word frequency, word length, word semantics, word structural complexity, and word pronunciation complexity. For example, vocabulary grades may include at least one of beginner, intermediate, and advanced levels, with the vocabulary difficulty increasing accordingly. For a video, different vocabulary grades correspond to the same, different, or not entirely identical words in the video's subtitles. These words are also called highlighted words (highlight_text) and need to be highlighted during the playback of the first video and the display of its subtitles. For example, the vocabulary levels include beginner and intermediate levels. If a Chinese subtitle in a video is "I like to eat apples", the highlighted word in the subtitle for the beginner level is "like". In this case, "like" should be highlighted while the first video is playing and this Chinese subtitle is displayed. In the intermediate level, the highlighted words in the subtitle are "like" and "apple". In this case, "like" and "apple" should be highlighted simultaneously while the first video is playing and this Chinese subtitle is displayed.

[0142] In some embodiments, the computer device displays a player interface, which includes a video display area, a subtitle display area, and a vocabulary level selection area. The video display area displays a first video to be played, the subtitle display area displays the subtitle content of the first video, and the vocabulary level selection area displays candidate vocabulary levels. Each candidate level also displays the number of words to be highlighted, for example, the candidate levels include: all vocabulary (28 words), beginner (5 words), intermediate (17 words), and expert (6 words). A first account performs a selection operation in the player interface, selecting the first video to be played and the first vocabulary level. Exemplarily, in response to the selection operation, the computer device determines the first video and the first vocabulary level.

[0143] Step 620: Send a playback request for the first video. The playback request is used to indicate the first video and the first vocabulary level. The first vocabulary level corresponds to the highlighted words to be highlighted in the subtitle content of the first video.

[0144] The playback request for the first video is generated and sent by the player. For example, the computer device sends a playback request for the first video, which indicates the first video and a first vocabulary level. The first vocabulary level is the vocabulary level selected by the first account, and the first vocabulary level corresponds to the highlighted words to be highlighted in the subtitle content of the first video.

[0145] Step 630: Receive the returned highlighted subtitle file, and, based on the highlighted subtitle file, highlight at least a portion of the highlighted words corresponding to the first word level during the playback of the first video and the display of the subtitle content of the first video; wherein, the highlighted subtitle file is synthesized based on the original subtitle file obtained from the playback request and the highlighted metadata, the original subtitle file is used to indicate the subtitle content of the first video, and the highlighted metadata is used to describe the highlighted information of the highlighted words corresponding to the first word level.

[0146] The highlighted subtitle file is synthesized based on the original subtitle file obtained from the playback request and highlighting metadata. Specifically, the highlighted subtitle file is a subtitle file synthesized from the original subtitle file of the first video and highlighting metadata. It is used to indicate that during the playback of the first video and the display of its subtitle content, at least a portion of the highlighted words corresponding to the first word level should be highlighted. It is understood that the first word level corresponds to multiple highlighted words. In the subtitle content of the first video, the same highlighted word may appear multiple times, for example, multiple times in the same subtitle line, or multiple times in the entire subtitle. This same highlighted word that appears multiple times can be highlighted every time it appears, or it can be selectively highlighted. The specific settings can be configured according to actual technical needs. For example, a highlighted word can be set to be highlighted every time a subtitle line appears, or a highlighted word can be set to be highlighted on the first appearance of a subtitle line and ignored on subsequent appearances.

[0147] For example, the computer device receives the returned highlighted subtitle file, and based on the highlighted subtitle file, during the playback of the first video and the display of the subtitle content of the first video, highlights at least a portion of the highlighted words corresponding to the first word level. In some embodiments, both the original subtitle file and the highlighted subtitle file are in VTT format. The highlighted subtitle file is returned in the form of a highlighted subtitle CDN address, which the computer device can access to download the highlighted subtitle file.

[0148] In summary, this application provides a scheme for synthesizing and distributing multi-version highlighted subtitle files. The complete subtitle file is divided into an original subtitle file and different highlighted metadata, which are stored separately. Based on the playback request corresponding to the first video and the first vocabulary level selected by the first account, a highlighted subtitle file synthesized based on the original subtitle file and the corresponding highlighted metadata is obtained. This allows at least a portion of the highlighted vocabulary corresponding to the first vocabulary level to be highlighted during the playback of the first video and the display of its subtitle content. On one hand, for the same video and the same language, only one original subtitle file is stored and maintained. When a playback request is received, a highlighted subtitle file is synthesized in real time based on different highlighted metadata, enabling the display of multiple versions of highlighted subtitles for the same video and the same language. This avoids generating multiple video sources for different vocabulary levels and also avoids storing multiple complete subtitle files for different vocabulary levels, reducing storage overhead and computing power costs. On the other hand, when adding a new vocabulary level or updating vocabulary, only the corresponding highlighted metadata needs to be changed, without changing the original subtitle file, significantly reducing storage and transcoding costs. Furthermore, it eliminates the need to modify the broadcast control system and player, reducing the scope of modification and the risk of going live.

[0149] It should be noted that the specific limitations of the video playback method embodiments, such as the specific steps for synthesizing the highlighted subtitle file, can be found in the specific limitations of the subtitle file synthesis method embodiments above, and will not be repeated here.

[0150] Next, the methods for synthesizing subtitle files and playing videos will be explained in detail with reference to the diagrams.

[0151] 1. Product realization.

[0152] The method described in this application is applicable to English learning scenarios on children's channels. Users can select bilingual videos (including Chinese and English) and vocabulary levels (beginner, intermediate, and advanced) on the video playback platform. The system retrieves the corresponding highlighted subtitle files based on the selected vocabulary level. When playing the bilingual video, the system automatically locates the English subtitles and highlights words matching the selected vocabulary level and appearing in a specified order with a specific color.

[0153] In some embodiments, this technology is also applicable to documentary playback scenarios. Documentaries often feature numerous place names, biological terms, and other specialized technical terms that appear repeatedly. These terms can be set to be highlighted the first time they appear in a given sentence's subtitle, allowing viewers to easily locate them. Furthermore, it is also suitable for real-time subtitle scenarios in conferences. Pre-defined business terms can be displayed in real-time during meetings, with the corresponding subtitles highlighted the first time these terms appear in a given sentence, facilitating quick information retrieval for attendees.

[0154] 2. Technical implementation.

[0155] The method in this application includes two links: an offline generation link for highlight metadata and a runtime distribution link for highlight subtitle files. The offline generation link for highlight metadata is used to generate highlight metadata offline, while the runtime distribution link for highlight subtitle files is used to synthesize and distribute highlight subtitle files at runtime.

[0156] Figure 7 This is a schematic diagram of an offline generation chain for highlighted metadata provided in an exemplary embodiment of this application. The original subtitle file is the subtitle file delivered by the film production company, containing only the timeline and subtitle content. The graded vocabulary library is a vocabulary set maintained according to learning age and grade level. For the original subtitle file and the graded vocabulary library, subtitle parsing, vocabulary matching, occurrence position recording, and highlighted metadata generation are performed based on the offline subtitle processing service. The operation management backend can review individual subtitles and adjust whether a certain subtitle, a certain word, or a certain occurrence frequency is highlighted. Highlighted metadata is stored according to multi-dimensional information corresponding to three dimensions: video identifier, language, and vocabulary level. When the original subtitle file is updated, the highlighted metadata can also be regenerated to keep the subtitle content text consistent with the highlighted information, avoiding the situation where the subtitle content changes but the highlighted metadata remains in the old version.

[0157] Figure 8 This is a schematic diagram of the runtime distribution chain of a highlighted subtitle file provided in an exemplary embodiment of this application. The player determines the vocabulary level, requests to play the video and obtain the highlighted subtitle file; the broadcast control system identifies the video scene and transmits the video identifier, language, and vocabulary level parameters to the subtitle proxy service; the subtitle proxy service generates a highlighted subtitle CDN address without reading the subtitle content; the highlighted subtitle CDN receives the playback request, and if the playback request hits the cache, it directly returns the highlighted subtitle file; if the playback request does not hit the cache, it obtains the original subtitle file and highlighted metadata, synthesizes the highlighted subtitle file, and stores the highlighted metadata according to the video identifier, language, and vocabulary level; the highlighted subtitle CDN returns the highlighted subtitle file to the player, which selects and displays it according to its existing capabilities.

[0158] Figure 9This is a schematic diagram of precise multi-dimensional cache invalidation provided by an exemplary embodiment of this application. Since the highlighted metadata is stored according to video identifier, language, and vocabulary level, the index key of the highlighted metadata is set as: {video_id, language, grade_level}. An update to the highlighted metadata of a certain vocabulary level only generates a change event for the corresponding index key. The change event is represented as: key=video_id; language=english; grade_level=advanced level. Multiple edge nodes are notified in the message queue to delete the old cache according to the index key. If the cache of a node in the highlighted subtitle CDN includes: cache item A: beginner level, cache item B: intermediate level, and cache item C: expert level, then only the hit cache item B is deleted. This method ensures that other vocabulary level caches continue to be hit, preventing them from being completely cleared. Popular videos reuse the synthesized results at edge nodes, balancing low latency, low back-to-origin pressure, and update consistency.

[0159] 2.1 Overall Architecture.

[0160] Figure 10 This is a schematic diagram of the overall architecture of the offline generation link of highlighted metadata and the runtime distribution link of highlighted subtitle files provided in an exemplary embodiment of this application.

[0161] For the offline generation of highlighted metadata: Based on a hierarchical vocabulary library and the original subtitle files, highlighted metadata is generated and stored according to video identifier, language, and vocabulary level. After the system automatically generates the highlighted metadata, operators can manually modify and edit it in real time on the subtitle editing page during the review process. Personalized adjustments can be made to the highlighted subscripts at the word level, supporting extremely fine-grained control such as "not highlighting a word that appears in a specific subtitle in a certain video," seamlessly synchronized to end users. Furthermore, if the original subtitle files are updated, the highlighted metadata will be regenerated synchronously, ensuring atomic consistency between the original subtitle files and the highlighted metadata.

[0162] For the runtime distribution chain of highlighted subtitle files: After the player initiates a playback request, the broadcast control system generates a highlighted subtitle CDN address. The player uses the proxy address to access the highlighted subtitle CDN. The highlighted subtitle CDN reads the original subtitle file and highlighted metadata, synthesizes the highlighted subtitle file in real time, and returns it to the player.

[0163] Specifically, the player initiates playback requests, downloads highlighted subtitle files, renders and displays subtitles according to standard VTT style rules. The broadcast control system processes playback requests and, upon recognizing video scene parameters, passes the video identifier, language, and vocabulary level to the subtitle proxy service to obtain the highlighted subtitle CDN address. The broadcast control system does not need to directly support multi-version subtitle scheduling in the same language. The subtitle proxy service generates a subtitle access address pointing to the highlighted subtitle CDN based on the video identifier, language, and vocabulary level. The subtitle proxy service only generates routing addresses and does not read the original subtitle file, highlight metadata, or perform highlighted subtitle file synthesis. The highlighted subtitle CDN receives playback requests, reads the original subtitle file and highlight metadata, and performs highlighted subtitle file synthesis at edge nodes, returning standard VTT subtitles with highlight tags. The original subtitle storage stores the original subtitle files without embedded highlight information. The highlight metadata storage stores the highlight metadata corresponding to different videos, languages, and vocabulary levels, with the index key: {video_id,language,grade_level}. The subtitle processing service is used to parse raw subtitle files offline and generate highlighted metadata using a hierarchical vocabulary database, while also supporting operations staff in editing the highlighted metadata. The hierarchical vocabulary database is used to maintain a set of target words according to their vocabulary level.

[0164] 2.2 Separate storage of original subtitle files and highlight metadata.

[0165] Figure 11 This is a schematic diagram illustrating the separate storage of the original subtitle file and highlight metadata, provided in an exemplary embodiment of this application. Specifically, the subtitle file is divided into two categories of data: "subtitle content body" and "highlight information description". The original subtitle file only stores basic information such as subtitle content, timeline, and subtitle sequence number; the highlight metadata only describes the words that need to be highlighted at a certain vocabulary level, as well as their occurrence position, color, and vocabulary identifier.

[0166] 2.3 Highlighting metadata structure.

[0167] Figure 12 This is a schematic diagram of the highlight metadata structure provided in an exemplary embodiment of this application. A complete highlight metadata set represents highlight information for a video, a language, and a vocabulary level. Its internal structure is a list, containing the following fields and their meanings.

[0168] (1) Complete data: composed of multiple single-sentence subtitle highlight information; (2) Single-sentence subtitle highlight information: subtitle number (index), corresponding to the original subtitle file number; subtitle start time (time), in milliseconds; set of highlighted words in the current subtitle (high_light_texts); (3) Highlighted word description information: the highlighted word to be highlighted (text); the highlight label (color) used to indicate the color, which is mapped to VTT style during composition; the position of the same highlighted word in the subtitle of the sentence (integer_array), used for selective highlighting; (4) A word identifier (text_id) used to indicate highlighted words, which is used to expand other resources such as word translation and pronunciation.

[0169] Figure 13 This is a schematic diagram illustrating the extraction of highlight metadata information according to an exemplary embodiment of this application. Highlight metadata is extracted based on the original subtitle file. Figure 14 This is a schematic diagram illustrating selective highlighting of highlighted words according to an exemplary embodiment of this application. The integer_array parameter controls the selective highlighting of at least a portion of the highlighted words; for example, in a subtitle, only the second identical word is highlighted, or all identical words in a subtitle are highlighted.

[0170] 2.4 Real-time synthesis of highlighted subtitle files.

[0171] Continue to refer to Figure 10 The broadcast control system sends a playback request, requesting the subtitle proxy service to obtain the highlighted subtitle CDN address. The highlighted subtitle CDN address includes video identifier, language, and vocabulary level parameters. The player uses the highlighted subtitle CDN address to obtain the highlighted subtitle file. The highlighted subtitle CDN requires the following steps.

[0172] 1. Parameter parsing: Parses the video identifier, language, and vocabulary level in the playback request; 2. Query cache; 3. Check for cache hits; if a cache hit occurs, proceed to step 4; otherwise, proceed to steps 5 and 6. 4. Return the matched highlighted subtitle file directly to the player; 5. Obtain the original subtitle file; 6. Retrieve highlighted metadata; 7. Inject VTT style tags to define color styles; 8. Align according to the subtitle numbering; 9. Insert a highlight label; 10. Assemble the highlighted subtitle file; 11. Return the highlighted subtitle file to the player.

[0173] It should also be noted that if the cache is not hit and multiple user accounts request the same vocabulary level at the same time, the cache penetration concept is used to select one playback request, execute the generation and caching of the highlighted subtitle file. After the highlighted subtitle file is cached, other playback requests will continue to retrieve the highlighted subtitle file from the cache.

[0174] The detailed logic steps for synthesizing highlighted subtitle files are briefly described below.

[0175] 1. Parse the original subtitle file; The VTT string in the original subtitle file is parsed into an array of subtitle items. Each subtitle item includes: subtitle number (index), subtitle start time, subtitle end time, and multiple lines of text (lines). Bilingual videos typically contain one line of English text and one line of Chinese text.

[0176] 2. Inject global style headers; Read the color configuration table {className→CSS color value} from the configuration center, write it into the style block (STYLE) of the original subtitle file, and generate a global style declaration in the form of ::cue(.className){color:...}.

[0177] 3. Align subtitle entries with highlighted metadata according to subtitle number; Scan the subtitle entry array and highlighted metadata sequentially using two pointers, aligning them according to the subtitle entries to identify the hit entries. For each hit entry, perform the following steps and move the highlighted metadata pointer forward.

[0178] 4. Identify the row to be highlighted; For a multi-line text of a matching entry, calculate the letter percentage line by line. The letter percentage is the number of letter characters divided by the total number of characters. Select the line or lines with the highest percentage as the lines to be highlighted. If all lines contain no letters, such as lines consisting only of Chinese characters or punctuation, skip the entry and return the original text.

[0179] 5. Locate the target word by its appearance number; Perform the following steps sequentially for each highlighted word in each line to be highlighted.

[0180] 5.1 Maintain the counter (foundCount) and the scanning cursor (start); the counter is used to record the number of times the highlighted word has been scanned in the line to be highlighted, with an initial value of 0; the scanning cursor is used to mark the starting position for finding the highlighted word from the text to be scanned in the line to be highlighted. 5.2 From the starting position of the line mark to be highlighted, search backwards for the next occurrence of the highlighted word; if not found, end the processing of the highlighted word; if found, proceed to 5.3; 5.3 Determine whether the current occurrence count of the highlighted word falls within the sequence number set (IntegerArray); if it does, divide the line to be highlighted into three segments: prefix (preserving the original style), target word (with the highlight tag attached), and remaining text; use the remaining text as the new text to be scanned, reset the scanning cursor to the beginning position of the remaining text, and continue scanning; if it does not hit, skip the occurrence, and jump the scanning cursor to the end of the highlighted word to continue scanning; 5.4 The counter increments, and steps 5.2 to 5.4 are repeated until the largest index in the set of indexes has been processed; 6. Generate highlighted tags; For the target words matched in 5.3, attach highlight tags.<c.{Color}> The target word, where {Color} is the class name of the style declared in step 2, thus completing the tag-style binding; 7. Avoid nesting and duplicate wrapping; In step 5.3, after a target word is hit, the prefix + target word is immediately physically removed from the text to be scanned, and subsequent scanning is performed only on the remaining text. Multiple highlighted words within the same subtitle entry are processed sequentially, with each step only cutting unlabeled plain text segments, while labeled segments are passed through as independent units, thus ensuring that any target word is ultimately only detected by one layer. <c>Wrapping tags avoids nesting conflicts; 8. Serialize and output the highlighted subtitle file; The processed subtitle entry array is reserialized into a VTT string, and the highlighted subtitle file is returned.

[0181] 3. Beneficial effects.

[0182] The method described in this application has at least the following beneficial effects.

[0183] 3.1 Reduce storage overhead and computing power costs.

[0184] The method in this application avoids generating multiple video sources for different vocabulary levels and storing multiple complete subtitle files for different vocabulary levels. Instead, it uses a single original subtitle file associated with different highlight metadata. When adding a new vocabulary level or updating vocabulary, only the corresponding highlight metadata needs to be changed, significantly reducing storage and transcoding costs. Specifically, if a multi-video hard compression scheme is used, assuming a language has 3 vocabulary levels and a video has 4 resolutions, each resolution has basic encodings like H.264 and H.265, resulting in 8 video files, 24 video files need to be transcoded to achieve subtitle highlighting. Updating a vocabulary requires re-transcoding all 24 video files. If a multi-subtitle file storage scheme is used, assuming a language has 3 vocabulary levels, 3 soft subtitle files need to be stored. Updating a vocabulary requires regenerating and re-storing 3 soft subtitle files, rendering the old subtitle files completely invalid, resulting in high storage and transcoding overhead.

[0185] 3.2. Zero modification to the single subtitle storage model for the same language.

[0186] The method in this application avoids the original subtitle system supporting multiple subtitle files of the same language. The original subtitles are still stored according to the existing model, and the highlight differences are carried by independent highlight metadata.

[0187] 3.3 Reduce the risks associated with upgrading the broadcast control system.

[0188] The method in this application generates a highlighted subtitle CDN address through a subtitle proxy service. The broadcast control system only needs to transmit video scene parameters and does not need to directly schedule multiple versions of subtitles or change the core subtitle distribution logic.

[0189] 3.4. Zero-service transformation of the player.

[0190] In the method of this application embodiment, the player ultimately downloads a standard VTT subtitle file. The synthesis of the highlighted subtitle file is completed in the highlighted subtitle CDN. The player only needs to display the subtitles according to its existing VTT rendering capabilities, without needing to understand the hierarchical vocabulary, highlighted metadata, or highlighted subtitle file synthesis algorithm.

[0191] 3.5. Supports precise highlighting in scenarios involving repeated words.

[0192] The method in this application embodiment records the position of the target word in the current subtitle sentence by indexing the subtitle, which can selectively highlight the same word that appears once or multiple times in the same subtitle sentence, avoiding false highlighting caused by global replacement.

[0193] 3.6. Supports caching and precise expiration by vocabulary level.

[0194] The method in this application embodiment caches and synthesizes results based on video identifier, language, and vocabulary level. Caches at different vocabulary levels are isolated from each other; when the highlighted metadata for a certain vocabulary level is updated, all nodes of the highlighted subtitle CDN are notified through a message queue, invalidating only the cache for the corresponding vocabulary level and re-triggering the cache update, thereby improving cache utilization and ensuring update consistency.

[0195] 3.7 Supports subsequent expansion of learning functions.

[0196] The highlighted metadata in this application includes fields such as vocabulary identifiers. This parameter can be seamlessly extended to external vocabulary resource interfaces to support learning functions such as vocabulary definitions, pronunciation audio, example sentences, and follow-along reading. Expanding new functions does not require modification of subtitle content or the player.

[0197] 3.8 Supports multiple application scenarios.

[0198] The highlighted metadata storage in this application includes a language dimension, making the method of this application fully applicable to other foreign language learning videos, vocabulary practice education platforms, professional terminology highlighting, and other related fields, with a wide range of application scenarios.

[0199] Figure 15 This is a block diagram of a subtitle file synthesis apparatus provided in an exemplary embodiment of this application. The subtitle file synthesis apparatus 800 includes at least some of the following modules: an acquisition module 810, a processing module 820, and a synthesis module 830.

[0200] The acquisition module 810 is used to acquire a playback request for a first video. The playback request is used to indicate the first video and a first vocabulary level. The first vocabulary level corresponds to a highlighted vocabulary to be highlighted in the subtitle content of the first video. The processing module 820 is used to obtain the original subtitle file and highlight metadata based on the playback request. The original subtitle file is used to indicate the subtitle content of the first video, and the highlight metadata is used to describe the highlight information of the highlighted words corresponding to the first word level. The synthesis module 830 is used to synthesize a highlighted subtitle file based on the original subtitle file and the highlighted metadata. The highlighted subtitle file is used to indicate that at least a portion of the highlighted words corresponding to the first word level are highlighted during the playback of the first video and the display of the subtitle content of the first video.

[0201] In some embodiments, the synthesis module 830 is configured to: The original subtitle file is parsed to obtain an array of subtitle entries; Based on the subtitle entry array and the highlight metadata, highlight tags are attached to the highlight words in the subtitle entry array that meet the highlight conditions, resulting in a processed subtitle entry array; The processed subtitle entry array is serialized into the highlighted subtitle file.

[0202] In some embodiments, the synthesis module 830 is configured to: Align the subtitle entry array with the highlighted metadata to determine the matched entries; Locate the highlighted words in each of the hit entries that satisfy the highlighted conditions; The highlighted words that meet the highlighted conditions are attached with the highlighted tags to obtain the processed subtitle entry array.

[0203] In some embodiments, the synthesis module 830 is configured to: For each hit entry containing multiple lines of text, identify the line to be highlighted; The highlighted words that meet the highlighted conditions are located sequentially in the row to be highlighted.

[0204] In some embodiments, the synthesis module 830 is configured to: For the j-th highlighted word in the i-th line to be highlighted, initialize a counter and a scanning cursor. The counter is used to record the number of times the j-th highlighted word has been scanned in the line to be highlighted. The scanning cursor is used to mark the starting position for finding the j-th highlighted word from the text to be scanned in the line to be highlighted. i and j are positive integers. Search for the j-th highlighted word from the starting position of the scanning cursor mark; If the j-th type of highlighted word is not found, the processing of the j-th type of highlighted word ends; If the j-th type of highlighted word is found, obtain the current occurrence count of the j-th type of highlighted word recorded by the counter, and perform processing on the j-th type of highlighted word based on the pre-set sequence number array and the current occurrence count; Update the value of the counter, and return to the step of searching for the j-th type of highlighted word from the starting position of the scanning cursor mark until the value of the counter exceeds the maximum index in the index array, and determine the j-th type of highlighted word that meets the highlighting condition from all the j-th type of highlighted words in the i-th row to be highlighted; The sequence number array is a set of occurrence numbers of highlighted words that meet the highlighting conditions.

[0205] In some embodiments, the synthesis module 830 is configured to: If the current occurrence count falls into the sequence number array, the j-th type of highlighted word with the current occurrence count is determined as the j-th type of highlighted word that satisfies the highlighting condition. Also, the remaining text in the i-th line to be highlighted is determined, the scanning cursor is reset to the starting position of the remaining text, and the process of searching for the j-th type of highlighted word from the starting position marked by the scanning cursor is returned. If the current occurrence count does not fall into the sequence number array, it is determined that the j-th type of highlighted word with the current occurrence count does not meet the highlighting condition. The scanning cursor is then moved to the end of the j-th type of highlighted word with the current occurrence count, and the process of searching for the j-th type of highlighted word from the beginning of the scanning cursor mark is returned.

[0206] In some embodiments, the synthesis module 830 is configured to: Based on the end position of the j-th highlighted word with the current occurrence frequency, the text to be scanned in the i-th line to be highlighted that is located after the end position is determined as the remaining text in the i-th line to be highlighted.

[0207] In some embodiments, the synthesis module 830 is configured to: For each of the multi-line text entries that are hit, the proportion of letters is calculated line by line. The line or lines of text in which the proportion of the letters is greater than the threshold are identified as the lines to be highlighted.

[0208] In some embodiments, the synthesis module 830 is further configured to: Read the color configuration table from the configuration center; Write the color configuration table into the style block of the original subtitle file to generate a global style declaration; The color configuration table is used to indicate the class name and the corresponding color value.

[0209] In some embodiments, the synthesis module 830 is configured to: The highlighted words that meet the highlighted conditions are attached with the highlighted tags corresponding to the class names in the global style declaration to obtain the processed subtitle entry array.

[0210] In some embodiments, the highlighted metadata includes at least one of the following: Single-sentence subtitle highlighting information; Highlighted words and their descriptions; Complete data consisting of multiple highlighted single-sentence subtitles.

[0211] In some embodiments, the single-sentence caption highlighting information includes at least one of the following: Subtitle number; Subtitle start time; The set of highlighted words in the caption.

[0212] In some embodiments, the highlighted word description information includes at least one of the following: Highlighted words to be highlighted; Highlighted labels used to indicate colors; The same highlighted words appear in the same subtitle; Vocabulary identifiers used to indicate highlighted words.

[0213] Figure 16 This is a block diagram of a video playback device provided in an exemplary embodiment of this application. The video playback device 900 includes at least some of the following modules: a determining module 910, a sending module 920, and a highlighting module 930.

[0214] Module 910 is used to determine the first video and the first vocabulary level in response to the selection operation; The sending module 920 is used to send a playback request for the first video. The playback request is used to indicate the first video and the first vocabulary level. The first vocabulary level corresponds to the highlighted vocabulary to be highlighted in the subtitle content of the first video. The highlighting module 930 is used to receive the returned highlighted subtitle file, and, based on the highlighted subtitle file, highlight at least a portion of the highlighted words corresponding to the first word level during the process of playing the first video and displaying the subtitle content of the first video; The highlighted subtitle file is synthesized based on the original subtitle file obtained from the playback request and the highlighted metadata. The original subtitle file is used to indicate the subtitle content of the first video, and the highlighted metadata is used to describe the highlighted information of the highlighted words corresponding to the first vocabulary level.

[0215] It should be noted that the specific limitations of the one or more subtitle file synthesis devices 800 and video playback devices 900 provided above can be found in the limitations of the subtitle file synthesis method and video playback method above, and will not be repeated here. Each module of the above-mentioned device can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in the processor of the computer device in hardware form or independent of the processor, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0216] This application also provides a computer device, which includes: a processor and a memory, wherein the memory stores a computer program; the processor is used to execute the computer program in the memory to implement the subtitle file synthesis method and / or video playback method provided in the above method embodiments.

[0217] Figure 17 This is a structural block diagram of a computer device provided in an exemplary embodiment of this application.

[0218] The computer device 1000 can be a terminal, such as a smartphone, tablet, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), unmanned reservation terminal, smart home appliance, smart voice interaction device, or unmanned vending terminal. The computer device 1000 may also be referred to as user equipment, portable terminal, portable mobile terminal, or other names.

[0219] Typically, computer device 1000 includes a processor 1001 and a memory 1002.

[0220] Processor 1001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0221] The memory 1002 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1002 are used to store at least one instruction, which is executed by the processor 1001 to implement the subtitle file synthesis method and / or video playback method provided in the embodiments of this application.

[0222] In some embodiments, the computer device 1000 may also optionally include: a peripheral device interface 1003 and at least one peripheral device. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1004, a touch display screen 1005, a camera assembly 1006, an audio circuit 1007, and a power supply 1008.

[0223] Peripheral device interface 1003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1001 and memory 1002. In some embodiments, processor 1001, memory 1002 and peripheral device interface 1003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1001, memory 1002 and peripheral device interface 1003 can be implemented on separate chips or circuit boards, and this application embodiment does not limit this.

[0224] The radio frequency (RF) circuit 1004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1004 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application embodiment.

[0225] The touch display screen 1005 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. The touch display screen 1005 also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to the processor 1001 for processing. The touch display screen 1005 is used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one touch display screen 1005, located on the front panel of the computer device 1000; in other embodiments, there may be at least two touch display screens, respectively located on different surfaces of the computer device 1000 or in a folded design; in some embodiments, the touch display screen 1005 may be a flexible display screen, located on a curved or folded surface of the computer device 1000. Furthermore, the touch display screen 1005 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The touch display screen 1005 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0226] The camera assembly 1006 is used to capture images or videos. Optionally, the camera assembly 1006 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is used for video calls or selfies, and the rear-facing camera is used for taking photos or videos. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, and a wide-angle camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, and panoramic shooting and VR (Virtual Reality) shooting by fusion of the main camera and the wide-angle camera. In some embodiments, the camera assembly 1006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash is a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0227] Audio circuit 1007 provides an audio interface between the user and computer device 1000. Audio circuit 1007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to processor 1001 for processing, or input to radio frequency circuit 1004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located at a different location within computer device 1000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker converts electrical signals from processor 1001 or radio frequency circuit 1004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, audio circuit 1007 may also include a headphone jack.

[0228] Power supply 1008 is used to supply power to the various components in computer device 1000. Power supply 1008 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 1008 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0229] In some embodiments, the computer device 1000 further includes one or more sensors 1009. The one or more sensors 1009 include, but are not limited to, an accelerometer 1010, a gyroscope 1011, a pressure sensor 1012, an optical sensor 1013, and a proximity sensor 1014.

[0230] Accelerometer 1010 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by computer device 1000. For example, accelerometer 1010 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1001 can control touchscreen 1005 to display the user interface in landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1010. Accelerometer 1010 can also be used for games or for acquiring user motion data.

[0231] The gyroscope sensor 1011 can detect the orientation and rotation angle of the computer device 1000. The gyroscope sensor 1011 can work in conjunction with the accelerometer sensor 1010 to acquire the user's 3D movements on the computer device 1000. Based on the data acquired by the gyroscope sensor 1011, the processor 1001 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.

[0232] The pressure sensor 1012 can be disposed on the side bezel of the computer device 1000 and / or on the lower layer of the touch display screen 1005. When the pressure sensor 1012 is disposed on the side bezel of the computer device 1000, it can detect the user's grip signal on the computer device 1000 and perform left / right hand recognition or quick operation based on the grip signal. When the pressure sensor 1012 is disposed on the lower layer of the touch display screen 1005, it can control operable controls on the UI interface based on the user's pressure operation on the touch display screen 1005. Operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0233] An optical sensor 1013 is used to collect ambient light intensity. In one embodiment, the processor 1001 can control the display brightness of the touch screen 1005 based on the ambient light intensity collected by the optical sensor 1013. Specifically, when the ambient light intensity is high, the display brightness of the touch screen 1005 is increased; when the ambient light intensity is low, the display brightness of the touch screen 1005 is decreased. In another embodiment, the processor 1001 can also dynamically adjust the shooting parameters of the camera assembly 1006 based on the ambient light intensity collected by the optical sensor 1013.

[0234] The proximity sensor 1014, also known as a distance sensor, is typically located on the front of the computer device 1000. The proximity sensor 1014 is used to detect the distance between the user and the front of the computer device 1000. In one embodiment, when the proximity sensor 1014 detects that the distance between the user and the front of the computer device 1000 is gradually decreasing, the processor 1001 controls the touchscreen display 1005 to switch from a screen-on state to a screen-off state; when the proximity sensor 1014 detects that the distance between the user and the front of the computer device 1000 is gradually increasing, the processor 1001 controls the touchscreen display 1005 to switch from a screen-off state to a screen-on state.

[0235] Those skilled in the art will understand that Figure 17 The structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0236] In an exemplary embodiment, this application also provides a chip, which includes programmable logic circuits and / or computer instructions. When the chip is run on a computer device, it is used to implement the subtitle file synthesis method and / or video playback method provided in the above method embodiments.

[0237] This application also provides a computer-readable storage medium storing a computer program, which is loaded and executed by a processor to implement the subtitle file synthesis method and / or video playback method provided in the above method embodiments.

[0238] This application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the processor of the computer device to load and execute the subtitle file synthesis method and / or video playback method provided in the above-described method embodiments.

[0239] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0240] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0241] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0242] The above description is only an optional embodiment of this application and is not intended to limit this application. Each embodiment can be implemented individually or in combination. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.< / c>

Claims

1. A method for synthesizing subtitle files, characterized in that, The method includes: Obtain a playback request for a first video, the playback request being used to indicate the first video and a first vocabulary level, the first vocabulary level corresponding to a highlighted vocabulary to be highlighted in the subtitle content of the first video; Based on the playback request, the original subtitle file and highlight metadata are obtained. The original subtitle file is used to indicate the subtitle content of the first video, and the highlight metadata is used to describe the highlight information of the highlighted words corresponding to the first word level. The original subtitle file is parsed to obtain an array of subtitle entries; Based on the subtitle entry array and the highlight metadata, highlight tags are attached to the highlight words in the subtitle entry array that meet the highlight conditions, resulting in a processed subtitle entry array; The processed subtitle entry array is serialized into a highlighted subtitle file, which is used to indicate that at least a portion of the highlighted words corresponding to the first word level are highlighted during the playback of the first video and the display of the subtitle content of the first video.

2. The method according to claim 1, characterized in that, Based on the subtitle entry array and the highlight metadata, highlight tags are attached to highlight words in the subtitle entry array that meet the highlighting conditions to obtain a processed subtitle entry array, including: Align the subtitle entry array with the highlighted metadata to determine the matched entries; Locate the highlighted words in each of the hit entries that satisfy the highlighted conditions; The highlighted words that meet the highlighted conditions are attached with the highlighted tags to obtain the processed subtitle entry array.

3. The method according to claim 2, characterized in that, The process of locating highlighted words that satisfy the highlighting conditions in each of the hit entries includes: For each hit entry containing multiple lines of text, identify the line to be highlighted; The highlighted words that meet the highlighted conditions are located sequentially in the row to be highlighted.

4. The method according to claim 3, characterized in that, The step of sequentially locating the highlighted words that satisfy the highlighting conditions in the row to be highlighted includes: For the j-th highlighted word in the i-th line to be highlighted, initialize a counter and a scanning cursor. The counter is used to record the number of times the j-th highlighted word has been scanned in the line to be highlighted. The scanning cursor is used to mark the starting position for finding the j-th highlighted word from the text to be scanned in the line to be highlighted. i and j are positive integers. Search for the j-th highlighted word from the starting position of the scanning cursor mark; If the j-th type of highlighted word is not found, the processing of the j-th type of highlighted word ends; If the j-th type of highlighted word is found, obtain the current occurrence count of the j-th type of highlighted word recorded by the counter, and perform processing on the j-th type of highlighted word based on the pre-set sequence number array and the current occurrence count; Update the value of the counter, and return to the step of searching for the j-th type of highlighted word from the starting position of the scanning cursor mark until the value of the counter exceeds the maximum index in the index array, and determine the j-th type of highlighted word that meets the highlighting condition from all the j-th type of highlighted words in the i-th row to be highlighted; The sequence number array is a set of occurrence numbers of highlighted words that meet the highlighting conditions.

5. The method according to claim 4, characterized in that, The process of processing the j-th highlighted word based on a pre-defined array of numbers and the current occurrence count includes: If the current occurrence count falls into the sequence number array, the j-th type of highlighted word with the current occurrence count is determined as the j-th type of highlighted word that satisfies the highlighting condition. Also, the remaining text in the i-th line to be highlighted is determined, the scanning cursor is reset to the starting position of the remaining text, and the process of searching for the j-th type of highlighted word from the starting position marked by the scanning cursor is returned. If the current occurrence count does not fall into the sequence number array, it is determined that the j-th type of highlighted word with the current occurrence count does not meet the highlighting condition. The scanning cursor is then moved to the end of the j-th type of highlighted word with the current occurrence count, and the process of searching for the j-th type of highlighted word from the beginning of the scanning cursor mark is returned.

6. The method according to claim 5, characterized in that, Determining the remaining text in the i-th line to be highlighted includes: Based on the end position of the j-th highlighted word with the current occurrence frequency, the text to be scanned in the i-th line to be highlighted that is located after the end position is determined as the remaining text in the i-th line to be highlighted.

7. The method according to claim 3, characterized in that, For each of the hit entries containing multiple lines of text, identifying the line to be highlighted includes: For each of the multi-line text entries that are hit, the proportion of letters is calculated line by line. The line or lines of text in which the proportion of the letters is greater than the threshold are identified as the lines to be highlighted.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Read the color configuration table from the configuration center; Write the color configuration table into the style block of the original subtitle file to generate a global style declaration; The color configuration table is used to indicate the class name and the corresponding color value.

9. The method according to claim 8, characterized in that, The process involves attaching highlight tags to highlighted words that meet the highlighting conditions to obtain the processed subtitle entry array, which includes: The highlighted words that meet the highlighted conditions are attached with the highlighted tags corresponding to the class names in the global style declaration to obtain the processed subtitle entry array.

10. The method according to any one of claims 1 to 7, characterized in that, The highlighted metadata includes at least one of the following: Single-sentence subtitle highlighting information; Highlighted words and their descriptions; Complete data consisting of multiple highlighted single-sentence subtitles.

11. The method according to claim 10, characterized in that, The single-sentence subtitle highlighting information includes at least one of the following: Subtitle number; Subtitle start time; The set of highlighted words in the caption.

12. The method according to claim 10, characterized in that, The highlighted vocabulary description information includes at least one of the following: Highlighted words to be highlighted; Highlighted labels used to indicate colors; The same highlighted words appear in the same subtitle; Vocabulary identifiers used to indicate highlighted words.

13. A video playback method, characterized in that, The method includes: In response to the selection action, determine the first video and the first vocabulary level; Send a playback request for the first video, the playback request being used to indicate the first video and the first vocabulary level, the first vocabulary level corresponding to the highlighted words to be highlighted in the subtitle content of the first video; Receive the returned highlighted subtitle file, and, based on the highlighted subtitle file, highlight at least a portion of the highlighted words corresponding to the first word level during the process of playing the first video and displaying the subtitle content of the first video; The method for synthesizing the highlighted subtitle file is as follows: based on the playback request, the original subtitle file and highlighted metadata are obtained; the original subtitle file is parsed to obtain a subtitle entry array; based on the subtitle entry array and the highlighted metadata, highlighted words in the subtitle entry array that meet the highlighted conditions are attached with highlighted tags to obtain a processed subtitle entry array; the processed subtitle entry array is serialized into a highlighted subtitle file. The original subtitle file is used to indicate the subtitle content of the first video, and the highlighted metadata is used to describe the highlighted information of the highlighted words corresponding to the first word level.

14. A subtitle file synthesis device, characterized in that, The device includes: The acquisition module is used to acquire a playback request for a first video. The playback request is used to indicate the first video and a first vocabulary level. The first vocabulary level corresponds to a highlighted vocabulary to be highlighted in the subtitle content of the first video. The processing module is used to obtain the original subtitle file and highlight metadata based on the playback request. The original subtitle file is used to indicate the subtitle content of the first video, and the highlight metadata is used to describe the highlight information of the highlighted words corresponding to the first word level. A synthesis module is used to parse the original subtitle file to obtain a subtitle entry array; based on the subtitle entry array and the highlight metadata, highlight tags are attached to the highlight words in the subtitle entry array that meet the highlight conditions to obtain a processed subtitle entry array; the processed subtitle entry array is serialized into a highlighted subtitle file, which is used to indicate that at least a portion of the highlighted words corresponding to the first word level are highlighted during the playback of the first video and the display of the subtitle content of the first video.

15. A video playback device, characterized in that, The device includes: The determination module is used to determine the first video and the first vocabulary level in response to the selection operation; The sending module is used to send a playback request for the first video. The playback request is used to indicate the first video and the first vocabulary level. The first vocabulary level corresponds to the highlighted vocabulary to be highlighted in the subtitle content of the first video. The highlighting module is used to receive the returned highlighted subtitle file, and, based on the highlighted subtitle file, highlight at least a portion of the highlighted words corresponding to the first word level during the process of playing the first video and displaying the subtitle content of the first video; The method for synthesizing the highlighted subtitle file is as follows: based on the playback request, the original subtitle file and highlighted metadata are obtained; the original subtitle file is parsed to obtain a subtitle entry array; based on the subtitle entry array and the highlighted metadata, highlighted words in the subtitle entry array that meet the highlighted conditions are attached with highlighted tags to obtain a processed subtitle entry array; the processed subtitle entry array is serialized into a highlighted subtitle file. The original subtitle file is used to indicate the subtitle content of the first video, and the highlighted metadata is used to describe the highlighted information of the highlighted words corresponding to the first word level.

16. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the subtitle file synthesis method as described in any one of claims 1 to 12, and / or, to implement the video playback method as described in claim 13.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is loaded and executed by a processor to implement the subtitle file synthesis method as described in any one of claims 1 to 12, and / or to implement the video playback method as described in claim 13.