Caption processing method and device based on editing template and computing equipment
By decomposing subtitles into minimum semantic units and aggregating them into semantic clusters based on semantic correlation, the problem of incomplete subtitle display is solved, and the semantic integrity and user experience are improved.
Patent Information
- Application Number
- CN202510696229.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, when subtitles are processed in the editing template, characters will be automatically intercepted after the number of words exceeds the threshold, resulting in incomplete subtitles display, affecting the effect and user experience.
Decompose the original subtitles into minimum semantic units, aggregate into semantic clusters according to the semantic correlation degree, and generate target subtitles based on the template word count threshold to ensure semantic integrity.
By generating subtitles composed of integer semantic clusters, subtitle semantic interruption is avoided, and the display effect and user experience are improved.
Smart Images

Figure CN120343353A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technologies, and in particular, to a subtitle processing method, apparatus, computing device, computer storage medium, and computer program product based on a clip template. Background Art
[0002] A clip template is a pre-designed standardized video framework that usually contains fixed elements and structures. Users can upload materials, add subtitles, etc. according to their own needs, so as to quickly produce video works that meet the user's needs.
[0003] In order to achieve good editing effects, clip templates usually limit the number of characters in subtitles, that is, ensure that the subtitles generated using the clip template do not exceed the corresponding character count threshold.
[0004] However, the inventor found in the implementation process that there are the following defects in the prior art: The prior art counts the number of characters in the input subtitle. When the number of characters in the input subtitle exceeds the character count threshold, the first N characters are automatically intercepted for display. For example, if the input subtitle is "I love mountains" and the character count threshold is 4, the prior art will intercept the first 5 characters for display, that is, the subtitle is displayed as "I lov". It can be seen from this that the subtitle processing method in the prior art will greatly reduce the integrity of the displayed subtitle, affect the subtitle display effect, and reduce the user experience. Summary of the Invention
[0005] In view of the above problems, this application is proposed to provide a subtitle processing method, apparatus, computing device, computer storage medium, and computer program product based on a clip template that overcomes the above problems or at least partially solves the above problems.
[0006] According to a first aspect of this application, there is provided a subtitle processing method based on a clip template, including:
[0007] Obtain any original subtitle input for a target clip template, and determine the template character count threshold corresponding to the original subtitle;
[0008] Decompose the original subtitle into multiple minimum semantic units;
[0009] Aggregate the multiple minimum semantic units according to the semantic association degree between the minimum semantic units, so as to split the original subtitle into at least one semantic cluster;
[0010] Extract a target semantic cluster from the semantic clusters corresponding to the original subtitle according to the template character count threshold, and generate a target subtitle according to the target semantic cluster.
[0011] In an optional implementation manner, the decomposing the original subtitle into multiple minimum semantic units includes:
[0012] Preprocess the original subtitles;
[0013] Identify the language corresponding to each character in the preprocessed original subtitles, and add the characters in the original subtitles to the corresponding language set;
[0014] For any language set, use the recognition algorithm matched by the language set to recognize the smallest semantic unit in the language set.
[0015] In an alternative embodiment, the aggregating the plurality of smallest semantic units according to the semantic association degree between the smallest semantic units includes:
[0016] Input the plurality of smallest semantic units into a pre-trained semantic aggregation model;
[0017] Obtain the semantic clusters output by the semantic aggregation model.
[0018] In an alternative embodiment, the aggregating the plurality of smallest semantic units according to the semantic association degree between the smallest semantic units includes:
[0019] For any smallest semantic unit, calculate the semantic association degree between this smallest semantic unit and the previous smallest semantic unit;
[0020] If the semantic association degree is greater than the association degree threshold, divide this smallest semantic unit into the semantic cluster corresponding to the previous smallest semantic unit;
[0021] If the semantic association degree is less than or equal to the association degree threshold, divide this smallest semantic unit into a new semantic cluster.
[0022] In an alternative embodiment, the extracting a target semantic cluster from the semantic clusters corresponding to the original subtitles includes:
[0023] Calculate the weights of each semantic cluster;
[0024] Extract a target semantic cluster from the semantic clusters corresponding to the original subtitles according to the weights of the semantic clusters.
[0025] In an alternative embodiment, the calculating the weights of each semantic cluster includes:
[0026] Determine the subtitle semantics of each semantic cluster;
[0027] Determine the target image corresponding to the original subtitles, and extract the image semantics of the target image;
[0028] For any semantic cluster, calculate the similarity between the subtitle semantics of the semantic cluster and the image semantics, and generate the weight of the semantic cluster according to the similarity.
[0029] In an alternative embodiment, the extracting the target semantic cluster from the semantic clusters corresponding to the original subtitle includes:
[0030] Extract the target semantic cluster from the semantic clusters corresponding to the original subtitle according to the position of the semantic cluster in the original subtitle.
[0031] According to the second aspect of the present application, there is provided a subtitle processing device based on a clip template, including:
[0032] An acquisition module, configured to acquire any original subtitle input for a target clip template and determine the template word count threshold corresponding to the original subtitle;
[0033] A decomposition module, configured to decompose the original subtitle into a plurality of minimum semantic units;
[0034] An aggregation module, configured to aggregate the plurality of minimum semantic units according to the semantic correlation degree between the minimum semantic units, so as to split the original subtitle into at least one semantic cluster;
[0035] A generation module, configured to extract a target semantic cluster from the semantic clusters corresponding to the original subtitle according to the template word count threshold, and generate a target subtitle according to the target semantic cluster.
[0036] In an alternative embodiment, the decomposition module is configured to: preprocess the original subtitle;
[0037] Identify the language corresponding to each character in the preprocessed original subtitle, and add the characters in the original subtitle to the corresponding language set;
[0038] For any language set, use the recognition algorithm matched by the language set to recognize the minimum semantic unit in the language set.
[0039] In an alternative embodiment, the decomposition module is configured to: input the plurality of minimum semantic units into a pre-trained semantic aggregation model;
[0040] Obtain the semantic clusters output by the semantic aggregation model.
[0041] In an alternative embodiment, the aggregation module is configured to: for any minimum semantic unit, calculate the semantic correlation degree between the minimum semantic unit and the previous minimum semantic unit;
[0042] If the semantic correlation degree is greater than the correlation degree threshold, divide the minimum semantic unit into the semantic cluster corresponding to the previous minimum semantic unit;
[0043] If the semantic correlation degree is less than or equal to the correlation degree threshold, divide the minimum semantic unit into a new semantic cluster.
[0044] In an alternative embodiment, the generation module is configured to: calculate the weights of each semantic cluster;
[0045] Extract a target semantic cluster from the semantic clusters corresponding to the original captions according to the weights of the semantic clusters.
[0046] In an alternative embodiment, the generation module is configured to: determine the caption semantics of each semantic cluster;
[0047] Determine a target image corresponding to the original caption, and extract the image semantics of the target image;
[0048] For any semantic cluster, calculate the similarity between the caption semantics of the semantic cluster and the image semantics, and generate the weight of the semantic cluster according to the similarity.
[0049] In an alternative embodiment, the generation module is configured to: extract a target semantic cluster from the semantic clusters corresponding to the original caption according to the position of the semantic cluster in the original caption.
[0050] According to a third aspect of the present application, there is provided a computing device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0051] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above-mentioned caption processing method based on a clip template.
[0052] According to a fourth aspect of the present application, there is provided a computer storage medium, and at least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to perform operations corresponding to the above-mentioned caption processing method based on a clip template.
[0053] According to a fifth aspect of the present application, there is provided a computer program product, including at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the above-mentioned caption processing method based on a clip template.
[0054] In the embodiments of the present application, the original subtitle is first decomposed into individual smallest semantic units, and then the smallest semantic units are aggregated according to the semantic correlation degree of the smallest semantic units, so as to split the original subtitle into at least one semantic cluster. Finally, the target subtitle is generated according to the semantic clusters of the original subtitle. Thus, the target subtitle generated by the embodiments of the present application is composed of an integer number of semantic clusters, and the semantic clusters have complete semantics, so as to ensure the semantic integrity of the displayed target subtitle, avoid the occurrence of subtitle semantic interruption phenomena, improve the subtitle display effect, and improve the user experience.
[0055] When processing the original subtitle, the embodiments of the present application first preprocess the original subtitle to remove invalid characters in the original subtitle, etc., and then identify the languages corresponding to each character in the preprocessed original subtitle, and use an identification algorithm matching the language to identify the smallest semantic units of the corresponding language, so as to improve the extraction accuracy of the smallest semantic units.
[0056] The embodiments of the present application use a semantic aggregation model constructed and trained based on a machine learning algorithm to perform the aggregation of the smallest semantic units, so as to improve the aggregation accuracy of the smallest semantic units and the generation accuracy of the semantic clusters.
[0057] The embodiments of the present application calculate the semantic correlation degree between any smallest semantic unit and its previous smallest semantic unit, and use the semantic correlation degree and the correlation degree threshold to divide the semantic clusters, so as to greatly improve the division efficiency of the semantic clusters.
[0058] The embodiments of the present application assign corresponding weights to different semantic clusters, and screen out the target semantic clusters according to the weights of the semantic clusters, so as to ensure that the generated target subtitle can not only meet the template word limit requirements, but also completely represent the user's intention, and improve the user experience.
[0059] The embodiments of the present application combine the subtitle semantics of the semantic clusters and the image semantics of the target images corresponding to the subtitles to obtain the weights of the semantic clusters, so as to improve the determination accuracy of the weights of the semantic clusters, ensure the matching degree of the extracted target semantic clusters with the user's intention, and improve the matching degree of the subtitles and the images, and improve the editing effect.
[0060] The embodiments of the present application extract the target semantic clusters from the semantic clusters corresponding to the original subtitle according to the positions of the semantic clusters in the original subtitle, so as to ensure the continuity of the obtained target semantic clusters, simplify the extraction process of the target semantic clusters, and improve the overall subtitle processing efficiency.
[0061] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified. Brief Description of the Drawings
[0062] Upon reading the following detailed description of the preferred embodiments, various other advantages and benefits will become apparent to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. Moreover, throughout the drawings, the same reference numerals are used to denote the same components. In the drawings:
[0063] Figure 1 A schematic diagram of an operating environment provided to implement at least one embodiment of the present application is shown;
[0064] Figure 2 A schematic flowchart of a subtitle processing method based on a clip template provided in Embodiment 1 of the present application is shown;
[0065] Figure 3 A schematic flowchart of a method for obtaining the smallest semantic unit provided in Embodiment 1 of the present application is shown;
[0066] Figure 4 A schematic flowchart of a subtitle processing method based on a clip template provided in Embodiment 2 of the present application is shown;
[0067] Figure 5 A schematic flowchart of a method for calculating the weight of a semantic cluster provided in Embodiment 2 of the present application is shown;
[0068] Figure 6 A schematic structural diagram of a subtitle processing device based on a clip template provided in Embodiment 3 of the present application is shown;
[0069] Figure 7 A schematic structural diagram of a computing device provided in Embodiment 4 of the present application is shown. Detailed Embodiments
[0070] The exemplary embodiments of the present application will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be completely conveyed to those skilled in the art.
[0071] First, a brief introduction to the relevant technical terms involved in the embodiments of the present application is given:
[0072] Clip template, a pre-designed standardized video framework, which usually contains fixed elements and structures, and users can upload materials, add subtitles, etc. according to their own needs.
[0073] Subtitle refers to the non-video content such as dialogue in a TV program, movie, or stage work displayed in text form, and also generally refers to the text in the post-processing of a video work.
[0074] Template word count threshold, which is the threshold for the word count of subtitles in the editing template.
[0075] It should be noted that the data related to users involved in the embodiments of this application are all information and data that have been authorized or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0076] Figure 1 FIG. shows a schematic diagram of an operating environment provided to implement at least one embodiment of this application. This application can be applied to an application environment including, but not limited to, client 2, server 4, and network 6.
[0077] Wherein:
[0078] Server 4 can be composed of a single or multiple computing devices. The multiple computing devices can include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing devices can load virtual machines based on virtual images and / or other data that define specific software (e.g., operating systems, dedicated applications, servers) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices. A hypervisor can be implemented to manage the use of different virtual machines on the same computing device.
[0079] Server 4 can be configured to communicate with client 2, etc. via network 6. Network 6 includes various network devices, such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. Network 6 can include physical links, such as coaxial cable links, twisted pair cable links, fiber optic links and their combinations, etc., or wireless links, such as cellular links, satellite links, Wi-Fi links, etc.
[0080] Server 4 can provide services such as storage, reading, downloading, writing, querying, deleting, etc., such as providing a download service for static resources for the client through multiple domain names.
[0081] Client 2 can be an electronic device running operating systems such as Windows, Android TM ) or IOS, such as a smart phone, a tablet device, a laptop computer, a virtual reality device, a gaming device, a set-top box, an in-vehicle terminal, a smart TV. Based on the above operating systems, various application programs, such as browsers, can be run.
[0082] Embodiment 1
[0083] Figure 2 FIG. 1 shows a schematic flowchart of a subtitle processing method based on a clip template provided in Embodiment 1 of the present application. The subtitle processing method provided in the embodiments of the present application can be executed by the above-mentioned client and / or server.
[0084] Specifically, as Figure 2 shown, the method includes the following steps:
[0085] Step S201: Obtain any original subtitle input for the target clip template, and determine the template word count threshold corresponding to the original subtitle.
[0086] The target clip template is the clip template adopted in the current clip process by the user. The target clip template contains at least one subtitle entry point through which the user can input the corresponding subtitle, and the subtitle originally input by the user is called the original subtitle. The target clip module usually configures corresponding word count limits for each subtitle. Therefore, in the present application, for any original subtitle, the template word count threshold corresponding to the original subtitle is determined.
[0087] In the specific implementation process, for any original subtitle input by the user, subsequent steps are executed to obtain the corresponding target subtitle, which is the subtitle to be displayed after being processed by the present application.
[0088] Step S202: Decompose the original subtitle into multiple minimum semantic units.
[0089] For each original subtitle, the original subtitle is decomposed into individual minimum semantic units. Among them, the minimum semantic unit is the smallest language unit that can independently express meaning and can be independently used grammatically in the corresponding language. For example, Chinese characters are the minimum semantic units in Chinese, and English words are the minimum semantic units in English, and so on. In the specific implementation process, the original subtitle can be split into multiple minimum semantic units by means of regular matching, named entity recognition, etc.
[0090] In an optional implementation manner, the minimum semantic unit can be specifically obtained by Figure 3 the steps S2021-S2023 shown:
[0091] S2021: Preprocess the original subtitle.
[0092] Among them, the preprocessing includes but is not limited to: removing HTML tags, URLs, invalid symbols (such as tab characters, line breaks, consecutive spaces), etc. After preprocessing the original subtitle, the preprocessed original subtitle is obtained.
[0093] S2022: Identify the language corresponding to each character in the preprocessed original subtitle, and add the characters in the original subtitle to the corresponding language set.
[0094] The original subtitles are composed of characters. The existing language recognition algorithm can be used to determine the language corresponding to each character and store it in the corresponding language set. That is, the language set corresponds to the language involved in the characters in the original subtitles one by one, and the characters contained in each language set correspond to the same language. For example, if the original subtitles involve two languages, Chinese and English, two sets, a Chinese set and an English set, are generated. The Chinese set is used to store Chinese characters, and the English set is used to store English characters. If the original subtitles only involve a single language, it only corresponds to one language set.
[0095] Optionally, the characters are specifically added to the language set in the order in the original subtitles. In the specific implementation process, when it is determined that the current character needs to be added to the corresponding language set, it is determined whether there is a gap between the current character and the last added character of the current language set (for example, if there is a separator such as a space between the current character and the last added character in the original subtitles, or there are other language characters between the current character and the last added character in the original subtitles, then it is determined that there is a gap between the current character and the last added character of the current language set). If there is a gap, the current character is added to the corresponding language set as a new set element. If there is no gap, the current character is merged with the last added character of the current language set (that is, the current character and the last added character of the current language set are merged into one set element). For example, the original subtitle is "我love大sea". If the current character is "v", the current English set is {lo}. Since "v" corresponds to the English set and there is no gap between "v" and the last added character "o" of the current English set, {lov} is obtained after "v" is added to the English set, and "lov" is a set element); if the current character is "s", the current English set is {love}. Since "s" corresponds to the English set and there is a gap between "v" and the last added character "e" of the current English set, {love, s} is obtained after "s" is added to the English set, that is, "love" is a set element and "s" is another set element.
[0096] S2023: For any language set, use a recognition algorithm that matches the language set to identify the smallest semantic unit in the language set.
[0097] Different extraction algorithms are configured for different languages, and the extraction algorithms can be flexibly configured and issued, thereby improving the scalability of the solution. For example, different regular expressions can be configured for different languages, and the minimum semantic unit of the corresponding language can be extracted through the regular expressions of the corresponding language.
[0098] Step S203: Aggregate multiple minimum semantic units according to the semantic correlation degree between the minimum semantic units, so as to split the original subtitle into at least one semantic cluster.
[0099] Although a minimum semantic unit can express meaning independently, there are many cases where a minimum semantic unit cannot express the semantics completely. For example, the Chinese character "ke" is a minimum semantic unit, but it cannot express the semantics completely. In view of this, after splitting the original subtitle into minimum semantic units, the embodiments of the present application further aggregate the minimum semantic units to obtain semantic clusters, and each semantic cluster can independently and completely express the corresponding semantics. For example, the minimum semantic unit "ke" cannot express the semantics completely, but semantic clusters containing "ke" such as "cola" and "irresistible" can express complete semantics.
[0100] In a specific implementation process, after decomposing the original subtitle to obtain each minimum semantic unit, arrange each minimum semantic unit in order according to the order in the original subtitle. There is a corresponding semantic correlation degree between two adjacent minimum semantic units. The higher the semantic correlation degree, the higher the probability that the two minimum semantic units are combined to express complete semantics. For example, the semantic correlation degree between two adjacent minimum semantic units in idioms, fixed phrases, allegorical sayings, etc. is high. That is, the semantic correlation degree between the minimum semantic units in the present application specifically refers to the semantic correlation degree between any two adjacent minimum semantic units. The obtained semantic clusters can be assigned corresponding labels to identify that the minimum semantic units in the semantic cluster are a whole, and can also be stored in the form of a map or the like.
[0101] Perform an aggregation process on multiple minimum semantic units arranged in order after splitting the original subtitle according to the semantic correlation degree. After aggregation, at least one semantic cluster is obtained, and each semantic cluster contains one or more minimum semantic units. Moreover, the multiple minimum semantic units included in the same semantic cluster are adjacent, and each semantic cluster can express complete semantics. Thus, the original subtitle can be split into at least one semantic cluster through the aggregation process.
[0102] In an alternative implementation manner, one or more of the following methods can be used to aggregate the minimum semantic units:
[0103] Aggregation method 1: Input multiple minimum semantic units into a pre-trained semantic aggregation model to obtain the semantic clusters output by the semantic aggregation model. Specifically, a semantic aggregation model is constructed in advance based on a machine learning algorithm. For example, a semantic aggregation model can be constructed based on a deep learning neural network algorithm. Moreover, each historical caption in the editing platform is obtained. The historical caption can be a test caption created by the editing platform tester or a caption obtained from the user's editing process after obtaining full authorization from the user. Further, the historical captions are labeled with semantic clusters according to the semantic correlation degree between adjacent minimum semantic units in the historical captions. Then, sample data is generated based on the historical captions and the corresponding semantic cluster labels, and the constructed semantic aggregation model is trained using the sample data. After meeting the training end condition, a trained semantic aggregation model is obtained. Input the minimum semantic units of the original caption obtained in step S202 into the semantic aggregation model. The semantic aggregation model outputs the semantic clusters corresponding to the original caption through learning and processing the semantic correlation degree of the minimum semantic units. By adopting this method, semantic clusters can be accurately divided, and the semantic integrity of the subsequent target caption can be improved.
[0104] Aggregation method 2: For any minimum semantic unit, calculate the semantic correlation degree between this minimum semantic unit and the previous minimum semantic unit. If the semantic correlation degree is greater than the correlation degree threshold, divide this minimum semantic unit into the semantic cluster corresponding to the previous minimum semantic unit. If the semantic correlation degree is less than or equal to the correlation degree threshold, divide this minimum semantic unit into a new semantic cluster. By adopting this method, the semantic correlation degree between semantic units can be quantified, and then the division of semantic clusters can be quickly realized according to the correlation degree threshold, improving the division efficiency of semantic clusters. Moreover, the correlation degree threshold can be set statically or adjusted dynamically, so as to adapt to caption processing in different scenarios and improve the scalability of this solution.
[0105] Further optionally, the semantic correlation degree between this minimum semantic unit and the previous minimum semantic unit can be specifically calculated by one or more of the following methods:
[0106] Association degree calculation method 1: A word library is pre-generated, which contains multiple word groups, and each word group consists of multiple minimum semantic units. For example, idioms, fixed phrases, proverbs, etc. can be added to the word library. Calculate the co-occurrence frequency of this minimum semantic unit and the previous minimum semantic unit in the word library, and determine the semantic association degree according to the co-occurrence frequency. Among them, the semantic association degree is positively correlated with the co-occurrence frequency. In addition, different weights can be assigned to different types of word groups in the word library. For example, the weight of idioms can be greater than the weight of two-character common phrases, and so on. Then, during the implementation of this method, calculate the target word group that matches this minimum semantic unit and the previous minimum semantic unit in the word library, and determine the weight of the target word group. Furthermore, combine the co-occurrence frequency of this minimum semantic unit and the previous minimum semantic unit in the word library and the weight of the target word group to comprehensively determine the semantic association degree. This calculation method is simple and easy to implement, does not require relying on a large computing model, and has low costs.
[0107] Association degree calculation method 2: Use a pre-trained model to calculate the semantic association degree between this minimum semantic unit and the previous minimum semantic unit. For example, the open-source BERT (Bidirectional Encoder Representations from Transformers) model can be used to calculate the similarity between this minimum semantic unit and the previous minimum semantic unit, and use this similarity as the semantic association degree. Using this method can accurately calculate the semantic association degree between this minimum semantic unit and the previous minimum semantic unit.
[0108] Step S204, according to the template word count threshold, extract the target semantic cluster from the semantic clusters corresponding to the original subtitle, and generate the target subtitle according to the target semantic cluster.
[0109] Taking the template word count threshold as a constraint condition, extract at least one target semantic cluster from the semantic clusters corresponding to the original subtitle, arrange the target semantic clusters in order to obtain the target subtitle, and display the target subtitle. Among them, the number of target semantic clusters ≤ the template word count threshold.
[0110] Since the target subtitle is composed of at least one semantic cluster, and each semantic cluster can express complete semantics, the generated target subtitle can achieve complete semantic expression.
[0111] It can be seen that in the subtitle processing method based on a clip template provided by the embodiments of the present application, the original subtitle is first decomposed into individual smallest semantic units, and the smallest semantic units are aggregated according to the semantic correlation degree of the smallest semantic units to obtain semantic clusters, so as to decompose the original subtitle into at least one semantic cluster. Finally, a target subtitle is generated according to the semantic clusters of the original subtitle. Thus, the target subtitle generated by the embodiments of the present application is composed of an integer number of semantic clusters, and the semantic clusters have complete semantics, thereby ensuring the semantic integrity of the displayed target subtitle, avoiding the occurrence of subtitle semantic interruption phenomena, improving the subtitle display effect, and improving the user experience.
[0112] Embodiment 2
[0113] Figure 4 Fig. shows a schematic flowchart of a subtitle processing method based on a clip template provided by Embodiment 2 of the present application. The subtitle processing method provided by the embodiments of the present application can be executed by the above-mentioned client and / or server.
[0114] Specifically, as Figure 4 shown, the method includes the following steps:
[0115] Step S401, obtain any original subtitle input for the target clip template, and determine the template word count threshold corresponding to the original subtitle.
[0116] Step S402, decompose the original subtitle into multiple smallest semantic units.
[0117] Step S403, determine whether the number of the smallest semantic units exceeds the template word count threshold; if so, execute Step S404; if not, execute Step S407.
[0118] Count the number of the smallest semantic units of the original subtitle, that is, obtain the number of the smallest semantic units. Compare the number of the smallest semantic units with the template word count threshold corresponding to the original subtitle. Different target subtitle generation methods are adopted according to different comparison results.
[0119] Specifically, if the number of the smallest semantic units exceeds the template word count threshold, it indicates that if all the smallest semantic units are added to the template, the template word count threshold will be exceeded, and thus the displayed subtitle will block other display areas. In view of this, in the embodiments of the present application, by executing subsequent Step S404 - Step S406 to obtain the target subtitle according to the semantic clusters, the smallest semantic units included in the displayed subtitle are streamlined to avoid blocking other display areas; if the number of the smallest semantic units does not exceed the template word count threshold, execute subsequent Step S407 to obtain the target subtitle according to the smallest semantic units, so as to display all the smallest semantic units through the target subtitle.
[0120] Step S404: Aggregate multiple minimum semantic units according to the semantic correlation degree between the minimum semantic units, so as to split the original subtitle into at least one semantic cluster.
[0121] Among them, the specific implementation processes of steps S401, S402, and S404 can refer to the description in Embodiment 1 and will not be elaborated here.
[0122] Step S405: Calculate the weights of each semantic cluster, and extract the target semantic cluster from the semantic clusters corresponding to the original subtitle according to the template word count threshold and the weights of the semantic clusters.
[0123] When the number of minimum semantic units obtained by decomposing the original subtitle exceeds the corresponding template word count threshold, the subtitle area cannot accommodate all semantic clusters, so it is necessary to screen the semantic clusters, and the screened semantic clusters are the target semantic clusters.
[0124] In the specific screening process, first calculate the weights of each semantic cluster, and then screen the target semantic cluster according to the weights of the semantic clusters.
[0125] In an alternative implementation, the semantic cluster weight can be specifically calculated through Figure 5 the steps S4051 - S4054 shown as follows:
[0126] S4051: Determine the subtitle semantics of each semantic cluster.
[0127] For any semantic cluster, determine the semantic keyword of this semantic cluster, and this semantic keyword is the subtitle semantics of the semantic cluster. Among them, the subtitle semantics of the semantic cluster can include: entity keyword, emotion keyword. Specifically, the entity keyword of the semantic cluster can be extracted through a named entity recognition algorithm, such as extracting entity names, scenes, etc. that appear in the semantic cluster; the emotion keyword of the semantic cluster can also be extracted through an emotion extraction algorithm, such as extracting the emotion keyword of the semantic cluster as positive, sad, etc. The subtitle semantics of the semantic cluster can also be analyzed by calling a machine learning model.
[0128] S4052: Determine the target image corresponding to the original subtitle.
[0129] Specifically, the frame where the original subtitle is located can be used as the target image, or the frame where the original subtitle is located, as well as the first N frames and / or the last N frames of the frame can be used as the target image together.
[0130] S4053: Extract the image semantics of the target image.
[0131] An image semantic extraction model is pre-trained based on a machine learning algorithm. This image semantic extraction model can analyze the image semantics of an image. Among them, the image semantic extraction model can be deployed on the user side or the server side according to the performance of the user side. This application does not limit the specific structure and training method of the image semantic extraction model, etc.
[0132] In the actual implementation process, the target image is input into the image semantic extraction model, and the semantic keywords output by the image semantic extraction model are obtained. Here, the semantic keywords are the image semantics of the target image. The image semantics can include: entity keywords and / or emotion keywords, etc. Among them, the entity keywords can include the element names, scene names, etc. included in the target image; the emotion keywords can include the emotion characteristics expressed by the target image.
[0133] S4054, for any semantic cluster, calculate the similarity between the subtitle semantics and the image semantics of this semantic cluster, and generate the weight of this semantic cluster according to this similarity.
[0134] For any semantic cluster, calculate the similarity between the subtitle semantics of this semantic cluster and the image semantics of the target image, and then calculate the weight of this semantic cluster according to this similarity. Among them, the weight of the semantic cluster is positively correlated with the corresponding similarity.
[0135] After obtaining the weight of the semantic cluster, extract the target semantic cluster in descending order of weight.
[0136] Optionally, the extraction process of the target semantic cluster can adopt one or a combination of the following extraction methods:
[0137] Extraction method 1: From the semantic clusters that have not been used as target semantic clusters, extract the semantic cluster with the highest current weight as the target semantic cluster, and judge whether the total number of the smallest semantic units included in the extracted target semantic clusters exceeds the template word count threshold; if so, end the semantic cluster extraction; if not, execute again the step of extracting the semantic cluster with the highest current weight as the target semantic cluster from the semantic clusters that have not been used as target semantic clusters. The total number of the smallest semantic units included in the target semantic cluster extracted by this method is equal to or slightly exceeds the template word count threshold.
[0138] Extraction method 2: From the semantic clusters that have not yet been the target semantic clusters, extract the semantic cluster with the highest current weight, and calculate the total number of the smallest semantic units of the semantic cluster with the highest current weight and the extracted target semantic cluster; if the total number of the smallest semantic units exceeds the template word count threshold, discard the semantic cluster with the highest current weight and end the extraction of the target semantic cluster; if the total number of the smallest semantic units does not exceed the template word count threshold, use the semantic cluster with the highest current weight as the target semantic cluster, and then execute the step of extracting the semantic cluster with the highest current weight from the semantic clusters that have not yet been the target semantic clusters again. The total number of the smallest semantic units included in the target semantic cluster extracted in this way does not exceed the template word count threshold.
[0139] Step S406, generate the target caption according to the target semantic cluster.
[0140] The target caption is obtained by concatenating the target semantic clusters in order according to their positions in the original caption.
[0141] In an alternative implementation, steps S405 - S406 can be replaced with: According to the positions of the semantic clusters in the original caption, extract the target semantic clusters from the semantic clusters corresponding to the original caption. Specifically, extract the semantic clusters in order from front to back. After extracting a semantic cluster, calculate whether the total number of the smallest semantic units included in the currently extracted semantic cluster and the target semantic cluster exceeds the template word count threshold; if it exceeds, discard the currently extracted semantic cluster and end the extraction of the semantic clusters; if it does not exceed, use the currently extracted semantic cluster as the target semantic cluster and extract the next semantic cluster in order. The target semantic clusters obtained in this way are adjacent in order, and the total number of the smallest semantic units of all the extracted target semantic clusters does not exceed the template word count threshold. Or, extract the semantic clusters as the target semantic clusters in order from front to back, and calculate whether the total number of the smallest semantic units included in the currently obtained target semantic cluster exceeds the template word count threshold; if it exceeds, end the extraction of the semantic clusters; if it does not exceed, extract the next semantic cluster in order. The target semantic clusters obtained in this way are adjacent in order, and the total number of the smallest semantic units of all the extracted target semantic clusters is equal to or slightly greater than the template word count threshold.
[0142] In an alternative implementation, steps S404 - S406 can be replaced with: Provide the original caption to a pre - set large - language model, generate a prompt word according to the template word count threshold (such as: Condense the following caption into a caption with a word count less than or equal to the template word count threshold M), and obtain the condensed result output by the pre - set large - language model as the target caption. Further optionally, the target image or the image semantics of the target image can also be provided to the pre - set large - language model, and the pre - set large - language model refines the original caption in combination with the target image.
[0143] In an alternative embodiment, if the total number of the minimum semantic units of the target subtitle obtained in the above manner slightly exceeds the template word count threshold, the font size or spacing of the target subtitle can be adjusted to ensure that the target subtitle does not exceed the display range of the subtitle template.
[0144] Further optionally, the number of minimum semantic units of each target subtitle generated by the target clip module within a preset time period and the corresponding template word count threshold can be obtained, and the difference between the number of minimum semantic units of each target subtitle and the template word count threshold is calculated. If the number of minimum semantic units of the target subtitle is greater than the template word count threshold, the difference is a positive value; if the number of minimum semantic units of the target subtitle is less than or equal to the template word count threshold, the difference is a negative value or zero. The proportion of subtitles with a positive difference is counted, and it is determined whether this proportion of subtitles exceeds a preset proportion threshold; if so, it indicates that the proportion of subtitles actually exceeding the template word count threshold is relatively large, and a corresponding template word count threshold adjustment prompt message is generated for the relevant processing end to adjust the template word count threshold according to this template word count threshold adjustment prompt message.
[0145] In an alternative embodiment, when generating a target subtitle according to a target semantic cluster, the target semantic clusters can be directly concatenated, and corresponding connecting words, spaces, etc. can also be added to ensure the coherence between the target semantic clusters. Specifically, the target semantic clusters are arranged in sequence according to their positions in the original subtitle. For any target semantic cluster, if this target semantic cluster is connected to the previous target semantic cluster in the original subtitle, this target semantic cluster is concatenated with the previous target semantic cluster; if this target semantic cluster is not connected to the previous target semantic cluster in the original subtitle, a connecting symbol such as a small-size space character or "|" is inserted between this target semantic cluster and the previous target semantic cluster.
[0146] Step S407, generate a target subtitle according to the minimum semantic units.
[0147] If the number of minimum semantic units included in the original subtitle does not exceed the template word count threshold, they are arranged in order according to their positions in the original subtitle, thereby generating a target subtitle. That is, in this case, there is no need to crop the minimum semantic units.
[0148] It can be seen that the subtitle processing method based on a clip template provided in the embodiments of the present application, in the case where the number of minimum semantic units of the original subtitle exceeds the template word count threshold, performs semantic aggregation to obtain semantic clusters, and then filters target semantic clusters according to the weights of the semantic clusters, ensuring that the generated target subtitle can not only meet the word count limit requirements but also completely represent the user's intention, improving the user experience.
[0149] Embodiment III
[0150] Figure 6The figure shows a schematic structural diagram of a subtitle processing device based on a clip template provided in the third embodiment of the present application. As Figure 6 shown, the device 600 includes: an acquisition module 610, a decomposition module 620, an aggregation module 630, and a generation module 640.
[0151] The acquisition module 610 is configured to acquire any original subtitle input for the target clip template and determine the template word count threshold corresponding to the original subtitle;
[0152] The decomposition module 620 is configured to decompose the original subtitle into multiple minimum semantic units;
[0153] The aggregation module 630 is configured to aggregate the multiple minimum semantic units according to the semantic association degree between the minimum semantic units, so as to split the original subtitle into at least one semantic cluster;
[0154] The generation module 640 is configured to extract a target semantic cluster from the semantic clusters corresponding to the original subtitle according to the template word count threshold, and generate a target subtitle according to the target semantic cluster.
[0155] In an optional implementation manner, the decomposition module 620 is configured to: preprocess the original subtitle;
[0156] Identify the language corresponding to each character in the preprocessed original subtitle, and add the characters in the original subtitle to the corresponding language set;
[0157] For any language set, use the recognition algorithm matched by the language set to recognize the minimum semantic units in the language set.
[0158] In an optional implementation manner, the decomposition module 620 is configured to: input the multiple minimum semantic units into a pre-trained semantic aggregation model;
[0159] Obtain the semantic clusters output by the semantic aggregation model.
[0160] In an optional implementation manner, the aggregation module 630 is configured to: calculate the semantic association degree between any minimum semantic unit and the previous minimum semantic unit;
[0161] If the semantic association degree is greater than the association degree threshold, divide the minimum semantic unit into the semantic cluster corresponding to the previous minimum semantic unit;
[0162] If the semantic association degree is less than or equal to the association degree threshold, divide the minimum semantic unit into a new semantic cluster.
[0163] In an optional implementation manner, the generation module 640 is configured to: calculate the weights of each semantic cluster;
[0164] Extract a target semantic cluster from the semantic clusters corresponding to the original subtitles according to the weight of the semantic cluster.
[0165] In an alternative embodiment, the generation module 640 is configured to: determine the subtitle semantics of each semantic cluster;
[0166] Determine a target image corresponding to the original subtitle, and extract the image semantics of the target image;
[0167] For any semantic cluster, calculate the similarity between the subtitle semantics of the semantic cluster and the image semantics, and generate the weight of the semantic cluster according to the similarity.
[0168] In an alternative embodiment, the generation module 640 is configured to: extract a target semantic cluster from the semantic clusters corresponding to the original subtitle according to the position of the semantic cluster in the original subtitle.
[0169] It can be seen that the subtitle processing device based on the clip template provided by the embodiment of the present application first decomposes the original subtitle into individual smallest semantic units, and aggregates the smallest semantic units according to the semantic correlation degree of the smallest semantic units to obtain semantic clusters, so as to decompose the original subtitle into at least one semantic cluster, and finally generate a target subtitle according to the semantic clusters of the original subtitle. Thus, the target subtitle generated by the embodiment of the present application is composed of an integer number of semantic clusters, and the semantic cluster has a complete semantics, thereby ensuring the semantic integrity of the displayed target subtitle, avoiding the occurrence of subtitle semantic interruption phenomenon, improving the subtitle display effect, and improving the user experience.
[0170] Embodiment Four
[0171] Figure 7 Fig. shows a schematic structural diagram of a computing device provided by the fourth embodiment of the present application. The specific implementation of the computing device is not limited in the specific embodiments of the present application.
[0172] As Figure 7 shown, the computing device may include: a processor 702, a communication interface 704, a memory 706, and a communication bus 708.
[0173] Wherein: the processor 702, the communication interface 704, and the memory 706 communicate with each other through the communication bus 708. The communication interface 704 is used to communicate with network elements of other devices such as clients or other servers. The processor 702 is used to execute the program 710, and specifically may execute the relevant steps in the foregoing embodiment of the subtitle processing method based on the clip template for the computing device.
[0174] Specifically, the program 710 may include program code, which includes computer operation instructions.
[0175] The processor 702 may be a central processing unit (CPU), or a specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0176] The memory 706 is used to store the program 710. The memory 706 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory. The program 710 is specifically configured to cause the processor 702 to perform the operations in the above-described subtitle processing method embodiments based on the clip template.
[0177] Embodiment Five
[0178] Embodiment Five of the present application provides a non-volatile computer storage medium, which stores at least one executable instruction or computer program, and the executable instruction or computer program can cause a processor to perform the operations corresponding to the subtitle processing method based on the clip template in any of the above method embodiments.
[0179] Embodiment Six
[0180] Embodiment Six of the present application provides a computer program product, which includes at least one executable instruction or computer program, and the executable instruction or computer program can cause a processor to perform the operations corresponding to the subtitle processing method based on the clip template in any of the above method embodiments.
[0181] In summary, according to the computing device, computer storage medium, and computer program product provided in this embodiment, the original subtitle is first decomposed into individual smallest semantic units, and the smallest semantic units are aggregated according to the semantic association degree of the smallest semantic units to obtain semantic clusters, so that the original subtitle is decomposed into at least one semantic cluster, and finally the target subtitle is generated according to the semantic clusters of the original subtitle. Thus, the target subtitle generated by the embodiments of the present application is composed of an integer number of semantic clusters, and the semantic clusters have complete semantics, thereby ensuring the semantic integrity of the displayed target subtitle, avoiding the occurrence of subtitle semantic interruption phenomena, improving the subtitle display effect, and enhancing the user experience.
[0182] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. A variety of general-purpose systems can also be used in conjunction with the teachings based herein. The structure required to construct such systems will be apparent from the above description. Additionally, the embodiments of the present application are not directed to any particular programming language. It should be understood that the content of the present application described herein can be implemented using a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present application.
[0183] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0184] Similarly, it should be understood that, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, the various features of the embodiments of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed present application requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single embodiments disclosed previously. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present application.
[0185] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from those of the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0186] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of this application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0187] Each component embodiment of this application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of this application. This application can also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0188] It should be noted that the above embodiments illustrate rather than limit this application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.
Claims
1. A subtitle processing method based on a clip template, characterized in that including: obtaining any original subtitle input for a target clip template, and determining a template word count threshold corresponding to the original subtitle; decomposing the original subtitle into a plurality of minimum semantic units; aggregating the plurality of minimum semantic units according to the semantic correlation degree between the minimum semantic units, so as to split the original subtitle into at least one semantic cluster; extracting a target semantic cluster from the semantic clusters corresponding to the original subtitle according to the template word count threshold, and generating a target subtitle according to the target semantic cluster.
2. The method according to claim 1, wherein The decomposing the original subtitle into a plurality of minimum semantic units includes: performing preprocessing on the original subtitle; identifying the language corresponding to each character in the preprocessed original subtitle, and adding the characters in the original subtitle to the corresponding language set; for any language set, using an identification algorithm matched with the language set to identify the minimum semantic units in the language set.
3. The method according to claim 1 or 2, characterized in that, The aggregating the plurality of minimum semantic units according to the semantic correlation degree between the minimum semantic units includes: inputting the plurality of minimum semantic units into a pre-trained semantic aggregation model; obtaining the semantic clusters output by the semantic aggregation model.
4. The method according to claim 1 or 2, characterized in that, The aggregating the plurality of minimum semantic units according to the semantic correlation degree between the minimum semantic units includes: for any minimum semantic unit, calculating the semantic correlation degree between this minimum semantic unit and the previous minimum semantic unit; if the semantic correlation degree is greater than the correlation degree threshold, classifying this minimum semantic unit into the semantic cluster corresponding to the previous minimum semantic unit; if the semantic correlation degree is less than or equal to the correlation degree threshold, classifying this minimum semantic unit into a new semantic cluster.
5. The method according to any one of claims 1-4, characterized in that The extracting a target semantic cluster from the semantic clusters corresponding to the original subtitle includes: calculating the weights of the semantic clusters; extracting a target semantic cluster from the semantic clusters corresponding to the original subtitle according to the weights of the semantic clusters.
6. The method according to claim 5, wherein The calculating the weights of the semantic clusters includes: determining the subtitle semantics of each semantic cluster; determining a target image corresponding to the original subtitle, and extracting the image semantics of the target image; for any semantic cluster, calculating the similarity between the subtitle semantics of this semantic cluster and the image semantics, and generating the weight of this semantic cluster according to the similarity.
7. The method according to any one of claims 1 to 4, characterized in that The extracting a target semantic cluster from the semantic clusters corresponding to the original subtitle includes: extracting a target semantic cluster from the semantic clusters corresponding to the original subtitle according to the position of the semantic cluster in the original subtitle.
8. A subtitle processing device based on a clip template, characterized in that, including: an obtaining module, configured to obtain any original subtitle input for a target clip template, and determine a template word count threshold corresponding to the original subtitle; a decomposing module, configured to decompose the original subtitle into a plurality of minimum semantic units; an aggregating module, configured to aggregate the plurality of minimum semantic units according to the semantic correlation degree between the minimum semantic units, so as to split the original subtitle into at least one semantic cluster; a generating module, configured to extract a target semantic cluster from the semantic clusters corresponding to the original subtitle according to the template word count threshold, and generate a target subtitle according to the target semantic cluster.
9. A computing device, characterized in that, including: A processor, a memory, a communication interface, and a communication bus, where the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the subtitle processing method based on a clip template according to any one of claims 1-7.
10. A computer storage medium, characterized in that, At least one executable instruction is stored in the storage medium, and the executable instruction causes the processor to perform operations corresponding to the subtitle processing method based on a clip template according to any one of claims 1-7.
11. A computer program product, characterized in that, It includes at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the subtitle processing method based on a clip template according to any one of claims 1-7.