A video text processing method, apparatus, electronic device, and storage medium
By analyzing scene information and identifying industry characteristics of video ringback tones, and using a generative pre-trained model to generate personalized video scripts, the video ringback tone generation needs of users in different industries have been addressed, improving the accuracy of video scripts and user experience.
Patent Information
- Application Number
- CN202411333038.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-09-24
AI Technical Summary
In existing technologies, video ringback tone generation methods cannot meet the personalized needs of users in different industries, resulting in the text and video materials being out of sync, the text content not being smooth enough, and affecting the user experience.
By analyzing scene information and identifying industry characteristics of the video ringback tones to be processed, a generative pre-trained model is used to generate video scripts that conform to industry characteristics. Combined with voice and text data, text generation processing is performed, background music matching is optimized, and personalized video ringback tones are generated.
It improves the accuracy and efficiency of video script generation, ensuring that the generated scripts are relevant to the merchant's industry and the video scenario, thereby enhancing the user experience and promotional effectiveness.
Smart Images

Figure CN119316522B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a video text processing method, apparatus, electronic device, and storage medium. Background Technology
[0002] Video ringback tones refer to the video content displayed to users before a voice or video call is connected. One technology involves generating video ringback tones through proxy channels, using generic text schemes and video templates with modifications to key information. However, in practical applications, it has been found that user groups across different industries vary significantly, and generic text schemes cannot meet the needs of diverse sectors. This results in inconsistencies between the text and video materials, less fluid text content, and incomplete promotional information, negatively impacting the user experience. In conclusion, the technical issues in these technologies require improvement. Summary of the Invention
[0003] The main objective of this application is to provide a video text processing method, apparatus, electronic device, and storage medium that can improve the user experience.
[0004] To achieve the above objectives, one aspect of this application proposes a video text processing method, the method comprising:
[0005] Get the video ringback tone to be processed;
[0006] The video ringback tone to be processed is converted into text to obtain a set of voice and text data.
[0007] The video ringback tone to be processed is subjected to scene information analysis and processing to obtain a video scene data set;
[0008] The video scene data set is subjected to industry feature recognition processing to obtain industry feature information;
[0009] The target video script is obtained by performing text generation processing on the voice and text data set based on the video scene data set and the industry feature information.
[0010] In some embodiments, the text conversion processing of the video ringback tone to be processed to obtain a voice-text data set includes the following steps:
[0011] The audio file of the video ringback tone to be processed is subjected to background music filtering to obtain dry text audio;
[0012] The dry text is processed by speech recognition to obtain the speech-text data set.
[0013] In some embodiments, the step of performing scene information analysis and processing on the video ringback tone to obtain a video scene data set includes the following steps:
[0014] The video ringback tone to be processed is subjected to video object and human action recognition processing to obtain a first scene data set;
[0015] The video ringback tone to be processed is subjected to scene time span and screen dwell length recognition processing to obtain a second scene data set;
[0016] The video scene data set is obtained based on the first scene data set and the second scene data set.
[0017] In some embodiments, the process of performing industry feature recognition processing on the video scene data set to obtain industry feature information includes the following steps:
[0018] The video scene dataset is classified using a generative pre-trained model to obtain the main industry label data.
[0019] The primary industry tag data is further subdivided by level and type to obtain secondary industry tag data;
[0020] The industry characteristic information is obtained based on the main industry label data and the secondary industry label data.
[0021] In some embodiments, the step of performing text generation processing on the speech-text data set based on the video scene data set and the industry feature information to obtain the target video text includes the following steps:
[0022] Based on the industry characteristic information, the video scene data set is processed to generate scene text, resulting in a video scene text set.
[0023] The target video script is obtained by performing natural language optimization on the audio-text data set based on the video scene script set.
[0024] In some embodiments, the step of performing scene text generation processing on the video scene data set based on the industry feature information to obtain a video scene text set includes the following steps:
[0025] The primary industry label data and secondary industry label data are determined based on the industry characteristic information.
[0026] The set of industry copywriting factors is determined based on the main industry tag data and the secondary industry tag data;
[0027] Determine the first scene data set and the second scene data set based on the video scene data set;
[0028] The industry copywriting factor set is processed by similarity calculation based on the first scene data set to obtain video scene information;
[0029] The video scene information is segmented based on the second scene data set to obtain the video scene text set.
[0030] In some embodiments, after obtaining the target video text, the method further includes the following steps:
[0031] Based on the industry characteristic information, the target video text is subjected to background music matching processing to obtain the industry video background music;
[0032] The background audio of the industry video is processed by speech synthesis to obtain an audio file;
[0033] The video file and the audio file to be processed are combined to obtain the target video ringback tone.
[0034] To achieve the above objectives, another aspect of this application provides a video text processing apparatus, the apparatus comprising:
[0035] The first module is used to obtain the video ringback tone to be processed;
[0036] The second module is used to perform text conversion processing on the video ringback tone to be processed, and obtain a set of voice text data.
[0037] The third module is used to perform scene information analysis and processing on the video ringback tone to be processed, and to obtain a video scene data set.
[0038] The fourth module is used to perform industry feature recognition processing on the video scene data set to obtain industry feature information;
[0039] The fifth module is used to perform text generation processing on the voice and text data set based on the video scene data set and the industry feature information to obtain the target video text.
[0040] In some embodiments, the second module is configured to perform text conversion processing on the video ringback tone to be processed, obtaining a voice-text data set, including:
[0041] The first unit is used to perform background music filtering on the audio file of the video ringback tone to be processed, so as to obtain dry text audio.
[0042] The second unit is used to perform speech recognition processing on the dry text to obtain the speech text data set.
[0043] In some embodiments, the third module is used to perform scene information analysis and processing on the video ringback tone to be processed, to obtain a video scene data set, including:
[0044] The third unit is used to perform video object and human action recognition processing on the video ringback tone to be processed, and obtain the first scene data set;
[0045] The fourth unit is used to identify the scene time span and screen dwell length of the video ringback tone to be processed, and obtain the second scene data set;
[0046] The fifth unit is used to obtain the video scene data set based on the first scene data set and the second scene data set.
[0047] In some embodiments, the fourth module is configured to perform industry feature recognition processing on the video scene data set to obtain industry feature information, including:
[0048] The sixth unit is used to perform industry tag classification processing on the video scene data set through a generative pre-trained model to obtain the main industry tag data;
[0049] The seventh unit is used to perform level and type subdivision processing on the main industry label data to obtain secondary industry label data;
[0050] The eighth unit is used to obtain the industry characteristic information based on the main industry label data and the secondary industry label data.
[0051] In some embodiments, the fifth module is configured to perform text generation processing on the voice-text data set based on the video scene data set and the industry feature information to obtain the target video text, including:
[0052] The ninth unit is used to perform scene text generation processing on the video scene data set based on the industry characteristic information to obtain a video scene text set;
[0053] The tenth unit is used to perform natural language optimization processing on the voice text data set based on the video scene text set to obtain the target video text.
[0054] In some embodiments, the ninth unit is configured to perform scene text generation processing on the video scene data set based on the industry characteristic information to obtain a video scene text set, including:
[0055] The first subunit is used to determine the primary industry label data and the secondary industry label data based on the industry characteristic information.
[0056] The second subunit is used to determine the set of industry copywriting factors based on the main industry tag data and the secondary industry tag data;
[0057] The third subunit is used to determine the first scene data set and the second scene data set based on the video scene data set;
[0058] The fourth subunit is used to perform similarity calculation on the industry copywriting factor set based on the first scene data set to obtain video scene information;
[0059] The fifth subunit is used to segment the video scene information according to the second scene data set to obtain the video scene text set.
[0060] In some embodiments, the apparatus further includes a sixth module, comprising:
[0061] The eleventh unit is used to perform background music matching processing on the target video text based on the industry feature information to obtain the industry video background music;
[0062] The twelfth unit is used to perform speech synthesis processing on the background audio of the industry video to obtain a speech file;
[0063] The thirteenth unit is used to synthesize the video file and the audio file of the video ringback tone to obtain the target video ringback tone.
[0064] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.
[0065] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.
[0066] The embodiments of this application include at least the following beneficial effects: This application provides a video text processing method, apparatus, electronic device, and storage medium. This solution obtains a video scene data set by analyzing and processing scene information of the video ringback tone to be processed. This allows for analysis of the video scene content, facilitating subsequent generation of video text with industry tags and other data, thus improving the accuracy of video text generation. Furthermore, this solution performs industry feature recognition processing on the video scene data set to obtain industry feature information, enabling industry segmentation and association processing of the video ringback tone to be processed, making the industry tag data more complete and improving the efficiency of video text generation. In addition, this solution performs text generation processing on the voice text data set based on the video scene data set and industry feature information to obtain the target video text. This generates video text that conforms to the merchant's industry and closely matches the video scene, thereby enhancing the video promotion effect and improving the user experience. Attached Figure Description
[0067] Figure 1 This is a flowchart of a video text processing method provided in an embodiment of this application;
[0068] Figure 2 This is a schematic diagram illustrating an industry type distribution provided in an embodiment of this application;
[0069] Figure 3 yes Figure 1 The flowchart of step S102 in the document;
[0070] Figure 4 This is a schematic diagram of the framework of a recognition model provided in an embodiment of this application;
[0071] Figure 5 yes Figure 1 The flowchart of step S103 in the process;
[0072] Figure 6 yes Figure 1 The flowchart of step S104 in the process;
[0073] Figure 7 This is a schematic diagram illustrating the calculation of similarity between industry copywriting and video scenes provided in an embodiment of this application;
[0074] Figure 8 This is a flowchart illustrating a method for regenerating video, as provided in an embodiment of this application.
[0075] Figure 9 This is a schematic diagram of the structure of a video text processing device provided in an embodiment of this application;
[0076] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0077] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0078] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”
[0079] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.
[0080] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0081] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0082] 1) Video ringback tone: This is a video message that a user sees before the call is connected when making a voice or video call over a VoLTE network. Users can create or upload personalized video content or choose from the operator's video library. Different video content can also be set for different callers.
[0083] 2) Wav2Vec 2.0 model: This model utilizes self-supervised learning to extract information from audio. It can better understand the contextual information in audio and provides rich audio feature representations. It can be applied to various fields such as language learning, speech analysis, and audio indexing.
[0084] 3) PyTorchVideo Models: This is a deep learning library focused on video understanding. PyTorchVideo Models provides reusable, modular, and efficient components needed to accelerate video understanding research. It supports various deep learning video components, such as video models, video datasets, and video-specific transformations, enabling applications in various fields such as video action recognition, video classification, video augmentation, and video object recognition.
[0085] 4) Generative Pre-trained Transformer (GPT): Also known as a text-to-text model, it is a powerful language model that, through training on large amounts of text data, can generate high-quality, coherent, and context-sensitive text. GPT has a wide range of applications, from automatic content generation and dialogue systems to code generation.
[0086] In related technologies, there are two main methods for generating video ringback tones for commercial clients. One method involves generation through agency channels, using a general text scheme and video template, modifying only key information such as address and merchant name to create the video ringback tone. The other method involves the merchant directly providing the video. However, due to the significant differences in user groups across different industries, it is necessary to design copy tailored to the characteristics of each industry to attract user attention and generate demand. General text schemes cannot meet the needs of different industries. In addition, although clients understand the nature of their industry, they may lack professional copywriting skills, which may result in the video failing to accurately reflect industry characteristics, leading to a disconnect between the copy and video materials, an unsmooth copywriting style, and an incomplete promotional content.
[0087] In view of this, this application provides a video text processing method, apparatus, electronic device, and storage medium. This solution obtains a video scene data set by analyzing scene information of the video ringback tone to be processed. This allows for analysis of the video scene content, facilitating subsequent generation of video text with industry tags and other data, thus improving the accuracy of video text generation. Furthermore, this solution performs industry feature recognition processing on the video scene data set to obtain industry feature information, enabling industry segmentation and association processing of the video ringback tone to be processed, making the industry tag data more complete and improving the efficiency of video text generation. In addition, this solution performs text generation processing on the voice text data set based on the video scene data set and industry feature information to obtain the target video text. This generates video text that conforms to the merchant's industry and closely matches the video scene, thereby enhancing the video promotion effect and improving the user experience.
[0088] This application provides a video text processing method, relating to the field of computer technology, which can improve the quality of video ringback tones, highlight customer personality and characteristics, and thus enhance the effectiveness of video promotion. The solution utilizes speech-to-text recognition technology and a large language model to extract speech-to-text data sets from video files, enabling the extraction of merchant introduction information. It also leverages the PyTorchVideo tool model to analyze video scene information. By integrating video scene information, customer industry characteristics, and speech-to-text content, the solution optimizes the merchant introduction text using a text-to-text large model. This solution can save labor costs, improve efficiency, enhance text quality, and can be applied to industries such as catering, retail, and tourism, helping customers improve the effectiveness of video promotion. The video text processing method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal may be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited thereto; the server may be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server may also be a node server in a blockchain network; the software may be an application that implements a video text processing method, but is not limited to the above forms.
[0089] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0090] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0091] Figure 1 This is an optional flowchart of a video text processing method provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105.
[0092] Step S101: Obtain the video ringback tone to be processed;
[0093] Step S102: Perform text conversion processing on the video ringback tone to be processed to obtain a voice text data set;
[0094] Step S103: Perform scene information analysis and processing on the video ringback tone to be processed to obtain a video scene data set;
[0095] Step S104: Perform industry feature recognition processing on the video scene data set to obtain industry feature information;
[0096] Step S105: Perform text generation processing on the voice text data set based on the video scene data set and the industry feature information to obtain the target video text.
[0097] Steps S101 to S105, as illustrated in this embodiment, involve converting the video ringback tone to text to obtain a speech-text data set. This optimizes the original video text to obtain the target video script, thereby improving the quality of the video ringback tone. Specifically, the ringtone content of the video ringback tone to be processed is converted into text data, and key information is extracted to obtain the speech-text data set. Then, scene information analysis is performed on the video ringback tone to be processed. By recognizing the video ringback tone's frame, information such as objects, actions, people, and lighting in the scene is obtained, and the duration of object actions is recorded to determine the video scene data set. Please refer to [link to relevant documentation]. Figure 2 , Figure 2This is a schematic diagram illustrating the industry type distribution provided in this application embodiment, showing the distribution of video ringback tones across various industries, including IT / Internet, individual professions, transportation, and leisure / entertainment. Since the user groups of different industries vary greatly, it is necessary to design copywriting tailored to the characteristics of different industries. Therefore, this application embodiment uses industry feature recognition processing based on video scene data sets to obtain industry feature information, which can identify the customer's primary industry type and the probability of related industries. The identified industry feature information is then used to optimize the video copywriting. Based on the above information, a text generation model can be used to regenerate new solution text, resulting in the target video copywriting. This application embodiment optimizes video copywriting content through speech recognition, text generation, and other processing, saving labor costs, improving efficiency, and enhancing copywriting quality.
[0098] In step S101 of some embodiments, the video ringback tone to be processed can be obtained by the user uploading a video ringback tone, or it can be obtained by other means, such as obtaining a video from a video library as the video ringback tone to be processed, and is not limited thereto.
[0099] Please see Figure 3 In step S102 of some embodiments, the text conversion processing of the video ringback tone to be processed to obtain a voice-text data set includes the following steps:
[0100] Step S301: Perform background music filtering on the audio file of the video ringback tone to be processed to obtain dry text audio;
[0101] Step S302: Perform speech recognition processing on the dry text to obtain the speech text data set.
[0102] In this embodiment, by reading the audio file of the video ringback tone to be processed and performing text conversion processing, it is possible to analyze customer product information, business scope, and industry sector to obtain a set of audio-text data. This information can provide highly valuable reference for subsequent optimization of the audio solution and language refinement. Since the audio file of the video ringback tone to be processed typically contains both text and background music, to avoid interference from background music in the audio analysis, this embodiment will remove the background music and use only the text portion for analysis. Please refer to... Figure 4In this embodiment, an optimized wav2vec 2.0 speech recognition model is used to process the dry text speech to obtain the speech text data set. The original speech waveform is input into the speech recognition model, and processed by a convolutional layer (CNN) to obtain quantization and hidden layer representations. Then, it is processed by a masking module and a transformer module to obtain the context representation. Finally, the representation is learned by maximizing the similarity between positive examples and minimizing the similarity between negative examples to obtain the final contrastive loss. This application embodiment specifically extracts key content from the text portion of a video ringtone by reading the audio, thereby obtaining the key terms. This key content will be used for subsequent video industry analysis and text copy optimization. Simultaneously, this application embodiment saves specific key information, such as phone numbers, merchant addresses, merchant names, main products or services, etc., forming a voice-text data set. By performing text conversion processing on the video ringtone to be processed, this application embodiment enables the processed voice-text data set to be reflected in subsequent optimization processes, ensuring the integrity and effectiveness of the video information.
[0103] Please see Figure 5 In step S103 of some embodiments, the process of analyzing and processing scene information of the video ringback tone to obtain a video scene data set includes the following steps:
[0104] Step S501: Perform video object and human action recognition processing on the video ringback tone to be processed to obtain a first scene data set;
[0105] Step S502: Perform scene time span and screen dwell length recognition processing on the video ringback tone to be processed to obtain a second scene data set;
[0106] Step S503: Obtain the video scene data set based on the first scene data set and the second scene data set.
[0107] In this embodiment, the video ringback tone to be processed is input into the PyTorchVideo tool model. This model, through understanding and calculation, obtains the scene information and spatial temporal sequence of the video to obtain a video scene data set. Specifically, by identifying light, objects, people, and actions in the video frame, a series list containing specific item and action names is generated, and the data is saved as the first scene data set. For example, if the video frame sequentially shows a building, a room, bed sheets, room items, room lighting atmosphere, and finally sales information, this embodiment will record the continuous changes in the scene and the specific content of each object, and save it as the first scene data set. The first scene data set mainly stores the objects, people, actions, and events that appear in the video scene, and can correctly reflect the video's ownership, information, promotional content, and specific events through items and actions. Then, the video ringback tone to be processed is subjected to scene time span and screen dwell length recognition processing by recording the time span and screen dwell length of each scene appearing in the video frame. Different time lengths can reflect which content is highlighted and help determine the font length required for the promotional copy. For example, the video sequence might show a building (3 seconds) -> a room (8 seconds) -> bed sheets (3 seconds) -> room items (5 seconds) -> room lighting and ambiance (2 seconds) -> the final sale (4 seconds). This embodiment of the application treats this time information as the time attribute of the scene and saves it to a second scene data set. The second scene data set stores the duration of each type of item's actions, and length analysis can determine the key content to be promoted in the video, the frequency of scene transitions, etc. This embodiment of the application provides a data foundation for subsequent video analysis by analyzing the scene information of the video ringback tone to be processed.
[0108] Please see Figure 6 In step S104 of some embodiments, the process of performing industry feature recognition processing on the video scene data set to obtain industry feature information includes the following steps:
[0109] Step S601: The video scene data set is classified by industry tags using a generative pre-trained model to obtain the main industry tag data;
[0110] Step S602: Subdivide the main industry label data by level and type to obtain secondary industry label data;
[0111] Step S603: Obtain the industry feature information based on the main industry label data and the secondary industry label data.
[0112] In this embodiment, a generative pre-trained model is used to classify video scene data into industry labels to obtain primary industry label data. This generative pre-trained model is a pre-trained GPT model. The training method can be training through extensive industry feature recognition and manual error correction, with the main training direction set as the sales industry. Since each industry has its own primary and secondary features, this embodiment calculates industry probabilities by comparing primary features and merging secondary features. For example, if a video shows bed sheets, it can be initially identified as a hotel, bedding store, supermarket, or real estate. Further comparison of the specific content next to the bed sheets, such as corridors or restrooms, further increases the probability of it being a hotel, while excluding bedding stores and supermarkets. This embodiment can calculate label classification using the GPT model, outputting industry label division probabilities as: 70% hotel industry, 20% clothing industry, and 10% construction industry. This embodiment can also calculate by comparing a lower threshold limit, i.e., setting a threshold and comparing it with the maximum probability value. For example, setting a threshold of 60% can determine the hotel industry as the customer's primary industry label, with other probabilities considered as related industries. This application embodiment can further subdivide the main industry tag data by level and type. Based on the main industry information from the previous step, it further subdivides the same industry by level and type. For example, the hotel industry can be subdivided into budget hotels, business hotels, tourist hotels, hot spring hotels, and themed hotels. By combining voice text information and video scene information, and then calculating using the GPT model, and comparing the subdivided features of the same industry, the specific subdivided type probabilities are obtained as follows: 80% business hotels, 12% tourist hotels, and 8% themed hotels. Thus, business hotels are identified as the customer's secondary industry information, and the others are identified as related industries. This application embodiment obtains industry feature information by performing industry feature recognition processing on the video scene data set, which can determine the industry types involved in the business scope of the video ringback tone to be processed, providing a basis for subsequent calculation of the main customer groups targeted by the industry. This ensures that the customer's industry information and areas of involvement are accurate and comprehensive, thereby generating more accurate video scripts and avoiding omissions.
[0113] In step S105 of some embodiments, the step of performing text generation processing on the voice-text data set based on the video scene data set and the industry feature information to obtain the target video text includes the following steps:
[0114] Based on the industry characteristic information, the video scene data set is processed to generate scene text, resulting in a video scene text set.
[0115] The target video script is obtained by performing natural language optimization on the audio-text data set based on the video scene script set.
[0116] In this embodiment, because the target audience for video ringback tones varies across industries, with differences in income levels, education levels, and hobbies, different copywriting content needs to be developed for different industries. Furthermore, the product and service types of each industry also differ, requiring the use of specialized vocabulary to accurately describe the products and services featured in the video ringback tones. Therefore, different vocabulary needs to be selected based on different industries, and scene text generation processing is performed on the video scene data set based on industry characteristic information to obtain a set of video scene copywriting. Then, the video scene copywriting set is optimized through verb and adjective modification, word replacement, and other optimization processes. The length of the copywriting is also limited to a minimum character limit. Simultaneously, natural language optimization processing is performed on the voice text data set, ensuring that the third data set remains in the same position within the newly generated text scheme scene, thus generating new copywriting content. This embodiment can use the GPT model for natural language optimization, making the text more natural, elegant, and fluent while maintaining the overall number and length of characters. This embodiment obtains the target video copywriting by performing text generation processing on the voice text data set based on the video scene data set and industry characteristic information. This results in a copywriting that covers all information about the service content, highlights featured services, and improves the efficiency of video copywriting generation and promotional effectiveness.
[0117] In some embodiments, the step of performing scene text generation processing on the video scene data set based on the industry feature information to obtain a video scene text set includes the following steps:
[0118] The primary industry label data and secondary industry label data are determined based on the industry characteristic information.
[0119] The set of industry copywriting factors is determined based on the main industry tag data and the secondary industry tag data;
[0120] Determine the first scene data set and the second scene data set based on the video scene data set;
[0121] The industry copywriting factor set is processed by similarity calculation based on the first scene data set to obtain video scene information;
[0122] The video scene information is segmented based on the second scene data set to obtain the video scene text set.
[0123] In this embodiment, primary industry tag data and secondary industry tag data are obtained from industry characteristic information. By selecting a primary industry tag data, a set of special copywriting factors based on the primary industry type can be obtained. For example, based on the hotel industry, a set of special copywriting factors including words such as "clean," "grand," and "high cost-performance ratio" can be obtained. By selecting a secondary industry tag data, a set of special copywriting factors based on the secondary industry type can be obtained. An industry copywriting factor set is obtained based on the combination of the primary and secondary industry tag data. A first scene data set and a second scene data set can be obtained from the video scene data set. Please refer to [link to relevant documentation]. Figure 7 Using the industry copywriting factor set as the x-axis and the first scenario dataset as the y-axis, a user similarity calculation coordinate system is constructed. Based on the discrete distribution of data on the x-axis and y-axis coordinate systems, the similarity between the two sets of feature factors is calculated. By comparing the similarity values, a higher similarity indicates that the copywriting in the industry copywriting factor set is more consistent with the video scene information V(X). The expression for the video scene information is as follows:
[0124]
[0125] In the formula, A represents any copywriting factor in the industry copywriting factor set, and B represents the scene data factor in the first scene data set. This application embodiment uses the copywriting with the highest calculated video scene information as the key copywriting content of the video scene. Then, based on the main objects and time span appearing in the second scene data set, the key copywriting content of the video is divided into one-sentence promotional content of different lengths. The number of words in the content is determined by the time span of the main object's appearance, thus calculating the video scene copywriting set for each segment of main promotional content and word count limit. This application embodiment obtains the video scene copywriting set by performing scene text generation processing on the video scene data set based on industry characteristic information. This allows for the development of personalized text solutions for different industries and different video content, thereby more effectively promoting the service content, service scope, service features, and merchant advantages of video ringback tones.
[0126] In some embodiments, after obtaining the target video text, the method further includes the following steps:
[0127] Based on the industry characteristic information, the target video text is subjected to background music matching processing to obtain the industry video background music;
[0128] The background audio of the industry video is processed by speech synthesis to obtain an audio file;
[0129] The video file and the audio file to be processed are combined to obtain the target video ringback tone.
[0130] In this embodiment, background music matching can be performed on the target video text based on industry characteristic information to select industry-specific background music. This embodiment can also select suitable professional voice broadcasters from a broadcaster database and generate audio files using multi-tone speech synthesis technology. Finally, the video file to be processed and the newly synthesized audio text are combined to regenerate a new video file, which can then be loaded to obtain the new video ringback tone, completing the update. This embodiment can generate new video ringback tones, improving the personalized display of video ringback tones and enhancing the user experience.
[0131] The solutions of this application embodiment will be described in detail and explained below with reference to specific application examples:
[0132] This application's embodiments can be applied to the computer field, enabling the automatic generation of introductory text for videos or video ringback tones for users, and also allowing the automatic generation of video ringback tones based on video text. Please refer to... Figure 8 This application embodiment converts the ringtone file provided by the customer into text, extracting the ringtone text content. Based on the video ringtone uploaded by the customer, PyTorchVisual is used to obtain the customer's video scene information and industry. Then, combining the video scene and ringtone text content with industry characteristic information, the customer's industry copywriting content is derived. Finally, the optimized text copy is used to generate speech from a suitable anchor, synthesizing an audio file to regenerate the video. This application embodiment can effectively mine promotional copywriting that matches the merchant's industry and closely resembles the video scene. By using intelligent speech generation methods to cover industry characteristics and generate personalized video copywriting, it reduces manual processing costs, improves customer promotional effectiveness, and thus increases conversion rates.
[0133] Please see Figure 9 This application also provides a video text processing apparatus that can implement the above-described video text processing method. The apparatus includes:
[0134] The first module 901 is used to acquire the video ringback tone to be processed;
[0135] The second module 902 is used to perform text conversion processing on the video ringback tone to be processed, and obtain a set of voice text data.
[0136] The third module 903 is used to perform scene information analysis and processing on the video ringback tone to be processed, and to obtain a video scene data set.
[0137] The fourth module 904 is used to perform industry feature recognition processing on the video scene data set to obtain industry feature information;
[0138] The fifth module 905 is used to perform text generation processing on the voice text data set based on the video scene data set and the industry feature information to obtain the target video text.
[0139] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0140] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described video text processing method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0141] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0142] Please see Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:
[0143] The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0144] The memory 1002 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 using the video text processing method of the embodiments of this application.
[0145] Input / output interface 1003 is used to implement information input and output;
[0146] The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0147] Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004);
[0148] The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0149] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video text processing method.
[0150] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0151] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0152] This application provides a video text processing method, apparatus, electronic device, and storage medium. It analyzes and processes scene information of the video ringback tone to obtain a video scene data set, enabling analysis of video scene content and facilitating subsequent generation of video text with industry tags and other data, thus improving the accuracy of video text generation. Furthermore, the solution performs industry feature recognition processing on the video scene data set to obtain industry feature information, allowing for industry segmentation and association processing of the video ringback tone, making the industry tag data more complete and improving the efficiency of video text generation. Additionally, the solution performs text generation processing on the voice text data set based on the video scene data set and industry feature information to obtain the target video text, generating video text that matches the merchant's industry and closely relates to the video scene, thereby enhancing the video promotion effect and improving the user experience.
[0153] This application embodiment constructs a model that distinguishes industry information characteristics and the merchant's target user group based on the industry information of the video ringback tone. It can combine the scene information of the customer-uploaded video with video text copy optimization calculations to generate an audio file that integrates the video context and recommends it to the customer. This solves the problem in existing technologies that cannot accurately generate copy and synthesize audio by combining video scenes and industry characteristics. This application embodiment, through intelligent optimization of the voice copy, can better introduce and promote customer products and services, ensuring the personalization, automation, and rationality of the generated copy. Furthermore, this application embodiment can effectively utilize feedback information from other merchants' industries for video ringback tone generation, accelerating the personalized generation effect and improving conversion rates.
[0154] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0155] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0157] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0158] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0159] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0160] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0161] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0162] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0163] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0164] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for processing video text, characterized in that, The method includes the following steps: Get the ringback tone of the video to be processed; The video ringback tone to be processed is converted into text to obtain a set of voice and text data. The video ringback tone to be processed is subjected to scene information analysis and processing to obtain a video scene data set; The video scene data set is subjected to industry feature recognition processing to obtain industry feature information; Based on the video scene data set and the industry feature information, the voice text data set is processed to generate text, thereby obtaining the target video text; The step of generating text from the speech-text data set based on the video scene data set and the industry feature information to obtain the target video text includes the following steps: Based on the industry characteristic information, the video scene data set is processed to generate scene text, resulting in a video scene text set. The target video text is obtained by performing natural language optimization on the voice text data set based on the video scene text set. The step of generating scene text from the video scene data set based on the industry characteristic information to obtain a video scene text set includes the following steps: The primary industry label data and secondary industry label data are determined based on the industry characteristic information. The set of industry copywriting factors is determined based on the main industry tag data and the secondary industry tag data; Determine the first scene data set and the second scene data set based on the video scene data set; The industry copywriting factor set is processed by similarity calculation based on the first scene data set to obtain video scene information; The video scene information is segmented based on the second scene data set to obtain the video scene text set.
2. The method according to claim 1, characterized in that, The text conversion process for the video ringback tone to be processed, resulting in a voice-text data set, includes the following steps: The audio file of the video ringback tone to be processed is subjected to background music filtering to obtain dry text audio; The dry text is processed by speech recognition to obtain the speech-text data set.
3. The method according to claim 1, characterized in that, The process of analyzing and processing scene information of the video ringback tone to obtain a video scene data set includes the following steps: The video ringback tone to be processed is subjected to video object and human action recognition processing to obtain a first scene data set; The video ringback tone to be processed is subjected to scene time span and screen dwell length recognition processing to obtain a second scene data set; The video scene data set is obtained based on the first scene data set and the second scene data set.
4. The method according to claim 1, characterized in that, The process of performing industry feature recognition processing on the video scene data set to obtain industry feature information includes the following steps: The video scene dataset is classified using a generative pre-trained model to obtain the main industry label data. The primary industry tag data is further subdivided by level and type to obtain secondary industry tag data; The industry characteristic information is obtained based on the main industry label data and the secondary industry label data.
5. The method according to any one of claims 1 to 4, characterized in that, After obtaining the target video text, the method further includes the following steps: Based on the industry characteristic information, the target video text is subjected to background music matching processing to obtain the industry video background music; The background audio of the industry video is processed by speech synthesis to obtain an audio file; The video file and the audio file to be processed are combined to obtain the target video ringback tone.
6. A video text processing device, characterized in that, The device includes: The first module is used to obtain the video ringback tone to be processed; The second module is used to perform text conversion processing on the video ringback tone to be processed, and obtain a set of voice text data. The third module is used to perform scene information analysis and processing on the video ringback tone to be processed, and to obtain a video scene data set. The fourth module is used to perform industry feature recognition processing on the video scene data set to obtain industry feature information; The fifth module is used to perform text generation processing on the voice and text data set based on the video scene data set and the industry feature information to obtain the target video text. The fifth module is used to perform text generation processing on the voice-text data set based on the video scene data set and the industry feature information to obtain the target video text, including: Based on the industry characteristic information, the video scene data set is processed to generate scene text, resulting in a video scene text set. The target video text is obtained by performing natural language optimization on the voice text data set based on the video scene text set. The step of performing scene text generation processing on the video scene data set based on the industry characteristic information to obtain a video scene text set includes: The primary industry label data and secondary industry label data are determined based on the industry characteristic information. The set of industry copywriting factors is determined based on the main industry tag data and the secondary industry tag data; Determine the first scene data set and the second scene data set based on the video scene data set; The industry copywriting factor set is processed by similarity calculation based on the first scene data set to obtain video scene information; The video scene information is segmented based on the second scene data set to obtain the video scene text set.
7. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Processing method and device, electronic equipment and medium
CN114697760A