Video polyphonic ringtone generation method and device, equipment and storage medium
Through the deep learning model combining user data and real-time information to generate personalized video ringtones, the problem of lack of personalized and dynamic adaptation of video ringtones in the existing technology is solved, and the user experience and content customization is improved.
Patent Information
- Application Number
- CN202510651761.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-22
AI Technical Summary
The existing video ringtone technology lacks personalized and dynamic adaptability, has low generation efficiency and single content, which makes it difficult to meet the personalized needs of users, and has a general user experience.
By obtaining the basic video data, festival information, user tag information and hotspot information uploaded by users, using the deep learning model to generate personalized video data, and combining blessing information and background music for multi-modal fusion to generate personalized video ringtones.
Improve the customization and personalization of video ringtones, improve user experience, reduce user editing operations, and realize real-time dynamic adaptation and personalized generation of content.
Smart Images

Figure CN120529017A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video ringback tone, and in particular to a method, apparatus, device and storage medium for generating a video ringback tone. Background Art
[0002] With the advancement of information technology, more and more applications are being developed to provide services to users. For example, video ringback tone (CRBT) is a value-added service based on communication networks. When a user makes a voice call, it plays a short video to the called party, replacing the traditional ringback tone. When the caller calls the called party, while waiting on hold, they can view personalized video content pre-set by the called party, such as personal videos, corporate promotions, or holiday greetings.
[0003] In related technologies, traditional video ringtones mainly rely on static preset content or simple template generation, lacking personalization and dynamic adaptation capabilities. Users can only select content from a limited preset library. There are generally problems such as low generation efficiency, single content, and lack of real-time data drive. It is difficult to meet users' personalized needs, and users' actual usage experience is average.
[0004] In summary, the problems existing in related technologies need to be solved urgently. Summary of the Invention
[0005] The purpose of this application is to solve one of the technical problems existing in the related art to at least a certain extent.
[0006] To this end, an object of the embodiments of the present application is to provide a method, apparatus, device and storage medium for generating a video ringback tone.
[0007] In order to achieve the above technical objectives, the technical solutions adopted in the embodiments of the present application include:
[0008] In one aspect, an embodiment of the present application provides a method for generating a video ringback tone, the method comprising:
[0009] In response to a user's video ringback tone production request, obtaining basic video data uploaded by the user;
[0010] Obtaining festival information at the current time, user tag information of the user, and hot spot information at the current time;
[0011] Using the festival information, the user tag information, and the hot spot information as tags, a deep learning model is used to generate corresponding personalized video data based on the basic video data;
[0012] Generate corresponding greeting information according to the festival information, and generate background music corresponding to the user according to the user tag information;
[0013] The personalized video data, the blessing information and the background music are multimodally fused to obtain a video ringback tone corresponding to the user.
[0014] In addition, the method for generating a video ringback tone according to the above embodiment of the present application may also have the following additional technical features:
[0015] Furthermore, in one embodiment of the present application, before using the holiday information, the user tag information, and the hot spot information as tags, the method further includes:
[0016] Preprocessing the festival information, the user tag information, and the hot spot information;
[0017] The preprocessing includes at least one of data cleaning processing, format conversion processing, context supplementation processing, and vectorization processing.
[0018] Furthermore, in one embodiment of the present application, generating corresponding greeting information according to the holiday information includes:
[0019] Randomly extracting a predetermined number of festival keywords from a pre-established keyword database based on the festival information;
[0020] The holiday keywords and preset prompts are input into a blessing generation model, and the corresponding blessing information is generated by the blessing generation model; wherein the preset prompts are used to instruct the blessing generation model to generate blessing information based on the holiday keywords.
[0021] Furthermore, in one embodiment of the present application, the blessing generation model is built based on the GPT model or the Deepseek model.
[0022] Furthermore, in one embodiment of the present application, the multimodal fusion of the personalized video data, the blessing information, and the background music includes:
[0023] Using the audio stream of the background music as a reference time axis, aligning the image frames of the personalized video data and the display timing corresponding to the greeting information to the same time coordinate system;
[0024] Detecting a beat point in the background music, and determining the beat point as a synchronization anchor point for switching the image frame or the blessing message;
[0025] The background music and the personalized video data are synchronized with music animation, and the blessing information and the personalized video data are adapted to text scenes.
[0026] Furthermore, in one embodiment of the present application, the text scene adaptation of the blessing information and the personalized video data includes:
[0027] Detecting a core area of an image frame of the personalized video data by an image recognition algorithm;
[0028] The display position of the blessing information in the image frame is dynamically determined based on the core area.
[0029] Furthermore, in one embodiment of the present application, the method further includes:
[0030] Obtaining predetermined channel setting parameters;
[0031] Determining specification restriction information corresponding to the video ringback tone according to the channel setting parameters;
[0032] The video ring back tone is compressed according to the specification restriction information, and the compressed video ring back tone is sent to the ring back tone platform corresponding to the user.
[0033] On the other hand, an embodiment of the present application provides a device for generating a video ringback tone, the device comprising:
[0034] A response unit, configured to respond to a user's video ringback tone production request and obtain basic video data uploaded by the user;
[0035] An acquisition unit, configured to acquire festival information at a current time point, user tag information of the user, and hot spot information at the current time point;
[0036] A first generating unit is configured to use the holiday information, the user tag information, and the hot spot information as tags to generate corresponding personalized video data based on the basic video data through a deep learning model;
[0037] A second generating unit is configured to generate corresponding greeting information according to the festival information, and generate background music corresponding to the user according to the user tag information;
[0038] The fusion unit is used to perform multimodal fusion on the personalized video data, the blessing information and the background music to obtain the video ringtone corresponding to the user.
[0039] In another aspect, an embodiment of the present application provides an electronic device, including:
[0040] at least one processor;
[0041] at least one memory for storing at least one program;
[0042] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned method for generating a video ringback tone.
[0043] On the other hand, an embodiment of the present application further provides a computer-readable storage medium storing a program executable by a processor. When the program executable by the processor is executed, it is used to implement the above-mentioned method for generating a video ringback tone.
[0044] The advantages and benefits of this application will be partially given in the following description, and partially become apparent from the following description, or learned through practice of this application:
[0045] The embodiment of the present application discloses a method, device, equipment and storage medium for generating a video ringback tone. In response to a user's video ringback tone production request, the method obtains the basic video data uploaded by the user; obtains the festival information at the current time point, the user's user tag information and the hot information at the current time point; uses the festival information, the user tag information and the hot information as tags, and generates corresponding personalized video data based on the basic video data through a deep learning model; generates corresponding blessing information according to the festival information, and generates background music corresponding to the user according to the user tag information; performs multimodal fusion on the personalized video data, the blessing information and the background music to obtain the video ringback tone corresponding to the user. The method can combine the festival information and hot information when the user initiates the video ringback tone production request with the user tag information to generate personalized video data, and further optimize the personalized video data to obtain the corresponding video ringback tone, which can improve the customization and personalization of the video ringback tone without the user having to perform complex editing operations, which is conducive to improving the user's video ringback tone usage experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following introduction is made to the drawings of the embodiments of the present application or the related technical solutions in the prior art. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solutions of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without any creative work.
[0047] Figure 1 A schematic diagram of an implementation environment of a method for generating a video ringback tone provided in an embodiment of the present application;
[0048] Figure 2 A flowchart of a method for generating a video ringback tone provided in an embodiment of the present application;
[0049] Figure 3 A schematic diagram of the structure of a device for generating a video ringback tone provided in an embodiment of the present application;
[0050] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The present application is further described below in conjunction with the accompanying drawings and specific embodiments. The described embodiments should not be considered as limiting the present application. All other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0052] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0054] 1) Video Ringback Tone: refers to a video played instead of the traditional audio ringback tone during a mobile phone call. It is usually a short video content selected by the ringback tone user.
[0055] 2) Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0056] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0057] 3) Machine Learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.
[0058] 4) Deep learning is a subfield of machine learning that simulates the way the human brain processes information by building and training multi-layer neural networks. Deep learning models can automatically extract complex features from large amounts of data and apply them to a variety of tasks, such as image recognition, natural language processing, speech recognition, and recommendation systems.
[0059] 5) Image recognition algorithms are a core task in computer vision, enabling computers to understand and interpret image content. These algorithms range from traditional methods to deep learning and are widely used in scenarios such as object detection, classification, and segmentation.
[0060] With the advancement of information technology, more and more applications are being developed to provide services to users. For example, video ringback tone (CRBT) is a value-added service based on communication networks. When a user makes a voice call, it plays a short video to the called party, replacing the traditional ringback tone. When the caller calls the called party, while waiting on hold, they can view personalized video content pre-set by the called party, such as personal videos, corporate promotions, or holiday greetings.
[0061] In related technologies, traditional video ringtones mainly rely on static preset content or simple template generation, lacking personalization and dynamic adaptation capabilities. Users can only select content from a limited preset library. There are generally problems such as low generation efficiency, single content, and lack of real-time data drive. It is difficult to meet users' personalized needs, and users' actual usage experience is average.
[0062] In view of this, an embodiment of the present application provides a method for generating a video ringback tone, which responds to a user's video ringback tone production request, obtains the basic video data uploaded by the user; obtains the holiday information at the current time point, the user's user tag information, and the hot spot information at the current time point; uses the holiday information, the user tag information, and the hot spot information as tags, and generates corresponding personalized video data based on the basic video data through a deep learning model; generates corresponding blessing information based on the holiday information, and generates background music corresponding to the user based on the user tag information; performs multimodal fusion on the personalized video data, the blessing information, and the background music to obtain the video ringback tone corresponding to the user. This method can combine the holiday information and hot spot information when the user initiates the video ringback tone production request, and combine the user tag information to generate personalized video data, and further optimize the personalized video data to obtain the corresponding video ringback tone, which can improve the customization and personalization of the video ringback tone, and does not require the user to perform complex editing operations, which is conducive to improving the user's video ringback tone usage experience.
[0063] Please refer to Figure 1 , Figure 1 The following is a schematic diagram showing an implementation environment of a method for generating a video ringback tone provided in an embodiment of the present application. In this implementation environment, the main hardware and software components involved include a terminal device 110 and a backend server 120. The terminal device 110 and the backend server 120 are in communication connection with each other.
[0064] Specifically, a method for generating a video ringback tone provided in an embodiment of the present application can be executed solely on the terminal device 110 side, or solely on the background server 120 side, or based on data interaction between the terminal device 110 and the background server 120.
[0065] The terminal device 110 in the above embodiment may include, but is not limited to, a mobile phone, a computer, a smart wearable device, a PDA device, an intelligent voice interaction device, a smart home appliance, an in-vehicle terminal, etc. The backend server 120 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0066] The terminal device 110 and the backend server 120 may establish a communication connection via a wireless network or a wired network. The wireless network or wired network uses standard communication technologies and / or protocols, and the network may be the Internet or any other network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network.
[0067] Of course, it is understandable that Figure 1 The implementation environment in the embodiment of the present application is only a method for generating a video ring back tone provided in some optional application scenarios, and the actual application is not fixed. Figure 1 The hardware and software environment shown.
[0068] Next, in conjunction with the introduction to the aforementioned implementation environment, a method for generating a video ringback tone provided in an embodiment of the present application is introduced and explained.
[0069] Please refer to Figure 2 , Figure 2 : is a schematic diagram of a method for generating a video ringback tone provided by an embodiment of the present application, the method for generating a video ringback tone including but not limited to:
[0070] Step 210: In response to the user's video ringback tone production request, obtain the basic video data uploaded by the user;
[0071] Step 220: In response to the user's video ringback tone production request, obtain the basic video data uploaded by the user;
[0072] Step 230: Using the holiday information, the user tag information, and the hot spot information as tags, generate corresponding personalized video data based on the basic video data through a deep learning model;
[0073] Step 240: Generate corresponding greeting information based on the holiday information, and generate background music corresponding to the user based on the user tag information;
[0074] Step 250: Perform multimodal fusion on the personalized video data, the blessing information, and the background music to obtain a video ringback tone corresponding to the user.
[0075] In an embodiment of the present application, a method for generating a video ringback tone is provided. The method can combine holiday information and hot spot information when a user initiates a video ringback tone production request with user tag information to generate personalized video data, and further optimize the personalized video data to obtain a corresponding video ringback tone. This method can improve the customization and personalization of the video ringback tone without requiring the user to perform complex editing operations, which is conducive to improving the user's video ringback tone usage experience.
[0076] The embodiments of this application aim to at least partially address the problems of insufficient personalization, poor real-time performance, and a limited interactive experience in existing video ringback tone technologies. By building an intelligent generation system based on AI and real-time data, the following innovative goals are achieved:
[0077] First, it breaks through the limitations of traditional static templates of video ringback tone, and uses multimodal technology to dynamically generate personalized video content that deeply matches user tag information (such as interest preferences, social behavior) and real-time scenarios (such as festivals, hot events). Second, through edge computing and 5G network optimization, it solves the low-latency requirements of content generation and distribution, ensuring real-time rendering and second-level push of video ringback tone. In addition, it designs a multimodal fusion algorithm to intelligently combine elements such as text, images, and audio to enhance the fun and interactivity of ringback tone. Finally, it provides operators and Internet platforms with a set of low-cost, high-efficiency video ringback tone value-added service solutions, which not only improves user experience but also expands commercial application scenarios, and promotes the technological integration and industrial upgrading of communication services and AI-generated content (AIGC).
[0078] Specifically, in an embodiment of the present application, when generating a personalized video ringback ring, a user can initiate a video ringback ring production request. This request can be triggered by the user clicking a related icon button, such as by providing a function entry for video ringback ring production on a front-end client interface or H5 interface, through which the user can initiate the video ringback ring production request. In an embodiment of the present application, when initiating a video ringback ring production request, the user can upload basic video data. The basic video data can be the original material for video ringback ring production. Subsequently, a video ringback ring related to the user's personalization will be generated based on the basic video data.
[0079] When generating a video ringback tone corresponding to the user, the festival information at the current time point, the user tag information of the user, and the hot information at the current time point can be obtained. Among them, festival information and hot information can be obtained through distributed crawler technology. For example, the corresponding festival information is determined based on calendar data, and hot information can be obtained through news data from news media. In the embodiment of the present application, hot information may include but is not limited to news data such as festival customs, social anecdotes, celebrity dynamics, and sports events. It should be noted that in the embodiment of the present application, when obtaining hot information, it can be obtained based on the user's current location, so that the obtained hot information is more in line with the user and better meets and adapts to the user's personalized needs.
[0080] For user tag information, a multi-dimensional user tag system (e.g., "music lover + football fan") can be constructed by integrating multiple sources of information such as operator CRM data, social behavior logs, and call records using a clustering algorithm to extract and store each user's user tag information. It is understood that user tag information can record user preferences. In the embodiments of this application, user tag information can be updated in real time, providing accurate personalized input for AI content generation and supporting dynamic identification of content needs in specific scenarios (e.g., birthdays and holidays).
[0081] After obtaining holiday information, user tag information, and hot spot information, these can be used as label data and input into a deep learning model to generate personalized video data. In the embodiments of this application, the relevant tasks are handled by a deep learning multimodal content generation engine. The deep learning multimodal content generation engine is the core component of the application, responsible for converting inputs such as user data and real-time scene information into high-quality, personalized video ringtone content. The deep learning model is part of the deep learning multimodal content generation engine.
[0082] Specifically, the functions of the deep learning multimodal content generation engine may include but are not limited to the following:
[0083] Cross-modal content understanding and alignment: Through a unified semantic encoder, the input data (such as holiday information and user tag information) is mapped into a vector representation, and contrastive learning is used to achieve semantic matching of text, images, and audio (for example, "Spring Festival" corresponds to lantern animation + festive music), ensuring high consistency of multimodal content in theme, style, and timing.
[0084] Dynamic Content Generation: As the core component of this system, deep learning models leverage multimodal AI technologies to achieve real-time conversion from raw data to personalized video data. Using a parallel processing architecture, it simultaneously runs text generation (using an optimized DeepSeek model to output greetings), image rendering (using an improved Stable Diffusion to generate scenes and character animations), and audio synthesis (using an enhanced WaveNet architecture to create speech and music), while incorporating real-time data streams (such as user location) for dynamic adjustments.
[0085] Multimodal Fusion and Synchronization: As a key component of content generation, the module utilizes dynamic time warping (DTW) algorithms and spatial adaptation technology to achieve precise spatiotemporal alignment of text, images, audio, and other elements. This module establishes a unified timeline coordinate system to ensure strict synchronization between voice announcements and character lip movements, music tempo, and animation effects (such as fireworks). It also intelligently adjusts text positioning to avoid visual occlusion (e.g., automatically avoiding facial areas). An adaptive rendering engine based on terminal screen characteristics (resolution, aspect ratio) optimizes the display parameters of each modal element in real time, ensuring smooth playback of generated videos even on low-end devices.
[0086] In this embodiment of the application, in addition to generating corresponding personalized video data based on a deep learning model, corresponding greeting information is also generated based on holiday information, and background music corresponding to the user is generated based on user tag information. Here, the greeting information can be text information related to the current holiday information, and the background music can be used in the personalized video data. Then, the personalized video data, the greeting information, and the background music can be multimodally fused to obtain the user's corresponding video ringtone.
[0087] It can be understood that a method for generating a video ringback tone provided in an embodiment of the present application, in response to a user's video ringback tone production request, obtains the basic video data uploaded by the user; obtains the festival information at the current time point, the user tag information of the user, and the hot information at the current time point; uses the festival information, the user tag information, and the hot information as tags, and generates corresponding personalized video data based on the basic video data through a deep learning model; generates corresponding blessing information according to the festival information, and generates background music corresponding to the user according to the user tag information; performs multimodal fusion on the personalized video data, the blessing information, and the background music to obtain the video ringback tone corresponding to the user. This method can combine the festival information and hot information when the user initiates the video ringback tone production request, and combine the user tag information to generate personalized video data, and further optimize the personalized video data to obtain the corresponding video ringback tone, which can improve the customization and personalization of the video ringback tone, and does not require the user to perform complex editing operations, which is conducive to improving the user's video ringback tone usage experience.
[0088] Specifically, in some embodiments, before using the holiday information, the user tag information, and the hot spot information as tags, the method further includes:
[0089] Preprocessing the festival information, the user tag information, and the hot spot information;
[0090] The preprocessing includes at least one of data cleaning processing, format conversion processing, context supplementation processing, and vectorization processing.
[0091] In an embodiment of the present application, holiday information, user tag information, and hot spot information may be preprocessed before being input as labels into a deep learning model. Here, the preprocessing operation may include, but is not limited to, at least one of data cleaning processing, format conversion processing, context supplementation processing, and vectorization processing. Data cleaning processing can remove noise, correct errors, and ensure label quality; format conversion processing can unify the label format and reduce ambiguity; context supplementation processing can be used to supplement implicit information and enhance label semantics; vectorization processing can convert labels into numerical features for model understanding. It is understandable that the actual preprocessing operation may also include other processes, and this application does not limit this.
[0092] Specifically, in some embodiments, generating corresponding greeting information according to the holiday information includes:
[0093] Randomly extracting a predetermined number of festival keywords from a pre-established keyword database based on the festival information;
[0094] The holiday keywords and preset prompts are input into a blessing generation model, and the corresponding blessing information is generated by the blessing generation model; wherein the preset prompts are used to instruct the blessing generation model to generate blessing information based on the holiday keywords.
[0095] In an embodiment of the present application, when generating corresponding greeting information based on holiday information, a predetermined number of holiday keywords can be randomly extracted from a pre-established keyword database. The predetermined number here can be one or more, and the present application does not impose any restrictions on this. Then, the holiday keywords and preset prompts can be input into the greeting generation model, and the corresponding greeting information can be generated by the greeting generation model. Here, the greeting generation model can be a large language model, such as one built based on a GPT model or a Deepseek model. The preset prompt is used to prompt the greeting generation model to process related tasks, that is, to generate greeting information based on holiday keywords. It may include content such as "Please generate greeting information based on the input holiday keywords, and the style of the greeting information is the style of ancient poetry", and the present application does not impose any restrictions on this.
[0096] Specifically, in some embodiments, the multimodal fusion of the personalized video data, the blessing information, and the background music includes:
[0097] Using the audio stream of the background music as a reference time axis, aligning the image frames of the personalized video data and the display timing corresponding to the greeting information to the same time coordinate system;
[0098] Detecting a beat point in the background music, and determining the beat point as a synchronization anchor point for switching the image frame or the blessing message;
[0099] The background music and the personalized video data are synchronized with music animation, and the blessing information and the personalized video data are adapted to text scenes.
[0100] In an embodiment of the present application, when performing multimodal fusion of personalized video data, greeting information, and background music, the audio stream of the background music can be used as the reference time axis to align the display timing of the image frames of the personalized video data and the corresponding greeting information to the same time coordinate system. In addition, beat points can be detected in the audio waveform of the background music as synchronization anchor points for the image frame / greeting information transformation. Subsequently, music animation synchronization is performed on the background music and personalized video data, and text scene adaptation is performed on the greeting information and personalized video data.
[0101] Specifically, in an embodiment of the present application, when performing music animation synchronization on background music and personalized video data, FFT technology can be used to analyze the spectrum energy of the background music, extract strong rhythm points, and control the release timing of the particle system (such as fireworks special effects) to coincide with the rhythm points. When performing text scene adaptation on the blessing information and personalized video data, relevant image detection algorithms (such as OCR) can be used to identify the core area of the image frame, such as the subtitle area or the face area, and dynamically adjust the display position of the blessing information in the image frame (such as avoiding the face, etc.). In this way, the display effect of the generated video ringback tone can be improved.
[0102] Specifically, in some embodiments, the method further includes:
[0103] Obtaining predetermined channel setting parameters;
[0104] Determining specification restriction information corresponding to the video ringback tone according to the channel setting parameters;
[0105] The video ring back tone is compressed according to the specification restriction information, and the compressed video ring back tone is sent to the ring back tone platform corresponding to the user.
[0106] In the embodiments of the present application, due to channel constraints on the playback platform, the file size of the video ringback tone played to users in real time must meet certain limits. Therefore, in the embodiments of the present application, predetermined channel setting parameters can be obtained. These channel setting parameters can be pre-set technical parameters for transmitting the video ringback tone. These parameters may include bandwidth, transmission rate, encoding format, resolution, frame rate, etc. Based on the channel setting parameters, the corresponding specification limit information for the video ringback tone, such as resolution, encoding format, file size, etc., can be determined. This specification limit information ensures that the video file meets network conditions during transmission, can be played normally on the user device, and maintains the highest possible quality.
[0107] In an embodiment of the present application, the video ringback tone can be compressed according to the specification restriction information and sent to the user's corresponding ringback tone platform. Specifically, during compression, the video resolution can be reduced from a higher resolution (such as 1080p, 4K) to a lower resolution (such as 720p or lower), which can significantly reduce the file size and improve the loading speed.
[0108] It can be understood that in the embodiments of the present application, GPU accelerated rendering and intelligent encoding technology can be used to achieve real-time synthesis and optimization of multimodal content: a lightweight rendering pipeline is used to ensure strict synchronization of dynamic elements, and the level of detail is automatically adjusted based on terminal performance; compression algorithms and special encoding are used to compress the video volume, while supporting segmented transmission and edge computing, so that the ringtone can quickly complete the entire process from generation to adaptation, significantly improving playback smoothness and reducing bandwidth costs.
[0109] In the embodiment of the present application, the processed video ringback tone will be sent to the user's corresponding ringback tone platform. As the core hub connecting content generation and terminal presentation, the ringback tone platform can achieve three core functions based on 5G edge computing and intelligent routing technology: ensuring low-latency distribution through dynamic edge node selection; combining real-time scene data (such as caller relationship and time period) to achieve precise push notifications for thousands of people; and integrating two-way interactive interfaces (such as barrage interaction) with data embedding systems to form a "generation-distribution-feedback" closed loop, ultimately achieving the commercial value of improved distribution efficiency and increased user interaction rate, and providing operators with a real-time visual content operation dashboard. When a user answers or makes a call, the ringback tone platform can play the video ringback tone for the user (caller and called ringback tone).
[0110] The technical solutions in the embodiments of the present application have at least the following advantages:
[0111] 1. Significantly improve user personalization: By dynamically integrating user portraits (interest tags, social behaviors) with real-time scene data (holiday information, hot information), highly customized video ringtones are generated, significantly improving user satisfaction.
[0112] 2. Video Ringback Tone Content Updates and Cultural Identity: The system can capture holiday and hot spot data in real time and generate corresponding video ringback tone content, effectively enhancing users' sense of freshness and identification with the video ringback tone product. Furthermore, traditional festival-themed ringback tones (such as the Dragon Boat Festival dragon boat race animation) can enhance young people's cultural identity.
[0113] 3. Intelligent processing and generation efficiency optimization: Use AI technology to automatically generate and update video content, reduce manual intervention costs and improve processing efficiency.
[0114] 4. Emotional transmission and social interaction: Through the emotion mapping algorithm and personalized parameter fusion mechanism, video ringtones become a new way for users to transmit emotions and interact socially, enhancing the emotional connection between users.
[0115] Reference Figure 3 In an embodiment of the present application, a device for generating a video ringback tone is further provided, comprising:
[0116] The responding unit 310 is configured to respond to a user's video ringback tone production request and obtain basic video data uploaded by the user;
[0117] An acquisition unit 320 is configured to acquire festival information at a current time point, user tag information of the user, and hot spot information at the current time point;
[0118] A first generating unit 330 is configured to use the holiday information, the user tag information, and the hot spot information as tags to generate corresponding personalized video data based on the basic video data through a deep learning model;
[0119] The second generating unit 340 is configured to generate corresponding greeting information according to the festival information, and generate background music corresponding to the user according to the user tag information;
[0120] The fusion unit 350 is used to perform multimodal fusion on the personalized video data, the blessing information and the background music to obtain the video ringback tone corresponding to the user.
[0121] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0122] Reference Figure 4 , an embodiment of the present application provides an electronic device, including:
[0123] at least one processor 410;
[0124] at least one memory 420, for storing at least one program;
[0125] When the at least one program is executed by the at least one processor 410, the at least one processor 410 implements the above-mentioned method for generating a video ringback tone.
[0126] Similarly, the contents of the above method embodiments are applicable to the present electronic device embodiment. The functions specifically implemented by the present electronic device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0127] The embodiment of the present application further provides a computer-readable storage medium, in which a program executable by the processor 410 is stored. When the program executable by the processor 410 is executed by the processor 410, it is used to perform the above-mentioned method for generating a video ringback tone.
[0128] Similarly, the contents of the above method embodiments are applicable to the computer-readable storage medium embodiments. The functions specifically implemented by the computer-readable storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0129] In some optional embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, the two boxes shown in succession may actually be executed substantially simultaneously or the boxes can sometimes be executed in reverse order. In addition, the embodiments presented and described in the flow chart of the present application are provided in an exemplary manner for the purpose of providing a more comprehensive understanding of the technology. The disclosed method is not limited to the operations and logic flows presented herein. Optional embodiments are contemplated in which the order of the various operations is changed and the sub-operations described as a part of a larger operation are performed independently.
[0130] In addition, although the present application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It is also understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. More specifically, given the properties, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the module will be understood within the routine skills of an engineer. Therefore, a person skilled in the art can implement the present application as set forth in the claims using ordinary techniques without undue experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0131] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0132] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0133] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.
[0134] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0135] In the above description of this specification, reference to the terms "one embodiment / example," "another embodiment / example," or "certain embodiments / examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any appropriate manner in any one or more embodiments or examples.
[0136] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and intent of the present application, and that the scope of the present application is defined by the claims and their equivalents.
[0137] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present application, and these equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A method for generating a video ringback tone, characterized in that: The method comprises: In response to a user's video ringback tone production request, obtaining basic video data uploaded by the user; Obtaining festival information at the current time, user tag information of the user, and hot spot information at the current time; Using the festival information, the user tag information, and the hot spot information as tags, a deep learning model is used to generate corresponding personalized video data based on the basic video data; Generate corresponding greeting information according to the festival information, and generate background music corresponding to the user according to the user tag information; The personalized video data, the blessing information and the background music are multimodally fused to obtain a video ringback tone corresponding to the user.
2. The method for generating a video ringback tone according to claim 1, wherein: Before using the festival information, the user tag information, and the hot spot information as tags, the method further includes: Preprocessing the festival information, the user tag information, and the hot spot information; The preprocessing includes at least one of data cleaning processing, format conversion processing, context supplementation processing, and vectorization processing.
3. The method for generating a video ringback tone according to claim 1, wherein: Generating corresponding greeting information according to the festival information includes: Randomly extracting a predetermined number of festival keywords from a pre-established keyword database based on the festival information; The holiday keywords and preset prompts are input into a blessing generation model, and the corresponding blessing information is generated by the blessing generation model; wherein the preset prompts are used to instruct the blessing generation model to generate blessing information based on the holiday keywords.
4. The method for generating a video ringback tone according to claim 3, wherein: The greeting generation model is built based on the GPT model or the Deepseek model.
5. The method for generating a video ringback tone according to claim 1, wherein: The multimodal fusion of the personalized video data, the blessing information, and the background music includes: Using the audio stream of the background music as a reference time axis, aligning the image frames of the personalized video data and the display timing corresponding to the greeting information to the same time coordinate system; Detecting a beat point in the background music, and determining the beat point as a synchronization anchor point for switching the image frame or the blessing message; The background music and the personalized video data are synchronized with music animation, and the blessing information and the personalized video data are adapted to text scenes.
6. The method for generating a video ringback tone according to claim 5, wherein: The text scene adaptation of the blessing information and the personalized video data includes: Detecting a core area of an image frame of the personalized video data by an image recognition algorithm; The display position of the blessing information in the image frame is dynamically determined based on the core area.
7. A method for generating a video ringback tone according to any one of claims 1 to 6, characterized in that: The method further comprises: Obtaining predetermined channel setting parameters; Determining specification restriction information corresponding to the video ringback tone according to the channel setting parameters; The video ring back tone is compressed according to the specification restriction information, and the compressed video ring back tone is sent to the ring back tone platform corresponding to the user.
8. A device for generating a video ringback tone, characterized in that: The device comprises: A response unit, configured to respond to a user's video ringback tone production request and obtain basic video data uploaded by the user; An acquisition unit, configured to acquire festival information at a current time point, user tag information of the user, and hot spot information at the current time point; A first generating unit is configured to use the holiday information, the user tag information, and the hot spot information as tags to generate corresponding personalized video data based on the basic video data through a deep learning model; A second generating unit is configured to generate corresponding greeting information according to the festival information, and generate background music corresponding to the user according to the user tag information; The fusion unit is used to perform multimodal fusion on the personalized video data, the blessing information and the background music to obtain the video ringtone corresponding to the user.
9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a method for generating a video ringback tone according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to implement a method for generating a video ringback tone according to any one of claims 1 to 7 when executed by the processor.