Video ringtone adjusting method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202610801066.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-25
AI Technical Summary
整个过程未对视频内容进行任何深度解析,无法识别其中的主体对象(如人物、动物、场景)、行为轨迹、情感倾向或语义主题,导致视频以“黑盒”形式进入系统,丧失后续智能处理的可能性
[0021]通过主动规避负偏好内容,并融入正偏好主题,有效降低用户跳过、关闭或投诉视频彩铃的概率,提升被叫用户的使用体验。
Smart Images

Figure CN122824844A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer vision technology, and in particular to a video ringback tone adjustment method, apparatus, electronic device and storage medium. Background Technology
[0002] Video ringback tones, as an important value-added communication service in the 5G era, have been widely deployed in operator networks. Current mainstream video ringback tones systems employ a highly static, one-way push-based technical architecture. Their core process can be summarized as "single content creation—original file storage—indiscriminate push," primarily aiming only to achieve basic playback of preset videos before the call is established, lacking the ability to understand content semantics and respond to user-differentiated needs.
[0003] In the content creation stage, users typically upload self-made videos through operator clients or partner platforms, or select template videos from the media library. The system only performs basic format checks (such as whether the resolution, duration, bitrate, and file size meet the specifications) and considers the content complete. The entire process does not involve any in-depth analysis of the video content, failing to identify the main subjects (such as people, animals, or scenes), behavioral patterns, emotional tendencies, or semantic themes. This results in the video entering the system as a "black box," losing the possibility of subsequent intelligent processing. Summary of the Invention
[0004] This disclosure aims to at least partially address one of the technical problems in the related art.
[0005] Therefore, one objective of this disclosure is to propose a method for adjusting video ringback tones.
[0006] The second objective of this disclosure is to provide a video ringback tone adjustment device.
[0007] The third objective of this disclosure is to propose an electronic device.
[0008] The fourth objective of this disclosure is to provide a non-transitory computer-readable storage medium.
[0009] The fifth objective of this disclosure is to provide a computer program product.
[0010] To achieve the above objectives, a first aspect of this disclosure provides a video ringback tone adjustment method, comprising: acquiring a first video ringback tone uploaded by a calling user, and acquiring positive preference data and negative preference data corresponding to each called user in a set of called users; performing content analysis on the first video ringback tone to determine the main trajectory data and main category data of the first video ringback tone; selecting users to be changed from the set of called users based on the main category data and the negative preference data of each called user; and adjusting the main category data based on the positive preference data and main trajectory data of the users to be changed to generate a second video ringback tone, wherein the second video ringback tone is used as a ringback tone when the calling user calls the user to be changed.
[0011] According to one embodiment of this disclosure, the step of performing content analysis on the first video ringback tone to determine the subject trajectory data and subject category data of the first video ringback tone includes: extracting features from the image frames of the first video ringback tone to obtain visual features; determining candidate subjects based on the visual features and calculating the reliability score of the candidate subjects; determining the target subject from the candidate subjects based on the reliability score, and performing cross-frame tracking of the target subject in the first video ringback tone to extract the subject category data and subject trajectory data of the target subject.
[0012] According to one embodiment of this disclosure, obtaining the subject trajectory data of the target subject includes: determining the boundary data of the target subject based on the image frames of the first video ringback tone; and performing target tracking on the first video ringback tone based on the boundary data to determine the subject trajectory data of the target subject.
[0013] According to one embodiment of this disclosure, the step of selecting users to be changed from the set of called users based on the subject category data and the negative preference data of each called user includes: performing upper-level processing on the subject category data to generate upper-level category information; matching the upper-level category information with the negative preference data; and determining the called user corresponding to the target negative preference data as the user to be changed in response to the existence of a successfully matched target negative preference data.
[0014] According to one embodiment of this disclosure, matching the hyperordinate category information with the negative preference data includes: generating a first word vector and a second word vector based on the hyperordinate category information and the negative preference data, respectively; calculating the cosine similarity value between the first word vector and the second word vector; and determining whether the hyperordinate category information matches the negative preference data based on the cosine similarity value.
[0015] According to one embodiment of this disclosure, adjusting the subject category data based on the positive preference data and subject trajectory data of the user to be changed to generate a second video ringback tone includes: determining target positive preference data from the positive preference data of the user to be changed based on the target negative preference data; and adjusting the subject category data based on the target positive preference data to generate a second video ringback tone.
[0016] According to one embodiment of this disclosure, adjusting the subject category data based on the target positive preference data to generate a second video ringback tone includes: generating video text based on the subject trajectory data and subject category data of the first video ringback tone; adjusting the video text based on the target positive preference data to generate target video text; and generating the second video ringback tone based on the target video text.
[0017] To achieve the above objectives, a second aspect of this disclosure provides a video ringback tone adjustment device, comprising: an acquisition module for acquiring a first video ringback tone uploaded by a calling user and acquiring positive preference data and negative preference data corresponding to each called user in a set of called users; an analysis module for performing content analysis on the first video ringback tone to determine the main trajectory data and main category data of the first video ringback tone; a filtering module for filtering users to be changed from the set of called users based on the main category data and the negative preference data of each called user; and an adjustment module for adjusting the main category data based on the positive preference data and main trajectory data of the users to be changed to generate a second video ringback tone, wherein the second video ringback tone is played as a ringback tone when the calling user calls the user to be changed.
[0018] To achieve the above objectives, a third aspect of this disclosure provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to implement the video ringback tone adjustment method as described in the first aspect of this disclosure.
[0019] To achieve the above objectives, a fourth aspect of this disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to implement the video ringback tone adjustment method as described in the first aspect of this disclosure.
[0020] To achieve the above objectives, a fifth aspect of this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the video ringback tone adjustment method as described in the first aspect of this disclosure.
[0021] By proactively avoiding content with negative preferences and incorporating themes with positive preferences, the probability of users skipping, turning off, or complaining about video ringback tones is effectively reduced, thus improving the user experience for called users. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of a video ringback tone adjustment method according to one embodiment of the present disclosure;
[0023] Figure 2 This is a schematic diagram of another video ringback tone adjustment method according to one embodiment of the present disclosure; Figure 3 This is a schematic diagram of another video ringback tone adjustment method according to one embodiment of the present disclosure; Figure 4 This is a schematic diagram of a video ringback tone adjustment device according to one embodiment of the present disclosure; Figure 5 This is a schematic diagram of an electronic device according to one embodiment of the present disclosure. Detailed Implementation
[0024] Embodiments of this disclosure are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0025] The acquisition, storage, use, and processing of data in this disclosed technical solution all comply with the relevant provisions of relevant laws and regulations.
[0026] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0027] In current technology, platforms store video ringback tones as raw media files in a centralized storage system, establishing a simple mapping relationship only through basic identifiers such as the called user's phone number. They fail to perform structured decomposition of the video content (such as subject segmentation, keyframe extraction, and text tag generation) or build a correlation model between content features and user profiles, making it impossible to semantically retrieve or reassemble massive amounts of video resources on demand.
[0028] During the transmission and playback phase, when the caller initiates a call, the network side recognizes that the called user has activated the video ringback tone service, and directly retrieves the unique video file preset by the user from the storage system, and pushes it to the caller's terminal via the transmission channel. This process only includes basic terminal adaptation operations (such as converting the encoding format or resolution according to the device's capabilities), and has no content awareness, no context judgment, and no dynamic adjustment capabilities throughout the entire process.
[0029] Ultimately, all callers, regardless of their interests, social relationships, geographical location, or historical behavior, receive the exact same video feed. This "one-size-fits-all" approach not only fails to improve user engagement and satisfaction but also fails to meet the growing demand for personalized communication experiences. In particular, the communicative value and emotional connection potential of video ringback tones are severely limited in scenarios such as social networking, marketing, and public welfare.
[0030] Figure 1 This is a schematic diagram of a video ringback tone adjustment method according to one embodiment of the present disclosure, as shown below. Figure 1 As shown, the method for adjusting video ringback tones includes the following steps: S101, obtain the first video ringback tone uploaded by the calling user, and obtain the positive preference data and negative preference data corresponding to each called user in the called user set.
[0031] The video ringback tone adjustment method of this application embodiment can be applied to the scenario of personalized generation of user video ringback tones. The execution subject of the video ringback tone adjustment in this application embodiment can be the video ringback tone adjustment device of this application embodiment, which can be installed on an electronic device.
[0032] In this embodiment of the disclosure, the first video ringback tone can be generated by the calling user uploading a custom video file through a client (such as a video ringback tone application), social media plugin, or API interface provided by the operator, or it can be generated by selecting a template video from the platform's content library and performing light editing (such as adding text or filters).
[0033] It should be noted that the first video ringback tone can be assigned a unique identifier and bound to the caller's number, serving as the original material for subsequent content parsing and personalized generation.
[0034] The called user set refers to the collection of all called user numbers that the current calling user has called or planned to call in the past or within a certain period of time (such as the past 30 days). It can be dynamically constructed through call records, address book authorization, or social relationship graphs.
[0035] Positive preference data represents the content characteristics that the called user is interested in or likes. Sources include: Explicit behaviors: liking, saving, and actively setting video ringback tones; Implicit behavior: Video types with high completion rates and repeated playback; User profile tags: age, gender, region, interest tags (such as "pet lover" or "tech enthusiast").
[0036] Negative preference data characterizes content features that the called user rejects or is not interested in, and its sources include: Skip / turn off video ringback tone operation history; Report or mark content categories as "not interested"; Historical feedback (e.g., setting "block advertising ringback tones"); Sensitive scene identification (such as automatic avoidance rules for religious, political, and violence-related content).
[0037] All preference data can be uniformly vectorized or labeled to form a structured user preference profile, supporting efficient matching and inference.
[0038] In one possible implementation, the platform retrieves the called user's personal information from the user information management database based on the called user's number, focusing on obtaining their pre-set sensitive word list and age information. The sensitive word list is set by the called user during platform registration or use and contains keywords that may cause them discomfort, such as "feline animals" or "hot dance"; the age information is used for subsequent age-appropriate modification of video content.
[0039] The retrieval process is achieved through efficient database indexing technology, ensuring that the required information is obtained in a short time.
[0040] Search process: When the word "feline" is detected in the video text, the inverted index can quickly locate all users who are sensitive to this word and predict the user range that needs to be replaced (to assist in batch processing scenarios).
[0041] Database-driven technology is implemented as follows: Based on the text information, let's assume the platform stores the following database table structure for user-sensitive words (using a relational database like MySQL as an example):
[0042] In this embodiment, an inverted index method that supports bidirectional fast matching between sensitive words and users can be used. That is, through the mapping relationship of "keyword → document (record)," all records containing a specific keyword can be quickly queried. It can also be used to reverse query "which users are sensitive to a certain sensitive word," but more importantly, it can assist in "fast matching of user sensitive word lists with video text."
[0043] S102, perform content analysis on the first video ringback tone to determine the main trajectory data and main category data of the first video ringback tone.
[0044] This process transforms raw video ringback tones from unstructured media streams into structured semantic representations, providing computable and editable content metadata for core personalized adjustments. Specifically, it includes the following two key technical sub-processes: It should be noted that the extraction of subject category data can identify the main semantic entities appearing in the video and their category labels (such as "person", "cat", "car", "landscape", "text", etc.).
[0045] In this embodiment of the disclosure, a multimodal deep learning model can be used to perform frame-by-frame or keyframe sampling analysis on the video frame sequence, and output data such as category label and confidence score for the foreground object detected in each frame.
[0046] The generation of subject trajectory data is used to track the movement path of the main subject in the video in the spatiotemporal dimension and to depict the pattern of its position change.
[0047] In this embodiment of the disclosure, a multi-target tracking algorithm can be used to perform cross-frame identity association on the aforementioned detected subjects, generating a spatiotemporal trajectory sequence for each subject.
[0048] In one possible implementation, higher-level motion features can be further extracted, such as direction of motion (left→right, stationary, rotation, etc.); speed changes (uniform speed, acceleration, sudden stop); interaction relationships (such as "a person pets a dog" or "a hand taps the screen"); and key action segments (such as waving, jumping, or mouth movements).
[0049] S103, based on the subject category data and the negative preference data of each called user, select the users to be changed from the set of called users.
[0050] The system filters out users from the set of called users who may have a negative experience due to the current video content and marks them as "users to be changed" for subsequent content adaptation.
[0051] In this embodiment of the disclosure, the subject category data extracted from the first video ringback tone can be matched with the negative preference list of each called user, supporting both exact matching and semantic generalization matching.
[0052] In this embodiment of the disclosure, any subject category that matches the negative preference label is considered incompatible.
[0053] S104, based on the positive preference data and subject trajectory data of the user to be changed, the subject category data is adjusted to generate a second video ringback tone, wherein the second video ringback tone is used as a ringback tone when the calling user calls the user to be changed.
[0054] Based on the positive preference data and subject trajectory data of the user to be changed, the subject category data is adjusted to perform semantic-level, non-destructive intelligent reconstruction of the first video ringback tone according to the target called user's interests and the original subject movement pattern, generating a highly adapted "second video ringback tone" to ensure that the called user's viewing experience is improved while maintaining the caller's creative intent.
[0055] The system uses the positive preference data of the user to be changed as a content guidance signal, and extracts candidate replacement / enhancement elements that match the preference from the platform's material library, AI generation model, or the original video. For example, if the user prefers "cats" but the original video contains "dogs", the "pet subject replacement" strategy is triggered.
[0056] In this embodiment, the process first acquires a first video ringback tone uploaded by the calling user, and then acquires positive and negative preference data for each called user in the called user set. Next, content analysis is performed on the first video ringback tone to determine its main trajectory data and main category data. Then, based on the main category data and the negative preference data of each called user, users to be changed are selected from the called user set. Finally, based on the positive preference data and main trajectory data of the users to be changed, the main category data is adjusted to generate a second video ringback tone. This second video ringback tone is played as a ringback tone when the calling user calls the user to be changed. By actively avoiding negative preference content and incorporating positive preference themes, the probability of users skipping, turning off, or complaining about the video ringback tone is effectively reduced, improving the user experience for called users.
[0057] In the above embodiments, content analysis is performed on the first video ringback tone to determine the main trajectory data and main category data of the first video ringback tone. It can also be achieved through... Figure 2 To further explain, the method includes: S201, Feature extraction is performed on the image frames of the first video ringback tone to obtain visual features.
[0058] S202, determine candidate subjects based on visual features and calculate the reliability score of the candidate subjects.
[0059] S203, based on the reliability score, the target subject is determined from the candidate subjects, and the target subject is tracked across frames in the first video ringback tone to extract the subject category data and subject trajectory data of the target subject.
[0060] The preprocessed video can be split into consecutive static video frames according to the frame rate (e.g., 30 frames / second), and the frame images can be standardized, including unifying the resolution (e.g., adjusting to 1080P) and normalizing the pixel values (compressing the pixel values to the range of 0-1), reducing the impact of lighting and scale differences on model detection. Multi-scale feature extraction of frame images is performed using a deep convolutional neural network (composed of an input layer, multiple convolutional layers, pooling layers, and residual connection layers), capturing visual features from low-level (edges, textures) to high-level (shapes, semantics). On the feature map, the anchor box mechanism is used to predict the subject category data (such as "cat", "person" and "sofa") and the subject trajectory data, and the confidence score (a metric for measuring the accuracy of prediction) is calculated. Multiple overlapping bounding boxes of the same subject are filtered, and the bounding box with the highest confidence is retained. Redundant predictions are removed, and finally the precise position of each subject in the current frame is determined, thus completing the subject segmentation. In this embodiment of the disclosure, to obtain the subject trajectory data of the target subject, the boundary data of the target subject can be determined first based on the image frames of the first video ringback tone, and then the target tracking of the first video ringback tone can be performed based on the boundary data to determine the subject trajectory data of the target subject.
[0061] In one possible implementation, the segmented subject can be tracked across frames based on the DeepSORT algorithm: Appearance feature extraction: The bounding box region of each subject is encoded using a convolutional neural network to generate a high-dimensional appearance feature vector, which is used to distinguish the visual differences between different subjects. Motion state prediction: The Kalman filter algorithm is used to predict the possible position of the subject in the current frame based on the motion parameters such as position and velocity of the subject in the previous few frames. Data association: Calculate the Mahalanobis distance (measuring motion consistency) between the predicted position and the detection position in the current frame, and the cosine distance (measuring appearance similarity) between the appearance feature vectors. Use the Hungarian algorithm to achieve cross-frame subject matching and association, and determine the correspondence of the same subject in consecutive frames. Trajectory Update and Recording: Assign a unique ID to each subject and update its motion trajectory data in real time, including parameters such as timestamp, position coordinates, and motion speed for each frame, to form complete subject trajectory data.
[0062] In one possible implementation, video frame preprocessing can be performed before analyzing the video frames of the first video ringback tone. For example, if the video resolution is 1920×1080 and the frame rate is 25 frames per second, each frame after splitting can be uniformly adjusted to 1080P and the pixel values can be normalized. In the 100th frame, the YOLOv8 model identifies three subjects—"cat," "sofa," and "floor"—through feature extraction: the bounding box coordinates of the "cat" are (500, 300, 700, 500), with a confidence score of 0.98; the bounding box coordinates of the "sofa" are (200, 400, 1000, 800), with a confidence score of 0.95; and the bounding box coordinates of the "floor" are (0, 600, 1920, 1080), with a confidence score of 0.99. After nonmaximum suppression, the bounding boxes of the three main subjects are retained, the subject segmentation is completed, and the position of each subject in the current frame is determined. Appearance feature extraction: Feature encoding is performed on the bounding box region (500, 300, 700, 500) of the "cat" to generate a 128-dimensional appearance vector, which is used to distinguish the cat from other objects. Motion state prediction: Based on the position of the "cat" in frames 95-99 (e.g., position in frame 99: (480, 290, 680, 490), velocity approximately 20 pixels / frame), the Kalman filter predicts the position in frame 100 as (500, 300, 700, 500), consistent with the detection result. The Mahalanobis distance between the predicted and detected positions is calculated to be 0.5 (less than the threshold of 1.0), and the cosine distance of the appearance features is 0.1 (less than the threshold of 0.3), confirming that the "cat" in frame 100 is the same subject as in frame 99. Tracking: Continuously tracked up to frame 200, recording the trajectory data of "obj_001": Frame 100 (timestamp 00:00:04, position (500,300,700,500)) → Frame 150 (timestamp 00:00:06, position (800,350,1000,550), speed 30 pixels / frame) → Frame 200 (timestamp 00:00:08, position (600,400,800,600), speed 25 pixels / frame), forming a complete trajectory record of "the cat moving from the left floor to the right sofa".
[0063] In the above embodiments, based on subject category data and negative preference data of each called user, users to be changed are selected from the set of called users. Furthermore, [the system can also be modified / modified]. Figure 3 To further explain, the method includes: S301, perform higher-level processing on the main category data to generate higher-level category information.
[0064] S302, match the upper category information with the negative preference data, and in response to the existence of a successfully matched target negative preference data, determine the called user corresponding to the target negative preference data as the user to be changed.
[0065] In this embodiment of the disclosure, a text similarity algorithm (e.g., cosine similarity algorithm) can be invoked to perform matching analysis between the called user's sensitive word library and the text description of user A's video ringback tone.
[0066] First, a first word vector and a second word vector are generated based on the hyperordinate category information and negative preference data, respectively. Then, the cosine similarity value of the first word vector and the second word vector is calculated. Finally, the cosine similarity value is used to determine whether the hyperordinate category information and the negative preference data match.
[0067] If the cosine similarity value is greater than the determination threshold, the matching is determined to be successful; otherwise, it is determined that the superordinate category information does not match the negative preference data.
[0068] In the embodiments of the present disclosure, before the subject category data is subjected to hypernym processing to generate superordinate category information, a video text may also be generated based on the subject trajectory data and subject category data of the first video ringtone, and then the video text is cleaned, for example, the following cleaning operations may be performed: Clean the video text descriptions and user-sensitive words, remove punctuation marks, special characters (such as "@" "#") and meaningless stop words (such as "de", "le", "zai"), and retain core semantic words. Word segmentation processing: use a Chinese word segmentation tool (such as Jieba word segmentation) to split the video text into independent words or phrases. For example, split "a white pet cat walking on the indoor floor" into "white pet cat indoor floor walking", and split the sensitive word "feline" into "feline". Lexical standardization: standardize the words after word segmentation, such as unifying capitalization (Chinese has no difference in capitalization, mainly for possible English words), synonym replacement (such as unifying "kitty" to "cat"), so as to reduce semantic ambiguity. In the embodiments of the present disclosure, generating the first word vector and the second word vector based on the superordinate category information and the negative preference data respectively may use the pre-trained word vector model Word2Vec to convert each word after word segmentation into a fixed-dimensional vector (such as 300 dimensions). The numerical distribution of the vector reflects the semantic features of the word, and the vector distance of words with similar semantics is closer. Weighted average or pooling processing is performed on the word vectors of video text descriptions and sensitive words to generate a vector representation of the entire text. For example, the vector of the video text description is obtained by weighted average of all the word vectors contained therein according to the occurrence frequency, and the vector of the sensitive word list is obtained by synthesizing the vectors of each sensitive word, so that the two paragraphs of text are converted into computable numerical vectors. The cosine similarity calculation can calculate the cosine value of the included angle between the video text description vector and the sensitive word vector through the cosine similarity algorithm, and the formula is: Cosθ = (A·B) (∥A∥∥B∥) Wherein, A is the first word vector, B is the second word vector, A·B is the dot product of the vectors, and ∥A∥ and ∥B∥ are the modulus lengths of the vectors respectively. The range of the cosine value is between [-1, 1], and the closer the value is to 1, the more similar the semantics of the two paragraphs of text are. Threshold judgment: set a similarity threshold (such as 0.6). When the calculated cosine similarity value is greater than or equal to the threshold, it is determined that there is a semantic association between the video text description and the sensitive words, that is, there is inappropriate content in the video; otherwise, it is determined that there is no inappropriate content. Taking the case where the called user presets the sensitive keyword "feline" as an example, and the video text description is "a white pet cat walking on the indoor floor, sometimes stopping to lick its paws, sometimes jumping onto the sofa," the specific process is as follows: The video text description, after cleaning, reads "white pet cat walking on the indoor floor, stopping to lick its paws, and jumping on the sofa"; the sensitive words, after cleaning, are "feline." Standardizing "pet cat" to "cat" leaves the sensitive words unchanged. In the above embodiments, the subject category data is adjusted based on the target positive preference data to generate the second video ringback tone. It can also be achieved through... Figure 3 To further explain, the method includes: S301, Generate video text based on the main trajectory data and main category data of the first video ringback tone.
[0069] S302, adjust the video text based on the target positive preference data to generate the target video text.
[0070] S303, Generate a second video ringback tone based on the target video text.
[0071] When a subject category that matches the negative preference data is detected in the video text description, the large language model GPT-4 can be invoked to replace and optimize the video text corresponding to the inappropriate subject.
[0072] In this embodiment of the disclosure, target positive preference data can be determined from the positive preference data of the user to be changed based on target negative preference data, and then the subject category data can be adjusted based on the target positive preference data to generate a second video ringback tone.
[0073] In this way, while avoiding negative content, the system intelligently filters out the "target positive preference" that is most relevant to the current video context and has the strongest operability from the user's diverse positive interests, and uses it as a guiding signal for content reconstruction, thereby generating a second video ringtone that is both safe and attractive.
[0074] It should be noted that the selection criteria for determining the target positive preference data from the positive preference data of the user to be changed based on the target negative preference data can be various, and no limitations are made here. For example, the target positive preference data can be determined from the positive preference data based on three dimensions: relevance (compatibility with the original video theme or scene), editability (ability to be achieved through AI generation / matching with material library), and conflict avoidance (ensuring that the target positive preference does not imply conflict).
[0075] The model generates appropriate replacement text based on the overall video scene, other main content, and the called user's sensitivities. For example, if the inappropriate subject is "cat," it is replaced with animals that the called user is less sensitive to, such as "dog" or "rabbit," while ensuring that the replaced text description is logically consistent with the overall context of the video. Simultaneously, the model optimizes the language of the replaced text to make it more natural, fluent, and in line with the expressive needs of the video content.
[0076] The platform inputs the optimized new text description into the Phenaki text-to-video generation model to generate new video content. This model, through semantic understanding of the text description and combined with temporal modeling capabilities for video generation, gradually generates a sequence of video frames that conform to the text description, starting from random noise. During generation, the model references parameters such as the original video's style, tone, and resolution to ensure that the newly generated video content maintains a consistent overall style with the original video. For content involving the subject's movement trajectory, the model generates coherent and natural motion images based on the trajectory information in the text description, making the new video content realistic and believable. Deep semantic analysis is performed on the replaced new text description to extract key information, including the subject category (e.g., "dog"), subject characteristics (e.g., "yellow fur, small dog"), activity context (e.g., "running on the indoor floor, fetching a toy"), environmental information (e.g., "living room, with sofa and carpet"), and movement trajectory (e.g., "running from left to right, then jumping onto the sofa"). Then, a pre-trained text encoder (such as BERT or CLIP text encoder) is used to convert the parsed text information into a high-dimensional text feature vector, which contains the semantic features and contextual information of the text, and serves as the input condition for the video generation model. In one possible approach, the style parameters of the first video ringback tone can also be extracted. Style feature parameters are extracted from the original video, including color tone (e.g., warm tone, cool tone), color saturation, brightness, contrast, resolution (e.g., 1080P), frame rate (e.g., 25 frames / second), and image texture (e.g., high definition, cartoon style, realistic style). The keyframes of the original video are processed by a feature extraction network to generate style feature vectors, ensuring that the newly generated video is consistent with the original video in terms of visual style. Text-to-video generation models (such as Phenaki) receive text feature vectors and style feature vectors, and gradually generate an initial sequence of video frames from random noise through a diffusion process. During diffusion, the model progressively optimizes the visual details of the video frames based on the temporal relationships and spatial layout described in the text.
[0077] The generated initial video frame sequence is processed for temporal coherence. Optical flow estimation algorithm and inter-frame interpolation technology are used to optimize the smoothness of the main motion trajectory, avoid inter-frame jumps or jitter, and ensure that the video picture is continuous and natural. In one possible scenario, super-resolution technology can also be used to improve the clarity of video images, enhancing details such as the edges and textures of the subject, making the generated video content more realistic and believable. The generated video content is compared with the text description to check whether the subject is accurate, whether the activity scenario matches, and whether the environmental information is consistent. If there are any discrepancies, they are fed back to the generation model for fine-tuning. By comparing the style parameters of the newly generated video with those of the original video, if there are significant differences (such as large color tone deviation), adjustments are made using a color mapping algorithm to ensure style consistency. For example, if the called user's sensitive keyword is "feline," the new text description is "A small yellow dog walks on the indoor floor, sometimes stopping to wag its tail, sometimes jumping onto the sofa." The original video is "White pet cat's indoor activities." The specific process is as follows: Semantic analysis: Extracting key information: The subject is a "small yellow dog", the activity is "walking on the indoor floor, stopping to wag its tail, and jumping onto the sofa", and the environment is "indoors, with a floor and a sofa". Text feature encoding: The text is converted into a 768-dimensional feature vector by CLIP text encoder, highlighting key semantics such as "yellow", "small dog", "walking", "wagging tail" and "sofa". The original video is in a realistic style with a resolution of 1920×1080, a frame rate of 25 frames per second, and a warm color tone (RGB mean value is [240,230,210]). The style vector is extracted to constrain the generation of the new video. Initial video frame generation: The Phenaki model generates 60 video frames (2.4 seconds long) based on text and style vectors. In the initial frame, the "small yellow dog" is located on the floor on the left side of the screen. Subsequent frames gradually show the dog walking, wagging its tail, and jumping onto the sofa. Timing coherence optimization: The canine movement trajectory is calculated using an optical flow algorithm, and the transition between frames is smoothed to make the jumping action natural and without stuttering. Enhanced details: Improved clarity of dog fur texture and optimized sofa fabric texture, bringing the image details closer to the clarity of the original video. Verification revealed that the "sofa color" in the generated video deviated from the original video (brown) (leaning towards a lighter yellow). By adjusting the sofa's RGB values to [180, 120, 60] through color mapping, it became consistent with the original style.
[0078] In another possible approach, user information of the user to be modified can be obtained, and video editing algorithms can be invoked to refine the newly generated video content. For underage users, the retouching operations include adjusting the screen colors to make them fresher and softer, reducing the screen transition speed to avoid visual fatigue, and filtering out potentially undesirable visual elements (such as overly bright and glaring colors, or suggestive violent scenes). For adult users, the main focus is on fine-tuning parameters such as brightness, contrast, and saturation to improve viewing comfort. The retouching process is implemented frame-by-frame, ensuring a uniform and natural retouching effect.
[0079] This achieves a balance between age-appropriateness and stylistic consistency. All editing algorithms are constrained by the stylistic features of the original video (such as tone, resolution, and texture) to ensure that the edited video is "stylistically consistent with the original content but provides a suitable user experience."
[0080] For example: if the original video is a "cartoon animation", color adjustments for minors will retain the vibrancy of the cartoon style and only reduce saturation; edge enhancement for adults will strengthen cartoon lines rather than convert to a realistic style; if the original video is a "realistic documentary", frame rate adjustments for scenes with minors will maintain the narrative rhythm of the documentary and only slow down violent motion segments; contrast optimization for scenes with adults will respect real light and shadow and avoid excessive beautification that leads to distortion.
[0081] Through the above algorithm, the platform can achieve "differentiated presentation of the same video source to different age groups," ensuring both content suitability and improving the viewing experience. The following examples illustrate the core algorithm and its effects for both minors and adults: In one possible approach, given that minors are more sensitive to the color, rhythm, and content safety of images, algorithms are needed to optimize visual comfort and content health.
[0082] Algorithm principle: Based on the HSV (Hue, Saturation, Brightness) color space, the algorithm adjusts the colors of video frames using a preset "child-friendly color matrix." It reduces the saturation of high-saturation colors (such as bright red and yellow) by 15%-20% and increases the brightness of low-brightness colors (such as light blue and light green) by 10%-15%, resulting in a soft, fresh warm or low-contrast overall image.
[0083] For example, the original video showed a brightly colored "red toy car" (80% saturation). After algorithm processing, the saturation was reduced to 65%, and the brightness was increased to 1.1 times the original, making the picture softer and avoiding stimulation to children's eyes. The overall color tone was changed from "high contrast cool color" to "low contrast warm color", which is more in line with children's visual preferences.
[0084] It can also combine target detection and image restoration technologies to automatically identify and replace elements in videos that may be inappropriate for minors (such as sharp objects, overly revealing decorations, and blurry text signs).
[0085] For example, if a "beer bottle" appears in the background of the original video, the algorithm will detect it and replace it with a "juice glass" using repair technology. The replacement area will blend seamlessly with the surrounding tablecloth texture, and the modification will be difficult to detect with the naked eye. The advertisement URL in the corner of the screen will be blurred so as not to affect the viewing of the main content.
[0086] Corresponding to the video ringback tone adjustment methods provided in the above embodiments, one embodiment of this disclosure also provides a video ringback tone adjustment device. Since the video ringback tone adjustment device provided in this disclosure corresponds to the video ringback tone adjustment methods provided in the above embodiments, the implementation methods of the above video ringback tone adjustment methods are also applicable to the video ringback tone adjustment device provided in this disclosure, and will not be described in detail in the following embodiments.
[0087] Figure 4 This is a schematic diagram of a video ringback tone adjustment device according to one embodiment of the present disclosure. As shown in FIG4, the video ringback tone adjustment device 400 includes: The acquisition module 410 is used to acquire the first video ringback tone uploaded by the calling user, and to acquire the positive preference data and negative preference data corresponding to each called user in the set of called users.
[0088] Analysis module 420 is used to perform content analysis on the first video ringback tone to determine the main trajectory data and main category data of the first video ringback tone.
[0089] The filtering module 430 is used to filter out users to be changed from the set of called users based on the subject category data and the negative preference data of each called user.
[0090] The adjustment module 440 is used to adjust the subject category data based on the positive preference data and subject trajectory data of the user to be changed, so as to generate a second video ringback tone, wherein the second video ringback tone is used as a ringback tone when the calling user calls the user to be changed.
[0091] According to one embodiment of this disclosure, content analysis is performed on a first video ringback tone to determine the subject trajectory data and subject category data of the first video ringback tone, including: extracting features from the image frames of the first video ringback tone to obtain visual features; determining candidate subjects based on the visual features and calculating the reliability score of the candidate subjects; determining the target subject from the candidate subjects based on the reliability score, and performing cross-frame tracking on the target subject in the first video ringback tone to extract the subject category data and subject trajectory data of the target subject.
[0092] According to one embodiment of this disclosure, obtaining the subject trajectory data of a target subject includes: determining the boundary data of the target subject based on the image frames of a first video ringback tone; and performing target tracking on the first video ringback tone based on the boundary data to determine the subject trajectory data of the target subject.
[0093] According to one embodiment of this disclosure, based on subject category data and negative preference data of each called user, users to be changed are selected from the set of called users, including: performing upper-level processing on the subject category data to generate upper-level category information; matching the upper-level category information with the negative preference data; and determining the called user corresponding to the target negative preference data as the user to be changed in response to the existence of a successfully matched target negative preference data.
[0094] According to one embodiment of this disclosure, matching hyperordinate category information with negative preference data includes: generating a first word vector and a second word vector based on the hyperordinate category information and the negative preference data, respectively; calculating the cosine similarity value of the first word vector and the second word vector; and determining whether the hyperordinate category information and the negative preference data match based on the cosine similarity value.
[0095] According to one embodiment of this disclosure, adjusting subject category data based on the positive preference data of the user to be changed and subject trajectory data to generate a second video ringback tone includes: determining target positive preference data from the positive preference data of the user to be changed based on target negative preference data; and adjusting subject category data based on the target positive preference data to generate a second video ringback tone.
[0096] According to one embodiment of this disclosure, adjusting subject category data based on target positive preference data to generate a second video ringback tone includes: generating video text based on subject trajectory data and subject category data of a first video ringback tone; adjusting the video text based on target positive preference data to generate target video text; and generating a second video ringback tone based on the target video text.
[0097] According to one embodiment of this disclosure, the video ringback tone adjustment device 400 is further configured to: acquire video style parameters of the first video ringback tone and / or user information of the user to be changed; and adjust the second video ringback tone based on the video style parameters and / or user information.
[0098] By proactively avoiding content with negative preferences and incorporating themes with positive preferences, the probability of users skipping, turning off, or complaining about video ringback tones is effectively reduced, thus improving the user experience for called users.
[0099] To implement the above embodiments, this disclosure also proposes an electronic device 500. Figure 5 This is a schematic diagram of an electronic device according to one embodiment of the present disclosure, such as... Figure 5As shown, the electronic device 500 includes: a processor 501 and a memory 502 communicatively connected to the processor. The memory 502 stores instructions executable by at least one processor. The instructions are executed by at least one processor 501 to achieve the functions described in this disclosure. Figures 1-3 The video ringback tone adjustment method of the embodiment.
[0100] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to implement the present disclosure. Figures 1-3 The video ringback tone adjustment method of the embodiment.
[0101] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program, which, when executed by a processor, implements the features of this disclosure. Figures 1-3 The video ringback tone adjustment method of the embodiment.
[0102] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0103] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.
[0104] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0105] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0106] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0107] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that contains, stores, communicates, propagates, or transmits programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0108] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0109] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0110] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0111] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for adjusting video ringback tones, characterized in that, include: Get the first video ringback tone uploaded by the calling user, and get the positive and negative preference data corresponding to each called user in the set of called users; Content analysis is performed on the first video ringback tone to determine the main trajectory data and main category data of the first video ringback tone; Based on the subject category data and the negative preference data of each called user, users to be changed are selected from the set of called users; Based on the positive preference data and subject trajectory data of the user to be changed, the subject category data is adjusted to generate a second video ringback tone, wherein the second video ringback tone is used as a ringback tone when the calling user calls the user to be changed.
2. The method according to claim 1, characterized in that, The content analysis of the first video ringback tone to determine the main trajectory data and main category data of the first video ringback tone includes: Feature extraction is performed on the image frames of the first video ringback tone to obtain visual features; Candidate subjects are identified based on the visual features, and the reliability scores of the candidate subjects are calculated. Based on the confidence score, the target subject is determined from the candidate subjects, and the target subject is tracked across frames in the first video ringback tone to extract the subject category data and subject trajectory data of the target subject.
3. The method according to claim 2, characterized in that, Obtaining the subject trajectory data of the target subject includes: The boundary data of the target subject is determined based on the image frames of the first video ringback tone; Based on the boundary data, target tracking is performed on the first video ringback tone to determine the subject trajectory data of the target.
4. The method according to any one of claims 1-3, characterized in that, The step of selecting users to be changed from the set of called users based on the subject category data and the negative preference data of each called user includes: The main category data is processed to generate higher-level category information; The higher-level category information is matched with the negative preference data. In response to the existence of a successfully matched target negative preference data, the called user corresponding to the target negative preference data is determined to be the user to be changed.
5. The method according to claim 4, characterized in that, The step of matching the higher-level category information with the negative preference data includes: First word vector and second word vector are generated based on the hyperclass information and the negative preference data, respectively. Calculate the cosine similarity value between the first word vector and the second word vector; The cosine similarity value is used to determine whether the hyperclassification information matches the negative preference data.
6. The method according to claim 4, characterized in that, The step of adjusting the subject category data based on the positive preference data and subject trajectory data of the user to be changed to generate a second video ringback tone includes: Target positive preference data is determined from the positive preference data of the user to be changed based on the target negative preference data; The subject category data is adjusted based on the target positive preference data to generate a second video ringback tone.
7. The method according to claim 6, characterized in that, The step of adjusting the subject category data based on the target positive preference data to generate the second video ringback tone includes: Video text is generated based on the subject trajectory data and subject category data of the first video ringback tone; The video text is adjusted based on the target positive preference data to generate target video text; The second video ringback tone is generated based on the target video text.
8. The method according to claim 1, characterized in that, The method further includes: Obtain the video style parameters of the first video ringback tone and / or the user information of the user to be changed; The second video ringback tone is adjusted based on the video style parameters and / or the user information.
9. A video ringback tone adjustment device, characterized in that, include: The acquisition module is used to acquire the first video ringback tone uploaded by the calling user, and to acquire the positive and negative preference data corresponding to each called user in the set of called users; The analysis module is used to perform content analysis on the first video ringback tone to determine the main trajectory data and main category data of the first video ringback tone; The filtering module is used to filter out users to be changed from the set of called users based on the main category data and the negative preference data of each called user; The adjustment module is used to adjust the subject category data based on the positive preference data and subject trajectory data of the user to be changed, so as to generate a second video ringback tone, wherein the second video ringback tone is used as a ringback tone when the calling user calls the user to be changed.
10. An electronic device, characterized in that, Including memory and processor; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.