A video processing method, apparatus, electronic device, and storage medium

CN122845867APending Publication Date: 2026-09-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510393194.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]然而,采用上述方式生成对象推荐文案时,可能由于缺乏对待推荐对象的理解,使得生成的对象推荐文案的质量不高,并且,生成的对象推荐文案难以与上述视频片段的画面内容相匹配,进而严重降低了最终生成的对象推荐视频的质量

Benefits of technology

[0065]本申请实施例提供了一种视频处理方法、装置、电子设备和存储介质,首先,对待推荐对象的视频素材进行分幕处理,获得多个视频片段;接着,分别对多个视频片段进行画面识别,获得多个视频片段各自的画面描述文本,并从设定素材库中查询出用于描述待推荐对象的对象描述文本;然后,基于该对象描述文本和上述多个视频片段各自的画面描述文本,生成对象推荐视频。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845867A_ABST
    Figure CN122845867A_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology, and more particularly to a video processing method, apparatus, electronic device, and storage medium. The method includes: segmenting video footage of an object to be recommended to obtain multiple video clips; performing image recognition on each of the multiple video clips to obtain image description text for each clip, each image description text describing the image content of the corresponding video clip; retrieving object description text from a predefined media library to describe the object to be recommended, and generating corresponding object recommendation text based on the object description text and the image description texts of the multiple video clips; selecting at least one video clip from the multiple video clips that semantically matches the object recommendation text, and generating an object recommendation video based on the selected at least one video clip and the object recommendation text. This application can improve the quality of generated object recommendation videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly to a video processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] In recent years, automatic video editing technology has been increasingly applied to various fields, especially object recommendation. In the object recommendation process, this technology can be used to automatically identify and edit video footage of the object to be recommended, generating a recommended video. This recommended video typically includes video clips of the object to be recommended and accompanying recommendation text; for example, when the object to be recommended is a game, the generated recommended video could include exciting game clips and game recommendation text.

[0003] Under the relevant technology, for video materials of the objects to be recommended, the existing automatic video mixing technology usually selects one or more video segments from the video materials; and generates corresponding object recommendation copy based on fixed copy generation rules (e.g., the copy must include the object name, object characteristics, target audience, etc.); finally, based on the selected one or more video segments, combined with the object recommendation copy, the corresponding object recommendation video is generated.

[0004] However, when generating object recommendation copy using the above method, the lack of understanding of the recommended object may result in low quality of the generated object recommendation copy. Furthermore, the generated object recommendation copy may be difficult to match with the visual content of the video clips, which in turn seriously reduces the quality of the final object recommendation video.

[0005] In view of this, it is necessary to redesign an automatic video editing scheme to overcome the above-mentioned defects. Summary of the Invention

[0006] This application provides a video processing method, apparatus, electronic device, and storage medium to improve the generation quality of object recommendation videos.

[0007] On one hand, an embodiment of this application provides a video processing method, including:

[0008] The video footage of the recommended target is split into multiple video segments;

[0009] Each of the multiple video segments is subjected to image recognition to obtain image description text for each video segment; each image description text is used to describe the image content of the corresponding video segment.

[0010] The object description text used to describe the object to be recommended is retrieved from the set material library, and corresponding object recommendation copy is generated based on the object description text and the screen description text of each of the multiple video clips;

[0011] From the plurality of video clips, at least one video clip is selected that semantically matches the object recommendation copy, and an object recommendation video is generated based on the selected at least one video clip and the object recommendation copy.

[0012] On one hand, an embodiment of this application provides a video processing apparatus, comprising:

[0013] The segmentation unit is used to segment the video material of the recommended object into multiple video clips.

[0014] The recognition unit is used to perform image recognition on the plurality of video segments respectively, and obtain image description text for each of the plurality of video segments; each image description text is used to describe the image content of the corresponding video segment;

[0015] The copywriting generation unit is used to query the object description text used to describe the object to be recommended from the set material library, and generate corresponding object recommendation copy based on the object description text and the screen description text of each of the multiple video clips;

[0016] The video generation unit is configured to select at least one video segment from the plurality of video segments that semantically matches the object recommendation text, and generate an object recommendation video based on the selected at least one video segment and the object recommendation text.

[0017] Optionally, the identification unit is specifically used for:

[0018] For each of the multiple video segments, perform the following operations:

[0019] Perform image recognition on a video clip to obtain the corresponding image recognition text, as well as obtain the image description text associated with the video clip;

[0020] The introductory text of the video clip is integrated with the recognized text of the video clip to obtain the introductory text of the video clip.

[0021] Optionally, each video segment has a corresponding theme scene; when performing image recognition on a video segment to obtain corresponding image recognition text, the recognition unit is specifically used for:

[0022] Perform image content recognition on a video clip to obtain the image content of the video clip;

[0023] The video clip is classified into thematic scene categories to obtain the thematic scene category of the video clip.

[0024] Based on the content and theme / scene category of a video clip, corresponding image recognition text is generated.

[0025] Optionally, when obtaining the on-screen description text associated with the video segment, the recognition unit is specifically used for:

[0026] Based on at least one of the image annotation text and image background text contained in the video clip, obtain the image description text associated with the video clip;

[0027] The image annotation text is obtained by performing text recognition on at least one video frame in the video segment, and the image background text is obtained by performing speech recognition on the image background speech associated with the video segment.

[0028] Optionally, when retrieving object description text from the set material library to describe the object to be recommended, the copy generation unit is specifically used for:

[0029] When the set material library includes multiple different types of material libraries, the content of objects that match the object to be recommended is queried from each of the multiple different types of material libraries;

[0030] The content of multiple retrieved objects is integrated to generate the object description text.

[0031] Optionally, when generating corresponding object recommendation text based on the object description text and the respective scene description texts of the plurality of video segments, the text generation unit is specifically used for:

[0032] The object description text and the screen description text of each of the multiple video segments are combined to generate a text generation prompt message;

[0033] Using a pre-trained copywriting generation model, and based on the copywriting generation prompts and the copywriting structure of the reference recommended copywriting, recommended copywriting for the target is generated.

[0034] Optionally, when the copywriting generation model generates the recommended copywriting for the target based on the copywriting generation prompt information and the copywriting structure of the reference recommended copywriting, the copywriting generation unit is specifically used for:

[0035] Based on the copy generation prompts and the copy structure of the reference recommended copy, the copy generation model generates candidate recommended copy and presents the candidate recommended copy in the copy preview interface.

[0036] In response to the adjustment operation on the candidate recommendation copy, the adjusted candidate recommendation copy is obtained, and the adjusted candidate recommendation copy is used as the object recommendation copy.

[0037] Optionally, when presenting the candidate recommended text in the text preview interface, the text generation unit is specifically used for:

[0038] The candidate recommendation text is broken down and parsed to obtain its structured information; wherein, the structured information includes the content attributes of each part of the candidate recommendation text.

[0039] The candidate recommended texts and the structured information are presented in the text preview interface.

[0040] Optionally, when selecting at least one video segment from the plurality of video segments that semantically matches the object recommendation text, the video generation unit is specifically used for:

[0041] The video features of each of the multiple video segments are extracted respectively, and the object recommendation text is divided into at least one text segment, and the text features of each of the at least one text segment are extracted respectively.

[0042] For the at least one text segment, perform the following operations respectively: based on the text features of a text segment, and the similarity between the text features and the video features of the plurality of video segments respectively, select at least one video segment that matches the text segment from the plurality of video segments;

[0043] Based on at least one video segment selected for the at least one text segment, at least one video segment that semantically matches the recommended text for the object is obtained.

[0044] Optionally, when extracting the video features of each of the plurality of video segments, the video generation unit is specifically used for:

[0045] For each of the multiple video segments, perform the following operations:

[0046] Text features are extracted from the descriptive text of a video clip to obtain text features, and image features are extracted from at least one video frame in the video clip to obtain at least one image feature;

[0047] The text features and at least one image feature are fused together to obtain the video features of the video segment.

[0048] Optionally, when obtaining at least one video segment that semantically matches the recommended text based on at least one video segment selected for the at least one text segment, the video generation unit is specifically used for:

[0049] When multiple video segments are selected for the at least one text segment, and the set video generation restrictions are received, at least one video segment that meets the video generation restrictions is selected again from the multiple selected video segments.

[0050] At least one video segment that is selected again will be considered as at least one video segment that semantically matches the recommended text for the object.

[0051] Optionally, when generating an object recommendation video based on the selected at least one video segment and the object recommendation text, the video generation unit is specifically used for:

[0052] For the at least one text segment, perform the following operations respectively: obtain at least one video segment that matches a text segment, and the corresponding cleaned video segment; convert the text segment into text-to-speech; and according to the duration of the text-to-speech, splice and edit the at least one matched cleaned video segment to obtain the corresponding edited video.

[0053] Based on at least one video clip and the corresponding text / voice for each of the at least one video clips, the recommended video for the target is obtained.

[0054] Optionally, the device further includes an adjustment unit for:

[0055] The recommended video for the object is displayed in the video preview interface;

[0056] In response to the adjustment operation for the recommended video of the object, the adjusted recommended video of the object is obtained.

[0057] Optionally, the screen splitting unit is specifically used for:

[0058] The video material is split into multiple segments to obtain multiple segmented clips; each video frame in each segment has the same theme scene.

[0059] When the duration of at least one of the multiple segment segments is not greater than a set duration, each segment segment in the at least one segment segment is spliced ​​with the adjacent segment segment.

[0060] The multiple video segments are obtained based on each spliced ​​segment and each unspliced ​​segment.

[0061] On one hand, an electronic device provided in this application includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of any of the above-described video processing methods.

[0062] On one hand, embodiments of this application provide a computer-readable storage medium including a computer program, which, when run on an electronic device, causes the electronic device to perform the steps of any of the above-described video processing methods.

[0063] On one hand, embodiments of this application provide a computer program product, the computer program product including a computer program stored in a computer-readable storage medium; when the processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any of the above-described video processing methods.

[0064] The solution in this application embodiment has at least the following beneficial effects:

[0065] This application provides a video processing method, apparatus, electronic device, and storage medium. First, the video material of the object to be recommended is processed into multiple video segments. Then, screen recognition is performed on each of the multiple video segments to obtain screen description text for each video segment, and object description text for describing the object to be recommended is retrieved from a set material library. Then, based on the object description text and the screen description text of each of the multiple video segments, an object recommendation video is generated.

[0066] By combining the aforementioned object description text with the individual scene description texts of multiple video clips, the object to be recommended can be understood more accurately, thereby generating high-quality object recommendation copy. Next, from the multiple video clips, at least one video clip that semantically matches the object recommendation copy is selected, ensuring a high degree of matching between the selected video clip and the object recommendation copy. Finally, based on the selected at least one video clip and the object recommendation copy, a high-quality object recommendation video can be generated. Therefore, the embodiments of this application improve the quality of generated object recommendation videos.

[0067] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0068] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0069] Figure 1 This is a schematic diagram illustrating an application scenario of a video processing method according to an embodiment of this application;

[0070] Figure 2 This is a flowchart of a video processing method according to an embodiment of this application;

[0071] Figure 3 This is a schematic diagram illustrating the generation process of screen description text in an embodiment of this application;

[0072] Figure 4 This is a schematic diagram illustrating the process of understanding video material in an embodiment of this application;

[0073] Figure 5 This is a schematic diagram illustrating the generation process of object recommendation copy in an embodiment of this application;

[0074] Figure 6 This is a schematic diagram of the interface for adjusting a candidate recommendation text in an embodiment of this application;

[0075] Figure 7 This is a schematic diagram illustrating the presentation of game recommendation text and structured information in an embodiment of this application;

[0076] Figure 8 This is a schematic diagram of the feature extraction process of a feature extraction model in an embodiment of this application;

[0077] Figure 9 This is a schematic diagram of a preview interface for recommending videos to an object in an embodiment of this application;

[0078] Figure 10 This is a schematic diagram of the composition structure of an automatic game video mixing and editing system according to an embodiment of this application;

[0079] Figure 11 This is a schematic diagram of a game video editing process in an embodiment of this application;

[0080] Figure 12A This is a schematic diagram of a game video editing interface in an embodiment of this application;

[0081] Figure 12B This is a schematic diagram of another game video editing interface in an embodiment of this application;

[0082] Figure 13 This is a schematic diagram of the composition structure of a video processing device according to an embodiment of this application;

[0083] Figure 14 This is a schematic diagram of the composition structure of an electronic device according to an embodiment of this application;

[0084] Figure 15 This is a schematic diagram of the composition structure of another electronic device according to an embodiment of this application. Detailed Implementation

[0085] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0086] The following describes some of the concepts involved in the embodiments of this application.

[0087] Large Language Models (LLMs) are an artificial intelligence technique that automatically learns language patterns and generates natural language sentences and paragraphs in the field of natural language processing. LLMs utilize unsupervised or semi-supervised learning on massive corpora, employing deep learning techniques to encode each word, phrase, and sentence as a numerical variable, and then train and optimize the model based on the existing corpus. Currently, LLMs have been applied to various natural language processing problems. Examples include, but are not limited to, Generative Pre-trained Transformer (GPT) models, Bidirectional Encoder Representations from Transformers (BERT) models, and Text-to-Text Transfer Transformer (T5) models.

[0088] Multimodal Large Language Models (MLLMs) are large language models capable of processing and understanding multiple types of data, such as text, images, audio, and video. These models are designed to combine information from different modalities to analyze and generate data in a more comprehensive and accurate manner. By learning and understanding data from different modalities, they can perform complex tasks such as cross-modal information retrieval, generation, and transformation. In practical applications, MLLMs can be used in fields such as autonomous driving, medical image analysis, intelligent assistants, and content generation, leveraging their cross-modal understanding capabilities to provide more intelligent and accurate services.

[0089] Retrieval-Augmented Generation (RAG): A technique that combines information retrieval and generative models to improve the performance of generative artificial intelligence systems. The working steps of a RAG model typically include: (1) Retrieval: Retrieving text fragments or information relevant to the input query from a large knowledge base or document collection. This step usually uses information retrieval techniques to find the most relevant content from a large amount of pre-stored data; (2) Generation: Using generative models (such as neural network language models) to combine the retrieved information with the input to generate more accurate and context-sensitive output.

[0090] Automatic Speech Recognition (ASR) is a technology that converts human speech into editable text, aiming to achieve automatic transcription from speech to text by analyzing speech signals. ASR technology is widely used in various fields, such as voice assistants, customer service, caption generation, voice search, and dictation systems. It promotes seamless communication and information retrieval by improving the convenience and efficiency of human-computer interaction. With the development of deep learning and natural language processing technologies, the accuracy and application scope of ASR systems are constantly improving.

[0091] Optical Character Recognition (OCR) is a technology that extracts and converts printed or handwritten text from images into editable text. OCR technology is widely used in document digitization, automated data entry, e-book creation, license plate recognition, and ticket processing. By converting physical text into digital format, OCR improves the efficiency of information processing, reduces errors from manual input, and facilitates the storage, retrieval, and sharing of information, especially in situations requiring rapid processing and analysis of large amounts of text data.

[0092] Video segmentation refers to the process of identifying and dividing different thematic scenes in a video stream during video processing. Each "thematic scene" or "segment" refers to a continuous segment of video with relatively consistent content, usually composed of a set of consecutive video frames that share similar visual features, backgrounds, or themes.

[0093] The word “exemplary” as used below means “serving as an example, embodiment, or illustration.” Any embodiment illustrated as an “exemplary” need not be construed as superior to or better than other embodiments.

[0094] The terms "first" and "second" used in this document are for descriptive purposes only and should not be construed as indicating relative importance or implying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0095] The design concept of the embodiments of this application will be briefly explained below.

[0096] Under the relevant technology, for video materials of the object to be recommended, the existing automatic video mixing technology usually selects one or more video segments from the video materials; and generates corresponding recommendation text based on fixed text generation rules (e.g., the text must include the object name, object characteristics, target audience, etc.); finally, based on the selected one or more video segments, combined with the recommendation text, the corresponding object recommendation video is generated.

[0097] However, when generating object recommendation copy using the above method, the lack of understanding of the recommended object may result in low quality of the generated object recommendation copy. Furthermore, the generated object recommendation copy may be difficult to match with the visual content of the video clips, which in turn seriously reduces the quality of the final object recommendation video.

[0098] In view of this, embodiments of this application provide a video processing method, apparatus, electronic device, and storage medium. First, video materials of the object to be recommended are segmented to obtain multiple video clips. Next, the screen description text of each of the multiple video clips is identified, and object description text describing the object to be recommended is retrieved from a set material library. Then, based on the object description text and the screen description texts of the multiple video clips, an object recommendation text for the object to be recommended is generated. By combining the object description text and the screen description texts of the multiple video clips, the object to be recommended can be accurately understood, thereby generating a high-quality object recommendation text. Next, at least one video clip that semantically matches the object recommendation text is selected from the multiple video clips, thereby ensuring a high degree of matching between the selected video clip and the object recommendation text. Finally, based on the selected at least one video clip and the object recommendation text, a high-quality object recommendation video can be generated. Therefore, embodiments of this application improve the generation quality of object recommendation videos.

[0099] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0100] like Figure 1 The diagram shown illustrates an application scenario of an embodiment of this application. The application scenario diagram includes a terminal device 110 and a server 120. The terminal device 110 and the server 120 can communicate via a communication network.

[0101] In one alternative implementation, the communication network can be a wired network or a wireless network. Therefore, the terminal device 110 and the server 120 can be directly or indirectly connected via wired or wireless communication. This application embodiment does not impose specific limitations here.

[0102] In this embodiment, the terminal device 110 includes, but is not limited to, mobile phones, tablets, laptops, desktop computers, e-book readers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. The terminal device may have a video processing-related client installed; this client can be software, a webpage, a mini-program, etc. The server 120 can be a backend server corresponding to the software, webpage, mini-program, etc., or a server specifically used for video processing; this application does not impose specific limitations. The server 120 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0103] It should be noted that the video processing methods in the various embodiments of this application can be executed by the terminal device 110 or the server 120 alone, or by the terminal device 110 and the server 120 working together.

[0104] In some embodiments, taking server 120 as an example, server 120 can receive video material of the object to be recommended uploaded by terminal device 110, then use the video processing method of this application embodiment to process the video material, finally generate object recommendation video, and return the object recommendation video to terminal device 110.

[0105] In other embodiments, taking the terminal device 110 as an example, the terminal device 110 can receive video material of the object to be recommended uploaded by the terminal device 110, then use the video processing method of the present application embodiment to process the video material, finally generate the object recommendation video, and return the object recommendation video to the terminal device 110.

[0106] In other embodiments, taking the collaborative execution of terminal device 110 and server 120 as an example, server 120 can receive video material of the object to be recommended uploaded by terminal device 110, then use the video processing method of the present application embodiment to process the video material, finally generate the object recommendation video, and return the object recommendation video to terminal device 110.

[0107] It should be noted that, Figure 1 The examples shown are merely illustrative; in reality, the number of terminal devices and servers is unlimited and is not specifically limited in the embodiments of this application.

[0108] In the field of object recommendation, the video processing method adopted in the embodiments of this application can greatly improve the generation quality and efficiency of object recommendation videos, and can be applied to any object recommendation scenario. The following are some specific application scenario examples:

[0109] 1. Game Scene

[0110] When a new game is about to be launched, it is usually necessary to release a game promotional video. The video processing method of this application embodiment can process the video materials of the game (such as exciting game clips) to quickly generate an attractive game promotional video that highlights the core gameplay and unique selling points of the game.

[0111] 2. Film and television scenes

[0112] When a new film or television work (such as a movie or TV series) is about to be released, it is usually necessary to release a trailer. Taking a movie as an example, the video processing method of this application embodiment can process the video material of the movie (such as exciting movie plots) to automatically generate a movie trailer, which can arouse the audience's interest.

[0113] 3. Tourism Scenarios

[0114] To promote tourist attractions, promotional videos can be generated to attract visitors. The video processing method described in this application can process video footage of tourist attractions (such as videos of beautiful scenery and delicious food) to automatically generate promotional videos.

[0115] 4. E-commerce platform scenario

[0116] For products on e-commerce platforms, generating product recommendation videos can introduce the products to consumers. The video processing method described in this application can process product video materials (such as product feature introduction videos, usage method introduction videos, etc.) to automatically generate product recommendation videos.

[0117] It should be noted that the above application scenarios are merely exemplary, and the video processing methods of this application embodiment are not limited to the above application scenarios.

[0118] The video processing method provided by the exemplary embodiments of this application will be described below with reference to the accompanying drawings and the application scenarios described above. It should be noted that the application scenarios described above are only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this respect.

[0119] See Figure 2The diagram shown is a flowchart of an implementation of a video processing method provided in this application. Taking a server as the execution entity as an example, the specific implementation process of this method includes the following S21-S24:

[0120] S21. Perform segmentation processing on the video materials of the recommended object to obtain multiple video clips.

[0121] The objects to be recommended can be of various types, including but not limited to games, items, works (such as films and music), and locations (such as tourist attractions).

[0122] Specifically, the server can receive video footage of objects to be recommended from terminal devices, or it can actively retrieve video footage of objects to be recommended from terminal devices; this video footage can be input from the user's terminal device. Video footage of objects to be recommended refers to the video materials used to create the object recommendation video, typically showcasing the characteristics, advantages, and appeal of the object; the content of the video footage varies depending on the type of object being recommended. For example, if the object is a game, the video footage usually includes exciting game clips to showcase the game's features, gameplay, characters, and other content that attracts players; similarly, if the object is a film or television work, the video footage usually includes exciting plot points to showcase the film's compelling content.

[0123] By segmenting the video footage of the recommended users into segments, different thematic scenes within the video footage can be distinguished. This allows consecutive video frames belonging to the same thematic scene to be grouped into a single video segment, thus giving each video segment a corresponding thematic scene. This facilitates the understanding and indexing of the video footage (e.g., locating specific thematic scenes), making video editing easier. Specifically, the server can employ various video segmentation methods to segment the video footage of the recommended users, primarily including threshold-based video segmentation methods, machine learning-based video segmentation methods, and deep learning-based video segmentation methods. These optional video segmentation methods are described below.

[0124] 1. Threshold-based video segmentation method: This method calculates the difference (such as color histogram difference, pixel difference, etc.) between adjacent video frames. When the difference exceeds a set threshold, it is considered that a scene switch has occurred between adjacent video frames, thus segmenting adjacent video frames belonging to different thematic scenes. This method is simple and direct.

[0125] 2. Machine Learning-Based Video Segmentation Methods: These methods utilize machine learning algorithms to automatically learn and distinguish different thematic scenes. For example, a classifier can be trained using models such as Support Vector Machines and Decision Trees. This classifier can extract feature representations from different video frames of the video material and determine whether a scene switch has occurred based on the differences in feature representations of adjacent video frames, thereby segmenting adjacent video frames belonging to different thematic scenes.

[0126] 3. Deep Learning-Based Video Segmentation Methods: In recent years, with the development of deep learning technology, many transition detection models based on structures such as convolutional neural networks and recurrent neural networks have emerged. These models can capture complex spatiotemporal features, providing more accurate scene boundary detection results. Specifically, transition detection models can extract deeper visual features from each video frame in the video footage. These visual features can capture information such as object shape, motion dynamics, and spatial layout. Then, they compare the differences in feature representations between adjacent video frames to determine if a scene transition has occurred, thereby segmenting adjacent video frames belonging to different thematic scenes. Furthermore, for gradual transition scenes, they can analyze the changing trends of the video frame sequence over a period of time to determine scene segmentation points.

[0127] For example, the transition detection model can use TransNet Version 2 (TransNetV2), a deep learning model used for video frame quality assessment and inter-frame difference detection.

[0128] It should be noted that different video segmentation methods can be used simultaneously to segment video footage. For example, the threshold-based video segmentation method mentioned above can be used first to segment the video footage once; then, a machine learning-based video segmentation method or a deep learning-based video segmentation method can be used to perform a second segmentation on the result of the first segmentation, thereby further improving the accuracy of video segmentation.

[0129] In some embodiments, considering that after segmenting the video material, there may be excessively short segment segments, these segments can be merged with adjacent segments to avoid flickering in the subsequently generated recommended video. In this case, the segmentation process described in S21, which involves processing the video material of the object to be recommended to obtain multiple video segments, may include the following steps A1-A3:

[0130] A1. Perform segmentation processing on the video footage to obtain multiple segmented clips; each video frame in each segmented clip has the same theme scene.

[0131] This step can be implemented using the video splitting method described in the previous embodiments, and will not be repeated here.

[0132] It should be noted that before segmenting video footage, it can be cleaned to remove unnecessary elements, such as subtitles, audio (e.g., background music, narration, noise), and logos (e.g., brand logos). This cleaned video footage makes segmentation easier. Cleaning video footage can be understood as extracting the image sequence corresponding to each video frame; then, using the aforementioned video segmentation method, the extracted image sequence is segmented to obtain multiple image subsequences; and each video frame corresponding to each image subsequence is treated as a segment.

[0133] The following example, using a game scene, illustrates the segmentation of video footage.

[0134] After segmenting a game's video footage, multiple segmented clips with corresponding thematic scenes can be obtained. Each thematic scene typically represents a specific game event, level, or activity. For example, taking an adventure game titled "Mysterious Ruins Adventure," the segmented clips from the game's video footage would include:

[0135] The first scene in the split sequence is an entrance exploration, featuring the protagonist standing in front of the entrance to an ancient ruin at the edge of a lush tropical rainforest.

[0136] The second segment, themed around solving a mystery in a large hall, features scenes of a massive stone hall within the ruins, surrounded by ancient murals and unsolved mysteries.

[0137] The third segment, themed around navigating an underground river, features a scene depicting a calm surface but surrounded by unknown dangers.

[0138] Scene four, the theme of which is the confrontation with the Guardians, includes a magnificent but trap-filled hall with a huge statue standing in the center. This is the territory of the Guardians.

[0139] Scene 5: The main scene unveils the treasure room, featuring a magnificent treasure room with walls inlaid with countless gems and a legendary treasure displayed in the center.

[0140] A2. When the duration of at least one segment in a plurality of segmented segments is not greater than the set duration, each segment in the at least one segmented segment is spliced ​​with the adjacent segmented segment.

[0141] The duration can be set as needed, such as 2 seconds, 3 seconds, etc., without limitation. If the duration of a segment is less than or equal to the set duration, the segment is considered too short. In this case, the segment can be spliced ​​with the adjacent preceding or following segment. Furthermore, if the spliced ​​segment is still less than or equal to the set duration, it can be spliced ​​with the adjacent preceding or following segment again, and so on.

[0142] A3. Based on each spliced ​​segment and each unspliced ​​segment, multiple video segments are obtained.

[0143] When a video segment is a candidate segment after splicing, the video segment can contain multiple thematic scenes.

[0144] In this embodiment of the application, after the video material is processed into segments, if there are excessively short segments, the excessively short segments can be merged with adjacent segments to avoid flickering in the subsequently generated object recommendation video.

[0145] After dividing the video footage of the target audience into multiple video segments, the next step is to perform image recognition on each video segment so that the image recognition results can be combined to generate the target audience recommendation copy.

[0146] S22. Perform image recognition on multiple video segments to obtain image description text for each video segment; each image description text is used to describe the image content of the corresponding video segment.

[0147] In this embodiment, a pre-trained image recognition model can be used to perform image recognition on each video segment to generate image description text for the video segment, including but not limited to descriptions of people, actions, events, environment (such as locations, buildings, etc.), emotions, etc. Optionally, the above image recognition model can be a multimodal large model, such as Large Language and Vision Assistant-OneVision (LLaVA-OneVision) or Generative Pre-trained Transformer4Omni (GPT4o).

[0148] Specifically, for each video segment, the image recognition model can extract the static image content of each video frame in the video segment. At the same time, it captures the dynamic changes of the static image content of each video frame through time series modeling, that is, it arranges the static image content of each video frame in chronological order and analyzes the changes in the static image content, such as changes in character actions and skill effects in a game scene. In this way, it can recognize the image content of the video segment and generate image description text to describe the image content.

[0149] In some embodiments, to gain a more accurate and in-depth understanding of each video segment, after recognizing the image recognition text of each video segment based on the image recognition model (generated based on the recognized image content), the image recognition text of each video segment can be combined with the image description text associated with that video segment to generate a more accurate image description text. Therefore, when performing image recognition on multiple video segments in S22 above to obtain the image description text for each video segment, the following steps A1-A3 can be performed separately for each video segment:

[0150] A1. Perform image recognition on a video clip to obtain the corresponding image recognition text.

[0151] In one optional implementation, considering that each video segment has a corresponding theme scene, combining the identified video segment's content with the theme scene can lead to a more accurate understanding of the video segment. Therefore, when identifying the image recognition text of each video segment, the video segment's content can be identified to obtain its content; simultaneously, the video segment can be classified into theme scene categories to obtain its category; and then, based on the video segment's content and theme scene category, corresponding image recognition text can be generated.

[0152] Specifically, the image recognition model described in the aforementioned embodiments is used to identify the image content of each video segment, and a pre-trained theme scene classification model is used to classify the theme scene of each video segment to obtain the theme scene category of the video segment. The theme scene classification model can be obtained by fine-tuning a multimodal large model, which includes, but is not limited to, LLaVA-OneVision and GPT4o. The fine-tuning process of the multimodal large model is described below.

[0153] Specifically, taking games as the type of object to be recommended as an example, a large number of game-related video clips can be collected. Each game video clip has a corresponding theme scene. These game video clips are then cleaned and labeled with theme scene categories, that is, the true theme scene category of each game video clip is labeled, resulting in the category labeling result for each game video clip. Finally, a first training sample set that meets the requirements is obtained. Each first training sample includes a game video clip and its corresponding category labeling result. Based on the above first training sample set, the large language model is fine-tuned to obtain the theme scene classification model.

[0154] The fine-tuning process of the multimodal large model includes multiple rounds of iterative training. In each round, one or more first training samples are selected from the first training sample set and input into the multimodal large model to obtain the category prediction result for each first training sample. This category prediction result includes the predicted topic scene category of the corresponding first training sample. Using a set loss function (e.g., cross-entropy loss function), a first loss value is calculated based on the difference between the category prediction result and the corresponding category label result for each first training sample. Based on this first loss value, the parameters of the multimodal large model are updated using the gradient descent method. When the iteration stopping condition is met, the currently fine-tuned multimodal large model is used as the topic scene classification model. The iteration stopping condition can be that the first loss value is not greater than a first set threshold, or the number of iterations reaches a first set number, etc.

[0155] The specific theme scene categories are related to the type of the object to be recommended. For example, if the type of the object to be recommended is a combat game, the theme scene categories include, but are not limited to, combat scenes, welfare scenes, victory scenes, and defeat scenes. As another example, if the type of the object to be recommended is a film or television work, the theme scene categories include, but are not limited to, exterior scenes (i.e., scenes filmed in natural environments or outdoors), interior scenes (i.e., scenes filmed in indoor environments), close-up scenes (i.e., focusing on the details of a specific object or person's face to highlight emotions or important items), action scenes, dialogue scenes, and fantasy scenes.

[0156] Understanding video clips by combining their visual content with thematic scene categories can provide a more comprehensive and in-depth understanding, because thematic scene categories provide rich contextual information.

[0157] Taking the example of a game as the type of recommended object, for instance, a video clip's content includes: a character using a skill to attack an enemy, a health bar decreasing on the screen, and special effects animations demonstrating the skill's effect. The main scene category of this video clip is a battle scene. In this case, combining the screen content and the main scene category, the generated screen recognition text for this video clip is: "The character used a skill to attack the enemy in battle, which consumed a certain amount of the character's energy, while simultaneously displaying the skill's special effects animations."

[0158] For example, a video clip might contain the following: a character stands in the center of a battlefield, the sky clears, and sunlight shines down. The main scene category of this video clip is a victory scene. In this case, by combining the content of the scene and the theme scene category, the generated image recognition text for this video clip would be: After a battle, the protagonist successfully defeated the enemy. With the arrival of victory, the originally dark and oppressive environment becomes bright and open, symbolizing the end of the crisis.

[0159] For example, a video clip might contain the following: a pop-up on the screen displaying the name, icon, and attributes of an item. If the main theme of the video clip is a "welfare" scene, then by combining the content of the video clip with the theme scene category, the generated text for the video clip would be: "We're giving players a welfare gift," meaning the item displayed on the screen is being given to the player as a reward.

[0160] In the above embodiments of this application, based on the image content and theme scene category of each video segment, the video segment can be understood more accurately, thereby generating more accurate image recognition text.

[0161] A2. Obtain the descriptive text associated with a video clip.

[0162] The descriptive text for the video clips can include text contained within the video content, specifically on-screen annotations, including but not limited to text annotating on-screen elements and subtitles. For example, in a game scenario, text is typically used to annotate on-screen elements to describe them, such as characters, skills, and environments. Annotations for characters could include character names, skill names, and location names. Similarly, in film and television scenarios, on-screen annotations often include subtitles. The descriptive text can also include text converted from audio within the video clips, specifically background audio, including but not limited to background narration and voice-over.

[0163] In an optional implementation, if a video clip contains at least one of on-screen annotation text and on-screen background audio, then when obtaining the on-screen description text associated with a video clip in step A2 above, the on-screen description text associated with the video clip can be obtained based on at least one of the on-screen annotation text and on-screen background text contained in the video clip; wherein, the on-screen annotation text is obtained by performing text recognition on the screen of at least one video frame in the video clip, and the on-screen background text is obtained by performing speech recognition on the on-screen background audio contained in the video clip.

[0164] Specifically, if a video clip contains on-screen annotation text or on-screen background audio, the converted on-screen background text can be used as the on-screen description text. If a video clip contains both on-screen annotation text and on-screen background audio, the converted on-screen background text can be combined to form the on-screen description text. Furthermore, keyword matching and semantic matching can be used to identify identical content in the on-screen annotation text and on-screen background text, and duplicate content can be removed.

[0165] When obtaining the image annotation text contained in a video clip, one or more video frames can be extracted from the video clip, and various extraction methods can be used. For example, a video frame can be extracted at fixed time intervals, or at fixed frame intervals; keyframes (such as I-frames) in the video clip can also be extracted. An I-frame is an independent frame that does not depend on other frames in video compression.

[0166] Specifically, OCR technology can identify the text contained in each extracted video frame. If different video frames contain the same text, duplicate text can be removed. For example, in a game scene, the text identified from the video frames of a video clip includes, but is not limited to, character names and skill names. ASR technology can convert background audio into background text.

[0167] A3. Integrate the on-screen description text with the on-screen recognition text to obtain a video clip's on-screen description text.

[0168] The accompanying descriptive text provides contextual information for the image recognition text, thus integrating the image recognition text and the descriptive text to achieve a more accurate and in-depth understanding of the video clip and obtain a more precise image description. For example, if the image recognition text is: "A dog suddenly stopped," and the descriptive text is: "The owner called the dog," it can be understood that the dog stopped running at the owner's command.

[0169] Specifically, a pre-trained large language model can be used to integrate the introductory text and the recognized text of a video clip to generate a descriptive text for that clip. Large language models include, but are not limited to, GPT and BERT series models. These models utilize natural language understanding techniques to perform contextual understanding on the introductory and recognized texts, extracting key information (such as location, people, and events), and searching for keywords in the introductory text. They then find corresponding descriptions in the recognized text, connecting different textual contents based on chronological order or causal relationships to combine the introductory and recognized texts into a coherent overall description. Through its powerful natural language understanding, contextual understanding, and generation capabilities, the large language model can transform multiple textual contents into a coherent and logically clear whole, efficiently integrating multiple textual contents.

[0170] For example, such as Figure 3 As shown, suppose the input text for the large language model is: "In a mysterious ancient forest, the protagonist's party encounters a group of enemies," and the image recognition text is: "tall trees, characters in battle, skill effects." The integrated image description text is: "In a mysterious ancient forest, the protagonist's party is facing a challenge from a group of powerful enemies. As the battle intensifies, the protagonists unleash their unique skills."

[0171] In this embodiment of the application, by integrating the image recognition text and image description text of each video segment, each video segment can be understood more deeply and accurately, thereby obtaining the image description text of each video segment more accurately.

[0172] The following is combined Figure 4 The process of understanding the video materials in the embodiments of this application will be described by way of example.

[0173] Assuming the above description text includes both image annotations and background text, for example, such as... Figure 4 As shown, the understanding process of video material in this application embodiment includes: after segmenting the video material into multiple video segments; for each video segment, the following operations are performed: each video segment is input into a multimodal large model for image recognition to obtain the image recognition text of the video segment. Specifically, the multimodal large model may include the image recognition model and the scene classification model in the aforementioned embodiment. At this time, the image content of the video segment is recognized based on the image recognition model, and the theme scene category of the video segment is recognized based on the scene classification model. Based on the image content and theme scene category of the video segment, image recognition text is generated.

[0174] The OCR model is used to perform text recognition on at least one frame of each video segment to obtain the labeled text. An ASR model is used to perform speech recognition on the background audio associated with each video segment to obtain the background text. Finally, based on a large language model, the labeled text, text, and background text of each video segment are integrated to generate the corresponding descriptive text. It should be noted that the labeled text and descriptive text can also be integrated using the aforementioned labeled text model to generate the descriptive text for the video segment.

[0175] This application embodiment uses a multimodal large model to perform in-depth analysis and understanding of video clips, which can more accurately identify and understand the content of video clips. Furthermore, by integrating the image recognition text extracted by the multimodal large model, the image annotation text recognized by the OCR model, and the image background text recognized by the ASR model, the image description text of the video clip can be generated more accurately.

[0176] The following description uses the "Mysterious Ruins Adventure" game from the aforementioned embodiments as an example to illustrate the content of the screen description text.

[0177] For example, the image recognition text for a video clip from the game "Mysterious Ruins Adventure" is: In a mysterious ancient forest, the protagonist's party encounters a group of enemies. After a fierce battle, they discover a hidden entrance leading to a forgotten temple.

[0178] The accompanying text for the aforementioned video clip reads: "At this crucial moment, our heroes not only face powerful enemies but also unravel the secrets of the ancient forest. As the battle concludes, a new challenge awaits them—exploring the legendary temple."

[0179] The accompanying text, including phrases like "critical moment" and "new challenge," emphasizes key points in the story's development, increasing tension and anticipation.

[0180] Next, the image recognition text will be combined with the image description text to form a more vivid and detailed image description text:

[0181] "In this mysterious ancient forest, our heroes not only face the challenge of a group of powerful enemies, but also have to unravel the secrets of the ancient forest. As the battle unfolds fiercely, the protagonists finally defeat the enemy. After the battle, they soon discover a stone gate, behind which lies the legendary forgotten temple, foreshadowing a new challenge that awaits them."

[0182] S23. Retrieve the object description text from the set material library to describe the object to be recommended, and generate corresponding object recommendation copy based on the object description text and the screen description text of each of the multiple video clips.

[0183] This involves retrieving relevant object content from a set material library based on query content related to the object to be recommended, such as content containing the object name, and then generating object description text based on the retrieved object content.

[0184] Specifically, to make the generated object description text more fluent and richer, Retrieval Augmentation (RAG) technology can be used. This involves retrieving object content related to the query from a pre-defined content library, then integrating the retrieved content into a pre-trained language generation model (e.g., a Transformer-based language model) to generate the object description text. RAG technology can automatically generate object description text highly relevant to the object to be recommended, facilitating accurate understanding of the recommended object.

[0185] The designated resource library can be a single library or multiple libraries of different types, such as the official website of the object to be recommended, social media platforms, third-party review websites, etc. It should be noted that access permission must be obtained for that resource library before querying its description text. The following describes the process of querying object description text from multiple different types of resource libraries.

[0186] In some embodiments, when the material library includes multiple different types of material libraries, object content matching the object to be recommended can be queried from each of the multiple different types of material libraries; the queried object content is then integrated to generate object description text.

[0187] The process involves retrieving relevant object content from multiple different types of content libraries based on query content related to the target object. Then, the object content from these libraries is integrated, including but not limited to removing duplicate content and sorting by relevance, which can be achieved through keyword matching, semantic matching, and other methods. Finally, a pre-trained language generation model is used to generate descriptive text for the integrated object content.

[0188] For example, taking games as the recommended target, first determine the different types of resource libraries that need to be accessed. These resource libraries include, but are not limited to: official game databases, which contain basic information about the game, character introductions, items, etc.; social media platforms, which contain player discussions about the game and can provide information about the game experience; and third-party review websites, which provide professional game reviews and analysis to help understand the game's unique selling points.

[0189] Based on the game query content, such as the game name, the system retrieves game-related content from multiple different types of resource libraries. For example, it obtains basic gameplay introductions and character skill parameters from official game databases, searches for trending player discussions on social media platforms, and retrieves professional game reviews from third-party review websites. The game content from these different resource libraries (i.e., the object content in the aforementioned embodiments) is then integrated. Based on the integrated game content, which includes information such as gameplay, characters, items, skills, and selling points, a pre-trained language generation model is used to generate an accurate and engaging game description text (i.e., the object description text in the aforementioned embodiments) based on the integrated game content.

[0190] In this embodiment, the multi-path collaborative retrieval method allows for a more comprehensive and accurate description of the object to be recommended. A pre-trained language generation model is used to generate a more fluent and natural description based on the retrieval results. Furthermore, the language style, depth, and other attributes of the generated content can be adjusted according to the characteristics of the target audience to meet personalized customization needs.

[0191] Based on the object description text queried in the above embodiments, and combined with the image description text of each of the multiple video segments identified in the aforementioned embodiments, object recommendation copy can be generated. The following embodiments will further describe the process of generating object recommendation copy.

[0192] In some embodiments, generating corresponding object recommendation text based on object description text and the respective scene description texts of multiple video segments in S23 of the above embodiments may include the following implementation steps:

[0193] The object description text and the individual scene description texts of multiple video clips are combined to generate copy generation prompts. Based on the copy generation prompts and the copy structure of the reference recommended copy, a pre-trained copy generation model is used to generate object recommendation copy.

[0194] In this embodiment of the application, to avoid the generated object recommendation copy lacking the visual content of each video clip, the visual description text of each video clip is extracted and included as part of the copy generation prompt information. Simultaneously, to avoid a lack of accurate and in-depth understanding of the recommended object due to a lack of background knowledge, the object description text retrieved in real-time is also included as part of the copy generation prompt information.

[0195] Optionally, depending on the object type of the object to be recommended, one or more existing recommendation texts with excellent performance (i.e., high volume) can be obtained from the object recommendation system for that object type as reference recommendation texts. In other words, the reference recommendation texts have a better recommendation effect in the object recommendation system. Different existing recommendation texts may have the same text structure or different text structures. Each text structure usually includes multiple structural tags. For example, taking the object type of the object to be recommended as a game, the text structure of an existing recommendation text can be: "Character -> Skill -> Benefit -> Call to Action", where "Character", "Skill", "Benefit", and "Call to Action" each represent a structural tag.

[0196] When the reference recommendation includes an existing recommendation, the copy generation prompts and the structure of the reference recommendation are input into the copy generation model. The model then parses the structure of the reference recommendation and uses natural language understanding to perform contextual understanding of the prompts. Based on these prompts, it generates a recommendation that matches the given structure. Specifically, for each structural tag in the copy structure, the model retrieves matching content from the prompts and generates the recommendation based on the retrieved content.

[0197] When the reference recommendation text includes multiple existing recommendation texts, the text generation model can integrate the text structures of the multiple existing recommendation texts. For example, it can remove duplicate structure tags to generate the final text structure, or it can select a text structure from the text structures of the multiple existing recommendation texts as the final text structure, and generate object recommendation text that conforms to the text structure based on the text generation prompt information.

[0198] Optionally, the copywriting generation model can employ a large language model. By referencing the copywriting structure of high-performing reference recommendation copy in the object recommendation system, the generated object recommendation copy can better meet the needs of object recommendation.

[0199] For example, taking the game "Pocket Monsters" as the target for recommendation, suppose the descriptive text of multiple video clips is integrated as follows:

[0200] "The first thing you see is the Pocket Warriors logo, accompanied by rousing background music, foreshadowing a thrilling adventure. The sound of 'flexible combination of land, sea, and air forces' accompanies the spectacular scene of the three forces appearing one after another. ... After a fierce battle, the player's side emerges victorious, and the word 'Victory' appears on the screen, accompanied by cheers and victory sound effects."

[0201] Suppose the game description text (i.e., object description text) for "Pocket Monsters" retrieved is as follows:

[0202] "In the game 'Pocket Warriors,' the main characters are divided into commanders and various troop units. Commanders can gain experience points and rewards by completing missions and campaigns, unlocking new skills and equipment. ... Players can freely combine land, sea, and air forces according to the actual situation to create a unique tactical system. High-quality game graphics and realistic sound effects bring players an immersive combat experience."

[0203] Let's assume the recommended copywriting structure is: "Role -> Skill -> Benefits -> Call to Action". For example... Figure 5 As shown, the copywriting generation prompt information is formed by combining the screen description text of each of the above video clips, the object description text obtained from the query, and the copywriting structure of the reference recommended copywriting. After inputting these into the copywriting generation model, the final generated game recommendation copywriting (i.e., object recommendation copywriting) is: "Experience Pocket Heroes in 10 seconds! Flexible combination of land, sea and air forces to showcase superior strategy and instantly become a commander! Exciting battlefields await your challenge!"

[0204] In this embodiment, the queried object description text and the screen description text of each of the multiple video clips are combined to generate copywriting generation prompts. Based on the copywriting generation prompts generated by the pre-trained copywriting generation model and combined with the copywriting structure of the reference recommended copywriting, object recommendation copywriting that meets the recommendation requirements can be generated quickly and accurately.

[0205] In some embodiments, to meet the customization and personalization needs of object recommendation copy, after generating candidate recommendation copy through a pre-trained copy generation model, the candidate recommendation copy can be presented in a copy preview interface to facilitate adjustments. In response to adjustments to the candidate recommendation copy, the adjusted candidate recommendation copy is obtained and used as the object recommendation copy. These adjustments include, but are not limited to, updating content and deleting content.

[0206] Specifically, after generating the object recommendation copy, the server can return it to the terminal device. The terminal device can respond to a preview operation by displaying a copy preview interface, which presents the object recommendation copy. This preview interface allows for adjustment options. Users can adjust the object recommendation copy by triggering these options. In response to the adjustment operation for the object recommendation video, the terminal device obtains the adjusted candidate recommendation copy and sends it to the server. The server then finalizes the candidate recommendation copy into the object recommendation copy to generate the object recommendation video.

[0207] For example, such as Figure 6 The image shown is a schematic of the text preview interface, displaying the candidate recommendation text "Experience Pocket Heroes in 10 seconds! Flexible combinations of land, sea, and air forces showcase superior strategy, instantly becoming a commander! Exciting battlefields await your challenge!" Editing and submit controls are also displayed. By triggering the editing control, the candidate recommendation text can be edited. After editing, the submit control can be triggered, at which point the terminal device receives the adjusted candidate recommendation text.

[0208] In this embodiment of the application, after generating candidate recommendation text, the candidate recommendation text can be presented to facilitate adjustment of the candidate recommendation text, and the adjusted candidate recommendation text can be used as the object recommendation text, thereby meeting the customization and personalization needs of object recommendation text.

[0209] In one optional implementation, in order to inform the user of the content composition of the generated candidate recommendation text and to make it easier to understand the content logic of the candidate recommendation text, the server can perform content decomposition and parsing on the generated candidate recommendation text to obtain the structured information of the candidate recommendation text. The structured information includes: the content attributes of each part of the candidate recommendation text; furthermore, the structured information of the candidate recommendation text can also be presented at the same time as the candidate recommendation text.

[0210] The server can employ a pre-trained text content segmentation model to analyze and segment the generated candidate recommendation texts. For example, the text content segmentation model can be obtained by fine-tuning a pre-trained large language model. Specifically, taking games as the type of object to be recommended, a large number of game-related recommendation texts can be collected, and these texts can be cleaned and structurally annotated. This involves annotating the true structured information of each game recommendation text (i.e., the content attributes of each part of the existing object recommendation text), obtaining the structural annotation results for each game recommendation text, and finally obtaining a second training sample set that meets the requirements. Each second training sample includes a game recommendation text and its corresponding structural annotation results. Based on the above second training sample set, the large language model is fine-tuned to obtain the text content segmentation model. The fine-tuning process of the large language model is described below.

[0211] The fine-tuning process of the large language model involves multiple rounds of iterative training. In each round, one or more second training samples are selected from the second training sample set and input into the large language model to be fine-tuned. This yields a structural prediction result for each second training sample, containing the predicted structural information of that sample. A predetermined loss function (e.g., cross-entropy loss) is used to calculate a second loss value based on the difference between the structural prediction result and the corresponding structural annotation result for each second training sample. Based on this second loss value, the parameters of the large language model are updated using gradient descent. When the iteration stopping condition is met, the currently fine-tuned large language model is used as the text content splitting model. The iteration stopping condition can be that the second loss value is not greater than a second predetermined threshold, or that the number of iterations reaches a second predetermined number.

[0212] After generating structured information for candidate recommendation texts, the server can send the candidate recommendation texts and their corresponding structured information to the terminal device. The terminal device can then display the candidate recommendation texts and their corresponding structured information in a text preview interface, according to its designated presentation method. Alternatively, the server can send the candidate recommendation texts and their corresponding structured information to the terminal device only upon receiving a text presentation request from the terminal device.

[0213] It should be noted that the structured information of the candidate recommendation text is different from the text structure of the reference recommendation text in the aforementioned embodiments. The division of this structured information is more detailed than the division of the text structure, and is used to show the content logic of the candidate recommendation text to the user.

[0214] The following is combined Figure 7 An example is provided to illustrate the structured information of candidate recommendation texts.

[0215] For example, such as Figure 7 As shown, taking the game "Endless Winter" as an example, the generated candidate recommendation text is: "Endless Winter, an adventure paradise for brave adventurers. Here, you will experience a thrilling survival challenge in the icy wilderness. Giant logs, large stoves, saws, chainsaws—all kinds of tools and props are at your disposal. With a wealth of resource management and survival strategies, you'll need to cut wood to build shelters, use fuel to maintain the stove's temperature, and fight off the cold. Teamwork is key, and the experience is intense and exciting. Build a warm town with your friends! If you want to experience the thrill of adventure, come to Endless Winter and conquer the polar world together!" Figure 7 The text uses rounded rectangles of different lines to divide the content of the candidate recommendation text, showcasing the content breakdown. Each type of rounded rectangle represents a content attribute, such as... Figure 7 The text includes phrases such as "game name," "selling points," "items," "others," and "call to action."

[0216] It should be noted that the content breakdown results of the candidate recommendation copy can also be displayed in other ways, such as using rounded rectangles (or other shapes) with different colored lines, or using different colored fonts to divide the content of the candidate recommendation copy into different parts. There are no limitations on this.

[0217] In this embodiment, by splitting and parsing the generated candidate recommendation text, the structured information of the candidate recommendation text can be presented at the same time, ensuring that the candidate recommendation text is logically clear and structurally reasonable. This allows users to intuitively understand the structure of the candidate recommendation text, making it easier to make targeted adjustments to the candidate recommendation text and further improve the user experience.

[0218] S24. Select at least one video clip from multiple video clips that matches the object recommendation copy, and generate an object recommendation video based on the selected at least one video clip and the object recommendation copy.

[0219] In this application, considering that the generation of the object recommendation copy references the visual description text of multiple video clips—meaning the object recommendation copy may contain visual content from multiple video clips—to accurately select at least one video clip that semantically matches the object recommendation copy, the object recommendation copy can be divided into at least one copy fragment. Then, each copy fragment is semantically matched with the multiple video clips; that is, the semantic similarity between each copy fragment and the multiple video clips is calculated, and at least one video clip whose semantic similarity meets the set conditions is selected. The semantic matching process will be described in detail in subsequent embodiments of this application.

[0220] It should be noted that, in addition to semantically matching each fragment of the object recommendation copy with multiple video clips, the entire object recommendation copy can also be semantically matched with multiple video clips. That is, the semantic similarity between the entire object recommendation copy and multiple video clips can be calculated, and at least one video clip whose semantic similarity meets the set conditions can be selected from the multiple video clips.

[0221] Typically, multiple video clips can be selected, but it's possible to select only one. Taking multiple video clips as an example, to generate a recommended video, we can obtain the cleaned video clips corresponding to each of these clips. The cleaned video clips remove unnecessary elements from the original video clips, such as the original audio, subtitles, and badges. The multiple cleaned video clips are then spliced ​​together in a predetermined order to obtain a spliced ​​video. The recommended text is then converted into corresponding recommended audio. Finally, the recommended audio is combined with the spliced ​​video to generate the recommended video.

[0222] In this embodiment, the retrieved object description text and the individual scene description texts of multiple video clips are combined to accurately understand the object to be recommended, thereby generating high-quality object recommendation copy. Next, at least one video clip is selected from the multiple video clips that semantically matches the object recommendation copy, ensuring a high degree of matching between the selected video clip and the object recommendation copy. Finally, based on the selected at least one video clip and the object recommendation copy, a high-quality object recommendation video can be generated. Therefore, this embodiment improves the quality of generated object recommendation videos.

[0223] The following example provides a detailed description of the matching process between video clips and recommended content.

[0224] In some embodiments, the process of selecting at least one video clip from multiple video clips that matches the recommended text for the object in step S24 above may include the following steps B1-B3:

[0225] B1. Extract the video features of each of the multiple video segments, and divide the object recommendation copy into at least one copy segment, and extract the copy features of each of the at least one copy segment.

[0226] In this process, semantic features can be extracted from the descriptive text of each video segment using a feature extraction model, and the obtained text features can be used as the video features of the video segment. For example, feature extraction models include, but are not limited to, the following text feature extraction models: Word to Vector (Word2Vec) and BERT.

[0227] In some embodiments, to extract video features of each video segment more accurately, one or more video frames can be extracted from each video segment while extracting text features of each video segment, and image features of one or more video frames can be extracted. The text features and at least one image feature of each video segment can then be fused. In this case, when extracting video features of multiple video segments in step B1 above, the following steps B11-B12 can be performed separately for each video segment:

[0228] B11. Extract text features from the descriptive text of a video clip to obtain text features, and extract image features from at least one video frame in the video clip to obtain at least one image feature.

[0229] In this step, the feature extraction model mentioned in the previous embodiments can be used to extract text features and at least one image feature from the video clip. The feature extraction model includes, but is not limited to, Contrastive Language-Image Pre-training (CLIP) and Vision-and-Language BERT (ViLBERT). CLIP uses contrastive learning to jointly train the image encoder and text encoder, ensuring that similar images and the text describing them are close in the embedding space, thus enabling the simultaneous extraction of text and image features. ViLBERT is a two-stream BERT architecture that can process image regions and text sequences separately, then combine them through a cross-modal attention mechanism, thereby simultaneously extracting text and image features.

[0230] B12. The text features and at least one image feature corresponding to the above video segment are fused to obtain the video features of the video segment.

[0231] The fusion processing method can be selected according to needs. Several optional fusion processing methods are introduced below.

[0232] For example, text features and image features are typically in vector form. The text features and at least one image feature are directly summed or weighted summed. In this case, it's necessary to ensure that the text features and image features have the same number of dimensions. This allows for the summation or weighted summation of feature values ​​of the same dimensions from the text features and image features to obtain the video features of the corresponding video segment. When weighted summing of text features and at least one image feature, the weights of the text features and each image feature can be set as needed. For example, the weight of the text feature can be 0.7, and the weight of the image feature can be 0.3. The weights of different image features can be the same or different; there is no limitation on this.

[0233] For example, text features and at least one image feature can be directly concatenated, or text features and at least one image feature can be multiplied by their respective weights and then concatenated, with the concatenated features serving as the video features of the corresponding video segment.

[0234] In this embodiment of the application, for each video segment, when extracting the video features of the video segment, the text features of the screen description text of the video segment and the image features of at least one video frame in the video segment can be extracted simultaneously; and the text features and at least one image features are fused together. By fusing multimodal information (i.e., text features and image features), the video features of the video segment can be obtained more accurately, so as to make the video segment and the text segment more accurately matched in the future.

[0235] Furthermore, in step B1 above, a predefined method can be used to divide the object recommendation text into at least one text segment. For example, the object recommendation text can be input into a pre-trained text segmentation model, which performs semantic analysis on the object recommendation text and divides it into at least one text segment based on the semantic analysis results; wherein, the text segmentation model can be a large language model.

[0236] For example, taking the game "Uncharted Adventure" from the aforementioned embodiments as an example, the game's recommendation text (i.e., the target recommendation text) is divided into four text fragments using a text fragment segmentation model. The text attributes of these fragments are: introduction, game plot, benefits introduction, and player call. The introduction text reads: "In the distant depths of the ocean, countless undiscovered secrets and treasures lie hidden. Begin your legendary journey!" The game plot text reads: "As a young explorer, you will traverse dense rainforests, delve into ancient underground ruins, and face various challenges." The benefits introduction text reads: "To celebrate the arrival of new players, 'Uncharted Adventure' has prepared a series of valuable benefits, with rich rewards awaiting you." The player call text reads: "Whether you explore alone or with partners, 'Uncharted Adventure' can meet your needs. Gather your team and set off together!"

[0237] It should be noted that if the above-mentioned object recommendation copy has already been content-segmented in the foregoing embodiments, such as using the copy content segmentation model in the foregoing embodiments to segment the above-mentioned object recommendation copy, the content segmentation result, i.e. the structured information of the object recommendation copy in the foregoing embodiments, can be directly used to divide the object recommendation copy into at least one copy fragment.

[0238] For each text segment, semantic features can also be extracted using the feature extraction model described in the previous embodiments to obtain the corresponding text features. In other words, to more accurately match text segments and video segments semantically, the same feature extraction model can be used to map text segments and video segments to the same feature space, such as... Figure 8As shown, the images of each video frame, the descriptive text, and the text fragments of a video clip can be input into the same feature extraction model for feature extraction, thereby obtaining the image features, text features, and text features respectively. The image features and text features are then fused to obtain the video features of the aforementioned video clip.

[0239] Furthermore, to improve the efficiency of text feature extraction and avoid irrelevant content in text segments interfering with the semantic matching of subsequent text and video segments, one or more keywords can be extracted from each text segment. Then, the semantic features of one or more keywords are extracted using the aforementioned feature extraction model, serving as the text features of that text segment. Specifically, after dividing the object recommendation text into at least one text segment using the text segment segmentation model of the aforementioned embodiment, one or more keywords can be extracted from each text segment. For example, taking a game scene as an example, the keywords of a text segment may include: character name, skills, game scene, and suitable storyboard category, etc. Among them, the suitable storyboard category indicates the selection of the visual style and narrative technique that best enhances the effect based on different types of text content. For example, suitable storyboard categories include narrative, action, emotional rendering, etc.

[0240] B2. For at least one text fragment, perform the following operations: Based on the text features of a text fragment, compare it with the video features of multiple video fragments, and select at least one video fragment that matches the text fragment.

[0241] The text features and video features are usually in the form of vectors. Therefore, vector similarity calculation methods can be used to calculate the similarity between the text features and each video feature. For example, similarity calculation methods include, but are not limited to, cosine similarity, Euclidean distance, Manhattan distance, etc.

[0242] Specifically, at least one video clip that meets a first set condition in terms of similarity can be selected from multiple video clips as at least one video clip that matches a text clip. For example, the first set condition can be: the top N with the highest similarity, where N is an integer greater than or equal to 1, or it can be: video clips with a similarity reaching a similarity threshold. This similarity threshold can be set as needed and is not specifically limited.

[0243] It should be noted that when the similarity between the video features of a certain video clip and different text features all meet the first set condition, the text clip corresponding to the text feature with the highest similarity can be used as the text clip that matches the video clip.

[0244] B3. Based on at least one video segment selected for at least one text segment, obtain at least one video segment whose semantics match the recommended text of the object.

[0245] In this embodiment of the application, the object recommendation text is divided into at least one text segment, and for each text segment, based on the text features of the text segment, the similarity between the text features of the text segment and the video features of multiple video segments is used to select at least one video segment that matches the text segment from the multiple video segments, thereby obtaining each video segment that matches the object recommendation text more accurately.

[0246] In an optional implementation, to further meet the customization requirements of the object recommendation video, the terminal device can also receive input video generation constraints and then send the video generation constraints to the server. When the server receives the video generation constraints, if it determines that multiple video segments are selected for the above-mentioned at least one text fragment, it further selects at least one video segment that meets the video generation constraints from the selected multiple video segments; and the at least one video segment selected again is taken as at least one video segment that semantically matches the object recommendation text.

[0247] The video generation constraints can be set arbitrarily, including but not limited to: video aspect ratio, video size, video duration, number of video segments, and object information related to the object to be recommended. Video aspect ratio refers to the ratio between the width and height of the video frame, such as 16:9 or 9:16; video size refers to the actual resolution of the video frames. For example, if the object to be recommended is a game, the above object information could include selling points and other relevant information.

[0248] In this embodiment of the application, at least one video segment is selected by filtering out video generation constraints, which can meet the customized needs of object-recommended videos and improve user experience.

[0249] In other embodiments, besides semantically matching each fragment of the object recommendation text with multiple video fragments, the entire object recommendation text can also be semantically matched with multiple video fragments. That is, the semantic similarity between the entire object recommendation text and multiple video fragments is calculated. Specifically, a feature extraction model can be used to extract video features from each video fragment and perform semantic feature extraction on the entire object recommendation text to obtain overall text features. The similarity (i.e., semantic similarity) between the video features of each of the multiple video fragments and the overall text features is calculated. At least one video fragment whose similarity satisfies a second predetermined condition is then selected from the multiple video fragments. For example, the second predetermined condition could be the top M most similar fragments, where M is an integer greater than or equal to 1, or it could be video fragments whose similarity reaches a similarity threshold. This similarity threshold can be set as needed and is not specifically limited.

[0250] In some embodiments, when at least one matching video segment is selected from multiple video segments for each of the above text segments, in the above embodiment S24, when generating an object recommendation video based on the selected at least one video segment and the object recommendation text, in order to achieve audio-visual synchronization of the object recommendation video, that is, to align the voice-over of the object recommendation text with each video segment, the following steps C1-C2 can be performed for each at least one text segment:

[0251] C1. Obtain at least one video segment that matches a text clip, and the corresponding cleaned video segments. Convert a text clip into text-to-speech, and splice and edit the at least one matching cleaned video clip according to the duration of the text-to-speech, to obtain the corresponding edited video.

[0252] If the video material has already been cleaned in the aforementioned embodiments, i.e., unnecessary elements such as original subtitles, audio, and corner marks have been removed from the video material, then after selecting at least one video segment for each text segment, the cleaned video segments corresponding to these video segments can be obtained directly.

[0253] Specifically, this involves converting text fragments into spoken text based on existing speech synthesis methods. For example, speech synthesis methods can employ deep learning models to transform the input text fragments into natural and fluent human speech (i.e., spoken text). The conversion process typically includes the following steps: text preprocessing, cleaning and standardizing the input text, including removing irrelevant characters and converting numbers into word forms; front-end processing, performing text analysis such as word segmentation, part-of-speech tagging, and prosodic prediction; acoustic modeling, generating acoustic features of the speech (such as Mel spectrograms) based on the results of front-end processing; back-end processing, converting the acoustic features into actual speech waveforms; and post-processing, adjusting the generated speech, such as adding appropriate pauses and adjusting the volume.

[0254] After converting each text segment into speech, at least one matching cleaned video segment can be spliced ​​together. If a text segment corresponds to multiple cleaned video segments, these video segments can be arranged according to their similarity to the text segment, such as in descending order of similarity, and then spliced ​​together in that order.

[0255] Furthermore, if the thematic scenes of each video segment were categorized in the aforementioned embodiments, then when splicing the multiple cleaned video segments, the order of these video segments can be determined based on their similarity to the matching text segments and their respective thematic scene categories. For example, the display priority of different thematic scene categories can be preset, placing video segments corresponding to higher-priority thematic scene categories first and those corresponding to lower-priority thematic scene categories later. If different video segments have the same thematic scene category or the same priority, they can be arranged according to their similarity, such as in descending order of similarity.

[0256] For each text segment, multiple cleaned video segments are spliced ​​together. The length of the resulting spliced ​​video is then edited to match the length of the corresponding text / audio, achieving more precise audio-visual synchronization. Specifically, when the length of the spliced ​​video exceeds the length of the text / audio, it can be trimmed to match the length, for example, by deleting unnecessary or redundant video frames to shorten the video's length. When the length of the spliced ​​video is shorter than the length of the text / audio, it can be extended to match the length, for example, by adding transition effects or inserting visual effects between different video segments.

[0257] C1. Based on at least one video clip and the corresponding text / voice for each of the at least one video clips, obtain recommended videos for the target audience.

[0258] Specifically, at least one edited video is spliced ​​together to obtain the final spliced ​​video, and the spliced ​​video is dubbed based on the above-mentioned text and voice. In addition, specific visual elements, subtitles, background music, transitions (i.e., the effect of smoothly transitioning from one theme scene to another) can be added to the final spliced ​​video as needed, and the final output is a recommended video.

[0259] In this embodiment, for each text segment, the text segment is converted into text-to-speech, and according to the duration of the text-to-speech, at least one matching cleaned video segment is spliced ​​and edited to obtain a corresponding edited video. This facilitates the alignment of the edited video with the corresponding text-to-speech, and then, based on the obtained at least one edited video and their respective corresponding text-to-speech, an object recommendation video is generated, which can achieve accurate audio-visual synchronization, enhance the attractiveness of the object recommendation video, and improve the user's viewing experience.

[0260] In some embodiments, after obtaining the object-recommended video, the object-recommended video can also be presented in the video preview interface. In response to the adjustment operation of the object-recommended video, the adjusted object-recommended video can be obtained, thereby meeting different video creative needs and improving the personalization of the object-recommended video.

[0261] After obtaining the recommended video, the server can return it to the terminal device. The terminal device can then preview the video, displaying it in a preview interface. This preview interface allows for editing options, such as adjusting video segments, changing background music, and adding text. Upon receiving these adjustments, the terminal device obtains the adjusted recommended video and sends it back to the server for distribution.

[0262] like Figure 9 The image shows a schematic diagram of the video preview interface for a recommended video. In this preview interface, the recommended video can be played frame by frame. The "Edit" control allows for cropping the recommended video, while the "Audio," "Text," and "Effects" controls allow for adding audio, text, and effects. After editing the recommended video, the "Done" control can be triggered, at which point the terminal device can export the adjusted recommended video. It should be noted that the above video preview interface is only an example; different interface presentation methods can be set as needed.

[0263] The video processing method in this application embodiment enhances the attractiveness and recommendation effect of object recommendation videos by automatically generating object recommendation copy, and can efficiently generate high-quality object recommendation videos for rapid deployment.

[0264] The following is combined Figure 10 The system architecture for implementing the video processing method of the embodiments of this application is described.

[0265] like Figure 10 As shown, the system architecture for implementing the video processing method of this application mainly includes three modules: a video material understanding module, a text generation module, and an audio-visual matching module. The video material understanding module segments the video material into scenes, then performs in-depth understanding on each video segment to obtain the understanding result of each video segment (i.e., the image description text in the aforementioned embodiment). The text generation module can query relevant content of the object to be recommended in real time to obtain query results (i.e., the object description text in the aforementioned embodiment), and generate object recommendation text based on the query results and the understanding results of each video segment. Finally, the audio-visual matching module matches the object recommendation text with each video segment after segmentation to obtain video segments that match the object recommendation text, ultimately generating the object recommendation video.

[0266] The following uses a game scenario as an example to illustrate the usage process of a game video editing tool that applies the video processing method of this application.

[0267] like Figure 11 As shown, users can upload game video footage that needs to be edited using the game video editing tool. The tool will automatically generate a game video edit (i.e., the object recommendation video in the aforementioned embodiment). The specific operation steps are as follows:

[0268] 1. User-uploaded video materials and input video generation restrictions.

[0269] The video generation restrictions include certain limitations on the final generated game montage video, such as selling points, video aspect ratio, video size, and video duration.

[0270] 2. Automatic processing of video footage, specifically including the following workflow:

[0271] (1) Clean the video footage, such as removing subtitles, audio, and corner marks from the video footage;

[0272] (2) Divide the cleaned video material into segments to obtain multiple video clips with corresponding thematic scenes;

[0273] (3) Understanding video materials. Perform image recognition on multiple video clips to obtain image description text for each video clip; specifically, identify the image content and theme / scene category of each video clip to generate image description text; the theme / scene category of each video clip can also be labeled.

[0274] 3. Game script (i.e., the object recommendation text in the aforementioned embodiments) generation.

[0275] The individual scene description texts of multiple video segments, after understanding the video material, are combined with the game introduction text retrieved based on RGA technology (i.e., the object description text in the aforementioned embodiment) to form a prompt (i.e., the text generation prompt information in the aforementioned embodiment). The text generation model can then form a game script based on the prompt and the script structure of the high-quality material script (i.e., the text structure of the reference recommended text in the aforementioned embodiment). The above process includes understanding the high-quality material script to extract its script structure.

[0276] 4. Video Segment Selection. First, select video segments from the multiple video segments behind the scenes that match the semantics of the game script. Then, select video segments that meet the video generation constraints from these video segments, thereby controlling the size, length, and other information of the game montage video.

[0277] 5. Output a game mashup video. Based on the generated game script, the final selected video clips (here, the cleaned video clips) are spliced ​​together, and specific elements, subtitles, voice-overs, transitions, etc., are added to output a game mashup video.

[0278] 6. Preview and Adjustment. Users can preview the generated game mashup video, make manual adjustments, and finally export the adjusted game mashup video and upload it to the recommended platform.

[0279] For example, such as Figure 12A The image shows the initial interface of the aforementioned game video mashup tool. On this initial interface, users can upload video footage (one or more clips) to be mashed up by clicking the "Upload Video" control. Clicking "Quickly Generate Script" displays the generated game script, and users can select the aspect ratio of the final game mashup video. Additionally, other restrictions can be set for the game mashup video (not shown in the image). Clicking the "Generate Video" control displays the final game mashup video. Figure 12B The image shown is a schematic diagram illustrating the result of a game video mashup generated by a game video mashup tool, including the generated game script and the game mashup video.

[0280] It should be noted that, Figure 12A and Figure 12B This is merely an example; the presentation of the interface can be customized according to actual needs, and no limitations are imposed.

[0281] The aforementioned game video mashup tool employs the video processing method described in this application embodiment, enabling intelligent editing. This involves automatically identifying and editing exciting video clips from video materials, and integrating game scripts and background music as needed to achieve precise beat-sync mashup editing, ensuring audio-visual synchronization and further enhancing the appeal and viewing experience of the game mashup video. The generated game script can be adjusted through the script preview interface, and the generated game mashup video can be fine-tuned and customized through the video preview interface to meet specific video creative needs. Furthermore, the final generated game mashup video can be automatically uploaded to a recommendation platform.

[0282] Based on the same inventive concept, this application also provides a video processing device. The principle of the device in solving the problem is similar to the method of the above embodiments. Therefore, the implementation of the device can refer to the implementation of the above method, and the repeated parts will not be described again.

[0283] like Figure 13 As shown, this is a structural schematic diagram of the video processing device 1300, which may include:

[0284] The segmentation unit 1301 is used to perform segmentation processing on the video material of the object to be recommended, and obtain multiple video segments.

[0285] The recognition unit 1302 is used to perform image recognition on multiple video segments respectively and obtain image description text for each video segment; each image description text is used to describe the image content of the corresponding video segment;

[0286] The copywriting generation unit 1303 is used to retrieve object description text from the set material library to describe the object to be recommended, and generate corresponding object recommendation copy based on the object description text and the screen description text of each of the multiple video clips.

[0287] The video generation unit 1304 is used to select at least one video segment from multiple video segments that semantically matches the object recommendation copy, and generate an object recommendation video based on the selected at least one video segment and the object recommendation copy.

[0288] In this embodiment, the retrieved object description text and the individual scene description texts of multiple video clips are combined to accurately understand the object to be recommended, thereby generating high-quality object recommendation copy. Next, at least one video clip is selected from the multiple video clips that semantically matches the object recommendation copy, ensuring a high degree of matching between the selected video clip and the object recommendation copy. Finally, based on the selected at least one video clip and the object recommendation copy, a high-quality object recommendation video can be generated. Therefore, this embodiment improves the quality of generated object recommendation videos.

[0289] Optionally, the identification unit 1302 is specifically used for:

[0290] Perform the following operations on each of the multiple video clips:

[0291] Perform image recognition on a video clip to obtain the corresponding image recognition text, as well as the image description text associated with the video clip;

[0292] By integrating the on-screen description text with the on-screen recognition text, a visual description text for a video clip is obtained.

[0293] Optionally, each video segment has a corresponding theme scene; when performing image recognition on a video segment to obtain the corresponding image recognition text, the recognition unit 1302 is specifically used for:

[0294] Perform image content recognition on a video clip to obtain the image content of the video clip;

[0295] Classify a video clip into thematic scene categories to obtain the thematic scene category of the video clip;

[0296] Based on the content and theme / scene category of a video clip, generate corresponding image recognition text.

[0297] Optionally, when acquiring the on-screen description text associated with a video clip, the recognition unit 1302 is specifically used for:

[0298] Based on at least one of the on-screen annotation text and on-screen background text contained in a video clip, obtain the on-screen description text associated with the video clip;

[0299] The on-screen annotation text is obtained by performing text recognition on at least one video frame in the video clip, while the on-screen background text is obtained by performing speech recognition on the on-screen background audio associated with the video clip.

[0300] Optionally, when retrieving object description text from the set material library to describe the object to be recommended, the copy generation unit 1303 is specifically used for:

[0301] When the material library includes multiple different types of material libraries, the content of the object that matches the object to be recommended is queried from each of the multiple different types of material libraries;

[0302] The content of multiple retrieved objects is integrated to generate object description text.

[0303] Optionally, when generating corresponding object recommendation copy based on object description text and the individual scene description texts of multiple video clips, the copy generation unit 1303 is specifically used for:

[0304] Combine the object description text and the individual scene description texts of multiple video clips to generate copy prompts;

[0305] By using a pre-trained copywriting generation model, based on copywriting generation prompts and the copywriting structure of reference recommended copywriting, target-specific recommended copywriting is generated.

[0306] Optionally, when generating recommended copywriting based on copywriting generation prompts and the copywriting structure of reference recommended copywriting using a pre-trained copywriting generation model, the copywriting generation unit 1303 is specifically used for:

[0307] Using a copy generation model, based on copy generation prompts and the copy structure of reference recommended copy, candidate recommended copy is generated and presented in the copy preview interface.

[0308] In response to the adjustment operation on the candidate recommendation copy, the adjusted candidate recommendation copy is obtained and used as the target recommendation copy.

[0309] Optionally, when presenting candidate recommended texts in the text preview interface, the text generation unit 1303 is specifically used for:

[0310] The candidate recommendation texts are broken down and analyzed to obtain their structured information; the structured information includes the content attributes of each part of the candidate recommendation text.

[0311] The preview interface displays candidate recommended texts and structured information.

[0312] Optionally, when selecting at least one video clip from multiple video clips that semantically matches the object recommendation text, the video generation unit 1304 is specifically used for:

[0313] Extract video features from multiple video segments respectively, and divide the object recommendation text into at least one text segment, and extract text features from at least one text segment respectively;

[0314] For at least one text fragment, perform the following operations respectively: based on the text features of a text fragment, compare it with the video features of multiple video fragments respectively, and select at least one video fragment that matches a text fragment from the multiple video fragments.

[0315] Based on at least one video segment selected for at least one text fragment, obtain at least one video segment whose semantics match the recommended text for the object.

[0316] Optionally, when extracting video features from multiple video segments separately, the video generation unit 1304 is specifically used for:

[0317] Perform the following operations on each of the multiple video clips:

[0318] Text features are extracted from the descriptive text of a video clip to obtain text features, and image features are extracted from at least one video frame in a video clip to obtain at least one image feature.

[0319] By fusing text features and at least one image feature, the video features of a video segment are obtained.

[0320] Optionally, when obtaining at least one video segment that semantically matches the recommended text based on at least one video segment selected for at least one text segment, the video generation unit 1304 is specifically used for:

[0321] When multiple video clips are selected for at least one text segment, and the set video generation constraints are received, at least one video clip that meets the video generation constraints is selected again from the multiple selected video clips.

[0322] At least one video clip that is selected again will be considered as at least one video clip that semantically matches the object recommendation copy.

[0323] Optionally, when generating an object recommendation video based on at least one selected video clip and combined with the object recommendation text, the video generation unit 1304 is specifically used for:

[0324] For at least one text fragment, perform the following operations respectively: obtain at least one video fragment that matches a text fragment, the corresponding cleaned video fragments, convert a text fragment into text speech, and according to the duration of the text speech, splice and edit the at least one matching cleaned video fragment to obtain the corresponding edited video.

[0325] Based on at least one video clip and the corresponding text / voice for each video clip, recommended videos for the target audience are obtained.

[0326] Optionally, the device also includes an adjustment unit for:

[0327] The video preview interface displays recommended videos for the target audience;

[0328] In response to the adjustment operation for the recommended video of the object, obtain the adjusted recommended video of the object.

[0329] Optionally, the screen splitting unit 1301 is specifically used for:

[0330] The video footage is split into multiple segments; each video frame in each segment has the same theme and scene.

[0331] When the duration of each segment in at least one of the multiple segment segments is not greater than the set duration, each segment in at least one segment segment is spliced ​​with the adjacent segment segment.

[0332] Multiple video segments are obtained based on the spliced ​​segment segments and the unspliced ​​segment segments.

[0333] For ease of description, the above sections are divided into modules (or units) according to their functions and described separately. Of course, in implementing this application, the functions of each module (or unit) can be implemented in one or more software or hardware components.

[0334] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0335] Having introduced the video processing method and apparatus according to exemplary embodiments of this application, we will now introduce an electronic device according to another exemplary embodiment of this application.

[0336] Based on the same inventive concept as the above-described method embodiments, this application also provides an electronic device. In one embodiment, the electronic device may be a server, such as... Figure 1 The server 120 is shown. In this embodiment, the structure of the electronic device can be as follows: Figure 14 As shown, it includes a memory 1401, a communication module 1403, and one or more processors 1402.

[0337] The memory 1401 is used to store computer programs executed by the processor 1402. The memory 1401 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0338] Memory 1401 may be volatile memory, such as random-access memory (RAM); memory 1401 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1401 may be any other medium capable of carrying or storing a desired computer program having the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1401 may be a combination of the above-described memories.

[0339] The processor 1402 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 1402 is used to implement the aforementioned video processing method when it calls a computer program stored in the memory 1401.

[0340] The communication module 1403 is used to communicate with other devices.

[0341] This application embodiment does not limit the specific connection medium between the memory 1401, communication module 1403, and processor 1402. This application embodiment... Figure 14 The memory 1401 and the processor 1402 are connected via a bus 1404, and the bus 1404 is in Figure 14 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 1404 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 14 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.

[0342] The memory 1401 stores a computer-readable storage medium containing a computer program for implementing the video processing method of the embodiments of this application. The processor 1502 is used to execute the aforementioned video processing method, such as... Figure 2 As shown.

[0343] In another embodiment, the electronic device can also be other electronic devices, such as... Figure 1 The terminal device 110 is shown. In this embodiment, the electronic device can be structured as follows: Figure 15 As shown, it includes components such as: communication component 1510, memory 1520, display unit 1530, camera 1540, sensor 1550, audio circuit 1560, Bluetooth module 1570, processor 1580, etc.

[0344] The communication component 1510 is used to communicate with the server. In some embodiments, it may include a Circuit-Based Wireless Fidelity (WiFi) module, which is a short-range wireless transmission technology. Electronic devices can use the WiFi module to help users send and receive information.

[0345] The memory 1520 can be used to store software programs and data. The processor 1580 executes various functions of the terminal device 110 and performs data processing by running the software programs or data stored in the memory 1520. The memory 1520 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1520 stores an operating system that enables the terminal device 110 to run. In this application, the memory 1520 may store the operating system and various application programs, and may also store computer programs that execute the video processing method of the embodiments of this application.

[0346] The display unit 1530 can also be used to display information input by the user or information provided to the user, as well as various menus of the terminal device 110, in a graphical user interface (GUI). Specifically, the display unit 1530 may include a display screen 1532 disposed on the front of the terminal device 110. The display screen 1532 may be configured as a liquid crystal display, a light-emitting diode, or the like. The display unit 1530 can be used to display object recommendation videos and object recommendation text, etc., as described in the embodiments of this application.

[0347] The display unit 1530 can also be used to receive input digital or character information and generate signal inputs related to user settings and function control of the terminal device 110. Specifically, the display unit 1530 may include a touch screen 1531 disposed on the front of the terminal device 110, which can collect touch operations of the user on or near it, such as clicking buttons, dragging scroll boxes, etc.

[0348] The touchscreen 1531 can be placed on top of the display screen 1532, or the touchscreen 1531 and the display screen 1532 can be integrated to realize the input and output functions of the terminal device 110. After integration, it can be referred to as a touch display screen. In this application, the display unit 1530 can display the application and the corresponding operation steps.

[0349] Camera 1540 can be used to capture still images, which users can then share via an application. There can be one or multiple cameras 1540. An object is projected onto a photosensitive element through a lens, generating an optical image. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the processor 1580 for conversion into a digital image signal.

[0350] The terminal device may also include at least one sensor 1550, such as an accelerometer 1551, a proximity sensor 1552, a fingerprint sensor 1553, and a temperature sensor 1554. The terminal device may also be equipped with other sensors such as a gyroscope, barometer, hygrometer, thermometer, infrared sensor, light sensor, and motion sensor.

[0351] Audio circuitry 1560, speaker 1561, and microphone 1562 provide an audio interface between the user and terminal device 110. Audio circuitry 1560 converts received audio data into electrical signals, which are then transmitted to speaker 1561, where they are converted into sound signals for output. Terminal device 110 may also be equipped with volume buttons for adjusting the volume of the sound signal. Conversely, microphone 1562 converts collected sound signals into electrical signals, which are then received by audio circuitry 1560, converted back into audio data, and output to communication component 1510 for transmission to, for example, another terminal device 110, or to memory 1520 for further processing.

[0352] The Bluetooth module 1570 is used to interact with other Bluetooth devices that also have a Bluetooth module via the Bluetooth protocol. For example, a terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smartwatch) that also has a Bluetooth module through the Bluetooth module 1570, thereby exchanging data.

[0353] The processor 1580 is the control center of the terminal device, connecting various parts of the terminal through various interfaces and lines. It executes various functions and processes data by running or executing software programs stored in the memory 1520 and calling data stored in the memory 1520. In some embodiments, the processor 1580 may include one or more processing units; the processor 1580 may also integrate an application processor and a baseband processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the baseband processor mainly handles wireless communication. It is understood that the baseband processor may not be integrated into the processor 1580. In this application, the processor 1580 can run an operating system, applications, user interface display and touch response, and the video processing method of this embodiment. Furthermore, the processor 1580 is coupled to the display unit 1530.

[0354] In some possible implementations, various aspects of the video processing method provided in this application can also be implemented in the form of a computer program product, which includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps of the video processing method according to the various exemplary embodiments of this application described above. For example, the electronic device can perform actions such as... Figure 2 The steps are shown in the figure.

[0355] Computer program products may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0356] The computer program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on an electronic device. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with a command execution system, apparatus, or device.

[0357] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.

[0358] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0359] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The computer program can execute entirely on the user's electronic device, partially on the user's electronic device, as a standalone software package, partially on the user's electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In cases involving remote electronic devices, the remote electronic device can be connected to the user's electronic device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external electronic device (e.g., via the Internet using an Internet service provider).

[0360] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0361] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0362] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing a computer-usable computer program.

[0363] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0364] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A video processing method, characterized in that, The method includes: The video footage of the recommended target is split into multiple video segments; Each of the multiple video segments is subjected to image recognition to obtain image description text for each video segment; each image description text is used to describe the image content of the corresponding video segment. The object description text used to describe the object to be recommended is retrieved from the set material library, and corresponding object recommendation copy is generated based on the object description text and the screen description text of each of the multiple video clips; From the plurality of video clips, at least one video clip is selected that semantically matches the object recommendation copy, and an object recommendation video is generated based on the selected at least one video clip and the object recommendation copy.

2. The method according to claim 1, characterized in that, The step of performing image recognition on the plurality of video segments respectively to obtain image description text for each of the plurality of video segments includes: For each of the multiple video segments, perform the following operations: Perform image recognition on a video clip to obtain the corresponding image recognition text, as well as obtain the image description text associated with the video clip; The introductory text of the video clip is integrated with the recognized text of the video clip to obtain the introductory text of the video clip.

3. The method according to claim 2, characterized in that, Each video segment has a corresponding theme scene; the step of performing image recognition on a video segment to obtain corresponding image recognition text includes: Perform image content recognition on a video clip to obtain the image content of the video clip; The video clip is classified into thematic scene categories to obtain the thematic scene category of the video clip. Based on the content and theme / scene category of a video clip, generate corresponding image recognition text.

4. The method according to claim 2, characterized in that, The step of obtaining the on-screen description text associated with the video segment includes: Based on at least one of the image annotation text and image background text contained in the video clip, obtain the image description text associated with the video clip; The image annotation text is obtained by performing text recognition on at least one video frame in the video segment, and the image background text is obtained by performing speech recognition on the image background speech associated with the video segment.

5. The method according to claim 1, characterized in that, The step of retrieving the object description text from the designated material library to describe the object to be recommended includes: When the set material library includes multiple different types of material libraries, the content of objects that match the object to be recommended is queried from each of the multiple different types of material libraries; The content of multiple retrieved objects is integrated to generate the object description text.

6. The method according to any one of claims 1 to 5, characterized in that, The step of generating corresponding object recommendation text based on the object description text and the respective scene description texts of the multiple video clips includes: The object description text and the screen description text of each of the multiple video segments are combined to generate a text generation prompt message; The pre-trained copywriting generation model generates recommended copywriting for the target based on the copywriting generation prompts and the copywriting structure of the reference recommended copywriting.

7. The method according to claim 6, characterized in that, The pre-trained copywriting generation model generates the target recommendation copy based on the copywriting generation prompt information and the copywriting structure of the reference recommendation copy, including: Based on the copy generation prompts and the copy structure of the reference recommended copy, the copy generation model generates candidate recommended copy and presents the candidate recommended copy in the copy preview interface. In response to the adjustment operation on the candidate recommendation copy, the adjusted candidate recommendation copy is obtained, and the adjusted candidate recommendation copy is used as the object recommendation copy.

8. The method according to claim 7, characterized in that, Presenting the candidate recommended text in the text preview interface includes: The candidate recommendation text is broken down and parsed to obtain its structured information; wherein, the structured information includes the content attributes of each part of the candidate recommendation text. The candidate recommended texts and the structured information are presented in the text preview interface.

9. The method according to any one of claims 1 to 5, characterized in that, The step of selecting at least one video segment from the plurality of video segments that semantically matches the recommended text for the object includes: The video features of each of the multiple video segments are extracted respectively, and the object recommendation text is divided into at least one text segment, and the text features of each of the at least one text segment are extracted respectively. For the at least one text segment, perform the following operations respectively: based on the text features of a text segment, and the similarity between the text features and the video features of the plurality of video segments respectively, select at least one video segment that matches the text segment from the plurality of video segments; Based on at least one video segment selected for the at least one text segment, at least one video segment that semantically matches the recommended text for the object is obtained.

10. The method according to claim 9, characterized in that, The step of extracting the video features of each of the multiple video segments includes: For each of the multiple video segments, perform the following operations: Text features are extracted from the descriptive text of a video clip to obtain text features, and image features are extracted from at least one video frame in the video clip to obtain at least one image feature; The text features and at least one image feature are fused together to obtain the video features of the video segment.

11. The method according to claim 9, characterized in that, The step of obtaining at least one video segment that semantically matches the recommended text of the object, based on at least one video segment selected for the at least one text segment, includes: When multiple video segments are selected for the at least one text segment, and the set video generation restrictions are received, at least one video segment that meets the video generation restrictions is selected again from the multiple selected video segments. At least one video segment that is selected again will be considered as at least one video segment that semantically matches the recommended text for the object.

12. The method according to claim 9, characterized in that, The step of generating an object recommendation video based on the selected at least one video segment and the object recommendation text includes: For the at least one text segment, perform the following operations respectively: obtain at least one video segment that matches a text segment, and the corresponding cleaned video segment; convert the text segment into text-to-speech; and according to the duration of the text-to-speech, splice and edit the at least one matched cleaned video segment to obtain the corresponding edited video. Based on at least one video clip and the corresponding text / voice for each of the at least one video clips, the recommended video for the target is obtained.

13. The method according to any one of claims 1 to 5, characterized in that, After obtaining the recommended video for the object, the process further includes: The recommended video for the object is displayed in the video preview interface; In response to the adjustment operation for the recommended video of the object, the adjusted recommended video of the object is obtained.

14. The method according to any one of claims 1 to 5, characterized in that, The video footage of the target audience is segmented to obtain multiple video clips, including: The video material is split into multiple segments to obtain multiple segmented clips; each video frame in each segment has the same theme scene. When the duration of at least one of the multiple segment segments is not greater than a set duration, each segment segment in the at least one segment segment is spliced ​​with the adjacent segment segment. The multiple video segments are obtained based on each spliced ​​segment and each unspliced ​​segment.

15. A video processing apparatus, characterized in that, include: The segmentation unit is used to segment the video footage of the object to be recommended into multiple video clips; each video clip has a corresponding theme scene. The recognition unit is used to perform image recognition on the plurality of video segments respectively, and obtain image description text for each of the plurality of video segments; each image description text is used to describe the image content of the corresponding video segment; The copywriting generation unit is used to query the object description text used to describe the object to be recommended from the set material library, and generate corresponding object recommendation copy based on the object description text and the screen description text of each of the multiple video clips; The video generation unit is configured to select at least one video segment from the plurality of video segments that semantically matches the object recommendation text, and generate an object recommendation video based on the selected at least one video segment and the object recommendation text.

16. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of any of the methods described in claims 1 to 14.

17. A computer-readable storage medium, characterized in that, It includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps of any of the methods described in claims 1 to 14.

18. A computer program product, characterized in that, The method includes a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any one of claims 1 to 14.