Video form generation method and device, equipment, storage medium and program product
By analyzing sample videos and descriptive information, a video form is generated, which solves the problem of information recommenders lacking professional skills, enables the efficient generation of high-quality videos, and improves the recommendation effect.
Patent Information
- Application Number
- CN202210071225.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-01-21
AI Technical Summary
The information recommenders lack the professional skills to produce high-quality videos, resulting in poor timeliness and effectiveness of recommendations.
By acquiring sample videos and descriptive information of the objects to be recommended, and using video fingerprinting and text vector analysis, a video form is generated. Video tags and text tags that match the characteristics of the objects to be recommended are selected to generate a high-quality video form.
It improves the timeliness and effectiveness of video recommendations, and the generated video forms more accurately represent the relevant characteristics of the objects to be recommended, thereby improving the accuracy of recommendations.
Smart Images

Figure CN116521937B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to artificial intelligence and recommendation technology, and in particular to a video form generation method and device, equipment, storage medium and program product. BACKGROUND
[0002] Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0003] Based on the recommendation system, the video is pushed to the user to spread the information of a specific object, which is a typical application of artificial intelligence in the field of recommendation technology. For example, the recommendation system identifies users interested in various objects such as goods and services, and sends video advertisements for recommending objects to the terminal device of the user, thereby helping the user to understand the relevant information of the object.
[0004] For the information recommendation party (a party with information recommendation needs, such as an advertiser), it often lacks professional skills to make videos, so it cannot make high-quality videos for recommending objects in a short time, which affects the timeliness of the recommendation and makes it difficult to achieve the expected recommendation effect of the recommended object. SUMMARY
[0005] The embodiments of the present application provide a video form generation method, device, electronic equipment and computer readable storage medium, and computer program product, which can accurately and efficiently generate a video form for making high-quality videos, thereby improving the timeliness and recommendation effect of the recommendation.
[0006] The technical solution of the embodiments of the present application is as follows:
[0007] The embodiments of the present application provide a video form generation method, comprising:
[0008] Obtain a sample video and description information of a to-be-recommended object;
[0009] Obtain a plurality of similar videos of the sample video based on the video fingerprint of the sample video, and obtain a plurality of video tags corresponding to the plurality of similar videos;
[0010] Obtain a plurality of similar texts of the description information based on the text vector of the description information, and obtain a plurality of text tags corresponding to the plurality of similar texts;
[0011] Select at least one label as a target label from the plurality of video labels and the plurality of text labels based on the heat value corresponding to the plurality of video labels and the plurality of text labels, respectively.
[0012] select at least one video field based on a screening index of each video field corresponding to the target label to generate a video form, wherein the video form is used to generate a video for recommending the to-be-recommended object.
[0013] The embodiment of the application provides a video form generation device.
[0014] a data acquisition module configured to acquire a sample video and description information of a to-be-recommended object;
[0015] a label acquisition module configured to acquire a plurality of similar videos of the sample video based on a video fingerprint of the sample video, and acquire a plurality of video labels corresponding to the plurality of similar videos; acquire a plurality of similar texts of the description information based on a text vector of the description information, and acquire a plurality of text labels corresponding to the plurality of similar texts;
[0016] The label acquisition module is further configured to select at least one label as a target label from the plurality of video labels and the plurality of text labels based on a heat value corresponding to each of the plurality of video labels and the plurality of text labels.
[0017] a form generation module configured to select at least one video field based on a screening index of each video field corresponding to the target label to generate a video form, wherein the video form is used to generate a video for recommending the to-be-recommended object.
[0018] The embodiment of the application provides an electronic device for generating a video form, and the electronic device comprises:
[0019] a memory configured to store executable instructions;
[0020] a processor configured to execute the executable instructions stored in the memory to implement the method of the embodiment of the application.
[0021] The embodiment of the application provides a computer readable storage medium storing executable instructions, and the executable instructions are executed by a processor to implement the video form generation method.
[0022] The embodiment of the application provides a computer program product comprising a computer program or instructions, and the computer program or instructions are executed by a processor to implement the video form generation method.
[0023] The embodiment of the application has the following beneficial effects:
[0024] By analyzing the sample video and the description information of the to-be-recommended object, relevant video tags and text tags are obtained, and a video field is obtained based on the tags, so that the video field is more in line with the related characteristics of the to-be-recommended object, and the generated video form can better represent the related characteristics of the to-be-recommended object, so that the video form can be used to generate a recommended video that more accurately recommends the to-be-recommended object, thereby improving the timeliness and recommendation effect of the recommendation. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 is a schematic diagram of an application scenario of a video form generation method provided by an embodiment of the present application;
[0026] Figure 2 is a structural schematic diagram of a video customization server for generating a video form provided by an embodiment of the present application;
[0027] Figure 3A is a flowchart of a video form generation method provided by an embodiment of the present application;
[0028] Figure 3B is a flowchart of a video form generation method provided by an embodiment of the present application;
[0029] Figure 3C is a flowchart of a video form generation method provided by an embodiment of the present application;
[0030] Figure 3D is a flowchart of a video form generation method provided by an embodiment of the present application;
[0031] Figure 3E is a flowchart of a video form generation method provided by an embodiment of the present application;
[0032] Figure 4 is a schematic diagram of the relationship between a tag and a field provided by an embodiment of the present application;
[0033] Figure 5 is a schematic diagram of the relationship between various databases provided by an embodiment of the present application;
[0034] Figure 6A is a flowchart of a video form generation method provided by an embodiment of the present application;
[0035] Figure 6B is a flowchart of a video form generation method provided by an embodiment of the present application;
[0036] Figure 6C is a schematic diagram of an initial form provided by an embodiment of the present application;
[0037] Figures 6D-6E is a schematic diagram of a video form provided by an embodiment of the present application. DETAILED DESCRIPTION
[0038] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings. The described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0039] In the following description, "some embodiments" are referred to, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0040] In the following description, the terms "first\second\third" are only to distinguish similar objects, and do not represent a specific order of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence as allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0042] It should be noted that in the embodiments of the present application, data related to user information, user feedback data, etc. When the embodiments of the present application are applied to specific products or technologies, the user's permission or consent is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region.
[0043] Before further detailing the embodiments of the present application, the terms and terms related to the embodiments of the present application are explained, and the terms and terms related to the embodiments of the present application are explained as follows.
[0044] 1) Video fingerprint, used to uniquely represent the characteristics of the video, the video fingerprint can be represented by a 1024-dimensional feature vector, that is, a 1024-dimensional video fingerprint.
[0045] 2) Term Frequency Inverse Document Frequency (TF-IDF) text vector: Term frequency is the frequency of a word appearing in a text. The more times a word appears in a text, the higher its term frequency. For example, if the total number of words in a text is C, and word D appears d times, then the term frequency of word D is TF = d / C. Inverse document frequency (IDF) can be represented by the logarithm of the inverse document frequency. Document frequency is the frequency of texts containing a word in a corpus. The more texts containing a word, the higher the document frequency, and the lower the IDF of that word. For example, if a corpus contains L texts, and word D appears in W texts, then the IDF of word D is IDF = lg(L / W). The TF-IDF is equal to the product of term frequency and IDF. The TF-IDF of each word in the text is a component; combining all components yields the TF-IDF text vector.
[0046] 3) Video form, or simply form, is used to describe different aspects of the video used for recommendation. It includes multiple form fields (referred to as fields). Each form field includes a parameter for a video type and the corresponding parameter value. The parameter value of each parameter has a certain range of values.
[0047] For example, in the form field "Video Length, 1 minute 30 seconds", "Video Length" is a parameter indicating the video type, and "1 minute 30 seconds" is the corresponding parameter value. Similarly, in the form field "Animation Scene, 3D Animation", "Animation Scene" is a parameter indicating the video type, and "3D Animation" is the corresponding parameter value.
[0048] 4) Recommendation performance data, which characterizes the recommendation effect achieved by the video. Taking video ads as an example, recommendation performance data is also ad performance data, such as: exposure rate, recall rate, influence on purchase intention, liking level, and second bounce rate. Exposure rate is the ratio of the actual number of people reached by the ad to the total number of people the ad can reach. Recall rate is the percentage of users who have seen the ad and can recall it. Influence on purchase intention is how many users the ad attracts to try the advertised item. Liking level is the percentage of users who liked the ad and their level of liking for it. Second bounce rate is the percentage of users who have made the first bounce and then made a second bounce. A first bounce occurs when a user accesses the ad video on the tested website via a link provided by an external website. A second bounce occurs when the user enters a deeper page of the ad video on the tested website.
[0049] 5) Popularity value, which represents the effectiveness of a tag (e.g., video tag, text tag) in at least one aspect of usage frequency and recommendation effect data. Popularity value is a measure of the popularity of a tag.
[0050] 6) Natural Language Processing (NLP) Intelligent Word Segmentation: Natural language processing intelligent word segmentation technology can use artificial intelligence to process natural language text and obtain the words in the text.
[0051] This application provides a method for generating a video form, a device for generating a video form, an electronic device and a computer-readable storage medium for generating a video form, and a computer program product, which can make the form fields more consistent with the characteristics of the object to be recommended, thereby generating a more accurate video form and improving the accuracy of video generation based on the video form.
[0052] See Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the video form generation method provided in this application embodiment. The servers involved include: a video customization server 201 (running a graphical front-end, i.e., a video customization platform) and a recommendation server 202 (belonging to a recommendation system, such as an advertising system), a network 300, and terminal devices (first terminal device 400A and second terminal device 400B). The video customization server 201 and the recommendation server 202 communicate via the network 300 or through other means. The terminal devices connect to the recommendation server 202 via the network 300, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.
[0053] The first user is the recommender (i.e., the party that needs recommendations, such as an advertiser), and the second user is a user who meets the targeted recommendation criteria for the target object. The target object can be a real object (e.g., food, daily necessities, electronic devices, or vehicles) or a virtual object (e.g., games, game items, online educational courses) or service (e.g., purchasing services, cleaning services, consulting services). The first user accesses the video customization platform (i.e., the graphical front-end of the video customization server 201) through the first terminal device 400A, uploads a sample video and description information of the target object, and the video customization server 201 analyzes the sample video and description information to generate a video form. Based on the video form, it generates a video (e.g., an advertising video) for recommending the target object. The video customization server 201 sends the advertising video to the recommendation server 202 via network 300. The recommendation server 202 has already stored the targeted recommendation criteria submitted by the first user through terminal 400B, and sends the video to the second terminal device 400B of the second user who meets the targeted recommendation criteria, allowing the second user to learn about the object of interest by watching the video.
[0054] This application embodiment can be implemented using database technology. A database, simply put, can be viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, capable of being shared by multiple users, having minimal redundancy, and being independent of application programs.
[0055] A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters or mobile devices; or according to the query language used, such as Structured Query Language (SQL) or XQuery; or according to performance priorities, such as maximum scale or maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, simultaneously supporting multiple query languages.
[0056] In this application embodiment, a single database can be deployed, or different databases can be deployed according to the type of data used, such as a corpus database, a video fingerprint database, a video tag database, a text vector database, a text tag database, and a tag field database (hereinafter, the databases in the above database names are referred to as "databases").
[0057] In some embodiments, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal device can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment of the invention.
[0058] This application embodiment can also be implemented through machine learning. Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.
[0059] This application embodiment can also be implemented using cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on cloud computing business models. It can form a resource pool, which can be used on demand, offering flexibility and convenience. Cloud computing technology will become an important support. The backend services of technical network systems require a large amount of computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the Internet industry, and driven by the demands of search services, social networks, mobile commerce, and open collaboration, in the future, every item may have its own hash code identification mark, which will need to be transmitted to the backend system for logical processing. Data of different levels will be processed separately, and various industry data will all require strong system support, which can only be achieved through cloud computing.
[0060] See Figure 2 , Figure 2This is a schematic diagram of the structure of a video customization server for generating video forms provided in an embodiment of this application, including: at least one processor 410, a memory 450, and at least one network interface 420. Various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 2 The general labeled all buses as Bus System 440.
[0061] Processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0062] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices physically located away from the processor 410.
[0063] The memory 450 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory.
[0064] In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0065] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks.
[0066] The network communication module 452 is used to reach other computing devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, WiFi, and Universal Serial Bus (USB).
[0067] In some embodiments, the video form generation device provided in this application can be implemented in software. Figure 2 A video form generation device 455 stored in memory 450 is shown. This device can be software in the form of programs or plugins, and includes the following software modules: a data acquisition module 4551, a tag acquisition module 4552, and a form generation module 4553. These modules are logically linked and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0068] See Figure 3A , Figure 3A This is a flowchart illustrating the video form generation method provided in this application embodiment. Figure 1 The video customization server in the middle is the main execution body, and will be combined with Figure 3A The steps shown are explained.
[0069] In step 101, the sample video and the description information of the object to be recommended are obtained.
[0070] For example, the sample video and description information can be from the recommender (e.g., the advertiser, corresponding to...). Figure 1 The first user in the process) through the terminal device ( Figure 1 The first terminal device (400A) sends the video to the video customization server. The recommended object can be a real object (e.g., food, daily necessities, electronic devices, or vehicles) or a virtual object (e.g., games, game items, online education courses), or a service (e.g., purchasing services, cleaning services, consulting services). The description information of the recommended object is presented in text form. For example, if the recommended object is a home appliance, the description information is the user manual for that appliance. The sample video and description information should reflect the first user's needs for recommended videos used to recommend the recommended object.
[0071] It should be pointed out that, although Figure 1 The order of execution of steps 102 to 104 is shown in the figure, but as can be understood from the description below, steps 102 and 103 can be executed sequentially or in parallel.
[0072] In step 102, multiple similar videos of the sample video are obtained based on the video fingerprint of the sample video, and multiple video tags corresponding to the multiple similar videos are obtained.
[0073] For example, video fingerprints are used to characterize the features of a video. In this embodiment, video fingerprints are represented in the form of feature vectors. Similar videos are those with a high degree of similarity to sample videos. For example, the similarity between multiple reference videos and sample videos can be determined (this can be represented by the similarity between video fingerprints). Reference videos with similarity within a similarity threshold range (e.g., similarity 0.9~1) are selected as similar videos, or multiple reference videos ranked first in descending order of similarity are selected as similar videos. The video tags corresponding to similar videos are related to the content of the similar videos. For example, if a similar video is a demonstration video of a game, the video tags for the similar video are: **** (where **** refers to the name of the game), game animation, XX (where XX refers to the name of a game character).
[0074] In some embodiments, see Figure 3B , Figure 3B This is a flowchart illustrating the video form generation method provided in this application embodiment; step 102 can be implemented through steps 1021 to 1023, as detailed below.
[0075] In step 1021, the video fingerprint of the sample video is obtained.
[0076] For example, in this embodiment of the application, video fingerprints are used as feature vectors for illustration. Step 1021 can be implemented in the following way: the sample video is segmented based on a preset duration (e.g., per second) to obtain multiple video segments, and a video frame (e.g., a keyframe) is extracted from each video segment; a deep learning convolutional neural network is called to extract features from each video frame to obtain the video frame features corresponding to each video frame; the video frame features corresponding to each video frame are combined to obtain the video fingerprint of the sample video.
[0077] For example, a video clip contains multiple video frames, such as a preset duration of 1 second. A keyframe is extracted from each video clip, and feature extraction is performed based on this keyframe. If the duration of the last video clip in the sample video is shorter than the preset duration, a video frame is still extracted from the last video clip (equivalent to rounding up the number of video clips, thus comprehensively reflecting the video's features and ensuring the accuracy of subsequent video fingerprint calculations). A deep learning convolutional neural network can be used to process all video clips, obtaining the video frame features corresponding to each clip. These video frame features are then combined to obtain the video fingerprint of the sample video. The video fingerprint of the sample video can be a multi-dimensional (e.g., 1024-dimensional) vector, meaning the video fingerprint is a multi-dimensional video fingerprint.
[0078] In step 1022, the similarity between the video fingerprint of each reference video and the video fingerprint of the sample video is determined. Multiple reference videos are selected from the top of the descending similarity ranking results as multiple similar videos of the sample video, or multiple reference videos with similarity greater than the similarity threshold are selected as multiple similar videos of the sample video.
[0079] For example, the correspondence between the video identifier and the video fingerprint of each reference video is stored in the video fingerprint database, and there is a one-to-one relationship between the video identifier and the video fingerprint of the reference video.
[0080] For example, video fingerprints can be represented by feature vectors, and the similarity between video fingerprints can be represented by the Euclidean distance between feature vectors. The shorter the Euclidean distance between feature vectors, the higher the similarity between video fingerprints. The following formula (1) is the Euclidean distance formula:
[0081]
[0082] Where X and Y are the feature vectors corresponding to the video fingerprint, x i y is the eigenvalue of the i-th position in the eigenvector X. i It is the eigenvalue of the i-th element in the eigenvector Y. dist(X, Y) is the Euclidean distance between eigenvector X and eigenvector Y. The smaller the Euclidean distance, the greater the similarity between the video corresponding to eigenvector X and the video corresponding to eigenvector Y.
[0083] For example, descending order sorting means that the higher the similarity, the higher the ranking. Correspondingly, the smaller the Euclidean distance between the feature vector corresponding to the video fingerprint of the reference video and the feature vector corresponding to the video fingerprint of the sample video, the higher the ranking of the reference video. A set number of reference videos (e.g., 10) are selected from the top of the descending order results as similar videos. Alternatively, multiple reference videos with a similarity greater than a preset similarity threshold (e.g., 0.9) are selected as similar videos.
[0084] In step 1023, based on the video identifiers of multiple similar videos, the correspondence between different reference videos and different video tags is queried to obtain multiple video tags corresponding to multiple similar videos.
[0085] Each similar video corresponds to at least one video tag.
[0086] For example, the mapping between video tags and reference video IDs is stored in the video tag database. The mapping types between video tags and reference video IDs include one-to-one, one-to-many, and many-to-one. Each video tag has a corresponding popularity value, which is determined based on at least one of the following: the frequency of use of the video tag in the video customization server and the corresponding video's recommendation performance data. The video recommendation performance data includes at least one of the following: impressions, clicks, and conversions (i.e., the number of users who placed orders, made purchases, or favorited the video after it was played). For example, if the recommended video used to recommend a target audience is an advertisement video, then the recommendation performance data can be reflected through the advertisement conversion rate. The ratio of the number of conversions by advertising users to the number of ad impressions is called the advertisement conversion rate.
[0087] For example, the popularity value can be obtained by taking the frequency of use of video tags in the video customization server and the corresponding video recommendation effect data as different weights, and then weighting and summing the frequency of use and recommendation effect data with their corresponding weights, and using the weighted sum as the popularity value.
[0088] In some embodiments, a weighted summation can be performed based on the popularity value of the video tag corresponding to the reference video and the similarity between the video fingerprint of the reference video and the video fingerprint of the sample video to obtain a weighted summation result. Based on the weighted summation result, all reference videos to be pushed in the recommendation server 202 are sorted in descending order, and multiple reference videos at the top of the descending sort are selected as similar videos, and the video tags corresponding to the similar videos are obtained.
[0089] In some embodiments, the video customization server 201 obtains the video fingerprint corresponding to the sample video, and determines similar videos corresponding to the sample video based on the similarity between the video fingerprint of the sample video and the video fingerprints in the video fingerprint database. The video customization server 201 then searches the video tag database based on the video identifiers of the similar videos to obtain multiple video tags corresponding to each similar video.
[0090] In this embodiment, obtaining similar videos of the sample video based on video fingerprints can improve the accuracy of obtaining similar videos. By obtaining video tags through the correspondence between similar videos and video tags, the obtained video tags are more consistent with the relevant features of the sample video, reducing the amount of computation required to obtain video tags.
[0091] In some embodiments, the video fingerprint of each reference video is stored in a video fingerprint database. Before step 102, data can also be written to the video fingerprint database and the video tag database in the following ways: obtain multiple reference videos, multiple video tags and the popularity value of each video tag, determine the video fingerprint corresponding to each reference video; and store the correspondence between the video identifier of each reference video and the video fingerprint of each reference video in the video fingerprint database.
[0092] For example, the reference video can be an advertising video, product introduction video, etc., crawled from the Internet. The video tags can be the title, topic, etc., corresponding to the video crawled from the Internet, or the video tags already associated with the reference video, or the video tags obtained by clustering analysis of the reference video. The initial popularity value of the video tag can be obtained based on the usage frequency of the video tag and the recommendation effect data of the video corresponding to the video tag.
[0093] For example, the initial popularity value of a video tag can be obtained in the following way: obtain the usage frequency of the video tag in the videos that capture the video tag, obtain the recommendation effect data corresponding to the videos that capture the video tag, and perform a weighted sum based on the corresponding usage frequency and recommendation effect data and the corresponding weights, and use the weighted sum as the popularity value.
[0094] In some embodiments, the video tag corresponding to each reference video and the popularity value of each video tag are stored in a video tag database. Before step 102, each reference video is processed as follows to write data to the video tag database: at least one video tag that matches the video content of the reference video is selected from multiple candidate video tags, and a correspondence is established between the video identifier of the reference video and at least one video tag; the correspondence between the video identifier of each reference video and at least one video tag, and the popularity value of each video tag are stored in the video tag database.
[0095] For example, the video fingerprint database and the video tag database can be separate databases or merged into a single database. The video fingerprint database can store each reference video, the video fingerprint of each reference video, and the correspondence between the video identifier of each reference video and the corresponding video fingerprint of each reference video.
[0096] For example, the video fingerprint corresponding to each reference video is determined as follows: the reference video is segmented based on a preset duration to obtain multiple video segments, and a video frame is extracted from each video segment; features are extracted from each video frame to obtain the video frame features corresponding to each video frame; the video frame features corresponding to each video frame are combined to obtain the video fingerprint corresponding to the reference video.
[0097] For example, the video customization server periodically (e.g., daily) updates the popularity value corresponding to each video tag in the video tag database.
[0098] In this embodiment, by storing the correspondence between video tags and video identifiers, and the correspondence between video fingerprints and video identifiers in a database, the corresponding video tags or video fingerprints can be quickly retrieved from the database based on the video identifiers, improving the computing efficiency of the video customization server and saving computing resources; furthermore, the periodic updating of the popularity value of the video tags ensures the timeliness of the popularity value corresponding to the video tags.
[0099] In step 103, multiple similar texts of the description information are obtained based on the text vector of the description information, and multiple text tags corresponding to the multiple similar texts are obtained.
[0100] Text vectors are used to represent the features of text and can be TF-IDF text vectors. Similar text is text that has a high degree of similarity to the descriptive information. The text vector library stores a large number of text identifiers for reference texts, as well as the correspondence between the text identifiers and text vectors of each reference text. The text tag library stores the correspondence between the text identifier (e.g., text ID) of each reference text and at least one text tag corresponding to each reference text.
[0101] In some embodiments, step 103 can be implemented as follows: determining the similarity between multiple reference texts and the sample text (which can be represented by the cosine similarity between text vectors), selecting reference texts with similarity within a similarity threshold range (e.g., similarity 0.9~1) as similar texts, or selecting multiple reference texts at the top of the similarity ranking in descending order as similar texts. The text tags corresponding to the similar texts are related to the content of the similar texts. For example, if the content of the similar text is a user manual for a garment steamer, the text tags for the similar texts might be: **** (where **** refers to the brand of the garment steamer), home appliances, garment steamer, portable, etc.
[0102] In some embodiments, see Figure 3C , Figure 3C This is a flowchart illustrating the video form generation method provided in this application embodiment; step 103 can be implemented through steps 1031 to 1033, as detailed below.
[0103] In step 1031, the text vector of the description information is obtained.
[0104] To facilitate explanation, the following is an example of a description of a mobile phone being recommended: "A mobile phone... supports wired super-fast charging and wireless super-fast charging; a wireless fast charger must be purchased separately." The following explanation will use this example description as a starting point.
[0105] For example, step 1031 can be implemented as follows: perform word segmentation on the description information to obtain multiple words included in the description information; perform the following processing on each word: determine the word frequency based on the number of times the word appears in the description information and the total number of words in the description information; determine the inverse document rate of the word based on the number of texts containing the word in the corpus and the total number of texts in the corpus; determine the text component corresponding to the word based on the word frequency and the inverse document rate; combine the text components corresponding to each word to obtain the text vector of the description information.
[0106] For example, word segmentation can be performed based on natural language processing technology. Based on the example description information, word segmentation is performed to obtain words such as "a", "mobile phone", "support", "wired", "super", etc. Assuming that the above description information is segmented into 100 words, the total number of words is 100. Among them, the word "fast charging" appears 3 times. Therefore, the word frequency of the word "fast charging" is the number of occurrences divided by the total number of words, 3 / 100 = 0.03.
[0107] For example, a corpus is a database that pre-stores a massive amount of text, such as 10 million texts. Continuing with the explanation based on the word "fast charging," assuming the corpus contains 1000 texts containing the word "fast charging," and the total number of texts containing the word "fast charging" is 1000, then the inverse document frequency of "fast charging" is lg(10000000 / 1000) = 4.
[0108] For example, multiplying the term frequency by the inverse document frequency (IVF) yields the text component corresponding to each word, which is the word's IVF inverse document rate. In the example above, the IVF for the word "fast charging" is 0.12. Combining the IVF inverse document rates of all words in the description information yields the text vector corresponding to the description information, which is the IVF inverse document rate vector. For example, the text vector of the description information is [0.5 0.2 ……0.12 0.3……].
[0109] In step 1032, the similarity between the text vector of each reference text and the text vector of the description information is determined. Multiple reference texts are selected from the head of the descending similarity sorting results as multiple similar texts of the description information, or multiple reference texts with similarity greater than the similarity threshold are selected as multiple similar texts of the description information.
[0110] For example, the text vectors of the reference texts, the text identifiers of each reference text, and the corresponding relationships between the text vectors can be stored in a text vector database. There is a one-to-one relationship between the text identifiers and the text vectors. The similarity between the text vectors can be obtained by the cosine similarity formula. The following formula (2) is the cosine similarity formula:
[0111]
[0112] Where A and B are different text vectors, A i B represents the value of the i-th position in text vector A. i This represents the i-th value in text vector B. cos θ It's cosine similarity. The higher the cosine similarity, the greater the similarity between the reference text and the descriptive information.
[0113] For example, descending order means that the higher the similarity, the higher the ranking. Correspondingly, the higher the cosine similarity, the higher the similarity between the reference video and the description information. A set number of reference texts (e.g., 10) are selected from the top of the descending order results as similar texts. Alternatively, multiple reference texts with a similarity greater than a similarity threshold (e.g., 0.9) are selected as similar texts.
[0114] In step 1033, the correspondence between different reference texts and different text tags is queried based on the text identifiers of multiple texts to obtain multiple text tags corresponding to multiple similar texts.
[0115] Here, each similar text corresponds to at least one text tag. For example, the correspondence between text tags and the text IDs of reference texts is stored in the text tag database. The types of correspondence between different reference texts and different text tags include: one-to-one, one-to-many, and many-to-one. Each text tag has a corresponding popularity value, which is determined based on at least one of the following: the frequency of use of the text tag in the video customization server and the recommendation effect data of the corresponding video. The recommendation effect data of the text includes at least one of the following: impressions, clicks, and conversions.
[0116] For example, the popularity value can be obtained by taking the frequency of use of text tags in the text customization server and the corresponding video recommendation effect data as different weights, and then weighting and summing the frequency of use and recommendation effect data with their corresponding weights, and using the weighted sum as the popularity value.
[0117] In some embodiments, a weighted summation can be performed based on the popularity value of the text tag corresponding to the reference text, the similarity between the text fingerprint of the reference text and the text vector of the description information, to obtain a weighted summation result. Based on the weighted summation result, all reference texts are sorted in descending order, and multiple reference texts at the top of the descending order are selected as similar texts, and the text tags corresponding to the similar texts are obtained.
[0118] In some embodiments, the video customization server 201 obtains the text vector corresponding to the description information. The text vector may be a term frequency inverse document rate vector. Based on the similarity between the text vector corresponding to the description information and the text vectors in the text vector, the video customization server 201 determines the similar text corresponding to the description information. The video customization server 201 searches a text tag library based on the text identifiers of the similar text to obtain multiple text tags corresponding to each similar text.
[0119] In this embodiment, obtaining similar text based on text vectors improves the accuracy of similar text acquisition. Obtaining text tags through the correspondence between similar text and text labels makes the acquired text tags more consistent with the content of the descriptive information, reducing the computational load required to obtain the text tags.
[0120] In some embodiments, the text vector of each reference text is stored in a text vector database, and the text tag corresponding to each reference text and the popularity value of each text tag are stored in a text tag database. Before step 103, data can also be written to the text vector database in the following ways: obtain multiple reference texts, multiple text tags and the popularity value of each text tag, determine the text vector corresponding to each reference text, and store the correspondence between the text identifier of each reference text and the text vector of each reference text in the text vector database.
[0121] For example, reference text can be advertising copy, product description text, etc., crawled from a corpus or the web, or descriptive information previously received by the video customization platform. Text tags can be keywords, titles, etc., of the text, or text tags already associated with the reference text, or text tags obtained through cluster analysis of the reference text. The initial popularity value of a text tag is determined based on its usage frequency. After the text tag is used in the video customization platform, its popularity value can be updated based on the recommendation performance data of the recommended video corresponding to the text tag.
[0122] For example, the initial popularity value of a text tag can be determined as follows: determine the total number of texts in the range of text tags to be crawled, obtain the number of times the text tag appears in the range of text tags to be crawled, obtain the frequency of occurrence of the text tag based on the number of occurrences and the total number of texts, and multiply the frequency of occurrence by the corresponding weight to obtain the popularity value of the text tag.
[0123] In some embodiments, the text tag corresponding to each reference text and the popularity value of each text tag are stored in a text tag database. Before step 103, each reference text may be processed as follows to write data to the text tag database: select at least one text tag that matches the text content of the reference text from a plurality of text tags, and establish a correspondence between the text identifier of the reference text and at least one text tag; store the correspondence between the text identifier of each reference text and at least one text tag, and the popularity value of each text tag, in the text tag database.
[0124] For example, the text vector database and the text label database can be separate databases or combined into a single database. The text vector database can store each reference text, the text vector of each reference text, and the correspondence between the text identifier of each reference text and its corresponding text vector.
[0125] For example, the text vector corresponding to each reference text is determined as follows: the reference text is segmented into words to obtain multiple words included in the reference text; each word is processed as follows: the word frequency is determined based on the number of times the word appears in the reference text and the total number of words in the reference text; the inverse document rate is determined based on the number of texts containing the word in the corpus and the total number of texts in the corpus; the text components corresponding to the word are determined based on the word frequency and the inverse document rate; and the text components corresponding to each word are combined to obtain the text vector of the reference text.
[0126] For example, the corpus used to calculate the text vector corresponding to the reference text is the same corpus used to calculate the text vector corresponding to the description information. When calculating the text vector, the amount of text in the corpus and the content corresponding to the text remain unchanged. That is, both the reference text and the description information calculate the inverse document rate of words based on the same reference basis to ensure the accuracy of the similarity between the reference text and the description information.
[0127] For example, the video customization server periodically (e.g., daily) updates the popularity value corresponding to each text tag in the text tag database.
[0128] In this embodiment, by storing the correspondence between text tags and text identifiers, and the correspondence between text fingerprints and text identifiers in a database, the corresponding text tags or text vectors can be quickly retrieved from the database based on the text identifiers, thereby improving the computational efficiency of the video customization server and saving computational resources. Furthermore, the periodic updating of the popularity value of the text tags ensures the validity of the popularity value corresponding to the text tags.
[0129] In step 104, based on the popularity values corresponding to multiple video tags and multiple text tags respectively, at least one tag is selected as the target tag from the multiple video tags and multiple text tags.
[0130] Here, the popularity value corresponding to the video tag is determined based on at least one of the following: the frequency of use of the video tag and the recommendation effect data of the corresponding video. The popularity value corresponding to the text tag is determined based on at least one of the following: the frequency of use of the text tag and the recommendation effect data of the corresponding video. The recommendation effect data of the video includes at least one of the following: number of impressions, number of clicks, and number of conversions.
[0131] For example, the popularity value of a tag can reflect the frequency of tag usage and the recommendation effect data of the video corresponding to the tag. The higher the usage frequency, exposure, clicks, and conversions, the higher the popularity value of the tag.
[0132] For example, step 104 can be implemented as follows: Obtain the popularity values corresponding to multiple video tags and multiple text tags respectively. Based on the popularity values corresponding to the multiple video tags and multiple text tags respectively, sort the multiple video tags and multiple text tags in descending order, and select at least one tag from the head of the descending sort result as the target tag, or select at least one tag with a popularity value greater than the popularity value threshold as the target tag.
[0133] For example, text tags and video tags are aggregated into a single sequence. Based on their popularity values, the text and video tags in this sequence are sorted in descending order. The sorted results are then selected from the beginning to the end of the descending order, with at least one tag chosen as the target tag. The popularity threshold can be determined based on the popularity values of the tags in the descending sorted results; for example, the popularity threshold could be the average popularity values of all tags in the descending sorted results.
[0134] In this embodiment, tags are evaluated based on popularity values to obtain tags that are used more frequently and can have a positive impact on video recommendation effects, thereby improving the accuracy of obtaining target tags and thus obtaining the corresponding video fields, which improves the accuracy of video form creation.
[0135] In step 105, at least one video field is selected to generate a video form based on the filtering criteria for each video field corresponding to the target tag.
[0136] The video form is used to generate videos for recommending objects. For example, the mapping between tags and video fields is stored in the tag field database. The relationship between tags and video fields can be many-to-many, one-to-one, or one-to-many. Each tag corresponds to at least one video field, and each video field includes a parameter indicating the video type and its corresponding value. Each parameter's value has a certain range of possible values.
[0137] refer to Figure 4 , Figure 4 This is a schematic diagram of the relationship between tags and fields provided in the embodiments of this application; the diagram includes multiple tags (tag 1, tag 2... tag N, the tag type can be text tag or video tag) and multiple fields (field 1, field 2... field N); tag 2 corresponds to field 1, tag 1 corresponds to field 1 and field 2, field 2 corresponds to multiple other tags in addition to tag 1, and there is a one-to-one relationship between field N and tag N, that is, the relationship between tags and fields can be many-to-many, one-to-one, or one-to-many.
[0138] For example, the filtering metric is obtained by weighted summation of multiple recommendation metrics and their corresponding weights. A higher filtering metric means that the parameters for a video type and their corresponding values in the video field better match the descriptive information of the sample video and the object to be recommended.
[0139] In some embodiments, see Figure 3D , Figure 3D This is a flowchart illustrating the video form generation method provided in this application embodiment; step 105 can be implemented through steps 1051 to 1054, as detailed below.
[0140] In step 1051, the correspondence between different tags and different video fields is queried based on the target tag to obtain the video field corresponding to each target tag.
[0141] For example, the correspondence between different tags (text tags or video tags) and different video fields is stored in the tag field database. The relationship between tags and video fields can be one-to-one, one-to-many, many-to-many, etc. By using the target tag as a search term, the correspondence between different tags and different video fields can be queried, resulting in multiple video fields corresponding to each target tag.
[0142] For example, let's illustrate the video field. For instance, if the target tag is "mobile game," the corresponding video fields for "mobile game" would be: Animation Scene, 3D Animation, Key Character, **** (**** refers to the name of the key character), etc. Here, "Animation Scene" is a parameter representing the video type, "3D Animation" is the parameter value corresponding to the "Animation Scene," "Key Character" is a parameter representing the video type, and the name of the key character is the parameter value corresponding to the key character.
[0143] In step 1052, the filtering criteria corresponding to each video field are determined.
[0144] For example, step 1052 can be achieved as follows: Obtain the weight values corresponding to the multiple recommendation metrics for each video field. For each video field, perform the following processing: Sum the multiple recommendation metrics for the video field based on their corresponding weight values to obtain the filtering metrics corresponding to the video field.
[0145] As an example, the types of recommendation metrics include: the number of times the video field is used, the number of impressions of the video corresponding to the video field, the number of clicks of the video corresponding to the video field, the number of conversions (or conversion rate) of the video corresponding to the video field, the recall rate, the degree of user liking, and the second bounce rate.
[0146] For example, multiple recommendation metrics for the video field can be obtained from the recommendation performance data corresponding to the recommended videos already created on the video customization server, and these metrics can be updated in real time or periodically. Each recommendation metric is multiplied by its corresponding weight value, and the sum of these multiplications is used as the filtering metric. For example, the weight for usage frequency is 0.8, the weight for impressions is 0.5, the weight for clicks is 0.8, and the weight for conversions is 1.2. The filtering metric would be: Filtering Metric = Usage Frequency * 0.8 + Impressions * 0.5 + Clicks * 0.8 + Conversions * 1.2.
[0147] In step 1053, based on the filtering criteria corresponding to each video field, multiple video fields are sorted in descending order, and at least one video field is selected as the target field from the head of the descending sort result of the filtering criteria.
[0148] For example, the higher the filtering index of the video field, the better the effect of the video field. That is, the video field can bring more positive recommendation effects in video form creation and video generation based on video forms, and the videos generated based on video forms can better meet the needs of sample videos and descriptive information.
[0149] In some embodiments, steps 1051 to 1053 can be implemented based on a field recommendation model. Multiple video fields and corresponding recommendation metrics for each video field are pre-acquired to train the field recommendation model. Field recommendations are then performed based on the trained model to obtain at least one video field as the target field. The data for the multiple recommendation metrics corresponding to the video fields can be periodically updated, and the field recommendation model can be updated based on this data to improve its performance in recommending fields more accurately.
[0150] In step 1054, a video form is generated based on at least one target field.
[0151] For example, the parameters of a video type corresponding to each target field and the corresponding data of the parameter values are summarized and stored in a form format to obtain a video form.
[0152] In this embodiment, the target field is selected from all video fields corresponding to the target tag based on the screening index, and the video form is generated based on the target field, which improves the accuracy of the video form generation. That is, the recommended video obtained by making the video based on the video form can effectively recommend the recommended object.
[0153] In some embodiments, see Figure 3E , Figure 3E This is a flowchart illustrating the video form generation method provided in this application embodiment; step 105 can be implemented through steps 1051 to 1052 and steps 1055 to 1057, as detailed below.
[0154] For example, the filtering indicators corresponding to each video field were obtained through steps 1051 to 1052.
[0155] In step 1055, multiple video fields are sorted in descending order based on the filtering criteria corresponding to each video field, and at least a portion of the video fields in the header of the descending sort result are displayed.
[0156] For example, the video field at the top of the descending sort results has a higher filtering index, and at least part of the video field at the top can be used as a field to be recommended. The video customization server sends the field to be recommended to the first user's terminal device, and the first user's terminal device displays the field to be recommended.
[0157] In step 1056, the target field is obtained by at least one of the following methods: in response to a selection operation for any video field among at least some video fields, the selected video field is used as the target field; in response to a custom field input operation, the input custom video field is used as the target field.
[0158] For example, after the recommended fields are displayed to the first user, the first user can select from the recommended fields according to their needs. The recommended field selected by the first user becomes the target field. Alternatively, custom video fields entered by the first user can also be used as target fields. For example, if the first user enters custom video fields such as "video length 1 minute 30 seconds" or "core selling points," using these custom video fields as target fields allows the first user to express their needs for recommended videos.
[0159] In step 1057, a video form is generated based on the target field.
[0160] For example, the parameters of a video type corresponding to each target field and the corresponding data of the parameter values are summarized and stored in a form format to obtain a video form.
[0161] In this embodiment, by recommending fields to be recommended to the first user, the first user can use more standardized video fields to describe their needs for recommended videos during the editing process of the video form; by allowing the first user to customize video fields, the video form can be made closer to the first user's needs, so that the recommended videos produced based on the video form can meet the first user's needs for recommending objects, resulting in better recommendation effects.
[0162] In some embodiments, the video field is stored in a tag field database; prior to step 105, the video field is stored in the tag field database by: obtaining multiple video fields and at least one tag corresponding to each video field, wherein the tag type includes text tags and video tags; and storing the correspondence between each video field and at least one tag corresponding to each video field in the tag field database.
[0163] In some embodiments, in response to a custom field input operation, after the input custom video field is used as the target field, the data in the tag field library is updated in the following ways: a correspondence is established between each custom video field and each target tag, and the correspondence between each custom video field and each target tag is stored in the tag field database; the popularity value of each target tag is updated based on the recommendation effect data of the video corresponding to each custom video field.
[0164] For example, if the first user-defined video field is a new video field not found in the tag video field library, then a correspondence between the new video field and the target tag is established, the popularity value of the target tag is updated accordingly, and the new video field is stored in the tag field database, which helps enrich the data content of the tag field database. If the first user-defined video field is a video field that exists in the tag field library, but a correspondence has not been established with the target tag, or the filtering index of the video field is too low to be selected as a recommended field, then the relevant data for the video field is updated to improve the filtering index corresponding to the video field.
[0165] In some embodiments, after step 105, a video is generated by: obtaining the video parameters and corresponding video parameter values included in each video field of the video form; obtaining video materials that match the video parameters and corresponding video parameter values; wherein the video materials include any one of images, text, audio, and video; and generating a video of the recommended object based on the obtained video materials.
[0166] For example, to facilitate explanation, let's take a mobile game as the target for recommendation, with the video fields being "Animation Scene," "3D Animation," "Key Character," and "****" (where **** refers to the name of the key character). For instance, "Animation Scene" and "Key Character" are the video parameter types, "3D Animation" is the parameter value corresponding to the animation scene, and "Key Character Name" is the parameter value corresponding to the key character. Based on these video fields, we can obtain materials such as 3D animation scene models, images of key characters, and related audio (e.g., character theme songs, character voice acting). Artificial intelligence can then be used to generate videos based on these materials to produce the recommended videos for the target game.
[0167] Continue to combine Figure 1 The following explanation is provided. The video customization server 201 creates a video based on the generated video form and sends the completed recommended video to the recommendation server 202. Continuing with the example of an advertising video, the recommendation server 202 sends the advertising video to the second terminal device 400B. The recommendation server 202 analyzes user data (e.g., user age, gender, and interests) from the big data dataset to identify the second user who meets the advertiser's specified recommendation criteria, and then sends the advertising video to the second user's second terminal device 400B. For example, if the advertising video recommends a mobile game, and the advertiser specifies the recommendation criteria as "targeting young people with spending power," the recommendation server 202 can determine the age range and spending power range of the user group that meets these criteria. The recommendation server 202 analyzes the user data, locates the second user who meets the above recommendation criteria, and pushes the advertising video to the second user's second terminal device.
[0168] In some embodiments, the recommendation server 202 may also place the advertising video on a designated placement platform (e.g., online e-commerce platform, video software, etc.) or a designated placement location (e.g., advertising screens in the real environment, electronic billboards, etc.).
[0169] In some embodiments, the video production process can be performed by a third user (the user responsible for creating the recommended video). The video customization server 201 sends a video form via the network to the third user's third-party terminal device or a third-party video production platform. The third user creates the video based on the video form and sends the completed recommended video to the video customization server 201 via the third-party terminal device. The video customization server 201 then sends the recommended video to the recommendation server 202, which recommends the video to the second terminal device 400B. The video customization server 201 can also send the recommended video to the first terminal device 400A so that the first user can accept and review the recommended video.
[0170] For example, the video form can be an advertising video order. The video customization server 201 of the advertising video customization platform sends the video form to the third-party terminal device of the party responsible for producing the advertisement. The party produces the video and sends the completed advertising video to the first-party terminal device 400A of the advertiser (i.e., the first user) through the video customization server 201, so that the advertiser can review the advertising video. After the advertising video passes the advertiser's review, it can be placed in the real environment (e.g., advertising screens, electronic billboards, etc.) or online platform (e.g., online e-commerce platforms, video platforms, etc.) as required by the advertiser through the recommendation server 202.
[0171] In some embodiments, the second terminal device can collect user feedback data from the second user regarding the advertising video (e.g., user purchases or favorites of products corresponding to the advertising video, user clicks on the advertising video to watch it, user dislikes the advertising video and blocks its push, etc.), and send the user feedback data to the recommendation server 202. The recommendation server 202 calculates the recommendation effect data of the advertising video based on the user feedback data (number of clicks, number of plays, number of conversions, second-bounce rate, memory rate, degree of influence on purchase intention, degree of liking, etc.). The recommendation server 202 synchronizes the recommendation effect data to the video customization server 201, so that the video customization server 201 can update the popularity value of the tags corresponding to the video field corresponding to the advertising video and update the various recommendation indicators in the filtering indicators of the video field corresponding to the advertising video based on the recommendation effect data, to ensure the timeliness of the tags, video fields, and other data stored in each database and improve the accuracy of the generated video forms.
[0172] In this embodiment, tags are obtained based on similar videos and similar text in the descriptive information of the sample videos. This reduces the computational load required for tag acquisition and saves computational resources. Video fields are obtained based on the correspondence between tags and video fields, and video fields used to generate video forms are selected based on filtering indicators. By mining the core content in the sample videos and descriptive information in a data-driven manner, the accuracy of video form generation is improved. This facilitates the creation of recommended videos that better meet the needs of the sample videos and descriptive information, thereby improving the timeliness and effectiveness of recommendations.
[0173] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0174] The video form generation method provided in this application can be applied in the following practical scenarios: Advertisers are users who need to create advertising videos. For the products they want to recommend, the advertiser has relevant data about the product (e.g., sample videos, description information). The advertiser can submit the relevant data and video fields (including the video parameter types of the advertising video and the parameter values corresponding to each parameter type) to a professional (responsible for video production) for video production. However, when specifying video fields, advertisers may use non-standard terminology, affecting the effectiveness of the advertising video produced by the professional. The video form generation method provided in this application can obtain video fields through data mining based on the sample videos and product description information provided by the advertiser, and generate corresponding video forms based on these video fields. Therefore, video production based on video forms can produce advertising videos that more accurately recommend products, improving the recommendation effect of advertising videos for products.
[0175] For example, a video form is a form used to create recommendation videos for items to be recommended (e.g., products), and can be a video transaction order. The video field in the video form includes various parameter types corresponding to the advertising video, and parameter values for each parameter type. Based on the various video fields in the video form, video production can be performed. In this embodiment, the object to be recommended is a product, and the recommendation video used to recommend the object is an advertising video.
[0176] refer to Figure 6A , Figure 6A This is a flowchart illustrating the video form generation method provided in this application embodiment. Figure 6A The execution entity for each step is the video customization server 201 of the video customization platform.
[0177] In step 600A, the sample video and description information are received.
[0178] For example, the description information specifically refers to the product description information. The advertiser uploads the sample video and description information to the video customization server 201 via the first terminal device 400A. The video customization server 201 is the server of the video production platform. For example, the product description information is represented in text form. Products can be physical products (e.g., food, daily necessities, vehicles, electronic devices, etc.) or virtual products (e.g., games, game props, online courses, etc.) or services (e.g., consulting services, purchasing services, intermediary services, cleaning services, etc.). The sample video is a video related to the content of the product description information.
[0179] In this embodiment, a game is used as an example for illustration. The product description information is as follows: "**** is a 5V5 team-based fair competitive mobile game, a national MOBA masterpiece! 5V5 fair battles, restoring the classic MOBA experience... Five-army battles, border breakouts, and more bring diverse combat fun! 10-second real-time cross-region matchmaking, team up with friends to rank up! Multiple heroes to choose from, first blood, pentakill, godlike, crush your opponents, and dominate the field!" **** in the above text refers to the game name.
[0180] After step 600A, steps 601A to 604A and steps 605A to 608A are also included. Steps 601A and 605A are not executed in any order and can be performed simultaneously.
[0181] In step 601A, multiple video frames are extracted from the sample video.
[0182] For example, the sample video is segmented (e.g., by seconds) into multiple video clips, and one frame is extracted from each video clip as the corresponding video frame. For instance, if the sample video is 26.5 seconds long, and it is segmented into 27 video clips in 1-second increments, assuming the length of the 26th video clip is 1 second and the length of the 27th video clip is 0.5 seconds, then one video frame (e.g., a keyframe) is extracted from each of these 27 video clips.
[0183] In step 602A, a neural network model is invoked based on the video frame to obtain the video fingerprint of the sample video.
[0184] For example, obtaining the video fingerprint of a sample video can be achieved as follows: Feature extraction is performed on each video frame using a deep learning convolutional neural network model to obtain the feature vector for each video frame. The feature vectors of all video frames are then combined to obtain the video fingerprint of the sample video. The feature vector can be a 1024-dimensional feature vector, meaning the video fingerprint is a 1024-dimensional video fingerprint.
[0185] In step 603A, the similarity between the sample video and the reference video is calculated based on the Euclidean distance formula to obtain similar videos.
[0186] For example, a video fingerprint database pre-stores the video fingerprints of a large number of reference videos, where each video ID is associated with a specific video fingerprint. Video fingerprints can be represented as feature vectors, and the similarity between sample videos and reference videos can be characterized by the Euclidean distance between these feature vectors.
[0187] For example, similarity is represented by Euclidean distance between vectors; the smaller the value, the higher the similarity. The following formula (3) is the Euclidean distance formula:
[0188]
[0189] Where X and Y are the feature vectors corresponding to the video fingerprint, x i y is the eigenvalue of the i-th position in the eigenvector X. i It is the eigenvalue of the i-th element in the feature vector Y. All videos in the video fingerprint database are sorted from highest to lowest similarity, with higher similarity videos ranking higher.
[0190] Example, Figure 5 This is a diagram showing the relationships between the various databases provided in this application embodiment. The video fingerprint database stores video fingerprints and their corresponding video identifiers. The video tag database stores a large number of video identifiers and at least one tag (video tag) corresponding to each video identifier. The video tag database and the video fingerprint database store the same video identifiers. The product text TF-IDF database stores a large number of text identifiers for product text (i.e., reference text) and their corresponding TF-IDF vectors. The product text tag database stores a large number of text identifiers for product text and at least one tag (text tag) corresponding to each text identifier. The product text tag database and the product text TF-IDF database store the same text identifiers. The tag mapping database (i.e., the tag field database) stores tags and their corresponding fields. Fields can be found in the tag mapping database using tags.
[0191] In step 604A, video tags are obtained from the video tag library based on the video IDs of similar videos.
[0192] For example, by searching the video tag library based on the video IDs of similar videos, the video tags corresponding to each similar video can be obtained.
[0193] In some embodiments, the top a videos in the similarity ranking can be considered as similar videos. Based on the popularity value of the tags, the video tags corresponding to all similar videos are sorted, and the top b tags in the ranking are obtained as the tag data (including the tags and their corresponding popularity values) for the sample video, resulting in multiple (b) video tags. The video customization platform updates the popularity value of each video tag in each database every preset time interval (e.g., 24 hours). Here, a and b are both positive integers.
[0194] For example, the selection of video tags corresponding to sample videos can also be done by obtaining video tags whose popularity values are within the popularity value range, or by obtaining the tag data of video tags whose popularity values are within the popularity value range and whose popularity values are topj, as the tag data of sample videos.
[0195] To facilitate explanation, the following example illustrates the concept: The product is a mobile game, and the sample video is a 26-second game video. The sample video is divided into frames, second by second, resulting in 26 video frame files. A trained neural network model extracts video fingerprint features from each video frame file, obtaining the video frame fingerprint information for each file. These video frame fingerprints are then combined to form the sample video's video fingerprint information, which can be represented as a feature vector. Similar videos to the sample video can be other game videos. Based on these game videos, video tags can be game character names, game names, or game competition names. The popularity value of the video tags for similar videos is obtained, and corresponding video tags are selected based on the popularity value, resulting in multiple video tags.
[0196] In step 605A, intelligent word segmentation is performed on the description information to obtain multiple words in the description information.
[0197] For example, text fields in product descriptions that can form Chinese words, idioms, or trending terms are segmented to obtain each word in the product description. Based on the product description example above, word segmentation can yield words such as "5V5," "team," "fair," "competitive," and "mobile game."
[0198] In step 606A, the TF-IDF of each word in the description information is calculated to obtain the TF-IDF vector of the description information.
[0199] For example, the more times a word appears in the text of product description information, the higher its word frequency; if a word appears more often in multiple paragraphs of text, the lower its inverse document rate (IDF). The specific formula for the TF-IDF of a word is TF-IDF = TF * IDF (word frequency multiplied by inverse document rate).
[0200] For example, we count the frequency of each word in the product description information, the total number of words in the product description information, the number of texts containing the word in the corpus, and the total number of texts in the corpus. For each word, we perform the following processing: divide the frequency of occurrence by the total number of words to obtain the word frequency; based on the number of texts containing the word and the total number of texts, we obtain the frequency of occurrence of the word in the corpus, and take the logarithm of the reciprocal of the frequency of occurrence as the inverse document rate (IDF) of the word. We multiply the IDF and the word frequency to obtain the TF-IDF of the word. We combine the TF-IDF of each word to obtain the TF-IDF vector of the product description information. Each element of the vector corresponds to the TF-IDF of a word in the product description information text.
[0201] To facilitate understanding, the following explanation is based on the product description information exemplified above. For example, let's calculate the TF-IDF value of each word in the product description information. Taking "5V5" from the example product description information as an example, let's assume the total number of words in the above product description information is 100. "5V5" appears twice in the text, and with a total word count of 100, the term frequency of "5V5" is 2 / 100 = 0.02. Assuming there are ten million texts in the corpus, and "5V5" appears in 1000 of them, then the inverse document rate of "5V5" is lg(10000000 / 1000) = 4. Multiplying the inverse document rate by the term frequency, we get the TF-IDF value of "5V5" as 0.02 * 4 = 0.08. We then calculate the TF-IDF value of each word in the product description information sequentially, and combine the TF-IDF values of each word according to the order of words in the text to generate the TF-IDF vector of the product description information.
[0202] In step 607A, the similarity between the descriptive information and the product text is calculated based on the cosine similarity formula to obtain similar text.
[0203] For example, the product text TF-IDF library (also known as the text vector library) stores a large number of text identifiers for product text (i.e., reference text) and a TF-IDF vector corresponding to each text identifier.
[0204] For example, the similarity between product text and product description information can be obtained by calculating the cosine similarity between the TF-IDF vectors of the product text and the TF-IDF vectors of the product description information.
[0205] For example, the following formula (4) is the cosine similarity formula:
[0206]
[0207] Where A and B are different TF-IDF vectors, A iB represents the value of the i-th element in the TF-IDF vector A. i Let θ represent the i-th value in the TF-IDF vector B, and cosθ be the cosine similarity. The higher the cosine similarity, the more similar the product text and product description information are.
[0208] In step 608A, text tags are obtained from the product text tag library based on the text ID of similar text.
[0209] For example, according to the above formula (4), the cosine similarity between the TF-IDF vector of each text in the product text tag library and the TF-IDF vector of the sample text is calculated. The similarity is sorted, and the top e product texts in the sort are taken as similar texts. Based on the text ID of the similar texts, the product text tag library is searched to obtain the text tags corresponding to the similar texts. Based on the popularity value of the text tags, the text tags are sorted, and the top f tags with the highest popularity value are obtained as the tag data of the product description information (including the popularity value of the tags and the tags). Where e and f are both positive integers.
[0210] For example, it can also retrieve text tags that are within the popularity range, or text tags that are within the popularity range and have a popularity value of topf.
[0211] For example, the product description information describes the game product, and the similar text contains game-related content. The text tags corresponding to the similar text can be mobile games, competitive games, and role-playing games.
[0212] In step 609A, target tags are selected based on popularity values, and video fields are obtained based on the correspondence between tags and fields.
[0213] For example, the relationship between tags and fields is many-to-many, one-to-many, or one-to-one. Based on the tag field mapping table stored in the tag field library, all fields corresponding to all tags can be obtained, and a field list can be formed based on these fields. Continuing with the above example, for instance: game names are mapped to animation scenes; competitions are mapped to product features; and character names are mapped to key characters.
[0214] For example, recommending and sorting a list of fields can be achieved by rating each field in the list (the rating is also a filtering metric), and then sorting all fields in the list in descending order based on the rating.
[0215] For example, based on advertising performance data and the number of times the video field corresponding to the tag was used to generate video forms, each parameter corresponding to the score is determined, and the weight value corresponding to each parameter (e.g., usage count, ad impressions, clicks, conversions) is analyzed. Based on advertising performance data and order usage counts, the parameters such as usage count, ad impressions, clicks, and conversions corresponding to the candidate fields are determined, and weighted calculations are performed to obtain the score for the candidate fields. The scoring formula can be: Score = Usage Count * 0.8 + Ad Impressions * 0.5 + Ad Clicks * 0.8 + Ad Conversions * 1.2. The weight values in the scoring formula can be adjusted according to advertising performance data and other data.
[0216] For example, after obtaining the score of each candidate field in the field list, the fields are sorted based on the scores, and the top-scoring candidate fields are selected as recommended video fields.
[0217] For example, the video field includes the video's parameter type and the corresponding parameter value. Based on the sample video and product description information given above, the final recommended fields could be animation scene, product features, and key characters. These video fields correspond to the video's parameter type.
[0218] In some embodiments, the parameter type and corresponding parameter value of the video can be generated directly without the intervention of the advertiser. The video field can be obtained automatically, and a video form can be generated based on the video field, and the video can be generated based on the video form.
[0219] In some embodiments, advertisers can also customize video fields. In step 610A, the custom video field is received. Advertisers can customize video fields of parameter types (i.e., customize video fields), such as key selling points. Based on the parameter types exemplified above, advertisers can also customize parameter values.
[0220] In step 611A, a video form is generated based on the video field.
[0221] For example, to facilitate explanation, the following is combined with the appendix. Figures 6C-6E To explain, Figure 6C This is a schematic diagram of the initial form provided in an embodiment of this application; Figures 6D-6E This is a schematic diagram of the video form provided in an embodiment of this application.
[0222] Example, Figure 6CIn the original form 601C, the initial form is generated before the video fields are obtained, and "Video Customization Form" is the form's title. After the advertiser uploads the sample video and product description information to the video customization server 201, the initial form 601C can be sent to the advertiser's first terminal device 400A for display. The video customization server 201 obtains multiple video fields based on the sample video and product description information and sends these fields to the first terminal device 400A. The human-computer interaction interface on the first terminal device displays the initial form 601C with multiple video fields filled in. (See reference...) Figure 6D The initial form 601C was filled with video field 604C ("Animation scene...", "Product features...", "Key characters...", "Video duration 1 minute", etc., where the ellipsis part refers to the specific parameter value corresponding to the parameter type), and was converted into video form 602C.
[0223] In some embodiments, advertisers can edit the video field in the video form 602C, or add a custom video field to the video form 602C. (See reference) Figure 6E , Figure 6E This demonstrates video form 602C after the advertiser added the custom video field 605C. The custom video field 605C includes "Video Duration 30 seconds," "Key Selling Points...", and "Platform Targeting...". The advertiser changed the video field "Video Duration 1 minute" to the custom video field "Video Duration 30 seconds" and added the custom video fields "Key Selling Points..." and "Platform Targeting...". The "Platform Targeting..." field indicates "The ad video will be placed on the specified platform."
[0224] In step 612A, an advertising video is created based on the video form.
[0225] For example, the video customization server 201 retrieves corresponding video materials (materials can be video clips, 3D models, music, images, text, etc.) based on the parameter types included in the video field of the video form and the parameter values corresponding to each parameter type. For example, based on the video field "Animation Scene: 3D Virtual Scene", it retrieves 3D models and animated images as video materials; based on the video field "Key Character: ***", it retrieves the character portrait, voice-over, and character theme song as video materials. Based on the video materials, parameter types, and parameter values corresponding to the parameter types, it performs video editing through artificial intelligence to generate advertising videos.
[0226] In some embodiments, the video customization server 201 may send the video form to the terminal device of the order taker (the user responsible for creating the video, i.e., the third user mentioned above), who then creates the video. After completing the video creation, the order taker uploads the advertising video to the video customization server 201. The video customization server 201 may also send the completed advertising video to the first terminal device 400A, allowing the advertiser to review the advertising video and provide feedback to improve its content.
[0227] In some embodiments, based on the completed advertising video, advertisements can be pushed to a second user who watches the advertising video, as shown in the reference. Figure 6B , Figure 6B This is a flowchart illustrating the video form generation method provided in this application embodiment. The following is a further explanation... Figure 6B The steps in the process will be explained.
[0228] In step 601B, the first terminal device 400A acquires the sample video and description information, and sends the sample video and description information to the video customization server 201.
[0229] For example, the steps for obtaining sample videos and descriptive information can be found in step 600A above.
[0230] In step 602B, the video customization service 201 generates a video form based on the sample video and description information.
[0231] For example, step 602B can be achieved by steps 600A to 611A above.
[0232] In step 603B, the video customization service 201 produces an advertising video based on the video form.
[0233] For example, the production process of an advertising video can be referred to step 612A above.
[0234] In step 604B, the recommendation server 202 pushes the advertising video to the second terminal device 400B.
[0235] For example, the second terminal device 400B corresponds to the second user watching the advertising video. The second user watching the advertising video can be a potential consumer of the products recommended in the video. The recommendation server 202 analyzes user data to identify second users who may be interested in the advertising video and pushes the video to their terminal device. Alternatively, when customizing the advertising video, the advertiser may specify recommendation criteria (e.g., placing the ad on a specific video platform or targeting a specific user group). The advertiser can add these criteria to the video form as custom video fields, for example... Figure 6E The custom video field "Platform..." (which is a recommendation condition) determines the second user who meets the recommendation condition, and pushes the advertisement video to the terminal device of the second user who meets the recommendation condition.
[0236] In some embodiments, continue to refer to Figure 6A In step 612A, the recommendation effect data of the advertising video is obtained, and the tag popularity value is updated.
[0237] For example, recommendation server 202 obtains feedback data from the second terminal device 400B regarding the advertising video from the second user (e.g., purchase records of the product corresponding to the advertising video, number of clicks, number of views, blocking of the advertising video, etc.). Based on the feedback data, it calculates the advertising performance data of the advertising video (e.g., click-through rate, impression rate, second-bounce rate, etc.). Recommendation server 202 synchronizes the advertising performance data to video customization server 201. Based on the advertising performance data corresponding to the completed advertising video, video customization server 201 can calculate the performance of the video field in the video form corresponding to the advertising video and the popularity value of the tags corresponding to the video field of the advertising video. This updates the popularity values of tags (including video tags and text tags) stored in various databases.
[0238] In some embodiments, if the advertiser customizes a field, the video customization server 201 of the video customization platform will record the mapping relationship between the customized video field and the tag information, store it in the tag field library accordingly, and update the popularity value of the corresponding tag of the customized video field according to the advertising performance data (the video customization server 201 of the video customization platform updates the popularity value of the tags stored in each database in real time, or updates it once every preset time interval), for the next form generation.
[0239] This application's embodiments can analyze sample videos and descriptive information to obtain corresponding tags, acquire corresponding video fields based on the relationship between tags and fields, and support custom fields. It performs data feedback and analysis based on advertising performance data, completing the intelligent generation of fields in the video form in a closed loop. It intelligently scores video fields to recommend higher-scoring video fields to advertisers, enabling them to select more effective video fields (video fields that more clearly express the advertiser's needs for various parameters of the advertising video), thus improving advertising effectiveness. By combining sample videos and descriptive information input by advertisers with video fingerprint libraries, text vector libraries, various tag libraries, tag field libraries, and advertising performance data, it effectively expands the scope of descriptive information for creating advertising videos. High-quality information fusion enables advertising videos created based on video forms to more accurately recommend products.
[0240] The following description continues to illustrate the exemplary structure of the video form generation device 455 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules in the video form generation device 455 stored in the memory 440 may include: a data acquisition module 4551, used to acquire sample videos and descriptive information of the objects to be recommended; a tag acquisition module 4552, used to acquire multiple similar videos of the sample video based on the video fingerprint of the sample video, and acquire multiple video tags corresponding to the multiple similar videos; and to acquire multiple similar texts of the descriptive information based on the text vector of the descriptive information, and acquire multiple text tags corresponding to the multiple similar texts; the tag acquisition module 4552 is used to select at least one tag as a target tag from the multiple first tags and multiple second tags based on the popularity values corresponding to the multiple first tags and multiple second tags respectively; and a form generation module 4553 is used to select at least one video field to generate a video form based on the filtering index of each video field corresponding to the target tag, wherein the video form is used to generate videos for recommending the objects to be recommended.
[0241] In some embodiments, the tag acquisition module 4552 is used to acquire the video fingerprint of the sample video; determine the similarity between the video fingerprint of each reference video and the video fingerprint of the sample video; select multiple reference videos as multiple similar videos of the sample video from the head of the descending similarity sorting results, or select multiple reference videos with similarity greater than a similarity threshold as multiple similar videos of the sample video; query the correspondence between different reference videos and different video tags based on the video identifiers of multiple similar videos to obtain multiple video tags corresponding to multiple similar videos, wherein each similar video corresponds to at least one video tag.
[0242] In some embodiments, the tag acquisition module 4552 is further configured to segment the sample video based on a preset duration to obtain multiple video segments, extract a video frame from each video segment, perform feature extraction on each video frame to obtain video frame features corresponding to each video frame, and combine the video frame features corresponding to each video frame to obtain the video fingerprint of the sample video.
[0243] In some embodiments, the tag acquisition module 4552 is further configured to acquire the text vector of the description information; determine the similarity between the text vector of each reference text and the text vector of the description information; select multiple reference texts as multiple similar texts of the description information from the head of the descending similarity sorting results, or select multiple reference texts with similarity greater than the similarity threshold as multiple similar texts of the description information; query the correspondence between different reference texts and different text tags based on the text identifiers of multiple texts to obtain multiple text tags corresponding to multiple similar texts, wherein each similar text corresponds to at least one text tag.
[0244] In some embodiments, the tag acquisition module 4552 is further configured to perform word segmentation on the description information to obtain multiple words included in the description information; and to perform the following processing on each word: determine the word frequency corresponding to the word based on the number of times the word appears in the description information and the total number of words in the description information; determine the inverse document rate corresponding to the word based on the number of texts containing the word in the corpus and the total number of texts in the corpus; determine the text component corresponding to the word based on the word frequency and the inverse document rate; and combine the text components corresponding to each word to obtain the text vector of the description information.
[0245] In some embodiments, the tag acquisition module 4552 is further configured to acquire the popularity values corresponding to multiple video tags and multiple text tags respectively, wherein the popularity value corresponding to a video tag is determined based on at least one of the usage frequency of the video tag and the recommendation effect data of the corresponding video, and the popularity value corresponding to a text tag is determined based on at least one of the usage frequency of the text tag and the recommendation effect data of the corresponding video, wherein the recommendation effect data of the video includes at least one of the following: number of exposures, number of clicks, and number of conversions; based on the popularity values corresponding to multiple video tags and multiple text tags respectively, the multiple video tags and multiple text tags are sorted in descending order, and at least one tag is selected as the target tag from the head of the descending sort result, or at least one tag with a popularity value greater than the popularity value threshold is selected as the target tag.
[0246] In some embodiments, the form generation module 4553 is further configured to query the correspondence between different tags and different video fields based on the target tags to obtain the video field corresponding to each target tag; determine the filtering index corresponding to each video field; sort multiple video fields in descending order based on the filtering index corresponding to each video field, and select at least one video field as the target field from the head of the descending sort result of the filtering index; and generate a video form based on at least one target field.
[0247] In some embodiments, the form generation module 4553 is further configured to obtain the weight values corresponding to multiple recommendation metrics for each video field, wherein the types of recommendation metrics include: the number of times the video field is used, the number of exposures of the video corresponding to the video field, the number of clicks of the video corresponding to the video field, and the number of conversions of the video corresponding to the video field; and to perform the following processing on each video field: to perform a weighted summation of the multiple recommendation metrics of the video field based on the corresponding weight values to obtain the filtering metrics corresponding to the video field.
[0248] In some embodiments, the form generation module 4553 is further configured to query the correspondence between different tags and different video fields based on the target tags to obtain the video field corresponding to each target tag; determine the filtering index corresponding to each video field, sort the multiple video fields in descending order based on the filtering index corresponding to each video field, and display at least a portion of the video fields at the top of the descending sort result; obtain the target field through at least one of the following methods: in response to a selection operation for any video field among the at least a portion of the video fields, use the selected video field as the target field; in response to a custom field input operation, use the input custom video field as the target field; and generate a video form based on the target field.
[0249] In some embodiments, the tag acquisition module 4552 is further configured to establish a correspondence between each custom video field and each target tag, and store the correspondence between each custom video field and each target tag in the video field database; and update the popularity value of each target tag based on the recommendation effect data of the video corresponding to each custom video field.
[0250] In some embodiments, the video fingerprint of each reference video is stored in a video fingerprint database, and the video tag corresponding to each reference video and the popularity value of each video tag are stored in a video tag database. The tag acquisition module 4552 is further configured to acquire multiple reference videos, multiple video tags, and the popularity value of each video tag, and determine the video fingerprint corresponding to each reference video; store the correspondence between the video identifier of each reference video and the video fingerprint of each reference video in the video fingerprint database; and perform the following processing on each reference video: select at least one video tag that matches the video content of the reference video from multiple video tags, establish the correspondence between the video identifier of the reference video and at least one video tag; and store the correspondence between the video identifier of each reference video and at least one video tag, and the popularity value of each video tag, in the video tag database.
[0251] In some embodiments, the text vector of each reference text is stored in a text vector database, and the text tag corresponding to each reference text and the popularity value of each text tag are stored in a text tag database. The tag acquisition module 4552 is further configured to acquire multiple reference texts, multiple text tags, and the popularity value of each text tag, and determine the text vector corresponding to each reference text; store the correspondence between the text identifier of each reference text and the text vector of each reference text in the text vector database; and perform the following processing on each reference text: select at least one text tag from multiple text tags that matches the text content of the reference text, and establish a correspondence between the text identifier of the reference text and at least one text tag; and store the correspondence between the text identifier of each reference text and at least one text tag, and the popularity value of each text tag, in the text tag database.
[0252] In some embodiments, the video field is stored in a tag field database; the form generation module 4553 is further configured to obtain multiple video fields and at least one tag corresponding to each video field, wherein the tag type includes text tags and video tags; and store the correspondence between each video field and at least one tag corresponding to each video field in the tag field database.
[0253] In some embodiments, the form generation module 4553 is further configured to obtain the video parameters and corresponding video parameter values included in each video field of the video form; obtain video materials that match the video parameters and corresponding video parameter values; wherein the video materials include any one of images, text, audio and video; and generate videos of recommended objects based on the obtained video materials.
[0254] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video form generation method described in this application.
[0255] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to execute the video form generation method provided in this application. For example, ... Figure 3A The method for generating video forms is shown.
[0256] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0257] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0258] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0259] In summary, this application's embodiments, by determining video tags based on similar videos of sample videos and text tags based on similar text of the descriptive information of the objects to be recommended, reduce the computational load required to obtain video and text tags, saving computational resources. Video fields are obtained based on the correspondence between tags and video fields, and video fields used to generate video forms are selected based on filtering indicators. By mining the core content from sample videos and descriptive information in a data-driven manner, video forms for generating recommended videos are accurately and efficiently generated. This allows the video forms to be used to generate more accurate recommended videos for the objects to be recommended, improving the timeliness and effectiveness of recommendations.
[0260] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for generating a video form, characterized in that, The method includes: Obtain descriptions of sample videos and objects to be recommended; Based on the video fingerprint of the sample video, obtain multiple similar videos of the sample video; Based on the video identifiers of the multiple similar videos, the correspondence between the reference video and the video tag is queried to obtain multiple video tags corresponding to the multiple similar videos, wherein each of the multiple similar videos corresponds to at least one of the video tags; Based on the text vector of the description information, obtain multiple similar texts of the description information; Based on the text identifiers of the multiple similar texts, the correspondence between the reference text and the text tags is queried to obtain multiple text tags corresponding to the multiple similar texts, wherein each of the multiple similar texts corresponds to at least one of the text tags; Based on the popularity values corresponding to the multiple video tags and the multiple text tags respectively, at least one tag is selected as the target tag from the multiple video tags and the multiple text tags; Based on the filtering criteria for each video field corresponding to the target tag, at least one of the video fields is selected to generate a video form, wherein the video form is used to generate videos for recommending the object to be recommended, and the video field includes the video parameter type and the corresponding video parameter value.
2. The method as described in claim 1, characterized in that, The step of obtaining multiple similar videos based on the video fingerprint of the sample video includes: Obtain the video fingerprint of the sample video; Determine the similarity between the video fingerprint of each reference video and the video fingerprint of the sample video. Select multiple reference videos from the top of the descending similarity ranking results as multiple similar videos of the sample video, or select multiple reference videos with similarity greater than a similarity threshold as multiple similar videos of the sample video.
3. The method as described in claim 2, characterized in that, The process of obtaining the video fingerprint of the sample video includes: The sample video is segmented based on a preset duration to obtain multiple video segments, and a video frame is extracted from each video segment. Feature extraction is performed on each video frame to obtain the video frame features corresponding to each video frame; The video frame features corresponding to each video frame are combined to obtain the video fingerprint of the sample video.
4. The method as described in claim 1, characterized in that, The step of obtaining multiple similar texts based on the text vector of the description information includes: Obtain the text vector of the description information; Determine the similarity between the text vector of each reference text and the text vector of the description information. Select multiple reference texts from the top of the descending similarity sorting results as multiple similar texts of the description information, or select multiple reference texts with similarity greater than a similarity threshold as multiple similar texts of the description information.
5. The method as described in claim 4, characterized in that, The process of obtaining the text vector containing the description information includes: The description information is segmented into words to obtain multiple words included in the description information; For each word, the following processing is performed: the word frequency is determined based on the number of times the word appears in the description information and the total number of words in the description information; the inverse document rate is determined based on the number of texts in the corpus that include the word and the total number of texts in the corpus; and the text component corresponding to the word is determined based on the word frequency and the inverse document rate. The text components corresponding to each word are combined to obtain the text vector of the descriptive information.
6. The method as described in claim 1, characterized in that, The step of selecting at least one tag as the target tag from among the multiple video tags and multiple text tags based on their respective popularity values includes: The popularity values corresponding to the plurality of video tags and the popularity values corresponding to the plurality of text tags are obtained respectively. The popularity value corresponding to the video tag is determined based on at least one of the usage frequency of the video tag and the recommendation effect data of the corresponding video. The popularity value corresponding to the text tag is determined based on at least one of the usage frequency of the text tag and the recommendation effect data of the corresponding video. The recommendation effect data of the video includes at least one of the following: number of exposures, number of clicks, and number of conversions. Based on the popularity values corresponding to the multiple video tags and the multiple text tags, the multiple video tags and the multiple text tags are sorted in descending order. At least one tag is selected from the head of the descending sort result as the target tag, or at least one tag with a popularity value greater than the popularity value threshold is selected as the target tag.
7. The method as described in claim 1, characterized in that, The step of selecting at least one of the video fields to generate a video form based on the filtering criteria for each video field corresponding to the target tag includes: Based on the target tag, query the correspondence between different tags and different video fields to obtain the video field corresponding to each target tag; Determine the filtering criteria corresponding to each of the video fields; Based on the filtering criteria corresponding to each video field, the multiple video fields are sorted in descending order, and at least one video field is selected as the target field from the head of the descending sort result of the filtering criteria. Generate a video form based on at least one of the target fields.
8. The method as described in claim 7, characterized in that, The step of determining the filtering criteria corresponding to each of the video fields includes: Obtain the weight values corresponding to multiple recommendation metrics for each video field, wherein the types of recommendation metrics include: the number of times the video field is used, the number of exposures of the video corresponding to the video field, the number of clicks of the video corresponding to the video field, and the number of conversions of the video corresponding to the video field; For each video field, the following processing is performed: the multiple recommendation indicators of the video field are weighted and summed based on their corresponding weight values to obtain the filtering indicators corresponding to the video field.
9. The method as described in claim 1, characterized in that, The step of selecting at least one of the video fields to generate a video form based on the filtering criteria for each video field corresponding to the target tag includes: Based on the target tag, query the correspondence between different tags and different video fields to obtain the video field corresponding to each target tag; Determine the filtering criteria corresponding to each video field, sort the multiple video fields in descending order based on the filtering criteria corresponding to each video field, and display at least a portion of the video fields at the top of the descending sort result; The target field is obtained through at least one of the following methods: in response to a selection operation for any video field among the at least some video fields, the selected video field is used as the target field; in response to a custom field input operation, the input custom video field is used as the target field. A video form is generated based on the target field.
10. The method as described in claim 9, characterized in that, The step of responding to a custom field input operation, after using the input custom video field as the target field, further includes: Establish a correspondence between each custom video field and each target tag, and store the correspondence between each custom video field and each target tag in a video field database; Based on the recommendation performance data of the videos corresponding to each of the custom video fields, the popularity value of each target tag is updated.
11. The method as described in claim 2, characterized in that, The video fingerprint of each reference video is stored in a video fingerprint database, and the video tag corresponding to each reference video and the popularity value of each video tag are stored in a video tag database; Before finding multiple similar videos based on the video fingerprint of the sample video and obtaining multiple video tags corresponding to the multiple similar videos, the method further includes: Acquire multiple reference videos, multiple video tags, and the popularity value of each video tag, and determine the video fingerprint corresponding to each reference video; The correspondence between the video identifier of each reference video and the video fingerprint of each reference video is stored in the video fingerprint database; For each of the reference videos, the following processing is performed: at least one of the video tags that matches the video content of the reference video is selected from the plurality of video tags, and a correspondence is established between the video identifier of the reference video and at least one of the video tags; The correspondence between the video identifier of each reference video and at least one video tag, and the popularity value of each video tag, are stored in the video tag database.
12. The method as described in claim 4, characterized in that, The text vector of each reference text is stored in a text vector database, and the text tag corresponding to each reference text and the popularity value of each text tag are stored in a text tag database; Before finding multiple similar texts based on the text vector of the description information and obtaining multiple text tags corresponding to the multiple similar texts, the method further includes: Obtain multiple reference texts, multiple text tags, and the popularity value of each text tag, and determine the text vector corresponding to each reference text; The correspondence between the text identifier of each reference text and the text vector of each reference text is stored in a text vector database; For each of the reference texts, the following processing is performed: at least one of the text tags that matches the text content of the reference text is selected from the plurality of text tags, and a correspondence is established between the text identifier of the reference text and at least one of the text tags; The correspondence between the text identifier of each reference text and at least one text tag, and the popularity value of each text tag, are stored in the text tag database.
13. The method as described in claim 1, characterized in that, The video field is stored in the tag field database; Before generating the video form by selecting at least one of the video fields based on the filtering criteria corresponding to the target tag, the method further includes: Obtain multiple video fields and at least one tag corresponding to each video field, wherein the tag type includes text tags and video tags; The correspondence between each video field and at least one tag corresponding to each video field is stored in the tag field database.
14. The method as described in claim 1, characterized in that, After generating the video form by selecting at least one of the video fields based on the filtering criteria corresponding to the target tag, the method further includes: Obtain the video parameter type and corresponding video parameter value for each video field in the video form; Obtain video material that matches the video parameter type and the corresponding video parameter value; wherein, the video material includes any one of image, text, audio and video; Based on the acquired video materials, the recommended videos of the objects to be recommended are generated.
15. A device for generating video forms, characterized in that, The device includes: The data acquisition module is used to acquire sample videos and descriptive information about the objects to be recommended; The tag acquisition module is used to acquire multiple similar videos of the sample video based on the video fingerprint of the sample video, query the correspondence between the reference video and the video tag based on the video identifier of the multiple similar videos, and obtain multiple video tags corresponding to the multiple similar videos, wherein each of the multiple similar videos corresponds to at least one video tag; and to acquire multiple similar texts of the description information based on the text vector of the description information, query the correspondence between the reference text and the text tag based on the text identifier of the multiple similar texts, and obtain multiple text tags corresponding to the multiple similar texts, wherein each of the multiple similar texts corresponds to at least one text tag; The tag acquisition module is further configured to select at least one tag as a target tag from the plurality of video tags and the plurality of text tags based on the popularity values corresponding to the plurality of video tags and the plurality of text tags respectively; The form generation module is used to select at least one of the video fields to generate a video form based on the filtering indicators of each video field corresponding to the target tag. The video form is used to generate videos for recommending the object to be recommended. The video field includes the video parameter type and the corresponding video parameter value.
16. The apparatus according to claim 15, characterized in that, The tag acquisition module is further configured to acquire the popularity values corresponding to the plurality of video tags and the popularity values corresponding to the plurality of text tags respectively. The popularity value corresponding to the video tag is determined based on at least one of the usage frequency of the video tag and the recommendation effect data of the corresponding video. The popularity value corresponding to the text tag is determined based on at least one of the usage frequency of the text tag and the recommendation effect data of the corresponding video. The recommendation effect data of the video includes at least one of the following: number of exposures, number of clicks, and number of conversions. Based on the popularity values corresponding to the multiple video tags and the multiple text tags, the multiple video tags and the multiple text tags are sorted in descending order. At least one tag is selected from the head of the descending sort result as the target tag, or at least one tag with a popularity value greater than the popularity value threshold is selected as the target tag.
17. The apparatus according to claim 15, characterized in that, The form generation module is also used to query the correspondence between different tags and different video fields based on the target tags, so as to obtain the video field corresponding to each target tag; Determine the filtering criteria corresponding to each video field, sort the multiple video fields in descending order based on the filtering criteria corresponding to each video field, and display at least a portion of the video fields at the top of the descending sort result; The target field is obtained by at least one of the following methods: in response to a selection operation for any video field among the at least some video fields, the selected video field is taken as the target field; In response to a custom field input operation, the input custom video field is used as the target field; A video form is generated based on the target field.
18. An electronic device for generating video forms, characterized in that, The electronic device includes: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the video form generation method according to any one of claims 1 to 14.
19. A computer-readable storage medium storing executable instructions, characterized in that, When the executable instructions are executed by the processor, they implement the video form generation method according to any one of claims 1 to 14.
20. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the video form generation method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Method and apparatus for generating information
CN109325148A
Multimedia data processing method and device and storage medium
CN110598014A
Video generation method and device, electronic equipment and storage medium
CN113099267A