A video frame positioning method, an electronic device and a computer readable storage medium
By constructing a tag feature library, the system automatically expands and matches tags based on the similarity between video tags, thus solving the problem of low intelligence caused by manual tag expansion and improving the intelligence and accuracy of video frame localization.
Patent Information
- Application Number
- CN202110776555.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-09
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-10-08
AI Technical Summary
In existing technologies, tag expansion during video frame localization relies on manual operation, resulting in low intelligence and easy omissions, affecting the comprehensiveness and accuracy of localization.
By constructing a tag feature library, multimedia tags are automatically expanded using the similarity between video tags, and then matched in the video frame text sequence to achieve intelligent positioning of video frames.
It improves the intelligence, efficiency, comprehensiveness and accuracy of video frame positioning, and can automatically expand tags and accurately locate video frames, thus improving the flexibility of video push.
Smart Images

Figure CN113821679B_ABST
Abstract
Description
Technical Field
[0001] This application relates to video processing technology in the field of computer internet, and more particularly to a video frame positioning method, an electronic device, and a computer-readable storage medium. Background Technology
[0002] A tag is a semantic description of information. To improve computational efficiency and recall, tags are often used for information processing. For example, for information to be pushed, the corresponding tags are usually used to locate the video frame in the video that is used to push the information, and then the information is pushed at the located video frame to achieve the push of the information.
[0003] Generally, to locate more video frames in a video that can be used to push push information based on the tags corresponding to the information to be pushed, the tags of the information to be pushed are usually expanded manually, and then the video frames used to push the information to be pushed are located in the video based on the expanded tags. However, in the process of locating the video frames used to push push information in the video, the expansion of the tags of the information to be pushed is done manually, and the intelligence of video frame localization is low. Summary of the Invention
[0004] This application provides a video frame positioning method, apparatus, electronic device, and computer-readable storage medium, which can improve the intelligence of video frame positioning.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides a video frame localization method, including:
[0007] Obtain the multimedia tag corresponding to the multimedia to be pushed, wherein the multimedia to be pushed is used to be presented during the playback of the video to be located;
[0008] The multimedia tags are expanded based on the tag feature library to obtain the tags to be matched. The tag feature library includes tag features corresponding to each video tag, and each tag feature is obtained based on the similarity between the video tags.
[0009] Based on the tag to be matched, text matching is performed on the video frame text sequence corresponding to the video to be located, and the video frame to be pushed in the video to be located is located based on the video frame corresponding to the matched video frame text.
[0010] This application provides a video frame positioning device, including:
[0011] The tag acquisition module is used to acquire the multimedia tags corresponding to the multimedia to be pushed, wherein the multimedia to be pushed is used to be presented during the playback of the video to be located;
[0012] The tag extension module is used to extend the multimedia tags based on the tag feature library to obtain the tags to be matched. The tag feature library includes tag features corresponding to each video tag, and each tag feature is obtained based on the similarity between the video tags.
[0013] The tag positioning module is used to perform text matching in the video frame text sequence corresponding to the video to be located based on the tag to be matched, and to locate the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text.
[0014] In this embodiment, the video frame localization device further includes a feature library acquisition module, configured to acquire at least one sample video and at least one sample tag corresponding to each sample video; construct each video tag corresponding to at least one sample video based on the at least one sample tag corresponding to each sample video, wherein one video tag is one sample tag; acquire a set of sample videos corresponding to each video tag based on the at least one sample tag corresponding to each sample video in the at least one sample video; determine the similarity between each video tag based on the set of sample videos corresponding to each video tag; determine the tag features corresponding to each video tag based on the similarity between each video tag, and determine the determined tag features corresponding to each video tag as the tag feature library.
[0015] In this embodiment of the application, the feature library acquisition module is further configured to perform sentiment analysis on each video tag to obtain a sentiment category; construct a target sentiment tag and a negative sentiment tag, wherein the target sentiment tag includes at least one of a positive sentiment tag and a neutral sentiment tag; based on the sentiment category, determine the associated sentiment tag corresponding to each video tag from the target sentiment tag and the negative sentiment tag; and determine the sentiment similarity between each video tag and the associated sentiment tag.
[0016] In this embodiment of the application, the feature library acquisition module is further configured to determine the tag features corresponding to each video tag based on the similarity between the various video tags and the sentiment similarity.
[0017] In this embodiment of the application, the feature library acquisition module is further configured to construct at least one sample tag set corresponding to each sample video based on at least one sample tag corresponding to each sample video; count the occurrence frequency of each sample tag in the sample tag set; filter the video tags whose occurrence frequency is greater than a frequency threshold from the sample tag set; and construct each video tag from the filtered video tags.
[0018] In this embodiment of the application, the feature library acquisition module is further configured to traverse each of the video tags, and obtain the similarity between the first traversed video tag and the second to the first i-th video tags, where i is the number of video tags in each video tag; and perform the following processing through iteration i: obtain the similarity between the i-th traversed video tag and the (i+1)-th to the first i-th video tags, where... And i is a positive integer variable with increasing value; the similarity between the first video tag obtained by iteration i and the second to the first video tags, and the similarity between the (I-1)th video tag and the first video tag, are determined as the similarity between each video tag.
[0019] In this embodiment of the application, the feature library acquisition module is further configured to perform the following processing through iteration j: when the sample video set corresponding to the i-th video tag and the sample video set corresponding to the j-th video tag include common sample videos, the tag weights that are positively correlated with the number of common sample videos and negatively correlated with the number of sample videos corresponding to both the i-th and j-th video tags are determined as the similarity between the i-th video tag and the j-th video tag, wherein, , and j is a positive integer variable with increasing value; when there are no common sample videos between the sample video set corresponding to the i-th video tag and the sample video set corresponding to the j-th video tag, the minimum similarity threshold is determined as the similarity between the i-th video tag and the j-th video tag; obtain the similarity between the i-th video tag obtained by iteration j and the (i+1)-1-1 video tags respectively.
[0020] In this embodiment of the application, the feature library acquisition module is further configured to, based on the similarity between the various video tags, determine the video tags whose similarity to the first-order video tags corresponding to each video tag is greater than a first similarity threshold as first-order similar video tags corresponding to each video tag; based on the similarity between the various video tags, determine the video tags whose similarity to the first-order similar video tags is greater than a second similarity threshold as second-order similar video tags corresponding to each video tag as second-order similar video tags corresponding to each video tag; and determine the tag features corresponding to each video tag based on the first-order similar video tags and the second-order similar video tags.
[0021] In this embodiment of the application, the tag expansion module is further configured to extract features to be expanded from the multimedia tags; obtain the feature similarity between the features to be expanded and each tag feature in the tag feature library; among the obtained feature similarities between the features to be expanded and each tag feature, combine the video tags corresponding to the feature similarities greater than a third similarity threshold into an expanded tag; and combine the expanded tag and the multimedia tag to obtain the tag to be matched.
[0022] In this embodiment of the application, the video frame localization device further includes a text recognition module, which is used to extract a video frame sequence from the video to be located based on video frame extraction information, wherein the video frame extraction information includes at least one of frame rate, number of bullet comments and keyframes; and to extract video frame text information from the video frame image corresponding to each video frame in the video frame sequence to obtain the video frame text information sequence corresponding to the video frame sequence.
[0023] In this embodiment of the application, the text recognition module is further configured to obtain the playback duration corresponding to the video to be located; when the playback duration is greater than the playback duration threshold, the video segment is extracted from the video to be located to obtain the video to be extracted.
[0024] In this embodiment of the application, the text recognition module is further configured to extract the video frame sequence from the video to be extracted based on the video frame extraction information.
[0025] In this embodiment of the application, the tag positioning module is further configured to obtain the character length corresponding to each sub-tag to be matched in the tags to be matched; sort the sub-tags to be matched in the tags to be matched based on the character length; and select the sub-tag to be matched with the longest character length from the sorted tags to be matched in turn for text matching in the video frame text sequence corresponding to the video to be located.
[0026] In this embodiment of the application, the video frame positioning device further includes a video frame merging module, which is used to obtain the frame interval between adjacent video frames in the video frame corresponding to the matched video frame text; merge adjacent video frames whose frame interval is less than the interval threshold, and determine the merged video frame as the video frame to be pushed in the video to be positioned for the multimedia to be pushed.
[0027] In this embodiment, the video frame positioning device further includes an information push module for playing the video to be positioned; when the playback progress of the video to be positioned reaches the video frame to be pushed, the multimedia to be pushed is presented.
[0028] This application provides an electronic device for video frame positioning, comprising:
[0029] Memory, used to store executable instructions;
[0030] The processor, when executing executable instructions stored in the memory, implements the video frame positioning method provided in the embodiments of this application.
[0031] This application provides a computer-readable storage medium storing executable instructions, which, when executed by a processor, implement the video frame positioning method provided in this application.
[0032] The embodiments of this application have at least the following beneficial effects: by pre-collecting various video tags and constructing a tag feature library using the tag features of each video tag obtained from the similarity between various video tags, after obtaining the multimedia tags corresponding to the multimedia to be pushed (the information to be pushed), the multimedia tags can be automatically expanded based on the tag feature library, and the video frame to be pushed can be located in the video to be located based on the expanded matching tags; thus, the expansion of multimedia tags is carried out automatically during the video frame location process. Therefore, the video frame location method provided by the embodiments of this application can improve the intelligence of video frame location. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of an optional architecture of the video frame positioning system provided in the embodiments of this application;
[0034] Figure 2 This is another optional architecture diagram of the video frame positioning system provided in the embodiments of this application;
[0035] Figure 3 This is provided by the embodiments of this application. Figure 1 A schematic diagram of an optional server structure;
[0036] Figure 4This is an optional flowchart illustrating the video frame positioning method provided in an embodiment of this application;
[0037] Figure 5 This is another optional flowchart illustrating the video frame positioning method provided in the embodiments of this application;
[0038] Figure 6 This is a schematic diagram of an exemplary second-order similarity video tag provided in an embodiment of this application;
[0039] Figure 7 This is a schematic diagram illustrating an exemplary video frame positioning process provided in an embodiment of this application;
[0040] Figure 8 This is a schematic diagram illustrating an exemplary similarity between various video tags provided in an embodiment of this application;
[0041] Figure 9 This is a schematic diagram illustrating an exemplary similarity between various video tags provided in an embodiment of this application;
[0042] Figure 10 This is a schematic diagram illustrating an exemplary method for obtaining a thesaurus provided in an embodiment of this application;
[0043] Figure 11 This is an exemplary schematic diagram of presenting multimedia to be pushed during the playback of a video to be located, provided by an embodiment of this application;
[0044] Figure 12 This is another exemplary schematic diagram of presenting multimedia to be pushed during the playback of a video to be located, provided by an embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0046] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0047] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0048] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0049] In the implementation of this application, the collection and processing of relevant data should be strictly in accordance with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0050] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0051] 1) Artificial Intelligence (AI) is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Therefore, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. In other words, AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities. In the embodiments of this application, AI can be used to locate video frames.
[0052] Furthermore, artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning. Moreover, with the research and advancement of AI technology, it has been researched and applied in numerous fields; for example, common applications include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, smart transportation, vehicle networking, drones, robots, smart healthcare, and smart customer service. As technology develops, AI will be applied in even more fields and play an increasingly important role. The application of AI in video processing as described in this application's embodiments will be explained later.
[0053] 2) Graph embedding is used to obtain the embedding features of each node in the graph. Embedding features are vectorized feature descriptions of each node in the graph, such as the label features in this embodiment. The distance between the embedding features of two nodes measures the semantic similarity between the two nodes; the more similar the two nodes are, the smaller the distance between their embedding features. In this embodiment, graph embedding can be used to obtain the label features corresponding to each video label.
[0054] 3) Graph, a computer data structure. A graph G=(V,E) is composed of nodes V and edges E between nodes. In the embodiments of this application, the similarity between video tags can be represented in the form of a graph.
[0055] 4) Optical Character Recognition (OCR) is used to recognize text information in images. In this embodiment, text information such as subtitles and bullet comments in a video can be extracted based on OCR to obtain the video frame text sequence corresponding to the video to be located.
[0056] 5) Blockchain is an encrypted, chain-like storage structure for transactions formed by blocks.
[0057] 6) Blockchain Network: A collection of nodes that incorporate new blocks into a blockchain through consensus.
[0058] Generally, locating video frames solely based on keywords of the information to be pushed to will miss some frames, resulting in a limited number of frames that can be located. Since more video frames increase the flexibility of push notifications, to locate more video frames for pushing information based on tags corresponding to the information to be pushed, the tags for the information to be pushed are usually expanded manually, and then the video frames for pushing the information are located based on the expanded tags. However, in the process of locating video frames for pushing information in the video, the expansion of the tags for the information to be pushed is done manually, resulting in low intelligence in video frame location. Furthermore, manual tag expansion is prone to omissions, leading to low comprehensiveness of the expanded tags and consequently low comprehensiveness and accuracy of the located video frames.
[0059] Based on this, embodiments of this application provide a video frame positioning method, apparatus, electronic device, and computer-readable storage medium, which can improve the intelligence, efficiency, comprehensiveness, and accuracy of video frame positioning. The following describes exemplary applications of the electronic device for video frame positioning (hereinafter referred to as the video frame positioning device) provided in embodiments of this application. The video frame positioning device provided in embodiments of this application can be implemented as various types of terminals such as smartphones, smartwatches, laptops, tablets, desktop computers, smart TVs, set-top boxes, smart in-vehicle devices, portable music players, personal digital assistants, dedicated messaging devices, and portable gaming devices, or it can be implemented as a server. The following will describe exemplary applications when the device is implemented as a server.
[0060] See Figure 1 , Figure 1 This is a schematic diagram of an optional architecture of the video frame positioning system provided in an embodiment of this application; as shown... Figure 1 As shown, to support a video frame positioning application, in the video frame positioning system 100, terminal 400 (terminal 400-1 and terminal 400-2 are shown as examples) connects to server 200 via network 300. Network 300 can be a wide area network (WAN), a local area network (LAN), or a combination of both. Figure 1 The example shown illustrates a scenario where the database 500 is independent of the server 200. Alternatively, the database 500 can also be integrated into the server 200, but this embodiment does not limit this.
[0061] Terminal 400 is used to obtain the video to be located, the multimedia to be pushed, and the video frame to be pushed from server 200 via network 300; play the video to be located, and when the playback progress of the video to be located reaches the video frame to be pushed (see the video frame image with the caption "I want to ride a roller coaster" shown in the graphical interface of terminal 400-1, and the video frame image with the caption "The supermarket next to our house is having a big sale" shown in the graphical interface of terminal 400-2), present the multimedia to be pushed (see the card with "XX Supermarket Anniversary Celebration" shown in the graphical interface of terminal 400-2).
[0062] Server 200 is used to obtain multimedia tags corresponding to the multimedia to be pushed, wherein the multimedia to be pushed is used to present during the playback of the video to be located; expand the multimedia tags based on the tag feature library to obtain the tags to be matched, wherein the tag feature library includes the tag features corresponding to each video tag, and each tag feature is obtained based on the similarity between the video tags; based on the tags to be matched, perform text matching in the video frame text sequence corresponding to the video to be located, and locate the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text. Server 200 is also used to send the video to be located, the multimedia to be pushed, and the video frame to be pushed to the terminal 400 via network 300.
[0063] In some embodiments of this application, server 200 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be directly or indirectly connected via wired or wireless communication, and this application does not impose any limitations on this.
[0064] The tag feature library involved in the video frame positioning method provided in this application embodiment, as well as the push multimedia and push video frames corresponding to the video to be located, can be stored on the blockchain.
[0065] Furthermore, in the video frame localization method provided in this application embodiment, the video frame localization device can act as a node on a blockchain; see also Figure 2 , Figure 2 This is another optional architecture diagram of the video frame positioning system provided in the embodiments of this application. Figure 2 In the video frame positioning system 100 shown, the server 200 locates the video frame to be pushed, and can also push the video frame to multiple terminals via the server 200. Figure 2The example shows terminals 400-1 and 400-2 sending push multimedia and push video frames corresponding to the video to be located.
[0066] In some embodiments, server 200, terminal 400-1, and terminal 400-2 can join the blockchain network 600 and become nodes within it. The type of blockchain network 600 is flexible and diverse; for example, it can be any type of public blockchain, private blockchain, or consortium blockchain. Taking a public blockchain as an example, any electronic device of any business entity can access the blockchain network 600 without authorization to act as a consensus node. For instance, terminal 400-1 is mapped to consensus node 600-1 in the blockchain network 600, server 200 is mapped to consensus node 600-2, and terminal 400-2 is mapped to consensus node 600-3.
[0067] Taking blockchain network 600 as a consortium blockchain as an example, server 200, terminal 400-1, and terminal 400-2 can become nodes after obtaining authorization and access blockchain network 600. Server 200 can determine the video frame to be pushed in the video to be located based on the tag feature library by executing a smart contract, and send the video frame to be pushed in the video to be located to blockchain network 600 for consensus. When the consensus is passed, the server then sends the video frame to be pushed in the video to be located to terminals 400-1 and 400-2. It can be seen that by having multiple nodes in the blockchain network reach a consensus to confirm the video frame to be pushed in the video to be located before sending it to terminals 400-1 and 400-2, the reliability and accuracy of video frame positioning can be improved.
[0068] See Figure 3 , Figure 3 This is provided by the embodiments of this application. Figure 1 A schematic diagram of an optional server architecture. Figure 3 The server 200 shown includes at least one processor 210, memory 250, and at least one network interface 220. In some embodiments of this application, the server 200 also includes a user interface 230. The various components of the server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to implement communication between these components. In addition to a data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 3 The general labeled all buses as Bus System 240.
[0069] Processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0070] User interface 230 includes one or more output devices 231 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 230 also includes one or more input devices 232, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0071] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices physically located away from the processor 210.
[0072] The memory 250 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 250 described in this application embodiment is intended to include any suitable type of memory.
[0073] In some embodiments of this application, memory 250 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, as illustrated below.
[0074] Operating system 251 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0075] The network communication module 252 is used to reach other computer devices via one or more (wired or wireless) network interfaces 220, exemplary network interfaces 220 including: Bluetooth, Wi-Fi, and Universal Serial Bus (USB), etc.
[0076] Presentation module 253 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 231 associated with user interface 230 (e.g., a display screen, a speaker, etc.).
[0077] The input processing module 254 is used to detect and translate one or more user inputs or interactions from one or more input devices 232.
[0078] In some embodiments of this application, the video frame positioning device provided in this application can be implemented in software. Figure 3 A video frame localization device 255 stored in memory 250 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a tag acquisition module 2551, a tag expansion module 2552, a tag localization module 2553, a feature library acquisition module 2554, a text recognition module 2555, a video frame merging module 2556, and an information push module 2557. These modules are logically connected and can therefore be arbitrarily combined or further separated according to their implemented functions. The functions of each module will be described below.
[0079] In some other embodiments of this application, the video frame positioning device provided in this application can be implemented in hardware. As an example, the video frame positioning device provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the video frame positioning method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0080] The following will describe the video frame positioning method provided in this application embodiment, with reference to exemplary applications and implementations of the video frame positioning device provided in the embodiments of this application.
[0081] See Figure 4 , Figure 4 This is an optional flowchart illustrating the video frame localization method provided in this application embodiment, which will be combined with... Figure 4 The steps shown are explained.
[0082] S401. Obtain the multimedia tag corresponding to the multimedia to be pushed.
[0083] In this embodiment of the application, when the video frame positioning device performs video frame positioning in the video to be positioned, it is based on the tag of the multimedia to be pushed. Thus, the video frame positioning device obtains the tag of the multimedia to be pushed, and thus obtains the multimedia tag; that is, the multimedia tag is the tag of the multimedia to be pushed.
[0084] It should be noted that the multimedia to be pushed is a type of multimedia information to be pushed, such as advertisements or articles; and the multimedia to be pushed can be any form of information such as videos, pictures, cards, tags, etc., and this application embodiment does not specifically limit this. Furthermore, corresponding tags are pre-set for the multimedia to be pushed. These pre-set tags for the multimedia to be pushed are multimedia tags, which include at least one tag. For example, when the multimedia to be pushed is an advertisement video for an amusement park, the multimedia tags would be "amusement park, tourism, party, happiness, joy, weekend, holiday, relaxation, go out to play, play and entertainment," etc. In addition, the video to be located is a pushing medium for the multimedia to be pushed. The multimedia to be pushed is pushed by being inserted into the video to be located; that is, the multimedia to be pushed is used to be presented during the playback of the video to be located.
[0085] S402. Expand the multimedia tags based on the tag feature library to obtain the tags to be matched.
[0086] In this embodiment, the video frame localization device pre-obtains a tag feature library, which includes various tag features corresponding to each video tag. Each tag feature is obtained based on the similarity between various video tags. Based on the features of the multimedia tag, the video frame localization device matches the tag features corresponding to the tag features with semantic similarity greater than a threshold in the tag feature library, and determines the matched tag and the multimedia tag together as the tag to be matched.
[0087] It should be noted that the tags to be matched include multimedia tags, as well as tags that are semantically similar to multimedia tags and matched based on a tag feature library. These semantically similar tags are obtained by expanding upon the multimedia tags. Furthermore, in the tag feature library, each video tag corresponds one-to-one with each tag feature.
[0088] It should also be noted that each label feature in the label feature library can be obtained based on at least one-order similar labels of the label corresponding to each label feature.
[0089] S403. Based on the tag to be matched, perform text matching in the video frame text sequence corresponding to the video to be located, and locate the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text.
[0090] In this embodiment, the video frame localization device pre-obtains a sequence of video frame texts corresponding to the video to be localized. Each video frame text in the sequence is text information in a video frame. Here, since the tag to be matched is also text information, the video frame localization device matches the tag to be matched with each video frame text. If it is determined that the video frame text includes the tag to be matched or a sub-tag of the tag to be matched, the video frame text is determined to be a video frame text that matches the tag to be matched. The video frame localization device determines all matched video frame texts as the matched video frame texts.
[0091] It should be noted that since each video frame text in the video frame text sequence corresponds to one video frame, the video frame positioning device can determine the video frame corresponding to each video frame text in the matched video frame text as the video frame to be pushed in the video to be located, or it can integrate the video frames corresponding to each video frame text in the matched video frame text and determine the integrated video frame as the video frame to be pushed in the video to be located. This application embodiment does not specifically limit this.
[0092] Understandably, the video frame positioning device uses a pre-obtained tag feature library to achieve automatic expansion of multimedia tags, which also increases the number and comprehensiveness of the expanded tags. As a result, when the number of video frames corresponding to the video frames to be pushed matched based on the tags to be matched is also large, the flexibility is high when inserting the multimedia to be pushed into the video to be positioned based on the matched video frames.
[0093] See Figure 5 , Figure 5 This is another optional flowchart illustrating the video frame positioning method provided in the embodiments of this application; as shown below. Figure 5 As shown in the embodiment of this application, S404 to S408 are included before S402; that is, before the video frame positioning device expands the multimedia tags based on the tag feature library to obtain the tag to be matched, the video frame positioning method also includes S404 to S408. Each step is described below.
[0094] S404. Obtain at least one sample video and at least one sample label corresponding to each sample video.
[0095] In this embodiment, the video frame localization device locates the tag feature library based on the video tags. Here, the video frame localization device first collects at least one sample video and collects the tag corresponding to each sample video in the at least one sample video, and the tag corresponding to each sample video is at least one sample tag; wherein, the sample tag is a tag for each sample video, and the number of sample tags contained in different sample videos can be the same or different.
[0096] S405. Based on at least one sample tag corresponding to each sample video, construct video tags corresponding to at least one sample video.
[0097] In this embodiment, the video frame localization device integrates at least one sample tag corresponding to each sample video in at least one sample video, thereby integrating all sample tags corresponding to at least one sample video. Here, the video frame localization device can determine all sample tags as individual video tags, or it can extract video tags from all sample tags to obtain individual video tags; this embodiment does not specifically limit this. A video tag is a sample tag.
[0098] S406. Based on at least one sample label corresponding to each sample video in at least one sample video, obtain the set of sample videos corresponding to each video label.
[0099] It should be noted that the sample video set corresponding to each video tag refers to the set of sample videos corresponding to each video tag. For each video tag, the video frame localization device acquires all sample videos corresponding to that video tag and combines these acquired sample videos into a sample video set.
[0100] S407. Based on the sample video set corresponding to each video tag, determine the similarity between each video tag.
[0101] In this embodiment, when the video frame localization device calculates the similarity between any two video tags, it determines the similarity based on sample videos that co-occur between the two video tags, and the similarity between any two video tags is positively correlated with the number of sample videos that co-occur between them. Here, when the video frame localization device obtains the similarity between any two video tags, it also obtains the similarity between all the video tags.
[0102] S408. Based on the similarity between various video tags, determine the tag features corresponding to each video tag, and define the determined tag features corresponding to each video tag as a tag feature library.
[0103] It should be noted that the video frame localization device determines the feature representation of each video tag based on the similarity between various video tags, for video tags with at least one order of similarity corresponding to each video tag, and then determines the tag feature corresponding to each video tag. Once the video frame localization device obtains the tag feature corresponding to each video tag, it also obtains the tag feature library corresponding to each video tag.
[0104] In this embodiment, after obtaining the tag feature library, the video frame positioning device can update the tag feature library based on newly collected sample videos, and then expand the multimedia tags based on the updated tag feature library. The process of updating the tag feature library involves combining the newly collected sample videos with at least one previous sample video, based on the process described in S404 to S408, to obtain the updated tag feature library.
[0105] Understandably, video frame localization devices collect video tags from sample videos to build a tag feature library, resulting in a high degree of coverage for each video tag in the library. Compared to building a tag feature library through model training, the construction efficiency is higher, and the resulting tag feature library is real-time. Therefore, when hotspots in videos have short durations and update rapidly, it is more suitable for pushing multimedia in the video scene, resulting in higher accuracy of the matching tags and improving the accuracy of pushing multimedia.
[0106] In addition, since labels are usually short string sequences such as words and phrases, the differences between labels determined based on length are not large. Therefore, video frame localization devices determine the corresponding similarity based on the number of sample videos that appear in both video labels. This is more accurate than similarity determined by the distance between labels (such as edit distance or longest common subsequence).
[0107] In this embodiment of the application, S407 is followed by S409 to S412; that is, after the video frame localization device determines the similarity between each video tag based on the sample video set corresponding to each video tag, the video frame localization method further includes S409 to S412. Each step is described below.
[0108] S409. Perform sentiment analysis on each video tag to obtain the sentiment category.
[0109] It should be noted that when acquiring the tag features corresponding to each video tag within each video tag, the video frame localization device can also combine the sentiment of each video tag with at least first-order similarity. Thus, the video frame localization device performs sentiment analysis on each video tag to determine the sentiment category to which each video tag belongs. The sentiment category can be at least one of positive sentiment, negative sentiment, and neutral sentiment.
[0110] S410. Construct target sentiment labels and negative sentiment labels.
[0111] In this embodiment of the application, the video frame positioning device constructs virtual tags in each video tag: a target emotion tag and a negative emotion tag; wherein, the target emotion tag includes at least one of a positive emotion tag and a neutral emotion tag, so that the target emotion tag corresponds to at least one of a positive emotion and a negative emotion, and the negative emotion tag corresponds to a negative emotion tag.
[0112] S411. Based on the sentiment category, determine the associated sentiment tag corresponding to each video tag from the target sentiment tag and the negative sentiment tag.
[0113] It should be noted that when the video frame localization device determines the associated sentiment tag for each video tag based on the sentiment category, if the target sentiment tag includes a positive sentiment tag, then the associated sentiment tag is the target sentiment tag when the sentiment category of the video tag is positive; if the sentiment category of the video tag is neutral, then there is no associated sentiment tag and no corresponding sentiment similarity. If the target sentiment tag includes both positive and neutral sentiment tags, then the associated sentiment tag is the target sentiment tag when the sentiment category of the video tag is either positive or neutral; if the sentiment category of the video tag is negative, then the associated sentiment tag is the negative sentiment tag.
[0114] S412. Determine the sentiment similarity between each video tag and its associated sentiment tag.
[0115] It should be noted that since the associated sentiment tag is a virtual tag corresponding to the sentiment category of each video tag, the video frame localization device determines the sentiment similarity threshold (e.g., 1) as the sentiment similarity between each video tag and the associated sentiment tag.
[0116] Accordingly, S408 can be implemented through S4081; that is, the video frame localization device determines the tag features corresponding to each video tag based on the similarity between each video tag, including S4081, which will be explained below.
[0117] S4081. Based on the similarity and sentiment similarity between various video tags, determine the tag features corresponding to each video tag.
[0118] In this embodiment, the video frame localization device combines the similarity between various video tags and the sentiment similarity to determine the tag features corresponding to each video frame tag.
[0119] Understandably, video frame positioning devices use the emotional category of video tags as information to determine tag features, widening the distance between tags representing positive or neutral emotions and tags representing negative emotions. This allows the multimedia to be pushed to the video frames representing positive or neutral emotions, thus improving the push effect.
[0120] In this embodiment, S405 can be implemented through S4051 to S4054; that is, the video frame positioning device constructs at least one video tag corresponding to each sample video based on at least one sample tag corresponding to each sample video, including S4051 to S4054. Each step is described below.
[0121] S4051. Based on at least one sample label corresponding to each sample video, construct a sample label set corresponding to at least one sample video.
[0122] It should be noted that the video frame localization device acquires all the sample tags corresponding to at least one sample video and determines all the sample tags corresponding to at least one sample video as the sample tag set; that is, the sample tag set is all the sample tags corresponding to at least one sample video.
[0123] S4052. Count the frequency of occurrence of each sample label in the sample label set.
[0124] In this embodiment of the application, the video frame positioning device counts the number of times each sample tag in the sample tag set appears in at least one sample video, and determines the occurrence frequency of each sample tag as the number of times each sample tag appears in at least one sample video.
[0125] S4053. Filter video tags that appear more frequently than the frequency threshold from the sample tag set.
[0126] In this embodiment of the application, the video frame positioning device compares the occurrence frequency of each sample tag with a frequency threshold, and deletes the sample tags whose occurrence frequency is less than or equal to the frequency threshold from the sample tag set, while selecting the sample tags whose occurrence frequency is greater than the frequency threshold as video tags.
[0127] S4054. The selected video tags are used to construct individual video tags.
[0128] It should be noted that the video frame positioning device combines all video tags that appear more frequently than a frequency threshold, thus obtaining each video tag; therefore, the frequency of each video tag in each video tag is greater than the frequency threshold.
[0129] Understandably, video frame localization devices reduce the number of video tags and improve computational efficiency by filtering video tags based on their frequency of occurrence in at least one sample video, thereby improving the efficiency of building a tag feature library.
[0130] In this embodiment, S407 can be implemented through S4071 to S4073; that is, the video frame positioning device determines the similarity between each video tag based on the sample video set corresponding to each video tag, including S4071 to S4073. Each step is described below.
[0131] S4071. Traverse each video tag and obtain the similarity between the first video tag and the second to ith video tags respectively.
[0132] It should be noted that the video frame localization device obtains the similarity between video tags by traversing each video tag and calculating the similarity between any two video tags within each video tag. Here, for the first video tag encountered, the video frame localization device obtains the similarity between the first video tag and each of the following video tags, from the second to the i-th video tags. Here, i represents the number of video tags in each video tag.
[0133] S4072. Perform the following processing through iteration i: obtain the similarity between the i-th video tag and the (i+1)-th to the 1-th video tags respectively.
[0134] It should be noted that the video frame localization device obtains the similarity between each video tag through iteration i. During each iteration i, the similarity between the i-th video tag and each of the (i+1)-1-1 video tags is obtained. , where i is a positive integer variable whose value increases and whose increment step is 1.
[0135] In this embodiment of the application, when the video frame localization device obtains the similarity between the traversed i-th video tag and the (i+1)-th to i-th video tags, it can achieve this through iteration j, where, And j is a positive integer variable with an increasing value and an increment step size of 1; the video frame localization device performs the following processing through iteration j: when the sample video set corresponding to the i-th video tag and the sample video set corresponding to the j-th video tag include common sample videos, the tag weights that are positively correlated with the number of common sample videos and negatively correlated with the number of sample videos corresponding to the i-th video tag and the j-th video tag are determined as the similarity between the i-th video tag and the j-th video tag; when the sample video set corresponding to the i-th video tag and the sample video set corresponding to the j-th video tag do not include common sample videos, the minimum similarity threshold (e.g., 0) is determined as the similarity between the i-th video tag and the j-th video tag; finally, the video frame localization device obtains the similarity between the i-th video tag obtained by iteration j and the (i+1)-th to the i-th video tags respectively.
[0136] It should be noted that the similarity between any two video tags is positively correlated with the number of common sample videos between the two video tags and negatively correlated with the number of sample videos corresponding to each of the two video tags.
[0137] For example, the similarity between the i-th video tag and the j-th video tag can be achieved by equation (1), which is:
[0138] (1)
[0139] in, Let be the similarity between the i-th video tag and the j-th video tag, with a value range of . , Let i be the set of sample videos corresponding to the i-th video tag. Let j be the set of sample videos corresponding to the j-th video tag. The number of public sample videos, Let be the number of sample videos corresponding to the i-th video tag. denoted as the number of sample videos corresponding to the j-th video tag.
[0140] Understandably, the video frame localization device determines the similarity between the i-th and j-th video tags by combining the number of common sample videos between the i-th and j-th video tags, the number of sample videos corresponding to the i-th video tag, and the number of sample videos corresponding to the j-th video tag. This results in a high accuracy of the obtained similarity score. Furthermore, it also ensures that the obtained similarity score ranges within a certain range. This is beneficial for obtaining label features.
[0141] S4073. The similarity between the first video tag obtained in iteration i and the second to the first video tags, and the similarity between the (I-1)th video tag and the first video tag, are determined as the similarity between each video tag.
[0142] It should be noted that the video frame localization device can obtain the similarity between the second video tag and the third to the first video tags through iterative processing i, and so on, until the similarity between the (I-1)th video tag and the first video tag. Then, by combining the similarity between the first video tag and the second to the first video tags, the similarity between each video tag is obtained.
[0143] In this embodiment, S408 can also be implemented through S4082 to S4084; that is, the video frame positioning device determines the tag features corresponding to each video tag based on the similarity between each video tag, including S4082 to S4084. Each step is described below.
[0144] S4082. Based on the similarity between various video tags, among the first-order video tags corresponding to each video tag, the video tags with a similarity greater than the first similarity threshold are determined as the first-order similar video tags corresponding to each video tag.
[0145] It should be noted that the first-order video tag of a video label is the video tag adjacent to another video tag, and the adjacent video tags are those that share common sample videos with that video tag. Here, the video frame localization device filters video tags with a similarity greater than a first similarity threshold from the first-order video tags, and combines the filtered video tags with a similarity greater than the first similarity threshold into first-order similar video tags corresponding to each video tag.
[0146] S4083. Based on the similarity between various video tags, among the second-order video tags corresponding to each video tag, the video tags whose similarity to the first-order similar video tags is greater than the second similarity threshold are determined as the second-order similar video tags corresponding to each video tag.
[0147] It should be noted that the second-order video tag of a video tag is the first-order video tag corresponding to that video tag's first-order video tag. Here, the video frame localization device selects video tags from the second-order video tags whose similarity to first-order similar video tags is greater than a second similarity threshold, and combines the selected video tags with similarity greater than the second similarity threshold into a second-order similar video tag corresponding to each video tag. The first similarity threshold and the second similarity threshold can be the same, different, or the first similarity threshold can be less than the second similarity threshold, etc., and this embodiment does not limit this.
[0148] S4084. Based on first-order similar video labels and second-order similar video labels, determine the label features corresponding to each video label.
[0149] In this embodiment, the video frame localization device uses first-order similar video tags and second-order similar video tags as contextual information for US video tags to determine the tag features corresponding to each video tag. For example, the video frame localization device uses embedding feature algorithms such as "LINE", "SDNE", or "GCN" to obtain tag features.
[0150] For example, see Figure 6 , Figure 6 This is a schematic diagram of an exemplary second-order similarity video tag provided in an embodiment of this application; as shown... Figure 6 As shown, the structure Figure 6-1 In the diagram, each circle represents a video tag, and the thickness of the edges between video tags indicates their similarity. Video tag 6-11 is a first-order video tag of video tag 6-12. Since the edge between video tags 6-11 and 6-12 is thicker than the standard edge, their similarity is greater than the first similarity threshold, making video tag 6-11 a first-order similar video tag of video tag 6-12. Video tag 6-13 is a second-order video tag of video tag 6-12. Since video tags 6-13 and 6-12 share common first-order video tags 6-14 to 6-17, video tag 6-13 is a second-order similar video tag of video tag 6-12.
[0151] Understandably, since at least one sample label corresponding to a sample video is obtained from at least one dimension, and at least one sample label corresponds to at least one dimension, the expansion effect is poor when the label feature library built based on the first-order similarity labels of video labels is expanded. The video frame localization device determines the label features corresponding to each video label based on the first-order similar video labels and the second-order similar video labels, and then builds a label feature library. When expanding the labels using the label feature library, the comprehensiveness of the expanded labels to be matched can be improved, thereby improving the expansion effect.
[0152] In this embodiment, S402 can be implemented through S4021 to S4024; that is, the video frame positioning device expands the multimedia tags based on the tag feature library to obtain the tags to be matched, including S4021 to S4024. Each step is described below.
[0153] S4021. Extract the features to be expanded from the multimedia tags.
[0154] It should be noted that the video frame positioning device extracts features from the multimedia tags, and the extracted features are the features to be expanded; that is, the features to be expanded are the features of the multimedia tags.
[0155] S4022. Obtain the feature similarity between the feature to be expanded and each label feature in the label feature library.
[0156] In this embodiment of the application, the video frame positioning device calculates the feature similarity between the feature to be expanded and each tag feature in the tag feature library, thereby obtaining the feature similarity corresponding to each tag feature library; wherein, each feature similarity corresponds one-to-one with each tag feature, and each feature similarity corresponds one-to-one with each video tag.
[0157] For example, the feature similarity between the feature to be expanded and each tag feature in the tag feature library can be achieved by equation (2), which is:
[0158] (2)
[0159] in, For feature similarity, Features to be expanded These are label features.
[0160] S4023. Among the feature similarities between the obtained features to be expanded and the features corresponding to each label feature, the video labels corresponding to the feature similarities greater than the third similarity threshold are combined into expanded labels.
[0161] It should be noted that the video frame localization device filters out feature similarities greater than the third similarity threshold from various feature similarities, and combines the video tags corresponding to the filtered feature similarities greater than the third similarity threshold to obtain extended tags; where the extended tags are tags that are similar to the multimedia tags based on the tag feature library, that is, the extended tags are the extended tags corresponding to the multimedia tags.
[0162] S4024. Combine extended tags and multimedia tags to obtain the tags to be matched.
[0163] In this embodiment, the video frame positioning device combines the extended tags and multimedia tags into a tag to be matched.
[0164] In this embodiment of the application, S403 is preceded by S413 and S414; that is, before the video frame positioning device performs text matching in the video frame text sequence corresponding to the video to be located based on the tag to be matched, the video frame positioning method further includes S413 and S414, and each step is described below.
[0165] S413. Based on video frame extraction information, extract video frame sequences from the video to be located.
[0166] It should be noted that the video frame extraction information includes at least one of the following: frame rate, number of bullet comments, and keyframes; where frame rate refers to the number of video frames extracted within the frame extraction period; number of bullet comments is the number of bullet comments corresponding to a video frame, and frame extraction can be performed if the number of bullet comments is greater than the number threshold; keyframes are the content representation frames of the video to be located.
[0167] In this embodiment of the application, the video frame localization device performs frame extraction from the video to be located before matching tags, which can reduce the amount of matching calculation and improve text matching efficiency while ensuring the coverage of video frames in the video to be located.
[0168] S414. Extract video frame text information from the video frame image corresponding to each video frame in the video frame sequence to obtain a video frame text information sequence corresponding to the video frame sequence.
[0169] It should be noted that the video frame localization device acquires the corresponding video frame image and extracts the text information for each video frame in the extracted video frame sequence, thus obtaining the text information of a video frame corresponding to the video frame; thereby, it obtains a sequence of video frame text information corresponding to the video frame sequence. Here, one video frame in the video frame sequence corresponds to one video frame text information in the video frame text information sequence.
[0170] In this embodiment, steps S415 and S416 are included before step S413; that is, before the video frame positioning device extracts the video frame sequence from the video to be positioned based on the video frame extraction information, the video frame positioning method further includes steps S415 and S416. Each step is described below.
[0171] S415. Obtain the playback duration of the video to be located.
[0172] It should be noted that the playback duration obtained by the video frame positioning device refers to the duration of the video to be located; the video frame positioning device determines the frame extraction range of the video to be located based on this playback duration.
[0173] S416. When the playback duration exceeds the playback duration threshold, extract video segments from the video to be located to obtain the video to be extracted.
[0174] In this embodiment, the video frame localization device can obtain a playback duration threshold (e.g., 10 minutes), which is used to determine the frame extraction range of the video to be located. Here, the video frame localization device compares the playback duration with the playback duration threshold. When the playback duration is greater than the playback duration threshold, it can extract the beginning and end of the video to be located based on the beginning duration threshold (e.g., 5 minutes) and the end duration threshold (e.g., 5 minutes) to obtain the intermediate video segment. The obtained intermediate video segment is the video segment extracted from the video to be located: the frame extraction video. The frame extraction video is the frame extraction range of the video to be located. When the playback duration is less than or equal to the playback duration threshold, the video to be located can be determined as the frame extraction video, that is, the video frame localization device directly extracts the video frame sequence from the video to be located.
[0175] Understandably, video frame positioning devices combine the playback duration of the video to be positioned to determine the video segments used for frame extraction, so as to present the media to be pushed in video segments with a high probability of playback, thereby improving the push effect of the media to be pushed.
[0176] Accordingly, in this embodiment of the application, the video frame positioning device in S413 extracts a video frame sequence from the video to be positioned based on the video frame extraction information, including: the video frame positioning device extracts a video frame sequence from the video to be extracted based on the video frame extraction information.
[0177] It should be noted that after the video frame positioning device extracts video segments from the video to be positioned, it extracts frames from the extracted video segments, i.e., the video to be extracted.
[0178] In this embodiment, S403 can be implemented through S4031 to S4033; that is, the video frame positioning device performs text matching in the video frame text sequence corresponding to the video to be located based on the tag to be matched, including S4031 to S4033. Each step is described below.
[0179] S4031. Obtain the character length corresponding to each sub-tag to be matched in the tag to be matched.
[0180] It should be noted that the tag to be matched includes at least one sub-tag to be matched. The video frame positioning device obtains the character length of each sub-tag to be matched in the tag to be matched, and obtains the character length corresponding to each sub-tag to be matched.
[0181] S4032. Sort the sub-tags to be matched in the tag to be matched based on the character length.
[0182] It should be noted that the video frame positioning device sorts the sub-tags to be matched in the tags to be matched based on the character length; here, the sorted tags to be matched can be sorted in ascending order based on the character length or in descending order based on the character length, and this application embodiment does not specifically limit this.
[0183] S4033. From the sorted tags to be matched, select the sub-tag with the longest character length in turn and perform text matching in the text sequence of the video frame corresponding to the video to be located.
[0184] It should be noted that the video frame positioning device is based on the principle of maximizing matching. In the video frame text sequence, it prioritizes text matching of the sub-label with the longest character length among the sorted labels to be matched.
[0185] In this embodiment of the application, S403 is followed by S417 and S418; that is, after the video frame positioning device locates the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text, the video frame positioning method further includes S417 and S418, and each step is described below.
[0186] S417. Obtain the frame interval between adjacent video frames in the video frame corresponding to the matched video frame text.
[0187] It should be noted that the video frame localization device can integrate the video frames corresponding to the obtained matching video frame text before pushing the multimedia to be pushed. During integration, the video frame localization device determines whether to merge based on the frame interval between adjacent video frames in the video frames corresponding to the matching video frame text. Adjacent video frames include a sequence consisting of at least two adjacent video frames.
[0188] S418. Merge adjacent video frames with a frame interval less than the interval threshold, and determine the merged video frame as the video frame to be pushed in the video to be located for the multimedia to be pushed.
[0189] It should be noted that the video frame positioning device compares the obtained frame interval with the interval threshold (e.g., 2 minutes). When it is determined that the frame interval is less than the interval threshold, it performs integration processing, merging adjacent video frames into one video frame, and determining the merged video frame as the video frame to be pushed in the video to be located for the multimedia to be pushed.
[0190] See also Figure 5In this embodiment of the application, S403 is followed by S419 and S420; that is, after the video frame positioning device locates the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text, the video frame positioning method further includes S419 and S420. Each step is described below.
[0191] S419. Play the video to be located.
[0192] It should be noted that the playback of the video to be located can be on a video frame positioning device, i.e., the video frame positioning device described in S419 plays the video to be located, or it can be on other playback devices (e.g., Figure 1 The video frame positioning device will send the video to be positioned to the other playback device so that the positioning video can be played on the other playback device.
[0193] S420: When the playback progress of the video to be located reaches the frame of the video to be pushed, the multimedia to be pushed is displayed.
[0194] It should be noted that, during the playback of the video to be located, if the video frame positioning device determines that the playback progress of the video to be located has reached the frame to be pushed, it will present the multimedia to be pushed. The presentation mode of the multimedia to be pushed can be accompanying the video to be located, or it can be parallel to the video to be located, etc., and this application embodiment does not specifically limit this; here, the presentation mode of the multimedia to be pushed can match the form of the multimedia to be pushed.
[0195] In some embodiments of this application, the video frame positioning method provided in the embodiments of this application can be used to locate the video frame to be pushed in multiple videos to be located, so as to realize the push of one multimedia to be pushed in multiple videos to be located; or the video positioning method provided in the embodiments of this application can be used to locate the video frames to be pushed corresponding to multiple multimedia to be pushed, so as to realize the push of multiple multimedia to be pushed in one video to be located.
[0196] In this embodiment, S402 is preceded by S421; that is, before the video frame localization device expands the multimedia tags based on the tag feature library to obtain the tag to be matched, the video frame localization method further includes S421. Each step is described below.
[0197] S421. Expand the multimedia tags based on the synonym tag model to obtain the initial tags to be matched.
[0198] It should be noted that, in order to obtain more tags to be expanded, video frame localization devices can first use a synonym tag model to expand multimedia tags with synonym tags. The synonym tag model includes at least one of a synonym tag library and a network model for determining synonyms. The synonym tag library includes various sets of synonym tags, and the network model for determining synonyms can be a neural network model from artificial intelligence.
[0199] Accordingly, in this embodiment of the application, in S402, the video frame positioning device expands the multimedia tags based on the tag feature library to obtain the tags to be matched, including: the video frame positioning device expands the initial tags to be matched based on the tag feature library to obtain the tags to be matched.
[0200] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0201] See Figure 7 , Figure 7 This is a schematic diagram illustrating an exemplary video frame positioning process provided in an embodiment of this application; as shown... Figure 7 As shown, this exemplary video frame localization process includes a training phase 7-1 and a prediction phase 7-2.
[0202] Training phase 7-1 includes the following steps:
[0203] S701, Collect videos 7-11 (at least one sample video).
[0204] Here, you can obtain Video 7-11 by collecting short videos, TV series clips, and trailer videos.
[0205] S702, Building Tags Figure 7-12 (Similarity between various video tags).
[0206] It should be noted that each video (each sample video) in Videos 7-11 corresponds to multiple tags (at least one sample tag), and the number of tags varies among different videos in Videos 7-11.
[0207] For example, video 7-11 includes video ,video and video ,video The corresponding tags include tags ,Label and tags ,video The corresponding tags include tags and tags ,video The corresponding tags include tags and tags As shown in Table 1:
[0208] Table 1
[0209]
[0210] use Indicates label Given the video set (sample video set), the video set corresponding to each tag can be obtained based on Table 1, as shown in Equation (3):
[0211] (3)
[0212] Based on formula (3) to construct tags Figure 7-12 ,Label Figure 7-12 It can be expressed by equation (4), which is shown below:
[0213] (4)
[0214] in, For tags Figure 7-12 , It is the set of nodes, that is, the set of all tags (each video tag), so V is as shown in equation (5):
[0215] (5)
[0216] E is an edge of V, and the process of determining it is as follows:
[0217] when hour, and There exists an edge between them, and and The weights (similarity) of the edges between them can be obtained through equation (1). See also... Figure 8 , Figure 8 This is an exemplary schematic diagram illustrating the similarity between various video tags provided in an embodiment of this application; as shown... Figure 8 As shown, the results obtained based on equation (3) Figure 7 Tags in Figure 7-12 : and The similarity between them is 0.50. and The similarity between them is 0.71. and The similarity between them is 0.50. and The similarity between them is 0.71. and The similarity between them is 0.50.
[0218] It should also be noted that in constructing tags Figure 7-12 At the same time, in order to improve computational efficiency and reduce labels Figure 7-12 The number of nodes (labels) in the sample can be determined by using Equation (6) to filter the label set (sample label set), removing labels whose frequency is below a threshold (frequency threshold), and then re-labeling. Figure 7-12 The construction is as follows: Equation (6) is shown below:
[0219] (6)
[0220] Where, τ threshold, For tags Frequency of occurrence.
[0221] S703, Constructing Emotional Labels Figure 7-1 3.
[0222] It should be noted that in the label Figure 7-12 Two virtual nodes are added: a positive sentiment node (target sentiment tag) and a negative sentiment node (negative sentiment tag). Simultaneously, the tags are... Figure 7-12 Each tag in the data undergoes sentiment analysis to determine its corresponding sentiment category: positive, neutral, or negative. Figure 7-12 When the sentiment category of a tag is positive, an edge is established between the tag and the positive sentiment node, and the weight of the edge between the tag and the positive sentiment node is 1 (sentiment similarity); when the tag... Figure 7-12 When the sentiment category of a tag is negative, an edge is created between the tag and the negative sentiment node, with a weight of 1 (sentiment similarity); thus constructing the sentiment tag. Figure 7-1 3.
[0223] See Figure 9 , Figure 9 This is an exemplary schematic diagram illustrating the similarity between various video tags provided in an embodiment of this application; as shown... Figure 9 As shown, based on Figure 8 tags Figure 7-12 Emotional tags constructed Figure 7-1 3: Added positive sentiment label 9-1 and negative sentiment label 9-2, due to... and All of these belong to positive emotions, therefore, and All of them have a similarity of 1 to the positive sentiment label 9-1; It belongs to negative emotions, therefore, The similarity between this label and the negative sentiment label 9-2 is 1. It belongs to neutral sentiment, and there is no corresponding edge between it and the positive sentiment label 9-1 and the negative sentiment label 9-2.
[0224] Understandably, through tags Figure 7-12 In this update, new positive and negative sentiment tags have been added. High similarity scores have been established between tags belonging to positive sentiment and between tags belonging to negative sentiment, ensuring that sentiments are similar not only semantically but also emotionally. For example, in the tags... Figure 7-12 In this context, both the tags "love" and "breakup" are related to dating, resulting in a high degree of similarity between them. However, when distributing multimedia information, it is often placed on video frames with positive emotional tags. Here, by adding sentiment analysis, the obtained emotional tags... Figure 7-1 3. This widens the gap between the labels "dating" and "breakup".
[0225] S704, Construct the label embedding feature library 7-14 (label feature library).
[0226] It's important to note that since a video's content involves multiple dimensions, and typically each dimension corresponds to a single tag. For example, a short video might be tagged with "novel adaptation | mainland Chinese TV series | amusement park | missing | TV series trailer." Therefore, relying solely on videos with shared tags has limited effect on expanding the tags. Here, by combining second-order similarity between tags, we can uncover different tags that share the same semantic meaning for the same video scene. For instance, for the same type of video, the corresponding tags might be "celebrity fan-taken video," "fan-taken video," or "fan-taken video."
[0227] Here, graph embedding algorithms such as "LINE", "SDNE", or "GCN" can be used to base sentiment tags on... Figure 7-1 Using second-order similarity in 3 to obtain sentiment tags Figure 7-1 The embedding features (label features) of each label in step 3 are used to obtain the label embedding feature library 7-14.
[0228] S705, Save the label embedding feature library 7-14.
[0229] Here, the embedded tag feature library 7-14 can be saved to the hard drive.
[0230] The prediction phase 7-2 includes the following steps:
[0231] S706. Expand the keyword 7-21 (multimedia tag) with synonyms.
[0232] It should be noted that for the keyword 7-21 of the information to be pushed (multimedia), it is first expanded using synonyms in the thesaurus 7-22 (synonym tag model) to obtain expanded keywords 7-23 (initial tags to be matched); for example, the thesaurus 7-22 is used to obtain the synonym "amusement park" for the keyword "amusement park". The thesaurus can be obtained from books or documents, or from the internet; see [link to relevant documentation]. Figure 10 , Figure 10 This is an exemplary schematic diagram of obtaining a thesaurus provided in an embodiment of this application; as shown... Figure 10 As shown in the search results interface 10-1, for tag 10-11 (cucumber), the corresponding aliases 10-12 (cucumber, prickly cucumber, king cucumber, diligent cucumber, green cucumber, Tang cucumber, hanging cucumber) and tag 10-11 are grouped together as a set of synonyms.
[0233] S707, Expand the tags for extended keywords 7-23.
[0234] Here, the tag embedding feature library 7-14 stored on the hard disk is used to expand the tag of extended keyword 7-23. First, the feature representation of extended keyword 7-23 (feature to be expanded) is obtained; then, the cosine similarity (feature similarity) between the feature representation of extended keyword 7-23 and the embedding feature of each tag in the tag embedding feature library 7-14 is calculated using equation (2), and tags with a cosine similarity greater than the threshold (third similarity threshold) are selected as synonym tags (extended tags) of extended keyword 7-23; finally, keyword 7-24 including extended keyword 7-23 and synonym tags of extended keyword 7-23 is obtained.
[0235] S708. Perform frame extraction on video 7-25 (the video to be located).
[0236] Here, considering both computational complexity and recognition coverage, the video (video to be located) is processed at a rate of one frame per second (video frame extraction information, frame rate) to obtain a video frame sequence.
[0237] Additionally, when the playback duration of video 7-25 exceeds ten minutes (the playback duration threshold), the first 5 minutes and the last 5 minutes are skipped (i.e., the intro and outro are skipped) before frame extraction.
[0238] S709. Perform text recognition on video frame sequences.
[0239] Here, OCR technology can be used to identify text information such as subtitles in the video frame images corresponding to the video frame sequence, and obtain the video text information sequence (video frame text sequence).
[0240] S710, point (video frame to be pushed) recognition.
[0241] It should be noted that, based on the matching of video text information in the video text information sequence with keyword 7-24, the video frames corresponding to the successfully matched video text information are determined as points, thus obtaining point 7-25.
[0242] Here, the matching is based on the maximum matching principle, prioritizing the matching of the longest string of keywords (sub-tags to be matched) among keywords 7-24; for example, first matching the keyword "wedding ring", then matching the keyword "wedding".
[0243] S711, merged point 7-25.
[0244] Here, if the interval between two adjacent points in points 7-25 is less than two minutes (interval threshold), they are merged into one point. If the interval between the two adjacent points after merging is still less than two minutes, merging continues until the interval between the two adjacent points after merging is greater than or equal to two minutes, thus obtaining the point prediction result.
[0245] It should be noted that in prediction node 7-2, the process of obtaining keyword 7-24 can be implemented by the extension module, the process of obtaining video text information sequence can be implemented by the text extraction module, and the process of obtaining point prediction results can be implemented by the point prediction module.
[0246] See Figure 11 , Figure 11 This is an exemplary schematic diagram provided by an embodiment of the present application, illustrating the presentation of multimedia to be pushed during the playback of a video to be located; for example... Figure 11 As shown, video frame image 11-1 is the image corresponding to the video frame to be pushed. In video frame image 11-1, the word "roller coaster" in the subtitle 11-11 "I want to ride a roller coaster" is... Figure 7 The keyword "roller coaster" in keyword 7-24 is matched, and thus, at this time, the information to be pushed, representing the amusement park advertising video, can be presented.
[0247] See Figure 12 , Figure 12 This is another exemplary schematic diagram provided in this application embodiment of presenting multimedia to be pushed during the playback of the video to be located; such as Figure 12 As shown, video frame image 12-1 is the image corresponding to the video frame to be pushed. In video frame image 12-1, the word "supermarket" in subtitle 12-11 "The supermarket next to our house is having a big sale" matches the keyword "supermarket" in keyword 7-24. Therefore, at this time, the information to be pushed, representing the supermarket's promotional activities (XX Supermarket Anniversary Celebration), is presented on video frame image 12-1.
[0248] Understandably, using keywords to pinpoint locations within a video and then displaying the message at those locations is a common push notification method. This method, known as "Ruyi Tie," improves conversion rates due to the high relevance between the message and the video content. Furthermore, Ruyi Tie's method, which inserts the message at predetermined locations, allows for dynamic adjustments to the message itself, the push area, and the video location, offering high flexibility.
[0249] It is also understood that the video frame localization method provided in this application realizes accurate and automatic expansion of keywords for information to be pushed by constructing a tag embedding feature library based on video. In addition, when constructing the tag embedding feature library, the embedding features of the tags are determined by combining sentiment analysis and second-order similarity, making the obtained embedding features more useful for the accurate expansion of keywords. Furthermore, the construction efficiency of the tag embedding feature library is fast, real-time, and highly consistent with the current video scene. Moreover, when it is necessary to update the tag embedding feature library, it can be quickly updated based on the tags in the new video.
[0250] The following description continues to illustrate the exemplary structure of the video frame positioning device 255 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 3 As shown, the software module stored in the video frame positioning device 255 in the memory 250 may include:
[0251] The tag acquisition module 2551 is used to acquire the multimedia tag corresponding to the multimedia to be pushed, wherein the multimedia to be pushed is used to be presented during the playback of the video to be located.
[0252] The tag extension module 2552 is used to extend the multimedia tags based on the tag feature library to obtain the tags to be matched, wherein the tag feature library includes each tag feature corresponding to each video tag, and each tag feature is obtained based on the similarity between each video tag;
[0253] The tag positioning module 2553 is used to perform text matching in the video frame text sequence corresponding to the video to be located based on the tag to be matched, and to locate the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text.
[0254] In this embodiment, the video frame localization device 255 further includes a feature library acquisition module 2554, configured to acquire at least one sample video and at least one sample tag corresponding to each sample video; construct each video tag corresponding to at least one sample video based on the at least one sample tag corresponding to each sample video, wherein one video tag is one sample tag; acquire a set of sample videos corresponding to each video tag based on the at least one sample tag corresponding to each sample video in the at least one sample video; determine the similarity between each video tag based on the set of sample videos corresponding to each video tag; determine the tag features corresponding to each video tag based on the similarity between each video tag, and determine the determined tag features corresponding to each video tag as the tag feature library.
[0255] In this embodiment of the application, the feature library acquisition module 2554 is further configured to perform sentiment analysis on each video tag to obtain a sentiment category; construct a target sentiment tag and a negative sentiment tag, wherein the target sentiment tag includes at least one of a positive sentiment tag and a neutral sentiment tag; based on the sentiment category, determine the associated sentiment tag corresponding to each video tag from the target sentiment tag and the negative sentiment tag; and determine the sentiment similarity between each video tag and the associated sentiment tag.
[0256] In this embodiment of the application, the feature library acquisition module 2554 is further configured to determine the tag features corresponding to each video tag based on the similarity between each video tag and the sentiment similarity.
[0257] In this embodiment of the application, the feature library acquisition module 2554 is further configured to construct at least one sample tag set corresponding to each sample video based on at least one sample tag corresponding to each sample video; count the occurrence frequency of each sample tag in the sample tag set; filter the video tags whose occurrence frequency is greater than the frequency threshold from the sample tag set; and construct each video tag from the filtered video tags.
[0258] In this embodiment of the application, the feature library acquisition module 2554 is further configured to traverse each of the video tags, and obtain the similarity between the first traversed video tag and the second to the first video tags, where I is the number of video tags in each video tag; and perform the following processing through iteration i: obtain the similarity between the i-th traversed video tag and the (i+1)-first video tags, where, And i is a positive integer variable with increasing value; the similarity between the first video tag obtained by iteration i and the second to the first video tags, and the similarity between the (I-1)th video tag and the first video tag, are determined as the similarity between each video tag.
[0259] In this embodiment of the application, the feature library acquisition module 2554 is further configured to perform the following processing through iteration j: when the sample video set corresponding to the i-th video tag and the sample video set corresponding to the j-th video tag include common sample videos, the tag weights that are positively correlated with the number of common sample videos and negatively correlated with the number of sample videos corresponding to both the i-th and j-th video tags are determined as the similarity between the i-th video tag and the j-th video tag, wherein, , and j is a positive integer variable with increasing value; when there are no common sample videos between the sample video set corresponding to the i-th video tag and the sample video set corresponding to the j-th video tag, the minimum similarity threshold is determined as the similarity between the i-th video tag and the j-th video tag; obtain the similarity between the i-th video tag obtained by iteration j and the (i+1)-1-1 video tags respectively.
[0260] In this embodiment of the application, the feature library acquisition module 2554 is further configured to, based on the similarity between the various video tags, determine the video tags whose similarity to the first-order video tags corresponding to each video tag is greater than a first similarity threshold as first-order similar video tags corresponding to each video tag; based on the similarity between the various video tags, determine the video tags whose similarity to the first-order similar video tags is greater than a second similarity threshold as second-order similar video tags corresponding to each video tag as second-order similar video tags corresponding to each video tag; and determine the tag features corresponding to each video tag based on the first-order similar video tags and the second-order similar video tags.
[0261] In this embodiment of the application, the tag expansion module 2552 is further configured to extract features to be expanded from the multimedia tags; obtain the feature similarity between the features to be expanded and each tag feature in the tag feature library; among the obtained feature similarities between the features to be expanded and each tag feature, combine the video tags corresponding to the feature similarities greater than a third similarity threshold into an expanded tag; combine the expanded tag and the multimedia tag to obtain the tag to be matched.
[0262] In this embodiment of the application, the video frame positioning device 255 further includes a text recognition module 2555, which is used to extract a video frame sequence from the video to be located based on video frame extraction information, wherein the video frame extraction information includes at least one of frame rate, number of bullet comments and keyframes; and to extract video frame text information from the video frame image corresponding to each video frame in the video frame sequence to obtain the video frame text information sequence corresponding to the video frame sequence.
[0263] In this embodiment of the application, the text recognition module 2555 is further configured to obtain the playback duration corresponding to the video to be located; when the playback duration is greater than the playback duration threshold, video segments are extracted from the video to be located to obtain the video to be extracted.
[0264] In this embodiment of the application, the text recognition module 2555 is further configured to extract the video frame sequence from the video to be extracted based on the video frame extraction information.
[0265] In this embodiment of the application, the tag positioning module 2553 is further configured to obtain the character length corresponding to each sub-tag to be matched in the tag to be matched; sort the sub-tags to be matched in the tag to be matched based on the character length; and select the sub-tag to be matched with the longest character length from the sorted tags to be matched in turn for text matching in the video frame text sequence corresponding to the video to be located.
[0266] In this embodiment of the application, the video frame positioning device 255 further includes a video frame merging module 2556, which is used to obtain the frame interval between adjacent video frames in the video frame corresponding to the matched video frame text; merge adjacent video frames whose frame interval is less than the interval threshold; and determine the merged video frame as the video frame to be pushed in the video to be positioned for the multimedia to be pushed.
[0267] In this embodiment, the video frame positioning device 255 further includes an information push module 2557, which is used to play the video to be positioned; when the playback progress of the video to be positioned reaches the video frame to be pushed, the multimedia to be pushed is presented.
[0268] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. The processor of a video frame positioning device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the video frame positioning method described in this application.
[0269] This application provides a computer-readable storage medium storing executable instructions. When these executable instructions are executed by a processor, they cause the processor to execute the video frame positioning method provided in this application. For example, ... Figure 4 The video frame localization method is shown.
[0270] In some embodiments of this application, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a device that includes one or any combination of the above-mentioned memories.
[0271] In some embodiments of this application, executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0272] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0273] As an example, executable instructions can be deployed to execute on a single computer device (in this case, the single computer device is the video frame positioning device), or to execute on multiple computer devices located at one location (in this case, the multiple computer devices located at one location are the video frame positioning devices), or to execute on multiple computer devices distributed across multiple locations and interconnected via a communication network (in this case, the multiple computer devices distributed across multiple locations and interconnected via a communication network are the video frame positioning devices).
[0274] In summary, this application embodiment pre-collects various video tags and constructs a tag feature library using the tag features of each video tag obtained from the similarity between the various video tags. Therefore, after obtaining the multimedia tags corresponding to the multimedia to be pushed (the information to be pushed), the multimedia tags can be automatically expanded based on this tag feature library, and the video frame to be pushed can be located in the video to be located based on the expanded matching tags. Thus, the expansion of multimedia tags is automatic during the video frame location process. Therefore, the video frame location method provided by this application embodiment can improve the intelligence of video frame location. Furthermore, when pushing multimedia to be pushed based on the located video frame, the method provided by this application embodiment offers high flexibility and improves the accuracy of the push. Moreover, when constructing the tag feature library, combining sentiment analysis and second-order similarity video tags to determine tag features makes the obtained tag features more useful for the accurate expansion of multimedia tags. Additionally, the tag feature library is constructed quickly, has real-time performance, and is highly consistent with the current video scene. Furthermore, when the tag feature library needs to be updated, it can be quickly updated based on sample tags in new sample videos.
[0275] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A video frame localization method, characterized in that, include: Based on the similarity between video tags in the tag feature library, among the first-order video tags corresponding to each video tag, the video tags with a similarity greater than a first similarity threshold are determined as the first-order similar video tags corresponding to each video tag. The similarity is positively correlated with the number of common sample videos between the two video tags and negatively correlated with the number of sample videos corresponding to the two video tags respectively. In the second-order video tags corresponding to each video tag, the video tags whose similarity to the first-order similar video tags is greater than the second similarity threshold are determined as the second-order similar video tags corresponding to each video tag. Based on the sentiment category of each video tag, determine the associated sentiment tag corresponding to each video tag from positive sentiment tags, neutral sentiment tags, and negative sentiment tags; Determine the emotional similarity between each video tag and the associated emotional tag; The tag features corresponding to each video tag are determined based on the first-order similar video tags, the second-order similar video tags, the similarity between each video tag, and the sentiment similarity. Obtain the multimedia tag corresponding to the multimedia to be pushed, wherein the multimedia to be pushed is used to be presented during the playback of the video to be located; Based on the tag feature library, extended tags that are semantically similar to the multimedia tags are matched, wherein the tag feature library includes each tag feature corresponding to each video tag; Combine the extended tag and the multimedia tag to obtain the tag to be matched; Based on the character length of each sub-tag to be matched in the tags to be matched, the sub-tags to be matched in the tags to be matched are sorted. From the sorted tags to be matched, the sub-tag with the longest character length is selected in turn for text matching in the video frame text sequence corresponding to the video to be located. Based on the video frame corresponding to the matched video frame text, the video frame to be pushed of the multimedia to be pushed in the video to be located is located.
2. The method according to claim 1, characterized in that, Before matching extended tags semantically similar to multimedia tags based on the tag feature library, the method further includes: Obtain at least one sample video and at least one sample tag corresponding to each sample video; Based on at least one sample tag corresponding to each sample video, construct at least one video tag corresponding to each sample video, wherein one video tag is one sample tag; Based on at least one sample tag corresponding to each of the at least one sample video, obtain a set of sample videos corresponding to each video tag; Based on the sample video set corresponding to each video tag, the similarity between the various video tags is determined.
3. The method according to claim 2, characterized in that, After determining the similarity between video tags based on the sample video set corresponding to each video tag, the method further includes: Sentiment analysis is performed on each of the video tags to obtain the sentiment category; Construct target sentiment labels and negative sentiment labels, wherein the target sentiment labels include at least one of the positive sentiment labels and the neutral sentiment labels.
4. The method according to claim 2 or 3, characterized in that, The step of constructing at least one video tag corresponding to each of the sample videos based on at least one sample tag corresponding to each of the sample videos includes: Based on at least one sample tag corresponding to each sample video, construct at least one sample tag set corresponding to each sample video; Count the frequency of occurrence of each sample label in the sample label set; Filter the video tags that appear more frequently than a frequency threshold from the sample tag set; The selected video tags are then used to construct individual video tags.
5. The method according to claim 2 or 3, characterized in that, The step of determining the similarity between video tags based on the sample video set corresponding to each video tag includes: Iterate through each of the video tags and obtain the similarity between the first video tag and each of the second to the first video tags, where I is the number of video tags in each video tag; The following processing is performed by iterating through i: Obtain the similarity between the i-th video tag and each of the (i+1)-th to the ith video tags, where... And i is a positive integer variable whose value increases. The similarity between the first video tag obtained in iteration i and the second to the first video tags, and the similarity between the (I-1)th video tag and the first video tag, are determined as the similarity between each video tag.
6. The method according to claim 5, characterized in that, The step of obtaining the similarity between the i-th video tag and the (i+1)-1-1 video tags includes: The following processing is performed by iterating through j: When the sample video set corresponding to the i-th video tag and the sample video set corresponding to the j-th video tag include common sample videos, the tag weights that are positively correlated with the number of common sample videos and negatively correlated with the number of sample videos corresponding to both the i-th and j-th video tags are determined as the similarity between the i-th and j-th video tags. , and j is a positive integer variable whose value increases; When there are no common sample videos between the sample video set corresponding to the i-th video tag and the sample video set corresponding to the j-th video tag, the minimum similarity threshold is determined to be the similarity between the i-th video tag and the j-th video tag. Obtain the similarity between the i-th video tag obtained in iteration j and the (i+1)-1-1 video tags.
7. The method according to any one of claims 1 to 3, characterized in that, The method of matching extended tags that are semantically similar to multimedia tags based on the tag feature library includes: Extract the features to be expanded from the multimedia tags; Obtain the feature similarity between the feature to be expanded and each of the label features in the label feature library; Among the feature similarities obtained between the feature to be expanded and each of the tag features, the video tags corresponding to the feature similarities greater than the third similarity threshold are combined into expanded tags.
8. The method according to any one of claims 1 to 3, characterized in that, Before performing text matching on the video frame text sequence corresponding to the video to be located based on the tag to be matched, the method further includes: Based on video frame extraction information, a video frame sequence is extracted from the video to be located, wherein the video frame extraction information includes at least one of frame rate, number of bullet comments, and keyframes; Video frame text information is extracted from the video frame image corresponding to each video frame in the video frame sequence to obtain the video frame text information sequence corresponding to the video frame sequence.
9. The method according to claim 8, characterized in that, Before extracting the video frame sequence from the video to be located based on video frame extraction information, the method further includes: Obtain the playback duration of the video to be located; When the playback duration exceeds the playback duration threshold, video segments are extracted from the video to be located to obtain the video to be extracted. The step of extracting a video frame sequence from the video to be located based on video frame extraction information includes: Based on the video frame extraction information, the video frame sequence is extracted from the video to be extracted.
10. The method according to any one of claims 1 to 3, characterized in that, After locating the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text, the method further includes: Obtain the frame interval between adjacent video frames in the video frame corresponding to the matched video frame text; Adjacent video frames with a frame interval less than the interval threshold are merged, and the merged video frame is determined as the video frame to be pushed in the video to be located for the multimedia to be pushed.
11. The method according to any one of claims 1 to 3, characterized in that, After locating the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text, the method further includes: Play the video to be located; When the playback progress of the video to be located reaches the frame of the video to be pushed, the multimedia to be pushed is displayed.
12. A video frame positioning device, characterized in that, include: The tag acquisition module is used to determine the video tags with similarity greater than a first similarity threshold as first-order similar video tags corresponding to each video tag, based on the similarity between each video tag in the tag feature library. The similarity is positively correlated with the number of common sample videos between the two video tags and negatively correlated with the number of sample videos corresponding to the two video tags respectively. In the second-order video tags corresponding to each video tag, the video tags whose similarity to the first-order similar video tags is greater than the second similarity threshold are determined as the second-order similar video tags corresponding to each video tag. Based on the sentiment category of each video tag, determine the associated sentiment tag corresponding to each video tag from positive sentiment tags, neutral sentiment tags, and negative sentiment tags; Determine the emotional similarity between each video tag and the associated emotional tag; The tag features corresponding to each video tag are determined based on the first-order similar video tags, the second-order similar video tags, the similarity between each video tag, and the sentiment similarity. Obtain the multimedia tag corresponding to the multimedia to be pushed, wherein the multimedia to be pushed is used to be presented during the playback of the video to be located; A tag extension module is used to match extended tags that are semantically similar to the multimedia tags based on the tag feature library, wherein the tag feature library includes each tag feature corresponding to each video tag; Combine the extended tag and the multimedia tag to obtain the tag to be matched; The tag positioning module is used to sort the sub-tags in the tag to be matched based on the character length of each sub-tag to be matched in the tag to be matched, select the sub-tag with the longest character length from the sorted tags to be matched in turn to perform text matching in the video frame text sequence corresponding to the video to be located, and locate the video frame to be pushed in the video to be located based on the video frame corresponding to the matched video frame text.
13. The apparatus according to claim 12, characterized in that, The tag acquisition module is also used for: Before matching extended tags that are semantically similar to multimedia tags based on the tag feature library, at least one sample video and at least one sample tag corresponding to each sample video are obtained. Based on at least one sample tag corresponding to each sample video, construct at least one video tag corresponding to each sample video, wherein one video tag is one sample tag; Based on at least one sample tag corresponding to each of the at least one sample video, obtain a set of sample videos corresponding to each video tag; Based on the sample video set corresponding to each video tag, the similarity between the various video tags is determined.
14. An electronic device for video frame positioning, characterized in that, include: Memory, used to store executable instructions; A processor, when executing executable instructions stored in the memory, implements the video frame positioning method according to any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that, It stores executable instructions for implementing the video frame positioning method according to any one of claims 1 to 11 when executed by a processor.
16. A computer program product comprising executable instructions or a computer program, characterized in that, When the executable instructions or computer program are executed by a processor, they implement the video frame positioning method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Video tag extension method and device, computer equipment and storage medium
CN111368141A
Video playing method and device
CN111901668A