A video information acquisition method, device, apparatus and storage medium

By performing entity recognition on the initial text information of the video, obtaining entity information and generating supplementary text information, the problem of low video recall and relevance in the film and television library is solved, and more efficient video information acquisition and search are achieved.

CN115618060BActive Publication Date: 2026-03-31IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-17
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

The existing video library has low video recall and relevance, mainly due to the exaggerated names of user-uploaded videos, resulting in insufficient search recall and relevance.

Method used

Entity recognition is performed on the initial text information of the video to obtain entity information, and supplementary text information of the video is obtained using the entity information. The initial text information and supplementary text information are saved as related information in the video library for video search.

Benefits of technology

It improves the recall and relevance of videos in the film and television library, ensuring the accuracy and comprehensiveness of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115618060B_ABST
    Figure CN115618060B_ABST
Patent Text Reader

Abstract

The application discloses a video information acquisition method, device and equipment and a storage medium. The method comprises the following steps: acquiring video data; wherein the video data comprises a video and initial text information of the video; performing entity recognition on the initial text information to obtain entity information; obtaining supplementary text information of the video by using the entity information; taking the initial text information and the supplementary text information as associated information of the video; wherein the associated information is used for being saved into a movie and television library together with the video, and the initial text information and the supplementary text information are used for searching the video in the movie and television library. In the above manner, the recall rate and the relevance of the video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video information processing technology, and in particular to a video information acquisition method, apparatus, device, and storage medium. Background Technology

[0002] With the rise of video, more and more people are watching videos. Most existing video libraries enter information based on user-uploaded names, which are then stored in a storable database. However, many video resources use flashy and attention-grabbing names, and by simply entering information based on user-uploaded names, low recall and relevance rates occur when other users want to watch the desired video. Summary of the Invention

[0003] The main technical problem addressed by this application is to provide a video information acquisition method, apparatus, device, and storage medium that can improve the recall and relevance of video.

[0004] To address the aforementioned technical problems, the first aspect of this application provides a method for acquiring video information. This method includes: acquiring video data; wherein the video data includes a video and initial text information of the video; performing entity recognition on the initial text information to obtain entity information; using the entity information to obtain supplementary text information of the video; and using the initial text information and supplementary text information as associated information of the video; wherein the associated information is used to save the video along with the video to a video library, and the initial text information and supplementary text information are used to search for the video in the video library.

[0005] To address the aforementioned technical problems, a second aspect of this application provides a video information acquisition device, comprising: a first acquisition module for acquiring video data; wherein the video data includes a video and initial text information of the video; an entity recognition module for performing entity recognition on the initial text information to obtain entity information; a second acquisition module for obtaining supplementary text information of the video using the entity information; and a storage module for storing the initial text information and supplementary text information as associated information of the video; wherein the associated information is used to save the video together with the video into a video library, and the initial text information and supplementary text information are used to search for the video in the video library.

[0006] To address the aforementioned technical problems, a third aspect of this application provides a video information acquisition device, comprising: a memory and a processor coupled to each other, wherein the memory stores program instructions; and the processor executes the program instructions stored in the memory to implement the method described in the first aspect.

[0007] To address the aforementioned technical problems, a fourth aspect of this application provides a computer-readable storage medium for storing program instructions that can be executed to implement the method described in the first aspect.

[0008] The beneficial effects of this application are as follows: Unlike existing technologies, this application obtains the initial text information of a video, performs entity recognition on the initial text information to obtain entity information, and then uses the entity information to obtain supplementary text information of the video; the initial text information and supplementary text information are used as the video's associated information; wherein, the associated information is used to save the video along with the video to a video library, and the initial text information and supplementary text information are used to search for videos in the video library. By saving both the initial text information and supplementary text information in the video library, when a user searches for a video, the corresponding video can be returned based on the matching degree between the search information and the initial text information and supplementary text information. Therefore, compared to the scheme that only saves the initial text information to the video library, this application saves both the initial text information and supplementary text information to the video library, which can improve the video recall rate and relevance. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating the first embodiment of the video information acquisition method provided in this application;

[0010] Figure 2 This is a schematic diagram of one embodiment of the data mapping information table provided in this application;

[0011] Figure 3 This is a flowchart illustrating the second embodiment of the video information acquisition method provided in this application;

[0012] Figure 4 This is a flowchart illustrating the third embodiment of the video information acquisition method provided in this application;

[0013] Figure 5 This is a schematic diagram of one embodiment of the entity recognition model provided in this application;

[0014] Figure 6 This is a schematic diagram of the framework of one embodiment of the video information acquisition device provided in this application;

[0015] Figure 7 This is a schematic diagram of another embodiment of the video information acquisition device provided in this application;

[0016] Figure 8 This is a schematic diagram of the framework of one embodiment of the video information acquisition device provided in this application;

[0017] Figure 9 This is a schematic diagram of a framework of one embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0019] It should be noted that the embodiments of this application contain descriptions involving "first," "second," etc., which are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature.

[0020] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0021] Please see Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating the first embodiment of the video information acquisition method provided in this application. Figure 2 This is a schematic diagram of one embodiment of the data mapping information table provided in this application; the method includes:

[0022] S11: Acquire video data.

[0023] In one embodiment, the video data includes the video and its initial text information. The initial text information may be information obtained by the user who uploaded the video summarizing key information about the video. For example, the initial text information may include at least one of the following: video name, characters featured in the video, video category, video description, video release date, and video rating. Understandably, the initial text information may also include other information, which is not specifically limited here. The initial text information may also be information obtained by the administrator of the platform to which the video is to be uploaded, summarizing key information about the video. For example, if the video is to be uploaded to the Tencent Video platform, the administrator of the Tencent Video platform may summarize the key information to obtain the initial text information. The video may be a short video.

[0024] S12: Perform entity recognition on the initial text information to obtain entity information.

[0025] In one embodiment, an entity recognition model can be used to perform entity recognition on the initial text information to obtain entity information. The entity recognition model can consist of three layers: the first layer encodes the initial text information to obtain a first encoded feature, which can be the word segmentation features of several words contained in the initial text information; the second layer extracts features from the first encoded feature to obtain a second encoded feature, which can be the textual semantic features of the initial text information; the third layer uses the second encoded feature to perform entity prediction to obtain entity information. The entity information can include an entity name and an entity tag. The entity tag is used to represent the category of the entity and can include place name tags, person name tags, movie tags, etc. In other embodiments, other methods (such as algorithms) can also be used to perform entity recognition on the initial text information to obtain entity information; no specific limitations are made here.

[0026] In one embodiment, entity recognition can be performed directly on the initial text information. In other embodiments, to avoid low accuracy of entity recognition results due to anomalous information in the initial text information, the initial text information can be cleaned and filtered first to remove anomalous information, resulting in filtered initial text information; then, entity recognition is performed on the filtered initial text information. Anomalous information may include at least one of spaces or random symbols. Understandably, anomalous information may also include other information, which is not specifically limited here.

[0027] S13: Obtain supplementary text information for the video using entity information.

[0028] In one embodiment, the entity information includes an entity name, and supplementary text information matching the entity name can be retrieved from a knowledge base based on the entity name. The knowledge base can be pre-built by the user and stores at least one entity name and at least one supplementary text information matching the entity name. In one specific embodiment, the supplementary text information includes at least one of the following: video alias, video tags, people featured in the video, video category, and video description. Understandably, in other embodiments, the supplementary text information may also include other information, such as the video's update time, whether the video is free, and the video's popularity score, etc., which are not specifically limited here.

[0029] Obtaining supplementary text information aims to improve video recall; the more complete the supplementary and initial text information, the higher the recall rate. However, there are also drawbacks, such as the impact of obtaining incorrect supplementary text information on the final result. To avoid this problem, after obtaining the supplementary text information, the correlation between the supplementary and initial text information can be calculated. By controlling the correlation, the issue of low recalled videos and low relevance caused by incorrect supplementary information can be addressed.

[0030] S14: Use the initial text information and supplementary text information as the video's associated information; wherein, the associated information is used to save the video together with the video to the video library, and the initial text information and supplementary text information are used to search for the video in the video library.

[0031] In one embodiment, the video, its initial text information, and its supplementary text information can be directly stored in a video library. In other embodiments, the initial text information and supplementary text information of the video can be used to generate data mapping information for the video. That is, the data mapping information includes the initial text information and supplementary text information of the video; understandably, the data mapping information may include some or all of the initial text information and supplementary text information of the video. In a specific embodiment, the data mapping information is represented in the form of data mapping information, such as... Figure 2 As shown, the data mapping information includes key fields in the initial text information and supplementary text information, such as video name, video alias, video category, and video popularity value.

[0032] When saving initial text information, supplementary text information, and videos to the video library, the initial text information and supplementary text information can be isolated. When a user searches for a video, their search information can be compared with the initial text information and supplementary text information. If the user's search information matches at least one of the initial text information or supplementary text information, the corresponding video is returned. For example, if the initial text information of a video contains "Princess Pearl" and the supplementary text information contains "Little Swallow," and the user's search information is "Little Swallow," then that video is returned. This method can improve the video recall rate. Furthermore, since the initial text information is manually entered by the user, errors may occur, while the supplementary text information generally does not contain errors. Therefore, video searches based on supplementary text information can improve the accuracy of video recall.

[0033] In this embodiment, after obtaining the initial text information of the video, entity recognition is performed on the initial text information to obtain entity information, and then the entity information is used to obtain supplementary text information of the video. The initial text information and the supplementary text information are used as the video's association information. The association information is used to save the video along with the video to a video library, while the initial text information and the supplementary text information are used to search for the video in the video library. By saving both the initial text information and the supplementary text information in the video library, when a user searches for a video, the corresponding video can be returned based on the matching degree between the search information and the initial text information and the supplementary text information. Therefore, compared to the prior art scheme that only saves the initial text information to the video library, this application saves both the initial text information and the supplementary text information to the video library, which can improve the video recall rate and relevance.

[0034] Please see Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the video information acquisition method provided in this application. The method includes:

[0035] S31: Acquire video data.

[0036] The video data includes the video itself and its initial text information.

[0037] S32: Perform entity recognition on the initial text information to obtain entity information.

[0038] S33: Obtain the target popularity value of the video in the target source.

[0039] The target information source is the information source to which the video needs to be uploaded. For example, if the video needs to be uploaded to Tencent, then the target information source is Tencent. In one embodiment, a reference popularity value of the video is obtained from at least one other information source. Based on the reference popularity value and the correlation between the popularity values ​​of the target information source and at least one other information source, the target popularity value of the video in the target information source is obtained. That is, before the video is uploaded to the target information source, it has already been uploaded to other information sources and has generated a reference popularity value in those other information sources. The reference popularity value can be determined based on information such as the number of clicks on the video in other information sources, which is not limited here. The correlation between the popularity values ​​of the target information source and at least one other information source can be predetermined. Specifically, a first popularity value of several sample videos in at least one other information source and a second popularity value of several sample videos in the target information source can be obtained, wherein the sample videos in at least one other information source and the sample videos in the target information source are the same. The correlation between the popularity values ​​of the target information source and at least one other information source is obtained by linear fitting using the first popularity value and the second popularity value. In one specific embodiment, the correlation between the target source and at least one other source in terms of popularity values ​​is linear. After obtaining the reference popularity value and the correlation between the target source and at least one other source, the target popularity value of the video in the target source can be obtained. In this embodiment, the target popularity value can be updated after being stored in the video library. Specifically, the updated reference popularity value of the video can be obtained from at least one other source at preset intervals. Based on the updated reference popularity value and the correlation between popularity values, the updated target popularity value is obtained.

[0040] In another embodiment, a preset popularity value can be directly obtained as the target popularity value of the video on the target source. The preset popularity value is set by the user. For example, if the video is uploaded to the target source for the first time and there is no video identical to the video in the target source's database, the preset popularity value can be set to 0. In this embodiment, after the target popularity value is stored in the video library, it can be updated. Specifically, the number of clicks on the video on the target source is counted at preset intervals, and the target popularity value is updated using the number of clicks. In one specific embodiment, the value obtained by reducing the number of clicks by a preset factor can be used as the updated target popularity value. The preset factor is set by the user, for example, 1000 times. Furthermore, the Spark computing engine can be used to count the number of clicks on the video on the target source. Spark is a fast and general-purpose computing engine designed for large-scale data processing. Spark has the advantages of Hadoop MapReduce (a programming framework for distributed computing programs). Spark's output can be stored in memory, thus eliminating the need to read and write to HDFS (Hadoop Distributed File System). Therefore, Spark's performance and computing speed are higher than MapReduce.

[0041] S34: Obtain supplementary text information for the video using entity information.

[0042] S35: Use the initial text information, supplementary text information, and target popularity value as the video's associated information. The associated information is used to save the video along with the video to the video library, while the initial text information and supplementary text information are used to search for the video in the video library.

[0043] In one implementation, when a user searches for videos in a video library, the system can return corresponding videos based on the relevance of the user's search information with the initial and supplementary text information. When returning videos, if the user sets the number of videos to return, videos with higher target popularity values ​​can be returned first, based on the target popularity value.

[0044] In this embodiment, the execution order of steps S33 and S34 can be changed. That is, step S33 can be executed first, followed by step S34; or step S34 can be executed first, followed by step S33.

[0045] The specific implementation methods of steps S31-S32 and S34-S35 can be referred to the relevant description of the first implementation method of the video information acquisition method, and will not be repeated here.

[0046] Please see Figure 4-5 , Figure 4 This is a flowchart illustrating the third embodiment of the video information acquisition method provided in this application. Figure 5 This is a schematic diagram of an implementation method of the entity recognition model provided in this application; the method includes:

[0047] S41: Acquire video data.

[0048] The video data includes the video itself and its initial text information.

[0049] S42: Encode the initial text information to obtain the first encoded feature of the initial text information.

[0050] S43: Extract features from the first coding feature to obtain the second coding feature.

[0051] S44: Entity prediction is performed using the second coding feature to obtain entity information.

[0052] In one embodiment, steps S42-S44 are performed by an entity recognition model. The entity recognition model comprises three layers: a first feature extraction layer, a second feature extraction layer, and an entity prediction layer. In this embodiment, the initial text information is encoded in the first feature extraction layer to obtain the first encoded features of the initial text information. The first encoded features are word segmentation features of several words. The first feature extraction layer can be a BERT (Bidirectional Enoceder Representations from Transformers) model. The BERT model enhances the generalization ability of the word vector model by capturing the semantic relationships between characters, words, and sentences, and can perceive multi-granular semantic relationships. The BERT model achieves contextual word vector representation through encoding via transformation (Transformer layer). BERT can utilize information from both forward and backward directions simultaneously. The Transformer encoder is the main part of the BERT model, through which encoding is performed. Specifically, as shown... Figure 5 As shown, the BERT model processes the initial text information (i.e. Figure 5 The BERT model encodes the initial text data to obtain the first encoded feature ce of the initial text information. The encoding of the initial text information by the BERT model includes: dividing the initial text information into several short sentences, for example, using punctuation marks such as commas and periods as dividing points; then using word segmentation tools or single-character methods to segment the short sentences into several words; for each word, obtaining the position information of each word based on its natural word order in the initial text information; and using the word segments and their position information to extract features, obtaining the word segmentation features of the word segments, which include the text features of the context of the text in which the word segments are located.

[0053] In the second feature extraction layer, the first encoded features are extracted to obtain the second encoded features. The second encoded features can be the textual semantic features of the initial text information. For example... Figure 5 As shown, the second feature extraction layer can be a BiLSTM neural network layer. The BiLSTM neural network layer contains a two-layer LSTM (Long Short Term Memory) network. Several word segmentation features are input into the BiLSTM neural network layer. The forward and backward LSTM outputs a hidden state sequence H = (h1, h2, ..., hi, ..., hn). The matrix formed by concatenating the hidden state sequences H = (h1, h2, ..., hi, ..., hn) according to their positions is the second encoded feature. The BiLSTM neural network layer not only has the characteristic of capturing long sentence relationships but can also capture contextual relationships to extract text semantic features. Because this neural network is formed by combining forward and backward propagation, it can capture the semantics of sentences from two directions, overcoming the limitations of unidirectional LSTM.

[0054] Entity prediction is performed using the second encoded feature in the entity prediction layer to obtain entity information, which includes an entity name and an entity label. As Figure 5 shown, the entity prediction layer can be a CRF layer. The CRF layer annotates the second encoded feature transmitted by the BiLSTM neural network layer and finally outputs the corresponding entity name and entity label. The entity label can be a user-defined label, which can include a place name label, a person name label, a movie label, etc. In a specific embodiment, the BIO rule is used to annotate the entity words in the text. The BIO rule performs full-text annotation on the training text data, where "O" represents non-entity characters in the training text, "B" represents the starting character of the entity in the training text, and "I" represents the middle and ending characters of the entity in the training text. Generally, a custom entity category follows the "B" and "I" annotations, such as "LOC", "VID", "PER", "ORG", etc. "LOC" represents a place name entity, "VID" represents a movie entity, "PER" represents a person name entity, and "ORG" represents an organization or company. For example, if the entity name is "Maldives", then "B-LOC" can be annotated above the character "Ma", and "I-LOC" can be annotated above the characters "er", "dai", and "fu" respectively. It can be understood that in other embodiments, other annotation methods can also be adopted, which are not specifically limited herein.

[0055] The principle of CRF layer annotation is roughly as follows: Let X=(x1,x2,…,xi,…,xn) and Y=(y1,y2,…,yi,…,ym) be joint random variables. If the random variable Y forms a Markov random field (MRF), then the conditional probability distribution P(Y|X) is called a conditional random field. X is called the input variable, the observation sequence, and Y is called the output sequence, the label sequence, the state sequence. Part of the entire sequence is determined by the second encoded feature output by the BiLSTM, and the other part is determined by the transition matrix of the CRF. Finally, through function calculation, the corresponding optimal sequence result is obtained.

[0056] In other embodiments, the entity recognition model can also be other models, such as a model composed of CNN (Convolutional Neural Network) and CRF, which is not specifically limited herein.

[0057] S45: Obtain supplementary text information of the video using the entity information.

[0058] S46: Use the initial text information and the supplementary text information as the associated information of the video; where the associated information is used to be saved in the video library together with the video, and the initial text information and the supplementary text information are used to search for the video in the video library.

[0059] The specific implementation of steps S41 and S45-S46 in this embodiment can be found in the relevant description of the first embodiment of the video information acquisition method, and will not be repeated here.

[0060] In this embodiment, before executing steps S42-S44 using the entity recognition model, the entity recognition model can be trained. Specifically, a large amount of sample text information is acquired, and the entity information in the training sample text information is labeled using the BIO rule to obtain labeled entity information. The large amount of sample text information is input into the entity recognition model, enabling the entity recognition model to encode, extract features, and predict entities from the sample text information to obtain sample entity information. The sample entity information may include sample entity names and sample entity labels. Based on the sample entity information and the labeled entity information, the parameters of the entity recognition model are adjusted to improve the accuracy of the entity recognition model.

[0061] Please see Figure 6 , Figure 6 This is a schematic diagram of a framework of an embodiment of the video information acquisition device provided in this application. The video information acquisition device 60 includes: a first acquisition module 61, an entity recognition module 62, a second acquisition module 63, and a storage module 64. The first acquisition module 61 is used to acquire video data, wherein the video data includes the video and the initial text information of the video. The entity recognition module 62 is used to perform entity recognition on the initial text information to obtain entity information. The second acquisition module 63 is used to obtain supplementary text information of the video using the entity information. The storage module 64 is used to use the initial text information and the supplementary text information as associated information of the video. The associated information is used to save the video together with the video to a video library, and the initial text information and the supplementary text information are used to search for the video in the video library.

[0062] In one embodiment, the supplementary text information includes at least one of the following: video alias, video tag, people featured in the video, video category, and video description; and / or, the entity information includes the entity name. Obtaining the supplementary text information of the video using the entity information includes: retrieving supplementary text information that matches the entity name from a knowledge base.

[0063] In one embodiment, after performing entity recognition on the initial text information to obtain entity information, the method further includes: obtaining the target popularity value of the video in the target information source; and using the initial text information and supplementary text information as the video's associated information, including: using the initial text information, supplementary text information, and target popularity value as the video's associated information.

[0064] In one embodiment, obtaining the target popularity value of a video in a target source includes: obtaining a reference popularity value of the video from at least one other source; and obtaining the target popularity value of the video in the target source based on the reference popularity value and the popularity value correlation between the target source and at least one other source.

[0065] In one embodiment, the correlation of popularity values ​​is linear; before obtaining the target popularity value of the video in the target source based on the reference popularity value and the correlation of popularity values ​​between the target source and at least one other source, the method further includes: obtaining a first popularity value of several sample videos in at least one other source and a second popularity value of several sample videos in the target source; and performing linear fitting using the first popularity value and the second popularity value to obtain the correlation of popularity values ​​between the target source and at least one other source.

[0066] In one embodiment, obtaining the target popularity value of a video in a target information source includes: obtaining a preset popularity value as the target popularity value of the video in the target information source; after saving the video's association information to the video library, it further includes: counting the number of clicks on the video in the target information source at preset intervals; and updating the target popularity value using the number of clicks.

[0067] In one embodiment, after acquiring video data, the method further includes: filtering out abnormal information in the initial text information to obtain filtered initial text information; the abnormal information includes at least one of space symbols and random symbols; performing entity recognition on the initial text information to obtain entity information, including: performing entity recognition on the filtered initial text information to obtain entity information.

[0068] In one embodiment, entity recognition is performed on the initial text information to obtain entity information, including: encoding the initial text information to obtain a first encoded feature of the initial text information; extracting features from the first encoded feature to obtain a second encoded feature; and using the second encoded feature to perform entity prediction to obtain entity information.

[0069] In one embodiment, encoding the initial text information to obtain the first encoded feature of the initial text information includes: dividing the initial text information to obtain several words and the position information of several words; based on the several words and the position information of several words, performing feature extraction using a first feature extraction layer to obtain the word segmentation features of several words; the word segmentation features of several words include the text features of the context of the text in which the several words are located.

[0070] And / or, perform feature extraction on the encoded features to obtain second encoded features, including: using the second feature extraction layer to process the word segmentation features of several words to obtain the text semantic features of the initial text information;

[0071] And / or, using the second encoding features to perform entity prediction to obtain entity information, including: based on text semantic features, using the entity prediction layer to perform entity prediction to obtain entity information contained in the initial text, the entity information including entity name and entity label.

[0072] In one embodiment, after using the initial text information and supplementary text information as the video's associated information, the method further includes: generating data mapping information for the video, the data mapping information containing the video's associated information; and saving the video and the data mapping information to a video library.

[0073] Please see Figure 7 , Figure 7 This is a schematic diagram of another embodiment of the video information acquisition device provided in this application. The second video information acquisition device 70 includes a source data layer 71, an information extraction layer 72, an information processing layer 73, and an information mapping layer 74. The source data layer 71 is used to acquire video data uploaded by users. The video data includes video and initial text information. Users can selectively save the video data to the original database. The information extraction layer 72 is used to perform entity recognition on the initial text information to obtain entity information. The information processing layer 73 is used to obtain supplementary text information of the video using the entity information. The supplementary text information may include information such as the video's alias and tags. It can also be used to perform data cleaning on the initial text information to filter out abnormal information. It is also used to obtain the target popularity value of the video in the target information source, for example, by obtaining the target popularity value through the click volume of existing online entities. The information mapping layer 74 is used to generate data mapping information of the video. The data mapping information includes the video's association information.

[0074] For details on the specific implementation of the steps performed by the source data layer 71, information extraction layer 72, information processing layer 73, and information mapping layer 74, please refer to the relevant steps of any of the above video information acquisition method embodiments, which will not be repeated here.

[0075] Please see Figure 8 , Figure 8 This is a schematic diagram of a framework of one embodiment of the video information acquisition device provided in this application.

[0076] The video information acquisition device 80 includes a memory 81 and a processor 82 coupled to each other. The memory 81 stores program instructions, and the processor 82 executes the program instructions to implement the steps in any of the above method embodiments. Specifically, the video information acquisition device 80 may include, but is not limited to, desktop computers, laptops, servers, mobile phones, tablets, etc., and is not limited thereto.

[0077] Specifically, processor 82 controls itself and memory 81 to implement the steps in any of the above method embodiments. Processor 82 can also be referred to as a CPU (Central Processing Unit). Processor 82 may be an integrated circuit chip with signal processing capabilities. Processor 82 can also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. Furthermore, processor 82 can be implemented using integrated circuit chips.

[0078] Please see Figure 9 , Figure 9 This is a schematic diagram of a computer-readable storage medium according to one embodiment of the present application. The computer-readable storage medium 90 stores program instructions 91, which, when executed by a processor, are used to implement the steps in any of the above method embodiments.

[0079] The computer-readable storage medium 90 can specifically be a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or a medium that can store computer programs. Alternatively, it can be a server that stores the computer program, which can send the stored computer program to other devices for execution or can also run the stored computer program itself.

[0080] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0081] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0082] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0083] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0084] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A video information acquisition method characterized by comprising: The method comprises: acquiring video data; wherein the video data comprises a video and initial text information of the video; performing entity recognition on the initial text information to obtain entity information; acquiring a reference hotness value of the video from at least one other source; obtaining a target hotness value of the video in a target source based on the reference hotness value and a hotness value correlation relationship between the target source and the at least one other source; obtaining supplementary text information of the video by using the entity information; taking the initial text information, the supplementary text information and the target hotness value as associated information of the video; wherein the associated information is used to be saved into a movie and television library together with the video, and the initial text information and the supplementary text information are used for searching the video in the movie and television library.

2. The method of claim 1, wherein, The supplementary text information comprises at least one of a video alias, a video tag, a person included in a video, a video category and a video introduction. And / or, the entity information comprises an entity name, and the obtaining of the supplementary text information of the video by using the entity information comprises: acquiring supplementary text information matched with the entity name from a knowledge base.

3. The method of claim 1, wherein, The hotness value correlation relationship is a linear relationship. Before the obtaining of the target hotness value of the video in the target source based on the reference hotness value and the hotness value correlation relationship between the target source and the at least one other source, the method further comprises: acquiring first hotness values of a plurality of sample videos in the at least one other source and second hotness values of the plurality of sample videos in the target source; performing linear fitting by using the first hotness values and the second hotness values to obtain the hotness value correlation relationship between the target source and the at least one other source.

4. The method of claim 1, wherein, After the saving of the associated information of the video into the movie and television library, the method further comprises: acquiring an updated reference hotness update value from at least one other source at intervals of a preset time; updating the target hotness value based on the reference hotness update value and the hotness value correlation relationship.

5. The method of claim 1, wherein, After the acquiring of the video data, the method further comprises: filtering abnormal information in the initial text information to obtain filtered initial text information; the abnormal information comprises at least one of a space symbol and a messy symbol; The performing of entity recognition on the initial text information to obtain entity information comprises: performing entity recognition on the filtered initial text information to obtain the entity information.

6. The method of claim 1, wherein, The performing of entity recognition on the initial text information to obtain entity information comprises: performing encoding on the initial text information to obtain first encoding features of the initial text information; performing feature extraction on the first encoding features to obtain second encoding features; performing entity prediction by using the second encoding features to obtain the entity information.

7. The method of claim 6, wherein, The performing of encoding on the initial text information to obtain first encoding features of the initial text information comprises: dividing the initial text information to obtain a plurality of words and position information of the plurality of words; Based on the plurality of words and the position information of the plurality of words, feature extraction is performed by using a first feature extraction layer to obtain word features of the plurality of words; the word features of the plurality of words include text features of a context of a text in which the plurality of words are located; And / or, the feature extraction on the encoded features to obtain second encoded features includes: The word features of the plurality of words are processed by using a second feature extraction layer to obtain text semantic features of the initial text information; And / or, the entity prediction by using the second encoded features to obtain the entity information includes: Based on the text semantic features, entity prediction is performed by using an entity prediction layer to obtain the entity information contained in the initial text, the entity information including an entity name and an entity label.

8. The method of claim 1, wherein, After the initial text information, the supplementary text information, and the target hotness value are taken as the association information of the video, the method further includes: Generating data mapping information of the video, the data mapping information including the association information of the video; Saving the video and the data mapping information to the film and television library.

9. A video information acquisition apparatus characterized by comprising: The apparatus includes: A first obtaining module configured to obtain video data; the video data including a video and initial text information of the video; An entity recognition module configured to perform entity recognition on the initial text information to obtain entity information; after the entity information is obtained, the apparatus is further configured to obtain a reference hotness value of the video from at least one other source; based on the reference hotness value and a hotness value association relationship between a target source and the at least one other source, a target hotness value of the video in the target source is obtained; A second obtaining module configured to obtain supplementary text information of the video by using the entity information; A storage module configured to take the initial text information, the supplementary text information, and the target hotness value as association information of the video; the association information is used to be saved to a film and television library together with the video, and the initial text information and the supplementary text information are used for searching the video in the film and television library.

10. A video information acquisition apparatus characterized by comprising: The device includes a memory and a processor coupled to each other; The memory stores program instructions; The processor is configured to execute the program instructions stored in the memory to implement the method of any one of claims 1-8.

11. A computer readable storage medium, characterized in that, The computer-readable storage medium is configured to store program instructions that can be executed to implement the method of any one of claims 1-8. The computer-readable storage medium is configured to store program instructions that can be executed to implement the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Information pushing method, device and equipment based on video content

    CN110418193A

  • Label extension method and device, electronic equipment and storage medium

    CN115033738A