Content recommendation method and device based on OTT industry natural language analysis

By using natural language analysis technology in the OTT industry to extract video tags and match them with user preference models, the problem of difficult personalized content recommendation in the existing technology is solved, and efficient and accurate content recommendation services are achieved.

CN120104833APending Publication Date: 2025-06-06SHENZHEN SDMC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182169.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the OTT industry, it is difficult for the existing technology to achieve personalized content recommendations, especially the problem of large operational workloads and high cost of manual identification tags under rich content.

Method used

By obtaining the target video, obtaining the target tags in text, and matching these tags with the built user preference model to generate a recommendation list. The specific steps include using the Google Cloud Natural Language API to analyze the video text and building a deep learning model of multi-label scene loss function to train the user preference model.

Benefits of technology

It has achieved more accurate and personalized content recommendation services to users, reducing operational workload and reducing the cost of manual identification of tags, ensuring the real-time and accuracy of content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104833A_ABST
    Figure CN120104833A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a content recommendation method, device and equipment based on OTT industry natural language analysis and a computer readable storage medium. The method comprises the following steps: acquiring a target video; the target video comprises an on-demand video and / or a live event video; performing text extraction on the target video to obtain a target label; and matching the target tag with a constructed user preference model to generate a recommendation list. In this way, more accurate and personalized content recommendation service can be provided for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of content push, and in particular, to a content recommendation method, apparatus, device and computer-readable storage medium based on natural language analysis in the OTT industry. Background Art

[0002] With the rapid development of Internet technology, the OTT (Over The Top) industry has gradually emerged. OTT refers to providing users with various application services such as videos, music, games, etc. through the Internet without relying on traditional broadcast transmission channels such as cable TV and satellite TV.

[0003] In the OTT industry, a common business scenario is personalized recommendation of on-demand and event live broadcasts. A common solution is to manually tag on-demand and event live broadcasts when injecting content, and then recommend content based on the user's viewing behavior. However, rich content will lead to a large workload for operations, and the cost of manual tag identification is also relatively high. Summary of the invention

[0004] According to an embodiment of the present application, a content recommendation solution based on natural language analysis in the OTT industry is provided, which can provide users with more accurate and personalized content recommendation services.

[0005] In a first aspect of the present application, a content recommendation method based on natural language analysis in the OTT industry is provided. The method comprises: Acquire a target video; the target video includes a video on demand and / or a live event video; Performing text extraction on the target video to obtain a target label; The target tag is matched with the constructed user preference model to generate a recommendation list.

[0006] Furthermore, extracting text from the target video to obtain a target tag includes: If the target video is a live event video, recording a video on demand corresponding to the live event video according to the type of the current device and the start time and end time of the live event video; Extracting text from the on-demand video to obtain video text; The video text is analyzed through the Google Cloud Natural Language API to obtain a target label.

[0007] Furthermore, the target tags include sentiment analysis tags, entity recognition tags and / or topic extraction tags.

[0008] Furthermore, it also includes: The user preference model can be constructed in the following way: Obtain a video data sample set; Extracting text from the video in the video data sample set to obtain video text; Converting the video text into a semantic vector; The semantic vector is taken as input, the probability values ​​of all event labels are taken as output, a multi-label scene loss function is constructed through the first deep learning model, and the second deep model is used as the vector feature extraction representation to train the user preference model.

[0009] Furthermore, the user preference model may adopt the following loss function:

[0010] Wherein, N is a set of negative samples; P is a positive sample set; Said is the score of the proportion of category i in the positive samples; Said is the score of the proportion of category j in negative samples.

[0011] In a second aspect of the present application, a content recommendation device based on natural language analysis in the OTT industry is provided. The device includes: An acquisition module, used for acquiring a target video; the target video includes a video on demand and / or a live event video; An extraction module, used to extract text from the target video to obtain a target label; The matching module is used to match the target tag with the constructed user preference model to generate a recommendation list.

[0012] In a third aspect of the present application, an electronic device is provided, which includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method described above is implemented.

[0013] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to the first aspect of the present application is implemented.

[0014] The content recommendation method based on natural language analysis in the OTT industry provided in the embodiment of the present application obtains a target video; the target video includes a video on demand and / or a live event video; text is extracted from the target video to obtain a target tag; the target tag is matched with a constructed user preference model to generate a recommendation list, thereby providing users with a more accurate and personalized content recommendation service.

[0015] It should be understood that the contents described in the Summary of the Invention are not intended to limit the key or important features of the embodiments of the present application, nor are they intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The above and other features, advantages and aspects of the embodiments of the present application will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein: Figure 1 A schematic diagram showing an exemplary operating environment in which embodiments of the present disclosure can be implemented; Figure 2 A flowchart of a content recommendation method based on natural language analysis in the OTT industry according to an embodiment of the present application; Figure 3 The example code for extracting video text according to an embodiment of the present application; Figure 4 Example code for performing Google Cloud Natural Language API analysis on text according to an embodiment of the present application; Figure 5 This is an example of the result returned by Entities according to an embodiment of the present application; Figure 6 A block diagram of a content recommendation device based on natural language analysis in the OTT industry according to an embodiment of the present application; Figure 7 It is a schematic diagram of the structure of a terminal device or a server suitable for implementing the embodiments of the present application. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0018] In addition, the term "and / or" in this article is only a description of the association relationship between the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0019] Figure 1 An exemplary system architecture 100 is shown to which an embodiment of the content recommendation method based on natural language analysis in the OTT industry or the content recommendation device based on natural language analysis in the OTT industry of the present application can be applied.

[0020] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0021] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as model training applications, video recognition applications, web browser applications, social platform software, etc.

[0022] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Compression Standard Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Compression Standard Audio Layer 4) players, laptop computers and desktop computers, etc. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, multiple software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0023] When the terminals 101, 102, and 103 are hardware, a video acquisition device may be installed thereon. The video acquisition device may be any device capable of acquiring video, such as a camera, a sensor, etc. A user may use the video acquisition device on the terminals 101, 102, and 103 to acquire video.

[0024] The server 105 may be a server that provides various services, such as a background server that processes data displayed on the terminal devices 101, 102, and 103. The background server may analyze the received data and feed back the processing result (such as the matching result) to the terminal device.

[0025] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or it can be implemented as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, multiple software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.

[0026] It should be understood that Figure 1 The number of terminal devices, networks and servers in the above is only illustrative. According to the implementation requirements, there may be any number of terminal devices, networks and servers. In particular, when the target data does not need to be acquired remotely, the above system architecture may not include a network, but only include terminal devices or servers.

[0027] like Figure 2 , which is a flow chart of a content recommendation method based on natural language analysis in the OTT industry according to an embodiment of the present application. Figure 2 It can be seen from the figure that the training method of the multi-label classification model of this embodiment includes the following steps: S110, obtaining a target video.

[0028] In this embodiment, the execution subject (eg Figure 1 The server shown) can obtain the government data sample set through wired or wireless connection.

[0029] Furthermore, the execution subject may obtain an electronic device (such as Figure 1 The target video may be a target video sent by the terminal device as shown in the figure, or may be a target video pre-stored locally.

[0030] The target video includes on-demand video and / or live event video.

[0031] S120, extracting text from the target video to obtain a target tag.

[0032] In some embodiments, for different mobile devices, playback links of Dash and HLS protocols can be configured separately: HLS protocol: https: / / example.com / video.m3u8 Dash Protocol: https: / / example.com / video.mpd Preferably, iOS devices can use the HLS protocol, and Android devices can use the Dash protocol.

[0033] In some embodiments, the live broadcast of the event is real-time, so in the present disclosure, the live broadcast of the event can be first recorded as a single on-demand file (on-demand video) and then analyzed.

[0034] Specifically, record the video according to the start and end time of the event live broadcast: For HLS: ffmpeg -i "https: / / example.com / video.m3u8" -c copyoutput.mp4; For DASH: ffmpeg -i "https: / / example.com / video.mpd" -c copyoutput.mp4.

[0035] Through Figure 3 The method shown in FIG. 1 extracts the on-demand video into video text. Figure 4 As shown, the extracted text is sent to the Google Cloud Natural Language API for analysis, and sentiment analysis, entity recognition, and / or topic extraction categories are returned.

[0036] An example of Sentiment return result: Score = 0.8 Magnitude = 0.7 The returned results will vary depending on the content of the text being analyzed. If the text is negative, the score will be close to -1; if it is neutral, it will be close to 0; if it is positive, it will be close to 1. The magnitude value reflects the intensity of the emotion, and both positive and negative emotions may have a high magnitude value; Entities returns the result, please refer to Figure 5 As shown; it should be noted that there will be a large error in identifying objects based on audio text analysis, so the metadata is not stored in this solution; Categories returns results: If the video is a news report about international politics, the possible categories returned are ["News / International Politics", "Media / News Content"], indicating that it belongs to the International Politics News subcategory in the News category, and the Media Content in the News category as a whole; A video of a football match may be ["Sports / Football", "Media / SportsContent"], which indicates that the specific sport is football and that the video belongs to sports media content. A video clip of a comedy movie may be classified as ["Entertainment / Comedy Films", "Media / Entertainment Content"], which means it is a comedy movie in the entertainment category and belongs to the entertainment content category.

[0037] Further, Sentiment metadata storage rules: Create a new content score table, which is associated with the content metadata table through the content unique identifier. The main data stored include the content metadata unique identifier, Score, and Magnitude.

[0038] Categories metadata storage rules: Create a new content classification table, which is associated with the content metadata table through the content unique identifier. The main data stored include the content metadata unique identifier and the category name (CategoryName).

[0039] In summary, the automatic classification and labeling of on-demand and live content is realized. In other words, the automatic classification and labeling of target videos can be realized.

[0040] S130, matching the target tag with the constructed user preference model to generate a recommendation list.

[0041] In some embodiments, the user preference model can be constructed as follows: Obtain a video data sample set; Extracting text from the video in the video data sample set to obtain video text; refer to the above-mentioned video text extraction method; Converting the video text into a semantic vector; The semantic vector is taken as input, the probability values ​​of all event labels are taken as output, a multi-label scene loss function is constructed through the first deep learning model, and the second deep model is used as the vector feature extraction representation to train the user preference model.

[0042] Specifically, the video text is segmented, and a CLS tag ([CLS] tag) is connected at the beginning of each text data; The word segmentation process of the video text includes: The government data sample set may be segmented using Jieba, SnowNLP, PkuSeg, THULAC and / or HanLP.

[0043] The text data after word segmentation is represented by embedding vectorization to obtain the CLS semantic encoding vector. That is, each word after word segmentation is represented by a feature vector based on the BERT pre-training model, each sentence is represented by embedding vectorization, the relative position of the text words is represented by encoding vectors, and the three feature vectors are added together.

[0044] Construct a multi-label scene loss function based on the pre-trained first deep learning model and perform finetuning training. The first deep learning model can use the encoder module in the bidirectional transformer model (second deep model) as the vector feature extraction representation; Among them, the first deep learning model includes an attention mechanism, which can automatically mine the semantic association relationship between the current word in the text and other words in the context, and ignore the distance, so that the semantic vector representation of the word obtained can fully mine the association information of the context.

[0045] In some embodiments, a fully connected layer may be set, and the CLS semantic encoding vector may be used as the input of the fully connected layer, and the output dimension length may be the number of event categories.

[0046] In some embodiments, the following custom loss function is used as the optimization objective:

[0047] Wherein, N is a set of negative samples; P is a positive sample set; Said is the score of the proportion of category i in the positive samples; Said is the score of the proportion of category j in negative samples.

[0048] Compared with the existing loss function, the above custom loss function can take into account the multi-label scenario. In the multi-label scenario, the above loss function can be efficiently fitted and can avoid the problem of imbalanced label category samples in large-scale classification.

[0049] In some embodiments, the first deep learning model is subjected to parameter distillation processing through a preset distillation mechanism. The first deep learning model is divided into N modules, and the multi-layer parameter layers in the first module are replaced with a layer of transformer parameters initialized by normal distribution to obtain a first module replacement layer; The first module replacement layer and the remaining modules are fine-tuned for multi-label task training, and the replacement layer parameters of the first module are retained after the training is completed; Repeat the above steps, and after the N modules have completed the multi-label task fine-tuning training, integrate the replacement layer parameters of all modules to construct a multi-layer BERT pre-training parameter to complete the parameter distillation of the first deep model.

[0050] When fine-tuning downstream tasks, parameter distillation can be performed for BERT, and the X-layer (preferably, the transformer is set to 12 layers) BERT is divided into three parts, 1-4 layers are A modules, 5-8 layers are B modules, and 9-12 layers are C modules. The model is trained in four rounds. In the first round of training, the parameters of the 1-4 layers of the A module are replaced with the parameters of the transformer layer initialized by the normal distribution. At this time, BERT has a total of 9 layers (1 layer of the A module, 4 layers of the B module and the C module respectively), and participates in the downstream multi-label task fine-tuning training. After the training is completed, the replacement layer parameters of the A module in this round are retained, and the replacement layer parameter training of the second round of the B module and the third round of the C module is completed by analogy. In the fourth round of training, the parameters of the replacement layers of the three modules A, B, and C are taken to construct a three-layer BERT pre-training parameter, and participate in the fine-tuning training of the downstream multi-label task.

[0051] Based on the optimal distillation model after training, the inference function is performed. According to the input video text, the probability values ​​of all event labels are calculated, and those greater than the defined multi-label threshold are retained. The corresponding labels are converted. By adding the distillation mechanism of the pre-trained model in the fine-tuning training of the multi-label classification task, the inference speed of the classification model is improved, and more accurate user labels are obtained.

[0052] In some embodiments, the data of the video watched by the user is input into the user preference model trained in the above manner to obtain one or more user tags of the current user. Furthermore, the target tag and user tag obtained through step S120 are matched and scored; for example, the total score of the content is 10 points, and the score of the content = 5 points * (number of likes / average number of likes for the content) + 3 points * (number of collections / average number of collections for the content) + 2 points * score of the content, and finally content with a score greater than 5 is selected to generate a recommendation list.

[0053] According to the embodiments of the present disclosure, the following technical effects are achieved: Improves operational efficiency and reduces the heavy workload of operators in recommending content. Can analyze live events and recommend them to users, ensuring the real-time nature of the content. At the same time, through the user preference model training method disclosed in the present invention, a multi-label classification model for video classification is constructed, which realizes the classification modeling of large-scale multi-event videos, and can quickly and accurately classify videos.

[0054] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.

[0055] The above is an introduction to the method embodiment. The following is a further explanation of the scheme described in this application through an apparatus embodiment.

[0056] Figure 6 A block diagram of a content recommendation device based on natural language analysis in the OTT industry according to an embodiment of the present application is shown, Figure 6 Shown include: The acquisition module 610 is used to acquire a target video; the target video includes a video on demand and / or a live event video; An extraction module 620 is used to extract text from the target video to obtain a target tag; The matching module 630 is used to match the target tag with the constructed user preference model to generate a recommendation list.

[0057] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0058] Figure 7 A schematic diagram of the structure of a terminal device or server suitable for implementing an embodiment of the present application is shown.

[0059] like Figure 7As shown, the terminal device or server includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage part 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the terminal device or server are also stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0060] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that a computer program read therefrom is installed into the storage section 708 as needed.

[0061] In particular, according to an embodiment of the present application, the above method flow steps can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above-mentioned functions defined in the system of the present application are executed.

[0062] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0063] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of the code, and the aforementioned module, program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0064] The units or modules involved in the embodiments described in the present application may be implemented by software or hardware. The units or modules described may also be arranged in a processor. The names of these units or modules do not, in some cases, constitute limitations on the units or modules themselves.

[0065] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or may exist independently without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the method described in the present application.

[0066] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of application involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the aforementioned application concept. For example, the above features are replaced with (but not limited to) technical features with similar functions applied in the present application.

Claims

1. A content recommendation method based on natural language analysis in the OTT industry, characterized in that: include: Get the target video; The target video includes a video on demand and / or a live event video; Performing text extraction on the target video to obtain a target label; The target tag is matched with the constructed user preference model to generate a recommendation list.

2. The method according to claim 1, characterized in that: The extracting text from the target video to obtain the target tag comprises: If the target video is a live event video, recording a video on demand corresponding to the live event video according to the type of the current device and the start time and end time of the live event video; Extracting text from the on-demand video to obtain video text; The video text is analyzed through the Google Cloud Natural Language API to obtain a target label.

3. The method according to claim 2, characterized in that The target tags include sentiment analysis tags, entity recognition tags and / or topic extraction tags.

4. The method according to claim 3, characterized in that Also includes: The user preference model can be constructed in the following way: Obtain a video data sample set; Extracting text from the video in the video data sample set to obtain video text; Converting the video text into a semantic vector; The semantic vector is taken as input, the probability values ​​of all event labels are taken as output, a multi-label scene loss function is constructed through the first deep learning model, and the second deep model is used as the vector feature extraction representation to train the user preference model.

5. The method according to claim 4, characterized in that The user preference model may adopt the following loss function: Wherein, N is a set of negative samples; P is a positive sample set; Said is the score of the proportion of category i in the positive samples; Said is the score of the proportion of category j in negative samples.

6. A content recommendation device based on natural language analysis in the OTT industry, characterized in that: include: An acquisition module, used to acquire a target video; The target video includes a video on demand and / or a live event video; An extraction module, used to extract text from the target video to obtain a target label; The matching module is used to match the target tag with the constructed user preference model to generate a recommendation list.

7. The device according to claim 6, characterized in that The extracting text from the target video to obtain the target tag comprises: If the target video is a live event video, recording a video on demand corresponding to the live event video according to the type of the current device and the start time and end time of the live event video; Extracting text from the on-demand video to obtain video text; The video text is analyzed through the Google Cloud Natural Language API to obtain a target label.

8. The device according to claim 7, characterized in that The target tags include sentiment analysis tags, entity recognition tags and / or topic extraction tags.

9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.