Information processing method and device, electronic device, storage medium, and program product

By generating event topics, analyzing correlations and aggregating video events, the problems of single video event content and low generation efficiency are solved, and rich content is updated in a timely manner.

CN114491149BActive Publication Date: 2025-09-19TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210040341.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-13
Publication Date
2025-09-19
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

The existing video events have single content, low generation efficiency, and cannot keep up with the pace of event changes in a timely manner.

Method used

An event theme is generated by acquiring information associated with the video, analyzing the association relationship between the first video event and the second video event, and aggregating them to generate a video event containing the event association relationship.

Benefits of technology

It enriches the content of video events, improves generation efficiency, enables timely discovery of the latest status of events, reduces manual editing, and lowers labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114491149B_ABST
    Figure CN114491149B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose an information processing method and apparatus, electronic device, storage medium, and program product. The method includes: obtaining information associated with a video and generating an event theme based on the obtained information; generating a first video event corresponding to the event theme and obtaining a second video event associated with the first video event; analyzing the first video event and the second video event to obtain an event association relationship; and aggregating the first video event and the second video event based on the event association relationship to obtain a video event containing the event association relationship. The technical solutions of the embodiments of the present application can enrich the content of video events and improve the efficiency of video event generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and more specifically, to an information processing method and device, an electronic device, a storage medium, and a program product. Background Art

[0002] With the advancement of communication technology, users' demand for information is shifting from text to video. Video is poised to become one of the dominant forms of internet content, replacing content consumption to a certain extent and gradually gaining a dominant position in media such as news and social platforms. In some scenarios, video events need to be generated, through which users can view detailed descriptions of the video events, relevant characters, and other information. Currently, the content of video events is relatively simple, and is often selected by operators based on their own experience. This is inefficient and cannot keep up with the pace of event changes. Summary of the Invention

[0003] To solve the above technical problems, the embodiments of the present application provide an information processing method and device, an electronic device, a storage medium, and a program product.

[0004] According to one aspect of an embodiment of the present application, there is provided an information processing method, the method comprising:

[0005] Acquire information associated with the video and generate an event topic based on the acquired information;

[0006] Generate a first video event corresponding to the event theme, and obtain a second video event associated with the first video event;

[0007] Analyzing the first video event and the second video event to obtain an event association relationship;

[0008] The first video event and the second video event are aggregated according to the event association relationship to obtain a video event containing the event association relationship.

[0009] According to one aspect of an embodiment of the present application, there is provided an information processing device, the device comprising:

[0010] a generation module configured to obtain information associated with the video and generate an event topic based on the obtained information;

[0011] an acquisition module configured to generate a first video event corresponding to the event theme and acquire a second video event associated with the first video event;

[0012] an analysis module configured to analyze the first video event and the second video event to obtain an event association relationship;

[0013] The aggregation module is configured to aggregate the first video event and the second video event according to the event association relationship to obtain a video event containing the event association relationship.

[0014] According to one aspect of an embodiment of the present application, an electronic device is provided, including:

[0015] one or more processors;

[0016] The storage device is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the information processing method as described above.

[0017] According to one aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of an electronic device, the electronic device executes the information processing method as described above.

[0018] According to one aspect of an embodiment of the present application, a computer program product is provided, including a computer program, wherein when the computer instructions are executed by a processor, the information processing method as described above is implemented.

[0019] In the technical solution provided in the embodiments of the present application, on the one hand, the generated video events contain event association relationships, which enriches the content of the video events and enables users to better understand the event content; on the other hand, automatically generating event themes, analyzing event association relationships, and generating video events not only improves the efficiency of video event generation, but also can timely discover the latest status of an event and aggregate it with other states of the event to obtain a video event, so that users can obtain the latest development status of the event in a timely manner.

[0020] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, serving to explain the principles of the present application. It is obvious that the drawings described below are merely some embodiments of the present application, and a person of ordinary skill in the art can derive other drawings based on these drawings without inventive effort. In the drawings:

[0022] Figure 1 It is a schematic diagram of an implementation environment involved in this application;

[0023] Figure 2is a flowchart of an information processing method shown in an exemplary embodiment of the present application;

[0024] Figure 3 yes Figure 2 A flowchart of step S110 in the illustrated embodiment in an exemplary embodiment;

[0025] Figure 4 is a schematic diagram of a process for determining an event theme, shown in an exemplary embodiment of the present application;

[0026] Figure 5 yes Figure 2 A flowchart of step S110 in the illustrated embodiment in an exemplary embodiment;

[0027] Figure 6 is a flowchart of obtaining a video content vector shown in an exemplary embodiment of the present application;

[0028] Figure 7 yes Figure 2 A flow chart of step S120 in the illustrated embodiment in an exemplary embodiment;

[0029] Figure 8 is a flowchart illustrating an exemplary embodiment of the present application for generating a video event based on an event theme;

[0030] Figure 9 yes Figure 7 A flow chart of step S440 in the illustrated embodiment in an exemplary embodiment;

[0031] Figure 10 yes Figure 2 A flow chart of step S120 in the illustrated embodiment in an exemplary embodiment;

[0032] Figure 11 is a flowchart of an information processing method shown in an exemplary embodiment of the present application;

[0033] Figure 12 is a schematic diagram of another implementation environment involved in this application;

[0034] Figure 13 is a structural diagram of an information processing device shown in an exemplary embodiment of the present application;

[0035] Figure 14 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0036] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0037] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0038] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.

[0039] It should also be noted that the term "plurality" used in this application refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0040] Before introducing the technical solutions of the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are explained first. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0041] Blockchain is a new application model for computer technologies, including distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a series of data blocks linked using cryptographic methods. Each block contains information about a batch of online transactions, used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product and service layer, and the application service layer.

[0042] The platform's product service layer provides the basic capabilities and implementation framework for typical applications. Developers can build on these basic capabilities, overlay business features, and complete the blockchain implementation of business logic. The application service layer provides application services based on blockchain solutions for business participants to use.

[0043] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0044] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.

[0045] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0046] Social networks originated from online social networking, which began with email. The internet is essentially a network of computers connected to each other. Early email solved the problem of remote email transmission and remains the most popular application on the internet. BBS (Bulletin Board System) was an early platform for spontaneously generating internet content. It normalized "group posting" and "forwarding," theoretically enabling the ability to publish information and discuss topics to everyone. With the development of communication technologies, such as the widespread adoption of smartphones, the ubiquity of Wi-Fi (wireless network communication technology), the general reduction in 4G rates, and the advent of the 5G era, users' demand for information is gradually shifting from text to video. Video (especially short videos) will gradually become one of the dominant content forms on the mobile internet, replacing content consumption to a certain extent and gradually gaining a dominant position in media such as news and social platforms. Currently, the content of video events is relatively simple, and video event content is usually selected by operators based on their own experience, which is inefficient and unable to keep up with the pace of event changes. Based on this, the embodiments of the present application provide an information processing method and device, electronic equipment, storage medium, and program product that enrich the content of video events and improve the efficiency of video event generation.

[0047] See also Figure 1 , Figure 1 Schematic diagram of an implementation environment involved in this application. This implementation environment includes an information processing device 100, a platform 200, and a terminal 300. The platform 200 includes a video content library for storing videos and video metadata and other information. The information processing device 100, platform 200, and terminal 300 communicate with each other via a wired or wireless network.

[0048] It should be understood that Figure 1 The number of information processing devices 100, platforms 200, and terminals 300 in the figure is merely illustrative. Any number of information processing devices 100, platforms 200, and terminals 300 may be provided according to actual needs.

[0049] The information processing device 100 may be a server or other device. The server may be a server that provides various services. It may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This is not limited here.

[0050] Platform 200 is used to store and display videos. It can be an Internet platform. Platform 200 can be deployed on a server or other device, and the video content library is deployed in the storage system corresponding to platform 200. The storage system can be a storage system built based on cloud storage technology, or of course, other types of storage systems. Cloud storage is a new concept that extends and develops from the concept of cloud computing. A distributed cloud storage system refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to bring together a large number of different types of storage devices (storage devices are also called storage nodes) in the network through application software or application interfaces to work together and provide external data storage and service access functions. The storage system can also be a blockchain system, that is, the video content library can be stored in the blockchain system.

[0051] The terminal can be an electronic device such as a smartphone, tablet, laptop, computer, or vehicle-mounted terminal.

[0052] Users can upload videos to the platform 200 through the terminal. After receiving the video, the platform 200 can store the video and data such as the video's metadata in the video content library; the information processing device 100 can obtain information associated with the video from the video content library, and generate an event theme based on the obtained information, and then generate a first video event corresponding to the event theme, and obtain a second video event associated with the first video event. Then, the first video event and the second video event are analyzed to obtain an event association relationship. Finally, the first video event and the second video event are aggregated according to the event association relationship to obtain a video event containing an event association relationship. On the one hand, the generated video event contains an event association relationship, which enriches the content of the video event and enables users to better understand the event content. On the other hand, automatically generating event themes, analyzing event association relationships, and generating video events not only improves the efficiency of video event generation, but also can timely discover the latest status of an event, and aggregate it with other states of the event to obtain a video event, so that users can obtain the latest development status of the event in a timely manner.

[0053] In some embodiments, the process of a user uploading a video to the platform 200 through a terminal may include: the user can shoot a video using a shooting tool on the terminal (such as an instant messaging software with a video shooting function, a short video social software, etc.), and then upload the video to the platform through the terminal. During the video upload process, the video will be re-transcoded and the video file will be normalized, and the video metadata will be saved to improve the video's playback compatibility on various platforms. The video will then be manually reviewed. During the manual review, some auxiliary features of the video will be obtained through a machine algorithm, such as obtaining categories, tags, etc.; then, based on the machine algorithm processing, manual standardization annotation will be performed to fill the video with relevant information, such as filling in the video's tags, categories, or a text description. This process is called video content standardization. After the video standardization is passed, it will enter the platform's video content library. The video can then be distributed to the external network or a recommendation engine. The recommendation engine recommends based on the user's profile features through a recommendation algorithm. The recommendation algorithm includes but is not limited to collaborative recommendation, matrix decomposition, and deep learning-based models. Alternatively, the user can also actively search on the platform to obtain videos in the video content library, or the user can obtain videos through social platforms (such as official accounts).

[0054] See also Figure 2 , Figure 2 This is a flowchart of an information processing method shown in an exemplary embodiment of the present application. This method can be applied to Figure 1 The implementation environment shown, which can be Figure 1 The information processing device 100 in the illustrated implementation environment executes.

[0055] like Figure 2 As shown, in an exemplary embodiment, the information processing method may include steps S110 to S140, which are described in detail as follows:

[0056] Step S110: Acquire information associated with the video, and generate an event theme based on the acquired information.

[0057] It should be noted that video is a dynamic image, and its types include but are not limited to short videos, micro-movies, etc.

[0058] Short videos, specifically, refer to frequently pushed video content played on various new media platforms, suitable for viewing on the go and during short breaks. These videos range in length from a few seconds to several minutes and incorporate topics such as skill sharing, humor, fashion trends, social issues, street interviews, public welfare education, advertising creativity, and commercial customization. Due to their short duration, they can be produced as standalone films or as part of a series. Unlike micro-films and live broadcasts, short video production does not require the same specific expression format and team requirements. They offer a simple production process, low barriers to entry, and high levels of engagement, making them more valuable than live broadcasts. However, their ultra-short production cycle and engaging content present challenges to the copywriting and planning skills of short video production teams. Excellent short video production teams often rely on established self-media platforms or intellectual property (IP). The emergence of short videos has enriched the forms of native advertising in new media. Short video producers have evolved from user-generated content (UGC), professionally produced content (PGC), and user uploads to specialized short video production organizations, MCNs (multi-channel networks), and specialized short video apps. Short video has become a key dissemination method for content startups and social media platforms. The variety of short videos is increasing, and the content is becoming increasingly diverse. Both producers and consumers of short video content have become a vast community.

[0059] Videos can be provided to users in a feed-based manner. A feed is a source of information, a way to present content to users and continuously update it, disseminating the latest information to users through the feed.

[0060] The information associated with the video is relevant information of the video, including but not limited to at least one of the meta-information of the video and the keywords of the video; wherein the meta-information of the video includes but is not limited to the title of the video, the publisher of the video, the summary of the video, the cover image, the release time, the size of the video file, the format of the video file, whether it is original, whether it is the first release, classification information, text information obtained by performing text recognition on the video, text information obtained by performing voice recognition on the audio in the video, etc.; wherein the classification information can be information annotated for the video during the manual review process, and the classification information can include categories and tags. The categories can be classified according to multiple levels. For example, for a video explaining the mobile phones of Company A, the first-level category can be technology, the second-level category can be smart phones, the third-level category can be domestic mobile phones, and the tags can be Company A, mobile phone model, etc.; the method of performing text recognition on the video can be based on OCR (Optical Character Recognition) technology.

[0061] The event topic is the theme of the video event to be generated. To ensure topicality, the event topic can be a relatively short text that describes the main information of the event, for example, "Zhurong Mars Rover Landed on Mars."

[0062] In order to generate an event theme, in this embodiment, information associated with the video may be acquired, and the event theme may be generated based on the acquired information.

[0063] S120: Generate a first video event corresponding to the event theme, and obtain a second video event associated with the first video event.

[0064] It should be noted that the first video event is an event generated based on an event theme, which includes a corresponding video, event description information, and the like.

[0065] The second video event is a video event associated with the first video event. It may be an event belonging to the same theme as the first video event. The second video event may be a historical video event, that is, a video event generated before the first video event is generated.

[0066] After obtaining the event theme, a first video event can be generated according to the event theme. The title of the first video event can be the event theme, and the first video event can include content such as the corresponding video and event description information.

[0067] It should be understood that events with the same theme can include different stages of development. For example, a celebrity's love and marriage story includes stages such as marriage, marriage breakdown, and divorce; and "Mars rover landing on Mars" includes stages such as rocket launch, rocket mid-flight, trajectory change, Mars rover landing on Mars, and Mars rover patrolling the Martian surface. Therefore, after generating a first video event, in order to find a video event with the same theme as the first video event, you can also obtain video events associated with the first video event and use these video events as the second video event.

[0068] Step S130: Analyze the first video event and the second video event to obtain an event association relationship.

[0069] It should be noted that event correlation refers to the relationship between multiple events, including but not limited to the development context of events (for example, the time development context, i.e., the timeline), the causal relationship between events, the relationship between the main characters in multiple events, etc.

[0070] After obtaining the second video event associated with the first video event, the first video event and the second video event are analyzed to obtain an event association relationship.

[0071] The specific analysis method can be flexibly set according to actual needs. In one example, in order to let users understand the order of event development, the time sequence of the first video event and the second video event can be sorted out to obtain a timeline. The time sequence between the first video event and the second video event can be determined based on the generation time of the first video event and the second video event; or the time sequence between the first video event and the second video event can be determined based on the occurrence time of the first video event and the second video event. In another example, the themes of the first video event and the second video event can be analyzed to determine the causal relationship between the occurrence of the first video event and the second video event or the order of the occurrence of the first video event and the second video event based on natural laws and the obtained themes. The causal relationship or order of the occurrence of the events can be obtained by analyzing the themes of the first video event and the second video event using a machine learning model. For example, if the first video event is "A and B get married" and the second video event is "A and B get divorced", by analyzing the themes of the first video event and the second video event, it can be determined that the first video event occurred first and the second video event occurred later. In another example, the main characters of each of the multiple events can be obtained, and then the relationship between the main characters can be obtained, and the relationship between the main characters can be used as the event association relationship.

[0072] Step S140 : Aggregate the first video event and the second video event according to the event association relationship to obtain a video event containing the event association relationship.

[0073] After the event association relationship is obtained, the first video event and the second video event may be aggregated according to the event association relationship to obtain a video event containing the event association relationship.

[0074] In this embodiment, information associated with the video is obtained, and an event theme is generated based on the obtained information; a first video event corresponding to the event theme is generated, and a second video event associated with the first video event is obtained; the first video event and the second video event are analyzed to obtain an event association relationship; the first video event and the second video event are aggregated according to the event association relationship to obtain a video event containing an event association relationship. On the one hand, the generated video event contains the event association relationship, which enriches the content of the video event and enables users to better understand the event content; on the other hand, automatically generating an event theme, analyzing the event association relationship, and generating a video event not only improves the efficiency of video event generation, but also can timely discover the latest status of an event, and aggregate it with other statuses of the event to obtain a video event, so that users can timely obtain the latest development status of the event; and it can also effectively reduce the process of manual editing of video events, reduce labor costs, improve business response and improve user experience.

[0075] In an exemplary embodiment, in order to ensure that the generated video event is popular and can arouse extensive discussion among users, and improve the efficiency of video event generation, Figure 2 Step S110 in the illustrated embodiment, i.e., the process of obtaining information associated with the video and generating an event theme based on the obtained information, may include: crawling short texts whose information popularity meets preset popularity conditions from different platforms, and using the crawled short texts as event themes.

[0076] It should be noted that the platform can be various Internet service platforms, including but not limited to search engine information platforms, social platforms, etc., among which social platforms include but are not limited to short video social platforms, other social platforms with video playback functions, etc.

[0077] Information popularity is a parameter that can reflect the popularity of information among users. It can be expressed by parameters that can reflect the popularity of information among users, such as the number of clicks, searches, readings, reposts, comments, likes, and number of discussion participants.

[0078] Short text is usually composed of a few words, which is short in length and easy for users to remember.

[0079] The preset popularity condition is a pre-set judgment condition used to determine whether a short text can be used as an event topic.

[0080] In this embodiment, in order to ensure the comprehensiveness and popularity of the event, end texts whose information popularity meets preset conditions are crawled from different platforms and used as event topics; the specific implementation method can be flexibly set according to actual needs.

[0081] In one embodiment, in order to facilitate users to understand popular events, a hot list is usually set up on the platform. For example, a search engine information platform will set up a hot search list based on parameters such as search volume and whether the number of videos recalled based on the search text has suddenly increased. A social platform will set up a hot list based on click-through rate, reading volume, forwarding volume, number of topic citations, etc.; therefore, in order to improve crawling efficiency and accuracy, the process of crawling short texts whose information popularity meets preset popularity conditions from different platforms and using the crawled short texts as event topics can include: crawling hot lists from different platforms, and using the short texts in the hot lists as event topics.

[0082] In another embodiment, the process of crawling short texts from different platforms whose information popularity meets a preset popularity condition and using the crawled short texts as the event subject may include crawling short texts from different platforms whose information popularity exceeds a preset popularity threshold and using the crawled short texts as the event subject. The preset popularity threshold can be flexibly set according to actual needs, for example, it can be more than 1 million forwardings, more than 2 million comments, etc.

[0083] By crawling short texts whose information popularity meets the preset popularity conditions from different platforms and using the crawled short texts as event topics, video events can be discovered in a timely manner, ensuring the topicality and popularity of the generated events, effectively reducing the process of manual discovery of video events, reducing labor costs, and further improving the efficiency of video event generation. In addition, crawling information from different platforms can ensure the comprehensiveness of the generated events.

[0084] In another exemplary embodiment, see Figure 3 As shown, Figure 3 for Figure 2 The flowchart of step S110 in the embodiment shown is in an exemplary embodiment. Figure 3 As shown, in the case where the information associated with the video includes the video title, the process of generating the event theme based on the acquired information may include steps S210 to S240, which are described in detail as follows:

[0085] Step S210: segment the video title to obtain a first candidate topic.

[0086] In this embodiment, after obtaining the video title, the video title may be segmented to obtain the first candidate topic.

[0087] The method for obtaining video titles can be flexibly set according to actual needs. For example, the video title can be input by the operator or crawled from the Internet. To improve the quality of the title, video titles can be crawled from content published by mainstream media, official accounts, authoritative websites, etc.

[0088] The method of segmenting the video title can be flexibly set according to actual needs. For example, the video title can be segmented using punctuation marks as segmentation points. In an example, assuming the video title is "A certain animal has over-breeded and has spread across 16 provinces. Why does no one dare to eat it?", it can be segmented into three short texts: "A certain animal has over-breeded", "Has spread across 16 provinces", and "Why does no one dare to eat it", and these three short texts are used as the first candidate topics.

[0089] Step S220: cluster the video titles to obtain video title clusters, and generate second candidate topics corresponding to the video title clusters.

[0090] In this embodiment, after obtaining the video titles, the video titles can be clustered to obtain several video title sets, each of which is considered a video title cluster. Then, candidate topics corresponding to the video title clusters, i.e., second candidate topics, are generated. For each video title cluster, one candidate topic can be generated.

[0091] When clustering, the video titles can be clustered using a clustering algorithm to obtain video title clusters.

[0092] It should be noted that, in this embodiment, the order of step S210 and step S220 is not restricted, wherein step S210 may be executed first, and then step S220; or, step S220 may be executed first, and then step S210; or, step S210 and step S220 may be executed simultaneously.

[0093] Step S230 : clustering the first candidate topic and the second candidate topic to obtain a candidate topic cluster.

[0094] After obtaining the first candidate topic and the second candidate topic, the first candidate topic and the second candidate topic may be clustered to obtain several candidate topic sets, each candidate topic set being a candidate topic cluster.

[0095] In some embodiments, in order to improve the quality of candidate topics, after obtaining the first candidate topic and the second candidate topic, the first candidate topic and the second candidate topic can be filtered according to preset filtering rules, and the filtered first candidate topic and the second candidate topic can be clustered to obtain a candidate topic cluster.

[0096] The filtering rules can be flexibly set according to actual needs. For example, the filtering rules include but are not limited to at least one of the following methods:

[0097] The first method is to delete a candidate topic if it contains a violation word. A violation word set can be pre-set. If a candidate topic among the first and second candidate topics contains a word in the violation word set, the candidate topic is deleted.

[0098] The second method is to delete a candidate topic if it exceeds a preset length. The preset length can be flexibly set based on actual needs, for example, 15 or 10. If the event topic is too long, its topicality and popularity are low. Deleting candidate topics that exceed the preset length can improve the topicality and popularity of the candidate topic.

[0099] The third method is to delete the candidate topic if it does not include a named entity, thereby filtering out candidate topics without event content. Named entities include names of people, organizations, places, and all other entities identified by names. For example, if the sentence "A certain animal has over-proliferated and has spread across 16 provinces, why does no one dare to eat it" is segmented into three short texts: "A certain animal has over-proliferated," "has spread across 16 provinces," and "why does no one dare to eat it," "why does no one dare to eat it" does not include a named entity, and the content of the event cannot be known from this short text, so it can be deleted.

[0100] Step S240: determining the event topic based on the cluster center of the candidate topic cluster.

[0101] After the candidate topic clusters are obtained, the event topics are determined according to the cluster centers of the candidate topic clusters. The cluster centers of the candidate topic clusters can be directly used as the event topics, or the event topics can be generated according to the cluster centers of the candidate topic clusters.

[0102] In order to improve the quality of event topics, after obtaining the candidate topic cluster, before determining the event topic based on the cluster center of the candidate topic cluster, it is also possible to determine whether the candidate topic cluster is a text describing the event based on the preset event detection rules. If so, the event topic is determined based on the cluster center of the candidate topic cluster.

[0103] Among them, the specific method of determining whether a candidate topic cluster is a text describing an event based on the preset event detection rules can be flexibly set according to actual needs. In one embodiment, it can be determined based on at least one of the parameters such as the source of the video corresponding to each candidate topic in the candidate topic cluster, whether the candidate topic contains named entities, and whether the candidate topic contains words of a specific part of speech, wherein the specific part of speech can include nouns, verbs, etc.; for example, when the proportion of candidate topics whose corresponding video sources are authoritative websites in the candidate topic cluster reaches a certain value, the candidate topic cluster can be determined to be a text describing an event; for another example, when the proportion of candidate topics containing words of a specific part of speech in the candidate topic cluster reaches a certain value, the candidate topic cluster can be determined to be a text describing an event.

[0104] To better understand the solution of this embodiment, see Figure 4 , Figure 4 As an example, a process diagram for determining the event theme based on the video title is shown in FIG. Figure 4 As shown, video titles can be obtained from authoritative websites, etc., and the video titles can be segmented to obtain first candidate topics; and the video titles can be clustered to obtain video title clusters, and second candidate topics can be generated for each video title cluster. The first candidate topics and the second candidate topics can be filtered based on filtering rules. After filtering, the first candidate topics and the second candidate topics can be clustered to obtain candidate topic clusters. Event detection can be performed on the candidate topic clusters based on event detection rules, and the event topic can be determined according to the cluster center of the candidate topic cluster that has passed the detection.

[0105] In this embodiment, the video title is segmented to obtain the first candidate topic, the video title is clustered to obtain a video title cluster, and a second candidate topic corresponding to the video title cluster is generated. The first candidate topic and the second candidate topic are clustered to obtain a candidate topic cluster, and the event topic is determined according to the cluster center of the candidate topic cluster, thereby automatically generating the event topic and improving the generation speed of the event topic.

[0106] In another exemplary embodiment, see Figure 5 As shown, Figure 5 for Figure 2 The flowchart of step S110 in the embodiment shown is in an exemplary embodiment. Figure 5 As shown, in the case where the information associated with the video includes videos uploaded within a preset time period, the process of generating an event theme based on the acquired information may include steps S310 to S330, which are described in detail as follows:

[0107] Step S310: clustering the videos uploaded within a preset time period to obtain a plurality of video clusters.

[0108] The preset time period can be flexibly set according to actual needs, for example, it can be set to 10 minutes, 20 minutes, etc.

[0109] After the user has finished making the video, he can upload the video to the platform, and the video content library of the platform will store the video uploaded by the user. In this embodiment, the video uploaded within the preset time period can be obtained from the video content library.

[0110] After obtaining videos uploaded within a preset time period, the obtained videos can be clustered to obtain multiple video sets, each of which is a video cluster; this can determine the concentration level of similar videos, and then determine whether different media accounts have recently reported on the same event in a concentrated manner, thereby discovering hot events.

[0111] The specific clustering method can be flexibly set according to actual needs.

[0112] In one embodiment, videos may be clustered based on text information associated with the videos, such as video titles, keywords, video summaries, text information obtained by performing text recognition on the videos, and text information obtained by performing speech recognition on the audio in the videos.

[0113] In another embodiment, the video content uploaded within a preset time period can be analyzed by a video classification model to obtain a video content vector for each video, and then clustering can be performed based on the video content vector. The video classification model is a model established based on machine learning that can extract features from videos to obtain video content vectors. The specific structure of the video classification model can be flexibly set according to actual needs; the video content vector can be understood as an "implicit" feature based on the video content, which contains two layers of meaning: the first layer of meaning: representation learning, low-dimensional dense features, one-dimensional arrays (for example, the video content vector is 128 floats); the second layer of meaning: metric learning, a vector of similarity measurement, the "distance" between two vectors represents the "similarity" of the two objects. In one example, see Figure 6As shown, the process of analyzing the content of a video using a machine learning model to obtain a video content vector for the video may include: inputting the video into the machine learning model, the TSN (Temporal Segment Networks) included in the machine learning model extracting a video frame sequence from the video to obtain a number of video frames, the Xception module included in the machine learning model extracting image features from the several video frames extracted by the TSN, then obtaining an image feature vector using NeXtVLad included in the machine learning model, and finally performing a weighted average of the image feature vectors to obtain a video content vector. Xception is another improvement to Inception-v3 proposed by Google following Inception; NeXtVLad is an image feature extraction algorithm used to aggregate frame-level features of a video clip into a feature vector.

[0114] When clustering videos uploaded within a preset time period and obtaining multiple video clusters, the videos can be clustered based on a hierarchical clustering method. The hierarchical clustering method is used to hierarchically decompose a given set of data objects. Based on the decomposition strategy used in the hierarchical decomposition, the hierarchical clustering method can be divided into agglomerative (i.e., top-down) and divisive (i.e., bottom-up) hierarchical clustering. The clustering process of the divisive method can be as follows:

[0115] Input: video set D that needs to be clustered, end condition.

[0116] Output: clustering results.

[0117] Process: 1. Classify all samples in the video set D into a cluster;

[0118] 2. Calculate the distance between any two samples in the same cluster (denoted as c) and find the two samples a and b with the farthest distance.

[0119] 3. Assign samples a and b to different clusters c1 and c2;

[0120] 4. Calculate the distances between the remaining sample points in the original cluster (c) and a and b. If the distance to a (dis(a)) is less than the distance to b (dis(b), then the sample point is assigned to c1; otherwise, it is assigned to c2.

[0121] End: Repeat steps 2-4 until the entered "end condition" is reached.

[0122] The termination condition can be flexibly set according to actual needs. In one embodiment, the termination condition may include the number of clusters, which is the number of clusters finally obtained. In the process of repeating steps 2-4, if the number of clusters obtained reaches the number of clusters, the repetitive process is terminated and the result is output. In one example, if the number of clustered data is 5, the video set D is divided into 5 clusters. In another embodiment, the termination condition may include: the distance between different clusters is less than a preset threshold. The preset threshold can be flexibly set according to actual needs. The distance between different clusters can be the distance between the cluster centers of different clusters, or the minimum distance between any two samples in different clusters, etc.

[0123] Step S320 , selecting a target video cluster having a number of video items greater than a preset value from the plurality of video clusters.

[0124] The preset value can be flexibly set according to actual needs, for example, it can be set to 100, 1000, etc.

[0125] After obtaining multiple video clusters, if the number of videos contained in a certain video cluster is greater than a preset value, it indicates that different media accounts have concentrated on reporting on the same event. Therefore, the video cluster containing a number of videos greater than the preset value can be used as the target video cluster, so as to discover hot events before they ferment.

[0126] Step S330: determining the event theme according to the cluster center of the target video cluster.

[0127] After the target video cluster is screened, the event theme can be determined based on the cluster center of the target video cluster. The specific determination method can be flexibly set according to actual needs.

[0128] It should be understood that the cluster center of the target video cluster is a video. In one embodiment, the title of the video can be used as the event theme. In another embodiment, the event theme can be generated based on the keywords of the video. For example, the keywords of the video can be combined to generate a short text, and the short text can be used as the event theme. In another embodiment, after obtaining the video title of the video, the process can proceed to S210 to obtain the event theme.

[0129] In this embodiment, the videos uploaded within a preset time period are clustered to obtain multiple video clusters, and a target video cluster with a number of videos greater than a preset value is screened out from the multiple video clusters. The event theme is determined based on the cluster center of the target video cluster, and whether different media accounts have concentrated on reporting on the same event is determined based on whether a large number of similar videos appear in the video content library. If so, the event theme is determined based on the cluster center of the similar videos, so that hot events can be discovered before they ferment.

[0130] In another exemplary embodiment, Figure 2 In the illustrated embodiment, step S110 (i.e., the process of obtaining information associated with the video and generating an event theme based on the obtained information) may include: obtaining candidate phrases associated with the video, and determining the event theme based on the information entropy of each word in the candidate phrases.

[0131] The words in the candidate phrases include but are not limited to at least one of the keywords of the video, tags of the video, words in the video title, words in the video description, and query words crawled from the Internet.

[0132] Information entropy is used to measure the expected value of a random variable. The larger the information entropy of a variable, the more possible states it may appear in, the more uncertain it is, and the greater the amount of information it contains.

[0133] The method for determining the event topic based on the information entropy of each word in the candidate phrase can be flexibly set according to actual needs. For example, in one example, target words with information entropy greater than a preset value can be screened from the candidate phrase, and the event topic can be generated based on the target words; alternatively, target words with information entropy less than a preset value can be screened from the candidate phrase, and the event topic can be generated based on the target words.

[0134] Alternatively, in another example, the mutual information between the words in the candidate phrase and the left-right information entropy of the word group in the candidate phrase may be calculated, and the event topic may be determined based on the calculated mutual information and left-right information entropy.

[0135] It should be noted that mutual information is the amount of information about another random variable contained in one random variable. Alternatively, mutual information can be viewed as the reduction in uncertainty in one random variable due to the knowledge of another random variable. It can indicate the strength of the association between words. A word group is a combination of multiple words. The left and right information entropies include left information entropy and right information entropy, which can indicate the likelihood that a word group can become a semantically independent topic. The larger the left and right information entropy value of a word group, the higher the probability that its combination will serve as the topic of an event. Therefore, the topic of an event can be determined based on the calculated mutual information and left and right information entropies. Word groups with high mutual information and high left and right information entropies can be used as the topic of the event.

[0136] In this embodiment, candidate phrases associated with the video are obtained, and the event theme is determined based on the information entropy of each word in the candidate phrases, thereby automatically generating the event theme and improving the accuracy of the event theme.

[0137] In an exemplary embodiment, see Figure 7 As shown, Figure 7 for Figure 2 The flowchart of step S120 in the embodiment shown is in an exemplary embodiment. Figure 7 As shown, the process of generating the first video event corresponding to the event theme may include steps S410 to S440, which are described in detail as follows:

[0138] Step S410: Obtain the query term corresponding to the event topic.

[0139] After generating the event topic, it is necessary to generate a video event corresponding to the event topic. The video event includes the corresponding video. In order to obtain a video that matches the video event, in this embodiment, the query term corresponding to the event topic can be obtained first, so as to facilitate searching for related videos based on the query term.

[0140] The query terms corresponding to the event topic include but are not limited to at least one of the keywords of the event topic, the event topic itself, and the candidate topics in the candidate topic cluster to which the event topic belongs.

[0141] Step S420: Retrieve candidate videos matching the query term from the video content library.

[0142] The video content library is used to store videos and information related to the videos. An inverted index can be used to create an index table to increase the speed of searching for candidate videos based on query terms.

[0143] After obtaining the query term corresponding to the event topic, the video content library can be searched for videos matching the query term, and the searched videos are used as candidate videos.

[0144] Among them, you can search for videos that match the query terms, such as video titles, video meta-information (such as text obtained by OCR recognition of the video, text obtained by recognizing the audio contained in the video), and video keywords, and use the hit videos as candidate videos.

[0145] In some implementations, to improve search speed, candidate videos matching the query term can be retrieved from the video content library based on Faiss. Faiss is an open-source library for clustering and similarity search that provides efficient similarity search and clustering for dense vectors, supporting searches on billions of vectors.

[0146] In some embodiments, after recalling candidate videos, to avoid duplication of recalled videos, the candidate videos may be deduplicated based on at least one of the following parameters: video title, video URL (Uniform Resource Locator), video cover image, video content vector, etc. This prevents identical videos from appearing in the same video event, ensuring richness of event content. After deduplication, step S430 is performed. For an introduction to video content vectors, please refer to the previous description and will not be repeated here.

[0147] In one example, deduplication of candidate videos is performed based on at least one of parameters such as video title, video URL (uniform resource locator), video cover image, video content vector, etc., which may include: if there are several candidate videos, and at least one of their parameters such as video title, URL, cover image, video content vector, etc. is the same, then one of the several candidate videos is retained. For example, assuming that the URLs of video 1, video 2 and video 3 are the same, then only video 3 can be retained.

[0148] Step S430 , calculating the correlation between the candidate videos and the query term, and screening out candidate videos whose correlation exceeds a first threshold from the candidate videos to obtain the target video.

[0149] After recalling candidate videos, in order to avoid situations where the videos are irrelevant to the event, this embodiment can also calculate the relevance of the candidate videos with the query term, and screen out candidate videos whose relevance exceeds a first threshold, and use these screened videos as target videos. The first threshold can be flexibly set according to actual needs, for example, to 90%, 95%, etc.

[0150] In some embodiments, step S430 includes: calculating the similarity between the named entities, keywords, and video titles of the candidate videos and the query terms, selecting the maximum value from the obtained similarities as the relevance of the candidate video to the query term, and screening the candidate videos whose relevance exceeds a first threshold as target videos. For example, assuming the first threshold is 90%, the similarity between the named entities of video 4 and the query terms is 70%, the similarity between the keywords of video 4 and the query terms is 94%, and the similarity between the video title of video 4 and the query terms is 98%. Then, the similarity between video 4 and the query terms is 98%. Since 98% is greater than 90%, video 4 is selected as the target video.

[0151] To improve accuracy, before calculating the relevance between candidate videos and query terms, the query terms can be filtered. After filtering, the relevance between candidate videos and query terms can be calculated. The method for filtering query terms can be flexibly set according to actual needs. For example, query terms that only contain function words or numbers can be deleted.

[0152] It should be noted that, in this embodiment, the target video is screened out from the candidate videos based on the correlation between the candidate video and the query term. In another embodiment, the target video can also be screened out from the candidate videos based on the correlation between the candidate video and the event theme. The method of calculating the correlation between the candidate video and the event theme may include: calculating the similarity between the named entities, keywords, and video titles of the candidate video and the event theme respectively, and selecting the maximum value from the obtained similarities as the correlation between the candidate video and the event theme. Of course, the correlation between the candidate video and the event theme can also be calculated by other methods; the method of screening out the target video from the candidate videos based on the correlation between the candidate video and the event theme may include: screening out candidate videos whose correlation with the event theme exceeds a threshold from the candidate videos to obtain the target video. The specific value of the threshold can be flexibly set according to actual needs.

[0153] Step S440: Generate a first video event according to the event theme and the target video.

[0154] After obtaining the target video, a first video event may be generated according to the event theme and the target video, wherein the title of the first video event may be the event theme.

[0155] In some implementations, to avoid a situation where the number of videos is too small to constitute an event, after obtaining the target videos, the present embodiment may further determine whether the number of target videos reaches a preset number. If so, a first video event is generated based on the event theme and the target videos. The specific value of the preset number can be flexibly set according to actual needs.

[0156] In some embodiments, in order to avoid repeated generation of the same video event, after generating the first video event, the first video event generated this time and the video events generated in the past can be clustered, and based on the clustering result, it can be determined whether there is a video event that is the same as the first video event. If so, the first video event is deleted. Determining whether two video events are the same can be determining whether the videos contained in the two video events are the same. Of course, other methods can also be used to determine. In one example, see Figure 8 As shown, the query terms corresponding to the event topic can be obtained first, and then candidate videos can be recalled from the video content library. Based on the correlation between the candidate videos and the query terms, the candidate videos can be filtered to obtain the target video. A first video event can be generated according to the target video, and the first video event and the historically generated video events can be clustered to obtain multiple event clusters. The same video events can be filtered out according to the clustering results, and the filtered video events can be stored.

[0157] In some embodiments, to ensure that videos in a video event are related to each other, before generating the first video event based on the event theme and the target video, the target video may be clustered to exclude irrelevant target videos. However, clustering may not be performed for event themes from a popular list.

[0158] In this embodiment, a query term corresponding to the event topic is obtained, candidate videos matching the query term are recalled from the video content library, the correlation between the candidate videos and the query term is calculated, and candidate videos whose correlation exceeds a first threshold are screened out from the candidate videos to obtain a target video. A first video event is generated based on the event topic and the target video, thereby ensuring the correlation between the video contained in the first video event and the event.

[0159] In an exemplary embodiment, see Figure 9 As shown, Figure 9 for Figure 7 The flowchart of step S440 in the embodiment shown is in an exemplary embodiment. Figure 9 As shown, under the condition that there are multiple target videos, the process of generating the first video event according to the event theme and the target videos may include steps S441 to S444, which are described in detail as follows:

[0160] In step S441 , the target videos are clustered to obtain target video clusters, and quality assessments are performed on the target videos and the target video clusters to obtain first quality values ​​corresponding to the target videos and second quality values ​​corresponding to the target video clusters.

[0161] First, it should be noted that the clustering method for clustering multiple target videos to obtain multiple target video clusters can be flexibly set according to actual needs. For example, multiple target videos can be clustered based on video content vectors and / or video reposting status to obtain multiple video sets, each of which serves as a target video cluster. The video reposting status includes at least one of the number of reposts, the number of comments, and the number of likes.

[0162] Secondly, in this embodiment, it is necessary to perform quality assessment on multiple target videos to obtain a quality value corresponding to each target video, and record the quality value of the target video as the first quality value. The method of performing quality assessment on the target video can be flexibly set according to actual needs.

[0163] In one embodiment, a first quality value of the target video may be obtained by weighted summing three parameters: the quality of the video source, the relevance of the video to the event theme, and the quality of the video content. The higher the quality of the video source, the higher the relevance of the video to the event theme, and the higher the quality of the video content, the higher the first quality value.

[0164] The quality of the video source represents the authority of the video source. A video source quality library can be pre-set to store quality values ​​of different video sources. Then, the corresponding quality value can be found from the video source quality library based on the source of the target video.

[0165] The correlation between a video and an event theme can be determined based on at least one of the similarity between the entire video title and the event theme, the similarity between each of several short texts obtained by segmenting the video title and the event theme, whether there is a named entity of the event theme in the video title, and whether the video title includes keywords of the event theme. The higher the similarity between the entire video title and the event theme, the higher the correlation between the video and the event theme; the higher the similarity between each of several short texts obtained by segmenting the video title and the event theme, the higher the correlation between the video and the event theme; if there is a named entity of the event theme in the video title, the higher the correlation between the video and the event theme; if there are keywords of the event theme in the video title, the higher the correlation between the video and the event theme.

[0166] The quality of video content can be determined based on at least one of the following: video resolution, clarity, aesthetics of the video cover image, and professionalism. The higher the resolution, clarity, aesthetics of the video cover image, and professionalism of the video, the higher the quality of the video content. The aesthetics of the video cover image can be determined by whether it has a professional layout, and the professionalism of the video can be determined by the video's shooting template, filters, transitions, and soundtrack.

[0167] In this embodiment, each target video cluster needs to be quality evaluated to obtain a quality value of each target video cluster, which is recorded as a second quality value. The method for quality evaluation of the target video cluster can be flexibly set according to actual needs.

[0168] In one embodiment, the second quality value of the target video cluster may be determined according to the number of target videos included in the target video cluster, wherein the higher the number of target videos included, the higher the second quality value of the corresponding target video cluster.

[0169] In another embodiment, the second quality value of the target video cluster may be determined based on the reposting status of the target videos included in the target video cluster, wherein the higher the number of comments, reposts, likes, etc., the higher the second quality value of the target video cluster.

[0170] Step S442 : Taking the target video with the highest first quality value in each target video cluster as a representative video, to obtain a plurality of representative videos.

[0171] In this embodiment, for each target video cluster, a video needs to be selected as its representative video, wherein the representative video is the target video with the highest first quality value in the target video cluster to which it belongs.

[0172] Step S443 : determining third quality values ​​corresponding to the plurality of representative videos according to the first quality value of the representative video and the second quality value of the target video cluster to which the representative video belongs, and sorting the plurality of representative videos in descending order of the third quality values.

[0173] After determining the first quality value of each target video, the second quality value of each target video cluster, and selecting a representative video from each target video cluster, for each representative video, the third quality value of the representative video can be determined based on the first quality value of the representative video and the second quality value of the target video cluster to which the representative video belongs. The specific determination method can be flexibly set according to actual needs. For example, the third quality value can be obtained by performing a weighted summation of the first quality value and the second quality value.

[0174] After obtaining the third quality values ​​of the plurality of representative videos, the plurality of representative videos may be sorted in descending order of the third quality values.

[0175] Step S444: Generate a first video event corresponding to the event theme and including a representative video with a specified ranking.

[0176] The designated ranking can be flexibly set according to actual needs, for example, it can be the top 10, top 20, etc.

[0177] After ranking the multiple representative videos, a representative video with a specified ranking can be selected from the multiple representative videos, and a first video event corresponding to the event theme and including the representative video with the specified ranking can be generated. Within the first video event, the representative videos can also be displayed in descending order of the third quality value. In some embodiments, the title of the first video event can be the event theme.

[0178] In this embodiment, multiple target videos are clustered to obtain multiple target video clusters, and quality assessments are performed on the multiple target videos and the multiple target video clusters respectively to obtain first quality values ​​corresponding to each of the multiple target videos and second quality values ​​corresponding to each of the multiple target video clusters; the target video with the highest first quality value in each target video cluster is used as a representative video to obtain multiple representative videos; third quality values ​​corresponding to each of the multiple representative videos are determined based on the first quality value of the representative video and the second quality value of the target video cluster to which the representative video belongs, and the multiple representative videos are sorted in descending order of the third quality values; a first video event corresponding to the event theme and containing representative videos of specified rankings is generated, thereby ensuring the quality and comprehensiveness of the videos in the first video event.

[0179] In an exemplary embodiment, see Figure 10 , Figure 10 for Figure 2 The flowchart of step S120 in the embodiment shown in the figure is as follows: Figure 10 As shown, the process of obtaining the second video event associated with the first video event may include steps S510 to S530, which are described in detail as follows:

[0180] Step S510: Acquire candidate video events that match the keyword of the first video event.

[0181] After the first video event is generated, corresponding video events may be recalled according to keywords of the first video event, and the corresponding video events may be used as candidate video events.

[0182] Step S520: Calculate the similarity between the first video event and the candidate video events.

[0183] After obtaining the candidate video events, the similarity between the first video event and the candidate video events can be calculated. The similarity between the first video event and the candidate video events can be calculated using a classification model based on machine learning. For example, the similarity between the first video event and the candidate video events can be calculated using an XGBoost classification model. XGBoost is an optimized distributed gradient boosting library designed to be efficient, flexible, and portable.

[0184] The specific method for calculating the similarity between the first video event and the candidate video event can be flexibly set according to actual needs. In one embodiment, the similarity between the first video event and the candidate video event can be calculated based on the feature parameters of the first video event and the feature parameters of the second video event. For example, the similarity between the first video event and the candidate video event can be calculated based on at least one of the following parameters:

[0185] Similarity between the query terms corresponding to the first video event and the candidate video events;

[0186] Similarity between the title of the first video event and the title of the candidate video event;

[0187] Similarity between keywords in the title of the first video event and the titles of the candidate video events;

[0188] Similarity between the video title, video keywords, and video content vectors of the main video included in the first video event and the main video included in the candidate video event; the main video can be any video in the video event or the video ranked first in the video event;

[0189] Similarity between the subject of the video included in the first video event and the subject of the video included in the candidate video event;

[0190] The time interval between the maximum video included in the first video event and the maximum event included in the candidate video events;

[0191] The difference between the average publishing time interval of the first video event and the average publishing time interval of the candidate video events, where the average publishing time interval is the average publishing time interval of the videos included in the video event.

[0192] It should be noted that the parameters used to calculate the similarity between the first video event and the candidate video event include but are not limited to the above parameters.

[0193] Step S530 : Screen out candidate video events whose similarity exceeds a second threshold from the candidate video events, and use the screened candidate video events as second video events.

[0194] The specific value of the second threshold can be flexibly set according to actual needs.

[0195] After calculating the similarity between the first video event and each candidate video event, a candidate video event with a similarity greater than a second threshold is selected from the plurality of candidate video events as the second video event.

[0196] It should be noted that, in this embodiment, if a second video event corresponding to a first video event is obtained, it indicates that the first video event is an associated event of the second video event. For example, assuming that the second video event is the landing of a Mars rover on Mars, and the first video event is the patrolling of the surface of Mars by a Mars rover, then the patrolling of the surface of Mars by a Mars rover is a further development of the landing of a Mars rover on Mars; if the second video event corresponding to the first video event is not obtained, it indicates that the first video event is a new event.

[0197] In some embodiments, considering that the previously generated associated events have been aggregated into one event, the candidate video event with the highest similarity and a similarity exceeding a second threshold can be screened out from the candidate video events, and the screened candidate video event is used as the second video event.

[0198] In this embodiment, candidate video events that match the keywords of the first video event are obtained; the similarity between the first video event and the candidate video events is calculated; candidate video events whose similarity exceeds a second threshold are screened out from the candidate video events, and the screened out candidate video events are used as the second video event. In this way, after a new video event is generated, historical video events associated with the video event can be searched out, and the new video event can be aggregated with the associated historical video events subsequently, so that the user is aware of the latest developments of the event without the need for the user to follow up and discover the progress of the event himself, thereby improving the user experience.

[0199] In an exemplary embodiment, see Figure 11 , Figure 11 FIG. 1 is a flow chart showing an information processing method according to an exemplary embodiment of the present application. Figure 11 As shown, in Figure 1 After step S140 in the illustrated embodiment, the message processing method may further include steps S150 to S170, which are described in detail as follows:

[0200] Step S150: Add the video event containing the event association relationship to the set of events to be pushed, and obtain the information heat value of each video event in the multiple video events contained in the set of events to be pushed on different platforms.

[0201] It should be noted that the video events included in the set of events to be pushed are video events to be pushed to the user.

[0202] The information popularity value is a value that can reflect the popularity of information among users. It can be represented by parameters that can reflect the popularity of information among users, such as the number of clicks, searches, readings, reposts, comments, likes, and the number of discussion participants.

[0203] After a video event including an event association relationship is generated, in order to let the user know the video event, the video event needs to be pushed to the user. The video event including the event association relationship can be first added to a set of events to be pushed.

[0204] After adding the video events containing the event association relationship to the set of events to be pushed, for each video event in the set of events to be pushed, the information heat value of the video event on different platforms is obtained.

[0205] In one embodiment, a method for obtaining the information popularity value of each video event on different platforms includes but is not limited to at least one of the following two methods:

[0206] The first method is to crawl the information heat value of each video event from the platform.

[0207] Typically, the platform will count and display the information popularity value of video events, so the information popularity value of video events can be directly crawled from the platform. For example, for platforms other than the platform to which the information processing device belongs, this method can be used to obtain the information popularity value of each video event on that platform.

[0208] The second method: determine the information popularity value of each video event on different platforms based on the number of clicks on the query terms corresponding to each video event on different platforms.

[0209] Different video events correspond to different query terms. The information popularity value of each video event on different platforms can be determined based on the number of clicks on the query terms corresponding to each video event on different platforms. For example, for the platform to which the information processing device belongs, this method can be used to obtain the information popularity value of each video event on that platform.

[0210] The specific method for determining the information popularity value of each video event on different platforms based on the number of clicks on the query terms corresponding to each video event on different platforms can be flexibly set according to actual needs. For example, in one example, the formula for determining the information popularity value of each video event on different platforms based on the number of clicks on the query terms corresponding to each video event on different platforms can be as follows:

[0211]

[0212] Among them, Score b (e) is the information heat value of video event e on platform b, p b (q e ) is the number of clicks on query term q of video event e on platform b, and Q(e) is the set of query terms corresponding to video event e.

[0213] In some embodiments, the information popularity value of each video event on different platforms can be periodically obtained at preset time intervals. The preset time interval can be flexibly set according to actual needs, for example, it can be set to 1 hour. If the information popularity value of a video event on a certain platform is no longer updated, in order to improve the accuracy of the information popularity, the information popularity value can be decayed according to time. The specific decay method can be flexibly set according to actual needs. In one example, the decay method is as follows:

[0214] Score b (e) = Score b (e)*exp(-a*(hh′))

[0215] Among them, Score b(e) is the information heat value obtained after attenuating the information heat value of video event e on platform b, Score b (e) is the latest information heat value of the video event e on platform b, h is the current time point, h′ is the time point when the information heat value stops updating, and a is the time attenuation coefficient. Its specific value can be flexibly set according to actual needs. For example, it can be set to 0.1. In an example, assuming that the preset time interval is 1 hour, that is, the information heat value of each video event on different platforms is obtained every 1 hour, at 12 o'clock, the information heat value of a certain video event on a certain platform is 1000, at 13 o'clock, the information heat value of the video event on the platform is 2000, at 14 o'clock, the information heat value of the video event on the platform is 2000, at 15 o'clock, the information heat value of the video event on the platform is 2000, and the current time point is 15:35. It is found that the information heat value of the video event on the platform stops updating at 13 o'clock, then Score b (e) is 2000, and h′ is 15 points.

[0216] Step S160 , performing weighted summation on the information heat values ​​of each video event on different platforms to obtain the total heat value of each video event.

[0217] In this embodiment, different weights are set for different platforms. For each video event in the video events to be pushed, after obtaining the popularity value of the video event on different platforms, the popularity values ​​of the video event on different platforms can be weighted and summed according to the weights of different platforms to obtain the total popularity value of the video event.

[0218] In some embodiments, to improve the accuracy of the heat value, before performing weighted summation of the information heat values ​​of each video event on different platforms, the information heat values ​​of each video event on different platforms may be normalized, and then weighted summation is performed based on the normalized information heat values. The specific method of normalization can be flexibly set according to actual needs. For example, in one example, the normalization method is as follows:

[0219]

[0220] Among them, Score b (e) is the Score′ b (e) The value obtained after normalization, is the average popularity value of events on a certain preset platform (the preset platform may be the platform to which the information processing device belongs, or other platforms), is the average heat value of events on platform b.

[0221] In order to avoid the situation where the information heat value of a video event on a certain platform does not exist, resulting in unreasonable calculation, in this embodiment, a boundary value can be set for the platform. When the information heat value of the video event on the platform cannot be obtained, a value between the boundary value and the minimum heat value is randomly selected as the information heat value of the video event on the platform, where the minimum heat value can be a preset value or the heat value of the event with the least heat among multiple hot events on the platform, and the boundary value can be half of the minimum heat value.

[0222] Step S170 , sorting the multiple video events according to the obtained total popularity value, and pushing the multiple video events according to the sorting position.

[0223] After obtaining the total heat value of each of the multiple video events in the event set to be pushed, the multiple video events in the event set to be pushed are sorted according to the total heat value, and the multiple video events to be displayed are pushed according to the sorting position.

[0224] Among them, the specific process of pushing multiple video events according to the sorting position can be flexibly set according to actual needs. For example, it can be: displaying multiple video events on a popular list according to the sorting position; or determining the number of users corresponding to each video event in multiple video events according to the sorting position, and pushing the video event to the corresponding number of users. Among them, the higher the sorting position, the more corresponding users can be. For example, assuming that the number of users corresponding to a certain video event is 500, it will be pushed to 500 users.

[0225] In this embodiment, a video event containing an event association relationship is added to a set of events to be pushed, and the information heat value of each video event on different platforms in the multiple video events contained in the set of events to be pushed is obtained; the information heat value of each video event on different platforms is weightedly summed to obtain the total heat value of each video event; the multiple video events are sorted according to the obtained total heat value, and the multiple video events are pushed according to the sorting position. In this way, after a video event containing an event association relationship is generated, the video event can be pushed according to the information heat value of the video event, so that the user is aware of the event association relationship, and the user does not need to continue to follow up and discover the latest progress of the event on his own, thereby improving the user experience.

[0226] The following is a detailed description of a specific application scenario of the embodiment of this application. Figure 12 , Figure 12 Schematic diagram of an implementation environment involved in this application, such as Figure 12As shown in the figure, the implementation environment includes: content consumption end, content production end, content distribution export service, recommendation distribution system, content database, manual review system, scheduling center service, upstream and downstream content interface server, statistics server, duplicate removal service, video event discovery service, video event generation service, video event aggregation service, video event thematic database, and video event thematic interface service. The functions of each module are as follows:

[0227] Content production end: the source of video and other content, used to connect to upstream and downstream content servers through a mobile terminal or back-end interface (for example, an API system, where API stands for Application Programming Interface), and upload and upload videos and other content through upstream and downstream content servers; content production ends include but are not limited to PGC, UGC, MCN content producers, etc.

[0228] Content consumption end: (1) As a consumer, it connects to the upstream and downstream content interface servers and obtains index information and content from the content database through the upstream and downstream content interface servers. The obtained content includes content recommended by the recommendation distribution system, content of subscribed topics, and content obtained through active searches. (2) The content consumption end can also report the operation data with user permission to the statistics server, for example, the query terms entered by the user, click data on search results, content sharing data, collection operations, forwarding operations, like operations, video upload operations, etc. to the statistics server. The content consumption end can browse data through the feed stream, enter various content channels to browse content and subscribe to corresponding topic content, and view the context of the entire video event through the entrance of the video event topic. In addition, the content consumption end can also upload videos and other content as a content production end.

[0229] The uplink and downlink content interface servers are connected to the content production end, receive videos and other content and content metadata from the content production end, store the content and content metadata in the content database, and submit the content to the scheduling execution server. The content metadata includes but is not limited to the size of the video file, cover image link, title, release time, author and other information. It should be noted that in this application, the videos, video metadata, operation data and other user-related data involved, when the above embodiments of this application are applied to specific products or technologies, are all obtained with the user's permission or consent, and the extraction, use and processing of the relevant data comply with local safety standards and local laws and regulations.

[0230] Content database: The core database of content. The metadata of the content published by content producers is stored in the content database, such as the size of the video file, cover image link, bit rate, file format, title, release time, author, whether it is original, whether it is the first release, etc. The content database also stores the classification of content during the manual review process, including categories and tags. (1) The content database is connected to the manual review system. The manual review system will read the original content in the content database. At the same time, the manual review system will return the manual review results and status of the original content to the content database. (2) The content database is connected to the scheduling center service. The scheduling center service mainly processes content through machine processing and manual review. The core of machine processing here is to call the deduplication service. The deduplication results will be written to the content database. Completely duplicate content will not be manually processed again. (3) The content database is connected to the video event discovery service. The video event discovery service obtains data from the content database.

[0231] Scheduling center service: responsible for the entire scheduling process of content flow, controlling the scheduling order and priority, among which, the scheduling center service can receive the stored content through the upstream and downstream content interface servers, and then obtain the metadata of the content from the content database; the scheduling deduplication service deduplicates the content and filters out duplicate content. For content that does not meet the duplicate filtering, it can output the content similarity and similarity relationship chain for the recommendation distribution system to use; the scheduling manual review system manually reviews the filtered content, and the content that passes the manual review system can be provided to the content consumption end through the recommendation distribution system and the content distribution export service, for example, through the recommendation engine, search engine or display page to the content consumption end; the scheduling center service can also communicate with the video event special interface service to obtain the generated video events containing event association relationships; the scheduling center service can also determine whether the content requires manual review or is directly distributed to the content consumption end through the content distribution export based on the configuration information.

[0232] Manual review system: It is necessary to obtain the original content in the content database. The manual review system can be a system developed based on a web database, which uses manual labor to conduct preliminary filtering to determine whether the content complies with regulations. During the filtering process, the machine algorithm can assist with low-quality and problem prompts to improve manual efficiency.

[0233] Video event discovery service: Obtain information such as popular lists and hot topics from the Internet to generate event topics. It can also obtain user-permitted operation data from the statistics server and obtain information such as popular lists based on the operation data.

[0234] Video event generation service: Generates the first video event based on the event topic input by the video event discovery service.

[0235] Video event aggregation service: Analyzes the first video event and the second video event to obtain the event association relationship.

[0236] Video event thematic database: saves the event association relationship generated by the video event aggregation service, aggregates the first video event and the second video event according to the event association relationship, and obtains the video event containing the event association relationship; provides a data source for the video event interface service.

[0237] Video event service interface service: (1) read the content of video event thematic data, calculate the heat of video events, and sort video events; (2) communicate with the dispatch center service.

[0238] Deduplication service: mainly used for massive deduplication to avoid duplicate content.

[0239] Statistics server: accepts user-authorized operation data uploaded by content consumers, and provides data source support and services for subsequent video event discovery and statistical analysis.

[0240] Recommendation distribution system: connects to the content distribution export service, obtains content from the content database, and sends it to the content consumption end through the content distribution export service to push content to users.

[0241] Content distribution export service: connects to the recommendation distribution system to distribute content to content consumption terminals.

[0242] See also Figure 13 , Figure 13 FIG. 1 is a block diagram of an information processing device according to an exemplary embodiment of the present application. Figure 13 As shown, the device includes:

[0243] A generating module 1301 is configured to obtain information associated with the video and generate an event topic based on the obtained information;

[0244] An acquisition module 1302 is configured to generate a first video event corresponding to an event theme and acquire a second video event associated with the first video event;

[0245] An analysis module 1303 is configured to analyze the first video event and the second video event to obtain an event association relationship;

[0246] The aggregation module 1304 is configured to aggregate the first video event and the second video event according to the event association relationship to obtain a video event containing the event association relationship.

[0247] In another exemplary embodiment, under the condition that the information associated with the video includes a video title, the generating module 1301 includes:

[0248] The segmentation module is configured to segment the video title to obtain a first candidate topic.

[0249] The candidate topic generating module is configured to cluster the video titles to obtain video title clusters, and generate second candidate topics corresponding to the video title clusters.

[0250] The first clustering module is configured to cluster the first candidate topic and the second candidate topic to obtain a candidate topic cluster.

[0251] The first topic generation module is configured to determine the event topic according to the cluster center of the candidate topic cluster.

[0252] In another exemplary embodiment, under the condition that the information associated with the video includes videos uploaded within a preset time period, the generating module 1301 includes:

[0253] The second clustering module is configured to cluster the videos uploaded within a preset time period to obtain multiple video clusters.

[0254] The first screening module is configured to screen out a target video cluster having a number of video items greater than a preset value from the multiple video clusters.

[0255] The second topic generation module is configured to determine the event topic according to the cluster center of the target video cluster.

[0256] In another exemplary embodiment, the obtaining module 1302 includes:

[0257] The term acquisition module is configured to obtain the query term corresponding to the event topic.

[0258] The recall module is configured to recall candidate videos matching the query term from the video content library.

[0259] The second screening module is configured to calculate the relevance between the candidate videos and the query term, and screen out candidate videos whose relevance exceeds a first threshold from the candidate videos to obtain the target video.

[0260] The event generation module is configured to generate a first video event according to an event theme and a target video.

[0261] In another exemplary embodiment, when there are multiple target videos, the event generation module includes:

[0262] The quality assessment module is configured to cluster multiple target videos to obtain multiple target video clusters, and perform quality assessment on the multiple target videos and the multiple target video clusters respectively to obtain first quality values ​​corresponding to each of the multiple target videos and second quality values ​​corresponding to each of the multiple target video clusters.

[0263] The representative video determination module is configured to use the target video with the highest first quality value in each target video cluster as the representative video to obtain multiple representative videos.

[0264] The sorting module is configured to determine the third quality value corresponding to each of the multiple representative videos based on the first quality value of the representative video and the second quality value of the target video cluster to which the representative video belongs, and sort the multiple representative videos in descending order of the third quality values.

[0265] The video event generation module is configured to generate a first video event corresponding to an event theme and including a representative video with a specified ranking.

[0266] In another exemplary embodiment, the obtaining module 1302 includes:

[0267] The search module is configured to obtain candidate video events that match the keyword of the first video event.

[0268] The calculation module is configured to calculate the similarity between the first video event and the candidate video event.

[0269] The third screening module is configured to screen out candidate video events whose similarity exceeds a second threshold from the candidate video events, and use the screened candidate video events as the second video events.

[0270] In another exemplary embodiment, the apparatus further comprises:

[0271] The heat value acquisition module is configured to add the video event containing the event association relationship to the event set to be pushed, and obtain the information heat value of each video event in the multiple video events contained in the event set to be pushed on different platforms.

[0272] The summing module is configured to perform weighted summation on the information heat values ​​of each video event on different platforms to obtain the total heat value of each video event.

[0273] The push module is configured to sort the multiple video events according to the obtained total heat value, and push the multiple video events according to the sorting position.

[0274] It should be noted that the information processing device provided in the above embodiment and the information processing method provided in the above embodiment belong to the same concept, and the specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here.

[0275] An embodiment of the present application also provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, which, when executed by one or more processors, enables the electronic device to implement the information processing methods provided in the above-mentioned embodiments.

[0276] Figure 14 A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown.

[0277] It should be noted that Figure 14 The computer system 1400 of the electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.

[0278] like Figure 14 As shown, computer system 1400 includes a central processing unit (CPU) 1401, which can perform various appropriate actions and processes according to programs stored in read-only memory (ROM) 1402 or programs loaded from storage unit 1408 into random access memory (RAM) 1403, such as executing the methods described in the above embodiments. Various programs and data required for system operation are also stored in RAM 1403. CPU 1401, ROM 1402, and RAM 1403 are connected to each other via bus 1404. Input / output (I / O) interface 1405 is also connected to bus 1404.

[0279] The following components are connected to the I / O interface 1405: an input section 1406 including a keyboard, a mouse, and the like; an output section 1407 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 1408 including a hard disk; and a communication section 1409 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1409 performs communication processing via a network such as the Internet. A drive 1410 is also connected to the I / O interface 1405 as needed. Removable media 1411, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1410 as needed, so that computer programs read from the removable media can be installed in the storage section 1408 as needed.

[0280] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1409, and / or installed from a removable medium 1411. When the computer program is executed by the central processing unit (CPU) 1401, the various functions defined in the system of the present application are executed.

[0281] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable computer program. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0282] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0283] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.

[0284] Another aspect of the present application provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor of an electronic device, the electronic device implements the aforementioned method. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device.

[0285] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions that, when executed by a processor, implement the methods provided in the above embodiments. The computer instructions may be stored in a computer-readable storage medium; a processor of an electronic device may read the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, causing the electronic device to perform the methods provided in the above embodiments.

[0286] The above content is only a preferred exemplary embodiment of the present application and is not intended to limit the implementation scheme of the present application. Ordinary technicians in this field can easily make corresponding changes or modifications based on the main ideas and spirit of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection required by the claims.

Claims

1. An information processing method, characterized in that: The method comprises: Acquire information associated with the video and generate an event topic based on the acquired information; Generate a first video event corresponding to the event theme, and obtain a second video event associated with the first video event; Analyzing the first video event and the second video event to obtain an event association relationship; aggregating the first video event and the second video event according to the event association relationship to obtain a video event containing the event association relationship; Generating the first video event corresponding to the event theme includes: obtaining a query term corresponding to the event theme; recalling candidate videos matching the query term from a video content library; calculating the relevance between the candidate videos and the query term, and screening out candidate videos whose relevance exceeds a first threshold from the candidate videos to obtain a target video; generating the first video event based on the event theme and the target video; The obtaining of the second video event associated with the first video event includes: obtaining candidate video events that match the keywords of the first video event; calculating the similarity between the first video event and the candidate video events; wherein the similarity includes the similarity calculated based on the query terms corresponding to the first video event and the candidate video events respectively; screening out candidate video events whose similarity exceeds a second threshold from the candidate video events, and using the screened-out candidate video events as the second video event.

2. The method according to claim 1, wherein The information associated with the video includes the video title; and generating the event theme based on the acquired information includes: Segmenting the video title to obtain a first candidate topic; Clustering the video titles to obtain video title clusters, and generating second candidate topics corresponding to the video title clusters; Clustering the first candidate topic and the second candidate topic to obtain a candidate topic cluster; The event topic is determined according to the cluster center of the candidate topic cluster.

3. The method according to claim 1, wherein The information associated with the video includes videos uploaded within a preset time period; and generating an event topic based on the acquired information includes: Clustering the videos uploaded within the preset time period to obtain multiple video clusters; Filtering out a target video cluster having a number of video items greater than a preset value from the plurality of video clusters; An event theme is determined according to the cluster center of the target video cluster.

4. The method according to claim 1, wherein There are multiple target videos; generating the first video event according to the event theme and the target videos includes: Clustering the multiple target videos to obtain multiple target video clusters, and performing quality assessment on the multiple target videos and the multiple target video clusters to obtain first quality values ​​corresponding to each of the multiple target videos and second quality values ​​corresponding to each of the multiple target video clusters; The target video with the highest first quality value in each target video cluster is used as a representative video to obtain multiple representative videos; determining a third quality value corresponding to each of the plurality of representative videos according to the first quality value of the representative video and the second quality value of the target video cluster to which the representative video belongs, and sorting the plurality of representative videos in descending order of the third quality values; A first video event corresponding to the event theme and including a representative video with a specified ranking is generated.

5. The method according to claim 1, wherein After aggregating the first video event and the second video event according to the event association relationship to obtain a video event containing the event association relationship, the method further includes: Adding the video event containing the event association relationship to a set of events to be pushed, and obtaining the information popularity value of each video event in the plurality of video events contained in the set of events to be pushed on different platforms; Performing weighted summation on the information heat values ​​of each video event on different platforms to obtain the total heat value of each video event; The multiple video events are sorted according to the obtained total heat value, and the multiple video events are pushed according to the sorting position.

6. An information processing device, characterized in that The device comprises: a generation module configured to obtain information associated with the video and generate an event topic based on the obtained information; an acquisition module configured to generate a first video event corresponding to the event theme and acquire a second video event associated with the first video event; an analysis module configured to analyze the first video event and the second video event to obtain an event association relationship; an aggregation module configured to aggregate the first video event and the second video event according to the event association relationship to obtain a video event containing the event association relationship; Generating the first video event corresponding to the event theme includes: obtaining a query term corresponding to the event theme; recalling candidate videos matching the query term from a video content library; calculating the relevance between the candidate videos and the query term, and screening out candidate videos whose relevance exceeds a first threshold from the candidate videos to obtain a target video; generating the first video event based on the event theme and the target video; The obtaining of the second video event associated with the first video event includes: obtaining candidate video events that match the keywords of the first video event; calculating the similarity between the first video event and the candidate video events; wherein the similarity includes the similarity calculated based on the query terms corresponding to the first video event and the candidate video events respectively; screening out candidate video events whose similarity exceeds a second threshold from the candidate video events, and using the screened-out candidate video events as the second video event.

7. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, enables the electronic device to implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor of a computer, the computer is caused to execute the method according to any one of claims 1 to 5.

9. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Information pushing method based on machine learning and related device

    CN111368063A