An information prompting method and device, electronic equipment and storage medium

By acquiring feature information from multimedia content and user history behavior, and using machine learning models to predict the timing of information prompts and filter target objects, the ineffectiveness of information prompts in existing technologies is solved, and personalized and efficient information prompts are achieved.

CN116226416BActive Publication Date: 2026-04-28TENCENT TECH WUHAN
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECH WUHAN
Filing Date
2021-12-06
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, video prompting methods cannot provide personalized information prompts, resulting in invalid prompts and a decline in user experience, especially for users with different viewing habits.

Method used

By acquiring content feature information of multimedia content and statistical feature information of user historical behavior, machine learning models are used to predict the information prompt time point, and target objects are selected for information prompts based on content embedding vectors and object embedding vectors.

Benefits of technology

Personalized information prompts were implemented, which improved the reach and relevance of the prompts, reduced invalid prompts, and enhanced the user experience and effectiveness of the information prompts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116226416B_ABST
    Figure CN116226416B_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field, and in particular to an information prompting method and device, electronic equipment and a storage medium, to improve the pertinence and reach rate of information prompting. The method comprises the following steps: obtaining content feature information of each multimedia content; obtaining statistical feature information of each multimedia content based on historical behaviors of each object to each multimedia content; obtaining information prompting time points of each multimedia content based on the obtained content feature information and statistical feature information; screening out target objects matched with each multimedia content based on content embedding vectors of the multimedia content and object embedding vectors of the associated objects; and prompting target information to the corresponding target objects when each multimedia content plays to the corresponding information prompting time point. The application realizes personalized intelligent information prompting by object matching and information prompting time point prediction, and effectively improves the pertinence and reach rate of information prompting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly to the field of artificial intelligence technology, and provides an information prompting method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of smart terminal devices, people can use various applications (APPs) installed on these devices to fulfill various needs in their lives, such as watching videos. While watching a video, interactive behavior can be initiated by providing information prompts to the user.

[0003] In related technologies, taking video scenarios as an example, there are two main ways to mobilize the audience to interact:

[0004] Method 1: Content creators on the platform provide explicit guidance within the video, such as through relevant prompts in the foreground and background or in the narration. Method 2: Content creators can set a pop-up notification time point each time they publish content, adding a corresponding pop-up notification to the video's progress bar.

[0005] In summary, both of the above methods require manual setup and have fixed prompt times, resulting in all viewers receiving prompts during the video playback. This can easily disturb those already familiar with the interaction rules. Furthermore, users with shorter average viewing times may finish watching before the set time, rendering the prompts ineffective. Therefore, developing a personalized, intelligent information prompting method to provide more targeted information and effectively improve the reach of prompts is a pressing issue. Summary of the Invention

[0006] This application provides an information prompting method, apparatus, electronic device, and storage medium to improve the targeting of information prompts and the reach of prompting information.

[0007] An information prompting method provided in this application includes:

[0008] The content feature information of each multimedia content is obtained separately, and the statistical feature information of each multimedia content is obtained separately based on the historical behavior of each object towards each multimedia content.

[0009] Based on the obtained content feature information and statistical feature information, the information prompt time points of each multimedia content are obtained respectively;

[0010] Perform the following operations for each multimedia content: Based on the content embedding vector of a multimedia content and the object embedding vector of at least one of the objects associated with the multimedia content, filter out at least one matching target object;

[0011] When each multimedia content reaches the corresponding information prompt time point, a target information prompt is given to at least one target object.

[0012] An information prompting device provided in this application embodiment includes:

[0013] The feature acquisition unit is used to acquire the content feature information of each multimedia content, and to acquire the statistical feature information of each multimedia content based on the historical behavior of each object towards each multimedia content.

[0014] The time determination unit is used to obtain the information prompt time point of each multimedia content based on the obtained content feature information and statistical feature information;

[0015] The object filtering unit is configured to perform the following operations for each multimedia content: based on the content embedding vector of a multimedia content and the object embedding vector of at least one of the objects associated with the multimedia content, filter out at least one matching target object.

[0016] The information prompting unit is used to provide target information prompts to at least one target object when each multimedia content is played to the corresponding information prompting time point.

[0017] Optionally, the time determination unit:

[0018] Based on the content feature information and statistical feature information of each multimedia content, the predicted playback completion rate of each multimedia content is obtained respectively.

[0019] Based on the predicted playback completion rate and playback duration of each multimedia content, the respective information prompt time point is determined.

[0020] Optionally, the device further includes:

[0021] The content filtering unit is used to filter at least one target multimedia content from the multimedia content based on the obtained content feature information and statistical feature information before the information prompting unit provides target information prompts to at least one target object when each multimedia content is played to the corresponding information prompting time point. Each target multimedia content is a multimedia content in which the probability of the object performing the target operation meets the expected standard.

[0022] The information prompting unit is specifically used for:

[0023] When the multimedia content of each target is played to the corresponding information prompt time point, a target information prompt is given to at least one target object.

[0024] Optionally, the content feature information includes first content sub-feature information determined based on the title of the multimedia content, second content sub-feature information determined based on at least one level of classification information of the multimedia content, and third content sub-feature information determined based on the publication duration of the multimedia content; the statistical feature information includes first statistical sub-feature information determined based on the number of times each object performs various specified operations on the multimedia content.

[0025] The content filtering unit is specifically used for:

[0026] Based on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information of each multimedia content, at least one target multimedia content is selected from the multimedia content.

[0027] Optionally, the content filtering unit is specifically used for:

[0028] For each multimedia content, the following operations are performed: input the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information of a multimedia content into a trained classification model; based on the classification model, perform low-order feature cross and high-order feature cross on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information, respectively, to obtain the first predicted probability of the multimedia content output by the classification model;

[0029] The multimedia content whose first predicted probability is greater than a preset probability threshold is selected as the target multimedia content.

[0030] Optionally, the content feature information includes second content sub-feature information determined based on at least one level of classification information of multimedia content, and third content sub-feature information determined based on the publication duration of multimedia content; the statistical feature information includes second statistical sub-feature information determined based on the browsing duration of each object for the multimedia content;

[0031] The time determination unit is specifically used for:

[0032] Based on the second content sub-feature information, the third content sub-feature information, and the second statistical sub-feature information of each multimedia content, the predicted playback completion rate of each multimedia content is predicted.

[0033] Optionally, the time determination unit:

[0034] Perform the following operations on each of the aforementioned multimedia contents:

[0035] Input the second content sub-feature information, the third content sub-feature information, and the second statistical sub-feature information of a multimedia content into the trained classification model respectively;

[0036] Based on the classification model, low-order feature cross and high-order feature cross are performed on the second content sub-feature information, the third content sub-feature information and the second statistical feature information, respectively, to obtain the second predicted probability corresponding to the multimedia content output by the classification model, and the second predicted probability is used as the predicted playback completion rate of the multimedia content.

[0037] Optionally, the object filtering unit is specifically used for:

[0038] Based on the content embedding vector of the multimedia content and the object embedding vector of the associated at least one of the objects, the target content embedding vector of the multimedia content is determined.

[0039] Based on the target content embedding vector and the object embedding vector of each object, the matching degree between the multimedia content and each object is determined respectively.

[0040] Based on the obtained matching scores, at least one target object matching the multimedia content is selected.

[0041] Optionally, the object filtering unit is specifically used for:

[0042] The content embedding vector of the multimedia content and the object embedding vector of at least one associated object are respectively input into the trained deep model.

[0043] Based on the deep model, the dimensionality reduction of the content embedding vector and the object embedding vector of at least one associated object is performed to obtain the target content embedding vector of the multimedia content.

[0044] Optionally, the content embedding vector of a multimedia content includes at least one of the following:

[0045] Based on the trained classification model, feature extraction is performed on the content feature information and statistical feature information of the multimedia content to obtain the first content sub-embedding vector;

[0046] The second content sub-embedding vector is obtained by embedding the multimedia content based on at least one of computer vision technology and natural language technology.

[0047] Optionally, the object embedding vector of at least one of the associated objects includes at least one of the following:

[0048] The first object sub-embedding vector is determined based on the average of the embedding vectors of each object browsing the multimedia content within a certain historical period.

[0049] The second object sub-embedding vector is determined by averaging the embedding vectors of each object that performs a target operation on at least one multimedia content within a certain historical period.

[0050] The third object sub-embedding vector is determined by averaging the embedding vectors of each object that performed a specified operation on the multimedia content within a certain historical period but did not perform the target operation.

[0051] An electronic device provided in this application includes a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of any of the above-described information prompting methods.

[0052] This application provides a computer-readable storage medium including a computer program. When the computer program is run on an electronic device, the computer program is used to cause the electronic device to perform the steps of any of the above-described information prompting methods.

[0053] This application provides a computer program product, which includes a computer program stored in a computer-readable storage medium. When the processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any of the above-described information prompting methods.

[0054] The beneficial effects of this application are as follows:

[0055] This application provides an information prompting method, apparatus, electronic device, and storage medium. Because this application predicts the information prompting time point for each multimedia content based on its own content feature information and statistical feature information determined by the historical behavior of each object, it achieves personalized setting of information prompting time points for different multimedia content. By predicting the information prompting time point based on the historical behavior of each object towards the multimedia content, the target information prompt is delivered at the corresponding time point. This time point can represent the estimated viewing time of the multimedia content, effectively reducing the sending of invalid prompts and improving the reach rate of prompting information among users. Furthermore, based on the content embedding vector of each multimedia content and the object embedding vector of related objects, the target objects matching each multimedia content are filtered, selecting the most suitable group of people who need target information prompts for each multimedia content, achieving personalized and more targeted prompts. In addition, the information prompting method in this application does not require manual setting and significantly improves the relevant task indicators of target information prompts while achieving intelligent prediction.

[0056] Other features and advantages of this application will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0057] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0058] Figure 1 This is a schematic diagram of a coin-operated guidance process for a content platform in the related technologies of this application;

[0059] Figure 2 This is an optional schematic diagram of an application scenario in an embodiment of this application;

[0060] Figure 3 This is a mapping diagram of an application scenario in an embodiment of this application;

[0061] Figure 4 This is a flowchart illustrating the first information prompting method in the embodiments of this application;

[0062] Figure 5 This is a schematic diagram of a target prompt message in an embodiment of this application;

[0063] Figure 6AThis is a flowchart illustrating the second information prompting method in an embodiment of this application;

[0064] Figure 6B This is a system framework diagram corresponding to the second information prompting method in the embodiments of this application;

[0065] Figure 7A This is a flowchart illustrating the third information prompting method in the embodiments of this application;

[0066] Figure 7B This is a system framework diagram corresponding to the third information prompting method in the embodiments of this application;

[0067] Figure 8 This is a schematic diagram of the structure of a classification model in an embodiment of this application;

[0068] Figure 9 This is a schematic diagram illustrating the feature information used in predicting high-potential coin-operated content in an embodiment of this application;

[0069] Figure 10 This is a schematic diagram of the structure of an FM in an embodiment of this application;

[0070] Figure 11 This is a schematic diagram of the structure of a dense layer in an embodiment of this application;

[0071] Figure 12 This is a schematic diagram of the structure of a Deep section in one embodiment of this application;

[0072] Figure 13 This is a schematic diagram of the structure of a depth model in an embodiment of this application;

[0073] Figure 14 This is a schematic diagram of the feature information applied to a prediction information prompt time point in an embodiment of this application;

[0074] Figure 15 This is a schematic diagram illustrating the specific implementation process of a video content coin-operated information prompt in an embodiment of this application;

[0075] Figure 16 This is a schematic diagram of the composition structure of an information prompting device according to an embodiment of this application;

[0076] Figure 17 This is a schematic diagram of the composition structure of an electronic device according to an embodiment of this application;

[0077] Figure 18 This is a schematic diagram of the composition structure of another electronic device using an embodiment of this application. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0079] The following describes some of the concepts involved in the embodiments of this application.

[0080] Multimedia content: refers to digitally transmitted resources, including but not limited to: articles, videos, audio, news information, etc.

[0081] Content feature information: used to describe the inherent attributes of multimedia content, including but not limited to at least one of the following: first content sub-feature information determined based on the title of the multimedia content, such as tags extracted from video titles, which can be called content tags; second content sub-feature information determined based on at least one level of classification information of the multimedia content, such as video first-level classification (game video), second-level classification (shooting game), third-level classification (peaceful game), etc., where the third-level classification is a further refinement of the second-level classification, and similarly, the second-level classification is a further refinement of the first-level classification; third content sub-feature information determined based on the release duration of the multimedia content, such as the time difference between the video's release time and the present time.

[0082] Content embedding vector: A vector obtained by embedding multimedia content, including but not limited to at least one of the following: a first content sub-embedding vector obtained by extracting the content feature information and statistical feature information of a multimedia content based on a trained classification model; a second content sub-embedding vector obtained by embedding a multimedia content based on at least one of computer vision technology and natural language technology.

[0083] Statistical feature information refers to features obtained based on the analysis of users' historical behavior towards multimedia content, including but not limited to at least one of the following: a first statistical sub-feature information determined based on the number of times each user performs various specified operations on multimedia content; and a second statistical sub-feature information determined based on the browsing time of each user on multimedia content. These specified operations include, but are not limited to: liking, donating coins, forwarding, and commenting.

[0084] Object embedding vector: refers to the vector obtained by embedding a user account, including but not limited to at least one of the following: a first object sub-embedding vector determined by averaging the embedding vectors of each object that viewed a multimedia content within a certain historical period; a second object sub-embedding vector determined by averaging the embedding vectors of each object that performed a target operation on at least one multimedia content within a certain historical period; and a third object sub-embedding vector determined by averaging the embedding vectors of each object that performed a specified operation on a multimedia content within a certain historical period but did not perform the target operation.

[0085] The embodiments of this application relate to artificial intelligence (AI), natural language processing (NLP), and machine learning (ML) technologies, and are designed based on computer vision technology and machine learning in artificial intelligence.

[0086] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence.

[0087] Artificial intelligence (AI) studies the design principles and implementation methods of various intelligent machines, enabling them to perceive, reason, and make decisions. AI technology mainly includes computer vision, natural language processing, and machine learning / deep learning. With the research and advancement of AI technology, it is being researched and applied in multiple fields, such as smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, autonomous driving, robotics, and smart healthcare. It is believed that with further technological development, AI will be applied in even more fields and play an increasingly important role.

[0088] NLP is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. Natural Language Processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close connection with linguistic research. Natural Language Processing techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0089] Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Compared to data mining, which focuses on finding patterns in large datasets, machine learning emphasizes algorithm design, enabling computers to automatically "learn" patterns from data and use these patterns to predict unknown data.

[0090] Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. The classification model and deep model in the embodiments of this application are trained using machine learning or deep learning techniques. Based on the classification model in the embodiments of this application, the prediction of the information prompt time point corresponding to each multimedia content and the filtering of target multimedia information can be achieved. Based on the deep model in the embodiments of this application, target objects can be matched for each multimedia content.

[0091] Furthermore, the embedding vectors in this application embodiment can be learned not only based on classification models, but also based on computer vision and natural language processing technologies. For example, the second content embedding sub-vector in the content embedding vector in this application embodiment refers to an embedding learned from multimedia content using at least one of computer vision and natural language processing technologies.

[0092] Based on the above method, when the multimedia content is played to the corresponding information prompt time point, targeted information prompts can be sent to at least one relevant target object, thereby achieving intelligent and personalized information prompts.

[0093] The design concept of the embodiments of this application is briefly introduced below:

[0094] With the rapid development of network technology, people's demand for the network is reflected in every corner of life. With the development of smart terminal devices, people can use various applications installed on smart terminal devices to meet various needs in life, such as watching videos.

[0095] As the penetration rate of both short and long videos among users gradually increases, video-related features such as likes, comments, shares, and bullet comments are closely watched as important indicators for evaluating the information flow ecosystem. However, these features can only statistically indicate the popularity of a video and cannot quantify the revenue generated by positive user feedback for content creators. Therefore, some platforms have added a coin-giving function to short videos and text / image content, where users can earn valuable resource rewards by checking in on the platform and extending their time spent on the platform.

[0096] However, users often become accustomed to using a product because most current news feed products only offer basic functions such as liking, commenting, forwarding, and sharing. Few users notice the coin-giving function in the news feed, or even if they do, they don't consciously use it. Therefore, providing pop-up prompts at appropriate times is particularly important.

[0097] In related technologies, there are two main ways to mobilize user interaction, such as Figure 1 As shown: First, content platform account owners can add coin-donation prompts to the video content, such as adding relevant subtitles in the pre- and post-video clips, or content creators can directly encourage users to donate coins in their narration; second, the system can add coin-donation pop-up prompts, which are added to the video's progress bar.

[0098] The existing technologies have the following drawbacks: 1) Indiscriminate pop-ups. Pop-up prompts are applied to every video, but some videos are of insufficient quality to entice users to donate coins, impacting user experience and yielding little benefit. 2) Lack of prior knowledge support. All users watching the video will be prompted by pop-ups, disturbing those already familiar with the coin-donation rules, thus also affecting user experience. 3) Fixed pop-up timing. Different users have different video viewing habits. For users with shorter average viewing times, they may finish watching before the set time, rendering the pop-up prompt ineffective. Therefore, a personalized, intelligent pop-up prompt method is needed to increase coin donations while ensuring a good user experience.

[0099] In view of this, this application provides a smart pop-up method for video coin-insertion prompts based on deep learning networks. It uses artificial intelligence to match the video with the people who most need the coin-insertion pop-up prompts and personalizes the pop-up timing for different users to match their habits.

[0100] The advantages of this method are: 1) Intelligent pop-ups. It avoids the low payout ratio of pop-up notifications in indiscriminate distribution tasks because intelligent pop-ups utilize basic video information and users' historical video browsing behavior, continuously learning in a deep learning network to select the most suitable audience for each video requiring a coin-donation pop-up notification; 2) Personalized pop-up timing. Based on users' historical browsing habits of other videos, it predicts the user's viewing time for the current video and pops up a coin-donation pop-up notification at the corresponding time, increasing the reach of the coin-donation pop-up notification to users; 3) Improved user experience. The personalized user-video matching method increases the popularity of coin-donation operations among new users while reducing the acceptance rate of pop-up notifications among experienced users, increasing the number of coins donated to the platform while avoiding excessive disturbance to users; 4) Reduced manpower requirements. Related technologies require extensive experience to set parameters for coin-donation pop-up notifications, which is not user-friendly for novice content creators or operators. The method in this application not only reduces manpower requirements but also significantly improves the final coin-donation metrics through intelligent prediction.

[0101] The preferred embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments and features in the embodiments of this application can be combined with each other without conflict.

[0102] like Figure 2 The diagram shown is an application scenario illustration of an embodiment of this application. The application scenario diagram includes two terminal devices 210 and one server 220.

[0103] In this embodiment, the terminal device 210 includes, but is not limited to, mobile phones, tablets, laptops, desktop computers, e-book readers, smart voice interaction devices, smart home appliances, and in-vehicle terminals. The terminal device may have a client installed for displaying multimedia content-related information. This client can be software (e.g., a browser, short video software), a webpage, or a mini-program. The server 220 is the backend server corresponding to the software, webpage, or mini-program, or a server specifically used for displaying multimedia content-related information; this application does not impose specific limitations. The server 220 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0104] It should be noted that the information prompting method in this embodiment can be executed by an electronic device, which can be a server 220 or a terminal device 210. That is, the method can be executed by the server 220 or the terminal device 210 alone, or by both the server 220 and the terminal device 210. For example, when executed by both the server 220 and the terminal device 210, the server 220 can obtain the content feature information and statistical feature information of the multimedia content, as well as the content embedding vector of the multimedia content and the object embedding vector of the associated object. Then, based on the above feature information and embedding vector, prediction is made to obtain the information prompting time point and the matching target object for each multimedia content, and these prediction results are fed back to the terminal device 210. When the terminal device 210 detects that the target object is playing multimedia content to the corresponding information prompting time point, it will provide the target information prompt to the corresponding target object. In addition, the method when executed by the server 220 or the terminal device 210 alone is similar, and will not be repeated here. The following mainly uses the execution by the server 220 alone as an example for illustration.

[0105] It should be noted that in the specific implementation of this application, user information, such as object historical behavior, feature information and other related data, is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0106] In one alternative implementation, the terminal device 210 and the server 220 can communicate via a communication network.

[0107] In one alternative implementation, the communication network is a wired network or a wireless network.

[0108] It should be noted that, Figure 2 The examples shown are merely illustrative; in reality, the number of terminal devices and servers is unlimited and is not specifically limited in the embodiments of this application.

[0109] In this embodiment of the application, when there are multiple servers, the multiple servers can form a blockchain, and the servers are nodes on the blockchain; as disclosed in the information prompting method of this application, the content feature information, statistical feature information, content embedding vector, object embedding vector and other data related to the multimedia content involved can all be stored on the blockchain, and no specific limitation is made here.

[0110] In one optional implementation, the information prompting method proposed in this application can be applied to the coin-insertion pop-up prompt of short videos in short video software, and can also be applied to the interactive function prompts of all information flow products, such as text and audio-visual products.

[0111] See Figure 3 As shown, it is a mapping diagram of an application scenario in an embodiment of this application. Based on Figure 3 As can be seen, the multimedia content in this application embodiment includes, but is not limited to: articles, videos, audio, and information streams; the target prompt information includes, but is not limited to: like prompt information, coin donation prompt information, comment prompt information, and forwarding prompt information.

[0112] The "like" prompt message prompts users to like the corresponding multimedia content, while the "coin" prompt message prompts users to donate coins to the corresponding multimedia content. Other prompt messages are similar and will not be repeated here.

[0113] In this embodiment of the application, personalized intelligent information prompts can be used to guide appropriate users, thereby achieving good interaction between users, content, and even the platform, and realizing a positive cycle of the ecosystem.

[0114] It should be noted that the following text mainly uses the coin insertion prompt during video playback as an example for illustration. Other information prompts follow a similar principle, and repeated points will not be elaborated upon.

[0115] Furthermore, the embodiments of this application can be applied to various scenarios, including not only the information prompt scenarios listed above, but also scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0116] The following describes the information prompting method provided by exemplary embodiments of this application in conjunction with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this respect.

[0117] See Figure 4 The diagram shown is a flowchart illustrating an information prompting method provided in this application. The following description primarily uses a server as the execution entity. The specific implementation flow of this method is as follows:

[0118] S41: The server obtains the content feature information of each multimedia content, and based on the historical behavior of each object towards each multimedia content, obtains the statistical feature information of each multimedia content.

[0119] Content feature information refers to the inherent attributes of multimedia content, including but not limited to at least one of the following:

[0120] The first content sub-feature information is determined based on the title of the multimedia content, the second content sub-feature information is determined based on at least one level of classification information of the multimedia content, and the third content sub-feature information is determined based on the publication duration of the multimedia content.

[0121] For example, the title of a video can be "Family Fun" or "Celebrating xx Festival," and the first content sub-feature information can be the content tag determined based on the title; the second content sub-feature information of the video includes, but is not limited to: the video's primary category, secondary category, and tertiary category, for example, a game video has the primary category as: Game, the secondary category as: Shooting, and the tertiary category as: Peacekeeper; the third content sub-feature information of the video can refer to: the time difference between the video's release time and the present time, where the video's release time is the video's publication time.

[0122] Statistical feature information refers to features obtained based on the analysis of users' historical behavior towards multimedia content, including but not limited to at least one of the following:

[0123] The first statistical sub-feature information is determined based on the number of times each object performs various specified operations on multimedia content; the second statistical sub-feature information is determined based on the browsing time of each object on multimedia content.

[0124] For example, the first statistical sub-feature information of a video may specifically be: the like-to-vote ratio, the number of coins, the number of likes, the number of biu (likes), and the number of comments.

[0125] Among them, "biu" refers to the forwarding behavior within the content platform X, "biu count" is the number of forwards within the content platform X, and the like-to-vote ratio is calculated as: the number of likes for a video divided by the number of coins.

[0126] The second statistical sub-feature information may specifically include: the total historical viewing time of the video, the average historical viewing time, and the median historical viewing time.

[0127] S42: Based on the obtained content feature information and statistical feature information, the server obtains the information prompt time point for each multimedia content;

[0128] An alternative implementation method is to carry out step S42 according to the following process:

[0129] S421: Based on the content feature information and statistical feature information of each multimedia content, obtain the predicted playback completion rate of each multimedia content;

[0130] In this embodiment, the predicted playback completion rate of multimedia content can be estimated using machine learning: by inputting the second content sub-feature information, the third content sub-feature information, and the second statistical sub-feature information of each multimedia content into a trained classification model, the predicted playback completion rate of each multimedia content can be predicted. The predicted playback completion rate of a multimedia content represents its estimated viewing duration; therefore, based on the predicted playback completion rate and the completed playback duration of the multimedia content, the corresponding information prompt time point can be calculated.

[0131] One alternative implementation involves performing the following operations for each multimedia content:

[0132] S4211: Input the second content sub-feature information, the third content sub-feature information, and the second statistical sub-feature information of a multimedia content into the trained classification model respectively;

[0133] S4212: Based on the classification model, low-order feature crossing and high-order feature crossing are performed on the second content sub-feature information, the third content sub-feature information and the second statistical feature information, respectively, to obtain the second predicted probability corresponding to a multimedia content output by the classification model, and the second predicted probability is used as the predicted playback completion rate of a multimedia content.

[0134] In this embodiment of the application, the DeepFM model is used as a classification model for illustration. That is, the DeepFM model is used to predict the video playback completion rate. By inputting video features and statistical features based on user historical behavior into the DeepFM model, the predicted value of the video playback completion rate is finally output.

[0135] The DeepFM model is an end-to-end model that can extract features of varying complexity from raw features, eliminating the need for manual feature engineering. This model comprises two parts: FM and Deep. The FM part extracts low-order features and performs low-order feature cross-processing, while the Deep part extracts high-order features and performs high-order feature cross-processing. Furthermore, because the input is only the raw features, and FM and Deep share the underlying input vector features, the DeepFM model trains very quickly. It should be noted that the specific processing steps using the DeepFM model will be described in detail below and will not be repeated here.

[0136] It should also be noted that the DeepFM model is only used as an example for classification in this application embodiment. In fact, any model that can be used for classification can be applied to this application embodiment, and no specific limitation is made here.

[0137] S422: Determine the information prompt time point for each multimedia content based on its predicted playback completion rate and playback duration.

[0138] For example, for video 1, the playback duration is 1 minute, and the corresponding predicted playback completion rate is 90%; for video 2, the playback duration is 2 minutes, and the corresponding predicted playback completion rate is 50%.

[0139] In this embodiment, the information prompt time point can be determined based on the product of the playback duration and the predicted playback completion rate. For example, for video 1, the 54th second of the video, or a moment around the 54th second, can be used as the information prompt time point; for video 2, the 1st minute of the video, or a moment around the 1st minute, can be used as the information prompt time point, etc. No specific limitation is made here.

[0140] S43: The server performs the following operations for each multimedia content: based on the content embedding vector of a multimedia content and the object embedding vector of at least one object associated with the multimedia content, it filters out at least one matching target object.

[0141] S44: When the multimedia content reaches the corresponding information prompt time point, the server will send a target information prompt to at least one target object.

[0142] Specifically, the server can notify the terminal device to provide target information prompts to the target object. After receiving the notification, the terminal device will provide target information prompts to the corresponding target object at the corresponding information prompt time.

[0143] For example, for video 1, 10 target objects are matched and numbered according to lowercase English letters, which can be represented as target object a-target object j; for video 2, 5 target objects are matched and numbered as target object k-target object o.

[0144] For video 1, when a target object 'a' logs into terminal device A to watch the video, the server can notify terminal device A (or notify it in advance) based on the current playback progress of the video, just before the information prompt time. After receiving the notification message, terminal device A can display the target prompt information to target object 'a' when video 1 is played for 54 seconds. Similarly, for video 2, when a target object 'o' logs into terminal device B to watch the video, the server can notify terminal device B (or notify it in advance) based on the current playback progress of the video, just before the information prompt time. After receiving the notification message, terminal device B can display the target prompt information to target object 'o' when video 2 is played for 1 minute.

[0145] like Figure 5 The diagram shown is a schematic representation of a target prompt message in an embodiment of this application. Figure 5 The target message is "Feel free to toss a coin!", displayed as a pop-up near the coin toss control. This encourages user interaction and feedback, such as tossing a coin. Of course, it could also be a scrolling text bar, a floating layer, or other formats; no specific limitation is made here.

[0146] In addition, the target prompt can be displayed for a certain duration. This duration can be a preset fixed duration, such as 5 seconds; it can also be based on the remaining playback time of the currently playing video—for example, the longer the remaining playback time, the longer the display duration, and vice versa; or it can be set according to the behavioral habits of the target audience. In practical applications, it can be flexibly set according to specific circumstances, and no specific limitations are made here.

[0147] Based on the above implementation method, this application avoids the problem of manually setting pop-up parameters in traditional operating systems. By intelligently comparing videos with user groups, videos are targeted to the target audience, reducing harassment to irrelevant users. In addition, it can accurately predict the completion rate of video playback, so as to deliver coin-operated pop-up ads before users exit the video, ensuring the reach rate of coin-operated pop-up prompts.

[0148] It should be noted that the methods listed above can effectively improve the targeting of information prompts, that is, to provide information prompts only to the target audience. Additionally, the targeting of information prompts can be further enhanced from another perspective: intelligently filtering out target multimedia content with greater potential for information prompts, while other ordinary multimedia content does not require information prompts.

[0149] An alternative implementation involves filtering the multimedia content based on the obtained content feature information and statistical feature information before providing the target information to at least one target object at the corresponding information prompt time point when each multimedia content is played, and selecting the multimedia content whose probability of the object performing the target operation meets the expected criteria as the target multimedia content.

[0150] See Figure 6A The diagram shown is a flowchart of the second information prompting method in this application embodiment. The specific implementation process of this method is as follows:

[0151] S61: The server obtains the content feature information of each multimedia content, and based on the historical behavior of each object towards each multimedia content, obtains the statistical feature information of each multimedia content.

[0152] S62: Based on the obtained content feature information and statistical feature information, the server obtains the information prompt time point for each multimedia content.

[0153] S63: The server performs the following operations for each multimedia content: based on the content embedding vector of a multimedia content and the object embedding vector of at least one object associated with the multimedia content, it filters out at least one matching target object.

[0154] S64: Based on the obtained content feature information and statistical feature information, the server selects at least one target multimedia content from each multimedia content.

[0155] In this embodiment of the application, when filtering target multimedia content, it can also be achieved based on machine learning: by filtering at least one target multimedia content from each multimedia content based on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information and the first statistical sub-feature information of each multimedia content.

[0156] One alternative implementation involves performing the following operations for each multimedia content:

[0157] S641: Input the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information of a multimedia content into the trained classification model respectively;

[0158] S642: Based on the classification model, low-order feature crossing and high-order feature crossing are performed on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information and the first statistical sub-feature information, respectively, to obtain the first predicted probability of a multimedia content output by the classification model.

[0159] S643: Select multimedia content in the multimedia content set whose first predicted probability is greater than a preset probability threshold as target multimedia content.

[0160] In this embodiment, the DeepFM model can also be used to mine target multimedia content. This application uses video content with coin-operated prompts as an example; therefore, the target multimedia content can also be called video content with high coin-operated potential (referred to as high-potential coin-operated video). By inputting video features and statistical features based on user historical behavior into the DeepFM model, the final output determines whether the video is a high-potential coin-operated video. It should also be noted that the specific processing procedure using the DeepFM model will be described in detail below and will not be repeated here.

[0161] S65: When the multimedia content of each target is played to the corresponding information prompt time point, the server shall send a target information prompt to at least one target object.

[0162] It should be noted that the execution order of steps S62 to S64 can be flexibly adjusted, or they can be executed in parallel. The above is just an example and is not specifically limited in this article.

[0163] See Figure 6B As shown, it is Figure 6A The information prompting method shown corresponds to a system framework diagram. By intersecting the high-potential coin-operated videos obtained in step S64 and the information prompting time points of each video obtained in step S62 (determined based on the estimated video viewing time), the high-potential coin-operated videos and their corresponding information prompting time points are obtained. After matching the target objects in step S63, a file of {video: information prompting time point: target audience} is output for later deployment.

[0164] Based on the above implementation method, the problem of manually setting pop-up parameters in traditional operating systems is avoided. By intelligently selecting video content with high coin-operating potential, comparing the video with the user group, and targeting the video to the target audience, the system can reduce harassment to irrelevant users. In addition, it can accurately predict the completion rate of the video playback, so as to deliver the coin-operating pop-up advertisement before the user exits the video, thus ensuring the reach rate of the coin-operating pop-up prompt.

[0165] In addition, when prioritizing the selection of target multimedia content, it is also possible to perform target object matching and information prompt timing prediction only on the selected target multimedia content, reducing computational data and improving computational efficiency. For details on the implementation process, please refer to [link / reference needed]. Figure 7A .

[0166] like Figure 7A The diagram shown is a flowchart of the third information prompting method in this application embodiment. The specific implementation process of this method is as follows:

[0167] S71: The server obtains the content feature information of each multimedia content, and based on the historical behavior of each object towards each multimedia content, obtains the statistical feature information of each multimedia content.

[0168] S72: Based on the obtained content feature information and statistical feature information, the server selects at least one target multimedia content from each multimedia content.

[0169] S73: Based on the obtained content feature information and statistical feature information, the server obtains the information prompt time points for each target multimedia content.

[0170] S74: The server performs the following operations for each target multimedia content: based on the content embedding vector of a target multimedia content and the object embedding vector of at least one object associated with the target multimedia content, it filters out at least one matching target object.

[0171] S75: When the target multimedia content is played to the corresponding information prompt time point, the server will send a target information prompt to at least one target object.

[0172] That is, firstly, target video content with high coin-operated potential is screened out; then, target audience matching and information prompt timing are predicted for these target videos. This method is similar to... Figure 6A Compared to the method shown, this method can reduce the amount of data required for computation and improve computational efficiency.

[0173] See Figure 7B As shown, it is Figure 7A The information prompting method shown is a system framework diagram. Figure 7A The corresponding intelligent pop-up notification method in the video progress bar can be divided into the following three tasks: Task 1: Predict high-potential videos for video submission by utilizing users' historical video viewing behavior and basic video characteristics; Task 2: Match the target audience for the video; Task 3: Estimate the most appropriate time for the video pop-up notification. By combining these three objectives, the video pop-up delivery task can be executed intelligently.

[0174] The following will combine Figure 8 The model structure shown provides a detailed description of the information prompting method in the embodiments of this application. Figure 8 This is a schematic diagram of the structure of a classification model in an embodiment of this application. The classification model is the DeepFM model.

[0175] Specifically, in Figure 7B Two models were used in the three tasks shown: DeepFM classification to predict high-potential coin-operated videos, and DeepFM regression to estimate video completion rates. The FM and Deep modules performed second-order cross-learning and higher-order learning on the input features, respectively, ultimately outputting the probability of a high-potential coin-operated video or the video completion rate. In addition, a simple DNN was used to match the target audience of the video, outputting the matching probability value for each user. The above models demonstrate strong expressive power and achieved good results online.

[0176] The following is a detailed introduction to Task 1:

[0177] In this application embodiment, Task 1 may refer to the screening of high-potential coin-operated videos based on the DeepFM model. The following mainly focuses on... Figure 8The diagram shows four parts: the input layer and feature representation layer (including sparse and dense feature layers), the Factorization Machine (FM) layer, the deep layer, and the output unit. These will be described in detail below.

[0178] (I) Input layer and feature expression layer.

[0179] When screening high-potential coin-operated videos, the input features used (content features and statistical features of the video) can be specifically divided into the following as listed in step S64: first content sub-feature information, second content sub-feature information, third content sub-feature information and first statistical sub-feature information of the multimedia content.

[0180] Specifically, such as Figure 9 As shown, this is a schematic diagram of the feature information applied to predicting high-potential coin-operated content in an embodiment of this application. The first content sub-feature information is the content tag, the second content sub-feature information is the first-level content category, and the third content sub-feature information is the time difference from the time of release to the present. The first statistical sub-feature information includes the like-to-coin ratio, number of coins, number of likes, number of biu (user interactions), and number of comments. Additionally, Figure 9 The input features shown also include a content ID, which uniquely identifies a multimedia content.

[0181] In this embodiment, when training the model based on the training sample dataset with Task 1 as the objective, data from the first three days can be sampled for prediction and validation, while data from the last three days can be used to predict the result for the next day. Samples with 10 or more coins cast on a video, or a like-to-like ratio of 100 or more, are considered positive samples, and the rest are negative samples. Taking the time period from 0901 to 0907 as an example, the data and labels from 0901 to 0903 can be used for prediction and validation to generate model M; the data from 0904 to 0906 is input into model M, and the obtained prediction results are used as the output for 0907.

[0182] Based on the above sampling method, the model can be continuously updated to improve its accuracy.

[0183] Assuming there are n samples, each represented as (χ, y), as shown above, this task has a total of 9 features. Therefore, χ is a vector composed of 9 fields, including user features, content features, and statistical features. y∈{0,1}, y=1 indicates that the prediction result is a positive sample, i.e., a high-potential coin-operated video. Conversely, y=1 indicates that the predicted video's coin-operated potential is not high enough, meaning that after the video is shown to the user, the probability of the user interfering with the coin is not high, and the probability does not meet the expected standard.

[0184] From an input perspective, there are currently nine feature values ​​as input, which can be divided into two parts: categorical features and continuous features. In this task, categorical features include the content's primary classification and content tags, while continuous features include the like-to-vote ratio, number of coins, number of likes, number of biu (user engagement), number of comments, and the time difference between the time of release and the present. Categorical features are generally represented as a one-hot vector, while continuous features are often left unprocessed as their original values.

[0185] The input data x = [x1, x2, ..., x9] is transformed into a sparse representation after one-hot encoding: X = [X...]. field1 ,X field2 ,…,X field9 The vectors in χ are then mapped to dense vector representations χ = [e1, e2, ..., e9]. Each vector in χ corresponds one-to-one with a field. These inputs can then be used to build a binary classification prediction model.

[0186] (ii) FM layer.

[0187] In other methods, feature training can only be performed when both feature values ​​exist. However, in cases of sparse features, FM can more effectively acquire second-order features, such as... Figure 10 The diagram shown is a structural schematic of an FM implementation in this application. Specifically, regardless of whether the value corresponding to feature i exists, a latent vector V is trained for each feature i. i When features i and j intersect, through V i ·V j We calculate the weights after the interaction of two features. The FM calculation involved in the network is as follows:

[0188]

[0189]

[0190] Where X is the feature input of the sparse feature layer. <V i V j >X j1 ·X j2 The inner product operation representing the intersection of two features is represented by the vector expression e of the dense layer. i =V i ·X i For a detailed structure of the dense layer, please refer to Figure 11 The diagram illustrates the feature processing procedure for the Deep part.

[0191] It should be noted that, although X... fieldThe dimensions are quite different, but they will all ultimately map to the embedding layer with the same dimension k. In addition, the latent vector V learned in FM is also applied in the Deep part; both parts share the same underlying layer throughout the network and jointly learn the latent vector V. The dense layer structure is as follows: Figure 11 The input layer X is an n×d two-dimensional matrix, and V is a d×k two-dimensional matrix, which ultimately generates an n×d embedding vector for n samples.

[0192] (III) The Deep part, namely Figure 8 The hidden layers of deep neural networks (DNNs) in the image.

[0193] The Deep part is a classic feedforward network, capable of learning high-dimensional interactions between features. From Figure 12 As can be seen, all input features are transformed into dense layer embedding vectors. The embedding vector is a low-level dense mapping of high-dimensional sparse vectors, which not only fully learns the information in the features but also enhances the model's generalization ability.

[0194] (iv) Output unit.

[0195] The outputs of the FM and Deep parts are combined and fed into a sigmoid layer to obtain the final prediction result. That is, the first predicted probability, in the embodiments of this application, can be Videos with a preset probability threshold are considered high-potential coin-operated videos; otherwise, they are not.

[0196]

[0197] In addition, the ID and corresponding embedding vector of the high-potential coin-operated video will be output for use in the next stage.

[0198] Next, we will provide a detailed introduction to Task 2:

[0199] This section mainly introduces the specific implementation details of matching video content with the target audience. The network structure is as follows: Figure 13 As shown, it is a schematic diagram of the structure of a deep model in an embodiment of this application. As shown in step S43, the input of the model includes: the content embedding vector of multimedia content, and the object embedding vector of at least one associated object.

[0200] Optionally, a content embedding vector of multimedia content refers to the vector obtained by embedding the multimedia content. Since there are many ways to perform embedding representation, a content embedding vector in this embodiment includes, but is not limited to, at least one of the following:

[0201] (1) Based on the trained classification model, feature extraction is performed on the content feature information and statistical feature information of a multimedia content to obtain the first content sub-embedding vector;

[0202] Taking video content as an example, the first content sub-embedding vector is the video embedding obtained by the DeepFM module. The result e of the dense layer in DeepFM corresponding to Task 1 is selected as the embedding representation of the current video. This embedding is an embedding vector obtained based on user statistics over three days and some basic video features, and has sufficient expressive power.

[0203] (2) Based on at least one of computer vision technology and natural language technology, a multimedia content is embedded to obtain a second content sub-embedding vector.

[0204] Taking video content as an example, the second content sub-embedding vector is the content-based video embedding, such as the combination of embedding learned from the video using computer vision technology and embedding learned from video-related text information (such as video descriptions or video comments) using natural language technology. This vector is independent of user context behavior and time, and is a relatively pure embedding expression.

[0205] Optionally, the object embedding vector of at least one associated object includes at least one of the following:

[0206] (1) The first object sub-embedding vector is determined based on the average of the embedding vectors of each object browsing a multimedia content within a certain historical time.

[0207] Taking video content as an example, the first object's sub-embedding vector is the user's embedding for viewing the current video:

[0208] Users who viewed the current video within the past three days were selected, and the average of their embeddings was used as the final embedding. The purpose of selecting this feature was to find users who were relevant to the content.

[0209] (2) The second object sub-embedding vector is determined by averaging the embedding vectors of each object that performs target operations on at least one multimedia content within a certain historical period.

[0210] Taking video content as an example, the second object sub-embedding vector is the overall video coin-throwing user embedding, where the target operation is coin-throwing:

[0211] Users who have watched videos on content platform X in the past three days and have donated coins to the video pool are selected. The average embedding of these users is used as the final embedding. The rationale for selecting this feature is to find commonalities among users who donated coins.

[0212] (3) The third object sub-embedding vector is determined by averaging the embedding vectors of each object that performed a specified operation on a multimedia content within a certain historical period but did not perform the target operation.

[0213] Taking video content as an example, the third object sub-embedding vector is the embedding of users who liked but did not give coins in the current video, where the target operation is giving coins and the specified operation is giving a like:

[0214] Users who have watched the current video and liked it in the last three days but have not donated coins are selected. The average of the embeddings of these users is used as the final embedding. The purpose of selecting this feature is to find potential coin-donating users.

[0215] like Figure 13 As shown, taking video content as an example, the input features include user embedding and video embedding. The initial video embedding is an embedding vector learned from the video content based on computer vision and natural language processing techniques. The initial user embedding is expressed by summing the video embeddings viewed by the user over the past three days. Specifically, the five types of user embeddings—the three user-behavior-based embeddings listed above (i.e., the embedding of users who viewed the current video, the embedding of users who donated coins to the entire video, and the embedding of users who liked the current video but did not donate coins) and the two types of video embeddings obtained through DeepFM (video embedding obtained by the DeepFM module, and the current video embedding)—are concatenated after processing, undergo dimensionality reduction through three ReLU activation layers, and finally output to the softmax module.

[0216] It should be noted that the deep model in the embodiments of this application can be a DNN network, using a common tower-like design, with the bottom layer having the largest dimension, and the number of neurons in each layer above it being halved, until the softmax input layer has 256 layers.

[0217] The output unit of the depth model will be described in detail below:

[0218] In this embodiment, the target audience matching problem can be viewed as a large-scale multi-class classification problem, meaning that at time t, under the context relationship C (referring to information such as mobile phone location and reading time), video V is among the currently online user group U, and user u is... i The probability of a match, also known as the second prediction probability.

[0219]

[0220] In formula (4), u i represents the user embedding vector, and v represents the video embedding vector, i.e., the target content embedding vector.

[0221] It should be noted that the user embedding in formula (4) refers to the vector obtained by embedding the account of user i, which is different from the above. Figure 13 The listed user embeddings are different, as mentioned above Figure 13 The user embeddings listed refer to the average of user embeddings that meet certain conditions.

[0222] In this embodiment, the goal of the DNN is to output the video embedding from the last Rectified Linear Unit (ReLU) layer, given five embedding inputs. The output video embedding is multiplied by the user embedding, and a softmax classifier is used to obtain the probability of each video matching the user. Finally, the top N users with the highest scores are selected as the output of the model, which are the target groups that match the videos.

[0223] In one optional implementation, step S43 can be specifically divided into the following sub-steps:

[0224] S431: Based on the content embedding vector of a multimedia content and the object embedding vector of at least one associated object, determine the target content embedding vector of a multimedia content, i.e. v in formula (4);

[0225] Step S431 can be further divided into the following sub-steps:

[0226] S4311: Input the content embedding vector of a multimedia content and the object embedding vector of at least one associated object into the trained deep model, respectively.

[0227] S4312: Based on a deep model, the dimensionality of the content embedding vector and the object embedding vector of at least one associated object is reduced to obtain a target content embedding vector of multimedia content.

[0228] S432: Based on the target content embedding vector and the object embedding vector of each object, determine the matching degree between a multimedia content and each object respectively;

[0229] Specifically, this step can be implemented based on the above formula (4). For a multimedia content, the matching degree P(w) between the multimedia content and each object can be calculated based on the target content embedding vector of the multimedia content and the object embedding vector of each object. t =i|V,C).

[0230] S433: Based on the obtained matching scores, select at least one target object that matches a multimedia content.

[0231] In other words, for a multimedia content, the top N users with the highest matching degree with the content are selected as the target objects.

[0232] Similar to the DeepFM model training process, in the DNN training process, the data from the first three days are used for prediction and validation, and the data from the last three days are used to predict the samples for the next day. Users who watched the current video and made a coin donation are considered positive samples, labeled y=1, and the rest are negative samples. Since users who made a coin donation represent a small percentage of the total video exposure, negative samples are sampled so that the ratio of positive samples to negative samples equals the total number of coins donated to the video / the total number of unique video exposures.

[0233] Based on the above implementation method, the problem of low payout ratio caused by pop-up reminders in indiscriminate distribution tasks is avoided. This is because intelligent pop-ups utilize basic video information and users' historical video browsing behavior to continuously learn in a deep learning network. For each video, the most suitable audience for coin-donation pop-up reminders is selected, achieving personalized matching between videos and users. This method increases the popularity of coin-donation operations among novice users while reducing the acceptance rate of pop-up reminders among experienced users. It increases the number of coins donated on the platform while avoiding excessive disturbance to users.

[0234] Finally, a detailed introduction to Task 3 will be provided:

[0235] The model used in this task is the same Figure 8 It also uses the DeepFM network structure, but there are some differences in the input features and fitting target. The following will mainly focus on... Figure 8 The input layer and feature representation layer (including sparse and dense feature layers) and output unit are shown in detail. These two parts, which differ from Task 1, will be discussed in detail:

[0236] (I) Input layer and feature expression layer.

[0237] When estimating the time of entering the pop-up prompt, the input features (video content features and statistical features) used can be specifically divided into the following listed in step S4211: second content sub-feature information, third content sub-feature information and second statistical sub-feature information.

[0238] Specifically, such as Figure 14 As shown, the second content sub-feature information includes the first-level content classification, the second-level content classification, and the third-level content classification. The third content sub-feature information is the time difference from the time of release to the present. The second statistical sub-feature information includes: the total historical viewing time of the video, the average historical viewing time, and the median historical viewing time.

[0239] (ii) Output unit.

[0240] This part inputs the processed features into the FM and Deep networks to share the underlying learning. The outputs of the FM and Deep parts are combined and fed into a sigmoid layer to obtain the final prediction result. The target fitted in the output unit is the ratio of the median viewing time of all users for each video to the total video duration. Therefore, the output result is also the predicted value of the video playback completion rate.

[0241] Similar to the training process for the two tasks mentioned above, this method uses data from the first three days for prediction and validation, and data from the last three days for predicting the next day's sample. Since this task is a regression problem, aiming to predict the most likely playback duration of a video, there is no distinction between positive and negative samples; all samples can be used for training. Taking the time period from September 1st to 7th as an example, the data and labels from September 1st to 3rd are used for prediction and validation to generate model M; the data from September 4th to 6th are input into model M, and the resulting predictions are used as the output for September 7th.

[0242] Based on the above implementation method, the viewing time of the current video is predicted according to the user's historical browsing habits of other videos, and a coin-insertion pop-up prompt appears at the corresponding time point, which improves the reach of the coin-insertion pop-up prompt among users.

[0243] See Figure 15 The diagram shown illustrates a specific implementation process for providing coin-operated information prompts for video content in an embodiment of this application. The method specifically includes the following steps:

[0244] (i) Based on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information and the first statistical sub-feature information of the video content, calculate the first prediction probability corresponding to each video content in the content pool, and select at least one target video content based on the first prediction probability.

[0245] In one alternative implementation, the specific process of calculating the first predicted probability corresponding to a video content based on the classification model can be summarized as follows:

[0246] First, based on the feature sparse layer in the DeepFM model, the first content feature information, second content feature information, third content feature information, and first statistical feature information of a video content are sparsely represented to obtain the first sparse vector X1 for each video content. field Then, through the dense layer in the classification model, the first sparse vector of the video content is mapped to the first dense vector e1; subsequently, based on the FM layer in the classification model, second-order feature extraction is performed on the first sparse vector and the first dense vector of the video content to obtain the first low-order cross feature y1. FM Furthermore, based on the DNN hidden layer in the classification model, high-dimensional cross-feature extraction is performed on the first dense vector to obtain the first high-order cross-feature y1. DNN Finally, based on the first low-order cross feature y1 corresponding to the video content... FM and the first higher-order cross feature y1 DNN Referring to formula (3), determine the first predicted probability of the video content.

[0247] (ii) Based on the second content sub-feature information, the third content sub-feature information and the second statistical sub-feature information of the video content, calculate the second prediction probability corresponding to each video content, and determine the coin insertion information prompt time point corresponding to each video content based on the second prediction probability.

[0248] In one alternative implementation, the specific process of calculating the second predicted probability corresponding to a video content based on the classification model can be summarized as follows:

[0249] First, based on the feature sparse layer in the DeepFM model, the second content feature information, third content feature information, and second statistical feature information of a video content are sparsely represented to obtain the second sparse vector X2 for each video content. field Then, through the dense layer in the classification model, the second sparse vector of the video content is mapped to a second dense vector e2; subsequently, based on the FM layer in the classification model, second-order feature extraction is performed on the second sparse vector and the second dense vector of the video content to obtain the second low-order cross feature y2 of the video content. FMFurthermore, based on the DNN hidden layer in the classification model, high-dimensional cross-feature extraction is performed on the second dense vector of each video content to obtain the second high-order cross-feature y2 of each video content. DNN Finally, based on the second low-order cross feature y2 corresponding to the video content... FM Second higher-order cross feature y2 DNN Referring to formula (3), determine the second prediction probability of the video content.

[0250] (iii) Based on the content embedding vector of the video content and the object embedding vector of at least one associated object, the matching degree of each video content is predicted for target audience matching.

[0251] Specifically, at least one associated object is determined based on some historical behavior of the object in relation to the video. Furthermore, the prediction of this matching degree can be based on… Figure 13 For detailed implementations of the DNN networks listed above, please refer to the above embodiments; repeated details will not be repeated.

[0252] Finally, when each target video reaches the corresponding coin-donation prompt time, a coin-donation prompt will be sent to at least one target group, in the following manner: Figure 5 As shown, no specific limitations are made here.

[0253] It should be noted that, Figure 15 The two DeepFM models listed above may have the same structure, but the corresponding training samples and model parameters are actually different. For details, please refer to the above examples. Repeated parts will not be repeated.

[0254] Based on the same inventive concept, embodiments of this application also provide an information prompting device. For example... Figure 16 As shown, this is a structural schematic diagram of the information prompting device 1600, which may include:

[0255] The feature acquisition unit 1601 is used to acquire the content feature information of each multimedia content, and to acquire the statistical feature information of each multimedia content based on the historical behavior of each object towards each multimedia content.

[0256] The time determination unit 1602 is used to obtain the information prompt time point of each multimedia content based on the obtained content feature information and statistical feature information.

[0257] The object filtering unit 1603 is used to perform the following operations for each multimedia content: based on the content embedding vector of a multimedia content and the object embedding vector of at least one object associated with the multimedia content, filter out at least one matching target object.

[0258] The information prompting unit 1604 is used to provide target information prompts to at least one target object when each multimedia content is played to the corresponding information prompt time point.

[0259] Optional, time determination unit 1602:

[0260] Based on the content feature information and statistical feature information of each multimedia content, the predicted playback completion rate of each multimedia content is obtained respectively.

[0261] Based on the predicted completion rate and playback duration of each multimedia content, the respective information prompt time points are determined.

[0262] Optionally, the device also includes:

[0263] The content filtering unit 1605 is used to filter at least one target multimedia content from each multimedia content based on the obtained content feature information and statistical feature information before the information prompting unit 1604 prompts the target information to at least one target object when each multimedia content is played to the corresponding information prompting time point. Each target multimedia content is a multimedia content in which the probability of the object performing the target operation meets the expected standard.

[0264] Information prompting unit 1604 is specifically used for:

[0265] When the multimedia content of each target is played to the corresponding information prompt time point, a target information prompt is given to at least one target object.

[0266] Optionally, the content feature information includes first content sub-feature information determined based on the title of the multimedia content, second content sub-feature information determined based on at least one level of classification information of the multimedia content, and third content sub-feature information determined based on the publication duration of the multimedia content; the statistical feature information includes first statistical sub-feature information determined based on the number of times each object performs various specified operations on the multimedia content.

[0267] Content filtering unit 1605 is specifically used for:

[0268] Based on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information of each multimedia content, at least one target multimedia content is selected from each multimedia content.

[0269] Optionally, content filtering unit 1605 is specifically used for:

[0270] For each multimedia content, perform the following operations: input the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information of a multimedia content into the trained classification model; based on the classification model, perform low-order feature cross and high-order feature cross on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information, respectively, to obtain the first predicted probability of a multimedia content output by the classification model.

[0271] Multimedia content in the multimedia content set whose first predicted probability is greater than a preset probability threshold is selected as target multimedia content.

[0272] Optionally, the content feature information includes second content sub-feature information determined based on at least one level of classification information of multimedia content, and third content sub-feature information determined based on the publication duration of multimedia content; the statistical feature information includes second statistical sub-feature information determined based on the browsing duration of each object for multimedia content;

[0273] The time determination unit 1602 is specifically used for:

[0274] Based on the second content sub-feature information, the third content sub-feature information, and the second statistical sub-feature information of each multimedia content, the predicted playback completion rate of each multimedia content is predicted.

[0275] Optional, time determination unit 1602:

[0276] Perform the following operations for each multimedia content:

[0277] Input the second content sub-feature information, the third content sub-feature information, and the second statistical sub-feature information of a multimedia content into the trained classification model respectively;

[0278] Based on the classification model, low-order feature crossing and high-order feature crossing are performed on the second content sub-feature information, the third content sub-feature information and the second statistical feature information, respectively, to obtain the second predicted probability corresponding to a multimedia content output by the classification model, and the second predicted probability is used as the predicted playback completion rate of a multimedia content.

[0279] Optionally, object filtering unit 1603 is specifically used for:

[0280] Based on the content embedding vector of a multimedia content and the object embedding vector of at least one associated object, determine the target content embedding vector of the multimedia content.

[0281] Based on the target content embedding vector and the object embedding vector of each object, the matching degree between a multimedia content and each object is determined respectively.

[0282] Based on the obtained matching scores, at least one target object matching a multimedia content is selected.

[0283] Optionally, object filtering unit 1603 is specifically used for:

[0284] Input the content embedding vector of a multimedia content and the object embedding vector of at least one associated object into the trained deep model, respectively.

[0285] Based on a deep model, the dimensionality of the content embedding vector and the object embedding vector of at least one associated object is reduced to obtain a target content embedding vector of multimedia content.

[0286] Optionally, the content embedding vector of a multimedia content includes at least one of the following:

[0287] Based on the trained classification model, feature extraction is performed on the content feature information and statistical feature information of a multimedia content to obtain the first content sub-embedding vector;

[0288] A second content sub-embedding vector is obtained by embedding a multimedia content using at least one of computer vision technology and natural speech technology.

[0289] Optionally, the object embedding vector of at least one associated object includes at least one of the following:

[0290] The first object sub-embedding vector is determined based on the average of the embedding vectors of each object browsing a multimedia content over a certain historical period.

[0291] The second object sub-embedding vector is determined by averaging the embedding vectors of each object that performs a target operation on at least one multimedia content within a certain historical period.

[0292] The third object sub-embedding vector is determined by averaging the embedding vectors of each object that performed a specified operation on a multimedia content within a certain historical period but did not perform the target operation.

[0293] This application predicts the information prompting time for each multimedia content based on its own content characteristics and statistical features determined by the historical behavior of each object. This allows for personalized setting of information prompting time points for different multimedia content. By predicting the information prompting time points based on the historical behavior of each object towards the multimedia content, and delivering target information prompts at the corresponding time points, the reach of prompts to users is improved. Furthermore, by using the content embedding vector of each multimedia content and the object embedding vector of related objects, the target objects matching each multimedia content are selected, choosing the most suitable group of people who need target information prompts for each multimedia content, thus achieving personalized and more targeted prompts. In addition, the information prompting method in this application does not require manual settings and significantly improves the relevant task indicators of target information prompts while achieving intelligent prediction.

[0294] For ease of description, the above sections are divided into modules (or units) according to their functions and described separately. Of course, in implementing this application, the functions of each module (or unit) can be implemented in one or more software or hardware components.

[0295] Having described the information prompting method and apparatus according to exemplary embodiments of this application, we will now describe another exemplary embodiment of an information prompting apparatus according to this application.

[0296] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0297] Based on the same inventive concept as the above-described method embodiments, this application also provides an electronic device. In one embodiment, the electronic device may be a server, such as... Figure 2 The server 220 is shown. In this embodiment, the structure of the electronic device can be as follows: Figure 17 As shown, it includes a memory 1701, a communication module 1703, and one or more processors 1702.

[0298] The memory 1701 is used to store computer programs executed by the processor 1702. The memory 1701 may mainly include a program storage area and a data storage area. The program storage area may store the operating system and programs required to run instant messaging functions, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0299] Memory 1701 may be volatile memory, such as random-access memory (RAM); memory 1701 may also be non-volatile memory, such as read-only memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or memory 1701 may be any other medium capable of carrying or storing a desired computer program having the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1701 may be a combination of the above-described memories.

[0300] Processor 1702 may include one or more central processing units (CPUs) or digital processing units, etc. Processor 1702 is used to implement the above-mentioned information prompting method when calling computer programs stored in memory 1701.

[0301] The communication module 1703 is used to communicate with terminal devices and other servers.

[0302] This application embodiment does not limit the specific connection medium between the memory 1701, communication module 1703, and processor 1702. This application embodiment... Figure 17 The memory 1701 and the processor 1702 are connected via a bus 1704, and the bus 1704 is in Figure 17 The diagram uses thick lines to describe the connections between other components; these are for illustrative purposes only and should not be considered limiting. The 1704 bus can be divided into address bus, data bus, control bus, etc. For ease of description, Figure 17 It is described using only a thick line, but does not indicate that there is only one bus or one type of bus.

[0303] The memory 1701 stores a computer storage medium, which in turn stores computer-executable instructions. These instructions are used to implement the information prompting method of this application embodiment. The processor 1702 is used to execute the aforementioned information prompting method, such as... Figure 4 or Figure 6A or Figure 7A As shown.

[0304] In another embodiment, the electronic device may also be other electronic devices, such as... Figure 2 The terminal device 210 is shown. In this embodiment, the electronic device can be structured as follows: Figure 18As shown, it includes components such as: communication component 1810, memory 1820, display unit 1830, camera 1840, sensor 1850, audio circuit 1860, Bluetooth module 1870, processor 1880, etc.

[0305] The communication component 1810 is used to communicate with the server. In some embodiments, it may include a Circuit-Based Wireless Fidelity (WiFi) module, which is a short-range wireless transmission technology. Electronic devices can use the WiFi module to help users send and receive information.

[0306] The memory 1820 can be used to store software programs and data. The processor 1880 executes various functions of the terminal device 210 and performs data processing by running the software programs or data stored in the memory 1820. The memory 1820 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1820 stores an operating system that enables the terminal device 210 to run. In this application, the memory 1820 may store the operating system and various applications, and may also store code that executes the information prompting method of the embodiments of this application.

[0307] The display unit 1830 can also be used to display information input by the user or information provided to the user, as well as various menus of the terminal device 210, forming a graphical user interface (GUI). Specifically, the display unit 1830 may include a display screen 1832 disposed on the front of the terminal device 210. The display screen 1832 may be configured as a liquid crystal display, a light-emitting diode, or the like. The display unit 1830 can be used to display multimedia content and target prompt information, as described in the embodiments of this application.

[0308] The display unit 1830 can also be used to receive input digital or character information and generate signal inputs related to user settings and function control of the terminal device 210. Specifically, the display unit 1830 may include a touch screen 1831 disposed on the front of the terminal device 210, which can collect touch operations of the user on or near it, such as clicking a button, dragging a scroll box, etc.

[0309] The touchscreen 1831 can be placed over the display screen 1832, or the touchscreen 1831 and the display screen 1832 can be integrated to realize the input and output functions of the terminal device 210. After integration, it can be referred to as a touch display screen. In this application, the display unit 1830 can display the application program and the corresponding operation steps.

[0310] Camera 1840 can be used to capture still images, which users can then post comments on via the application. There can be one or multiple cameras 1840. An object is projected onto a photosensitive element through a lens, generating an optical image. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to the processor 1880 to be converted into a digital image signal.

[0311] The terminal device may also include at least one sensor 1850, such as an accelerometer 1851, a proximity sensor 1852, a fingerprint sensor 1853, and a temperature sensor 1854. The terminal device may also be equipped with other sensors such as a gyroscope, barometer, hygrometer, thermometer, infrared sensor, light sensor, and motion sensor.

[0312] Audio circuitry 1860, speaker 1861, and microphone 1862 provide an audio interface between the user and terminal device 210. Audio circuitry 1860 converts received audio data into electrical signals, which are then transmitted to speaker 1861, where they are converted into sound signals for output. Terminal device 210 may also be equipped with volume buttons for adjusting the volume of the sound signal. On the other hand, microphone 1862 converts collected sound signals into electrical signals, which are received by audio circuitry 1860, converted back into audio data, and then output to communication component 1810 for transmission to, for example, another terminal device 210, or to memory 1820 for further processing.

[0313] The Bluetooth module 1870 is used to interact with other Bluetooth devices that also have a Bluetooth module via the Bluetooth protocol. For example, a terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smartwatch) that also has a Bluetooth module through the Bluetooth module 1870, thereby exchanging data.

[0314] The processor 1880 is the control center of the terminal device, connecting various parts of the terminal through various interfaces and lines. It executes various functions and processes data by running or executing software programs stored in the memory 1820 and calling data stored in the memory 1820. In some embodiments, the processor 1880 may include one or more processing units; the processor 1880 may also integrate an application processor and a baseband processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the baseband processor mainly handles wireless communication. It is understood that the baseband processor may not be integrated into the processor 1880. In this application, the processor 1880 can run the operating system, applications, user interface display and touch response, as well as the information prompting method of the embodiments of this application. Furthermore, the processor 1880 is coupled to the display unit 1830.

[0315] In some possible implementations, various aspects of the information prompting method provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product is run on an electronic device, the computer program causes the electronic device to perform the steps of the information prompting method according to the various exemplary embodiments of this application described above. For example, the electronic device can perform actions such as... Figure 4 or Figure 6A or Figure 7A The steps are shown in the figure.

[0316] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0317] The program product of the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on a computing device. However, the program product of this application is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with a command execution system, apparatus, or device.

[0318] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. This propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with a command execution system, apparatus, or device.

[0319] Computer programs contained on readable media may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0320] Computer programs for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The computer program can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0321] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0322] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0323] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing a computer-usable computer program.

[0324] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, produce a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0325] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0326] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. An information prompting method, characterized in that, The method includes: The content feature information of each multimedia content is obtained separately, and the statistical feature information of each multimedia content is obtained separately based on the historical behavior of each object towards each multimedia content. Based on the content feature information and statistical feature information of each multimedia content, the predicted playback completion rate of each multimedia content is obtained; based on the predicted playback completion rate and playback duration of each multimedia content, the information prompt time point is determined. Perform the following operations for each multimedia content: Based on the content embedding vector of a multimedia content and the object embedding vector of at least one of the objects associated with the multimedia content, filter out at least one matching target object; When each multimedia content reaches the corresponding information prompt time point, a target information prompt is given to at least one target object.

2. The method as described in claim 1, characterized in that, Before providing the target information prompt to at least one target object when each multimedia content is played to the corresponding information prompt time point, the method further includes: Based on the obtained content feature information and statistical feature information, at least one target multimedia content is selected from the multimedia content, wherein each target multimedia content is: multimedia content in which the probability of the object performing the target operation meets the expected standard. The step of providing target information prompts to at least one target object when each multimedia content is played to the corresponding information prompt time point includes: When the multimedia content of each target is played to the corresponding information prompt time point, a target information prompt is given to at least one target object.

3. The method as described in claim 2, characterized in that, The content feature information includes first content sub-feature information determined based on the title of the multimedia content, second content sub-feature information determined based on at least one level of classification information of the multimedia content, and third content sub-feature information determined based on the publication duration of the multimedia content; the statistical feature information includes first statistical sub-feature information determined based on the number of times each object performs various specified operations on the multimedia content. The step of selecting at least one target multimedia content from the multimedia content based on the obtained content feature information and statistical feature information includes: Based on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information of each multimedia content, at least one target multimedia content is selected from the multimedia content.

4. The method as described in claim 3, characterized in that, The step of pre-screening at least one target multimedia content from the multimedia content based on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information of each multimedia content includes: For each multimedia content, the following operations are performed: input the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information of a multimedia content into a trained classification model; based on the classification model, perform low-order feature cross and high-order feature cross on the first content sub-feature information, the second content sub-feature information, the third content sub-feature information, and the first statistical sub-feature information, respectively, to obtain the first predicted probability of the multimedia content output by the classification model; The multimedia content whose first predicted probability is greater than a preset probability threshold is selected as the target multimedia content.

5. The method as described in claim 1, characterized in that, The content feature information includes second content sub-feature information determined based on at least one level of classification information of multimedia content, and third content sub-feature information determined based on the publication duration of multimedia content; the statistical feature information includes second statistical sub-feature information determined based on the browsing duration of each object for the multimedia content; The step of obtaining the predicted playback completion rate for each multimedia content based on its respective content feature information and statistical feature information includes: Based on the second content sub-feature information, the third content sub-feature information, and the second statistical sub-feature information of each multimedia content, the predicted playback completion rate of each multimedia content is predicted.

6. The method as described in claim 5, characterized in that, The prediction of the predicted playback completion rate for each multimedia content is obtained based on its second content feature information, third content feature information, and second statistical feature information, including: Perform the following operations on each of the aforementioned multimedia contents: Input the second content sub-feature information, the third content sub-feature information, and the second statistical sub-feature information of a multimedia content into the trained classification model respectively; Based on the classification model, low-order feature cross and high-order feature cross are performed on the second content sub-feature information, the third content sub-feature information and the second statistical feature information, respectively, to obtain the second predicted probability corresponding to the multimedia content output by the classification model, and the second predicted probability is used as the predicted playback completion rate of the multimedia content.

7. The method according to any one of claims 1 to 6, characterized in that, The step of filtering out at least one matching target object based on the content embedding vector of a multimedia content and the object embedding vector of at least one object associated with the multimedia content includes: Based on the content embedding vector of the multimedia content and the object embedding vector of the associated at least one of the objects, the target content embedding vector of the multimedia content is determined. Based on the target content embedding vector and the object embedding vector of each object, the matching degree between the multimedia content and each object is determined respectively. Based on the obtained matching scores, at least one target object matching the multimedia content is selected.

8. The method as described in claim 7, characterized in that, Determining the target content embedding vector of the multimedia content based on the content embedding vector of the multimedia content and the object embedding vector of the associated at least one of the objects includes: The content embedding vector of the multimedia content and the object embedding vector of at least one associated object are respectively input into the trained deep model. Based on the deep model, the dimensionality reduction of the content embedding vector and the object embedding vector of at least one associated object is performed to obtain the target content embedding vector of the multimedia content.

9. The method according to any one of claims 1 to 6, characterized in that, The content embedding vector of a multimedia content includes at least one of the following: Based on the trained classification model, feature extraction is performed on the content feature information and statistical feature information of the multimedia content to obtain the first content sub-embedding vector; The second content sub-embedding vector is obtained by embedding the multimedia content based on at least one of computer vision technology and natural language technology.

10. The method according to any one of claims 1 to 6, characterized in that, The object embedding vector of at least one of the associated objects includes at least one of the following: The first object sub-embedding vector is determined based on the average of the embedding vectors of each object browsing the multimedia content within a certain historical period. The second object sub-embedding vector is determined by averaging the embedding vectors of each object that performs a target operation on at least one multimedia content within a certain historical period. The third object sub-embedding vector is determined by averaging the embedding vectors of each object that performed a specified operation on the multimedia content within a certain historical period but did not perform the target operation.

11. An information prompting device, characterized in that, include: The feature acquisition unit is used to acquire the content feature information of each multimedia content, and to acquire the statistical feature information of each multimedia content based on the historical behavior of each object towards each multimedia content. The time determination unit is used to obtain the predicted playback completion rate of each multimedia content based on its content feature information and statistical feature information; and to determine the information prompt time point of each multimedia content based on its predicted playback completion rate and playback duration. The object filtering unit is configured to perform the following operations for each multimedia content: based on the content embedding vector of a multimedia content and the object embedding vector of at least one of the objects associated with the multimedia content, filter out at least one matching target object. The information prompting unit is used to provide target information prompts to at least one target object when each multimedia content is played to the corresponding information prompting time point.

12. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of any of the methods described in claims 1 to 10.

13. A computer-readable storage medium, characterized in that, It includes a computer program that, when run on an electronic device, causes the electronic device to perform the steps of any of the methods described in claims 1 to 10.

14. A computer program product, characterized in that, The method includes a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, causing the electronic device to perform the steps of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Content recommendation method and device, computer equipment and storage medium

    CN111708941A

  • Information prompting method, device and system, client, server and storage medium

    CN112347276A

  • Content recommendation method and device, electronic equipment and storage medium

    CN113641916A