Content push method, device and computer storage medium

By using long and short-term memory neural networks and deep neural networks to predict the playback volume of advertising resources and dynamically adjust the playback strategy, the problem of low accuracy of existing hypercast control algorithms is solved, and precise control of advertising resources playback is achieved, avoiding hypercasting.

CN112785328BActive Publication Date: 2025-05-23TENCENT TECH SHANGHAI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011135302.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-21
Publication Date
2025-05-23
Estimated Expiration
2040-10-21

AI Technical Summary

Technical Problem

The existing hypercast control algorithm relies on human-set thresholds, with low accuracy, and the threshold cannot be dynamically adjusted to adapt to the characteristics of different orders and real-time playback conditions, making it difficult to accurately control the hypercast of advertising resources.

Method used

By obtaining the playback data within a predetermined time period of the resource to be played, the resource playback feature vectors are obtained one by one with multiple consecutive time slices, and a pre-trained long and short-term memory neural network model and deep neural network model are input to predict the playback amount, and decide whether to stop playing based on the prediction results to avoid hyperplay.

Benefits of technology

It realizes dynamic adjustment of playback strategies based on real-time playback data, improves the playback accuracy of advertising resources, and avoids the occurrence of super broadcast phenomena.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112785328B_ABST
    Figure CN112785328B_ABST
Patent Text Reader

Abstract

The present application discloses a content push method, comprising: in response to a received resource acquisition request, obtaining a matching resource to be played; obtaining playback data within a predetermined time period of the resource to be played; processing the playback data to obtain a resource playback feature vector corresponding to a plurality of continuous time slices; inputting the plurality of resource playback feature vectors into a pre-trained long short-term memory neural network model to obtain a first input vector; extracting statistical sparse features from the playback data; processing the statistical sparse features to obtain a second input vector; inputting the first input vector and the second input vector into a pre-trained deep neural network model to obtain a predicted playback volume; if the sum of the playback volume of the resource to be played and the predicted playback volume is greater than or equal to the predetermined playback volume, then stopping playback of the resource to be played. The above method can avoid overbroadcasting of advertising resources. The present application also discloses a content push device, a server, and a computer storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a content push method, device and computer storage medium. Background Art

[0002] Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time frame. It is a massive, high-growth, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery, and process optimization capabilities. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process large amounts of data within a tolerable time frame. Technologies applicable to big data include large-scale parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.

[0003] Based on big data technology, many special forms of products and services can be realized. For example, the ad push service based on user interests can push content that users are interested in to users without affecting their use of the product. The media platform can estimate the available inventory of each ad space under various targeting conditions, and then reasonably sell ad playback resources to customers. Currently, most advertisements are sold through signed guaranteed contracts. Customers can purchase a fixed amount of exposure for contracted advertisements. When the exposure of the advertisement reaches the exposure required in the contract, the media platform can stop broadcasting the advertisement accordingly.

[0004] The current overplay control algorithm is relatively simple. It makes a rough estimate of the online playback situation through prior knowledge, and then sets a threshold, such as 95% of the maximum playable rate, and observes the playback progress. When the threshold is reached, the stop order flag is set to 1, i.e. the order is stopped.

[0005] There are a lot of problems with current technology: first, the threshold is given by prior knowledge, which is a subjective will of people and has a very low accuracy rate; second, the threshold is the same for all orders, but the number of contract advertising orders is tens of thousands, and the audience targeting and advertising position of each order are different. The same threshold is obviously unreasonable. Finally, online advertising playback is controlled by various factors, such as frequency, directional crowding, mixing algorithms, etc. The delay time and the playback volume during the delay period are changing almost every moment. It is difficult to accurately control overbroadcasting with an unchanging threshold.

[0006] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Summary of the invention

[0007] The embodiments of the present application provide a resource intelligent playback control method, device and computer storage medium.

[0008] In a first aspect, an embodiment of the present application provides a content push method, including:

[0009] In response to the received resource acquisition request, a matching resource to be played is acquired, where the resource to be played has a corresponding predetermined playback volume;

[0010] Obtaining playback data of the resource to be played within a predetermined time period;

[0011] Processing the playback data to obtain resource playback feature vectors corresponding to a plurality of continuous time slices;

[0012] Inputting the plurality of resource playback feature vectors into a pre-trained long short-term memory neural network model to obtain a first input vector;

[0013] Extracting statistically sparse features from the playback data;

[0014] Processing the statistical sparse feature to obtain a second input vector;

[0015] Inputting the first input vector and the second input vector into a pre-trained deep neural network model to obtain a predicted playback volume; and;

[0016] If the sum of the played amount of the resource to be played and the predicted played amount is greater than or equal to the predetermined played amount, the playing of the resource to be played is stopped.

[0017] In a second aspect, an embodiment of the present application provides a content push device, including:

[0018] A resource acquisition module, configured to acquire a matching resource to be played in response to a received resource acquisition request, wherein the resource to be played has a corresponding predetermined playback volume;

[0019] A playback data acquisition module, used to acquire playback data of the resource to be played within a predetermined time period;

[0020] A first data processing module, used for processing the playback data to obtain resource playback feature vectors corresponding to a plurality of continuous time slices;

[0021] A first vector extraction module is used to input the multiple resource playback feature vectors into a pre-trained long short-term memory neural network model to obtain a first input vector to process the time series data into a first input vector based on the time series;

[0022] A second data processing module, used for extracting statistical sparse features from the playback data;

[0023] A second vector extraction module, used for processing the statistical sparse feature to obtain a second input vector;

[0024] A prediction module, configured to input the first input vector and the second input vector into a pre-trained deep neural network model to obtain a predicted playback volume; and

[0025] The playback control module is used to stop playing the resource to be played if the sum of the playback amount of the resource to be played and the predicted playback amount is greater than or equal to the predetermined playback amount.

[0026] In a third aspect, an embodiment of the present application provides a smart device, comprising: a memory; one or more processors coupled to the memory; one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the content push method provided in the first aspect above.

[0027] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a program code is stored. The program code can be called by a processor to execute the content push method provided in the first aspect above.

[0028] According to the content push method, device and server provided by the above scheme, by converting discrete category feature codes into vectors, and further using a neural network model to predict the actual future playback volume, overbroadcasting of advertising resources can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0030] Figure 1A schematic diagram of the architecture of a content push system provided in an embodiment of the present application is shown.

[0031] Figure 2 A flow chart of a content push method provided by an exemplary embodiment of the present application is shown.

[0032] Figure 3 Shows Figure 2 A schematic flow chart of overcast control in the method shown.

[0033] Figure 4 Shows Figure 2 Schematic diagram of the input vectors processed in the method shown.

[0034] Figure 5 A schematic diagram of the structure of a DNN model provided by an exemplary embodiment of the present application is shown.

[0035] Figure 6 A schematic diagram of the structure of a DNN model provided by an exemplary embodiment of the present application is shown.

[0036] Figure 7 A schematic diagram of the structure of an LSTM model provided by an exemplary embodiment of the present application is shown.

[0037] Figure 8 A schematic diagram of the network structure of an LSTM model provided by an exemplary embodiment of the present application is shown.

[0038] Fig. 9 A training flowchart of an LSTM model provided by an exemplary embodiment of the present application is shown.

[0039] Fig.10 A training flowchart of a DNN model provided by an exemplary embodiment of the present application is shown.

[0040] Fig.11 A flow chart of a content push method provided by an exemplary embodiment of the present application is shown.

[0041] Fig.12 A flow chart of a content push method provided by an exemplary embodiment of the present application is shown.

[0042] Fig.13 A structural block diagram of a content push device provided by an exemplary embodiment of the present application is shown.

[0043] Fig.14 Shows Fig.13 The effect schematic diagram of the content push device is shown. DETAILED DESCRIPTION

[0044] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0045] See also Figure 1 , which is a schematic diagram of the architecture of a content push system provided in an embodiment of the present application. The content push system includes a content push server 10 and a media client 20.

[0046] The content push server 10 is used to return the content to be played according to the request of the media client 20. After receiving the content acquisition request sent by the media client 20, the content push server 10 will acquire the corresponding content to be played and return it to the media client 20.

[0047] The media client 20 may include, for example, a computer, a laptop, a tablet computer, a mobile phone, etc. Various applications such as a browser or a shopping application may be run in the media client 20. When these applications are running, they will request various content data from the content push server 10, that is, they will send various content acquisition requests to the content push server 10.

[0048] Specifically, the content push server 10 includes: a processor 11, a main memory 12, a non-volatile memory 13, and a network module 14. The processor 11 and the main memory 12 are connected via a first bus 17. It can be understood that the first bus 17 here is only for illustration and is not limited to a physical bus. Any hardware architecture and technology that can connect the main memory 12 to the processor 11 can be used.

[0049] The main memory 12 is generally a volatile memory, such as a dynamic random access memory (DRAM).

[0050] The non-volatile memory 13 and the network module 14 are connected to the first bus 17 via an input / output (IO) bus 18, and can interact with the processor 11. The IO bus can be, for example, a peripheral component interconnect (PCI) bus or a high-speed serial computer expansion bus (PCI-E).

[0051] The nonvolatile memory 13 may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read Only Memory), an EPROM, a hard disk, or a ROM.

[0052] The network module 14 can be connected to the content push server 10 via a network signal.

[0053] The non-volatile memory 13 stores a content engine 131 and an anti-overbroadcast model 132 .

[0054] The content engine 131 is responsible for receiving a content acquisition request sent by the media client 10, acquiring object features according to the content acquisition request, and retrieving all matching content resources in the resource library to be played according to the object features.

[0055] In a specific implementation, the anti-overcast model 132 is a neural network model.

[0056] In a specific implementation, the above-mentioned content resources refer to advertisements, for example. The advertisements may be in various forms such as text, pictures, videos, or animations.

[0057] In a content push system, there are several participants: resource playback service providers, resource playback service delivery entities, and users. Among them, resource playback service providers provide services to users and resource playback service delivery entities through the above-mentioned content push system. Taking the advertising system as an example, the resource playback service delivery entity is the advertiser. When the advertiser needs to deliver an advertisement, he can sign a contract order with the advertising resource provider and specify the advertisements to be delivered, the expected playback volume, time limit, targeting conditions, etc. in the contract order.

[0058] The targeting condition refers to the target of advertisement delivery specified by the advertiser, which may include any one of the following targeting conditions or a combination of multiple targeting conditions:

[0059] (1) Geographic targeting: For example, many advertisers’ businesses have regional characteristics;

[0060] (2) Demographic attribute targeting: mainly including age, gender, education level, etc.;

[0061] (3) Channel targeting: Suitable for vertical media that are closer to conversion needs, such as automobiles, maternity and baby products, with a narrow coverage;

[0062] (4) Contextual targeting: matching relevant ads based on the specific content of the web page (e.g., keywords, topics, etc.);

[0063] (5) Behavioral targeting: understanding user interests based on their historical visit behavior and placing advertisements accordingly;

[0064] (6) Precise location targeting: Advertising is delivered based on precise geographic location based on terminal device information (e.g. Global Positioning System (GPS) positioning information, Internet Protocol (IP) address, etc.).

[0065] See also Figure 2 , which is a flow chart of a content push method provided by an exemplary embodiment of the present invention. The method can be executed by the above-mentioned content push system, for example, to control the playback of resources. The method includes the following steps:

[0066] Step S1, receiving a resource acquisition request.

[0067] In a specific implementation, the above-mentioned media client 20 is installed with an application, and the application here may include various applications such as news client, browser, instant messaging software, video player software, novel reader, etc. In addition to requesting normal browsing content data from the background server, the application will also send a resource acquisition request to the content push server 10 to obtain the resources to be played, so as to display and play them in the media client 20. Correspondingly, the content push server 10 will receive the above-mentioned resource acquisition request.

[0068] In a specific embodiment, the resource acquisition request may include a user identifier such as a user token. The content push server 10 extracts the user token from the resource acquisition request, uses the user token to find the corresponding user in the user data, and further obtains the user's characteristic information. The user token here can be, for example, a Json web token (JWT) defined in the open standard RFC7519, or a session ID (Session), or any other identifier that can uniquely indicate the user's identity.

[0069] After receiving the above-mentioned resource acquisition request, the content push server 10 will obtain the corresponding multimedia resources and return them to the media client 20 for display or playback. The multimedia resources here refer to advertisements, for example. It can be understood that advertisements are defined from the perspective of content. In terms of presentation, multimedia resources may include any combination of text, pictures, animations, and videos, and may even be program scripts used to generate text, pictures, animations, and videos. Any technology that can be run by the media client 20 and display the corresponding content in its interface for users to watch can be used. The following will describe the process of the content push server 10 returning the corresponding multimedia resources according to the above-mentioned resource acquisition request in combination with more specific embodiments as follows.

[0070] Step S2: Recall resources according to the resource acquisition request.

[0071] In a specific implementation, firstly, user information may be obtained according to the user identifier included in the resource acquisition request. The object feature information here may include object feature data of different dimensions, such as the user's gender, age, education level, interests, geographic location, etc.

[0072] On the other hand, for all resources to be played, each resource to be played can set one or more object features as matching conditions, or subscription conditions. Therefore, for each current user, according to these set matching conditions, it can be determined whether the resource to be played matches the user.

[0073] For example, when a male user in Shanghai opens a music application, all advertisements will be retrieved. Assume that there are only three advertisements A, B, and C, among which A is reserved for Shanghai women, B is reserved for Shanghai men, and C is reserved for men. Then, advertisements B and C match the user, but advertisement A does not match the user.

[0074] The goal of this step is to obtain at least part of the resources to be played that match the current user. Generally speaking, a relational database can be used to store the matching relationship between different resources to be played and object features. Then, the object features can be used as a query condition to query in the relational database to obtain at least part of the above-mentioned resources to be played. It can be understood that the relational database here is taken as an example, and the storage method of the matching relationship between object features and resources to be played is not subject to any restrictions.

[0075] It is understandable that since the number of resources to be played is generally very large, even if object feature information is used for matching search, the number of resources to be played obtained in the end may be very large, and it is impossible to display them all at once in the user interface. Therefore, it is necessary to decide one or several resources to be played for display and playback based on these recalled resources to be played.

[0076] Step S3: Filter the recalled resources to be played to obtain resource matching results.

[0077] In a specific implementation, a user model can be established based on object features. On the other hand, for each resource to be played, keywords can be extracted from the text and image of the resource to be played, and the matching degree between the user and each resource to be played is calculated based on the user model and the keyword. The resources to be played obtained in step S2 can be sorted according to the matching degree, and then only the resources to be played whose matching degree is within a predetermined range are taken, and those whose matching degree is lower than the predetermined range can be discarded in the current request.

[0078] Furthermore, the above process directly searches for matching resources through static features of users such as gender and age. However, these features only direct the display of resources to a specific user group. This targeted object feature matching does not reflect the user's own interest characteristics. Therefore, in a content push system, it is also possible to collect user interest data of different dimensions, and then calculate the matching degree between the user's interest data and the resources to be played, so that the resources to be played can be sorted according to the user's interest, and then only the resources to be played whose user interest is within the predetermined range are taken, and those whose matching degree is lower than the predetermined range can be discarded in the current request and no longer returned.

[0079] In a specific implementation, the resources to be played can also be filtered according to the similarity between the content currently browsed by the user and the resources to be played. For example, keywords can be extracted from the content currently browsed by the user and the resources to be played, and then semantic analysis can be performed, and then the similarity between the content currently browsed by the user and the resources to be played can be determined based on the semantic analysis results. The higher the similarity, the higher the ranking of the corresponding resources to be played, and the more likely they are to be retained and returned to the media client 20.

[0080] It is understandable that even after the above sorting and screening process, there are still many resources left to be played, but the number of resources that can be displayed per request is limited. Therefore, in any case, only a specified number of resources to be played can be obtained and returned from the resources to be played that meet the standards.

[0081] In a specific implementation, a random function can be used to determine which resources to be played can be kept. For example, a probability value (Rate), such as 0.5, can be given to each resource to be played, and a random function with a uniform distribution of 0-1 is used to generate a random number for each resource to be played, or if the random number is greater than the Rate value, the resource to be played is kept, otherwise the resource to be played is not displayed this time. Therefore, by adjusting the rate value of each resource to be played, the probability of the resource to be played being displayed can be controlled.

[0082] Step S4, returning the above resource matching result.

[0083] After obtaining a sufficient number of resources to be played, the content push server 10 can return them to the media client 20. Accordingly, after receiving the resources to be played, the media client 20 displays them in its user interface, that is, displays and plays the above text, pictures, animations, and videos.

[0084] In a specific implementation, in the above step S3, in addition to the above sorting and screening process, the overbroadcast control process can also be performed on the recalled resources to be played. Figure 3 , the supercast control process includes the following steps:

[0085] Step S31, obtaining the playback data of each resource to be played.

[0086] exist Figure 1 In the content push system shown, each resource playback operation of the media client 20 can also be recorded. Specifically, the display playback time of a certain resource to be played in the media client 20 and the corresponding object feature information can be recorded. These playback data can be stored in a database for easy query and use. In the method provided in this embodiment, based on these playback data, it can be predicted whether a certain resource to be played will be overplayed within a predetermined time range in the future, and based on the prediction result, it can be controlled whether the resource to be played needs to continue to be pushed to the media client.

[0087] Overbroadcast here refers to the phenomenon that the display playback volume of a resource to be played exceeds the scheduled playback volume. This scheduled playback volume has been pre-set or dynamically acquired. Generally speaking, the resource to be played is associated with an order contract for the release of a resource, and the order contract can stipulate the scheduled playback volume. In this case, this scheduled playback volume can be booked in advance. But it can be understood that the pre-set here does not mean just a fixed value. For example, the scheduled playback volume can be set for each day, the scheduled playback volume can be set for each hour, the scheduled playback volume for users in different regions can be set, and the scheduled playback volume for users of different genders can be set. In other words, this scheduled playback volume can be related to object characteristics, time, or geographic location. But no matter what the specific setting strategy is, the scheduled playback volume will have an upper limit. If it exceeds this upper limit, it will be considered overbroadcast.

[0088] In another scenario, the scheduled playback volume is not directly agreed upon in advance by the order contract, but can be dynamically set by the computer system of the resource playback and delivery entity through the application programming interface. In this case, the resource playback and delivery entity can dynamically obtain the statistical data of resource playback and dynamically adjust the scheduled playback volume based on different user feedback. In this scenario, the content push system is required to open the application programming interface for the resource playback and delivery entity to call.

[0089] According to the specific statistical results of resource playback data, overall, the actual playback volume of resources changes relatively smoothly, but it will show different characteristics at different time points, which shows that the actual playback volume of resources is closely related to time. In addition, the actual playback volume of resources is also related to the object feature information related to the resources to be played. Suppose there are resources D1 and D2 to be played, and their scheduled playback volume is 1000, but the object features associated with them are different. Then, under the premise of not performing overplay control, the actual playback volume curve growth of the resources is also different. In summary, if you want to predict the playback volume of a resource to be played at a scheduled time in the future based on the existing playback data of the resource to be played, you need to consider both the time factor and the object feature information.

[0090] Step S32: Process the acquired playback data to obtain resource playback feature vectors corresponding to a plurality of continuous time slices.

[0091] The purpose of processing the acquired playback data is to establish a neural network model, so that the resource playback volume within a predetermined time in the future can be predicted based on the resource playback data in the past, and the overplay control of the resource playback can be further performed based on the prediction results. For the application scenario of resource playback control, the playback volume of the next 1 minute can generally meet the accurate overplay control requirements. In other words, the goal is to predict the resource playback volume within the next minute based on the resource playback data in the past. Of course, the 1 minute here is only for illustration, and the actual time is not subject to any restrictions.

[0092] For scenarios where the future is predicted based on past data, the Long Short-Term Memory (LSTM) model can be used to make the above predictions. For the LSTM model, the input features must meet two requirements: the features must be expressed as vectors, and the features must be stacked based on time slices to form stacked features that can be applied to the LSTM model.

[0093] For the LSTM model, its input feature is a three-dimensional matrix, which is organized according to the structure of (batch size, time step, feature). Specifically, the batch size refers to the number of time slices corresponding to a batch of playback data. The time slice refers to the minimum time unit for resource playback data statistics. The time step refers to the number of time slices to be stacked for features, which corresponds to a continuous time period. Features refer to the multidimensional vectors obtained according to the above steps.

[0094] In a specific implementation, the above multidimensional vector may include the following information: exposure, inventory, placement, order, etc. Among them, exposure refers to the amount of playback, and inventory refers to all available playbacks.

[0095] In a specific implementation, the batch size is 10, the time step is 5, and the length of the time slice is 1 minute. That is, the goal of this input feature construction method is to use the resource playback data of the past 10 minutes to predict the resource playback volume of the next 1 minute. It can be understood that the batch size, time step, and time slice length are not limited to the above exemplary values ​​and can be adjusted according to actual conditions.

[0096] Assume that there is resource playback data within 10 minutes, that is, the resource playback data of 10 time slices are marked as 1-10 respectively. If the time step is 5, the feature vectors corresponding to the time slices 1, 2, 3, 4, and 5 are stacked, the feature vectors corresponding to the time slices 2, 3, 4, 5, and 6 are stacked, the feature vectors corresponding to the time slices 3, 4, 5, 6, and 7 are stacked, the feature vectors corresponding to the time slices 4, 5, 6, 7, and 8 are stacked, and so on, until all the time slices are used. Assuming the batch size is N and the time step is S, the actual number of stacked features is N-S+1.

[0097] In a specific embodiment, see Figure 4 , which is a schematic diagram of resource playback feature vectors corresponding to multiple consecutive time slices after processing. The first dimension is the number of training data samples in each batch, which is 100 in this implementation, the time step is 5, and the feature dimension is 4, specifically the four features of exposure, inventory, layout ID, and order ID.

[0098] Step S33: acquiring a first input vector according to the resource playback feature vectors corresponding one-to-one to a plurality of continuous time slices.

[0099] First, the existing resource playback data is used to train the above LSTM model. That is, training samples are constructed based on the existing resource playback data, and the actual playback volume is used to verify the model, and the training is continued until the above LSTM model converges.

[0100] After the model is trained, the playback data in the recent period is used to construct a batch of input features, and the first input vector is obtained by inputting the trained model. It is worth noting that in this model, the output of the LSTM model is not the direct playback volume, and the first input vector of its output is a dense matrix.

[0101] Step S34: Process the sparse features in the acquired playback data to obtain a second input vector.

[0102] It can be understood that the input feature in the neural network model is a kind of expression of data, generally represented by a vector. Assuming its dimension is N*N, it means that the vector is composed of N*N numbers. If the number of 0 values ​​in the N numbers is greater than the number of non-zero values, the feature can be defined as a sparse feature. The higher the proportion of 0, the higher the sparsity.

[0103] In the content push system, category feature information such as placement, regional classification, crowd label, media label, basic attributes, etc. are sparse features.

[0104] The placement here refers to, for example, the location where the resource is played; the regional classification refers to the user's geographical area, which may generally be detailed to provinces, cities, districts, or even streets; the population label refers to the user's group label such as literary youth; the media label refers to, for example, different applications; and the basic attributes refer to, for example, the user's gender, age, etc.

[0105] In a specific implementation, the user's category feature information can be encoded, for example, using One-Hot encoding to convert discrete feature data into a vector. One-Hot will map the category feature to a vector based on the number of values. Each bit in the vector represents whether the sample has entered this feature. For example, gender has two values, male and female. After encoding, male will be mapped to (0,1) and female will be mapped to (1,0). For example, if the media is divided into application A, application B, application C, and application D, then application A will be mapped to (1,0,0,0). By analogy, the category feature data can be converted into a vector.

[0106] It is worth noting that if there are two category features, two vectors are obtained. If there are more category features, there will be more vectors. Multiple vectors can be embedded (Embedding) to map multiple vectors into a multidimensional vector.

[0107] Step S35, obtaining the predicted playback volume according to the first input vector and the second input vector.

[0108] In a specific implementation, the first input vector and the second input vector are concatenated to form a third input vector, and the third input vector is input into a pre-trained deep neural network model (Deep Neural Networks, DNN) to obtain the above-mentioned predicted playback volume.

[0109] In a specific embodiment, see Figure 5 , which is a schematic diagram of the structure of the adopted DNN model. The deep neural network includes: an input layer 210, a dense layer 220, a hidden layer 230, and an output 240.

[0110] The input layer 210 is used to process dense features and statistical sparse features respectively.

[0111] The dense features are actually the output results of the above LSTM model.

[0112] Statistically sparse features refer to all features except time-related features, which can include historical playback speed and directional features: age, gender, region, etc.

[0113] The upper DNN model is used to integrate time series and non-time series features and then make predictions. In a specific implementation, the DNN model is a multi-layer perceptron (MLP) model, which is composed of multiple common perceptrons, such as Figure 6 shown.

[0114] Figure 6 The input layer shown in the figure is a new input vector formed by concatenating the dense matrix output by the LSTM model and the vector obtained by the statistical sparse feature mapping. The middle hidden layer is three layers, and the output layer is the predicted future playback volume.

[0115] In a specific implementation, the above-mentioned splicing method is explained as follows: for example, the vector produced by the LSTM model is (0.52, 0.36, 0.25, 0.14, 0.896), and the features produced by the statistical sparse features are (0.89, 0.56, 0.25, 0.45, 0.578, 0.56). The result after splicing is (0.52, 0.36, 0.25, 0.14, 0.896, 0.89, 0.56, 0.25, 0.45, 0.578, 0.56).

[0116] Step S36: Perform overplay control based on the predicted play volume obtained above.

[0117] In a specific implementation, the above-mentioned overbroadcast control includes the following steps: determining whether the current playback volume plus the predicted playback volume is greater than or equal to the predetermined playback volume, if so, pausing the playback display of the resource to be played, if not, continuing the playback display of the resource to be played.

[0118] In a specific implementation, the resources whose playback is suspended can be directly stopped from being returned to the media client, that is, the delivery of the resources that are about to be overbroadcast is suspended.

[0119] In another specific embodiment, the subject that executes the overbroadcast control is not the same as the subject that ultimately returns the resources to the media client 20. In this case, a stop order flag for indicating the suspension of the playback display can be directly given to the resource to be played or its corresponding order contract. The stop order flag can be stored in a database or cache system to facilitate the call of the subject that executes the return of resources to the media client 20.

[0120] According to the content push method provided by this embodiment, the features related to the time series password are processed by the LSTM model, and input into the DNN model together with the statistical sparse features, and the existing data is used to predict the future playback volume, so that the overbroadcast of advertising resources can be avoided according to the predicted playback volume. Figure 5 The figure shows a comparison before and after the method of this embodiment is used. The left side shows the case where the method of this embodiment is not used. It can be seen that there is an obvious overcasting phenomenon. After the method of this embodiment is used, the overcasting phenomenon is basically eliminated. Figure 7 , which is a schematic diagram of the network structure of the LSTM model provided by an exemplary embodiment of the present application. Figure 7 As shown, the LSTM model includes: a data input layer 11, a coding layer 12, an embedding layer 13, a stacking layer 14 and a model layer 15.

[0121] The data input layer 101 is used to obtain input data features, where the data features generally refer to features related to the time password, such as the predetermined amount, the playback amount, the playback speed, etc.

[0122] The encoding layer 102 is used to encode the categories provided by the data input layer 101, for example, using One-Hot encoding to convert discrete feature data into vectors. One-Hot will map the category feature to a vector according to the number of values. Each bit in the vector represents whether the sample has entered this feature. For example, gender has two values, male and female. After encoding, male will be mapped to (0,1) and female will be mapped to (1,0). For example, if the media is divided into application A, application B, application C, and application D, then application A will be mapped to (1,0,0,0).

[0123] The embedding layer 103 is used to map the vector output by the encoding layer 102 and the numerical features provided by the data input layer 11 into a vector of a specified dimension. For example, it can be mapped into a 16-dimensional vector.

[0124] The output of the embedding layer 103 can be input into the stacking layer 104 for input feature stacking to obtain training samples.

[0125] LSTM belongs to the Recurrent Neural Network (RNN) model, which requires constructing samples into a three-dimensional matrix of (batch size, time step, feature) dimensions, namely, feature stacking, where the batch size is the size of a batch of features, and the features are the features of the order, orientation, and playback levels.

[0126] The construction method is as follows: for example, if there are 10 time slices and the time step is 5, then the data of time slices 1, 2, 3, 4, 5 are combined together, 2, 3, 4, 5, 6 are combined together, 3, 4, 5, 6, 7 are combined together, and so on. The length of the time slice can be, for example, 1 second, 5 seconds, 10 seconds, 20 seconds, 30 seconds, 1 minute, or longer.

[0127] See also Figure 8 , which is a schematic diagram of the structure of an LSTM model provided by an exemplary embodiment of the present application. The LSTM neural network model structure includes multiple memory blocks, each memory block includes a forget gate f t , input gate i t , output gate O t and a memory unit C t . The horizontal line represents the cell state (Cell State)

[0128] The LSTM neural network model includes a four-layer structure, specifically including a first neural network layer 103, a second neural network layer 104, a third neural network layer 105 and a fourth neural network layer 106. The first neural network layer 103, the second neural network layer 104 and the fourth neural network layer 106 are sigmoid neural network layers, and the third neural network layer 105 is a tanh neural network layer.

[0129] The reason why LSTM neural network has "memory" is that there are connections between the networks at different "time points", rather than feedforward or feedback at a single time point. Figure 6 The hidden layers shown are connected by arrows, where the arrows represent jump connections between neural units in a sequence of time steps.

[0130] The first step of LSTM is to decide what information can be passed through the cell state. This decision is controlled by the forget gate ft layer through the first neural network layer 103, which will be based on the output h of the previous moment. t-1 and the current input x t To generate an f between 0 and 1 t value, to decide whether to let the information C learned at the last moment t-1 Passed or partially passed. As follows:

[0131] f t =σ(Wf ·[h t-1 ,x t ]+b f )

[0132] The second step of LSTM is to generate new information that needs to be updated. This step consists of two parts. First, the input gate i t The second neural network layer 104 is used to determine which values ​​are used for updating, followed by the third neural network layer 105 to generate new candidate values, which may be added to the unit state as the candidate values ​​generated by the current layer. The values ​​generated by these two parts can be combined for updating.

[0133] i t =σ(W i ·[h t-1 ,x t ]+b i )

[0134]

[0135] Then the cell state is updated. First, the old cell state is multiplied by ft to forget the unnecessary information, and then Add them together to get the candidate value.

[0136]

[0137] The last step is to determine the output of the model. First, an initial output is obtained through the fourth neural network layer 106, and then the Ct value is scaled to between -1 and 1 using the third neural network layer 105, and then multiplied pairwise with the output obtained by the fourth neural network layer 106 to obtain the output of the model.

[0138] O t =σ(W o [h t-1 ,x t ]+b o )

[0139] h t =O t *tanh(C t )

[0140] As a deep model, LSTM must have dense input data. However, as mentioned above, the relevant parameters in the advertising playback process are all category features, which are not dense. Therefore, the existing LSTM model needs to be modified.

[0141] See also Fig. 9 , which is a flow chart of the training method of the above-mentioned LSTM model, the training method comprises the following steps:

[0142] Step S101, recording resource playback data.

[0143] exist Figure 1 In the content push system shown, after returning the resource to be played to the media client 20, the playing data of the resource can also be recorded. The playing data of the resource can be stored in the database system.

[0144] In a specific embodiment, the resource playback data mentioned above may include: category feature information and numerical data. Among them, the category feature information is a label used to describe the user classification attribute, and its value range will be one of multiple options. For example, taking the user's gender as an example, it may be one of the following options: male, female, unknown. In addition, the user's gender object feature will not have other options. Of course, the number of options here is only for illustration, and gender may also include different definitions, but in either case, the possibility of its options is certain. Taking the user's geographic location as an example, if it is subdivided into the user's province, for Chinese users, there are only more than 30 possible value options. By analogy, such options can only be to take one or more object features from a fixed list of options to define the above-mentioned category feature information.

[0145] In contrast, some other resource playback data, such as scheduled playback volume and actual playback volume, may have a continuously distributed distribution range. This type of data is defined as numerical data in the above technique.

[0146] In a specific implementation, recording the resource playback data includes adding a new resource playback record in a database, and the resource playback record may include an identifier of the resource being played and the resource playback data.

[0147] Step S102: Process the recorded resource playback data into a feature vector.

[0148] As mentioned above, to predict the playback volume of a resource to be played at a scheduled time in the future based on its existing playback data, it is necessary to consider both the time factor and the object feature information. The value of the category feature information in the object feature is discrete and cannot be used directly as a vector.

[0149] In a specific implementation, the user's category feature information can be encoded, for example, using One-Hot encoding to convert discrete feature data into a vector. One-Hot will map the category feature to a vector based on the number of values. Each bit in the vector represents whether the sample has entered this feature. For example, gender has two values, male and female. After encoding, male will be mapped to (0,1) and female will be mapped to (1,0). For example, if the media is divided into application A, application B, application C, and application D, then application A will be mapped to (1,0,0,0). By analogy, the category feature data can be converted into a vector.

[0150] It is worth noting that if there are two category features, two vectors are obtained. If there are more category features, there will be more vectors. These vectors cannot be directly applied to the LSTM model, and embedding processing is required to map multiple vectors into a multidimensional vector. At the same time, the data of the scheduled playback volume and the actual playback volume of the resource to be played can also be mapped to the multidimensional vector.

[0151] Step S103, stacking the processed feature vectors by time steps to obtain training samples.

[0152] For the LSTM model, its input feature is a three-dimensional matrix, which is organized according to the structure of (batch size, time step, feature). Specifically, the batch size refers to the number of time slices corresponding to a batch of resource playback data. The time slice refers to the minimum time unit for resource playback data statistics. The time step refers to the number of time slices to be stacked for features, which corresponds to a continuous time period. Features refer to the multidimensional vectors obtained according to the above steps.

[0153] In a specific implementation, the batch size is 10, the time step is 5, and the length of the time slice is 1 minute. That is, the goal of this input feature construction method is to use the resource playback data of the past 10 minutes to predict the resource playback volume of the next 1 minute. It can be understood that the batch size, time step, and time slice length are not limited to the above exemplary values ​​and can be adjusted according to actual conditions.

[0154] Assume that there is resource playback data within 10 minutes, that is, the resource playback data of 10 time slices are marked as 1-10 respectively. If the time step is 5, the feature vectors corresponding to the time slices 1, 2, 3, 4, and 5 are stacked, the feature vectors corresponding to the time slices 2, 3, 4, 5, and 6 are stacked, the feature vectors corresponding to the time slices 3, 4, 5, 6, and 7 are stacked, the feature vectors corresponding to the time slices 4, 5, 6, 7, and 8 are stacked, and so on, until all the time slices are used. Assuming the batch size is N and the time step is S, the actual number of stacked features is N-S+1.

[0155] Step S104: Use the stacked training samples to train the LSTM model.

[0156] In a specific implementation, the output parameter of the LSTM model can be set to a 1*1 vector, that is, a numerical value, and the meaning of the data value is the estimated future playback volume.

[0157] In a specific implementation, the square of the difference between the estimated playback volume and the actual playback volume is used as a parameter for model deviation (loss) evaluation.

[0158] In a specific embodiment, the gradient descent method is used for model training.

[0159] It can be understood that the training process of the neural network model is essentially the process of finding the optimal model parameters. The definition of the optimal model parameters here refers to making the difference between the model's predicted results and the actual values ​​less than the specified range, that is, the square of the difference between the estimated playback volume and the actual playback volume is less than the specified range. If this condition is met, the model is considered to have successfully converged and the training is successful.

[0160] It is worth noting that after the above LSTM model training is completed, its output can be set to output a feature vector instead of directly outputting the predicted playback volume.

[0161] See also Fig.10 , which is a schematic diagram of the training process of the above-mentioned DNN. The training process includes the following steps:

[0162] Step S201, recording the playback record and the corresponding feature data.

[0163] The playback records here refer to, for example, the playback time of each advertisement, as well as the corresponding order information, user information, etc. Various features can be proposed based on this information. Generally speaking, features can be divided into time sequence features and non-time series features. Among them, time series features refer to features that are closely related to time series, such as the amount of playback at different time points. Non-time series features, for example, refer to features that have no obvious correlation with time series, such as historical playback speed, age, gender, region, etc.

[0164] Step S202, mapping the non-time series features into a second input vector.

[0165] Specifically, the features are first mapped into vectors through one-hot encoding, and then embedded and concatenated into a dense vector together with the numerical features.

[0166] Step S203, processing the time series data into dense vectors and stacking them to input into the LSTM model.

[0167] Time series feature data refers to exposure, inventory, placement information, order attributes, etc. Time series data can be divided into categorical data and numerical data. The categorical data is encoded and mapped into a vector, and then concatenated with the numerical data into a dense vector. Furthermore, according to the time step, the data features are divided into multiple forms, and multiple samples are stacked to achieve the purpose of being output to the LSTM model.

[0168] Step S204, input the data generated in step S203 into the LSTM model to obtain a first input vector representing time series.

[0169] Step S205, concatenating the vectors generated in step S202 and step S204 to form a third input vector.

[0170] In a specific implementation, the above-mentioned splicing method is explained as follows: for example, the vector produced by the LSTM model is (0.52, 0.36, 0.25, 0.14, 0.896), and the features produced by the statistical sparse features are (0.89, 0.56, 0.25, 0.45, 0.578, 0.56). The result after splicing is (0.52, 0.36, 0.25, 0.14, 0.896, 0.89, 0.56, 0.25, 0.45, 0.578, 0.56).

[0171] Step S206, input the input vector into the DNN model.

[0172] In step S207, the DNN model outputs the estimated future playback volume, and the square of the difference between the estimated playback volume and the actual playback volume is used to evaluate the model loss. Repeat the above steps S204 to S207 until the model converges. After the converged model is obtained, the relevant data of the resources to be played to be predicted can be processed to obtain the input vector, and then the trained DNN model can be input to obtain the predicted playback volume. If the sum of the current playback volume plus the estimated playback volume is greater than or equal to the maximum daily playback volume, a stop order mark can be sent, otherwise continue to play.

[0173] In a specific implementation, the square of the difference between the estimated playback volume and the actual playback volume is used as a parameter for model deviation (loss) evaluation.

[0174] In a specific embodiment, the gradient descent method is used for model training.

[0175] It can be understood that the training process of the neural network model is essentially the process of finding the optimal model parameters. The definition of the optimal model parameters here refers to making the difference between the model's predicted results and the actual values ​​less than the specified range, that is, the square of the difference between the estimated playback volume and the actual playback volume is less than the specified range. If this condition is met, the model is considered to have successfully converged and the training is successful.

[0176] See also Fig.11 An exemplary embodiment of the present application provides a content push method, the method and Figure 3 The method shown is similar, except that step S36 comprises:

[0177] Step S361, obtaining a corresponding playback probability value according to the predicted playback volume.

[0178] As described above, in the process of retrieving resources, one request may retrieve a large number of matching resources to be played. In this case, one or more resources to be played need to be obtained from the matching resources and returned to the media client 20 .

[0179] In a specific implementation, a random function can be used to determine which resources to be played can be kept. For example, a probability value (Rate), such as 0.5, can be given to each resource to be played, and a random function with a uniform distribution of 0-1 is used to generate a random number for each resource to be played, or if the random number is greater than the Rate value, the resource to be played is kept, otherwise the resource to be played is not displayed this time. Therefore, by adjusting the rate value of each resource to be played, the probability of the resource to be played being displayed can be controlled.

[0180] In a specific embodiment, the above-mentioned playback probability value is positively correlated with the difference between the predicted playback volume and the actual playback volume, that is, the larger the difference is, the larger the playback probability value is, and a clear functional relationship can be formed between the playback probability value R and the difference, that is, R = f(d), wherein R represents the playback probability value of the resource to be played, and d represents the difference between the predicted playback volume and the actual playback volume. Of course, the positive correlation between the playback probability value R and the difference is not limited to forming a functional relationship. For example, it can also be through a simple mapping table, according to different intervals of the difference, directly giving different probability value coefficients, and the final playback probability value R needs to be multiplied by the probability value coefficient when calculating.

[0181] Step S362: Use the playback probability value to control resource filtering.

[0182] For each recalled resource, the above-mentioned playback probability value can be used to control whether to play. Specifically, a random function with a value that follows a uniform distribution of 0-1 is used to generate a random number for each resource to be played. If the random number is greater than the playback probability value, the resource to be played is kept, otherwise the resource to be played is not displayed this time.

[0183] It can be understood that since the difference between the predicted playback volume and the actual playback volume is positively correlated, when the difference is large, that is, a large proportion of the scheduled playback plan has not been completed, the resources to be played have a higher probability of playback, which can increase the playback speed of the resources to be played. When the difference is small, that is, the scheduled playback plan has been basically completed, the resources to be played have a smaller probability of playback, and their playback speed will decrease. In this way, the overall playback will gradually approach the scheduled playback volume at a relatively slow speed, thereby further reducing the probability of overbroadcasting.

[0184] See also Fig.12 An exemplary embodiment of the present application provides a content push method, the method and Figure 3 The method shown is similar, except that after step S36, it further includes:

[0185] Step S37, after waiting for a predetermined time, quantitative supplementary playback is performed according to whether the actual playback volume of the resource to be played is greater than or equal to the predetermined playback volume.

[0186] In the above-described embodiments, whether to stop playing a resource is determined based on the predicted playback volume. Specifically, when it is determined that the actual playback volume is about to exceed the playback volume, the playback of the resource is stopped. In actual application scenarios, there is still a probability that the actual playback volume of the resource still does not reach the predetermined playback volume, and this situation may be regarded as a breach of contract.

[0187] The above-mentioned predetermined time is not subject to any specific restrictions, as long as it is greater than the time from the issuance of the stop order instruction to the full implementation of the instruction and the complete cessation of playback of the specified resource. In a specific embodiment, the predetermined time may be, for example, 1 minute. That is, 1 minute after the issuance of the power outage sign, supplementary playback may be performed based on whether the actual playback is consistent with the predetermined playback amount.

[0188] Assuming that a special playback resource needs to be supplemented, the supplementary playback volume is the scheduled playback volume minus the actual playback volume. For the supplementary playback resources, a separate supplementary playback resource library can be set up, and each resource to be played stores the amount of supplementary playback required. Figure 1 In the content push system shown, resources in the supplementary playback resource library can be matched preferentially, so that the resource delivery contract for a certain time period can be completed in advance.

[0189] According to the content push method provided in this embodiment, after sending the stop order mark, quantitative supplementary playback is performed according to whether the actual playback volume of the resource to be played is greater than or equal to the predetermined playback volume. This can not only control the resource to be played from being overbroadcast, but also make its actual playback volume accurately equal to the actual playback volume, thereby improving the working efficiency of the content push system.

[0190] In a specific implementation, the above-mentioned media client 20 includes a video playback application. When the user operates the video playback application, the media client 20 terminal can be triggered to perform a video advertisement loading action in one or more interfaces. For example, when the user clicks to play a video that he wants to watch in the video playback application, the video advertisement loading action is triggered. In other words, the media client 20 will send a video advertisement acquisition request to the resource server 10.

[0191] After receiving the video advertisement acquisition request, the content push server 10 recalls all video advertisements matching the current object features from the video advertisement library. For the recalled original video advertisements, operations such as sorting and filtering can be performed to finally obtain a specified number of video advertisements, which are then sent to the media client 20. After receiving the video advertisement returned by the resource server 10, the media client 20 can play the video advertisement in the interface of the video playback application.

[0192] See also Fig.13 , which is a structural block diagram of a content push device provided in an exemplary embodiment of the present application. The device includes:

[0193] The resource acquisition module 310 is used to obtain a matching resource to be played in response to a received resource acquisition request, where the resource to be played has a corresponding predetermined playback volume.

[0194] In a specific implementation, the above-mentioned resource refers to, for example, an advertisement, which may be in the form of text, picture, video, or animation.

[0195] In a content push system, there are several participants: resource playback service providers, resource playback service delivery entities, and users. Among them, resource playback service providers provide services to users and resource playback service delivery entities through the above-mentioned content push system. Taking the advertising system as an example, the resource playback service delivery entity is the advertiser. When the advertiser needs to deliver an advertisement, he can sign a contract order with the advertising resource provider and specify the advertisements to be delivered, the expected playback volume, time limit, targeting conditions, etc. in the contract order.

[0196] The targeting condition refers to the target of advertisement delivery specified by the advertiser, which may include any one of the following targeting conditions or a combination of multiple targeting conditions:

[0197] (1) Geographic targeting: For example, many advertisers’ businesses have regional characteristics;

[0198] (2) Demographic attribute targeting: mainly including age, gender, education level, etc.;

[0199] (3) Channel targeting: Suitable for vertical media that are closer to conversion needs, such as automobiles, maternity and baby products, with a narrow coverage;

[0200] (4) Contextual targeting: matching relevant ads based on the specific content of the web page (e.g., keywords, topics, etc.);

[0201] (5) Behavioral targeting: understanding user interests based on their historical visit behavior and placing advertisements accordingly;

[0202] (6) Precise location targeting: Advertising is delivered based on precise geographic location based on terminal device information (e.g. Global Positioning System (GPS) positioning information, Internet Protocol (IP) address, etc.).

[0203] The playback data acquisition module 320 is used to acquire the playback data of the resource to be played within a predetermined time period.

[0204] exist Figure 1In the content push system shown, each resource playback operation of the media client 20 can also be recorded. Specifically, the display playback time of a certain resource to be played in the media client 20 and the corresponding object feature information can be recorded. These playback data can be stored in a database for easy query and use. In the method provided in this embodiment, based on these playback data, it can be predicted whether a certain resource to be played will be overplayed within a predetermined time range in the future, and based on the prediction result, it can be controlled whether the resource to be played needs to continue to be pushed to the media client.

[0205] Overbroadcast here refers to the phenomenon that the display playback volume of a resource to be played exceeds the scheduled playback volume. This scheduled playback volume has been pre-set or dynamically acquired. Generally speaking, the resource to be played is associated with an order contract for the release of a resource, and the order contract can stipulate the scheduled playback volume. In this case, this scheduled playback volume can be booked in advance. But it can be understood that the pre-set here does not mean just a fixed value. For example, the scheduled playback volume can be set for each day, the scheduled playback volume can be set for each hour, the scheduled playback volume for users in different regions can be set, and the scheduled playback volume for users of different genders can be set. In other words, this scheduled playback volume can be related to object characteristics, time, or geographic location. But no matter what the specific setting strategy is, the scheduled playback volume will have an upper limit. If it exceeds this upper limit, it will be considered overbroadcast.

[0206] In another scenario, the scheduled playback volume is not directly agreed upon in advance by the order contract, but can be dynamically set by the computer system of the resource playback and delivery entity through the application programming interface. In this case, the resource playback and delivery entity can dynamically obtain the statistical data of resource playback and dynamically adjust the scheduled playback volume based on different user feedback. In this scenario, the content push system is required to open the application programming interface for the resource playback and delivery entity to call.

[0207] According to the specific statistical results of resource playback data, overall, the actual playback volume of resources changes relatively smoothly, but it will show different characteristics at different time points, which shows that the actual playback volume of resources is closely related to time. In addition, the actual playback volume of resources is also related to the object feature information related to the resources to be played. Suppose there are resources D1 and D2 to be played, and their scheduled playback volume is 1000, but the object features associated with them are different. Then, under the premise of not performing overplay control, the actual playback volume curve growth of the resources is also different. In summary, if you want to predict the playback volume of a resource to be played at a scheduled time in the future based on the existing playback data of the resource to be played, you need to consider both the time factor and the object feature information.

[0208] The first data processing module 330 is used to process the playback data to obtain resource playback feature vectors corresponding to a plurality of continuous time slices.

[0209] As mentioned above, to predict the playback volume of a resource to be played at a scheduled time in the future based on the existing playback data of the resource to be played, it is necessary to consider both the time factor and the object feature information. For object feature information, category feature information such as placement, regional classification, crowd label, media label, basic attributes, etc. cannot be directly input into the LSTM model. First, the period needs to be processed into a vector.

[0210] The placement here refers to, for example, the location where the resource is played; the regional classification refers to the user's geographical area, which may generally be detailed to provinces, cities, districts, or even streets; the population label refers to the user's group label such as literary youth; the media label refers to, for example, different applications; and the basic attributes refer to, for example, the user's gender, age, etc.

[0211] In a specific implementation, the user's category feature information can be encoded, for example, using One-Hot encoding to convert discrete feature data into a vector. One-Hot will map the category feature to a vector based on the number of values. Each bit in the vector represents whether the sample has entered this feature. For example, gender has two values, male and female. After encoding, male will be mapped to (0,1) and female will be mapped to (1,0). For example, if the media is divided into application A, application B, application C, and application D, then application A will be mapped to (1,0,0,0). By analogy, the category feature data can be converted into a vector.

[0212] It is worth noting that if there are two category features, two vectors are obtained. If there are more category features, there will be more vectors. These vectors cannot be directly applied to the LSTM model, and embedding processing is required to map multiple vectors into a multidimensional vector. At the same time, the data of the scheduled playback volume and the actual playback volume of the resource to be played can also be mapped to the multidimensional vector.

[0213] For the LSTM model, its input feature is a three-dimensional matrix, which is organized according to the structure of (batch size, time step, feature). Specifically, the batch size refers to the number of time slices corresponding to a batch of playback data. The time slice refers to the minimum time unit for resource playback data statistics. The time step refers to the number of time slices to be stacked for features, which corresponds to a continuous time period. Features refer to the multidimensional vectors obtained according to the above steps.

[0214] In a specific implementation, the batch size is 10, the time step is 5, and the length of the time slice is 1 minute. That is, the goal of this input feature construction method is to use the resource playback data of the past 10 minutes to predict the resource playback volume of the next 1 minute. It can be understood that the batch size, time step, and time slice length are not limited to the above exemplary values ​​and can be adjusted according to actual conditions.

[0215] Assume that there is resource playback data within 10 minutes, that is, the resource playback data of 10 time slices are marked as 1-10 respectively. If the time step is 5, the feature vectors corresponding to the time slices 1, 2, 3, 4, and 5 are stacked, the feature vectors corresponding to the time slices 2, 3, 4, 5, and 6 are stacked, the feature vectors corresponding to the time slices 3, 4, 5, 6, and 7 are stacked, the feature vectors corresponding to the time slices 4, 5, 6, 7, and 8 are stacked, and so on, until all the time slices are used. Assuming the batch size is N and the time step is S, the actual number of stacked features is N-S+1.

[0216] The first vector extraction module 340 is used to input the plurality of resource playback feature vectors into a pre-trained long short-term memory neural network model to obtain a first input vector.

[0217] First, the existing resource playback data is used to train the above LSTM model. That is, training samples are constructed based on the existing resource playback data, and the actual playback volume is used to verify the model, and the training is continued until the above LSTM model converges.

[0218] After the model is trained, the playback data in the most recent period is used to construct a batch of input features, which are then input into the trained model to obtain the first input vector.

[0219] The second data processing module 350 is used to extract statistical sparse features from the playback data.

[0220] It can be understood that the input feature in the neural network model is a kind of expression of data, generally represented by a vector. Assuming its dimension is N*N, it means that the vector is composed of N*N numbers. If the number of 0 values ​​in the N numbers is greater than the number of non-zero values, the feature can be defined as a sparse feature. The higher the proportion of 0, the higher the sparsity.

[0221] Statistically sparse features are not time series features, including historical playback speed, and directional features: age, gender, region, playback speed, etc.

[0222] The second vector extraction module 360 ​​is used to process the statistical sparse feature to obtain a second input vector.

[0223] In a specific implementation, the category feature information of the user can be encoded, for example, using One-Hot encoding to convert discrete feature data into a vector. One-Hot will map the category feature to a vector based on the number of values. Each bit in the vector represents whether the sample has entered this feature. For example, gender has two values, male and female. After encoding, male will be mapped to (0,1) and female will be mapped to (1,0). For example, if the media is divided into application A, application B, application C, and application D, then application A will be mapped to (1,0,0,0). By analogy, the category feature data can be converted into a vector. In this way, the statistical sparse feature can be converted into a second input vector.

[0224] The prediction module 370 is used to input the first input vector and the second input vector into a pre-trained deep neural network model to obtain a predicted playback volume.

[0225] like Figure 4 As shown, in the deep neural network model, the first input vector is obtained based on time-series related features, and the second input vector is obtained based on statistical sparse features. Both are input into the DNN model as input vectors, and finally the predicted playback volume is output.

[0226] The playback control module 380 is used to stop playing the resource to be played if the sum of the playback amount of the resource to be played and the predicted playback amount is greater than or equal to the predetermined playback amount.

[0227] In a specific implementation, the above-mentioned overbroadcast control includes the following steps: determining whether the current playback volume plus the predicted playback volume is greater than or equal to the predetermined playback volume; if so, pausing the playback display of the resource to be played; if not, continuing the playback display of the resource to be played without performing additional operations.

[0228] In a specific implementation, the resources whose playback is suspended can be directly stopped from being returned to the media client, that is, the delivery of the resources that are about to be overbroadcast is suspended.

[0229] In another specific embodiment, the subject that executes the overbroadcast control is not the same as the subject that ultimately returns the resources to the media client 20. In this case, a stop order flag for indicating the suspension of the playback display can be directly given to the resource to be played or its corresponding order contract. The stop order flag can be stored in a database or cache system to facilitate the call of the subject that executes the return of resources to the media client 20.

[0230] According to the content push device provided in this embodiment, by embedding discrete category feature codes into vectors, and further using the LSTM model to predict the actual future playback volume, overplay of advertising resources can be avoided. Fig.14 The figure shows a comparison before and after the method of this embodiment is used, wherein the left side is the method without using this embodiment, and it can be seen that there is an obvious overcast phenomenon, while after the method of this embodiment is used, the overcast phenomenon is basically eliminated.

[0231] In a specific implementation, the above-mentioned content push device further obtains a corresponding playback probability value according to the predicted playback volume, and further controls the playback of the to-be-played resource according to the playback probability value.

[0232] In a specific embodiment, the above-mentioned playback probability value is positively correlated with the difference between the predicted playback volume and the actual playback volume, that is, the larger the difference is, the larger the playback probability value is, and a clear functional relationship can be formed between the playback probability value R and the difference, that is, R = f(d), wherein R represents the playback probability value of the resource to be played, and d represents the difference between the predicted playback volume and the actual playback volume. Of course, the positive correlation between the playback probability value R and the difference is not limited to forming a functional relationship. For example, it can also be through a simple mapping table, according to different intervals of the difference, directly giving different probability value coefficients, and the final playback probability value R needs to be multiplied by the probability value coefficient when calculating.

[0233] It can be understood that since the difference between the predicted playback volume and the actual playback volume is positively correlated, when the difference is large, that is, a large proportion of the scheduled playback plan has not been completed, the resources to be played have a higher probability of playback, which can increase the playback speed of the resources to be played. When the difference is small, that is, the scheduled playback plan has been basically completed, the resources to be played have a smaller probability of playback, and their playback speed will decrease. In this way, the overall playback will gradually approach the scheduled playback volume at a relatively slow speed, thereby further reducing the probability of overbroadcasting.

[0234] In a specific implementation, the above-mentioned content push device also waits for a predetermined time and then performs quantitative supplementary playback according to whether the actual playback volume of the resource to be played is greater than or equal to the predetermined playback volume. After sending the stop order mark, quantitative supplementary playback is also performed according to whether the actual playback volume of the resource to be played is greater than or equal to the predetermined playback volume, which can not only control the resource to be played from being overplayed, but also make its actual playback volume accurately equal to the actual playback volume, thereby improving the working efficiency of the content push system.

[0235] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A content push method, characterized in that, comprising: In response to a received resource acquisition request, acquiring a to-be-played resource that matches, where the to-be-played resource has a corresponding predetermined playback volume; Acquiring playback data of the to-be-played resource within a predetermined time period; Processing the playback data to obtain resource playback feature vectors corresponding one by one to a plurality of consecutive time slices; the resource feature vector corresponding to one of the time slices indicates exposure, inventory, position, and order attributes; the exposure refers to the played volume of the to-be-played resource, the inventory refers to all available playback volumes of the to-be-played resource, and the position refers to the position where the to-be-played resource is played; Inputting the plurality of resource playback feature vectors into a pre-trained long short-term memory neural network model to obtain a first input vector, including: stacking the resource playback feature vectors corresponding to any consecutive S time slices out of N time slices involved in the predetermined time period according to a specified time step S to obtain N-S+1 stacked features, where S is a positive integer less than N and N is a positive integer; combining the N-S+1 stacked features into a three-dimensional matrix; inputting the three-dimensional matrix into the long short-term memory neural network model, and outputting the first input vector by the long short-term memory neural network model; Extracting statistical sparse features from the playback data; the statistical sparse features include historical playback speed, age, gender, and region; Processing the statistical sparse features to obtain a second input vector; Inputting the first input vector and the second input vector into a pre-trained deep neural network model to obtain a predicted playback volume, where the deep neural network model is a multi-layer perceptron model; and If the sum of the played volume of the to-be-played resource and the predicted playback volume is greater than or equal to the predetermined playback volume, stop playing the to-be-played resource.

2. The content push method according to claim 1, characterized in that, the playback data includes categorical features and numerical features, and the processing the playback data to obtain resource playback feature vectors corresponding one by one to a plurality of consecutive time slices includes: Mapping the categorical features to corresponding vectors; and Combining the vectors corresponding to the categorical features with the numerical features into a dense vector to obtain the resource playback feature vector.

3. The content push method according to claim 2, characterized in that, the mapping the categorical features to corresponding vectors includes: Using one-hot encoding to map the categorical features to corresponding vectors.

4. The content push method according to claim 1, characterized in that, the stopping playing the to-be-played resource includes: marking the to-be-played resource as a blocked state.

5. The content push method according to any one of claims 1-4, characterized in that, the method further includes: Recording the playback record of the to-be-played resource and the corresponding feature data; Processing the playback record and the corresponding feature data into a training sample; and Training the deep neural network model using the training sample.

6. The content push method according to claim 5, characterized in that, the training the deep neural network model using the training sample includes: The deep neural network model is trained using the gradient descent method.

7. The content push method according to claim 6, It is characterized in that The use of the gradient descent method to train the deep neural network model includes: The deep neural network model is evaluated using the square of the difference between the estimated playback volume output by the deep neural network model and the actual playback volume corresponding to the training sample.

8. The content push method according to claim 1, It is characterized in that The method further comprises: Obtaining a playback probability value of the resource to be played according to the predicted playback volume; and Determine whether to stop returning the resource to be played based on the playback probability value.

9. The content push method according to claim 8, It is characterized in that The step of obtaining the playback probability value of the resource to be played according to the predicted playback volume includes: Obtaining the difference between the predicted playback volume and the actual playback volume of the resource to be played; and According to the difference, a probability value positively correlated with the difference is obtained as the playback probability value.

10. The content push method according to claim 1, It is characterized in that After stopping playing the resource to be played for a predetermined time, the method further includes: Get the actual playback volume of the resource to be played; If the actual playback amount is less than the predetermined playback amount, the resource to be played is played after receiving a matching user request until the actual playback amount is equal to the predetermined playback amount.

11. A content push device, It is characterized in that include: A resource acquisition module, configured to acquire a matching resource to be played in response to a received resource acquisition request, wherein the resource to be played has a corresponding predetermined playback volume; A playback data acquisition module, used to acquire the playback data of the resource to be played within a predetermined time period; A first data processing module is used to process the playback data to obtain resource playback feature vectors corresponding to a plurality of continuous time slices; the resource feature vector corresponding to a time slice indicates exposure, inventory, position and order attributes; the exposure refers to the amount of playback of the resource to be played, the inventory refers to all available playback amounts of the resource to be played, and the position refers to the location where the resource to be played is played; A first vector extraction module, used for inputting the plurality of resource playback feature vectors into a pre-trained long short-term memory neural network model to obtain a first input vector; A second data processing module, used for extracting statistical sparse features from the playback data; The statistical sparse features include historical playback speed, age, gender and region; A second vector extraction module, used for processing the statistical sparse feature to obtain a second input vector; A prediction module, used for inputting the first input vector and the second input vector into a pre-trained deep neural network model to obtain a predicted playback amount, wherein the deep neural network model is a multi-layer perceptron model; A playback control module, configured to stop playing the resource to be played if the sum of the played amount of the resource to be played and the predicted played amount is greater than or equal to the predetermined played amount; The method of inputting the multiple resource playback feature vectors into a pre-trained long short-term memory neural network model to obtain a first input vector includes: stacking the resource playback feature vectors corresponding to any S consecutive time slices of N time slices involved in a predetermined time period according to a specified time step S to obtain N-S+1 stacked features, where S is a positive integer less than N and N is a positive integer; combining the N-S+1 stacked features into a three-dimensional matrix; and inputting the three-dimensional matrix into the long short-term memory neural network model, so that the long short-term memory neural network model outputs the first input vector.

12. A server, It is characterized in that include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-10.

13. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Resource processing method and device and computer storage medium

    CN110544134A

  • Stock data fluctuation prediction method based on deep neural network and multi-task learning

    CN111476418A

  • Video processing method and device, computer equipment and storage medium

    CN111565316A