Advertisement push method, device, electronic device and medium based on large model
By analyzing user behavior and emotions and utilizing large language models to optimize ad push timing, we address the issues of insufficient user appeal and user experience disruption associated with traditional ad push methods, achieving higher click-through rates, conversion rates, and an immersive viewing experience.
Patent Information
- Application Number
- CN202411794886.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-12-06
AI Technical Summary
Traditional advertising push methods are difficult to effectively attract users, affect the user's viewing experience, and are difficult to improve click-through rate and conversion rate. They cannot dynamically adjust the advertising insertion time according to user behavior and emotions.
By analyzing the user's viewing time information, advertising preference information, emotional information and scene information during video playback, using a large language model to determine the ads to be pushed and their insertion time, and combining the user's interest distribution and emotional state to optimize ad push.
It improves the click-through rate and conversion rate of advertisements, reduces the interference of advertisements on the user's viewing experience, and ensures an immersive viewing experience for users.
Smart Images

Figure CN119648305B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the field of large language models (LLM) and personalized recommendation technology. Background Art
[0002] Pushing ads to users while they're watching a video is a common advertising push scenario. For example, after a user clicks play on a video, an ad plays before the video itself. Another example is playing ads during a fixed timeframe during video playback. This type of ad push increases ad volume and boosts advertisers' revenue. Summary of the Invention
[0003] The present disclosure provides a large-scale model-based advertising push method, device, electronic device, and medium.
[0004] A first aspect of the embodiments of the present disclosure provides an advertisement push method based on a large model, comprising:
[0005] During playback of a current video, determining, based on a first operation performed by a user on the current video, viewing time information of the user on the current video;
[0006] Acquiring advertisement preference information of the user for the pushed advertisement, where the advertisement preference information is determined based on a second operation performed by the user on the pushed advertisement;
[0007] Obtaining emotional information of the user during the playback of the pushed advertisement;
[0008] Obtaining scene information of multiple video clips included in the current video;
[0009] Based on the viewing time information, the advertisement preference information, the emotion information and the scene information, a large language model is used to determine the advertisement to be pushed and the time when the advertisement to be pushed is to be inserted during the playback of the current video.
[0010] A second aspect of the embodiments of the present disclosure provides an advertisement push device based on a large model, comprising:
[0011] a determination module, configured to determine, during playback of a current video, information about a user's viewing time of the current video based on a first operation performed by the user on the current video;
[0012] an acquisition module, configured to acquire advertisement preference information of the user for the pushed advertisement, wherein the advertisement preference information is determined based on a second operation performed by the user on the pushed advertisement;
[0013] The acquisition module is further configured to acquire the user's emotional information during the playback of the pushed advertisement;
[0014] The acquisition module is further configured to acquire scene information of a plurality of video clips included in the current video;
[0015] The determination module is further used to determine the advertisement to be pushed and the time to insert the advertisement to be pushed during the playback of the current video based on the viewing time information, the advertising preference information, the emotional information and the scene information using a large language model.
[0016] According to a third aspect of the present disclosure, an electronic device is provided, including:
[0017] at least one processor; and
[0018] a memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any one of the methods according to the first aspect.
[0020] According to a fourth aspect of an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to any one of the first aspects.
[0021] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of the first aspects is implemented.
[0022] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0024] Figure 1 This is a flow chart of a method for pushing advertisements based on a large model provided by an embodiment of the present disclosure;
[0025] Figure 2 is a flow chart of a method for determining advertising preference information provided by an embodiment of the present disclosure;
[0026] Figure 3is a flow chart of a method for determining emotional information provided by an embodiment of the present disclosure;
[0027] Figure 4 is an exemplary schematic diagram of a process for determining emotional information provided by an embodiment of the present disclosure;
[0028] Figure 5 is a flow chart of a method for determining scene information provided by an embodiment of the present disclosure;
[0029] Figure 6 is a flow chart of another large model-based advertising push method provided by an embodiment of the present disclosure;
[0030] Figure 7 This is an exemplary schematic diagram of an advertisement push process based on a large model provided by an embodiment of the present disclosure;
[0031] Figure 8 This is a structural diagram of an advertisement push device based on a large model provided by an embodiment of the present disclosure;
[0032] Figure 9 It is a block diagram of an electronic device used to implement the large model-based advertisement push method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0034] Traditional advertising strategies mostly rely on a fixed timeline insertion method, where ads are played during fixed periods of video playback. However, with the development of online video platforms, traditional advertising methods have become less effective in attracting users. Single ad placements can easily interrupt users' viewing experience, impacting their smooth viewing experience and making it difficult to significantly increase click-through and conversion rates.
[0035] In order to solve the above problems, the embodiments of the present disclosure provide an advertisement push method based on a large model, which is applied to electronic devices, such as servers, desktop computers, or laptop computers with data processing capabilities. Figure 1 As shown, the advertisement push method based on the large model provided by the embodiment of the present disclosure includes the following steps:
[0036] S101. During playback of a current video, determine a user's viewing time information for the current video based on a first operation performed by the user on the current video.
[0037] For ease of description, the embodiment of the present disclosure refers to the operation behavior performed by the user on the current video as the first operation behavior. The first operation behavior can be various operation behaviors supported by the video. The operation behavior has attributes such as operation type, operation time, and operation user. For example, operation types include: play, close, skip, fast forward, pause, and rewind. The operation time represents the playback time of the video and the system time of the playback terminal when the user performs the operation behavior. The operation user is the user identifier (ID) of the user who performs the operation behavior. For example, the operation user can be the user ID of the logged-in user of the video playback terminal.
[0038] Viewing time information may include: cumulative viewing time, continuous viewing time periods, skip time periods, fast-forward time periods, and rewind time periods. Each continuous viewing time period is the period between a play action and a stop action. Stop actions include closing, skipping, fast-forwarding, pausing, and rewinding. The cumulative viewing time is the sum of each continuous viewing time period.
[0039] Optionally, the playback terminal of the current video may send the first operation behavior to the electronic device whenever it receives a first operation behavior during playback of the current video, or periodically send the first operation behavior received in the previous period to the electronic device. Alternatively, the playback terminal may send the first operation behavior to the electronic device in other ways, which are not specifically limited in the embodiments of the present disclosure.
[0040] By using the user's viewing time information, we can analyze the distribution of the user's interest in the video content. For example, the user is more interested in the video clips that are watched continuously for a long time, and is less interested in the video clips that are skipped.
[0041] S102: Obtaining the user's advertising preference information for the pushed advertisement, wherein the advertising preference information is determined based on a second operation performed by the user on the pushed advertisement.
[0042] Pushed ads include: ads that have been inserted in the current video, and ads that have been inserted in the user's previous videos.
[0043] For the convenience of description, the embodiment of the present disclosure refers to the operation behavior performed by the user on the pushed advertisement as the second operation behavior. The second operation behavior can be various operation behaviors supported by the advertisement.
[0044] The secondary actions a user performs on a pushed ad can reflect their preference for the pushed ad. For example, if a user clicks pause or close during an ad, it can indicate low interest in the ad. Clicking a product link in an ad, which redirects to the product details page, or clicking a call to action (CTA) button in an ad, which engages with the ad, can both indicate high interest. Therefore, based on the secondary actions, we can determine the user's preference for pushed ads.
[0045] S103: Acquire the user's emotional information during the playback of the pushed advertisement.
[0046] Emotional information can include emotion type, emotion intensity, and the time period corresponding to the emotion type. A user's emotional information can reflect their level of focus while watching a video. For example, positive emotions include excitement, joy, and happiness, while negative emotions include anxiety, sadness, boredom, and fear. Other emotion types include neutral, doubt, and surprise. Neutral emotion indicates a lack of significant emotional fluctuations, such as relaxation and concentration.
[0047] S104: Obtain scene information of multiple video clips included in the current video.
[0048] The scene information may include: the video clip includes the beginning, development, climax or ending plot.
[0049] It should be noted that S101 to S104 can be executed in sequence or in parallel, and the execution order of S101 to S104 is not limited to Figure 1 in the order shown.
[0050] S105 : Based on the viewing time information, the advertisement preference information, the emotion information, and the scene information, the large language model is used to determine the advertisement to be pushed and the time when the advertisement to be pushed is to be inserted during the playback of the current video.
[0051] In the disclosed embodiment, viewing time information can reflect the distribution of user interests at different playback times when watching a video, scene information can represent the content of different video clips, and advertising preference information and emotional information can reflect the user's interest in different advertisements. Therefore, by combining viewing time information, advertising preference information, emotional information, and scene information, the user's preferences for advertisements and video content can be obtained, and the advertisements to be pushed can be determined, which can increase the possibility that the advertisements to be pushed will be liked by the user. Moreover, by combining viewing time information, advertising preference information, emotional information, and scene information to determine the time to be inserted, the possibility of the advertisement being inserted at the playback time of a video clip with a high degree of user interest can be reduced, and the possibility of the advertisement being inserted at the playback time of a video clip with a low degree of user interest can be increased, thereby reducing the interruption of advertisement playback to the user's video viewing and ensuring a more immersive video viewing experience for the user.
[0052] The following is a detailed description of the advertising push method based on the large model provided by the embodiment of the present disclosure:
[0053] In some embodiments of the present disclosure, the user's advertising preference information for pushed advertisements in S102 includes the user's preference for various advertisement types. Advertisement types can be divided according to advertisement content, for example, advertisement types include: electronic products, cosmetics, or sports equipment.
[0054] See also Figure 2 The electronic device determines the user's advertising preference information for the pushed advertisement based on the second operation performed by the user on the pushed advertisement in S102, including:
[0055] S201: Determine a first preference score of a user for a first pushed advertisement based on a type of a first operation performed by a user on the first pushed advertisement, wherein the first pushed advertisement is a pushed advertisement inserted during playback of a current video.
[0056] A data collection script can be pre-deployed in the terminal playing the current video. Through the data collection script, the terminal can transmit, during playback of the first pushed advertisement, a first user action performed on the first pushed advertisement to an electronic device via a communication protocol (web socket). The first action includes parameters such as an advertisement ID, an action ID, and an action time. Furthermore, the first action may include other parameters, such as advertisement type, an action score, and information about the scene of the video clip before the advertisement was inserted.
[0057] Each type of action is pre-assigned a corresponding action score. For example, clicking an ad link has an action score of 5, skipping an ad link has an action score of 0, pausing or closing an ad link has an action score of -5, and rewatching an ad has an action score of 3. For each first pushed ad, the sum or weighted sum of the action scores corresponding to the types of first actions performed on the first pushed ad can be used as the first preference score for the first pushed ad.
[0058] In addition, for each first pushed advertisement, if no negative operation behavior is received from the user during the playback of the first pushed advertisement, such as negative operation behaviors including pause, close and fast forward, it means that the user has watched the first pushed advertisement in its entirety, and the first preference score of the first pushed advertisement can be increased to a default value, for example, a default value of 2.
[0059] S202: Determine the advertisement type of the first pushed advertisement.
[0060] The electronic device may pre-store the advertisement type of each advertisement, and thus may obtain the advertisement type of the first pushed advertisement.
[0061] S203: Obtain the user's preference for the advertisement type.
[0062] The preference is determined based on a second preference score of the user for a second pushed advertisement, where the second pushed advertisement is a pushed advertisement of the same advertisement type that was inserted into a video that the user has previously watched.
[0063] S204: Update the user's preference for the advertisement type according to the first preference score to obtain updated user preferences for various advertisement types.
[0064] For this ad type, the weighted sum of the first preference score and the second preference score can be used as the user's preference for this ad type. The weight of each preference score is a time interval decay weight, that is, the further the ad push time is from the current moment, the lower the weight of the preference score. For example, for an ad that has been pushed more than one month from the current moment, the weight corresponding to its preference score is 0.2; for an ad that has been pushed more than one week but less than one month from the current moment, the weight corresponding to its preference score is 0.5; for an ad that has been pushed less than one week from the current moment, the weight corresponding to its preference score is 1. Through this calculation method, the user's preference for an ad type can better reflect the user's recent ad preferences.
[0065] Alternatively, for the advertisement type, the average of the first preference score and the second preference score may be used as the user's preference for the advertisement type. Alternatively, the user's preference for the advertisement type may be updated in other ways, which are not specifically limited in the present embodiment.
[0066] Through the above method, the embodiment of the present disclosure can update the user's preference for various types of advertisements based on the user's operation behavior on the advertisements inserted in the current video, so that the user's preference for various types of advertisements can be updated in real time, ensuring the accuracy of the user's advertising preferences reflected therein.
[0067] In some embodiments of the present disclosure, the emotional information in S103 includes the user's average emotional score while viewing the pushed advertisement. The electronic device can determine the user's emotional information in real time during playback of the pushed advertisement, before recommending the advertisement. The emotional information can be determined in two ways:
[0068] See also Figure 3 ,The first way to determine emotional information includes the following steps:
[0069] S301. During the playback of a pushed advertisement, obtain a user video shot of the user.
[0070] It should be noted that the electronic device obtains the emotional information of each pushed advertisement in the same way. Figure 3 The relevant steps are described using a pushed advertisement as an example.
[0071] The playback terminal to which the advertisement has been pushed has a built-in camera or can be connected to an external camera, so that the playback terminal can shoot the user through the camera during the playback of the pushed advertisement to obtain the user video. If the user has previously granted the user video transmission permission, the playback terminal can send the user video to the electronic device.
[0072] For example, the playback terminal can play videos and advertisements in the form of an application, web page, or mini-program. The playback terminal can display a permission acquisition agreement when the application, web page, or mini-program is opened for the first time. The permission acquisition agreement is used to apply for multiple permissions, including user video transmission permissions. When the user agrees to the permission acquisition agreement, it indicates that the user has granted various permissions including user video transmission permissions. Alternatively, the playback terminal can display a pop-up window for applying for user video transmission permissions when the application, web page, or mini-program is opened for the first time. When the user clicks "Agree" in the pop-up window, it indicates that the user has granted user video transmission permissions. Alternatively, the user video transmission permissions can also be obtained through other means, which is not specifically limited in the embodiments of the present disclosure.
[0073] Optionally, the playback terminal can send the captured user video to the electronic device every time a user video of a preset length is captured; or, after the playback of the pushed advertisement is completed, the playback terminal can send the user video captured during the time period from the start of the playback of the pushed advertisement to the completion of the playback to the electronic device.
[0074] S302 : for video frames included in the user video, identify the user's facial expressions and body language in the video frames using an emotion recognition model, and output the probability of the user in the video frames being in various emotion types based on the facial expressions and body language.
[0075] Among them, the emotion recognition model can be a visual algorithm model based on deep learning, such as Convolutional Neural Networks (CNN).
[0076] Each video frame can be input into an emotion recognition model. The model then identifies the facial and body regions within the video frame, extracts features from the facial regions, and uses these features to identify facial expressions. Furthermore, the emotion recognition model can also extract features from the body regions, identifying body language based on these features. Based on the facial expression and body language, the model then determines the probability of the user experiencing various emotions within the video frame. For example, a smile in the facial region may reflect excitement, while a frown in the facial region may reflect confusion. For another example, a sitting posture in the body region may reflect happiness, while a sitting posture with legs crossed may reflect confusion.
[0077] S303: Determine an emotion score of the video frame based on the probability of the user being in various emotion types in the video frame.
[0078] For each video frame included in the user video, the weighted sum of the probabilities of the user being in various emotion types in the video frame can be used as the emotion score of the video frame. The weights of the probability of the emotion type can be preset.
[0079] For example, suppose that within a video frame, the probability of the user being excited is 0.6, the probability of being happy is 0.25, the probability of being confused is 0.1, and the probability of being bored is 0.05; and the weight of the probability of excitement is 1, the weight of the probability of happiness is 0.5, the weight of the probability of confusion is -0.3, and the weight of boredom is -1. The resulting emotion score for this video frame is: 0.6 × 1 + 0.25 × 0.5 + 0.1 × (-0.3) + 0.05 × (-1) = 0.717.
[0080] The emotion score for each video frame ranges from 0 to 1. A score closer to 0 indicates a lower user concentration and a lower interest in the ad. A score closer to 1 indicates a higher user concentration and a higher interest in the ad.
[0081] S304: Smoothing the emotion scores of consecutive video frames to obtain an average emotion score.
[0082] The emotion scores of every N consecutive video frames included in the user video may be smoothed to obtain an average emotion score, where N is a preset positive integer, for example, N=5.
[0083] The average emotion score can be calculated using a time-weighted sliding average. Specifically, for N consecutive video frames, the sum of the emotion scores of each frame multiplied by a time decay factor is calculated, and the ratio of this sum to the sum of the time decay factors is used as the average emotion score. The time decay factor ranges from [0 to 1]; the earlier the video frame is played, the smaller the time decay factor corresponding to its emotion score.
[0084] The disclosed embodiments combine a user's facial expressions and body language to more comprehensively identify the user's emotions while watching a video, improving the accuracy of identifying emotion types. Furthermore, since the emotion recognition model's recognition results may contain errors, smoothing the emotion scores of consecutive video frames improves the accuracy of the average emotion score in reflecting the user's interest in the delivered ads. This can reduce the impact of the emotion recognition model's recognition errors on ad recommendations and improve the accuracy of ad recommendations based on emotion information.
[0085] The second method of determining emotional information includes the following steps:
[0086] Step 1: During the playback of the pushed advertisement, multiple groups of probabilities are received from the playback terminal of the pushed advertisement.
[0087] The single group probability represents the probability of the user being in various emotion types at a time. The playback terminal can determine the probability of the user being in various emotion types in each video frame included in the user video according to the above S302 method, and regard the probability corresponding to each video frame as a group.
[0088] Optionally, the playback terminal that has received the pushed advertisement can determine the probability of the user being in various emotion types within each video frame in real time during playback of the pushed advertisement, and send the set of probabilities determined in real time to the electronic device. Alternatively, after obtaining a certain number of sets of probabilities, the playback terminal can send the obtained multiple sets of probabilities to the electronic device. Alternatively, the playback terminal can also send multiple sets of probabilities to the electronic device through other methods, which are not specifically limited in the present embodiment.
[0089] Step 2: Based on the same set of probabilities, determine the user's emotion score at a moment.
[0090] The method of determining the emotion score in step 2 is the same as that in the above S303 , and reference may be made to the description of the above S303 , which will not be repeated here.
[0091] Step 3: Smooth the continuous emotion scores to obtain the average emotion score.
[0092] The method of determining the average emotion score in step 3 is the same as that in the above S304 , and reference may be made to the description of the above S304 , which will not be repeated here.
[0093] In this disclosed embodiment, the playback terminal can capture and analyze user videos to determine the probability of the user being in various emotions. Only the probabilities of the user being in each emotion type are transmitted to the electronic device, thus avoiding potential privacy leaks in the transmission of user videos and ensuring security. Furthermore, since emotion scores within a given moment may have errors, the emotion scores of consecutive video frames are smoothed, improving the accuracy of the average emotion score in reflecting the user's interest in the delivered ads, thereby improving the accuracy of ad recommendations based on emotion information.
[0094] The implementation process of the second method of determining emotional information can be found in Figure 4 The playback terminal, which has received a pushed ad, uses a camera to capture user video during playback of the pushed ad. Each video frame captured by the camera is then input into an emotion recognition model within the playback terminal. The emotion recognition model outputs the probability of the user experiencing various emotion types within each video frame, generating multiple sets of probabilities. The playback terminal then sends each set of probabilities output by the recognition model to a server. The server sends each received set of probabilities to its own configured advertising system (Advertise System, AdvertiseSys). The advertising system, based on the received sets of probabilities, determines the user's emotional information during playback of the pushed ad. The advertising system is a software module that implements the large-model-based advertising push method provided in the disclosed embodiments.
[0095] It should be noted that there may be cases where a user video includes multiple users. In this case, the facial areas and limb areas of different users in the user video can be segmented, and the emotional information of the user can be determined based on the facial area and limb area of the same user.
[0096] Afterwards, the average value of the emotion information of each user may be used as the emotion information of each user during the playback of the pushed advertisement.
[0097] Alternatively, the primary viewer can be identified from each user, and the emotional information of the primary viewer can be used as the emotional information of each user during the playback of the pushed advertisement. The primary viewer can be: the owner of the login account of the playback terminal, the user who mainly performs the operation, or the user with the highest concentration, etc.
[0098] Alternatively, advertisements may be recommended for each user, and the advertisements pushed to the user may be displayed at a designated position in the current video, wherein the designated position may be determined based on the relative position between the user and the display screen of the playback terminal.
[0099] Alternatively, when it is detected that the user video includes multiple users, the playback terminal may display inquiry information and receive interest information selected by the users as emotion information of each user.
[0100] In some embodiments of the present disclosure, the scene information in the above S104 includes: the emotion type and scene type of each video clip in the current video, the emotion transition time period and scene transition time period in the current video, and the playback time period of the specified object in the current video.
[0101] The scene information of the multiple video clips included in the current video can be predetermined, or determined in real time during the advertisement push process. Figure 5 , the scene information is determined by the following steps:
[0102] S501: Divide the current video into multiple video segments.
[0103] Optionally, the current video may be divided according to a preset fixed duration to obtain multiple video segments of fixed duration, for example, the fixed duration is 10 seconds or 30 seconds.
[0104] Alternatively, an image processing model can be used to identify the scene to which each video frame belongs, and consecutive video frames belonging to the same scene can be grouped into a video segment. For example, a shot cut model can be used to identify whether each video frame is a shot cut frame, and the video frames between adjacent shot cut frames can be grouped into a video segment.
[0105] Alternatively, an image recognition model may be used to identify whether each video frame is an emotion inflection point video frame, and the video frames between adjacent emotion inflection point video frames are taken as a video segment.
[0106] Alternatively, the current video may be divided in other ways, which are not specifically limited in the embodiment of the present disclosure.
[0107] S502. Using the emotion and content analysis model, perform image feature extraction on the video frames included in the video clip, perform text feature extraction on the text in the video frames, and perform audio feature extraction on the audio included in the video clip, and determine the emotion type of the video clip based on the feature extraction results.
[0108] The emotion and content parsing model is a multimodal model with image, text, and audio processing capabilities, enabling it to recognize input images, text, and audio. The emotion and content parsing model is an artificial intelligence (AI) model based on natural language processing (NLP) and CNN, such as a large language model. The large language model used here can be the same model as that used in S105 above, or different models.
[0109] The video frames included in the video clip can reflect the visual tone of the video clip, various elements in the scene, changes in the expressions and body language of people in the scene, and changes in the relationship between people and objects. Therefore, the video frames can reflect the emotions expressed in the video clip.
[0110] The text within a video frame includes subtitles, such as character dialogue and narration. Text can express the emotions of the characters and the tone of the plot, so the text within a video frame can reflect the emotions expressed in the video clip.
[0111] Audio includes the voice of the characters and background sound. Voice can reflect the tone of the user's speech, while background sound can include ambient sound and background music, which can reflect the emotional tone of the scene. Therefore, the audio included in the video clip can reflect the emotion expressed in the video clip.
[0112] For example, emotion types include: relaxed and cheerful, mildly tense, severely tense, conflicting, and peaceful, etc.
[0113] S503: Determine an emotion transition time period based on the emotion type of each video clip.
[0114] It can be determined whether the emotion types of each of the two adjacent video clips are the same. If they are different, the video frames within a preset time period before and after the dividing point between the adjacent video clips are used as the emotion transition period, or the two video clips are combined into the emotion transition period. Alternatively, the emotion transition period can be determined by other methods, which are not specifically limited in the present embodiment.
[0115] For example, assuming that video segments a and video segments b are adjacent video segments, the emotion type of video segment a is relaxed and cheerful, and the emotion type of video segment b is slightly nervous, then the last 10 seconds of video segment a and the first 10 seconds of video segment b constitute an emotion transition time period.
[0116] S504: Using the object detection and scene recognition model, perform image recognition on the video frames included in the video clip, and output the object detection result and the scene type of the video clip.
[0117] The object detection result indicates whether the video clip contains a preset object and the playback time period of the preset object. For example, the preset objects include: landscape, food, cars, electronic products, pets or cosmetics.
[0118] For example, the scene types of a video clip include climax, transition, and smooth scene.
[0119] The object detection and scene recognition model can be a deep learning-based visual algorithm model. For example, the object detection and scene recognition model can be a CNN. Alternatively, the object detection and scene recognition model can include an object detection sub-model and a scene recognition sub-model. The object detection sub-model identifies the object detection result, and the scene recognition sub-model identifies the scene type. The object detection sub-model can be a You Only Look Once (YOLO) object detection model or a Faster R-CNN model, and the scene recognition sub-model can be a Residual Network (ResNet).
[0120] S505: Determine a scene transition time period based on the scene type to which each video clip belongs.
[0121] The method for determining the scene transition time period is the same as the method for determining the emotion transition time period described above. Please refer to the relevant description of S503 above, and will not be repeated here.
[0122] During emotional transitions, the emotional atmosphere of a video changes, for example, from relaxed to tense, or from conflicting to relaxed. This often results in changes in the video's plot, meaning the video's content is generally discontinuous. Similarly, during scene transitions, the filming environment of a video changes, so the video's content is generally discontinuous. Therefore, inserting ads during emotional transitions or scene transitions generally doesn't disrupt users. Therefore, these emotional transitions and scene transitions provide valuable insights into determining which ads to push and when to insert them.
[0123] Furthermore, after obtaining the emotion type and scene type of the video clip, when recommending ads, we can avoid recommending ads that do not match the emotion type or scene type of the video clip. For example, we can insert humorous ads into sad video clips, or insert relaxing ads into tense video clips, or insert interactive ads that may cause pauses into slow-paced video clips, etc. This improves the matching degree between the recommended ads and the emotion type or scene type of the video clip. For example, we can insert travel ads into video clips that include natural scenery such as mountains and rivers, and insert game ads or ads for game hardware supporting game into video clips that include e-sports. Therefore, the emotion type and scene type of the video clip have certain reference value for the subsequent determination of the ads to be pushed and the time to insert them.
[0124] Furthermore, after obtaining the playback time period of a preset object included in the current video, ads related to the preset object can be retrieved as ads to be pushed during ad recommendations. For example, an ad for pet supplies can be played during the playback time period of pets. Therefore, the playback time period of the preset object provides a reference for determining the ads to be pushed and the insertion time.
[0125] In some embodiments of the present disclosure, see Figure 6 The electronic device in step S105 determines, based on the advertisement preference information, the emotion information, and the scene information, a method for determining the advertisement to be pushed and the time at which the advertisement to be pushed should be inserted during the playback of the current video using a large language model, including the following steps:
[0126] S601: Determine a plurality of recommended insertion times based on scene information, wherein the recommended insertion times include the play times of video clips belonging to non-climax scenes.
[0127] For example, non-climax scenes include transition scenes and flat scenes.
[0128] The highest priority can be set for emotional transitions and scene transitions, the lowest priority for the playback time period of the climax scene, and the intermediate priority for the playback time period of the remaining video clips containing the designated object. When executing S601, the video clips can be sorted in descending order of priority based on the scene information, and multiple time periods can be selected as recommended insertion times. The number of time periods can be preset based on actual needs.
[0129] S602: Based on the user's current emotion type, filter out an advertisement insertion triggering time from the recommended insertion times.
[0130] The user's emotion type within a preset time period from the current moment can be obtained as the current emotion type. If the current emotion type indicates too high or too low concentration, for example, excitement or boredom, and the time difference between the start time of the recommended insertion time and the current moment is less than the preset time difference, the recommended insertion time is filtered out, and the remaining recommended insertion times are retained.
[0131] S603: When the ad insertion trigger time arrives, the viewing time information, ad preference information, emotion information, and scene information are input into the large language model. The large language model is used to predict the user's click-through rate for each ad. If the click-through rate meets the preset conditions, candidate ads, candidate ad versions, and candidate insertion times are output based on the click-through rate. If the click-through rate does not meet the preset conditions, candidate ads, candidate ad versions, and candidate insertion times are output based on the alternative ads. The ad version indicates the display format of the ad content.
[0132] In the disclosed embodiments, viewing time information, ad preference information, emotion information, and scene information can be updated in real time. When the ad insertion trigger time arrives, the latest viewing time information, ad preference information, emotion information, and scene information can be input into the large language model. The large language model is used to predict the user's click-through rate for each ad and the degree of match between each ad and the current video. Based on the click-through rate and match degree, candidate ads, candidate ad versions, and candidate insertion times are output.
[0133] In addition to viewing time information, advertising preference information, emotional information and scene information, other information can also be input into the large language model, such as the first operation behavior performed by the user on the current video, the second operation behavior performed by the user on the pushed advertisement, and the advertisement type and content information of each advertisement, such as whether the content information contains preset objects and the degree of matching between the advertisement and the current video.
[0134] Ad versions refer to the display format of the ad content. Different versions of the same ad may share the same promotional objectives but differ in presentation format or style. For example, ad versions include short and long versions, emotional and data-driven, interactive and passive viewing.
[0135] The large language model can be a multimodal pre-trained model based on a transformer architecture. Through its internal attention mechanism, the large language model can efficiently estimate the user's appeal to each advertisement in the current video playback scenario, namely, the user's click-through rate (CTR). It then determines whether the predicted CTR meets preset conditions, such as the number of CTRs exceeding a preset value. If the CTR meets the preset conditions, the large language model can output candidate ads, ad versions, and insertion times based on the CTR. If the CTR does not meet the preset conditions, it outputs candidate ads, ad versions, and insertion times based on the alternative ads.
[0136] S604: Determine the advertisement to be pushed based on the candidate advertisements and the candidate advertisement versions, and determine the insertion time based on the candidate insertion time.
[0137] The disclosed embodiment can filter out the time period containing the climax plot of the current video based on the scene information. Since the climax plot is more exciting than other plots, the user needs to be undisturbed. Therefore, filtering out the time period containing the climax plot can reduce the disturbance to the user caused by playing advertisements in the video. The disclosed embodiment also filters the advertisement insertion trigger time based on the current emotion type, thereby avoiding inserting advertisements when the user's concentration is too high, causing disturbance to the user, or reducing the number of advertisements recommended to the user when the user's concentration is too low, thereby avoiding exacerbating the user's negative emotions. Moreover, the disclosed embodiment also makes advertisement recommendations based on viewing time information, advertisement preference information, emotion information and scene information, thereby increasing the possibility that the determined candidate advertisements and candidate advertisement versions are liked by the user and match the current video.
[0138] In some embodiments of the present disclosure, the method in which the above-mentioned S602 electronic device filters out the advertising insertion trigger time from each recommended insertion time based on the user's current emotion type can be implemented as follows: based on the user's current emotion type and the scene information of the video clip included in the recommended insertion time, determine the user's predicted emotion type at the recommended insertion time, and then filter out the advertising insertion trigger time from each recommended insertion time based on the predicted emotion type.
[0139] A deep learning emotion curve model can be used to predict the user's emotion type at the recommended insertion time based on the current emotion type and the scene information of the video clip included in the recommended insertion time. Based on the emotion type, a score is output indicating the suitability of the recommended insertion time for advertising. Multiple recommended insertion times can then be selected, ranked from highest to lowest scores, as the trigger times for ad insertion. The number of recommended insertion times can be pre-set based on actual needs.
[0140] Through the above method, the embodiment of the present disclosure can predict the user's emotions during the recommended insertion time, reduce the insertion of advertisements when the user's concentration on the video is too high or too low, reduce the impact of advertising push on the user's video viewing, ensure the user's immersive viewing experience, and reduce the aggravation of the user's negative emotions caused by advertising push.
[0141] In the embodiment of the present disclosure, the above-mentioned S604 determines the method of the advertisement to be pushed based on the candidate advertisements and the candidate advertisement versions, which can be implemented as follows: selecting a target candidate advertisement from each candidate advertisement, and then filtering out the version to be pushed from the candidate advertisement versions based on the total length of the current video, the user's current emotion type and the user's historical operation behaviors on each advertisement version, and using the target candidate advertisement of the pushed version as the advertisement to be pushed.
[0142] In one implementation, the candidate advertisements output by the large language model also have candidate scores, so the candidate advertisement with the highest candidate score can be selected as the target candidate advertisement.
[0143] Since the candidate score can reflect the possibility that the candidate ad is liked by the user and the degree of match between the candidate ad and the current video, selecting the candidate ad with the highest candidate score can increase the possibility that the ad recommended to the user will be liked by the user, thereby increasing the ad click-through rate, increasing advertiser revenue, and reducing the interruption caused by the recommended ad to the user's video viewing.
[0144] In another implementation, videos that are associated with the current video and videos of interest to the user may be obtained as related videos, and then target candidate advertisements may be selected based on the correlation between the related videos and the candidate advertisements.
[0145] The videos associated with the current video include: other videos in the same series as the current video, and videos with the same theme or similar plot as the current video. For example, if the current video is an episode of a TV series, then the videos associated with the current video include: other episodes of the TV series, as well as sequels and prequels of the TV series. For another example, if the current video is the first part of a movie, then the videos associated with the current video include other parts of the movie. For another example, if the current video is a video about pet training, then the videos associated with the current video include videos introducing pet training props.
[0146] The user's videos of interest include: videos that the user has watched more than a preset number of times within a unit historical time period, videos that the user has liked, and videos that the user has collected.
[0147] The relevance between each candidate advertisement and each related video may be calculated, and the candidate advertisement with the highest relevance may be selected as the target candidate advertisement.
[0148] For example, if the relevant video is a science fiction video, you can select science fiction-related advertisements as target candidate advertisements.
[0149] Because the current video shares content connections with other videos in the same series, recommending ads that are more relevant to related videos improves the coherence of ads within the series and increases click-through rates. Furthermore, ads that are more relevant to videos a user is interested in are more likely to be liked by users, leading to a higher likelihood of being clicked on after being pushed, increasing advertiser revenue.
[0150] It is understandable that when the total duration of the current video is short, it is not suitable to insert a long advertisement; on the contrary, when the total duration of the current video is long, a long advertisement can be inserted.
[0151] Therefore, after determining the target candidate advertisements, the electronic device can determine the maximum advertisement duration based on the total duration of the current video. For example, from multiple preset duration intervals, the target duration interval to which the total duration of the current video belongs is determined, and the advertisement duration corresponding to the target duration interval is then used as the maximum advertisement duration. Then, from each version of the target candidate advertisement, the electronic device filters out advertisement versions whose duration exceeds the maximum advertisement duration, and retains the remaining advertisement versions.
[0152] Furthermore, from each candidate ad version, those that do not match the user's current emotion are filtered out, while the remaining ad versions are retained. For example, if the current emotion type is positive, ad versions with negative emotions are filtered out; if the current emotion type is negative, ad versions with positive emotions are filtered out.
[0153] Furthermore, based on the historical operation behaviors performed by the user on each advertisement version, the user's preference for each advertisement version is determined, and then advertisement versions with a preference less than a preset threshold are filtered out, and the remaining advertisement versions are retained.
[0154] Through the above method, the embodiment of the present disclosure can filter out target candidate advertisements from the candidate advertisements output by the large language model, and select the advertisement version that best matches the current video, best matches the user's current emotion, and is most likely to be liked by the user based on the total length of the current video, the user's current emotion type, and the user's historical operation behaviors on each advertisement version.
[0155] In some embodiments of the present disclosure, after determining the push version, the electronic device may further determine the grade of the advertisement, wherein the grade indicates the degree of guidance on click behavior. For example, the grades of advertisements include: elementary, intermediate, and advanced.
[0156] Primary ads are less likely to lead to clicks, aiming to capture initial user attention, foster initial interest in the promotional item, and pave the way for subsequent levels of advertising. Primary ads are shorter, typically only 5-10 seconds, and offer a relatively brief introduction to the promotional objective. Examples include teaser ads, brand introduction ads, and ads with simple hooks.
[0157] Intermediate ads are designed to encourage clicks, typically featuring multiple selling points and brand details. These ads aim to deepen users' impressions of the product and encourage them to engage with it. Intermediate ads are of medium length, typically include a detailed description, and may feature a small number of interactive buttons or product links.
[0158] Advanced ads are highly targeted at clickbait. Their core goal is to drive clicks through a strong call to action, ultimately leading to purchases, registrations, or other events. Advanced ads typically feature longer durations, more comprehensive descriptions of the target audience, and more complete logic. They often include multiple product links and / or interactive buttons.
[0159] On this basis, the electronic device uses the push version of the target candidate advertisement as the advertisement to be pushed, including the following steps:
[0160] Step 1: When the user's current emotion type is positive, the user's historical operation behaviors on the target candidate advertisement are obtained as the target operation behaviors.
[0161] Before the user watches the current video, the user may have been pushed a target advertisement while watching other videos. Therefore, the user's historical operation behaviors on the target candidate advertisements can be obtained as the target operation behaviors.
[0162] Step 2: Obtain the historical rankings of the target candidate ads that the user has watched.
[0163] For the convenience of description, in the embodiment of the present disclosure, the level of the target candidate advertisements that the user has watched is referred to as a historical level.
[0164] In the embodiment of the present disclosure, the grade of an advertisement may be pre-set based on actual conditions such as the duration of the advertisement, the type of content, and the number of product links and interactive buttons included in the advertisement.
[0165] Step 3: Based on the target operation behavior and the historical level, the push level is determined, and the target candidate advertisement of the push version of the push level is used as the advertisement to be pushed.
[0166] It can be determined whether the number of active target operation behaviors has reached a preset number. If so, the previous level of the highest historical level can be used as the push level. If not, the highest historical level is used as the push level.
[0167] When the user's current emotion type is positive, it means that the user is in a good mood and may interact with the advertisement. Therefore, the push level is determined based on the target operation behavior and historical level. Higher-level advertisements can be pushed to the user. Through long-term preparation and guidance, the user will click on the advertisement, and then make product purchases or user registrations, thereby maximizing the interests of the advertiser.
[0168] In some embodiments of the present disclosure, the manner in which the electronic device determines the time to be inserted based on the candidate insertion times in S604 may be implemented as follows: determining the time to be inserted based on preset restriction parameters and the candidate insertion times.
[0169] The restriction parameters include at least one of the following: a minimum time interval between adjacent advertisements, a maximum number of advertisement insertions in a single video, a maximum number of advertisement insertions per unit time, and a minimum time interval for pushing the same advertisement to the user.
[0170] That is, obtain the time when the advertisement was last played during the current video playback, and determine whether the time interval between this time and the candidate insertion time exceeds the minimum time interval between adjacent advertisements. If so, the candidate insertion time can be deleted. Also, calculate the sum of the number of advertisements inserted in the current video and the number of candidate insertion times, and determine whether the sum exceeds the maximum number of advertisement insertions in a single video. If so, some candidate insertion times can be deleted so that the sum does not exceed the maximum number of advertisement insertions in a single video. Also, determine whether the time interval between the last time the advertisement to be pushed to the user and the candidate insertion time exceeds the minimum time interval for pushing the same advertisement to the user. If so, delete the candidate insertion time. Before deleting a candidate insertion time at any time, if there is only one candidate insertion time left, the candidate insertion time may not be deleted, and the candidate insertion time may be adjusted to meet the judgment conditions.
[0171] Optionally, the restriction parameters may also include other parameters, such as a minimum time interval for pushing advertisements of the same type to the user, etc., which is not specifically limited in the embodiment of the present disclosure.
[0172] Through the above method, the disclosed embodiments can control the maximum number of ad insertions within a video, reducing excessive interruptions to users. Furthermore, the time interval between pushed ads can be controlled, reducing the frequency of ads pushed to users. Furthermore, repeated ads can be reduced, ensuring that ads remain fresh for users.
[0173] In the above S105, after the electronic device determines the advertisement to be pushed and the time to be inserted based on the viewing time information, advertising preference information, emotional information and scene information using the large language model, the electronic device can send the advertisement to be pushed and the time to be inserted to the playback terminal of the current video, so that the playback terminal plays the advertisement to be pushed within the time to be inserted.
[0174] Afterwards, the electronic device may also update the advertisement recommendation strategy based on the user's feedback on the pushed advertisement, including the following steps:
[0175] Step I: Obtain the third operation behavior performed by the user on the pushed advertisement.
[0176] Step II: Based on the third operation behavior, adjust the ranking of the advertisement to be pushed among the alternative advertisements; and / or, based on the third operation behavior, adjust the minimum time interval for pushing the advertisement to be pushed to the user; and / or, based on the third operation behavior, adjust the ranking of related advertisements among the alternative advertisements, the related advertisements including advertisements of the same type as the advertisement to be pushed.
[0177] Alternative ads are ranked according to their push weights. Therefore, the ranking of the candidate ads can be adjusted by adjusting their push weights. When the first action is positive, the push weight of the candidate ad can be increased; when the first action is negative, the push weight of the candidate ad can be decreased. The large language model also considers the push weights of candidate ads when predicting click-through rates.
[0178] When the first operation behavior is of a positive type, the minimum time interval for pushing the advertisement to be pushed to the user can be reduced; when the first operation behavior is of a negative type, the minimum time interval for pushing the advertisement to be pushed to the user can be increased.
[0179] When the first operation behavior is of a positive type, the push weight of the related advertisement may be increased; and when the first operation behavior is of a negative type, the push weight of the related advertisement may be reduced.
[0180] In addition to the above adjustments, other parameters used in the advertisement recommendation process may also be adjusted based on the third operation behavior, which is not specifically limited in the embodiments of the present disclosure.
[0181] Through the above method, the advertising recommendation strategy can be continuously optimized based on the user's feedback behavior on the pushed advertisements, thereby increasing the possibility that the advertisements pushed to users will be liked by users and reducing the disturbance caused to users by pushing advertisements to users.
[0182] See also Figure 7 The following describes the overall process of the large-scale model-based advertising push method provided by the embodiment of the present disclosure:
[0183] Among them, the electronic device includes a user behavior monitoring module, an emotion recognition module, a video content analysis module, an advertising push and strategy linkage module, and an intelligent recommendation model module. These five modules are all software modules included in the electronic device.
[0184] During the playback of the current video, the playback terminal of the current video sends a first operation behavior performed by the user on the current video and a second operation behavior performed by the user on the pushed advertisement to the user behavior monitoring module.
[0185] The user behavior monitoring module determines the user's viewing time information for the current video based on the first operation behavior, and determines the user's advertising preference information for the pushed advertisement based on the second operation behavior, and sends the viewing time information and advertising preference information to the intelligent recommendation model module.
[0186] During the playback of the pushed advertisement, the playback terminal sends multiple sets of probabilities to the emotion recognition module.
[0187] The emotion recognition module determines the user's emotional information during the playback of the pushed advertisement based on the received multiple sets of probabilities, and sends the emotional information to the intelligent recommendation model module.
[0188] The video content analysis module determines scene information of multiple video clips included in the current video and sends the scene information to the intelligent recommendation model module.
[0189] The intelligent recommendation model module uses a large language model based on viewing time information, advertising preference information, emotional information and scene information to determine candidate advertisements, candidate advertisement versions and candidate insertion times, and sends the candidate advertisements, candidate advertisement versions and candidate insertion times to the advertising push and strategy linkage module.
[0190] The advertisement push and strategy linkage module determines the advertisement to be pushed based on the candidate advertisements and the candidate advertisement versions, and determines the time to be inserted based on the candidate insertion time.
[0191] The advertisement push and strategy linkage module sends the advertisement to be pushed and the time to be inserted to the playback terminal, so that the playback terminal plays the advertisement to be pushed within the time to be inserted.
[0192] The playback terminal sends the third operation behavior performed by the user on the advertisement to be pushed to the advertisement push and strategy linkage module.
[0193] The advertising push and strategy linkage module adjusts the advertising recommendation strategy based on the third operation behavior.
[0194] Figure 7 The specific implementation methods of each step can refer to the above description and will not be repeated here.
[0195] Compared with the method of pushing advertisements with the same theme as the current video to users, this method cannot optimize the timing of advertisement push. Even if the pushed advertisements may fall within the user's interest range, the push timing may still disrupt the user's viewing experience.
[0196] Compared with the method of pushing ads with high click-through rates to users, due to the different interest distributions of different users, using a unified method to push ads to all users will have low returns and it is difficult to significantly increase the click-through rate of ads.
[0197] The disclosed embodiment can simultaneously optimize the ads pushed and the timing of ad pushes while users are watching videos, increasing the likelihood that the pushed ads will be liked by users while reducing the disruption caused by ad pushes during video viewing. Furthermore, the disclosed embodiment can continuously learn user preferences during video playback, reducing the disruption caused by ads during video viewing and increasing ad click-through rates. This eliminates the need for additional manual adjustments and fundamentally improves ad recommendation effectiveness.
[0198] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of videos, advertisements and emotional information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0199] It should be noted that the large language model in this embodiment is not a model for a specific user and cannot reflect the personal information of a specific user.
[0200] It should be noted that the videos, advertisements, and emotional information in this embodiment may come from public datasets.
[0201] Based on the same inventive concept, corresponding to the above method embodiment, the embodiment of the present disclosure also provides an advertisement push device based on a large model, such as Figure 8 As shown, the device includes: a determination module 801 and an acquisition module 802;
[0202] Determining module 801, configured to determine, during playback of the current video, user viewing time information for the current video based on a first operation performed by the user on the current video;
[0203] An acquisition module 802 is configured to acquire user preference information for pushed advertisements, where the advertisement preference information is determined based on a second operation performed by the user on the pushed advertisements;
[0204] The acquisition module 802 is further used to obtain the user's emotional information during the playback of the pushed advertisement;
[0205] The acquisition module 802 is further configured to acquire scene information of multiple video clips included in the current video;
[0206] The determination module 801 is further configured to determine the advertisement to be pushed and the time to insert the advertisement to be pushed during the playback of the current video using a large language model based on viewing time information, advertisement preference information, emotion information, and scene information.
[0207] In some embodiments of the present disclosure, the determining module 801 is further configured to:
[0208] determining a first preference score of the user for the first pushed advertisement based on a type of a first operation performed by the user on the first pushed advertisement, where the first pushed advertisement is a pushed advertisement that has been inserted during playback of a current video;
[0209] Determining the advertisement type of the first pushed advertisement;
[0210] Obtaining a user preference for an advertisement type, where the preference is determined based on a second preference score of the user for a second pushed advertisement, where the second pushed advertisement is a pushed advertisement of a type that was inserted into videos previously viewed by the user;
[0211] The user's preference for the advertisement type is updated according to the first preference score to obtain updated user preferences for various advertisement types.
[0212] In some embodiments of the present disclosure, the determining module 801 is further configured to:
[0213] During the playback of the pushed advertisement, obtain the user video shot of the user;
[0214] For the video frames included in the user video, the emotion recognition model identifies the user's facial expressions and body language in the video frames, and outputs the probability of the user in the video frames being in various emotion types based on the facial expressions and body language;
[0215] Determine the emotion score of the video frame based on the probability of the user being in various emotion types in the video frame;
[0216] The emotion scores of consecutive video frames are smoothed to obtain the average emotion score.
[0217] In some embodiments of the present disclosure, the determining module 801 is further configured to:
[0218] During the playback of a pushed ad, the multiple sets of probabilities sent by the playback terminal of the pushed ad are received. Each set of probabilities represents the probability of the user being in various emotion types at a moment.
[0219] Based on the same set of probabilities, determine the user's emotion score at a moment;
[0220] The continuous sentiment scores are smoothed to obtain the average sentiment score.
[0221] In some embodiments of the present disclosure, the determining module 801 is further configured to:
[0222] Divide the current video into multiple video segments;
[0223] Using the emotion and content parsing model, image features are extracted from the video frames included in the video clip, text features are extracted from the text within the video frames, and audio features are extracted from the audio included in the video clip, and based on the feature extraction results, the emotion type of the video clip is determined;
[0224] Determine the emotional transition period based on the emotional type of each video clip;
[0225] Using the object detection and scene recognition model, perform image recognition on the video frames included in the video clip, and output the object detection result and the scene type of the video clip. The object detection result indicates whether the video clip contains a preset object and the playback time period of the preset object;
[0226] Based on the scene type of each video clip, the scene transition time period is determined.
[0227] In some embodiments of the present disclosure, the determining module 801 is specifically configured to:
[0228] Determining a plurality of recommended insertion times based on the scene information, the recommended insertion times including the playback time of the video clips belonging to non-climax scenes;
[0229] Based on the user's current mood type, filter the ad insertion trigger time from the recommended insertion times;
[0230] When the ad insertion trigger time arrives, the viewing time information, ad preference information, emotion information, and scene information are input into the large language model. The large language model is used to predict the user's click-through rate for each ad. If the click-through rate meets the preset conditions, candidate ads, candidate ad versions, and candidate insertion times are output based on the click-through rate. If the click-through rate does not meet the preset conditions, candidate ads, candidate ad versions, and candidate insertion times are output based on the alternative ads. The ad version represents the display format of the ad content.
[0231] Advertisements to be pushed are determined based on the candidate advertisements and the candidate advertisement versions, and the time to be inserted is determined based on the candidate insertion time.
[0232] In some embodiments of the present disclosure, the determining module 801 is specifically configured to:
[0233] Determining a predicted emotion type of the user at the recommended insertion time based on the user's current emotion type and scene information of the video clip included in the recommended insertion time;
[0234] Based on the predicted emotion type, the advertisement insertion trigger time is filtered out from each recommended insertion time.
[0235] In some embodiments of the present disclosure, the determining module 801 is specifically configured to:
[0236] selecting a target candidate advertisement from among the candidate advertisements;
[0237] Based on the total length of the current video, the user's current emotion type, and the user's historical operations on each ad version, the version to be pushed is filtered from the candidate ad versions, and the target candidate ad of the push version is used as the ad to be pushed.
[0238] In some embodiments of the present disclosure, the determining module 801 is specifically configured to:
[0239] The candidate ad with the highest candidate score is selected as the target candidate ad.
[0240] In some embodiments of the present disclosure, the determining module 801 is specifically configured to:
[0241] Obtain videos that are associated with the current video and videos that the user is interested in as related videos;
[0242] Based on the relevance between the relevant videos and the candidate ads, a target candidate ad is selected.
[0243] In some embodiments of the present disclosure, the determining module 801 is specifically configured to:
[0244] When the user's current emotion type is positive, the user's historical operation behavior on the target candidate advertisement is obtained as the target operation behavior;
[0245] Obtaining historical ratings of target candidate ads that have been viewed by the user, where the rating indicates the degree of guidance on click behavior;
[0246] Based on the target operation behavior and the historical level, the push level is determined, and the target candidate advertisement of the push version of the push level is used as the advertisement to be pushed.
[0247] In some embodiments of the present disclosure, the determining module 801 is specifically configured to:
[0248] Based on preset restriction parameters and candidate insertion times, the time to be inserted is determined. The restriction parameters include at least one of the following: the minimum time interval between adjacent advertisements, the maximum number of advertisement insertions in a single video, the maximum number of advertisement insertions per unit time, and the minimum time interval for pushing the same advertisement to users.
[0249] In some embodiments of the present disclosure, the apparatus further comprises:
[0250] The acquisition module 802 is further configured to obtain a third user operation on the pushed advertisement after determining the advertisement to be pushed and the time at which the advertisement to be pushed should be inserted during playback of the current video based on the viewing time information, the advertisement preference information, the emotion information, and the scene information using the large language model;
[0251] An adjustment module is used to adjust the order of the advertisement to be pushed among the alternative advertisements based on a third operation behavior; and / or, based on the third operation behavior, adjust the minimum time interval for pushing the advertisement to be pushed to the user; and / or, based on the third operation behavior, adjust the order of related advertisements among the alternative advertisements, the related advertisements including advertisements of the same type as the advertisement to be pushed.
[0252] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0253] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0254] like Figure 9As shown, electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of electronic device 900. Computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to bus 904.
[0255] Multiple components in the electronic device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0256] The computing unit 901 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as the large-model-based advertising push method. For example, in some embodiments, the large-model-based advertising push method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the large-model-based advertising push method described above can be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the large model-based advertisement push method in any other appropriate manner (eg, by means of firmware).
[0257] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0258] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0259] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0260] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0261] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0262] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0263] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0264] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for pushing advertisements based on a large model, comprising: During playback of a current video, determining, based on a first operation performed by a user on the current video, viewing time information of the user on the current video; Acquiring advertisement preference information of the user for the pushed advertisement, where the advertisement preference information is determined based on a second operation performed by the user on the pushed advertisement; Obtaining emotional information of the user during the playback of the pushed advertisement; Obtaining scene information of multiple video clips included in the current video; Based on the viewing time information, the advertising preference information, the emotion information, and the scene information, using a large language model, determining an advertisement to be pushed and a time at which the advertisement to be pushed is to be inserted during playback of the current video; The determining of the advertisement to be pushed and the time to insert the advertisement to be pushed during the playback of the current video using a large language model based on the viewing time information, the advertisement preference information, the emotion information, and the scene information includes: determining a plurality of recommended insertion times based on the scene information, the recommended insertion times including playback times of video clips belonging to non-climax scenes; Based on the current emotion type of the user, filtering out an advertisement insertion triggering time from each recommended insertion time; When the advertisement insertion trigger time is reached, the viewing time information, the advertisement preference information, the emotion information, and the scene information are input into the large language model, the large language model is used to predict the click-through rate of the user for each advertisement, and when the click-through rate meets a preset condition, candidate advertisements, candidate advertisement versions, and candidate insertion times are output based on the click-through rate; when the click-through rate does not meet the preset condition, candidate advertisements, candidate advertisement versions, and candidate insertion times are output based on the alternative advertisements, wherein the advertisement version indicates the display form of the advertisement content; The advertisement to be pushed is determined based on the candidate advertisements and the candidate advertisement versions, and the time to be inserted is determined based on the candidate insertion times.
2. The method according to claim 1, wherein the advertising preference information is determined by: determining a first preference score of the user for a first pushed advertisement based on a type of a first operation performed by the user on the first pushed advertisement, where the first pushed advertisement is a pushed advertisement that has been inserted during playback of the current video; Determining the advertisement type of the first pushed advertisement; Obtaining the user's preference for the advertisement type, where the preference is determined based on a second preference score of the user for a second pushed advertisement, where the second pushed advertisement is a pushed advertisement of the advertisement type that was inserted into a video previously watched by the user; The user's preference for the advertisement type is updated according to the first preference score to obtain updated user preferences for various advertisement types.
3. The method according to claim 1, wherein the emotion information is determined by: During the playback of the pushed advertisement, obtaining a user video shot of the user; For video frames included in the user video, identifying the user's facial expressions and body language in the video frames using an emotion recognition model, and outputting probabilities of the user being in various emotion types in the video frames based on the facial expressions and body language; determining an emotion score of the video frame based on probabilities of the user being in various emotion types within the video frame; The emotion scores of consecutive video frames are smoothed to obtain the average emotion score.
4. The method according to claim 1, wherein the emotion information is determined by: During the playback of the pushed advertisement, receiving multiple groups of probabilities sent by the playback terminal of the pushed advertisement, wherein each group of probabilities represents the probability of the user being in various emotion types at a moment; Determining an emotion score of the user at a moment based on the same set of probabilities; The continuous sentiment scores are smoothed to obtain the average sentiment score.
5. The method according to claim 1, wherein the scene information is determined by the following steps: Dividing the current video into multiple video segments; Using the emotion and content parsing model, performing image feature extraction on the video frames included in the video clip, performing text feature extraction on the text within the video frames, and performing audio feature extraction on the audio included in the video clip, and determining the emotion type to which the video clip belongs based on the feature extraction results; Determine the emotional transition period based on the emotional type of each video clip; Using an object detection and scene recognition model, perform image recognition on the video frames included in the video clip, and output an object detection result and a scene type of the video clip, wherein the object detection result indicates whether the video clip includes a preset object and a playback time period of the preset object; Based on the scene type of each video clip, the scene transition time period is determined.
6. The method according to claim 1, wherein The step of selecting an advertisement insertion trigger time from among the recommended insertion times based on the current emotion type of the user includes: determining a predicted emotion type of the user at the recommended insertion time based on the current emotion type of the user and scene information of the video clip included in the recommended insertion time; Based on the predicted emotion type, an advertisement insertion triggering time is filtered out from the recommended insertion times.
7. The method according to claim 1, wherein The determining the advertisement to be pushed based on the candidate advertisement and the candidate advertisement version includes: selecting a target candidate advertisement from among the candidate advertisements; According to the total duration of the current video, the current emotion type of the user and the historical operation behaviors performed by the user on each advertisement version, the version to be pushed is screened from the candidate advertisement versions, and the target candidate advertisement of the pushed version is used as the advertisement to be pushed.
8. The method according to claim 7, wherein: The candidate advertisements have candidate scores, and selecting a target candidate advertisement from the candidate advertisements includes: The candidate advertisement with the highest candidate score is selected as the target candidate advertisement.
9. The method according to claim 7, wherein: The step of selecting a target candidate advertisement from the candidate advertisements includes: Obtaining videos that are associated with the current video and videos of interest to the user as related videos; The target candidate advertisement is selected based on the relevance between the related video and the candidate advertisement.
10. The method according to any one of claims 7 to 9, wherein: The step of using the push version of the target candidate advertisement as the advertisement to be pushed includes: When the current emotion type of the user is positive, obtaining the historical operation behavior performed by the user on the target candidate advertisement as the target operation behavior; Obtaining a historical rating of the target candidate advertisement viewed by the user, wherein the rating indicates a degree of guidance on click behavior; Based on the target operation behavior and the historical level, a push level is determined, and the target candidate advertisement of the push version of the push level is used as the advertisement to be pushed.
11. The method according to any one of claims 1 to 9, wherein: The determining the to-be-inserted time based on the candidate insertion time includes: The time to be inserted is determined based on preset restriction parameters and the candidate insertion time, wherein the restriction parameters include at least one of the following: a minimum time interval between adjacent advertisements, a maximum number of advertisement insertions in a single video, a maximum number of advertisement insertions per unit time, and a minimum time interval for pushing the same advertisement to the user.
12. The method according to claim 11, after determining the advertisement to be pushed and the time to insert the advertisement to be pushed during playback of the current video using a large language model based on the viewing time information, the advertisement preference information, the emotion information, and the scene information, further comprising: Obtaining a third operation performed by the user on the advertisement to be pushed; Adjusting the order of the advertisement to be pushed among the candidate advertisements according to the third operation behavior; and / or, Adjusting a minimum time interval for pushing the advertisement to be pushed to the user according to the third operation behavior; and / or, According to the third operation behavior, the order of related advertisements in the candidate advertisements is adjusted, and the related advertisements include advertisements of the same advertisement type as the advertisement to be pushed.
13. An advertisement push device based on a large model, comprising: a determination module, configured to determine, during playback of a current video, information about a user's viewing time of the current video based on a first operation performed by the user on the current video; an acquisition module, configured to acquire advertisement preference information of the user for the pushed advertisement, wherein the advertisement preference information is determined based on a second operation performed by the user on the pushed advertisement; The acquisition module is further configured to acquire the user's emotional information during the playback of the pushed advertisement; The acquisition module is further configured to acquire scene information of a plurality of video clips included in the current video; The determination module is further configured to determine, based on the viewing time information, the advertising preference information, the emotion information, and the scene information, an advertisement to be pushed and a time at which the advertisement to be pushed is to be inserted during playback of the current video using a large language model; The determining module is specifically configured to: determining a plurality of recommended insertion times based on the scene information, the recommended insertion times including playback times of video clips belonging to non-climax scenes; Based on the current emotion type of the user, filtering out an advertisement insertion triggering time from each recommended insertion time; When the advertisement insertion trigger time is reached, the viewing time information, the advertisement preference information, the emotion information, and the scene information are input into the large language model, the large language model is used to predict the click-through rate of the user for each advertisement, and when the click-through rate meets a preset condition, candidate advertisements, candidate advertisement versions, and candidate insertion times are output based on the click-through rate; when the click-through rate does not meet the preset condition, candidate advertisements, candidate advertisement versions, and candidate insertion times are output based on the alternative advertisements, wherein the advertisement version indicates the display form of the advertisement content; The advertisement to be pushed is determined based on the candidate advertisements and the candidate advertisement versions, and the time to be inserted is determined based on the candidate insertion times.
14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 12.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-12.
16. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Advertisement data pushing method and device
CN113159836A
Information pushing method and electronic device utilizing method
US20220309534A1