Personalized video push optimization method and device, storage medium and program product
By collecting and analyzing user input data and matching video content changes, the problem of inaccurate pushing of short video platforms during the initial user use stage is solved, and more efficient video push and user satisfaction are achieved.
Patent Information
- Application Number
- CN202510512156.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
It is difficult for existing short video platforms to accurately push video content during the initial use stage of users, resulting in users wasting time browsing irrelevant videos.
By collecting and parsing user input data, obtaining user preference information, and performing screenshots and parsing when video content changes, matching preference information and video information to control the interaction between the agent and the video application.
It realizes that users can understand user interests and needs more accurately in the early stages of using, reduce the time of browsing unrelated videos, improve users' efficiency in obtaining valuable video content, and improve user satisfaction.
Smart Images

Figure CN120034673A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and in particular to a personalized video push optimization method, device, storage medium, and program product. Background Art
[0002] Currently, various short video tools mainly push video information based on users' personal preferences. Their push mechanism usually relies on in-depth analysis of users' past browsing history, likes, comments and other multi-dimensional behavioral data, and then builds an accurate user interest portrait, based on which videos with similar content are pushed.
[0003] However, this seemingly efficient and intelligent push mode has significant drawbacks in actual application, especially when users first use the short video platform, the platform will push a large number of videos that do not meet the actual needs of users. Users have to spend a lot of time providing feedback to the platform until the platform can provide recommended content that is relatively in line with their personalized needs, which undoubtedly causes a huge waste of users' precious time. Summary of the invention
[0004] In view of this, the embodiments of the present disclosure provide a personalized video push optimization method, device, storage medium, and program product, which can solve the problem of inaccurate push of the target video application in the initial stage of user use, guide it to accurately push videos, thereby reducing the time users spend browsing irrelevant videos.
[0005] In a first aspect, the embodiments of the present disclosure provide a personalized video push optimization method, which adopts the following technical solutions: Collect user input data, parse the input data, and obtain user preference information; After the target video application is started, the dynamic changes of the content of the currently playing video are tracked, and the currently playing video is captured based on the degree of change of the video content to obtain a video screenshot; Analyze the video screenshot to obtain video information; The preference information is matched with the video information, and based on the matching result, the intelligent agent is controlled to perform interactive operations with the target video application.
[0006] Optionally, collecting user input data, parsing the input data, and obtaining user preference information includes: Receive user input data through the front-end interface, or capture user input data from a target video application; Classify and store each user's input data into a user preference information database; The preset multimodal macro model is used to read the input data of each user from the user preference information database, and the preference information of each user is output.
[0007] Optionally, tracking the dynamic changes of the content of the currently playing video, taking screenshots of the currently playing video based on the degree of change of the video content, and obtaining the video screenshots includes: Obtain a video stream of the currently playing video, and extract video frames at preset time intervals; Extracting scene features and key point positions of characters from the video frame; Based on the scene features, obtaining a scene change degree value of the current video content; Based on the key point positions of the characters, obtaining a value of a degree of change of the character's actions in the current video content; When the scene change degree value is greater than a first threshold, or the character action change degree value is greater than a second threshold, a screenshot is taken of the currently playing video to obtain a video screenshot.
[0008] Optionally, the personalized video push optimization method further includes: Collect current hardware resource data of user's smart terminal devices according to preset hardware resource indicators; Acquire an adjustment coefficient based on the current hardware resource data; Based on the adjustment coefficient, a preset first basic threshold and a preset second basic threshold, the first threshold and the second threshold are acquired.
[0009] Optionally, matching the preference information with the video information, and controlling the agent to interact with the target video application based on the matching result, includes: Inputting the preference information and the video information into a preset large language model to obtain classified or graded matching results; Based on the matching results, the intelligent agent is controlled to perform different simulated user operations to achieve interaction with the target video application; The target video application optimizes a video push mechanism of the target video application based on the interactive operation.
[0010] Optionally, controlling the agent to perform different simulated user operations based on the matching result to achieve interaction with the target video application includes: Triggering the agent to generate an operation instruction based on the matching result, wherein the operation instruction includes a first identifier of a target video application, a second identifier of a currently playing video, and an interactive operation type; Sending the operation instruction to a corresponding adaptation module based on the first identifier; Using the adaptation module to convert the operation instruction into a native interaction instruction recognizable by the target video application, the native interaction instruction includes a standard identifier converted from the second identifier and a standard interaction type converted from the interaction operation type; The native interaction instruction is executed to perform corresponding interaction operations on the currently playing video.
[0011] Optionally, the personalized video push optimization method further includes: When the interactive operation performed by the agent on the target video application does not include pausing the playback, the agent continues to track the dynamic changes of the currently playing video content, and takes a screenshot of the currently playing video based on the degree of change of the new video content to obtain a new video screenshot; The video information is updated based on the video screenshots accumulated by the currently playing video.
[0012] In a second aspect, the embodiment of the present disclosure further provides a personalized video push optimization system, which adopts the following technical solutions: A user input module, used to collect user input data, parse the input data, and obtain user preference information; The video screenshot module is used to track the dynamic changes of the currently playing video content after the target video application is started, and take a screenshot of the currently playing video based on the degree of change of the video content to obtain a video screenshot; A video analysis module, used to analyze the video screenshot to obtain video information; The information matching module is used to match the preference information with the video information, and control the intelligent agent to interact with the target video application based on the matching result.
[0013] In a third aspect, the embodiments of the present disclosure further provide a computer device, which adopts the following technical solution: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any of the personalized video push optimization methods described above.
[0014] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable a computer to execute any of the personalized video push optimization methods described above.
[0015] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, including a computer program / instruction, which implements the steps of any of the above methods when executed by a processor.
[0016] The personalized video push optimization method provided by the disclosed embodiment can more accurately understand the user's interests and needs before the user uses the target video application by actively collecting and parsing input data, and provide a reliable basis for subsequent optimization of video push. By real-time tracking of the dynamic changes of video content, the timing of screenshots can be determined, key moments can be accurately captured, invalid screenshots can be reduced, and by parsing such video screenshots based on dynamic changes, the understanding of video content can be enhanced, and accurate video information can be obtained. The obtained preference information is matched with the video information obtained by parsing the video screenshots, and it can be judged whether the user is interested in the currently playing video according to the matching results, and then the intelligent agent is controlled to simulate the user browsing the video and perform different interactive behaviors on the target video application. With the help of these specific interactive behaviors, more detailed information is fed back to the target video application, and the problem of inaccurate push of the application in the early stage of user use is gradually improved, and it is guided to accurately push videos, thereby reducing the time for users to browse irrelevant videos, improving the efficiency of users to obtain valuable video content, making the pushed videos more in line with user expectations, and improving user satisfaction.
[0017] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the following preferred embodiments are specifically cited and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 A flowchart of a personalized video push optimization method provided by an embodiment of the present disclosure; Figure 2 A schematic diagram of a process flow of an intelligent screenshot method provided by an embodiment of the present disclosure; Figure 3 A schematic diagram of a flow chart of a method for obtaining a first threshold and a second threshold provided in an embodiment of the present disclosure; Figure 4 A flowchart of a method for performing an interactive operation provided by an embodiment of the present disclosure; Figure 5 A block diagram of the principle of a personalized video push optimization system provided by an embodiment of the present disclosure; Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] The embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0021] It should be clear that the following embodiments of the present disclosure are described by specific specific examples, and those skilled in the art can easily understand other advantages and effects of the present disclosure from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. The present disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that the following embodiments and features in the embodiments can be combined with each other in the absence of conflict. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present disclosure.
[0022] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein may be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present disclosure, it should be understood by those skilled in the art that an aspect described herein may be implemented independently of any other aspect, and two or more of these aspects may be combined in various ways. For example, any number of aspects described herein may be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein may be used to implement this device and / or practice this method.
[0023] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present disclosure. The drawings only show components related to the present disclosure rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0024] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the aspects described may be practiced without these specific details.
[0025] Reference Figure 1The present disclosure provides a personalized video push optimization method, comprising the following steps: S1: Collect user input data, parse the input data, and obtain user preference information; S2: After the target video application is started, the dynamic changes of the content of the currently playing video are tracked, and the currently playing video is captured based on the degree of change of the video content to obtain a video screenshot; S3: parse the video screenshot to obtain video information; S4: Match the preference information with the video information, and control the agent to interact with the target video application based on the matching results.
[0026] At present, it is difficult for target video applications to obtain sufficient and accurate behavioral data in the early stages of user use. Due to the lack of rich data support, the constructed user interest portraits are bound to be biased, making it difficult to accurately capture the user's real interests and needs. At the same time, in order to cover various types of content as widely as possible to meet the potential needs of different users, the algorithm design of the target video application may be overly broad, resulting in a large number of irrelevant videos being pushed in the early stages of user use. In addition, during the data processing and analysis process, the target video application may not fully consider the dynamic changes in user interests and the differences in needs in different scenarios, resulting in a low degree of match between the pushed content and the actual needs of users.
[0027] The personalized video push optimization method provided by the present disclosure can more accurately understand the user's interests and needs before the user uses the target video application by actively collecting and parsing input data, providing a reliable basis for subsequent optimization of video push. By tracking the dynamic changes of video content in real time to determine the timing of screenshots, it is possible to accurately capture key moments and reduce invalid screenshots, and by parsing such video screenshots based on dynamic changes, it is possible to enhance the understanding of video content and obtain accurate video information. The obtained preference information is matched with the video information obtained by parsing the video screenshots, and it can be determined whether the user is interested in the currently playing video according to the matching results, and then the intelligent agent is controlled to simulate the user's browsing of the video and perform different interactive behaviors on the target video application. With the help of these specific interactive behaviors, more detailed information is fed back to the target video application, and the problem of inaccurate push of the application in the early stage of user use is gradually improved, and it is guided to accurately push videos, thereby reducing the time for users to browse irrelevant videos, improving the efficiency of users in obtaining valuable video content, making the pushed videos more in line with user expectations, and improving user satisfaction.
[0028] In S1, users can pre-select the target video applications to be monitored in the front-end interface of the system, and set personal interest preferences and disliked content. In addition, users can also make relevant settings directly in the target video applications. Among them, the system supports diversified data input methods, providing users with a convenient and personalized operation experience. Specific data input methods include: first, text input, users can directly type in their favorite video types, content elements and other detailed information in the input box; second, voice input, users describe their preferences through voice, and the system uses advanced voice recognition technology to convert voice into text for processing; third, image upload, users can upload pictures to express their likes or dislikes for specific types of videos, for example, uploading landscape pictures represents a preference for natural scenery videos; fourth, video upload, users can upload video clips they like or dislike, so as to clearly express their likes or dislikes for such videos.
[0029] If the user inputs data reflecting personal preferences in the target video application, the preference data is obtained from the target video application by data crawling or extraction.
[0030] Based on the above, the user's input data includes at least one of text, voice, personal preference images and personal preference video clips. These input data of each user are automatically collected and stored in a preset user preference information database, and the data is classified and sorted according to different users. The multimodal large model will read the input data of each user from the database separately and process it. By using the user preference information database, the multimodal large model can directly read the classified and sorted user data without repeated collection and classification, saving processing time and computing resources, and speeding up model processing.
[0031] Among them, the multimodal large model includes a multimodal input layer, a feature fusion layer, a semantic understanding layer, and a text generation layer. The multimodal input layer is responsible for receiving different modal data from the user preference information database and preprocessing these data. The processing methods include preliminary cleaning and preprocessing of the text, removing special characters, stop words, etc.; using speech recognition technology to convert speech signals into text, and also performing text preprocessing operations; using image feature extraction algorithms, such as convolutional neural networks (CNN), to extract the visual features of personal preference images and convert image information into feature vectors; performing frame processing on personal preference video clips, extracting features from each frame of the image, and extracting the audio information of the video and converting it into text, and comprehensively processing to obtain the feature representation of the video.
[0032] The feature fusion layer uses early fusion, late fusion or hybrid fusion methods to fuse the different modal features output by the multimodal input layer to obtain fused features. The semantic understanding layer uses pre-trained language models (such as the GPT series, BERT, etc.) to further process the fused features, map them to the semantic space, mine the semantic information, classify and cluster the features, understand the user's favorite video types, content elements and other information, output the semantic understanding results of user needs and pass them to the text generation layer. The text generation layer uses natural language generation technology (such as the sequence to sequence (Seq2Seq) model) to convert the semantic understanding results into a comprehensive text description of user preferences. This text description is the user's preference information.
[0033] During the training phase of the multimodal large model, a large amount of text data related to the video is collected, including the video title, description, comments, etc., as well as the user's evaluation and preference description of the video, to form a text sample set; a large amount of voice information related to the video is recorded, such as the user's verbal evaluation and recommendation of the video, and the corresponding text content is annotated to form a voice sample set; various types of pictures are collected, including video covers, screenshots, etc., and the video type, content elements and other information represented by the picture are annotated to form an image sample set; different types of videos are collected, and the content, type, label, etc. of the video are annotated, and the audio information of the video is extracted and the corresponding text is annotated to form a video clip sample set. According to the preset ratio, the collected text sample set, voice sample set, image sample set and video clip sample set are split to generate a variety of single modality sample sets and multimodal sample sets. For example, 70% of the text sample set is split into a text modality sample set, 70% of the voice sample set is split into a voice modality sample set, 70% of the image sample set is split into an image modality sample set, 70% of the video clip sample set is split into a video modality sample set, and the remaining 30% of the text sample set, voice sample set, image sample set and video clip sample set are combined into a multimodal sample set.
[0034] Split the multimodal large model into dedicated sub-models adapted to different modalities, and use a single modality sample set to pre-train the corresponding dedicated sub-model. For example, use the text modality sample set to train the text-specific sub-model to enable it to learn the characteristics and patterns of text data; use the voice modality sample set to train the voice-specific sub-model to enable it to master the laws of voice signals; use the image modality sample set to train the image-specific sub-model to extract the key features of the image; and use the video modality sample set to train the video-specific sub-model to enable it to understand the characteristics of the video content.
[0035] The pre-trained exclusive sub-models are integrated to form a fusion model, and the fusion model is jointly fine-tuned using a multimodal sample set. During the fine-tuning training process, the back propagation algorithm is used to adjust the parameters of the fusion model. The back propagation algorithm backpropagates the error information from the text generation layer to the multimodal input layer based on the error between the output of the fusion model and the true label, and updates each parameter in the fusion model based on this. By continuously iterating and optimizing the parameters, the fusion model can better integrate information from different modalities, accurately understand user needs, and finally generate a trained multimodal large model.
[0036] In S2, the system function bar has an option of "Automatically take control after the video tool is started". If the user checks this option, when the user opens the target video application, the system will automatically start monitoring the video playback content; if the user does not check this option, the system will remind the user to start the function of automatically simulating user behavior. After the user starts this function, the system will immediately start monitoring the video playback content and determine the timing of the screenshot based on the content of the currently playing video, providing strong support for the subsequent acquisition of video information.
[0037] Reference Figure 2 The flow chart of the intelligent screenshot method shown in the figure, "tracking the dynamic changes of the content of the currently playing video, taking screenshots of the currently playing video based on the degree of change of the video content, and obtaining video screenshots" includes the following steps: S21: Obtain the video stream of the currently playing video and extract video frames at preset intervals; S22: extracting scene features and key point positions of characters from video frames; S23: Based on the scene features, obtaining a scene change degree value of the current video content; S24: based on the key point positions of the characters, obtaining a value of a change degree of the character's action in the current video content; S25: When the scene change degree value is greater than the first threshold value, or the character action change degree value is greater than the second threshold value, a screenshot of the currently playing video is taken to obtain a video screenshot.
[0038] In S21, after the target video application is started, the video stream is obtained through the interface provided by the operating system (Android or iOS, etc.) or the target video application. For the video stream, video frames are extracted at fixed time intervals. These video frames will serve as basic data for subsequent analysis, where the fixed time interval is the preset time, and the specific value is adjusted according to actual conditions, for example, 0.1 seconds or 5 seconds.
[0039] In S22-S24, for scene switching detection, the scene features include at least one of color features and texture features. The video frame is converted to a specific color space, such as RGB or HSV color space, and then the color features of the video frame are extracted. A color histogram is constructed based on these color features, and the color histogram can effectively reflect the overall color distribution of the video frame. The gray level co-occurrence matrix (GLCM) algorithm is used to extract the texture features of the video frame.
[0040] In one embodiment, the similarity between the color histograms of adjacent video frames is calculated using methods such as histogram intersection method and Bhattacharyya distance, which is defined as the first similarity. The similarity between the texture features of adjacent video frames is calculated using the Euclidean distance algorithm or the cosine similarity algorithm, which is recorded as the second similarity. At this time, the current video content includes two adjacent video frames, and the scene change degree value can be the first similarity, the second similarity, or the weighted sum of the first similarity and the second similarity. This method is suitable for scenes when the preset time is large.
[0041] In another embodiment, for the video frames within a preset time range, the average value of the color histogram is calculated to obtain the average color histogram, and the average texture feature is calculated. For example, the video frames within every 5 seconds can be averaged. The average color histogram of the video frames within two adjacent preset time ranges is obtained, and the first similarity between the two average color histograms is calculated. At the same time, the average texture features of the video frames within two adjacent preset time ranges are obtained, and the second similarity between the two average texture features is calculated. At this time, the current video content includes video frames within two adjacent preset time ranges, and the degree of change of the character's movements can be the first similarity, the second similarity, or the weighted sum of the first similarity and the second similarity. This method is suitable for scenes where the preset time is relatively small.
[0042] In the field of human pose estimation, key points of characters in video frames are usually detected. These key points are specific location points that describe the human body posture and structure. For example, 17 key points are defined in the COCO dataset, including nose, eyes, ears, shoulders, elbows, wrists, hips, knees, and ankles.
[0043] Regarding character action change detection, in one embodiment, the key point position of each character in a video frame is accurately detected, and based on the position information of the character's key points in adjacent video frames, the displacement of each character's key point between adjacent frames is calculated, and then the average value of the displacement of all key points in adjacent video frames is calculated. The average value is the scene change degree value of the current video content. At this time, the current video content includes two adjacent video frames. This method is suitable for scenes when the preset time is large.
[0044] In another embodiment, the position information of the key points of the characters in the video frames within each preset time range is extracted, and based on the position information, the average displacement of the key points of the characters in the video frames within two adjacent preset time ranges is obtained, and the average value is the scene change degree value of the current video content. For example, for each video frame within 5 seconds, the average position of each key point within this period is first calculated, and then the displacement between the average positions in two adjacent time periods is calculated. At this time, the current video content includes video frames within two adjacent preset time ranges, and this method is suitable for scenes with smaller preset times.
[0045] In S25, the scene change degree value is compared with a preset first threshold value, and the character action change degree value is compared with a second threshold value. When either the scene change degree value or the character action change degree value is greater than its corresponding threshold value, a screenshot of the currently playing video is immediately taken. By using this method, when the target video application is playing a video, the time when a screenshot is required can be accurately located to obtain a screenshot that can contain key elements of the currently playing video.
[0046] In S3, in addition to automatically taking screenshots of the currently playing video and obtaining video screenshots, the audio and title of the currently playing video are also extracted. The multimodal large model technology is used to comprehensively understand the currently playing video through video screenshots, audio, and titles, and output a text summary of the currently playing video. This summary is the video information. The technical principle of extracting video information here is the same as the technical principle of obtaining preference information, which will not be repeated here.
[0047] In order to prevent excessive impact on system performance caused by excessive screenshot frequency, the system will adaptively adjust the first threshold and the second threshold according to the hardware resource status of the smart terminal device, so as to control the screenshot frequency within an appropriate range. Figure 3 A flowchart of a method for obtaining a first threshold value and a second threshold value is shown, and the method for obtaining the first threshold value and the second threshold value includes the following steps: S261: Collecting current hardware resource data of the user's smart terminal device according to preset hardware resource indicators; S262: Acquire an adjustment coefficient based on current hardware resource data; S263: Obtain a first threshold and a second threshold based on the adjustment coefficient, a preset basic threshold and a second basic threshold.
[0048] In S261, a series of different screenshot frequencies are set on the test device with similar hardware configuration to the target intelligent terminal device, such as 1 time per minute, 5 times per minute, 10 times per minute, etc. Run for a period of time (such as 30 minutes) at each screenshot frequency, and use the system's own performance monitoring tools (such as Windows system task manager, Linux system top command, macOS system activity monitor, etc.) or third-party performance monitoring software (such as HWMonitor, Master Lu, etc.) to monitor various indicators, including but not limited to CPU usage, memory usage, disk I / O read and write rate, GPU usage, etc. Compare the changes in various indicators under different screenshot frequencies, find out the indicators that change significantly with the increase of screenshot frequency, and determine these indicators as hardware resource indicators affected by the screenshot frequency.
[0049] In S262 and S263, when the target video application is started, the user is requested to obtain necessary permissions such as access to system performance information. After the application is successful, the current hardware resource data is obtained through the API provided by the device operating system according to the determined hardware resource indicators. The current hardware resource data is subjected to weighted summation and other operations to obtain an adjustment coefficient, the adjustment coefficient is multiplied by the first basic threshold to obtain the first threshold, and the adjustment coefficient is multiplied by the second basic threshold to obtain the second threshold.
[0050] In order to further reduce the impact of screenshot operations on system performance, the system will adaptively adjust the resolution of captured video screenshots according to the hardware resource status of the smart terminal device, and divide different intervals according to the adjustment coefficient. Each interval corresponds to a fixed screenshot resolution.
[0051] In S4, the matching results of preference information and video information are output to the Agent control module, and the agent automatically interacts with the target video application according to the matching results. Among them, the matching results are specifically divided into classification results or graded results. The first case is to use a binary classification method to distinguish the matching results into two categories: success and failure. When the match is successful, the agent will automatically perform a series of positive feedback operations, including liking, commenting, collecting, sharing, and continuing to watch the currently playing video; when the match fails, the agent will automatically perform negative feedback operations, including negatively marking the currently playing video (for example, long pressing the screen and selecting "dislike") and terminating the current video playback. This binary classification method has a simple and efficient logical architecture. It only needs to determine the compliance of the matching results according to the preset standards, thereby achieving low computing cost and fast processing of large-scale data. By clearly setting the threshold, this method can accurately screen out video content that matches the user's preferences. The agent then feeds back the user's preferences to the target video application through interactive operations, so that it can accurately grasp the user's preferences in a short time, and optimize its own mechanism accordingly, effectively reducing the probability of invalid recommendations, and significantly improving the recommendation efficiency and accuracy.
[0052] The second case is to divide the matching results into multiple levels according to the matching degree between the preference information and the video information, for example, into 5 levels, namely "especially like", "like", "normal", "dislike" and "very dislike". The agent performs different simulated user operations according to the level to achieve interaction with the target video application. For example, if the level is "especially like", the agent will like the currently playing video with a high probability, post positive comments, collect videos, search for related keywords, and may trigger sharing operations; if the level is "like", the agent will like, comment and collect the currently playing video with a certain probability; if the level is "normal", the agent will check the comment area of the currently playing video with a small probability; if the level is "dislike", the agent will negatively mark the currently playing video; if the level is "very dislike", the agent will negatively mark the currently playing video and immediately terminate the current video playback. For example, when simulating a user watching a food video, the agent monitors the video content in real time and generates video information (such as "a food broadcaster named XX is enjoying steak in a luxurious environment, with elegant background music and gorgeous color matching"), and predicts the matching degree based on the user's preference information (such as "the user likes food, photography, music, and prefers high-end restaurants with elegant environments"). The prediction result shows that the user's preference for the video is "like", and the agent controls the background to perform operations such as likes, comments and collections of the video with a certain probability. After the current video is played, the next video is automatically played. This multi-level classification method not only evaluates whether the user likes the video, but also further subdivides the user's preference degree to achieve refined capture of user preferences. The agent feeds back this refined preference information to the target video application through interactive operations, helping it to gradually optimize the video push mechanism. In the optimization process, the target video application is able to consider more dimensional content, making the video push mechanism more flexible, thereby meeting the personalized needs of different users in diverse scenarios.
[0053] Among them, after the current video is finished playing, the following two methods can be selected to automatically play the next video. One is to trigger the "automatic continuous play" function of the target video application through user active opening or system call, and play the next video in the video list in sequence according to the preset logic of the target video application; the other is to use the intelligent agent to simulate user instructions (such as simulating the sliding event of the touch screen or sending the corresponding mouse wheel event) to trigger the sliding operation of the video playback interface, so as to load and play the new video content.
[0054] In a specific implementation scheme, the method for obtaining the matching result is as follows: input the preference information and video information into a preset large language model (such as ChatGPT), perform semantic comparison and analysis on the preference information and video information through the large language model, and output the matching result. Before the large language model is put into practical application, judgment cases for binary classification (success and failure) or five-level classification (especially like, like, general, dislike, very dislike) need to be embedded in its prompt word (Prompt) so that the model can accurately understand the specific classification criteria for the degree of matching. This matching method using a large language model can deeply understand the semantics of preference information and video information, understand the context, and process complex comparison logic, thereby providing more accurate results.
[0055] In another specific implementation scheme, the method for obtaining the matching result is as follows: extract multiple first keywords from the user preference information and determine the emotional tendency of each first keyword, assign weights to each first keyword according to the emotional tendency, for example, assign positive weights (such as +1, +2, etc.) to keywords with positive emotional tendencies according to the degree of liking, and assign negative weights (such as -1, -2, etc.) to keywords with negative emotional tendencies according to the degree of dislike. Extract multiple second keywords from the video information, obtain the similarity between each second keyword and the first keyword through the cosine similarity algorithm or the Jaccard similarity algorithm, and use the weight of the first keyword as the weighting coefficient of the similarity, perform weighted summation on the similarities based on the weighting coefficient, obtain the matching score, and determine the matching result based on the matching score. This keyword matching method is simple to implement and has fast calculation speed, and is suitable for rapid screening of large-scale data.
[0056] Reference Figure 4 The flowchart of the interactive operation method shown, "controlling the agent to perform different simulated user operations based on the matching results to achieve interaction with the target video application" includes the following steps: S41: triggering the agent to generate an operation instruction based on the matching result, where the operation instruction includes a first identifier of the target video application, a second identifier of the currently playing video, and an interactive operation type; S42: Sending the operation instruction to the corresponding adaptation module based on the first identifier; S43: using an adaptation module to convert the operation instruction into a native interaction instruction recognizable by the target video application program, where the native interaction instruction includes a standard identifier converted from the second identifier and a standard interaction type converted from the interaction operation type; S44: Execute native interaction instructions to perform corresponding interactive operations on the currently playing video.
[0057] In S41, the operation instruction is a formatted instruction that includes necessary parameters for indicating the operation, wherein the first identifier may be the name, package name or unique number of the application program, which is used to clarify the specific application targeted by the operation, and the second identifier includes the title, ID or URL of the video, etc., which is used to accurately point to the currently playing video. The interactive operation type may be represented by different numbers or symbols, for example: "1" represents a like operation, "2" represents a share operation, "3" represents a favorite operation, "4" represents a pause operation, etc.
[0058] In S42, for each video application that has reached cooperation, a corresponding adaptation module is pre-built. Since each video application has a unique interface and protocol, the main function of the adaptation module is to convert the general operation instruction into a format that can be recognized by the corresponding application. When the intelligent agent generates an operation instruction, it will send the instruction to the corresponding adaptation module according to the first identifier in the instruction.
[0059] In S43, the adaptation module parses the operation instruction according to the interface specification and protocol of the target video application, maps and converts the second identifier and interactive operation type therein into a standard identifier and standard interactive type that can be directly executed by the target video application, thereby generating a native interactive instruction.
[0060] In S44, the adaptation module sends the converted native interaction instructions to the target video application for execution. After receiving the native interaction instructions, the application performs corresponding interactive operations on the currently playing video according to the instruction requirements, such as like, share, collect, etc.
[0061] There are differences in the interfaces and protocols of different video applications. This method uses the adapter module to convert instructions, so that the general operation instructions generated by the agent can adapt to various applications, realize effective interaction between the agent and different video applications, and enhance the compatibility and versatility of the system. For each video application, only one adapter module needs to be built to handle the instruction conversion, without the need to develop different interaction logics for each application for the agent, reducing the development workload and cost. In addition, the adapter module converts the general operation instructions into native interaction instructions that can be directly executed by the application, avoiding the application from performing additional parsing and processing of the instructions, so that the interaction behavior can be executed more quickly and accurately, and improving the interaction efficiency of the system. When the interface or protocol of the video application changes, only the corresponding adapter module needs to be modified and adjusted, without affecting the normal operation of the agent and other applications, which is convenient for the maintenance and subsequent expansion of the system.
[0062] Furthermore, the matching result of preference information and video information determines whether the agent continues to watch the currently playing video. If the interactive operation performed by the agent on the target video application includes pausing the video, the dynamic change tracking of the currently playing video content will be terminated. At the same time, the agent controls the target video application to play the next video, and then the next video is tracked for dynamic changes. If the interactive operation performed by the agent on the target video application does not include pausing the video, when the interactive operation performed by the agent on the target video application does not include pausing the video, the agent continues to track the dynamic changes of the currently playing video content, takes a screenshot of the currently playing video based on the degree of change of the new video content, obtains a new video screenshot, and then based on the accumulated video screenshots of the currently playing video, it can obtain new and more accurate video information, realize the update of the video information, and continue to match the new video information with the preference information until the currently playing video is paused or completed. Through this method, the agent can simulate the viewing habits of the user, and when the video content is not interested in, the agent will pause the video in time, and watch the video content of interest in full, so as to better adapt to the interactive needs of the target video application, and make the interactive behavior of the target video application closer to the operating habits of real users.
[0063] The target video application comprehensively collects various interactive behavior data based on interactive behaviors, such as play, pause, fast forward, fast rewind, like, comment, favorite, share, search keyword, etc. At the same time, it records relevant information of the video, including title, tag, category, duration, release time, etc. Based on the interactive behavior data, it helps the target video application to build and optimize user interest portraits and improve the push effect of the target video application.
[0064] In summary, the personalized video push optimization method disclosed in the present invention, with the help of multimodal data processing technology, can accurately extract the user's preference information and the video information of the currently playing video, and provide two matching mechanisms of binary classification and grade division for users to choose from, so as to meet the needs of different users. Based on the matching results of preference information and video information, the intelligent agent can automatically perform related operations, replace the user's manual intervention, realize automatic optimization of the video push mechanism, and automatically update the video content without the user manually operating the intelligent terminal device (covering various common devices such as mobile phones, tablets, computers, etc.), thereby freeing the user's hands and effectively saving the user's time.
[0065] Reference Figure 5 The present disclosure provides a personalized video push optimization system, including: The user input module 101 is used to collect user input data, parse the input data, and obtain user preference information; The video screenshot module 102 is used to track the dynamic changes of the currently playing video content after the target video application is started, and take a screenshot of the currently playing video based on the degree of change of the video content to obtain a video screenshot; The video analysis module 103 is used to analyze the video screenshot and obtain video information; The information matching module 104 is used to match the preference information with the video information, and control the intelligent agent to interact with the target video application based on the matching result.
[0066] The various variations and specific examples in the personalized video push optimization method provided above are also applicable to the personalized video push optimization system provided in the present invention. Through the above detailed description of the personalized video push optimization method, those skilled in the art can clearly know the implementation method of the personalized video push optimization system. For the sake of brevity of the specification, it will not be described in detail here.
[0067] The computer device according to the embodiment of the present disclosure includes a memory and a processor. The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, a random access memory (RAM) and / or a cache memory (cache), etc. The non-volatile memory may include, for example, a read-only memory (ROM), a hard disk, a flash memory, etc.
[0068] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory, so that the computer device performs all or part of the steps of the personalized video push optimization method of each embodiment of the present disclosure.
[0069] Those skilled in the art should be able to understand that in order to solve the technical problem of how to obtain a good user experience, the present embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the protection scope of the present disclosure.
[0070] like Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure is shown, which is a schematic diagram of the structure of a computer device suitable for implementing the embodiment of the present disclosure. Figure 6 The computer device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0071] like Figure 6 As shown, the computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage device into a random access memory (RAM). In the RAM, various programs and data required for the operation of the computer device are also stored. The processor, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0072] Typically, the following devices can be connected to the I / O interface: input devices such as sensors or visual information acquisition devices; output devices such as display screens; storage devices such as tapes, hard disks, etc.; and communication devices. The communication device can allow the computer device to communicate with other devices (such as edge computing devices) wirelessly or by wire to exchange data. Figure 6 A computer device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0073] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the personalized video push optimization method of the embodiment of the present disclosure are executed.
[0074] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0075] According to the computer-readable storage medium of the embodiment of the present disclosure, non-transitory computer-readable instructions are stored thereon. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the personalized video push optimization method of each embodiment of the present disclosure are executed.
[0076] The above-mentioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card) and media with built-in ROM (e.g., ROM box).
[0077] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0078] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.
[0079] In the present disclosure, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. The block diagrams of the devices, devices, equipment, and systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagram. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open words, referring to "including but not limited to", and can be used interchangeably with them. The words "or" and "and" used here refer to the words "and / or" and can be used interchangeably with them, unless the context clearly indicates otherwise. The words "such as" used here refer to the phrase "such as but not limited to", and can be used interchangeably with them.
[0080] Additionally, as used herein, "or" used in a list of items beginning with "at least one" indicates a separate list, so that, for example, a list of "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not mean that the example described is preferred or better than other examples.
[0081] It should also be noted that in the system and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0082] Various changes, substitutions, and modifications of the techniques described herein may be made without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of the present disclosure is not limited to the specific aspects of the processes, machines, manufactures, compositions of events, means, methods, and actions described above. Currently existing or later to be developed processes, machines, manufactures, compositions of events, means, methods, or actions that perform substantially the same functions or achieve substantially the same results as the corresponding aspects described herein may be utilized. Thus, the appended claims include such processes, machines, manufactures, compositions of events, means, methods, or actions within their scope.
[0083] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
[0084] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.
Claims
1. A personalized video push optimization method, characterized in that: include: Collect user input data, parse the input data, and obtain user preference information; After the target video application is started, the dynamic changes of the content of the currently playing video are tracked, and the currently playing video is captured based on the degree of change of the video content to obtain a video screenshot; Analyze the video screenshot to obtain video information; The preference information is matched with the video information, and based on the matching result, the intelligent agent is controlled to perform interactive operations with the target video application.
2. The personalized video push optimization method according to claim 1, characterized in that: The collecting of user input data, parsing the input data, and obtaining user preference information includes: Receive user input data through the front-end interface, or capture user input data from a target video application; Classify and store each user's input data into a user preference information database; The preset multimodal macro model is used to read the input data of each user from the user preference information database, and the preference information of each user is output.
3. The personalized video push optimization method according to claim 1, characterized in that: The tracking of dynamic changes in the content of the currently playing video, taking screenshots of the currently playing video based on the degree of change in the video content, and obtaining video screenshots includes: Obtain a video stream of the currently playing video, and extract video frames at preset time intervals; Extracting scene features and key point positions of characters from the video frame; Based on the scene features, obtaining a scene change degree value of the current video content; Based on the key point positions of the characters, obtaining a value of a degree of change of the character's actions in the current video content; When the scene change degree value is greater than a first threshold, or the character action change degree value is greater than a second threshold, a screenshot is taken of the currently playing video to obtain a video screenshot.
4. The personalized video push optimization method according to claim 3 is characterized in that: Also includes: Collect current hardware resource data of user's smart terminal devices according to preset hardware resource indicators; Acquire an adjustment coefficient based on the current hardware resource data; Based on the adjustment coefficient, a preset first basic threshold and a preset second basic threshold, the first threshold and the second threshold are acquired.
5. The personalized video push optimization method according to claim 1 or 2, characterized in that: The step of matching the preference information with the video information and controlling the agent to interact with the target video application based on the matching result includes: Inputting the preference information and the video information into a preset large language model to obtain classified or graded matching results; Based on the matching results, the intelligent agent is controlled to perform different simulated user operations to achieve interaction with the target video application; The target video application optimizes a video push mechanism of the target video application based on the interactive operation.
6. The personalized video push optimization method according to claim 5, characterized in that: The controlling the intelligent agent to perform different simulated user operations based on the matching result to achieve interaction with the target video application program includes: Triggering the agent to generate an operation instruction based on the matching result, wherein the operation instruction includes a first identifier of a target video application, a second identifier of a currently playing video, and an interactive operation type; Sending the operation instruction to a corresponding adaptation module based on the first identifier; Using the adaptation module to convert the operation instruction into a native interaction instruction recognizable by the target video application, the native interaction instruction includes a standard identifier converted from the second identifier and a standard interaction type converted from the interaction operation type; The native interaction instruction is executed to perform corresponding interaction operations on the currently playing video.
7. The personalized video push optimization method according to claim 3, characterized in that: Also includes: When the interactive operation performed by the agent on the target video application does not include pausing the playback, the agent continues to track the dynamic changes of the currently playing video content, and takes a screenshot of the currently playing video based on the degree of change of the new video content to obtain a new video screenshot; The video information is updated based on the video screenshots accumulated by the currently playing video.
8. A computer device, characterized in that: The computer device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the personalized video push optimization method described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which are used to enable a computer to execute the personalized video push optimization method described in any one of claims 1-7.
10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Training data generation method and device, electronic equipment and storage medium
CN112699910A
Video recommendation method and device, computer equipment and storage medium
CN113420181A
Content recommendation method and device, electronic equipment and storage medium
CN116628313A
Video content recommendation method of IPTV (Internet Protocol Television)
CN117750134A
Video recommendation method and device, storage medium and computer program product
CN118250517A