Video intention understanding method and system based on recessive behavior entropy

By constructing a dynamic knowledge graph and calculating implicit behavior entropy using implicit user interaction behaviors, the edge weights of video clips and creative concepts are dynamically updated. This solves the problem of existing technologies being unable to capture users' implicit aesthetic preferences, and achieves a deep understanding of the value of video content and accurate recommendations.

CN121981126APending Publication Date: 2026-05-05GUANGZHOU TAIDONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU TAIDONG TECH CO LTD
Filing Date
2026-02-09
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing video recognition models cannot capture the real-time, implicit aesthetic preferences generated by users during interaction. Traditional video recommendation and mining algorithms ignore implicit behaviors, resulting in poor experience for video platform creators and inaccurate user recommendations.

Method used

A dynamic knowledge graph is constructed, and implicit behavior entropy is calculated by collecting users' implicit interaction behavior. The edge weights between video clips and creative concepts are dynamically updated, and semantic association abundance and creative value score are calculated based on implicit behavior entropy.

Benefits of technology

It enables precise mining of the deep creative core of videos, improves the quantification of video content value and the accuracy of recommendations, and enhances the experience for creators and users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121981126A_ABST
    Figure CN121981126A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a video intention understanding method and system based on recessive behavior entropy. The method comprises the following steps: for any video clip, collecting real-time interaction behaviors of a user on the video clip to obtain a hidden interaction sequence; calculating the recessive behavior entropy of the video clip according to the occurrence probability of various interaction behaviors in the recessive interaction sequence; dynamically updating the weight of the edge between the video clip node and the creative concept node based on the recessive behavior entropy; and for any video clip, calculating the semantic association abundance according to the weight of the edge connected with the video clip, and calculating the creative value score of the video clip according to the semantic association abundance of the video clip. According to the method, the value of the video can be deeply mined, and the experience of creators and users is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a video intent understanding method and system based on implicit behavioral entropy. Background Technology

[0002] In today's rapidly developing digital economy, short videos, micro-dramas, and AIGC (AI-generated content) are experiencing rapid growth. As this trend progresses, the way video content value is extracted has changed. Initially, the extraction of video content value focused primarily on the basic level, namely identifying physical tags such as objects, scenes, or people in the video—tags that were relatively intuitive and easy to obtain. However, with industry development and rising user demands, the focus of video content value extraction has now shifted to a cognitive-level analysis of deeper creative cores such as emotional tension, narrative logic, cultural metaphors, and aesthetic style. This shift reflects the market's increasing demands for a deeper and broader understanding of video content.

[0003] Despite the continuous development of video understanding technology, existing technologies still face numerous challenges when handling complex creative intentions. Existing video recognition models are inherently "static." These models rely heavily on predefined labeling systems during their construction. While these pre-defined labeling systems can meet basic video content recognition needs to some extent, they also have limitations. Because video recognition models are fixed and cannot be flexibly adjusted, they struggle to capture the instantaneous and implicit aesthetic preferences that users develop during interaction. For example, while watching a video, a user might suddenly develop a particular fondness for a color scheme or resonate with a unique narrative rhythm, but existing recognition models cannot promptly perceive and record these instantaneous, implicit aesthetic preferences.

[0004] In addition, traditional video recommendation and mining algorithms often fall into the "traffic trap," focusing excessively on explicit metrics such as completion rate and likes while neglecting user interaction processes. For example, a user's repeated replay of a specific camera movement clip or prolonged pauses at an artistic composition reveal a hidden desire for specific creative techniques. Existing knowledge graphs cannot dynamically adjust, resulting in simplistic recommendation results that fail to deliver truly valuable information to users. In summary, the insufficient depth of understanding of video intent in related technologies leads to a lack of exposure for some high-artistic-value content on video platforms, resulting in a poor experience for short video creators. Simultaneously, the difficulty for users to receive accurate video recommendations contributes to a poor user experience on short video platforms. Summary of the Invention

[0005] To deepen the understanding of video content and improve the experience for creators and users of short video platforms, this application provides a video intent understanding method and system based on implicit behavioral entropy.

[0006] Firstly, this application provides a video intent understanding method based on implicit behavioral entropy, employing the following technical solution: A video intent understanding method based on implicit behavioral entropy includes: constructing a dynamic knowledge graph, which includes at least video clip nodes and creative concept nodes, with edge connections between video clip nodes and creative concept nodes; For any video segment, the implicit interaction sequence is obtained by collecting the user's real-time interaction behavior with the video segment; the implicit behavior entropy of the video segment is calculated based on the occurrence probability of various interaction behaviors in the implicit interaction sequence. The weights of the edges between video clip nodes and creative concept nodes are dynamically updated based on implicit behavioral entropy. For any video segment, the semantic association abundance is calculated based on the weight of the edges connected to the video segment, and the creative value score of the video segment is calculated based on the semantic association abundance, where the semantic association abundance is positively correlated with the creative value score.

[0007] By collecting real-time user interactions with video clips and forming implicit interaction sequences, this approach more closely reflects users' actual attention distribution and cognitive interests compared to relying solely on explicit feedback metrics such as likes and completion rates. Subsequently, implicit behavioral entropy is calculated based on the probability of these interactions, enabling the system to quantify the intensity of user attention on specific video clips using information theory. This implicit behavioral entropy is then used to dynamically update the edge weights between video clip nodes and creative concept nodes, giving the knowledge graph continuous evolution capabilities. When a user shows higher attention to a particular video clip, the corresponding edge weight increases, thereby strengthening the semantic association between that video clip and related creative concepts. This user-behavior-driven dynamic update mechanism overcomes the static and fixed nature of existing knowledge graphs. Finally, by defining the sum of the weights of the edges connecting video clips as semantic association richness and calculating the creative value score accordingly, creative value assessment no longer relies solely on content similarity or manual rules but comprehensively reflects the semantic structure of the content and genuine user feedback. This improves the depth of value mining for video content and provides better feedback and experience for creators and users alike.

[0008] Optionally, the initial weight of the edge between any video segment node and the creative concept node is positively correlated with the similarity between the video features of the video segment and the semantic features of the creative concept.

[0009] By using the similarity between video features and semantic features to initialize edge weights, the graph structure already possesses basic semantic rationality before accumulating sufficient user interaction behavior, thus avoiding the instability caused by completely random or manually set weights.

[0010] Optionally, the weights of the edges between video clip nodes and creative concept nodes are dynamically updated based on the implicit behavioral entropy, including: for an edge between a video clip node and any creative concept node, calculating an update value based on the implicit behavioral entropy of the video clip node, weighting and fusing the current weight of the edge with the update value to obtain the latest weight, and using the latest weight to update the weight of the edge.

[0011] By weighting and fusing the original weights with the updated values ​​obtained based on implicit behavior entropy, the problem of drastic fluctuations in weights with single user behavior is avoided, so that the dynamic graph can reflect real-time changes in user interests while still retaining a certain amount of historically accumulated information.

[0012] Optionally, the update value is calculated based on the latent behavioral entropy of the video segment node, including: obtaining the feature similarity between the video segment node and the creative concept node, and linearly combining the feature similarity with the latent behavioral entropy to obtain the update value.

[0013] This ensures that the updated values ​​reflect both the objective similarity between the video content and the creative concept, and also incorporate the user's subjective attention during the actual viewing process, thereby reducing semantic bias caused by relying solely on behavioral data.

[0014] Optionally, the creative value score of a video segment can be calculated based on the semantic association abundance of the video segment, including: using the normalized result of the semantic association abundance as the creative value score.

[0015] By directly using the normalized result of semantic association abundance as the creative value score, the scoring results have a unified dimension and comparability. This technical feature simplifies the creative value assessment process and reduces computational complexity.

[0016] Optionally, the creative value score of a video segment is calculated based on the semantic association abundance of the video segment, including: using the normalized result of the semantic association abundance as the gain score, and using the weighted sum of the gain score and the base score as the creative value score; the steps for obtaining the base score include: calculating the similarity between the video features of the video segment and the semantic features corresponding to each benchmark word in the preset word library, and taking the maximum value of the similarity calculation results as the base score.

[0017] A weighted fusion mechanism of base score and gain score is introduced to make creative value scoring take into account both the quality of the content itself and user behavior feedback.

[0018] Optionally, the implicit behavior entropy of the video segment is calculated based on the occurrence probability of various interactive behaviors in the implicit interaction sequence, including: interactive behaviors are divided into positive interactive behaviors and negative interactive behaviors, positive interactive behaviors; all positive interactive behaviors in the implicit interaction sequence constitute a positive interaction set, and all negative interactive behaviors in the implicit interaction sequence constitute a negative interaction set; positive attention entropy is calculated based on the probability of each positive interactive behavior in the positive interaction set, and negative loss entropy is calculated based on the probability of each negative interactive behavior in the negative interaction set, and the difference between the positive attention entropy and the negative loss entropy is taken as the implicit behavior entropy.

[0019] Distinguishing between positive and negative interactive behaviors provides a more direct and accurate reflection of users' immediate cognitive responses and true attention distribution during viewing than explicit metrics such as likes and comments. Secondly, using information entropy to quantify the probability distribution of interactive behaviors identifies strong user exploratory interest in the deeper artistic details of the video. Finally, by subtracting positive attention entropy from negative churn entropy to offset redundant information interference, the true creative value of video clips can be accurately quantified.

[0020] Optionally, the similarity between the video features of the video clip and the semantic features corresponding to each benchmark word in the preset lexicon is calculated, including: extracting the video feature vector of the video clip and the semantic vector of each benchmark word in the preset lexicon, and using the cosine angle between the video feature vector and the semantic vector as the similarity between the video features of the video clip and the semantic features of the benchmark words in the preset lexicon.

[0021] Optionally, the steps for obtaining video clips include: segmenting the original video stream uploaded by the creator to obtain video clips.

[0022] Secondly, this application provides a video intent understanding system based on implicit behavioral entropy, employing the following technical solution: A video intent understanding system based on implicit behavioral entropy includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a video intent understanding method based on implicit behavioral entropy as described above is implemented.

[0023] The aforementioned video intent understanding method based on implicit behavioral entropy generates a computer program, which is then stored in memory for loading and execution by a processor. This allows for the creation of a system based on the memory and processor, making it convenient to use.

[0024] This application has the following technical effects: A dynamic knowledge graph is constructed, and implicit behavioral entropy is calculated by capturing implicit user interactions such as pausing, zooming, and rewinding. This entropy reflects the user's attention to the deep artistic details of the video. The system uses this entropy value to dynamically update the weights of the edges between the video and the creative concept, achieving accurate mining and value quantification of the deep creative core of the video, and solving the problem of insufficient depth in understanding complex creative intentions. Attached Figure Description

[0025] Figure 1 This is a flowchart of a video intent understanding method based on implicit behavioral entropy, according to an embodiment of this application. Detailed Implementation

[0026] This application discloses a video intent understanding method based on implicit behavioral entropy. By collecting users' implicit interactive behaviors such as pausing, zooming in, and rewinding in real time, a behavioral entropy model is constructed, and this entropy value is used to dynamically adjust the edge weights between video clips and creative concepts in a knowledge graph. This mechanism enables the system to keenly capture emerging creative needs in real time, breaking the limitations of static labels and achieving the beneficial effects of enhancing the accuracy of deep motivation recognition in videos and improving the objectivity of creative value assessment.

[0027] Reference Figure 1 A video intent understanding method based on latent behavioral entropy includes steps S1-S4.

[0028] S1: Construct a dynamic knowledge graph, which includes at least video clip nodes and creative concept nodes, with edge connections between video clip nodes and creative concept nodes.

[0029] In this embodiment, a knowledge graph that supports dynamic updates is first constructed. This graph contains at least three types of core nodes: the first type is video segment nodes. Each video clip has a distinct timestamp attribute. , This represents the start time of a video clip within the original video stream uploaded by the creator to the video platform; This represents the end time of a video clip within the original video stream uploaded by the creator to the video platform.

[0030] A video clip refers to a segment within the original video stream. In this embodiment, before interacting with the short video platform, computer vision algorithms (such as shot boundary detection) are used. The system automatically analyzes the raw video stream. When there are drastic pixel changes, histogram shifts, or scene transitions, the system automatically marks keyframes and divides them into different video segment nodes based on the continuity of the content. For example, a skiing video can be divided into multiple video segments with different timestamp ranges, such as "preparing to jump," "air flip," and "landing."

[0031] The second category is creative concept nodes. It is used to represent high-dimensional semantic concepts such as "loneliness", "industrial style" and "cyberpunk".

[0032] The association between any two nodes is determined by the edge between them, and the association between nodes is connected by the edge.

[0033] The initial weights of the edges are determined based on the pre-trained multimodal model. The weights of the edges mainly reflect the relationships between nodes. For example, the weight of the edge between a video clip and "loneliness" reflects the degree of association between the video clip and the emotion "loneliness".

[0034] In this embodiment, taking the initial weight of the edge between the video clip node and the creative concept node as an example, the following is adopted: ( The model extracts video features from video clip nodes and creative concept features from creative concept nodes. Specifically, video clip nodes extract video feature vectors through a visual encoder. Video features are high-dimensional mathematical vectors representing the visual content of the video, including physical attributes such as the geometric shape of objects, the composition and layout of the scene, color distribution, and light and shadow contrast.

[0035] The creative concept node is transformed into a computer-processable semantic vector using a text encoder. This represents the model's consensus understanding of the term after learning from massive amounts of internet text and image data. The cosine similarity between the video feature vector of the video clip node and the semantic vector of the creative concept node is calculated, and this similarity is used as the initial weight of the edge between the video clip node and the creative concept node.

[0036] S2: For any video segment, collect the user's real-time interactive behavior on the video segment to obtain the implicit interaction sequence; calculate the implicit behavior entropy of the video segment based on the occurrence probability of various interactive behaviors in the implicit interaction sequence.

[0037] The system collects implicit interaction sequences in real time through embedded points on the user's device while the user is watching the video. These interactive behaviors are not explicit likes or favorites, but rather deeper behaviors that reflect the distribution of user attention. Interactive behaviors can include: pausing video, zooming in on video, fast-forwarding video, and skipping video. This metric reflects the zoom level of users for specific screen details and video playback. This metric reflects users' repeated viewing behavior within the same timestamp range.

[0038] The video pause function is an interactive operation that allows users to stop the video from playing. After the video is paused, users can better observe the composition of elements in a certain video frame, reflecting the user's level of attention to the video.

[0039] Video zoom-in refers to the operation of enlarging the video image. This operation allows for a clearer observation of details in a certain part of the video, and also reflects the user's attention to the video.

[0040] Video playback refers to the user's ability to re-observe a specific segment of a video by adjusting the progress bar. For example, based on the timeline of a video segment, the video segment starts from... Play to At that time, drag the progress bar to The previous action constitutes an effective video playback interaction process. Repeatedly watching the video to learn elements and camera techniques also reflects the user's level of attention to the video segment.

[0041] Fast forwarding in video refers to playing a video at more than double the speed, reflecting that the user may not be interested in the current video.

[0042] Video skipping refers to the user's action of skipping the current video content using a progress bar. For example, based on the timeline of a video segment, the user skips from... Slide directly to The operation constitutes a valid video skip interaction process.

[0043] The implicit interaction sequence is formed by collecting the above interaction behaviors along the actual timeline. For ease of distinction, the real physical timeline is used as the actual timeline, and the timeline of the video clip is used as the video timeline. For example, assume there is a timestamp range of... Video clips within. On the actual timeline. The user plays a video clip, at which point the video timeline is... On the real timeline When the user pauses the video for the first time, the video timeline is as follows: Observe the background stretching effect; then on the real timeline. The video will continue playing. (On the actual timeline) At that time, video timeline At this point, to perform video playback interaction, drag the video timeline to... At this point, the actual timeline is... The video clip then continued playing. (The actual timeline is...) At that time, the user interacted with the video clip again by playing it back. (The actual timeline is...) At that time, the user zoomed in on the video clip. (On the actual timeline...) At that time, the user fast-forwarded the video clip. Based on the above description, it can be inferred that the user... , , , , If an interaction occurred with a video clip, the resulting implicit interaction sequence can be represented as: .

[0044] The above interactive behaviors can be divided into two categories: positive and negative. Positive interactive behaviors include pausing, zooming in, and rewinding the video; these behaviors reflect the user's level of attention to the video. Negative interactive behaviors include fast-forwarding and skipping the video; these behaviors reflect the user's level of disinterest in the video.

[0045] To measure the true level of user attention to a video segment, this embodiment introduces positive attention entropy and negative churn entropy. The difference between positive attention entropy and negative churn entropy is used as the implicit behavior entropy to characterize the true level of user attention to the video segment.

[0046] The formula for calculating the positive attention entropy can be expressed as: In the formula, Indicates video clip The corresponding positive attention entropy; Indicates video clip The set of all positive interaction behaviors in the corresponding implicit interaction sequence; This indicates a specific type of positive interaction behavior, such as pausing video, zooming in on video, or playing back video. Indicates interactive behavior The probability of appearing in the set of positive interaction behaviors.

[0047] It is important to note that if the set of all positive interaction behaviors is empty, the positive attention entropy is set to 0. If the set of all positive interaction behaviors contains only one positive behavior, the positive attention entropy is set to the preset base entropy, which is 1 in this embodiment.

[0048] Information entropy can effectively characterize the uncertainty and information content of random variables. When users engage in frequent and diverse interactive behaviors in a specific video segment, it indicates that the user pays more attention to the video segment and that the video segment contains implicit information beyond the conventional labels.

[0049] When users exhibit a higher frequency and more diverse range of specific interactive behaviors towards a particular video clip, such as repeatedly replaying a certain camera movement, it indicates that the probability distribution of the specific interactive behavior corresponding to that video clip in the overall sequence will tend towards a specific distribution characteristic, thus causing a change in latent behavior entropy. When the value of latent behavior entropy increases, it reflects a higher level of user attention to the deeper creative aspects of the video clip, suggesting that the video clip may contain high-value latent creative points.

[0050] The formula for calculating negative flow entropy can be expressed as: In the formula, Indicates video clip The corresponding negative loss entropy; Indicates video clip The set of all negative interaction behaviors in the corresponding implicit interaction sequence; This indicates a specific type of negative interactive behavior, such as pausing video, zooming in on video, or playing back video. Indicates interactive behavior The probability of appearing in the set of positive interaction behaviors.

[0051] It is important to note that if the set of all negative interaction behaviors is empty, the negative loss entropy is set to 0. If the set of all negative interaction behaviors contains only one type of negative behavior, the positive attention entropy is set to the preset base entropy, which is 1 in this embodiment.

[0052] Finally, the difference between the positive attention entropy and the negative loss entropy is calculated to obtain the implicit behavior entropy.

[0053] Assuming that there is only one replay in a certain implicit interaction sequence, the positive attention entropy is 1, the negative loss entropy is 0, and the final calculated implicit behavior entropy is 1.

[0054] Positive attention entropy indicates that the user's interest in the segment is more exploratory, involving both detailed viewing and repeated viewing. A higher negative churn entropy indicates a more decisive and diverse willingness to refuse to watch this type of video, suggesting higher redundancy for the user. A final calculated implicit behavior entropy greater than 0 indicates that while some parts of the segment were fast-forwarded, the user's in-depth exploration behavior was dominant, and the video segment possesses some creative value. A final calculated implicit behavior entropy less than 0 indicates that the intensity of negative churn behavior outweighs positive attention, suggesting lower value for the video segment.

[0055] S3: Dynamically update the weights of the edges between video clip nodes and creative concept nodes based on implicit behavioral entropy.

[0056] In order to enable the dynamic knowledge graph to perceive emerging creative needs in real time, for any given moment, the system calculates the latest weights based on the implicit behavioral entropy obtained in step S2, and uses the latest weights to dynamically adjust the weights of the edges between nodes.

[0057] For an edge between a video clip node and any creative concept node, calculate the update value based on the implicit behavioral entropy of the video clip node, weight the current weight of the edge and the update value to obtain the latest weight, and use the latest weight to update the weight of the edge.

[0058] The update value is calculated based on the latent behavioral entropy of the video clip nodes, including: obtaining the feature similarity between the video clip nodes and the creative concept nodes, and linearly combining the feature similarity with the latent behavioral entropy to obtain the update value.

[0059] Taking the latest weight of the edge between the video node and the creative concept node as an example, the formula for calculating the latest weight can be expressed as: In the formula, Indicates video clip With creative concepts The latest weight of the edges between them; Video clip With creative concepts The original weights of the edges between them; Indicates video clip The video feature vector; Represents creative concept nodes semantic feature vectors; Represents video feature vectors semantic feature vectors of creative concepts Cosine similarity between them; Indicates video clip The entropy of latent behavior; Represents a linear normalization function; This represents the weight adjustment factor, which determines the proportion of historically accumulated weights that are retained during the update process. In this embodiment, it is taken as 0.7 based on experience. This is a proportional fusion factor, mainly used to adjust the fusion priority between explicit content similarity and implicit behavioral contribution. In this embodiment, it is set to 0.6 based on experience.

[0060] In this embodiment, the entropy normalization process for latent behavior can be performed using... Standardization methods map latent behavioral entropy to Interval.

[0061] In the formula, The updated value reflects the increment in weights resulting from user interaction. The cosine similarity between video features and creative concept features reflects explicit content associations, while the normalized implicit behavioral entropy reflects the reinforcing contribution of user behavior to the creative association. When users show high attention to a video segment, the implicit behavioral entropy increases, thereby enhancing the dynamic edge weights between the video segment and related creative concepts during weight updates. This mechanism allows the video intent understanding system to discover creative ideas beyond cognitive understanding. For example, even if the video footage and the node "loneliness" have moderate similarity in the feature space, if a large number of users repeatedly replay the video segment, the system will confirm its deeper creative value by increasing the edge weights.

[0062] Here, when a user is watching a short video platform, the weight of the edges connecting to video clip nodes can be used to determine the next video to be recommended to the user. For example, if the edge between a video clip and the node "loneliness" has the largest weight, it means the user is currently interested in this type of video, and similar videos can be recommended to the user, thereby improving the user experience.

[0063] S4: For any video segment, calculate the semantic association abundance based on the weight of the edge connected to the video segment, and calculate the creative value score of the video segment based on the semantic association abundance, where the semantic association abundance is positively correlated with the creative value score.

[0064] After dynamically updating the edge weights, the system combines the topological attributes of the graph with a deep learning evaluation model to give a final creative value score to the video clip. .

[0065] In one embodiment, the formula for calculating the creative value score can be expressed as: In the formula, Indicates video clip Creative value score; Represents video clips in a dynamic knowledge graph The semantic association abundance of a node, i.e., its association with a video segment. The sum of the weights of all edges connected to a node; It is a linear normalization function.

[0066] Semantic association abundance reflects the richness of the creativity of a video clip. In other embodiments, semantic association abundance can also be calculated in the following way: count the weights of all edges connected to the video clip nodes, and record the number of creative concept nodes whose weights exceed a preset threshold, and use this number as the semantic association abundance.

[0067] The ratio of the semantic relevance abundance to the maximum value of a video clip node reflects the core position of that clip in the creative intent space; that is, the richer the creative concepts that the clip can express and the greater its edge weights, the higher its importance in the topological structure. Furthermore, because the edge weights are updated through implicit behavioral entropy, the creative value score can also reflect user approval.

[0068] For example, when a video clip is repeatedly replayed and zoomed in by users, its latent behavioral entropy increases significantly. Through the update in step S3, the edge weights between the video clip and related creative concept nodes increase. In the calculation of step S4, the semantic association abundance of the video clip node increases accordingly, and the final creative value score will also be at a higher level. This effectively solves the problem in traditional technologies that rely solely on mechanical indicators such as completion rate while neglecting the mining of humanistic value.

[0069] In another embodiment, the normalized result of semantic association abundance can be used as the gain score, and the result of the weighted sum of the gain score and the base score can be used as the creative value score. The steps for obtaining the base score include: calculating the similarity between the video features of the video segment and the semantic features corresponding to each benchmark word in the preset word library, and taking the maximum value in the similarity calculation results as the base score.

[0070] Specifically, video clips The formula for calculating the creative value score can also be expressed as: In the formula, Indicates video clip Creative value score; Represents video clips in a dynamic knowledge graph The semantic association abundance, i.e., the abundance of semantic association with video segments The sum of the weights of all edges connected to a node; It is a linear normalization function.

[0071] Indicates video clip The basic score is obtained through the following steps: First, a set of general terms representing high artistic value (bag of words) is preset, such as: {"cinematic", "exquisite composition", "light and shadow tension", "high definition", "professional photography"}; then, video clips are extracted using the CLIP visual encoder. Video features ;calculate The cosine similarity with the semantic vector of each benchmark word in the set is used as the base score, and the maximum value is taken as the base score.

[0072] By combining implicit behavioral entropy with dynamic graphs, the underlying motivations of videos are deeply explored, providing precise technical support for creative content creation in the AIGC field, thereby improving the exposure of creative videos. Alternatively, partial evaluation information can be pushed to creators, evaluating which segments of videos receive more attention from users and can be used as core material for secondary creation or other operations. In addition, on the user side, if a user generates a high creative value rating for a video segment, similar videos can be pushed to that user in subsequent recommendations.

[0073] This application also discloses a video intent understanding system based on implicit behavioral entropy, including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a video intent understanding method based on implicit behavioral entropy according to this application.

[0074] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.

[0075] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A video intent understanding method based on latent behavioral entropy, characterized in that, include: Construct a dynamic knowledge graph, which includes at least video clip nodes and creative concept nodes, with edge connections between video clip nodes and creative concept nodes. For any video segment, the implicit interaction sequence is obtained by collecting the user's real-time interaction behavior with the video segment; the implicit behavior entropy of the video segment is calculated based on the occurrence probability of various interaction behaviors in the implicit interaction sequence. The weights of the edges between video clip nodes and creative concept nodes are dynamically updated based on implicit behavioral entropy. For any video segment, the semantic association abundance is calculated based on the weight of the edges connected to the video segment, and the creative value score of the video segment is calculated based on the semantic association abundance, where the semantic association abundance is positively correlated with the creative value score.

2. The video intent understanding method based on latent behavioral entropy according to claim 1, characterized in that, The initial weight of the edge between any video segment node and the creative concept node is positively correlated with the similarity between the video features of the video segment and the semantic features of the creative concept.

3. The video intent understanding method based on latent behavioral entropy according to claim 1, characterized in that, The weights of the edges between video clip nodes and creative concept nodes are dynamically updated based on the implicit behavioral entropy. This includes: for an edge between a video clip node and any creative concept node, calculating an update value based on the implicit behavioral entropy of the video clip node, weighting and fusing the current weight of the edge with the update value to obtain the latest weight, and using the latest weight to update the weight of the edge.

4. The video intent understanding method based on latent behavioral entropy according to claim 3, characterized in that, The update value is calculated based on the latent behavioral entropy of the video clip nodes, including: obtaining the feature similarity between the video clip nodes and the creative concept nodes, and linearly combining the feature similarity with the latent behavioral entropy to obtain the update value.

5. The video intent understanding method based on latent behavioral entropy according to claim 1, characterized in that, The creative value score of a video clip is calculated based on the semantic association abundance of the video clip, including: using the normalized result of the semantic association abundance as the creative value score.

6. The video intent understanding method based on latent behavioral entropy according to claim 1, characterized in that, The creative value score of a video clip is calculated based on the semantic association abundance of the video clip, including: using the normalized result of the semantic association abundance as the gain score, and using the weighted sum of the gain score and the base score as the creative value score; the steps to obtain the base score include: calculating the similarity between the video features of the video clip and the semantic features corresponding to each benchmark word in the preset word library, and taking the maximum value of the similarity calculation results as the base score.

7. The video intent understanding method based on latent behavioral entropy according to claim 1, characterized in that, The implicit behavior entropy of a video segment is calculated based on the probability of occurrence of various interactive behaviors in the implicit interaction sequence. This includes: interactive behaviors are divided into positive and negative interactive behaviors; all positive interactive behaviors in the implicit interaction sequence constitute a positive interaction set, and all negative interactive behaviors in the implicit interaction sequence constitute a negative interaction set; positive attention entropy is calculated based on the probability of each positive interactive behavior in the positive interaction set, and negative loss entropy is calculated based on the probability of each negative interactive behavior in the negative interaction set; the difference between the positive attention entropy and the negative loss entropy is taken as the implicit behavior entropy.

8. A video intent understanding method based on latent behavioral entropy according to claim 6, characterized in that, Calculate the similarity between the video features of a video clip and the semantic features corresponding to each benchmark word in a preset lexicon, including: extracting the video feature vector of the video clip and the semantic vector of each benchmark word in the preset lexicon, and using the cosine angle between the video feature vector and the semantic vector as the similarity between the video features of the video clip and the semantic features of the benchmark words in the preset lexicon.

9. The video intent understanding method based on latent behavioral entropy according to claim 1, characterized in that, The steps for obtaining video clips include: segmenting the original video stream uploaded by the creator to obtain video clips.

10. A video intent understanding system based on latent behavioral entropy, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, implement a video intent understanding method based on implicit behavioral entropy according to any one of claims 1-9.