Dynamic resource scheduling method and system based on reinforcement learning and knowledge graph fusion

CN122736697APending Publication Date: 2026-09-11BEIJING HONGTU XINDA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610880350.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

这类方法存在以下显著不足:一是缺乏对剧情发展动态变化的理解能力,无法感知当前剧集的情节张力、情绪走向等细粒度上下文信息;二是缺乏对用户实时情绪状态的有效感知,无法判断用户在观看过程中的情感波动和心理状态;三是缺乏广告展示时序的动态优化能力,无法根据剧情情绪曲线和用户情绪状态选择最佳的广告介入时机,导致广告展示与剧情节奏和用户情绪错位,造成强烈的打扰感和用户体验下降

Benefits of technology

[0016] The beneficial effects of this invention are as follows: By using a lightweight on-device natural language processing model to analyze dialogue, background music emotion tags, and bullet screen sentiment in real time, and combining this with micro-behaviors such as user playback speed adjustments and repeated viewings, a "plot-emotion" temporal vector is constructed. This achieves fine-grained real-time perception of the dynamic development of the plot and the user's participation in the plot, overcoming the limitations of existing technologies that rely solely on coarse-grained content tags. Utilizing the on-device sensor capabilities available through quick apps, such as the front-facing camera or touch pressure sensors, and with user authorization, the invention perceives the user's facial micro-expressions or the force of their actions, inferring their immediate emotional state. This achieves real-time, seamless, and accurate perception of the user's emotional state, providing reliable user-side contextual features for ad emotion matching. Through a dynamic learning model, plot emotions and user emotions are used as strong contextual features for emotional resonance matching of ad creatives. Relaxing or motivating ads are matched after tense plot points, and heartwarming brand ads are matched after touching moments, significantly improving the emotional consistency between ads and content and reducing ad intrusion. By dynamically adjusting the timing of emotional engagement in ad display—such as delaying display after emotional climaxes and displaying retention ads earlier during emotional lows—the timing of ad placement is dynamically adapted to the plot rhythm and user emotional state, further enhancing ad conversion depth and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122736697A_ABST
    Figure CN122736697A_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent advertising recommendation technology, and in particular to a dynamic resource scheduling method and system based on the fusion of reinforcement learning and knowledge graphs. The method includes: analyzing the dialogue text, background music emotion tags, and bullet screen sentiment in real time using a lightweight natural language processing model on the device side to extract the emotional features of the plot; simultaneously collecting the user's micro-behavioral data; constructing a plot-emotion time-series vector based on the plot emotional features and micro-behavioral data; obtaining sensor perception data from the device side through quick apps to infer the user's immediate emotional state; inputting the plot-emotion time-series vector and the user's immediate emotional state into a dynamic learning model to generate an advertising matching strategy; and selecting advertising materials from the advertising library that have an emotional resonance matching relationship with the current plot emotion and the user's emotion according to the timing of emotional intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent advertising recommendation technology, and in particular to a dynamic resource scheduling method and system based on the fusion of reinforcement learning and knowledge graphs. Background Technology

[0002] With the rapid development of the mobile internet, short dramas, as an emerging form of content consumption, have quickly gained widespread attention due to their fast pace, tight plots, and immersive experience. Quick apps, as lightweight applications that can be used without downloading or installation, provide an efficient platform for the distribution and monetization of short dramas. Advertising is one of the core ways for short drama quick apps to achieve commercial monetization. How to organically integrate advertising content with the plot while ensuring a good user experience is a key technological challenge currently facing the industry.

[0003] Existing ad matching methods mostly rely on coarse-grained contextual features such as user profile tags, historical behavior statistics, and content category tags for ad recommendations. For example, they match ad categories using static profile data such as user age, gender, and interest tags, or perform contextual matching based on coarse-grained tags such as video content titles and categories. These methods have the following significant shortcomings: First, they lack the ability to understand the dynamic changes in plot development and cannot perceive fine-grained contextual information such as the tension and emotional trajectory of the current episode; second, they lack effective perception of the user's real-time emotional state and cannot judge the user's emotional fluctuations and psychological state during viewing; third, they lack the ability to dynamically optimize ad display timing, failing to select the optimal ad insertion time based on the plot's emotional curve and the user's emotional state, resulting in a misalignment between ad display and the plot's rhythm and the user's emotions, causing a strong sense of disturbance and a decline in user experience.

[0004] The information disclosed in this background section is intended only to enhance the understanding of the general background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0005] This invention provides a dynamic resource scheduling method and system based on the fusion of reinforcement learning and knowledge graphs, thereby effectively solving the problems in the background technology.

[0006] To achieve the above objectives, the technical solution adopted by this invention is: a dynamic resource scheduling method based on the fusion of reinforcement learning and knowledge graphs, comprising the following steps: The system uses a lightweight natural language processing model on the device to analyze the dialogue text, background music emotion tags, and bullet screen sentiment in real time to extract the emotional features of the plot; at the same time, it collects the user's micro-behavioral data, which includes at least one of the following: playback speed adjustment operation and repeated replay operation. Based on the plot emotion features and the micro-behavioral data, a plot-emotion temporal vector is constructed according to the time series. The user's real-time emotional state is inferred by acquiring data from edge sensors through quick apps; the edge sensors include at least one of a front-facing camera and a touch pressure sensor. The plot-emotion time sequence vector and the user's real-time emotional state are input into a dynamic learning model to generate an ad matching strategy. The ad matching strategy includes the selection of ad creatives to be matched and the timing of emotional intervention in ad display. According to the ad matching strategy, ad creatives that resonate with the current plot and user emotions are selected from the ad library, and the ads are displayed according to the emotional intervention timing.

[0007] Furthermore, the background music emotion label is obtained by: extracting the audio features of the background music on the device side, the audio features including at least rhythm, pitch and timbre features, and inputting the audio features into a pre-trained emotion classification model to output the emotion label corresponding to the background music.

[0008] Furthermore, the front-facing camera is activated after user authorization, and the lightweight facial expression recognition model on the device performs real-time analysis on the collected user facial images, identifies the user's facial micro-expression features, and maps them to emotion categories.

[0009] Furthermore, the touch pressure sensor collects the touch force value of the user on the quick app interface, and compares the force value with a preset threshold to determine the user's emotional state.

[0010] Furthermore, the dynamic learning model includes an edge-side online learning sub-model and a cloud-based deep reinforcement learning sub-model; The edge-side online learning sub-model is used for rapid incremental updates based on real-time user feedback, and the cloud-based deep reinforcement learning sub-model is used for global optimization training in the cloud and updates the edge-side online learning sub-model through model parameter distribution.

[0011] Furthermore, the timing of emotional intervention includes: delaying the display of advertisements after the emotional climax of the plot, or displaying retention advertisements in advance during the emotional low point of the plot.

[0012] Furthermore, the emotional resonance matching relationship is quantitatively matched by calculating the cosine similarity or Euclidean distance between the emotional vector of the advertising material and the plot-emotion time sequence vector.

[0013] This invention also includes a dynamic resource scheduling system based on the fusion of reinforcement learning and knowledge graphs, comprising: The edge data acquisition module is used to analyze the dialogue text, background music emotion tags, and bullet screen sentiment in real time through a lightweight edge natural language processing model to extract the emotional features of the plot; as well as collect the user's micro-behavioral data. The time-series vector construction module is used to construct a "plot-emotion" time-series vector based on the plot emotion features and the micro-behavioral data according to the time sequence. The user emotion perception module is used to obtain sensor data from the device side through the quick app and infer the user's real-time emotional state. The dynamic learning matching module is used to input the "plot-emotion" time-series vector and the user's real-time emotional state into the dynamic learning model to generate an advertising matching strategy. The ad display execution module is used to select ad creatives from the ad library and display ads according to the timing of emotional intervention, based on the ad matching strategy.

[0014] The present invention also includes a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described above.

[0015] The present invention also includes a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as described above.

[0016] The beneficial effects of this invention are as follows: By using a lightweight on-device natural language processing model to analyze dialogue, background music emotion tags, and bullet screen sentiment in real time, and combining this with micro-behaviors such as user playback speed adjustments and repeated viewings, a "plot-emotion" temporal vector is constructed. This achieves fine-grained real-time perception of the dynamic development of the plot and the user's participation in the plot, overcoming the limitations of existing technologies that rely solely on coarse-grained content tags. Utilizing the on-device sensor capabilities available through quick apps, such as the front-facing camera or touch pressure sensors, and with user authorization, the invention perceives the user's facial micro-expressions or the force of their actions, inferring their immediate emotional state. This achieves real-time, seamless, and accurate perception of the user's emotional state, providing reliable user-side contextual features for ad emotion matching. Through a dynamic learning model, plot emotions and user emotions are used as strong contextual features for emotional resonance matching of ad creatives. Relaxing or motivating ads are matched after tense plot points, and heartwarming brand ads are matched after touching moments, significantly improving the emotional consistency between ads and content and reducing ad intrusion. By dynamically adjusting the timing of emotional engagement in ad display—such as delaying display after emotional climaxes and displaying retention ads earlier during emotional lows—the timing of ad placement is dynamically adapted to the plot rhythm and user emotional state, further enhancing ad conversion depth and user experience. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a schematic diagram of the system structure of the present invention; Figure 3 This is a schematic diagram of the structure of the computer device of the present invention. Detailed Implementation

[0019] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0020] Example 1: like Figure 1 As shown: A dynamic resource scheduling method based on the fusion of reinforcement learning and knowledge graphs includes the following steps: The system uses a lightweight natural language processing model on the device to analyze the dialogue text, background music emotion tags, and bullet screen sentiment in real time to extract the emotional features of the plot; at the same time, it collects micro-behavioral data of users, including at least one of the following: playback speed adjustment operation and repeated replay operation. Based on the emotional characteristics of the plot and micro-behavioral data, a plot-emotion time-series vector is constructed according to the time series. The app acquires data from on-device sensors to infer the user’s real-time emotional state; the on-device sensors include at least one of the front-facing camera and a touch pressure sensor. The plot-emotion time sequence vector and the user's real-time emotional state are input into the dynamic learning model to generate an ad matching strategy. The ad matching strategy includes the selection of ad creatives to be matched and the timing of emotional intervention in ad display. Based on the ad matching strategy, ad creatives that resonate with the current plot and user emotions are selected from the ad library, and the ads are displayed according to the timing of emotional engagement.

[0021] By using a lightweight on-device natural language processing model to analyze dialogue, background music emotion tags, and bullet screen sentiment in real time, and combining this with micro-behaviors such as user playback speed adjustments and repeated viewings, a "plot-emotion" temporal vector is constructed. This enables fine-grained real-time perception of the dynamic development of the plot and user engagement, overcoming the limitations of existing technologies that rely solely on coarse-grained content tags. Leveraging on-device sensor capabilities such as the front-facing camera or touch pressure sensors available in quick apps, and with user authorization, the system perceives user facial micro-expressions or the force of their actions to infer their immediate emotional state. This achieves real-time, seamless, and accurate perception of user emotional states, providing reliable user-side contextual features for ad emotion matching. A dynamic learning model uses plot emotions and user emotions as strong contextual features to match ad creatives to emotional resonance. It matches soothing or motivating ads after tense scenes and heartwarming brand ads after touching moments, significantly improving the emotional consistency between ads and content and reducing ad intrusion. By dynamically adjusting the timing of emotional engagement in ad display—such as delaying display after emotional climaxes and displaying retention ads earlier during emotional lows—the timing of ad placement is dynamically adapted to the plot rhythm and user emotional state, further enhancing ad conversion depth and user experience.

[0022] In this embodiment, the background music emotion label is obtained in the following way: the audio features of the background music are extracted on the device side, including at least rhythm, pitch and timbre features, and the audio features are input into a pre-trained emotion classification model to output the emotion label corresponding to the background music.

[0023] The front-facing camera is activated after user authorization, and a lightweight facial expression recognition model on the device analyzes the collected user facial images in real time, identifies the user's facial micro-expression features, and maps them to emotion categories.

[0024] The touch pressure sensor collects the force value of the user's touch operation on the quick app interface, and compares the force value with a preset threshold to determine the user's emotional state.

[0025] The dynamic learning model includes an edge-side online learning sub-model and a cloud-based deep reinforcement learning sub-model. The edge-side online learning sub-model is used for rapid incremental updates based on real-time user feedback, while the cloud-based deep reinforcement learning sub-model is used for global optimization training in the cloud and updates the edge-side online learning sub-model through model parameter distribution.

[0026] As a preferred embodiment of the above, the timing of emotional intervention includes: delaying the display of advertisements after the emotional climax of the plot, or displaying retention advertisements in advance during the emotional low point of the plot.

[0027] In this embodiment, the emotional resonance matching relationship is quantitatively matched by calculating the cosine similarity or Euclidean distance between the emotional vector of the advertising material and the plot-emotion time sequence vector.

[0028] Example 2: The technical solution of the present invention will be described in detail below with reference to a specific implementation scenario. This embodiment takes an urban romance short drama as the application scenario. The short drama is played on a certain quick app platform, and the platform needs to insert advertisements during the playback of the drama in order to realize commercial monetization.

[0029] Step S1: Real-time data acquisition and extraction of emotional features from the device side Users open and watch a short urban romance drama through a quick app. During playback, a lightweight natural language processing model on the device (such as a lightweight text analysis model based on MobileBERT or TinyBERT) performs real-time semantic analysis on the dialogue text of the currently playing episode, extracting emotional tendency scores (such as emotional polarity values ​​between -1 and 1) from the dialogue. Simultaneously, the device processes the background music audio signal in real-time, extracting audio features such as rhythm (BPM value), pitch (fundamental frequency distribution), and timbre (Mel-frequency cepstral coefficients, MFCC), and inputs these features into a pre-deployed emotion classification model on the device (such as a lightweight audio emotion classification model based on MobileNet), outputting the emotion label corresponding to the current background music, such as "tense," "relaxed," "moved," "cheerful," or "sad." Taking the urban romance drama in this embodiment as an example, in a scene of misunderstanding and conflict between the male and female leads, the background music emotion label is identified as "tense," and in a scene of reconciliation after the misunderstanding is resolved, the background music emotion label is identified as "moved."

[0030] In addition, the device also collects real-time data on the bullet comments (danmu) of the current episode, using a lightweight sentiment analysis model to analyze the sentiment of the bullet comment text and calculate the ratio of positive to negative bullet comments as a reference feature of the audience's sentiment. Simultaneously, the device collects micro-behavioral data from users, including whether they adjust the playback speed (e.g., from 1.0x to 0.75x slow motion or 2.0x fast forward) and whether they repeatedly rewatch a particular segment (e.g., dragging the progress bar to replay the same segment more than twice). This data is continuously recorded and timestamped.

[0031] Step S2: Construct the "Plot-Emotion" temporal vector Based on the plot emotion features (dialogue emotion polarity, background music emotion tags, and bullet screen emotion tendencies) extracted in step S1 and the user's micro-behavioral data, a "plot-emotion" temporal vector is constructed by fusion encoding according to the time sequence. Specifically, multi-source data is aligned and aggregated using preset time windows (e.g., every 5 seconds). Each time window generates a feature vector containing the following dimensions: dialogue emotion polarity value, one-hot encoding of background music emotion tags, proportion of positive emotion in bullet screen comments, proportion of negative emotion in bullet screen comments, user playback speed change flag, and number of times the user repeatedly rewatches the clip. The feature vectors of all time windows are concatenated in chronological order to form a complete "plot-emotion" temporal vector sequence. Taking this embodiment as an example, within the 2-minute segment of the conflict scene between the male and female protagonists, the plot emotion temporal vector shows a clear peak of tension, and the user repeatedly rewatches the clip, further strengthening the emotional weight of that segment.

[0032] Step S3: Real-time user emotion perception With user authorization, the device uses the Quick App feature to access the front-facing camera and captures facial images of the user at a sampling rate of 1 frame per second during viewing. A lightweight facial expression recognition model deployed on the device (such as a lightweight model based on MobileFaceNet) analyzes the captured facial images in real time, identifying the user's facial action unit (AU) features, including micro-expression features such as raised eyebrows, upturned corners of the mouth, and tightened lips. These features are then mapped to basic emotion categories (happiness, sadness, anger, surprise, fear, and disgust) and complex emotional states (such as a mixture of emotion and sadness, or tension mixed with fear and focus). Simultaneously, the device uses a touch pressure sensor to collect the user's touch force on the Quick App interface. When an abnormally increased force is detected when the user slides the screen or clicks a button, the device, combined with the contextual emotional state, infers that the user is currently in a state of tension or excitement.

[0033] Taking this embodiment as an example, in a conflict scene between the male and female protagonists, the front-facing camera captured micro-expression features on the user's face, such as furrowed brows and slightly parted lips. The facial expression recognition model determined that the user was in a state of tension, with a confidence level of 87%. Combined with the touch pressure sensor detecting that the user's touch on the screen increased significantly (from the baseline value of 0.2N to 0.8N) while watching the clip, this further confirmed the user's tension.

[0034] Step S4: Dynamic learning model generates ad matching strategy The "plot-emotion" time-series vector constructed in step S2 and the user's real-time emotional state inferred in step S3 are input into the dynamic learning model. In this embodiment, the dynamic learning model includes an on-device online learning sub-model and a cloud-based deep reinforcement learning sub-model.

[0035] The on-device online learning sub-model adopts an online random forest or incremental logistic regression model, and performs fast incremental update according to the user's feedback after each advertisement impression (such as whether the user clicks the advertisement, whether the user skips the advertisement, and the advertisement viewing duration, etc.) to realize personalized real-time adaptation. The cloud deep reinforcement learning sub-model is deployed on a cloud server, and performs deep reinforcement learning training (using DQN or PPO algorithm) based on historical data of large-scale user groups, takes advertisement click-through rate, conversion rate and user retention rate as a comprehensive reward function, and globally optimizes the advertisement selection strategy and the display timing strategy. The model parameters in the cloud are periodically (such as daily or weekly) delivered to the end side after model compression and quantization, so as to update the base model parameters of the on-device online learning sub-model.

[0036] In this embodiment, the dynamic learning model generates an advertisement matching strategy according to the current "plot-emotion" time series vector (represented by a peak tension emotion) and the user's real-time emotional state (tense and highly focused): select soothing or incentive advertisement materials (such as tourism vacation advertisements and fitness incentive advertisements), and display the advertisement 15 seconds after the emotion peak to avoid interrupting the viewing experience when the user is highly engaged in the plot.

[0037] Step S5: Matching advertisement materials and displaying at emotional intervention timing According to the advertisement matching strategy generated in step S4, advertisement materials having an emotional resonance matching relationship with the current plot emotion and user emotion are selected from an advertisement library. Specifically, each advertisement material in the advertisement library is pre-labeled with an emotional vector, which includes emotional tone labels (such as soothing, incentive, warm, joyful, tense, etc.) of the advertisement and an emotional intensity value. The system calculates the cosine similarity between the emotional vector of the advertisement material and the "plot-emotion" time series vector, and selects the advertisement material with the highest similarity for delivery.

[0038] In this embodiment, after the conflict scene ends, the plot enters the reconciliation stage, at this time the plot emotion changes from tension to moving, and the user's emotion also changes to moving mixed with relaxation. The dynamic learning model selects a warm brand advertisement (such as a family-themed advertisement of a certain jewelry brand), and displays it at the 10th second after the moving emotion appears (the emotion fermentation point). After the advertisement is displayed, the user watches the advertisement completely and generates a click behavior, and the system records this feedback data and uses it for subsequent model update.

[0039] In another scenario, when it is detected that the plot enters a plain period and the user's emotion is in a low state (relaxed but distracted, and may have a tendency to exit viewing soon), the system dynamically selects retention advertisements (such as the preview of the next episode or interactive advertisements) and displays them in advance to reduce the user churn rate.

[0040] Example 3: As Figure 2As shown, this embodiment provides a dynamic resource scheduling system based on the fusion of reinforcement learning and knowledge graphs to implement the above method. The system includes: The edge data acquisition module is deployed on the user's terminal device and includes a lightweight natural language processing model, an audio emotion classification model, and a bullet screen sentiment analysis model. It is used to analyze the dialogue text, background music emotion tags, and bullet screen sentiment in real time to extract the emotional features of the plot. At the same time, it collects the user's micro-behavioral data, including playback speed adjustment operations and repeated replay operations.

[0041] Temporal vector construction module: Deployed on user terminal devices, it is used to construct a "plot-emotion" temporal vector by fusing and encoding the plot emotion features and the micro-behavioral data according to the time sequence.

[0042] User Emotion Perception Module: Deployed on the user's terminal device, it is used to acquire sensor data from the device side through quick apps, including facial micro-expression images captured by the front camera and user operation force data captured by the touch pressure sensor, and infer the user's real-time emotional state using a lightweight facial expression recognition model.

[0043] The dynamic learning matching module includes an on-device online learning sub-model deployed on the user's terminal device and a cloud-based deep reinforcement learning sub-model deployed on the cloud server. It is used to input the "plot-emotion" time-series vector and the user's real-time emotional state into the dynamic learning model to generate an advertising matching strategy.

[0044] Ad display execution module: Deployed on the user's terminal device, it is used to select ad creatives from the ad library according to the ad matching strategy and display ads according to the emotional intervention timing, and collect user feedback data on the ads.

[0045] The modules above communicate with each other through the internal communication mechanism of Quick App, and the device and cloud synchronize model parameters and aggregated data through network communication protocols.

[0046] Example 4: Please see Figure 3 The diagram shows a structural schematic of a computer device provided in an embodiment of this application. An embodiment of this application provides a computer device 400, including a processor 410 and a memory 420. The memory 420 stores a computer program executable by the processor 410. When the computer program is executed by the processor 410, it performs the method described above.

[0047] This application embodiment also provides a storage medium 430, on which a computer program is stored, and the computer program is executed by a processor 410 to perform the above method.

[0048] The storage medium 430 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0049] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. "A plurality of" means two or more, unless otherwise explicitly specified.

[0050] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0051] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0052] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0053] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0054] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0055] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0056] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A dynamic resource scheduling method based on the fusion of reinforcement learning and knowledge graph, characterized in that, Includes the following steps: The system uses a lightweight natural language processing model on the device to analyze the dialogue text, background music emotion tags, and bullet screen sentiment in real time to extract the emotional features of the plot; at the same time, it collects the user's micro-behavioral data, which includes at least one of the following: playback speed adjustment operation and repeated replay operation. Based on the plot emotion features and the micro-behavioral data, a plot-emotion temporal vector is constructed according to the time series. The user's real-time emotional state is inferred by acquiring data from edge sensors through quick apps; the edge sensors include at least one of a front-facing camera and a touch pressure sensor. The plot-emotion time sequence vector and the user's real-time emotional state are input into a dynamic learning model to generate an ad matching strategy. The ad matching strategy includes the selection of ad creatives to be matched and the timing of emotional intervention in ad display. According to the ad matching strategy, ad creatives that resonate with the current plot and user emotions are selected from the ad library, and the ads are displayed according to the emotional intervention timing.

2. The method according to claim 1, characterized in that, The background music emotion label is obtained in the following way: the audio features of the background music are extracted on the device side, the audio features include at least rhythm, pitch and timbre features, and the audio features are input into a pre-trained emotion classification model to output the emotion label corresponding to the background music.

3. The method according to claim 1, characterized in that, The front-facing camera is activated after user authorization, and the lightweight facial expression recognition model on the device performs real-time analysis on the collected user facial images, identifies the user's facial micro-expression features, and maps them to emotion categories.

4. The method according to claim 1, characterized in that, The touch pressure sensor collects the force value of the user's touch operation on the quick app interface, and compares the force value with a preset threshold to determine the user's emotional state.

5. The method according to claim 1, characterized in that, The dynamic learning model includes an edge-side online learning sub-model and a cloud-based deep reinforcement learning sub-model. The edge-side online learning sub-model is used for rapid incremental updates based on real-time user feedback, and the cloud-based deep reinforcement learning sub-model is used for global optimization training in the cloud and updates the edge-side online learning sub-model through model parameter distribution.

6. The method according to claim 1, characterized in that, The timing of emotional intervention includes: delaying the display of advertisements after the emotional climax of the plot, or displaying retention advertisements in advance during the emotional low point of the plot.

7. The method according to claim 1, characterized in that, The emotional resonance matching relationship is quantitatively matched by calculating the cosine similarity or Euclidean distance between the emotional vector of the advertising material and the plot-emotion time sequence vector.

8. A dynamic resource scheduling system based on the fusion of reinforcement learning and knowledge graphs, characterized in that, include: The edge data acquisition module is used to analyze the dialogue text, background music emotion tags, and bullet screen sentiment in real time through a lightweight edge natural language processing model to extract the emotional features of the plot. And collect users' micro-behavioral data; The time-series vector construction module is used to construct a "plot-emotion" time-series vector based on the plot emotion features and the micro-behavioral data according to the time sequence. The user emotion perception module is used to obtain sensor data from the device side through the quick app and infer the user's real-time emotional state. The dynamic learning matching module is used to input the "plot-emotion" time-series vector and the user's real-time emotional state into the dynamic learning model to generate an advertising matching strategy. The ad display execution module is used to select ad creatives from the ad library and display ads according to the timing of emotional intervention, based on the ad matching strategy.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-7.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method as described in any one of claims 1-7.