A video prompt word completion method, device, equipment and medium

CN122840044APending Publication Date: 2026-09-29MALANSHAN AUDIO & VIDEO LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611007834.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

传统方法通常采用一键生成整段提示词的交互方式,偏离了创作者边想边写、逐步完善的自然工作流,用户难以在生成过程中进行细粒度干预,生成后的修改成本较高

Benefits of technology

[0008]本发明的有益效果在于,本发明提供的上述视频提示词补全方法,首先获取用户输入的提示词前缀上下文并生成多个补全候选及其初始概率,为用户提供了可选择的续写方向,符合边想边写的自然交互习惯;然后采集用户对补全候选的交互反馈信号并生成对应的奖励值,将用户隐式反馈行为转化为可计算的奖励信号,充分利用了用户自然交互过程中产生的反馈信息;之后将奖励值分别输入至实时路径和定期路径,实时路径将奖励值写入反馈记忆库以生成用于下一次候选生成时调整初始概率的实时调整分数,可以使得用户每一次采纳或拒绝都能立即影响后续推荐,实现反馈的即时生效;定期路径基于累积的奖励值定期执行强化学习以更新策略网络,实现了深层模式的批量学习;最后将定期路径中策略网络对补全候选预测的概率分布与实时调整分数进行融合得到最终概率并输出补全结果,这样兼顾了策略网络的稳定性和实时反馈的敏捷性,既保证了策略的稳定性,又实现了反馈的实时生效,显著提升了视频提示词补全的准确性和用户体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840044A_ABST
    Figure CN122840044A_ABST
Patent Text Reader

Abstract

The application discloses a video prompt word completion method and device, equipment and medium, and relates to the technical field of video processing, which comprises the following steps: acquiring an input prompt word prefix context, generating a plurality of completion candidates and determining the initial probability of each completion candidate; collecting the interactive feedback signal corresponding to the completion candidate and generating the corresponding reward value; inputting the reward value into the real-time path and the regular path respectively, writing the reward value into the feedback memory bank by the real-time path to generate a real-time adjustment score for adjusting the initial probability when generating the next candidate, and performing reinforcement learning based on the accumulated reward value by the regular path to update the policy network; and fusing the probability distribution predicted by the policy network in the regular path with the real-time adjustment score to obtain the final probability of each completion candidate and output the completion result. In this way, the real-time feedback and immediate response are combined with the regular strategy update, the common knowledge is refined and the personalized preference is adapted, and the accuracy of the video prompt word completion and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, and in particular to a method, apparatus, device, and medium for video prompt word completion. Background Technology

[0002] With the rapid development of Artificial Intelligence Generated Content (AIGC) technology, video prompt tools based on large language models have been initially applied. These tools primarily rely on contextual learning or supervised fine-tuning of large models to directly generate complete video prompts based on user input, or they train completion models based on static general corpora. Traditional methods typically employ an interactive approach of generating entire prompts with a single click, deviating from the natural workflow of creators writing and refining their ideas step-by-step. Users find it difficult to intervene with fine-grained details during the generation process, and post-generation modifications are costly. Furthermore, traditional methods, based on static general models, cannot adapt to the different expression habits of users and lack effective adaptability, resulting in a low degree of matching between completion suggestions and actual user needs. Summary of the Invention

[0003] The purpose of this invention is to provide a video prompt word completion method, apparatus, device, and medium that can simultaneously extract common knowledge and adapt to personalized preferences, and ensure that user feedback can take effect in real time, thereby improving the accuracy of video prompt word completion and user experience.

[0004] To address the aforementioned technical problems, this invention provides a video prompt word completion method, comprising: Obtain the context of the input prompt word prefix, generate multiple completion candidates, and determine the initial probability of each completion candidate; Collect the interactive feedback signals corresponding to the completion candidates, and generate corresponding reward values ​​based on the interactive feedback signals; The reward value is input into the real-time path and the periodic path respectively; wherein, the real-time path writes the reward value into the feedback memory to generate a real-time adjustment score for adjusting the initial probability during the next candidate generation; the periodic path performs reinforcement learning periodically based on the accumulated reward value to update the policy network; The probability distribution of the prediction of the completion candidates by the policy network in the periodic path is fused with the real-time adjustment score to obtain the final probability of each completion candidate, and the completion result is output based on the final probability.

[0005] To address the aforementioned technical problems, the present invention also provides a video prompt word completion device, comprising: The candidate generation module is used to obtain the context of the input prompt word prefix, generate multiple completion candidates, and determine the initial probability of each completion candidate; The feedback acquisition module is used to acquire the interactive feedback signals corresponding to the completion candidates and generate corresponding reward values ​​based on the interactive feedback signals. The path distribution module is used to input the reward value into the real-time path and the periodic path respectively; wherein, the real-time path writes the reward value into the feedback memory to generate a real-time adjustment score for adjusting the initial probability during the next candidate generation; the periodic path performs reinforcement learning periodically based on the accumulated reward value to update the policy network; The fusion output module is used to fuse the probability distribution of the policy network's prediction of the completion candidates in the periodic path with the real-time adjustment score to obtain the final probability of each completion candidate, and output the completion result based on the final probability.

[0006] To address the aforementioned technical problems, the present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the video prompt word completion method described above.

[0007] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the aforementioned video prompt word completion method.

[0008] The beneficial effects of this invention are as follows: The video prompt word completion method provided by this invention first obtains the context of the prompt word prefix input by the user and generates multiple completion candidates and their initial probabilities, providing the user with selectable directions for continuation, which conforms to the natural interaction habit of thinking and writing simultaneously; then, it collects the user's interactive feedback signals on the completion candidates and generates corresponding reward values, transforming the user's implicit feedback behavior into calculable reward signals, making full use of the feedback information generated during the user's natural interaction; then, the reward values ​​are input into the real-time path and the periodic path respectively. The real-time path writes the reward values ​​into the feedback memory to generate a real-time adjustment score for adjusting the initial probability when generating the next candidate, so that each adoption or rejection by the user can immediately affect subsequent recommendations, achieving immediate feedback effectiveness; the periodic path periodically performs reinforcement learning based on the accumulated reward values ​​to update the policy network, realizing batch learning of deep modes; finally, the probability distribution of the prediction of the completion candidates by the policy network in the periodic path is fused with the real-time adjustment score to obtain the final probability and output the completion result. This balances the stability of the policy network and the agility of real-time feedback, ensuring both the stability of the policy and the real-time effectiveness of the feedback, significantly improving the accuracy of video prompt word completion and user experience.

[0009] In addition, the present invention also provides a corresponding video prompt word completion device, electronic device and computer-readable storage medium for the video prompt word completion method, which has the same or corresponding technical features as the video prompt word completion method mentioned above, and has the same effect. Attached Figure Description

[0010] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart of the video prompt word completion method provided in this embodiment of the invention; Figure 2 This is a schematic diagram of the video prompt word completion device provided in an embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0013] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0014] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0015] The specific application environment architecture or specific hardware architecture on which the video prompt word completion method depends is described here.

[0016] The embodiments of the present invention provide a video prompt word completion method, and the method is described in detail in conjunction with the execution flow of the video prompt word completion method. Figure 1The flowchart of the video prompt word completion method provided in the embodiments of the present invention is as follows: Figure 1 As shown, the method includes: S101. Obtain the context of the input prompt word prefix, generate multiple completion candidates, and determine the initial probability of each completion candidate.

[0017] It's important to note that video cue word completion refers to a task where, during the process of generating cue words from a video input by an interactive object (such as a user), the system predicts and recommends the next word, phrase, or semantic clause in real time based on the currently input context. This task differs from cue word generation, which generates an entire cue word from scratch. Completion emphasizes partial continuation based on existing input, similar to autocomplete in a code editor, but geared towards the professional semantic system of video generation. Video cue words contain multi-dimensional information such as camera language, scene description, and style parameters. The completion task requires precise semantic-level recommendations across these dimensions. In this context, the cue word prefix context refers to the video cue word segment currently input by the interactive object, such as a cat or slow-motion text. Completion candidates refer to possible continuations predicted based on the current context, which can be a single word, phrase, or semantic clause, such as "a cat is running on the lawn." The initial probability of a completion candidate refers to the predicted probability of each candidate by the base model under the current strategy, reflecting the model's confidence in each completion candidate without considering any user feedback.

[0018] In step S101, the present invention can receive the context of the prompt word prefix input by the user through the video creation platform, call the pre-trained base model to perform semantic understanding of the context, and generate multiple possible completion candidates through the decoder. Each candidate is assigned an initial probability value, which reflects the model's preliminary judgment on the rationality of the candidate content.

[0019] S102. Collect the interactive feedback signals corresponding to the candidates and generate the corresponding reward values ​​based on the interactive feedback signals.

[0020] The aforementioned interactive feedback signals refer to the behavioral responses of the interactive object after seeing the displayed completion candidates, including but not limited to clicking "accept," pressing the Tab key to confirm, continuing to type other content, ignoring the candidate, or remaining inactive for an extended period. The reward value is the result of quantifying the above interactive behaviors into numerical values, used to represent the interactive object's degree of preference for a specific completion candidate; a positive value indicates a positive preference, and a negative value indicates a negative preference.

[0021] In step S102, the present invention can display the candidate set generated in step S101 to the user and monitor the user's interaction behavior in real time. When the user adopts, rejects, ignores, or partially modifies the behavior, the behavior is mapped to a specific reward value according to a preset reward rule. This step transforms the user's implicit natural interaction behavior into a calculable and optimizable reward signal, providing a data foundation for subsequent reinforcement learning and making full use of the large amount of feedback information generated during user interaction.

[0022] S103. Input the reward value into the real-time path and the periodic path respectively; wherein, the real-time path writes the reward value into the feedback memory to generate a real-time adjustment score for adjusting the initial probability when the next candidate is generated; the periodic path performs reinforcement learning periodically to update the policy network based on the accumulated reward value.

[0023] In this invention, a real-time path refers to a path that processes each interaction feedback from an interactive object immediately, utilizing the feedback information in the next inference without waiting for parameter updates. A periodic path refers to a path that processes accumulated interaction data in batches at a preset time period (e.g., daily or weekly) and learns the preferences of the interactive object through parameter updates. A feedback memory is a database structure used to store interaction feedback records. The feedback memory can include two levels: a personal memory and a global memory. Each feedback record contains at least contextual information, candidate information, and a corresponding reward value. The real-time adjustment score is a score calculated based on historical records in the feedback memory, used to adjust the candidate probability during the current inference. The accumulated reward value refers to the set of all reward values ​​collected within a preset time window. The policy network refers to a neural network model used to predict and complete the candidate probability distribution, and in this invention, it includes two levels: a global policy network and a personalized policy network.

[0024] When executing step S103, the present invention can simultaneously distribute the reward value generated in step S102 to two independent processing paths. In the real-time path, the reward value is immediately written into the feedback memory, so that the feedback information can be retrieved and utilized in the user's next inference request, thereby achieving an immediate impact on subsequent completion recommendations. In the periodic path, the reward value is accumulated in the training data pool, and when a preset periodic condition is met (such as the accumulated interaction data volume reaching a preset threshold or the time interval since the last update reaching a preset time interval), the parameter update of the policy network is triggered. The two paths work together, with the real-time path ensuring the immediate response of feedback and the periodic path ensuring the stable learning of deep patterns. The time interval threshold in the preset periodic condition can be preset according to the needs of the actual application scenario, such as being set to 1 day, 3 days, or 7 days, or dynamically adjusted according to the user's interaction frequency; the data volume threshold can be preset according to the model training stability requirements, such as being set to 100, 500, or 1000 data points.

[0025] S104. The probability distribution of the prediction of the completion candidates by the policy network in the periodic path is fused with the real-time adjustment score to obtain the final probability of each completion candidate, and the completion result is output based on the final probability.

[0026] The final probability of the aforementioned completion candidates refers to the final predicted probability obtained by recalculating each completion candidate after comprehensively considering the prediction results of the policy network in the periodic path and the historical feedback information of the interactive objects in the real-time path. This probability is used for final sorting and display.

[0027] In step S104, this invention can obtain the probability distribution of the policy network's predictions for each completion candidate in the periodic path, and simultaneously obtain the real-time adjustment score calculated based on the feedback memory in the real-time path. These two are then fused. The candidates are then sorted according to their final probabilities from highest to lowest, and one or more candidates with the highest probabilities are output as the final completion result to the user. By combining the stable prediction capability of the policy network with the agile adjustment capability of real-time feedback, the completion result is both professional and reasonable, and can respond in real-time to changes in the user's immediate preferences.

[0028] In the video prompt word completion method provided by this invention, the first step is to obtain the context of the prompt word prefix input by the user and generate multiple completion candidates and their initial probabilities, providing the user with selectable directions for continuation, which conforms to the natural interaction habit of thinking and writing simultaneously. Then, the method collects the user's interactive feedback signals on the completion candidates and generates corresponding reward values, transforming the user's implicit feedback behavior into calculable reward signals, fully utilizing the feedback information generated during the user's natural interaction. Next, the reward values ​​are input into the real-time path and the periodic path respectively. The real-time path writes the reward values ​​into a feedback memory to generate a real-time adjustment score for adjusting the initial probability during the next candidate generation, ensuring that each acceptance or rejection by the user immediately affects subsequent recommendations, achieving immediate feedback effectiveness. The periodic path periodically performs reinforcement learning based on the accumulated reward values ​​to update the policy network, achieving batch learning of deep modes. Finally, the probability distribution of the policy network's prediction of the completion candidates in the periodic path is fused with the real-time adjustment score to obtain the final probability and output the completion result. This approach balances the stability of the policy network and the agility of real-time feedback, ensuring both policy stability and real-time feedback effectiveness, significantly improving the accuracy of video prompt word completion and user experience.

[0029] It should be noted that the video prompt completion method of this invention can be applied to various practical scenarios. In professional video creation platforms (such as AI video generation tools and film post-production software), creators need to write complex video generation prompts. This invention can provide professional completion suggestions in real time during the creator's input process, and each acceptance or rejection by the interactive object immediately affects the next completion recommendation, achieving a personalized experience that becomes increasingly intuitive with use. In short video platforms and AIGC content creation tools, many non-professional interactive objects lack experience in prompt writing. The global strategy network of this invention can extract high-quality prompt paradigms from all interactive objects, providing completion suggestions for new interactive objects and effectively lowering the creation threshold. In education, training, and teaching assistance scenarios, beginners are unfamiliar with the professional vocabulary of video prompts. The global strategy network can learn professional expressions from interactive objects and guide beginners to gradually learn and master professional vocabulary through completion recommendations.

[0030] Furthermore, in a specific implementation, in the video prompt word completion method provided in the embodiments of the present invention, step S101, which generates multiple completion candidates and determines the initial probability of each completion candidate, may specifically include: segmenting and encoding the context of the prompt word prefix to generate an input sequence and a context semantic embedding vector corresponding to the prompt word prefix context; inputting the input sequence into a pre-trained base model, generating a hidden representation of the context through the encoder of the base model, and generating a class probability distribution for the next time step based on the hidden representation through the decoder of the base model; selecting several completion candidates from the class probability distribution using a preset sampling strategy to form a candidate set, and extracting the candidate semantic embedding vector corresponding to each completion candidate; and determining the initial probability of each completion candidate under the current strategy according to the class probability distribution.

[0031] In implementation, the base model is based on the Transformer architecture and pre-trained on a large-scale video prompt corpus, possessing the ability to understand the professional semantics of video generation. Its vocabulary includes specialized units for professional terms such as camera language, style parameters, and scene descriptions. This invention first segments the context of the prompt words input by the user, converting each unit into a corresponding word embedding vector and adding positional encoding information to form an input sequence. The input sequence is processed by the encoder layer of the base model, capturing semantic associations in the context through a multi-head self-attention mechanism to generate a hidden representation vector of the context, while simultaneously extracting the semantic embedding vector e_C of the context for use by subsequent modules. The decoder generates the category probability distribution for the next time step based on the hidden representation of the context, and selects G completion candidates from this category probability distribution using a preset sampling strategy (e.g., Top-K sampling or Beam Search strategy), forming a candidate set Y={y_1,y_2,...,y_G}, where G is the group size, typically 4 to 8. Simultaneously, the semantic embedding vector e_yi of each completion candidate is extracted. For each candidate y_i, its initial probability under the current strategy is calculated: P_θ(y_i|C)=softmax(logits(y_i)); Where P represents the probability; θ represents the model parameters, indicating the parameter set of the current policy network; P_θ represents the probability under the current policy; y_i represents the i-th completion candidate; C represents the context prefix of the prompt word input by the interactive object; y_i|C represents the selection of y_i under the context C; softmax is the normalized exponential function; logits is the unnormalized prediction score; logits(y_i) is the logits value of candidate y_i, representing the original prediction score output by the model decoder for candidate y_i. After the candidate set is sorted according to the initial probability, deduplication can be performed, and the candidate can be displayed to the user in the form of gray text or a drop-down list. This allows multiple candidates and their semantic embedding vectors to be generated through a pre-trained base model, providing users with diverse choices and providing a data foundation for subsequent feedback collection and reinforcement learning. At the same time, the extracted semantic embedding vectors support the calculation of similarity for real-time score adjustment.

[0032] Furthermore, in a specific implementation, in the video prompt word completion method provided in the embodiments of the present invention, step S102 collects the interaction feedback signal corresponding to the completion candidate and generates a corresponding reward value based on the interaction feedback signal. Specifically, this may include: monitoring the interaction behavior corresponding to the completion candidate; if the interaction behavior is to adopt the completion candidate, a positive reward value is assigned; if the interaction behavior is to reject or ignore the completion candidate, a negative reward value is assigned; if the interaction behavior is to adopt the completion candidate and then delete or modify it, a negative reward value with an absolute value less than the absolute value of the negative reward value corresponding to rejection or ignoring is assigned.

[0033] In implementation, this invention can display a candidate set to the user and then monitor the user's interaction within a preset behavior window. Specifically, if the user presses the Tab key or clicks a candidate, it indicates acceptance of the candidate, assigning a reward value R=+1.0. This is the strongest positive feedback signal, indicating that the candidate highly matches the user's intent. If the user continues to type other content or remains inactive for an extended period (exceeding a preset timeout threshold T_ignore, which can be set based on experience in the application scenario, typically 3 seconds), it indicates rejection or ignoring of the candidate, assigning a reward value R=-0.5, indicating that the candidate does not match the user's intent. If the user accepts a candidate but immediately deletes or modifies the candidate content within a preset time window T_partial (which can be set based on experience in the application scenario, typically 2 seconds), it indicates partial acceptance, assigning a reward value R=-0.2, providing more granular negative feedback than complete rejection. By assigning differentiated rewards to these three interaction behaviors, the user's preference for different candidates can be captured more accurately, providing richer training signals for reinforcement learning.

[0034] Furthermore, in a specific implementation, in the video prompt word completion method provided in the embodiments of the present invention, step S103, where the real-time path writes the reward value into the feedback memory bank to generate a real-time adjustment score for adjusting the initial probability during the next candidate generation, may specifically include: constructing and maintaining a personal memory bank and a global memory bank; the personal memory bank is used to store the feedback records of the current interactive object, and the global memory bank is used to store the feedback records of all interactive objects; when the next candidate generation is triggered, a preset number of feedback records with semantic similarity to the current prompt word prefix context are retrieved from the personal memory bank and the global memory bank, respectively; based on the semantic similarity between the candidate semantic embedding vector in the retrieved preset number of feedback records and the candidate semantic embedding vector of the current completion candidate, and the reward value corresponding to the feedback record, a personalized real-time adjustment score and a collective real-time adjustment score are calculated, respectively; the personalized real-time adjustment score and the collective real-time adjustment score are weighted and fused to obtain the real-time adjustment score.

[0035] In implementation, this invention can maintain a personal memory bank (which can be called a personalized feedback memory bank) M_u for each interactive object, storing the N most recent feedback records of that interactive object (typically N=100), with each record being a triple: M_u={(e_C^j,e_y^j,r^j)}; Where j = 1, 2, ..., N, j represents the j-th similar feedback record in the search results; e_C^j is the semantic embedding vector of the context (output by the base model encoder), e_y^j is the semantic embedding vector of the candidate completion (output by the base model decoder), and r^j is the reward value. Simultaneously, a global feedback memory (M_global) is maintained to store feedback records within the most recent time window of all interactive objects.

[0036] When an interactive object generates new feedback, the new record is written to the memory bank in real time. If the memory bank is full, the oldest record is discarded (first-in, first-out strategy). When the interactive object triggers the next inference in context C, the base model generates G candidates and their embedding vectors. For each candidate y_i, the current context embedding e_C and the candidate embedding e_yi are calculated. Then, the K records most semantically similar to the current context are retrieved from M_u (based on the cosine similarity of e_C), and the personalized real-time adjustment score is calculated. S_rt^u(y_i)=Σ_{j∈TopK_u}sim(e_yi,e_y^j)·r^j / Σ_{j∈TopK_u}|sim(e_yi,e_y^j)|; Where S_rt^u(y_i) is the personalized real-time adjustment score for the current candidate y_i; Σ_{j∈TopK_u} represents the summation of the K most similar records retrieved from the personal memory bank; sim() is the cosine similarity function; sim(e_yi,e_y^j) represents the cosine similarity between the semantic embedding of the current candidate and the semantic embedding of the candidate in the j-th historical record; Σ_{j∈TopK_u}|sim(e_yi, e_y^j)| represents the summation of the absolute values ​​of the similarities, used as the normalization denominator. The meaning of this formula is: if a historical candidate semantically similar to the current candidate has been adopted by the interacting object (r^j is positive), then the real-time score of the current candidate is positive and should be increased; otherwise, it should be decreased.

[0037] Similarly, the collective real-time adjusted score S_rt^global(y_i) is calculated from M_global. Then, the personalized real-time adjusted score is weighted and merged with the collective real-time adjusted score: S_rt(y_i)=α(u)·S_rt^global(y_i)+(1-α(u))·S_rt^u(y_i); Wherein, α(u) is the fusion weight, with new interaction objects having α(u) ≈ 1, relying more on collective real-time feedback; old interaction objects have smaller α(u), relying more on individual real-time feedback; S_rt(y_i) is the real-time adjustment score of the current candidate y_i (the final real-time adjustment score after fusion); S_rt^global(y_i) is the collective real-time adjustment score calculated based on the global memory; 1-α(u) represents the fusion weight of the personalized part; S_rt^u(y_i) is the personalized real-time adjustment score calculated based on the individual memory. Through the above retrieval and calculation, even within the interval of periodic path parameter updates, each adoption or rejection of an interaction object can immediately affect the next completion recommendation through the real-time path, achieving immediate effect of feedback.

[0038] Furthermore, in a specific implementation, in the video prompt word completion method provided in the embodiments of the present invention, step S103, the periodic path, is based on the accumulated reward value and periodically performs reinforcement learning to update the policy network. Specifically, it may include: when the preset periodic conditions are met, obtaining the interaction data within the target time window from the training data pool and grouping and aggregating it according to the prompt word prefix context; for multiple completion candidates corresponding to the same prompt word prefix context, calculating the average and standard deviation of the reward value within the group, normalizing the reward value of each completion candidate to obtain the relative advantage within the group of each completion candidate; updating the parameters of the global policy network using the policy gradient algorithm based on the relative advantage within the group; mounting a low-rank adaptive network as a personalized policy network on the base model, and updating the parameters of the low-rank adaptive network based on the relative advantage within the group.

[0039] In implementation, the periodic path can employ the Group Relative Policy Optimization (GRPO) algorithm. This algorithm eliminates the need for additional training of the value network (Critic network), directly leveraging the multi-candidate concurrency inherent in the completion task to calculate the relative advantage within a group. The global policy network is updated via GRPO. Specifically, when preset periodic conditions are met (e.g., reaching a preset time interval threshold or cumulative data volume threshold), interaction data of all interacting objects within the target time window are retrieved from the training data pool and aggregated by context C. For feedback from multiple interacting objects under the same context C, the reward signals are merged using majority voting to eliminate noise in the feedback from individual interacting objects. For the G candidates corresponding to the same context C, the mean(R) and standard deviation(R) of the group reward are calculated, and the reward of each candidate is Z-score normalized to obtain the relative advantage within the group. A_i=(r_i-mean(R)) / (std(R)+ε); Where ε is a small constant to prevent division by zero errors; A_i is the relative advantage within the group of the i-th candidate for completion; mean(R) is the average reward value of all candidates within the group; std(R) is the standard deviation of the reward values ​​of all candidates within the group; ε is a constant used to prevent the denominator from being zero. A candidate advantage that is widely adopted is positive, and a candidate advantage that is generally rejected is negative.

[0040] Next, the global policy network π_global is updated using the policy gradient algorithm. The probability ratio is first calculated using the following formula: ρ_i(θ)=π_θ(y_i|C) / π_ref(y_i|C); Where ρ_i(θ) is the probability ratio of the i-th candidate, which is a function of the policy parameters; π_θ(y_i|C) is the probability that the current policy network selects candidate y_i in context C; and π_ref(y_i|C) is the probability that the reference policy network (the policy before the update) selects candidate y_i in context C.

[0041] Then, the cutting target is calculated using the following formula: L_clip=-E[min(ρ_i(θ)·A_i,clip(ρ_i(θ),1-ε_clip,1+ε_clip)·A_i)]; Where L_clip is the pruning loss function; E[·] represents the expected value (averaged over all training samples); ρ_i(θ) is the probability ratio of the i-th candidate under the current policy; A_i is the relative advantage of the i-th candidate within the group; clip(ρ_i(θ),1-ε_clip, 1+ε_clip) means pruning the probability ratio to the interval [1-ε_clip,1+ε_clip]; ρ_i(θ)·A_i represents the policy gradient term before pruning; clip(ρ_i(θ),1-ε_clip,1+ε_clip)·A_i) represents the policy gradient term after pruning; ε_clip is the pruning parameter, which controls the range of policy update magnitude (typical value 0.2).

[0042] Simultaneously incorporating KL divergence constraints, the final loss function is: L_global=L_clip+β·KL(π_θ||π_ref); Where L_global is the final loss function of the global policy network; L_clip is the pruning loss (policy gradient objective); β is the KL constraint coefficient (typically 0.04); KL(π_θ || π_ref) represents the KL divergence between the current policy π_θ and the reference policy π_ref. The collective module uses a large learning rate (typically 5e-5) and a large batch size (typically 64) to ensure stable updates of the global policy.

[0043] Meanwhile, this invention mounts a low-rank adaptive network (LoRA) as a personalized policy network on the base model. The formula for LoRA is W'=W+ΔW=W+B·A, where W∈R^{d×d} is the original weight matrix (frozen); W'∈R^{d×d} is the updated weight matrix; ΔW is the weight change; A∈R^{d×r} and B∈R^{r×d} are both low-rank decomposition matrices, and r is the LoRA rank (typically 4 to 16). During initialization, A is randomly initialized, and B is a zero matrix, ensuring that ΔW=0 in the initial state, and that the personalized policy is consistent with the global policy.

[0044] The personalized strategy network calculates the relative advantage within the group using only the data of the current interaction object u: A_i^u=(r_i^u-mean(R^u)) / (std(R^u)+ε); Where A_i^u is the relative advantage of the i-th completion candidate of the interaction object u within the group; r_i^u is the reward value of the interaction object u for the i-th completion candidate; mean(R^u) is the average of all reward values ​​of the interaction object u within the group; and std(R^u) is the standard deviation of all reward values ​​of the interaction object u within the group.

[0045] Then, the personalized policy is updated using A_i^u, with the loss function being: L_user=L_clip^u+β_user·KL(π_user||π_ref); Where L_user is the final loss function of the personalized policy network; L_clip^u is the pruning loss (policy gradient objective) corresponding to the interaction object u; β_user is the personalized KL divergence constraint coefficient, which controls the regularization strength (typical value 0.1); KL(π_user||π_ref) is the KL divergence between the personalized policy π_user and the reference policy π_ref.

[0046] A small learning rate (typically 1e-5) and small batch size (typically 8) are used to prevent overfitting. The KL constraint coefficient β_user is set to a relatively large value (typically 0.1) to ensure that the personalized strategy does not deviate too far from the global strategy. For cold starts, the LoRA parameter of new interactive objects is initialized to zero, and the personalized strategy is completely consistent with the global strategy. Through the above two-layer structure of the global strategy network and the personalized strategy network, common knowledge can be extracted from all interactive objects, while personalized preferences are adapted for each interactive object.

[0047] In the above steps, the low-rank adaptive network has a lifecycle management mechanism, which includes: adopting a lazy loading strategy, creating the parameters of the low-rank adaptive network for the current interactive object only when the first interaction feedback is generated; classifying interactive objects into different levels according to the interval between the interaction time and the current time, and configuring different parameter storage locations and update strategies for interactive objects of different levels; adopting a sliding window strategy when updating the parameters of the low-rank adaptive network, using only a preset number of interaction data for each update; and deleting the parameters of the low-rank adaptive network when no interaction is generated within a preset time threshold.

[0048] In implementation, due to the low-rank nature of LoRA (r is much smaller than d), the number of additional parameters for each interactive object is only a tiny fraction of the base model, typically not exceeding one percent. Therefore, even with a user base of millions, storage overhead remains manageable. To efficiently manage the personalized policy network for large-scale interactive objects, this invention designs the following lifecycle management mechanism: First, a lazy loading strategy is adopted, where low-rank adaptive network parameters θ_u are created only when an interactive object generates its first interaction feedback. During initialization, A_u is initialized using a random normal distribution, and B_u is initialized as a zero matrix, ensuring that ΔW = B_u·A_u = 0. That is, the newly created personalized policy network is completely consistent with the global policy network, and objects that have never interacted do not occupy any storage resources. The second approach categorizes interactive objects into three levels based on the interval between their interactions and the current moment, as well as the amount of interaction data: Active objects (interacting within the last N_active days, typically N_active = 7 days), whose low-rank adaptive network parameters reside in memory and participate in training during each periodic reinforcement learning update; Normal objects (interacting within the last N_normal days, typically N_normal = 30 days), whose low-rank adaptive network parameters are stored on disk and loaded into memory only during periodic reinforcement learning updates, then written back to disk after training; and Silent objects (not interacting for more than N_normal days), whose low-rank adaptive network parameters do not participate in periodic reinforcement learning updates but are retained on disk. When a silent object becomes active again, its low-rank adaptive network parameters are reloaded and included in the periodic update process. The third approach employs a sliding window strategy when updating low-rank adaptive network parameters. For active objects, each update uses only the latest preset number of interaction data points for parameter updates (typically W = 200 times), allowing the model to focus more on recent preference changes and capture preference drift. Fourthly, when an interactive object does not interact within a preset time threshold (typically 90 days), its low-rank adaptive network parameters are deleted from disk to free up storage resources. If the object subsequently becomes active again, the low-rank adaptive network parameters are recreated (initialized to zero, equivalent to the global policy network), essentially a cold start, because the preferences of long-inactive objects may have changed, and the original low-rank adaptive network parameters may no longer be applicable. Through the above lifecycle management mechanism, efficient storage and updating of personalized policy networks for millions of interactive objects can be achieved.

[0049] Furthermore, in a specific implementation, in the video prompt word completion method provided in the embodiments of the present invention, step S104 fuses the probability distribution of the prediction of completion candidates by the policy network in the periodic path with the real-time adjustment score to obtain the final probability of each completion candidate. Specifically, this may include: obtaining the probability distribution predicted by the policy network in the periodic path; the probability distribution includes the probability distribution predicted by the global policy network and the probability distribution predicted by the personalized policy network; determining the dynamic fusion weight according to the interaction history, and weighting and fusing the probability distribution predicted by the global policy network and the probability distribution predicted by the personalized policy network to obtain the probability distribution output by the periodic path; taking the logarithm of the probability distribution output by the periodic path and superimposing it with the real-time adjustment score to obtain the final logical value of each completion candidate; and normalizing the final logical value of each completion candidate to obtain the final probability of each completion candidate.

[0050] In implementation, this invention first dynamically weights and fuses the probability distribution predicted by the global policy network with the probability distribution predicted by the personalized policy network to obtain the probability distribution of the periodic path output: P_periodic(y|C,u)=α(u)·P_global(y|C)+(1-α(u))·P_user(y|C,u); Wherein, P_periodic(y|C,u) is the probability distribution of the final output of the periodic path, that is, the predicted probability of candidate y given the context C and the interaction object u; P_global(y|C) is the probability distribution predicted by the global policy network, that is, the policy trained based on all interaction objects; α(u) is the dynamic fusion weight, which is a function of the interaction history of the interaction object u; α(u) adopts the Sigmoid function form: α(u)=σ(k·(n_0-n(u)))=1 / (1+exp(k·(n(u)-n_0))), n(u) is the number of interaction history of the interaction object, n_0 is the transition midpoint (typical value 50), and k is the transition rate (typical value 0.1); P_user(y|C,u) is the probability distribution predicted by the personalized policy network, that is, the policy trained based on the personal historical data of the current interaction object u. New interactive objects α(u)≈1, relying entirely on collective common knowledge; old interactive objects α(u)→0.5 or lower, relying more on personalized preferences.

[0051] Subsequently, the real-time adjusted scores will be incorporated into the log probability of the probability distribution of the periodic path output: logit_final(y_i)=log(P_periodic(y_i|C,u))+η·S_rt(y_i); Where η is the real-time adjustment intensity coefficient (typically 0.5 to 2.0), controlling the degree of influence of real-time feedback on the final output; logit_final(y_i) is the final logical value of candidate y_i (unnormalized final score); log(P_periodic(y_i|C,u)) is the logarithm of the probability distribution of the periodic path output, i.e., the log probability; S_rt(y_i) is the real-time adjustment score, which can be calculated based on the personal memory bank and the global memory bank. When S_rt(y_i) is positive (similar candidates have been adopted historically), the final probability of the candidate is increased; when S_rt(y_i) is negative (similar candidates have been rejected historically), the final probability of the candidate is decreased. The final probability is obtained by softmax normalization: P_final(y_i|C,u)=softmax(logit_final(y_i))_i; Where P_final(y_i|C,u) is the final probability of candidate y_i, that is, the probability that candidate y_i will be output given the context C and the interaction object u; logit_final(y_i) is the final logical value of candidate y_i.

[0052] This invention sorts candidates by P_final and outputs the completion result with the highest probability. This logarithmic probability adjustment design has the following advantages: the probability distribution of the periodic path output provides a stable baseline probability, ensuring the rationality of the completion result; real-time score adjustment makes fine adjustments on the baseline, avoiding drastic changes; when real-time feedback data is insufficient, it naturally degenerates into periodic path output, ensuring robustness.

[0053] From the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0054] Embodiments of the present invention also provide a video prompt word completion device. Figure 2 This is a schematic diagram of the video prompt word completion device provided in an embodiment of the present invention. This embodiment is based on the perspective of functional modules, such as… Figure 2 As shown, the device includes: The candidate generation module 10 is used to obtain the context of the input prompt word prefix, generate multiple completion candidates, and determine the initial probability of each completion candidate; Feedback acquisition module 11 is used to acquire the interactive feedback signal corresponding to the candidate completion and generate the corresponding reward value based on the interactive feedback signal; The path distribution module 12 is used to input reward values ​​into the real-time path and the periodic path respectively; wherein, the real-time path writes the reward value into the feedback memory to generate a real-time adjustment score for adjusting the initial probability in the next candidate generation; the periodic path performs reinforcement learning periodically based on the accumulated reward value to update the policy network. The fusion output module 13 is used to fuse the probability distribution of the policy network's prediction of the completion candidates in the periodic path with the real-time adjusted scores to obtain the final probability of each completion candidate, and output the completion result based on the final probability.

[0055] In the video prompt word completion device provided in the embodiments of the present invention, the candidate generation module can provide users with selectable continuation directions, the feedback acquisition module can convert implicit user feedback into reward signals, the path distribution module can realize real-time effectiveness of feedback and periodic batch learning, and the fusion output module can balance the stability of the policy network and the agility of real-time feedback. This ensures both the stability of the policy and the real-time effectiveness of feedback, significantly improving the accuracy of video prompt word completion and user experience.

[0056] Since the embodiments of the video prompt word completion device and the video prompt word completion method correspond to each other, the description of the features in the embodiment corresponding to the video prompt word completion device can be found in the relevant description of the embodiment corresponding to the video prompt word completion method, and will not be repeated here. Furthermore, it has the same beneficial effects as the video prompt word completion method mentioned above.

[0057] Furthermore, in a specific implementation, in the video prompt word completion device provided in the embodiments of the present invention, the candidate generation module 10 can be specifically used to segment and encode the prompt word prefix context to generate an input sequence and a context semantic embedding vector corresponding to the prompt word prefix context; input the input sequence into a pre-trained base model, generate a hidden representation of the context through the encoder of the base model, and generate a class probability distribution for the next time step based on the hidden representation through the decoder of the base model; select several completion candidates from the class probability distribution using a preset sampling strategy to form a candidate set, and extract the candidate semantic embedding vector corresponding to each completion candidate; determine the initial probability of each completion candidate under the current strategy according to the class probability distribution.

[0058] Furthermore, in a specific implementation, in the video prompt word completion device provided in the embodiments of the present invention, the feedback acquisition module 11 can be used to monitor the interactive behavior corresponding to the completion candidate; if the interactive behavior is to adopt the completion candidate, a positive reward value is assigned; if the interactive behavior is to reject or ignore the completion candidate, a negative reward value is assigned; if the interactive behavior is to adopt the completion candidate and then delete or modify the completion candidate, a negative reward value with an absolute value less than the absolute value of the negative reward value corresponding to rejection or ignoring is assigned.

[0059] Furthermore, in a specific implementation, in the video prompt word completion device provided in the embodiments of the present invention, the path distribution module 12 can be specifically used to construct and maintain a personal memory bank and a global memory bank; the personal memory bank is used to store the feedback records of the current interactive object, and the global memory bank is used to store the feedback records of all interactive objects; when the next candidate generation is triggered, a preset number of feedback records with semantic similarity to the current prompt word prefix context are retrieved from the personal memory bank and the global memory bank, respectively; based on the semantic similarity between the candidate semantic embedding vector in the retrieved preset number of feedback records and the candidate semantic embedding vector of the current completion candidate, and the reward value corresponding to the feedback record, the personalized real-time adjustment score and the reward value are calculated respectively. The system performs real-time collective score adjustments; it then weights and fuses the individualized real-time adjusted scores with the collective real-time adjusted scores to obtain the final real-time adjusted score; when a preset periodic condition is met, it retrieves interaction data within the target time window from the training data pool and aggregates it by grouping according to the context of the prompt word prefix; for multiple completion candidates corresponding to the same prompt word prefix context, it calculates the average and standard deviation of the reward values ​​within the group, normalizes the reward values ​​of each completion candidate, and obtains the relative advantage of each completion candidate within the group; based on the relative advantage within the group, it updates the parameters of the global policy network using a policy gradient algorithm; and it mounts a low-rank adaptive network as an individualized policy network on the base model, updating the parameters of the low-rank adaptive network based on the relative advantage within the group.

[0060] Furthermore, in a specific implementation, in the video prompt word completion device provided in the embodiments of the present invention, the fusion output module 13 can be specifically used to obtain the probability distribution of the policy network prediction in the periodic path; the probability distribution includes the probability distribution of the global policy network prediction and the probability distribution of the personalized policy network prediction; the dynamic fusion weight is determined according to the interaction history, and the probability distribution of the global policy network prediction and the probability distribution of the personalized policy network prediction are weighted and fused to obtain the probability distribution of the periodic path output; the logarithm of the probability distribution of the periodic path output is taken and superimposed with the real-time adjustment score to obtain the final logical value of each completion candidate; the final logical value of each completion candidate is normalized to obtain the final probability of each completion candidate.

[0061] Embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described video prompt word completion method embodiments.

[0062] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described video prompt word completion method embodiments when running.

[0063] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0064] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described video prompt word completion method embodiments.

[0065] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described video prompt word completion method embodiments.

[0066] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-chip (SoC), a complex programmable logic device (CPLD), a microcontroller unit (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.

[0067] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0068] The foregoing has provided a detailed description of a video prompt word completion method, apparatus, device, and medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only intended to help understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the present invention.

Claims

1. A method for completing video prompts, characterized in that, include: Obtain the context of the input prompt word prefix, generate multiple completion candidates, and determine the initial probability of each completion candidate; Collect the interactive feedback signals corresponding to the completion candidates, and generate corresponding reward values ​​based on the interactive feedback signals; The reward value is input into the real-time path and the periodic path respectively; wherein, the real-time path writes the reward value into the feedback memory to generate a real-time adjustment score for adjusting the initial probability during the next candidate generation; the periodic path performs reinforcement learning periodically based on the accumulated reward value to update the policy network; The probability distribution of the prediction of the completion candidates by the policy network in the periodic path is fused with the real-time adjustment score to obtain the final probability of each completion candidate, and the completion result is output based on the final probability.

2. The video prompt word completion method according to claim 1, characterized in that, Generating multiple completion candidates and determining the initial probability of each completion candidate includes: The context of the prompt word prefix is ​​segmented and encoded to generate the input sequence and the context semantic embedding vector corresponding to the context of the prompt word prefix. The input sequence is input into a pre-trained base model, the encoder of the base model generates a hidden representation of the context, and the decoder of the base model generates a class probability distribution for the next time step based on the hidden representation. A preset sampling strategy is used to select several completion candidates from the category probability distribution to form a candidate set, and the candidate semantic embedding vector corresponding to each completion candidate is extracted; The initial probability of each completion candidate under the current strategy is determined based on the category probability distribution.

3. The video prompt word completion method according to claim 1, characterized in that, Collect the interaction feedback signal corresponding to the completion candidate, and generate the corresponding reward value based on the interaction feedback signal, including: Monitor the interactive behavior corresponding to the completion candidates; If the interaction behavior is to adopt the completion candidate, a positive reward value is assigned; If the interaction behavior is to reject or ignore the completion candidate, a negative reward value is assigned; If the interaction behavior is to adopt the completion candidate and then delete or modify it, then a negative reward value with an absolute value less than the absolute value of the negative reward value corresponding to the rejection or ignoring is assigned.

4. The video prompt word completion method according to claim 1, characterized in that, The real-time path writes the reward value into a feedback memory to generate a real-time adjustment score for adjusting the initial probability during the next candidate generation, including: Construct and maintain a personal memory bank and a global memory bank; the personal memory bank is used to store the feedback records of the current interactive object, and the global memory bank is used to store the feedback records of all interactive objects; When the next candidate generation is triggered, a preset number of feedback records that are semantically similar to the current prompt word prefix are retrieved from the personal memory bank and the global memory bank, respectively. Based on the semantic similarity between the candidate semantic embedding vectors in the preset number of feedback records and the candidate semantic embedding vectors of the current completion candidate, as well as the reward value corresponding to the feedback record, the personalized real-time adjustment score and the collective real-time adjustment score are calculated respectively. The personalized real-time adjustment score is weighted and fused with the collective real-time adjustment score to obtain the real-time adjustment score.

5. The video prompt word completion method according to claim 2, characterized in that, The periodic path, based on accumulated reward values, periodically performs reinforcement learning to update the policy network, including: When the preset periodic conditions are met, the interactive data within the target time window is obtained from the training data pool and grouped and aggregated according to the context of the prompt word prefix; For multiple completion candidates corresponding to the same prompt word prefix context, calculate the average and standard deviation of the reward value within the group, normalize the reward value of each completion candidate, and obtain the relative advantage of each completion candidate within the group; Based on the relative advantages within the group, the parameters of the global policy network are updated using the policy gradient algorithm; A low-rank adaptive network is mounted on the base model as a personalized policy network, and the parameters of the low-rank adaptive network are updated based on the relative advantages within the group.

6. The video prompt word completion method according to claim 5, characterized in that, The low-rank adaptive network has a lifecycle management mechanism, which includes: A lazy loading strategy is adopted, and the parameters of the low-rank adaptive network are created only when the first interaction feedback is generated for the current interaction object; Based on the interval between the interaction time of the interactive object and the current time, the interactive objects are divided into different levels, and different parameter storage locations and update strategies are configured for interactive objects of different levels. When updating the parameters of the low-rank adaptive network, a sliding window strategy is adopted, and the parameters are updated only using a preset number of interaction data for each update; If no interaction occurs within a preset time threshold, the parameters of the low-rank adaptive network are deleted.

7. The video prompt word completion method according to claim 5, characterized in that, The probability distribution of the policy network's predictions for the completion candidates in the periodic path is fused with the real-time adjusted score to obtain the final probability of each completion candidate, including: Obtain the probability distribution of the policy network prediction in the periodic path; the probability distribution includes the probability distribution of the global policy network prediction and the probability distribution of the personalized policy network prediction; Based on the interaction history, the dynamic fusion weights are determined, and the probability distribution predicted by the global policy network and the probability distribution predicted by the personalized policy network are weighted and fused to obtain the probability distribution of the periodic path output. The logarithm of the probability distribution of the periodic path output is then superimposed with the real-time adjustment score to obtain the final logical value of each of the completion candidates. The final logical values ​​of each of the completion candidates are normalized to obtain the final probability of each completion candidate.

8. A video prompt word completion device, characterized in that, include: The candidate generation module is used to obtain the context of the input prompt word prefix, generate multiple completion candidates, and determine the initial probability of each completion candidate; The feedback acquisition module is used to acquire the interactive feedback signals corresponding to the completion candidates and generate corresponding reward values ​​based on the interactive feedback signals. The path distribution module is used to input the reward value into the real-time path and the periodic path respectively; wherein, the real-time path writes the reward value into the feedback memory to generate a real-time adjustment score for adjusting the initial probability during the next candidate generation; the periodic path performs reinforcement learning periodically based on the accumulated reward value to update the policy network; The fusion output module is used to fuse the probability distribution of the policy network's prediction of the completion candidates in the periodic path with the real-time adjustment score to obtain the final probability of each completion candidate, and output the completion result based on the final probability.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the video prompt word completion method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the video prompt word completion method as described in any one of claims 1 to 7.