Training and processing method and device based on time sequence performance gain, equipment and medium

By generating video entropy factors to adjust reward intensity and constructing ordered frame content, the problem of insufficient temporal sensitivity in video reasoning training of multimodal large language models is solved, achieving more stable and generalized video understanding and reasoning capabilities, and improving the accuracy of video reasoning in fintech and healthcare businesses.

CN120953893APending Publication Date: 2025-11-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511246082.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing multimodal large language models based on reinforcement learning lack temporally sensitive reward signals during video reasoning training. This causes the models to tend to rely on single-frame information and fail to effectively learn reasoning capabilities across time dimensions, affecting the accuracy of video reasoning in fintech and healthcare businesses.

Method used

By acquiring the motion entropy and scene entropy of video training samples, a video entropy factor is generated. The reward intensity is adjusted and the training process is organized. Ordered frame content and shuffled frame content are constructed, a reward signal is generated, and the policy parameters are updated until the training termination condition is met, forming a reward signal for temporal performance gain and optimizing the model training process.

Benefits of technology

It enhances the model's ability to understand and reason about videos across time dimensions, improves the accuracy and generalization of video reasoning in fintech and healthcare businesses, and avoids the model's dependence on single-frame information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953893A_ABST
    Figure CN120953893A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a training and processing method, device, equipment and medium based on time sequence performance gains. Constructing ordered frame and disordered frame contents to generate an answer result, comparing the answer result to obtain a time sequence performance gain, generating a reward signal in combination with reward intensity and the time sequence performance gain, inputting a strategy optimization process to update strategy parameters, processing a training process to generate training output, and updating a training model until a target model is obtained. And completing the task by using the target model. According to the method, complex samples are highlighted through video entropy factors, and cross-frame reasoning is enhanced in combination with time sequence performance gain, so that reward signals simultaneously reflect sample difficulty and time sequence differences, a model is prevented from depending on a single frame, and the stability and generalization ability of video reasoning are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a training and processing method, apparatus, device, and storage medium based on time-series performance gain. Background Technology

[0002] In recent years, the performance of Large Language Models (LLMs) in reasoning tasks has gradually improved. Researchers have begun to explore incorporating reinforcement learning (RL) into the training of multimodal large language models (MLLMs) to enhance their understanding and reasoning abilities in joint image and text tasks. However, existing techniques still have significant shortcomings when extended to video reasoning scenarios. Traditional reinforcement learning methods such as GRPO (Group Relative Policy Optimization) lack explicit modeling and reward mechanisms for temporal information when processing video data. This leads to models relying heavily on single-frame information during training, neglecting reasoning across time dimensions. This limitation often prevents models from developing stable and accurate temporal reasoning capabilities when facing complex and dynamic scenes.

[0003] In the fintech sector, video inference technology has been increasingly applied to remote identity authentication, transaction behavior monitoring, and intelligent risk control. However, existing models, when faced with video data containing continuous actions, often rely solely on a single keyframe for judgment, neglecting the complete temporal sequence of transaction behavior. For example, in video authentication scenarios for remote account opening, if the model fails to fully capture the temporal characteristics of user actions, it may lead to misjudgments in identity verification or missed detections of fraudulent activities. This exposes the lack of an effective reward-driven mechanism in traditional video inference technology when processing complex temporal data, making it unable to strengthen the learning of long-term logical consistency during training.

[0004] In the healthcare field, video reasoning is also used in scenarios such as remote diagnosis, rehabilitation assessment, and surgical assistance. However, existing technologies often demonstrate insufficient understanding of temporal relationships when processing videos containing dynamic sequences of patient movements, rehabilitation training processes, or medical images. For example, the continuous movement sequences of patients in rehabilitation training contain rich information about the disease progression, but existing models often focus on the identification of certain local frames, neglecting the cross-frame movement evolution trends. This may lead to inaccurate rehabilitation assessment results, or even the omission of important dynamic pathological features in assisted diagnosis, thereby affecting the reliability of medical decisions. Summary of the Invention

[0005] The main objective of this invention is to provide a training and processing method, apparatus, device, and storage medium based on temporal performance gain, aiming to solve the technical problem that existing multimodal large language models based on reinforcement learning lack temporally sensitive reward signals in video reasoning training, causing the models to tend to rely on single-frame information and fail to effectively learn reasoning capabilities across time dimensions.

[0006] To achieve the above objectives, the present invention provides a training and processing method based on time-series performance gain, comprising:

[0007] Obtain video training samples, and generate a video entropy factor based on the motion entropy and scene entropy of the video training samples;

[0008] The reward intensity is adjusted and the training process is organized based on the video entropy factor;

[0009] Based on the training process, ordered frame content and shuffled frame content are constructed, and the answer result is generated based on the ordered frame content and the shuffled frame content through the training model.

[0010] The timing performance gain is determined based on the comparison of the response results.

[0011] A reward signal is generated based on the adjusted reward intensity and the time-series performance gain;

[0012] The reward signal is input into the policy optimization process to update the policy parameters, and the training process is processed based on the updated policy parameters to generate training output;

[0013] Update the training model based on the training output;

[0014] Repeat the reward intensity adjustment step, timing test step, gain determination step, signal generation step, policy update step, process processing step, and model update step until the training termination condition is met and the target model is obtained.

[0015] The target task is processed through the target model to generate the target task result.

[0016] Furthermore, to achieve the above objectives, the present invention provides a training and processing apparatus based on time-series performance gain, comprising:

[0017] The entropy factor generation module is used to acquire video training samples and generate video entropy factors based on the motion entropy and scene entropy of the video training samples.

[0018] The reward control module is used to adjust the reward intensity and organize the training process based on the video entropy factor.

[0019] The input construction and reasoning module is used to construct ordered frame content and shuffled frame content based on the training process, and to generate answer results based on the ordered frame content and shuffled frame content through the training model;

[0020] The timing gain evaluation module is used to determine the timing performance gain based on the comparison of the response results;

[0021] A reward signal generation module is used to generate a reward signal based on the adjusted reward intensity and the timing performance gain.

[0022] The strategy optimization module is used to input the reward signal into the strategy optimization process to update the strategy parameters, and to process the training process based on the updated strategy parameters and generate training output.

[0023] The model update module is used to update the training model based on the training output;

[0024] The loop control module is used to repeatedly execute the reward intensity adjustment step, the timing test step, the gain determination step, the signal generation step, the policy update step, the process processing step, and the model update step until the training termination condition is met and the target model is obtained.

[0025] The task execution module is used to process the target task through the target model and generate the target task result.

[0026] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a timing performance gain-based training and processing program stored in the memory and executable on the processor, wherein the timing performance gain-based training and processing program, when executed by the processor, implements the steps of the timing performance gain-based training and processing method as described above.

[0027] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a training and processing program based on timing performance gain, wherein when the training and processing program based on timing performance gain is executed by a processor, it implements the steps of the training and processing method based on timing performance gain as described above.

[0028] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a training and processing method, apparatus, device, and medium based on temporal performance gain, comprising: acquiring video training samples and generating a video entropy factor; adjusting the reward intensity and organizing the training process based on the video entropy factor; constructing ordered and shuffled frame content and generating response results through a training model; determining the temporal performance gain based on a comparison of the response results; generating a reward signal based on the adjusted reward intensity and temporal performance gain; inputting the reward signal into a policy optimization process to update policy parameters; processing the training process based on the updated policy parameters to generate training output; updating the training model based on the training output; repeating the above operations until the training termination condition is met to obtain the target model; and processing the target task through the target model to generate the target task result. This invention, by introducing a video entropy factor and temporal performance gain in the generation of the reward signal, enables the training process to dynamically adjust the reward intensity and explicitly guide the model to focus on inference capabilities across time dimensions. The video entropy factor ensures that video segments with high motion complexity and scene changes receive higher training weights, while the temporal performance gain supplements the time sensitivity of cross-frame inference. The combination of the two ensures that the generated reward signal considers both sample uncertainty and temporal rationality. Through this mechanism, the model avoids relying on single-frame information during training, gradually improving its ability to understand and reason about videos across time periods, thereby achieving a more stable and generalized target task processing effect. Attached Figure Description

[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0030] Figure 1 This is a schematic diagram of an application environment for a training and processing method based on time-series performance gain according to an embodiment of the present invention;

[0031] Figure 2 This is a flowchart illustrating an embodiment of the training and processing method based on time-series performance gain according to the present invention.

[0032] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the training and processing device based on time-series performance gain of the present invention;

[0033] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0034] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0035] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0036] The training and processing method based on time-series performance gain provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain video training samples from the user terminal and generate a video entropy factor. Based on the video entropy factor, it adjusts the reward intensity and organizes the training process, constructs ordered and shuffled frame content, and generates response results through the training model. Based on the comparison of response results, it determines the temporal performance gain, generates a reward signal based on the adjusted reward intensity and temporal performance gain, inputs the reward signal into the policy optimization process to update the policy parameters, processes the training process based on the updated policy parameters to generate training output, updates the training model based on the training output, and repeats the above operations until the training termination condition is met to obtain the target model. The target model processes the target task to generate the target task result. This invention introduces a video entropy factor and temporal performance gain into the reward signal generation, enabling the training process to dynamically adjust the reward intensity and explicitly guide the model to focus on cross-time dimension reasoning capabilities. The video entropy factor ensures that video segments with high motion complexity and scene changes receive higher training weights, while the temporal performance gain supplements the time sensitivity of cross-frame reasoning. The combination of the two ensures that the generated reward signal considers both sample uncertainty and temporal rationality. Through this mechanism, the model avoids relying on single-frame information during training, gradually improving its video understanding and reasoning capabilities across time periods, thereby achieving more stable and generalized target task processing results. The user end can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server end can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0037] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the training and processing method based on timing performance gain provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0038] like Figure 2 As shown, the training and processing method based on time-series performance gain proposed in this invention includes the following steps:

[0039] S10, Obtain video training samples, and generate a video entropy factor based on the motion entropy and scene entropy of the video training samples;

[0040] In this embodiment, acquiring video training samples is the starting point. These samples can originate from publicly available video datasets or from internal video data accumulated within business scenarios. Data acquisition methods include surveillance cameras, mobile terminal cameras, and video sequences exported from virtual simulation platforms. The characteristic of video training samples is that they contain continuous frame sequences and retain information on motion and scene changes over time. This type of data possesses both spatial and temporal dynamic characteristics, thus providing a data foundation for subsequent entropy calculations and complexity characterization.

[0041] After acquiring video training samples, motion-related descriptive quantities need to be extracted. Motion entropy calculation typically relies on optical flow feature vectors, where optical flow represents the velocity and direction of motion of pixels in adjacent frames. By performing differential analysis on consecutive frames, the motion vector of each pixel can be obtained. Then, the distribution of all pixels is statistically analyzed to calculate spatial motion entropy, directional motion entropy, and temporal motion entropy. Spatial motion entropy measures the diversity of motion in different regions of the image; directional motion entropy measures whether the distribution of motion directions is concentrated; and temporal motion entropy reflects the complexity of temporal evolution by analyzing the stability of motion patterns between frames. These entropy components can describe the uncertainty of motion in a video from multiple dimensions.

[0042] Calculating scene entropy requires extracting scene change features from video training samples. Specifically, this can be done by performing semantic segmentation or scene classification on each frame, mapping the object categories and scene categories in the frame to feature distributions, and then statistically analyzing their cross-frame changes. Feature distribution entropy reflects the uniformity of the distribution of different object categories in the frame, semantic transformation entropy measures the frequency of scene semantic changes in the video sequence, and visual complexity entropy is reflected through the diversity of image texture, color distribution, and the number of objects. These scene entropy components characterize the uncertainty at the video content level.

[0043] There is an interaction between the motion entropy component and the scene entropy component. For example, intense motion accompanied by complex scene changes further increases the modeling difficulty. To reflect this interaction, it is necessary to construct an interaction term and capture the coupling effect between motion and scene through methods such as joint probability distribution or mutual information calculation. Combining the various components with the interaction term yields an initial entropy value, which can comprehensively reflect the complexity of the video in terms of both motion and scene.

[0044] The initial entropy value often has a wide distribution range, and using it directly can lead to instability in training rewards. Therefore, nonlinear transformation is required. Common processing methods include using a logarithmic function to compress the high-value range, using a hyperbolic tangent function to smooth the gradient, or constraining the value to a finite range through normalization mapping. The video entropy factor generated after nonlinear processing can balance the outstanding representation of highly complex videos with the guarantee of numerical stability.

[0045] In the implementation process, different optical flow extraction methods can be used. Traditional dense optical flow methods can be used to calculate pixel-level motion vectors, while deep learning optical flow network models can be used to obtain more stable motion descriptions. Scene change features can be extracted through image classification using convolutional neural networks, object labels and region distribution can be extracted through semantic segmentation networks, or attention mechanisms can be combined to aggregate semantic change information across frames. The calculation of interaction terms can utilize mutual information, capture the dependencies between different components using correlation coefficient matrices, or establish the coupling relationship between motion and scene nodes using graph neural networks.

[0046] The nonlinear transformation of the initial entropy value can be performed using different function forms. A logarithmic function can be chosen for compression, ensuring that high values ​​are not excessively amplified; a hyperbolic tangent function can be chosen for smooth mapping, ensuring differentiability in gradient calculation; or a sigmoid function can be chosen with a temperature parameter set to adjust the distribution range of the curvature influence factor. Different function forms adapt to different training needs; for example, a logarithmic function is used in training sensitive to differences in difficulty, while a hyperbolic tangent function is used when smooth gradient propagation is required.

[0047] For adaptation to different scenarios, video data in the healthcare field often involves patient movements, rehabilitation training, or surgical procedures. Motion entropy extraction focuses on the amplitude and stability of movements, while scene entropy focuses on the device environment and human interaction. In the financial field, video data may be used to monitor over-the-counter transactions, ATM operations, or financial service scenarios. Motion entropy focuses on detecting the completeness of transaction actions, while scene entropy focuses on scene transitions, device layout, and changes in user group behavior.

[0048] Example explanation: In the field of healthcare, rehabilitation training videos contain the frequency and amplitude of different patients' movements. By calculating motion entropy, it can be found that video samples with large movement amplitude and multiple directions are more difficult to learn. Scene entropy reflects the complexity of rehabilitation equipment and movement environment. The resulting video entropy factor can be used to highlight these more difficult training samples, so that the model pays more attention to the stability and continuity of temporal movements during training.

[0049] In the fintech business, counter operation videos or automated terminal operation videos often contain multiple interactions in a short period of time. Motion entropy can capture rapid changes in the user's hand movements, and scene entropy can reveal the impact of multiple user switching and changes in the background environment. The generated video entropy factor can help optimize the identification and analysis capabilities of these complex transaction scenarios during the training process, thereby improving the system's ability to understand abnormal behaviors or complex processes in the video.

[0050] This embodiment acquires video training samples and generates a video entropy factor based on motion entropy and scene entropy, which can quantify the complexity of the video. Motion entropy enhances the measurement of cross-frame dynamics, scene entropy supplements the uncertainties at the semantic and content levels, and the interaction term reflects the coupling relationship between the two. The video entropy factor, after nonlinear processing, can serve as a stable and effective indicator, enabling subsequent training to better identify high-value samples and allocate appropriate reward intensity, thereby improving the model's learning performance in complex temporal scenarios.

[0051] S20, Adjust the reward intensity and organize the training process based on the video entropy factor;

[0052] In this embodiment, the key to adjusting the reward intensity and organizing the training process based on the video entropy factor lies in how to dynamically scale the reward signal using the video entropy factor, enabling the model to prioritize learning video samples with higher value during training. First, the video entropy factor for the current training stage needs to be obtained. Since the video entropy factor has already been generated in the preceding processing through a combination of motion entropy and scene entropy, the acquisition operation here is not a recalculation, but rather a direct call to the value or vector generated in the preceding steps. This entropy factor serves as the basic quantitative indicator for subsequent reward adjustment.

[0053] After obtaining the video entropy factor, a reward scaling factor needs to be determined based on this factor. The reward scaling factor is a continuous variable, and its value is positively correlated with the value of the video entropy factor. It is typically implemented through a function mapping, such as a linear mapping that directly scales the entropy factor value to a preset range, or a non-linear function such as a logarithmic function or a sigmoid function to avoid excessively large extrema. By setting the scaling factor, higher reward weights can be assigned to high-entropy videos during training, thereby improving the model's ability to learn complex temporal structures.

[0054] Once the reward scaling factor is determined, it needs to be multiplied by the baseline reward strength to generate the adjusted reward strength. The baseline reward strength is a fixed value set by the training system under unbiased conditions to ensure the stability of the training process. By multiplying it by the scaling factor, a personalized reward value can be formed, and different video samples will receive different training drives during training.

[0055] After generating the adjusted reward intensity, the training samples need to be prioritized according to their magnitude, forming a sample priority sequence. The ranking criterion for the sequence is usually the magnitude of the reward intensity value, with higher reward intensity corresponding to higher priority. In this way, video samples with high complexity and high learning value are placed at the forefront of the training sequence, thereby increasing their weight in model updates.

[0056] Finally, a subset of training data is selected based on the sample priority sequence, and these subsets are organized into a training process. The selection of training data subsets can be done through fixed-ratio sampling, such as the top 20% of high-priority samples, or through dynamic adjustment, gradually replacing subsets within a time window. Organizing the training process means feeding higher-priority subsets into the training model to ensure that gradient updates are more focused on video samples that pose a greater challenge to temporal complexity.

[0057] In practice, there are several ways to map the reward scaling factor. If the goal is to maintain the linear interpretability of the reward adjustment, a linear mapping can be chosen; if it is desirable to limit over-scaling under high entropy conditions, a logarithmic function can be chosen; if it is necessary to increase the sensitivity to differences in the early stages of training and then stabilize in the later stages, a sigmoid function can be chosen and the curvature can be adjusted by the temperature parameter.

[0058] The generation of sample priority sequences can also be adjusted according to different application scenarios. In healthcare scenarios, video clips containing a large number of cross-action variations can be prioritized, while video clips with simple actions can be given lower reward intensity. In financial scenarios, video samples containing multi-user interactions or complex transaction behaviors can be trained first, while the reward priority of video samples with high repetition and simple operations can be reduced.

[0059] The selection of training data subsets can employ a dynamic window mechanism. For example, during iteration, the reward intensity can be recalculated and the priority sequence updated every few rounds, thereby continuously changing the training subset and ensuring that the model is constantly exposed to the latest high-value samples.

[0060] Example Explanation: In the healthcare field, rehabilitation training videos fall into two categories: simple repetitive movements and complex coordinated movements. Adjusting the reward intensity based on the video entropy factor allows videos of complex coordinated movements to receive higher rewards, thus enabling the model to prioritize learning the temporal patterns of these videos. This helps the system more accurately analyze patients' rehabilitation progress and movement stability.

[0061] In the fintech business, some scenarios in counter operation videos or ATM operation videos only involve single-user, single-step operations, while others involve multi-user switching and rapid operation interactions. By using reward intensity generated based on video entropy factors, videos with complex interactions can be prioritized in the training process, thereby improving the model's robustness in detecting abnormal transaction behavior or identifying abnormal user operations.

[0062] This embodiment adjusts the reward intensity and organizes the training process based on the video entropy factor, prioritizing the introduction of high-entropy, high-complexity video samples into training. This exposes the model to samples rich in temporal information during the learning process, thereby improving the model's ability to perform cross-temporal reasoning and understand complex scenes. The reward scaling and priority sequence mechanism effectively avoids the problem of the model being dominated by low-value samples during training, ensuring that computational resources and training time are concentrated on improving temporal reasoning capabilities.

[0063] S30, construct ordered frame content and shuffled frame content based on the training process, and generate answer results based on the ordered frame content and shuffled frame content through the training model;

[0064] In this embodiment, ordered and shuffled frame content are constructed based on the training process, and the response results are generated through a trained model. The core issue lies in how to transform the video data in the training process into an input format that can distinguish temporal logic. First, video segment sequences need to be extracted from the training process. This operation typically relies on a time window segmentation mechanism. For example, extraction can be performed according to a fixed number of frames, or segmentation can be based on scene change detection results. The extracted segments need to maintain the temporal correlation between frames so that they can be used subsequently to generate both ordered and shuffled inputs.

[0065] Constructing ordered frame content involves arranging the extracted video segments in their original chronological order, thus preserving the video's temporal logic and event evolution. Shuffling frame content, on the other hand, involves randomly rearranging the same video segments. The shuffling operation uses a random number seed to generate a permutation index, ensuring that the shuffling order is different in each training round. In this way, the two inputs are identical in content, but their temporal structure is disrupted, creating a comparative experimental condition.

[0066] After constructing the ordered and shuffled frame content, further data consistency verification is required. This is typically done by calculating the content hash value of the input sequences to confirm that the two inputs indeed originate from the same video segment. Hash values ​​can be generated using MD5 or SHA family functions to obtain the first content hash value for the ordered frame content and the second content hash value for the shuffled frame content, respectively. If the comparison results match, it can be confirmed that the two sets of inputs originate from the same source, differing only in their temporal order.

[0067] After consistency verification, the first inference branch of the model needs to be trained with ordered frame content to obtain the ordered frame response; simultaneously, the second inference branch of the model needs to be trained with shuffled frame content to obtain the shuffled frame response. The first and second inference branches typically share parameters but execute independently on the computation graph, ensuring that the inference outputs for the two types of inputs can be directly compared. The response results are usually output as vectors, which can be either the confidence distribution of the classification labels or an embedding representation of free text, depending on the model's design goals.

[0068] In practical implementation, the method for extracting video segments can be adjusted according to the application scenario. If processing action recognition data, continuous segments can be extracted at a fixed frame rate; if it is an interactive scenario, segments can be extracted by combining key event trigger points. The method of constructing shuffled frame content can also be flexibly adjusted. In addition to completely random shuffling, a partial shuffling method can be used, changing only the order of some segments, to test the robustness of the model under partial temporal perturbations.

[0069] Besides using hash functions, content consistency verification can also utilize frame-level feature comparison. For example, features can be extracted from video segments to generate signature vectors, and then cosine similarity can be calculated. When the similarity exceeds a threshold, the input consistency is confirmed. The branching structure of the training model can also be implemented in various ways, such as using completely independent inference branches or sharing feature extraction layers and separating decision layers, thereby reducing computational resource consumption while ensuring fair comparison.

[0070] Example Explanation: In the healthcare field, sports rehabilitation videos contain a series of continuous movements. If the video clips are input in their original order, the model can learn the patterns of movement evolution. However, if the order is disrupted, the logic is broken. By comparing the responses generated from the two inputs, the system can help determine whether the patient's movements have reasonable temporal connections, thereby providing support for rehabilitation assessment.

[0071] In the fintech sector, videos of counter operations or self-service device operations often consist of multiple sub-operations, such as identity verification, amount input, and transaction confirmation. By constructing ordered and shuffled inputs, the model can learn the correct temporal relationship of these operations when comparing the response results, thereby improving its ability to detect abnormal operation sequences and contributing to risk identification and prevention.

[0072] This embodiment introduces a contrast signal during training by constructing ordered and shuffled frame content and generating corresponding responses for each. This forces the model to explicitly focus on temporal logic when faced with inputs from the same source but with different temporal structures. This not only enhances the model's sensitivity to cross-temporal information but also effectively prevents the model from taking shortcuts by obtaining inference results through local frame information, thereby improving its true temporal reasoning ability.

[0073] S40, determine the timing performance gain based on the comparison of the answer results;

[0074] In this embodiment, determining the temporal performance gain based on the comparison of response results requires first obtaining the response results from ordered frame content and shuffled frame content. These response results are typically presented in the form of probability distributions or embedding vectors, containing the model's inference outputs at different time sequences. The basis for comparison lies in extracting elements that characterize the model's stability and confidence level; therefore, the response results need to be converted into confidence feature vectors. Confidence feature vectors can be obtained by extracting the distribution information of the model's predicted probabilities for each category, or by normalizing the output embedding vectors, thus forming a comparable numerical representation.

[0075] After obtaining the confidence feature vectors of the ordered and shuffled frame responses, a difference metric needs to be calculated between them. This difference metric reflects the difference in inference stability and consistency of the model when faced with inputs that preserve temporal logic versus inputs that disrupt temporal logic. The difference can be calculated using Euclidean distance, Manhattan distance, inverse cosine similarity index, or other numerical measures that can quantify the difference. A larger difference metric indicates a more pronounced instability of the model under temporal perturbations, and therefore a higher sensitivity to temporal performance.

[0076] The difference metric needs to be further mapped into a gain signal that can be passed to the training process, thus requiring a differentiable mapping function. This function is typically designed to be smooth and continuous, such as the sigmoid, hyperbolic tangent, or linear normalization function. The differentiable mapping function ensures that the gain value can transmit the gradient signal during backpropagation, allowing the training process to perceive the impact of the difference metric on parameter updates. This function yields the original gain value, the magnitude of which directly corresponds to the importance of temporal performance.

[0077] After generating the initial gain value, it needs to be constrained within a range. The purpose of this constraint is to prevent the gain value from being too large, leading to gradient explosion, or too small, causing the signal to disappear. Constraints can be achieved by setting upper and lower thresholds, such as limiting the value to between 0 and 2, or by normalizing the value to a fixed interval. After constraining, the temporal performance gain is obtained, and this gain will be passed to the subsequent reward signal generation stage.

[0078] In practice, the method for extracting confidence feature vectors can be adjusted according to the task type. If the task is categorical output, the category probability distribution can be directly extracted as the confidence feature vector; if it is a text generation task, the word-level probabilities of the output sequence can be used and a weighted average can be formed into a vector representation. The calculation of the difference measure can be flexibly switched, and the measure method may be different for different tasks. For example, cosine similarity is used in semantic relevance tasks, while L2 norm is used in numerical output tasks.

[0079] The form of the differentiable mapping function can also be adjusted. For example, when it is necessary to enhance the importance of highly dissimilar samples, a hyperbolic tangent function can be used to amplify the dissimilar values ​​in a non-linear manner; when stability needs to be ensured, a linear function combined with threshold constraints can be used. The implementation of range constraints can also be diversified. Hard truncation can be used to limit the gain value between upper and lower limits, or soft constraints can be used, such as normalizing to the mean and standard deviation range. Different implementation methods correspond to different training objectives and can be adjusted according to the characteristics of the task.

[0080] Example Explanation: In the healthcare field, remote rehabilitation training videos often contain a series of continuous movements. For example, when a patient completes a full limb rehabilitation movement, multiple steps need to be performed sequentially. By inputting ordered video clips and shuffled video clips into the training model and comparing the responses, we can detect whether the model truly captures the sequential logic of the movements. A larger difference metric indicates that the model is more sensitive to the temporal sequence of movements and can be used for subsequent temporal enhancement training.

[0081] In the fintech business, video clips of users performing online operations may include steps such as identity verification, amount input, and confirmation. If the model's responses differ significantly between ordered and shuffled clips, it indicates that the model can capture the importance of the operation sequence in the task logic. Generating temporal performance gains through this mechanism can help the system strengthen its ability to learn the order of transaction steps during training, thereby improving the accuracy of anomaly detection and risk control.

[0082] This embodiment explicitly quantifies the model's sensitivity to temporal logic understanding by comparing ordered frame responses with shuffled frame responses and extracting the difference signals. By converting the differences into temporal performance gains, the training process can focus more optimization efforts on samples with significant temporal differences. This not only avoids the model's tendency to rely on single-frame shortcuts for inference but also enables it to gradually learn cross-temporal logical relationships during training, thereby enhancing its ability to understand and process complex video data.

[0083] S50, generate a reward signal based on the adjusted reward intensity and the timing performance gain;

[0084] In this embodiment, generating a reward signal based on the adjusted reward intensity and temporal performance gain requires combining signals from two different sources. This ensures the reward value reflects both the importance of the sample itself and the differences in temporal logical understanding. First, the adjusted reward intensity has been scaled using a video entropy factor, reflecting the complexity and learning value of different training samples. Simultaneously, the temporal performance gain, derived from the comparison of responses between ordered and shuffled frames, reflects the model's differences in temporal logic performance. These two quantities need to be combined in a unified function space to form a reward signal usable for gradient propagation.

[0085] In practical implementation, the temporal performance gain is input into a time correction function, generating a time correction term. The time correction function typically possesses differentiability and numerical smoothness, and its role is to transform the original gain value into a numerical range that fits the training objective. For example, the sigmoid function or the hyperbolic tangent function can be used to restrict the gain value to a specific range, thereby preventing outliers from interfering with subsequent training. The generated time correction term is essentially a correction factor used to dynamically adjust the reward intensity.

[0086] Subsequently, the adjusted reward intensity needs to be multiplied by the time correction term to obtain the composite reward value. This multiplication relationship reflects the interaction between reward intensity and temporal performance; that is, when the video sample complexity is high and the model performs well in temporal logic, the reward value can be amplified; conversely, when one part is weak, the reward value is suppressed. This combination ensures that the reward signal truly reflects the comprehensive training value of the samples.

[0087] After generating the synthetic reward value, its differentiability needs to be verified. The goal of the verification process is to ensure that the reward value can provide an effective gradient during backpropagation, without gradient vanishing or exploding. Differentiability verification can be achieved through numerical analysis, such as checking whether the derivative of the synthetic reward value is continuous in the function space, and ensuring its differentiability through numerical smoothing or truncation operations if necessary.

[0088] Once the synthesized reward value passes the differentiability verification, it can be used as a reward signal in the policy optimization process. The numerical range of the reward signal usually needs to remain stable to ensure that different training samples are compared on the same scale, thereby maintaining the balance of the training process.

[0089] In implementation, the choice of time correction function can be diverse. If it is necessary to enhance the model's sensitivity to subtle temporal differences, an exponential function can be used to amplify small differences; if it is necessary to maintain training stability, a linear function combined with upper and lower threshold control can be used. Besides multiplication, the synthetic reward value can also be generated through weighted summation, but multiplication better reflects the interdependencies.

[0090] In the differentiability verification stage, symbolic computation can be used to directly differentiate the function and check if the derivative is within the numerical range. Alternatively, random sampling verification can be used, which involves collecting local derivatives of the synthesized reward value within a certain range; if the derivative satisfies the continuity condition, it passes the verification. For environments with limited computational resources, an approximate verification method can be used, replacing analytical differentiation with local differences to improve computational efficiency.

[0091] The numerical range of the reward signal can be further optimized to meet different task requirements. For example, in scenarios where it is necessary to strongly distinguish the value of samples, a larger numerical range can be set; while in scenarios where it is necessary to maintain training balance, a smaller range can be set and normalization operations can be used.

[0092] Example Explanation: In the healthcare field, such as video training for rehabilitation motion analysis, the complexity of different video segments varies. Some segments contain simple movements, while others have complex and sequentially ordered movements. Through a reward signal generation mechanism, video segments with complex movement sequences will receive higher reward weights. Simultaneously, the model's performance when comparing temporal logic will be amplified or suppressed, thereby guiding the model to better learn the patterns of movement connections.

[0093] In the fintech field, for example, videos or interaction sequences used to monitor user transaction processes vary in complexity and temporal importance across different steps. Certain steps are crucial for risk control, such as identity verification and fund confirmation. Through a reward signal mechanism, the system can assign higher weights to video training samples corresponding to these critical steps. Combined with the model's response to the correctness of sequential logic, this encourages the model to strengthen the temporal dependencies of the transaction process during training, thereby improving the accuracy of anomaly detection and compliance review.

[0094] This embodiment generates a reward signal by combining the adjusted reward intensity with temporal performance gain, enabling the training process to simultaneously consider sample complexity and temporal logic performance. The reward signal not only guides the model to gain greater attention on complex video samples but also effectively corrects the model's shortcomings in temporal inference. Through a time correction function and a differentiability verification mechanism, the reward signal possesses numerical stability and gradient transitivity, thereby ensuring the optimization process is efficient and reliable, ultimately improving the model's inference ability across time dimensions.

[0095] S60, the reward signal is input into the policy optimization process to update the policy parameters, and the training process is processed based on the updated policy parameters to generate training output;

[0096] In this embodiment, inputting the reward signal into the policy optimization process to update the policy parameters, and then processing the training process and generating training output based on the updated policy parameters, is a key step in combining the reward signal formed in the previous stage with the policy parameter adjustment mechanism. The value of the reward signal is first input into the optimization process to drive the policy update. In implementation, the gradient update amount needs to be derived from the reward signal. This update amount can characterize the direction and magnitude of the reward signal's sensitivity to the policy parameters. Through the calculation and superposition of parameter gradients, the existing policy parameters can be gradually corrected to tend towards generating higher reward values.

[0097] After calculating the gradient update, the system adjusts the current policy parameters based on this update, resulting in updated policy parameters. This process involves not only numerical additions and subtractions but may also include adjusting the learning rate factor to ensure that the updated policy parameters reflect the positive guidance of the reward signal without losing stability due to excessive adjustments. The updated policy parameters need to be immediately deployed to the inference execution environment, i.e., loaded into the model inference engine. As the computation execution module, the inference engine is responsible for applying the new parameters to the actual input data to test its performance in specific tasks.

[0098] During inference execution, video segment data from the training process is sequentially input into the inference engine, which loads new parameters. This data undergoes feature extraction, temporal modeling, and semantic analysis through the model's computational layers, outputting corresponding inference results. These results are then encapsulated to form a training output data package containing semantic understanding. This training output data package is a direct product of the policy optimization process and a crucial input for subsequent model updates and performance verification. In this way, the reward signal-driven policy parameter adjustment and data processing within the training process form a closed loop, establishing a direct correlation between reward optimization and model output.

[0099] In the implementation process, different methods can be used to generate gradient update amounts. For example, derivative-based numerical calculations can be used to directly derive the partial derivatives of the reward signal with respect to the policy parameters, or a difference approximation method can be used to estimate the trend of change in a local numerical space. During parameter updates, a fixed learning rate or an adaptive learning rate can be used to improve training stability by dynamically adjusting the update magnitude.

[0100] During the inference engine loading phase, parameter deployment can be implemented either as a complete replacement (loading all updated policy parameters at once) or as an incremental update, replacing only the significantly changed parameters to reduce computational overhead. In video segment data processing, the feature extraction module can employ convolutional computation units to capture spatial information or attention computation structures to enhance temporal dependencies. The training output data package can be encapsulated to include semantic understanding results in vector form or structured label information, allowing for the use of different forms of supervision signals in subsequent model updates.

[0101] The strategy optimization process can be adapted to different application scenarios. For example, in training scenarios with high timeliness requirements, the interval between parameter updates can be shortened to improve the model's ability to respond quickly to reward signals; in scenarios with large-scale data, batch gradient calculation can be used to reduce the volatility of individual training sessions and improve overall stability.

[0102] Example Explanation: In the healthcare field, video clip data can be replaced with action sequences recorded during rehabilitation training. Reward signals guide parameter optimization, enabling the model to pay greater attention to temporal sequence and action details when analyzing complex movement sequences. The generated training output data package not only includes the recognition of individual actions but also the logical understanding of action combinations, thereby better supporting rehabilitation progress monitoring and personalized training program recommendations.

[0103] In the fintech business, the input to the training process can be replaced with user interaction videos or interface operation sequences. Reward signals, by optimizing strategy parameters, guide the model to more accurately identify abnormal logic or sequential deviations in user operations. The training output data package provides an understanding of the completeness of the transaction process, thereby providing a more robust basis for judgment in the risk control system and improving the accuracy of anomaly detection and risk identification.

[0104] This embodiment achieves direct coupling between the reward signal and model parameter adjustment by inputting the reward signal into the policy optimization process and updating the parameters, and then combining it with the training process to generate the output result. This process not only dynamically improves the model's performance on complex samples, but also ensures that the output result reflects the optimization direction of the reward mechanism. The closed-loop relationship between the training output and the reward signal enables the model to gradually converge to a state with stronger temporal reasoning ability and higher generalization performance.

[0105] S70, Update the training model based on the training output;

[0106] In this embodiment, updating the training model based on the training output is a data-driven weight correction process. The goal is to transform the supervision signals, trajectory information, and temporal reward-related elements carried in the training output into backpropagable numerical constraints, thereby improving the parameter set of the training model in one or more iterations. The training output has been encapsulated into a data packet in previous stages, typically containing sample identifiers, timestamp sequences, model-generated response text or intermediate representations, probability distributions for each generation position, confidence feature vectors, temporal performance gains, time correction terms, reward signals, and branch markers for ordered and shuffled frame content. The update process first unpacks and verifies the training output, checking the branch markers and hash consistency results to ensure that paired samples correspond to the same frame set in content, differing only in temporal arrangement. Subsequently, an objective function is established, with each component maintaining alignment with the previous stages in terms of name and origin. The reward-related component uses the reward signal and its composite value after interaction with the time correction term to apply to the corresponding generation position or segment level. The temporal contrast component constructs a metric-driven constraint based on the difference in confidence feature vectors between ordered and shuffled frame responses, enabling the model to distinguish cross-temporal dependencies. If the training output includes semantic understanding results or labels, cross-entropy or sequence-to-sequence consistency constraints can be introduced to ensure that updates do not weaken recognition and representation capabilities. To pass these constraints to the training model, the gradient direction and magnitude with respect to the current parameters are calculated, and one or more weight updates are completed using mechanisms such as learning rate, weight decay, gradient pruning, and gradient accumulation. The update objects are not limited to the generation head; they can also cover the visual encoder, temporal modeling layer, multimodal fusion layer, and value estimation head, but the update range needs to be controlled through masking or freezing strategies to avoid large drifts in the early convergence phase. After weight adjustment, the weights, optimizer state, gradient scale, and mixed precision scaling factor for this round are recorded with the same version number and written to a checkpoint to support rollback and breakpoint continuation training. If the training output is organized in batches, it is necessary to ensure that the alignment relationship between ordered and shuffled pairs of samples is maintained within the batch, and consistent partitioning and aggregation rules are adopted in a distributed environment to avoid target offset caused by inconsistencies in statistics between multiple machines.

[0107] A phased strategy can be adopted for parameter updates. The first phase updates only the language generation head and value estimation head, keeping the visual encoder and temporal modeling layer frozen, allowing reward-related gradients to preferentially shape the output distribution and temporally sensitive decoding behavior. The second phase gradually unfreezes the temporal modeling layer and reduces the learning rate, allowing temporal performance gains to have a more direct shaping effect on cross-frame dependent representations. The third phase unfreezes the high-level blocks of the visual encoder with smaller steps, prompting the discriminative features of motion patterns and scene transitions to align with upstream temporal constraints. This phased strategy can reduce the instability risk caused by large-scale synchronous parameter updates.

[0108] Weight scheduling can be introduced into the objective function combination. The reward term applies the reward signal to the logarithmic form of the generation probability, amplifying the advantageous direction; the temporal comparison term takes the difference in confidence feature vectors as input and uses a hyperbolic mapping with boundary intervals or temperature control to adaptively enhance the separation between ordered frame responses and shuffled frame responses as training progresses; the supervised consistency term is used to maintain text readability and task relevance, and its weight gradually decreases in the later stages to avoid suppressing temporal constraints. The weights of each term can be adjusted by the moving average of the temporal performance gain in the training output. When the gain increases, the weight of the temporal comparison term is moderately increased; when the gain stagnates, the weight of the reward term is increased to enhance exploration.

[0109] For numerical optimization, an adaptive optimizer with momentum and weight decay can be used, combined with gradient pruning to limit the impact of extreme samples on the update magnitude. To address memory pressure caused by long video sequences, block-based backpropagation and activation checkpoints can be employed, dividing the process into time windows for forward and backpropagation. Simultaneously, attention connections across windows are reused through cached key-value mappings to reduce redundant computation. In mixed-precision training, the loss scaling factor needs to be recorded, automatically backing down and reducing scaling when gradient overflow occurs to ensure safe updates.

[0110] In data-parallel or tensor-parallel environments, training outputs need to be synchronously aggregated across devices for reward-related statistics, such as the batch mean and variance of synthesized reward values ​​and temporal difference measures, to ensure the objective function has a consistent scale from a global perspective. During cross-device alignment, consistent slices of paired samples must be preserved to avoid splitting ordered and shuffled frames into different synchronization domains, which could lead to the invalidation of comparison terms. For version management, each update is assigned a monotonically increasing training model version number, which is archived along with the batch number of the training output generated for that version and a snapshot of the reward signal. In case of abnormal divergence, the system rolls back to the most recent stable version and reduces the learning rate or temporarily increases the weight of the supervision consistency term to restore readability.

[0111] In terms of robustness and generalization, small perturbations can be introduced into the training output to enhance performance. For example, low-amplitude random noise can be added to the confidence feature vector, or temperature jitter can be used in the time correction term to improve tolerance to distribution shifts without disrupting temporal relationships. For noisy videos or low frame rate samples, the tolerance range of the temporal comparison term can be increased during updates to avoid misjudging perceptual difficulties as temporal weakening. For samples with extremely high motion complexity, the normalization parameters of the temporal modeling layer can be fine-tuned locally first to make the attention distribution more concentrated, and then the updates of the generator head weights can be relaxed to prevent gradients from spreading out of control in high-entropy scenes.

[0112] Example Explanation: In the healthcare business domain, the training output originates from the reasoning process oriented towards action assessment, including textual explanations, confidence feature vectors, and temporal performance gains for each key action segment. When updating the training model, firstly, the probability distribution of the explanation of the complete action chain is amplified by the reward signal; secondly, the temporal comparison constraint is driven by the confidence difference between ordered and scrambled action segments; and thirdly, a small amount of supervision is used to maintain terminology standardization for consistency. After multiple rounds of updates, the weights can more stably cover the start, transition, and end stages when processing coherent action sequences, and output lower confidence for scrambled segments. In the rehabilitation guidance system, this is reflected in a more sensitive warning for incorrect action sequences.

[0113] In the fintech business, the training output reflects the understanding of transaction process videos or screen recordings, including key step identification text, confidence feature vectors for each step, and temporal performance gains calculated from process integrity. When updating the training model, reward signals are used to increase the probability of generating complete and legitimate paths, while the difference between ordered and shuffled processes is used to strengthen cross-step dependencies in the model's weights. After the update, the system generates significantly low-confidence judgments for missing or swapped operation segments, helping to detect abnormal processes earlier in risk control audits and reduce false positives and false negatives.

[0114] This embodiment updates the trained model based on the training output, directly projecting temporal rewards and contrastive differences into the weight space, forming continuous constraints around time dependencies. After updating, the weights tend to produce high-confidence, consistent responses across frames when faced with ordered frames, and maintain discriminative ability when faced with shuffled frames, thus reducing the tendency to rely on single-frame shortcuts. Through phased unfreezing and weight scheduling, the representation layer and generation layer are gradually aligned, expanding the training stability region and improving convergence speed. Distributed synchronization and versioned rollback reduce the failure probability of large-scale training, while long-sequence memory optimization ensures the feasibility of video-level training. The overall effect is reflected in improved temporal inference accuracy and enhanced robustness under abnormal temporal conditions.

[0115] S80, repeatedly execute the reward intensity adjustment step, timing test step, gain determination step, signal generation step, policy update step, process processing step and model update step until the training termination condition is met and the target model is obtained.

[0116] In this embodiment, the cyclic execution adopts a closed-loop structure with seven consecutive stages: reward intensity adjustment, timing test, gain determination, signal generation, policy update, process processing, and model update, with the training termination condition serving as the stopping threshold. At the beginning of each round, the expected distribution and state metric of the training output are generated from the training model obtained in the previous round and the current data queue; the reward intensity adjustment stage reads the video entropy factor and the baseline reward intensity, calculates the adjusted reward intensity, and refreshes the sample priority queue and batch composition accordingly; the timing test stage performs dual-branch inference paths on the ordered and shuffled frame contents of the same video segment, forming ordered frame response results and shuffled frame response results, and ensures the equivalence of frame content through hash consistency verification; the gain determination stage extracts the confidence feature vectors of the two types of responses, calculates the difference metric, obtains the original gain value through differentiable mapping, and then obtains the timing performance gain through range constraints; the signal generation stage... The reward generation stage feeds the temporal performance gain into a time correction function to obtain a time correction term, which is then multiplied by the adjusted reward intensity to obtain the synthetic reward value. After completing the differentiability verification, the reward signal is output. The policy update stage uses the reward signal for target construction and gradient derivation to obtain the gradient update amount and update the policy parameters to obtain the updated policy parameters. The process processing stage loads the updated policy parameters into the model inference engine, parses the video segment data in the training process, generates semantic understanding results, and encapsulates them into a training output data package. The model update stage unpacks and aligns the training output data package, combines the reward term, temporal comparison term, and necessary consistency term to complete weight correction, and writes it into the checkpoint. These seven stages form a complete loop. To avoid bias accumulation, multi-dimensional monitoring quantities are maintained within the loop, including the moving average and variance of the synthetic reward value, the median and quantiles of the temporal performance gain, language quality indicators, gradient norm and parameter drift, and the stability of the validation set task score and video entropy hierarchy. Training termination conditions employ multi-criteria gating, typically including maximum number of rounds or computational budget, a stagnation criterion where temporal performance gain increases below a threshold within a certain window, early stopping when validation set scores do not improve within a patience window, simultaneous entry of reward variance and gradient norm into a stable region, and safety thresholds such as numerical anomalies or divergence detection. At the end of each round, the loop controller refreshes monitoring data and evaluates the stopping flag. Once the stopping criteria are met, parameters are frozen and the target model is labeled. To ensure the stability of terminology and data dependencies, all metrics within the loop use the same version of paired samples of ordered and shuffled frame content. Sampling and sharding rules remain consistent across devices to avoid the invalidation of comparison terms. Numerically, the learning rate and temporal correction temperature are annealed with each round to reduce late-stage oscillations. Reward scaling and temporal comparison weights are adaptively scheduled according to the gain increase, maintaining a balance between exploration and convergence at different stages.

[0117] In a standalone environment, a fixed-length loop and fixed-threshold control method can be used. The number of samples, maximum number of rounds, and learning rate annealing plan are set for each round. A dual-threshold strategy using the moving average of the synthetic reward value and the validation set time-series metrics is employed to determine whether to stop early. This configuration is simple to implement and suitable for training tasks with moderate data scale and moderate video length.

[0118] In a multi-machine data parallel environment, synchronous aggregation and entropy hierarchical sampling can be employed. In each round, samples are divided into several levels based on the video entropy factor, and quotas are allocated between levels according to the adjusted reward intensity, ensuring that high-entropy videos receive more update opportunities. Global reduction is used for the statistics of reward signals and temporal performance gains to ensure consistent target scale; the same segmentation strategy is maintained for ordered and shuffled frame content to prevent paired samples from splitting across the synchronization domain. A globally consistent patience window and distribution offset detection are introduced as termination conditions to avoid false shutdowns triggered by single-machine fluctuations.

[0119] For long videos and high-resolution scenarios, temporal segmentation and block-based reverse processing can be employed. The video is divided into overlapping time windows, with timing tests and workflow processing performed within each window. Key-value caching maintains cross-window connectivity between multiple windows. Model updates utilize gradient accumulation and mixed precision, with gradient clipping to limit peak values ​​and loss scaling to prevent overflow. Memory usage and throughput monitoring are incorporated into the termination condition; if resource pressure is excessive, the time window length can be reduced or the number of accumulation steps increased to maintain stability.

[0120] Robust scheduling can be employed on training sets that are prone to noise and sudden scene changes. The time-corrected temperature uses a lower value at high entropy levels to reduce the risk of gradient bursts, and a higher value at low entropy levels to amplify weak temporal signals. The upper limit of the reward scaling is tightened with each round to prevent excessive bias towards extreme samples in later stages. The early stopping criterion adds a lower bound constraint on the reward variance to avoid spurious convergence.

[0121] In resource-constrained environments, intermittent evaluation and batch freezing can be adopted. In every few rounds, the validation metrics are evaluated only on a subset. Once the threshold is reached, the visual encoder and temporal modeling layer are frozen, and only the generation head and value head are fine-tuned to reduce computational overhead. The termination condition prioritizes the parallel logic of stagnation and budget threshold to ensure that a stable target model is obtained as soon as possible.

[0122] Example Description: In the healthcare field, for a dataset of chronic disease rehabilitation action videos, high-entropy segments are prioritized for batch processing during loop execution. The ordered and shuffled results obtained from bi-branch inference drive temporal gain. After policy updates, the same action chains are processed again. If the cross-frame consistency score and gain increase plateau within the monitoring window, and the accuracy of distinguishing the transitions between actions in the validation set reaches a threshold, the process stops and the target model is output. This target model produces higher-confidence interpretations of coherent action segments and significantly reduces the weighting of shuffled action sequences, which is beneficial for evaluating training progress videos and remote rehabilitation guidance videos.

[0123] In the fintech business, when faced with screen recordings of transaction processes or videos of counter operations, the video entropy factor is used for hierarchical sampling within the loop. Segments with frequent page switching and many operation objects are given higher priority. The temporal gain comes from the difference between the correct process and the disordered process. After the strategy is updated and implemented, the process records are processed again and the monitoring volume is accumulated. When the moving average of the synthetic reward value is stable and the consistency score of the validation set process does not improve significantly for several consecutive windows, and the budget threshold is reached, the process terminates, and the target model is obtained, which is used to improve the stability and recall rate of abnormal process identification.

[0124] This embodiment transforms the chain of reward intensity adjustment, temporal comparison, and policy update into a convergent iterative process through loop control. Combined with adaptive weights and temperature annealing, this makes temporally relevant gradients stronger in the early stages and more stable in the later stages, avoiding shortcuts that rely solely on single frames. Multi-criteria stopping criteria reduce overfitting and training oscillations, while checkpoints and alignment rules reduce target bias caused by distributed offsets, ultimately resulting in a more stable target model across time inference.

[0125] S90, the target task is processed through the target model to generate the target task result.

[0126] In this embodiment, after the target model enters the inference state, it first completes task access and routing, receiving the input object and task description of the target task. The input object can be a video, image sequence, or timestamp-aligned multimodal signal, and the task description is passed in as an instruction template or structured parameters. The access layer selects a task header and decoding strategy based on the task description. The task header covers forms such as classification, detection, segmentation, retrieval, question answering, summarization, and process compliance judgment. The decoding strategy switches between greedy, sampling, bundle search, and constraint decoding. The input standardization process includes frame rate resampling, resolution alignment, duration pruning or expansion, color space unification, regularization of speech and text channels, and normalization and labeling consistent with the training period. Long-term objects use sliding time windows and cross-window caching. Time windows share key-value caches and segment-level memory to ensure that cross-window causal dependencies are preserved. To reduce redundant calculations, frame embedding and intermediate attention key values ​​support cache hits and partial failure recalculation. The caching strategy is dynamically adjusted according to scene popularity and window coverage. The forward inference phase consists of four levels: visual encoding, temporal modeling, cross-modal alignment, and task header. Visual encoding extracts frame-level features, temporal modeling aggregates temporal dependencies, cross-modal alignment maps video, text, and speech to the same semantic subspace, and the task header generates intermediate results for task alignment. The result generation phase completes candidate structure construction, confidence estimation, and uncertainty quantification. Confidence is estimated using temperature scaling, Dirichlet calibration, or Monte Carlo sampling. For text-based output, punctuation-level and sentence-level confidence intervals are provided; for structured output, field-level confidence and consistency scores are provided. Post-processing performs rule-based repair and constraint satisfaction based on task type, including temporal order constraints, entity alignment constraints, bounding box validity, and process state machine consistency. To ensure traceability, the inference process records input summaries, version fingerprints, hashes of key intermediate quantities, and metric snapshots. To adapt to multiple scenarios, the service layer supports both batch parallelism and streaming return scheduling. Batch parallelism is used for high throughput, while streaming return is used for low-latency dialogue or monitoring. The final output of the target task results includes the main conclusions, auxiliary evidence index, time location information, confidence distribution and necessary structured fields. The output format is consistent with the task agreement so that it can be directly entered into the downstream business process.

[0127] In offline batch processing scenarios, a high-precision configuration using fixed windows and beam search can be adopted. The video is divided into fixed-length windows with overlapping areas, and reasoning is performed in window order. The beam search width and maximum generation length are set separately for each task type. Candidates from adjacent windows are merged, and deduplication and splicing are performed based on time consistency. This is suitable for needs such as data archiving, historical playback, and comprehensive auditing.

[0128] In online interactive scenarios, streaming time windows and incremental caching can be used. The inference engine progresses by rolling forward with shorter time windows, retaining key-value caches and semantic memories of the most recent windows. Text decoding adopts a greedy plus sampling hybrid strategy of temperature annealing, which maintains diversity while ensuring the first response latency. When the input source is a long video that is captured and uploaded simultaneously, the framework immediately provides a stage result after reaching the threshold number of frames, and performs incremental corrections when subsequent windows are reached, which is suitable for human-computer collaboration and online monitoring.

[0129] In a multi-task unified deployment scenario, a routing gateway and a task header factory can be used. The gateway selects the appropriate task header based on the task description or upstream tags and dynamically assembles the adaptation layer, such as a text constraint vocabulary, a domain entity dictionary, and layout priors. If no matching task header is found, a general question-and-answer header is enabled and an output structurer is attached to parse the free text into key-value pairs or time-series tags to ensure task coverage.

[0130] In high-reliability scenarios, dual-path consistency checks can be introduced. A single request is executed through two independent inference replicas, which use different random seeds or different temperature parameters to compare the consistency of key fields or time-series labels. If the consistency is below a threshold, re-inference or cooling recalculation is triggered, and a tagged result is returned. At the same time, inconsistent examples are recorded and added to the replay pool.

[0131] In resource-constrained deployments, distillation and hierarchical inference can be employed. A high-capacity target model is distilled to produce a medium-capacity execution unit for most requests, and inputs with low confidence or constraint conflicts are then escalated to high-capacity branch processing. Feature extraction and temporal modeling are hierarchically hosted on different devices, with the edge handling frame features and the central side handling cross-modal processing and generation, thus reducing the load on the edge devices.

[0132] In privacy-sensitive scenarios, anonymization and desensitization can be built-in. Before inference, sensitive areas such as faces, ID cards, barcodes, and accounts are automatically detected and blurred or replaced. Sensitive source tracing marks and processing records are attached during the output stage. The audit interface only exposes hashes and digests to protect the original content.

[0133] Example Description: In the healthcare business field, for home-based exercise rehabilitation guidance videos, the target model receives frame sequences and the task description "identify the sequence of actions and identify the erroneous steps". The input is a standardized completion frame rate aligned with the resolution. The temporal modeling outputs segmented action labels and timestamps. Post-processing is performed to repair local noise based on the trajectory sequence constraints of human key points. The results include the start and end times of each action, the correctness label, the confidence level, and the specific segments that need to be practiced, meeting the home rehabilitation platform's need for interpretable feedback.

[0134] In the fintech business field, for remote account opening process screen recording and photo sequence, the target model receives the task description "determine the consistency of process order and data", multimodal alignment unifies the encoding of page element text, voice prompts and operation trajectory, the task header outputs the step sequence and abnormal event list, post-processing removes illegal jumps according to process state machine constraints, and the result gives whether the process is passed or not, the type of abnormality, the corresponding time period and confidence level distribution, which facilitates the risk control system to perform automated review and manual spot checks.

[0135] This embodiment achieves stable, structurally consistent, temporally consistent, and traceable target task results by coordinating the target model with task routing, input standardization, long-term caching, task header assembly, confidence calibration, and constraint post-processing. This is accomplished across different durations and modalities. Sliding time windows and incremental caching reduce redundant computations while maintaining cross-window dependencies, calibration and consistency checks suppress occasional decoding drift, and hierarchical inference balances latency and cost, ultimately improving the reliability of temporal understanding and business availability in practical environments.

[0136] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a training and processing method, apparatus, device, and medium based on temporal performance gain, comprising: acquiring video training samples and generating a video entropy factor; adjusting the reward intensity and organizing the training process based on the video entropy factor; constructing ordered and shuffled frame content and generating response results through a training model; determining the temporal performance gain based on a comparison of the response results; generating a reward signal based on the adjusted reward intensity and temporal performance gain; inputting the reward signal into a policy optimization process to update policy parameters; processing the training process based on the updated policy parameters to generate training output; updating the training model based on the training output; repeating the above operations until the training termination condition is met to obtain the target model; and processing the target task through the target model to generate the target task result. This invention, by introducing a video entropy factor and temporal performance gain in the generation of the reward signal, enables the training process to dynamically adjust the reward intensity and explicitly guide the model to focus on inference capabilities across time dimensions. The video entropy factor ensures that video segments with high motion complexity and scene changes receive higher training weights, while the temporal performance gain supplements the time sensitivity of cross-frame inference. The combination of the two ensures that the generated reward signal considers both sample uncertainty and temporal rationality. Through this mechanism, the model avoids relying on single-frame information during training, gradually improving its ability to understand and reason about videos across time periods, thereby achieving a more stable and generalized target task processing effect.

[0137] In one embodiment, step S10 includes:

[0138] S101, Obtain the video training samples and extract the continuous frame optical flow feature vectors of the video training samples;

[0139] S102, determine the spatial motion entropy component, directional motion entropy component, and temporal motion entropy component based on the optical flow feature vector;

[0140] S103, Extract the scene change features of the video training samples;

[0141] S104, Based on the scene change characteristics, determine the feature distribution entropy component, semantic transformation entropy component, and visual complexity entropy component;

[0142] S105, determine the interaction terms between the spatial motion entropy component, directional motion entropy component, and temporal motion entropy component and the feature distribution entropy component, semantic transformation entropy component, and visual complexity entropy component;

[0143] S106, combine the spatial motion entropy component, the directional motion entropy component, the temporal motion entropy component, the feature distribution entropy component, the semantic transformation entropy component, the visual complexity entropy component, and the interaction term to generate an initial entropy value;

[0144] S107, Perform nonlinear transformation processing on the initial entropy value to generate a video entropy factor.

[0145] In this embodiment, the goal is to establish a time-sensitive and quantifiable complexity profile, enabling the training process to identify high-value segments and generate stable reward weights. First, the video training samples undergo temporal and spatial consistency processing: the frame rate is aligned to a preset sampling frequency, the dimensions are scaled proportionally to the shorter side and filled into the computational grid, pixel values ​​are normalized to a unified dynamic range, and the time axis is divided into windows of fixed duration or fixed frame number, with overlapping areas set for adjacent windows to preserve cross-window dependencies. Then, within each window, the optical flow feature vectors of consecutive frames are estimated. The optical flow is represented by a two-dimensional displacement and decomposed into amplitude and direction. If necessary, camera motion compensation is performed using global homography or affine estimation to reduce the interference of camera shake. To balance differentiability and robustness, the optical flow amplitude and direction are not directly hard-binned; instead, a kernel density smoothing or soft histogram strategy is used to generate a probability distribution, allowing subsequent entropy calculations based on the distribution to participate in gradient propagation.

[0146] Spatial motion entropy components characterize the discreteness of motion in spatial layout. The cumulative optical flow amplitude of a single or multiple frames is projected onto a predefined spatial grid. The amplitude probability is calculated and normalized in each grid cell, and the summation of the global grid distribution using the Shannon entropy formula yields a measure of spatial hierarchy complexity. Directional motion entropy components characterize directional diversity. Optical flow directions are mapped onto a unit circle and a soft distribution is formed with equal angular steps. A more uniform distribution indicates richer motion directions and a higher entropy value. Temporal motion entropy components focus on the uncertainty of amplitude changes over time. The amplitude sequence within a window is modeled using temporal kernel density or differential autocorrelation to obtain a temporal probability distribution and entropy measure. Robust statistics are used to suppress isolated spikes, and a lower threshold is introduced to avoid numerical degradation under extremely low motion conditions. The three motion entropy components can be calculated within the same window and aggregated temporally to ensure a response to both rapid and slow changes.

[0147] Scene change features are responsible for extracting perturbations in the content and appearance layers. To obtain the feature distribution entropy component, multi-scale keypoints or mid-level semantic features are extracted in each frame, and their density or activation intensity is projected onto a spatial grid to form a probability distribution, measuring the balance of information carried by different regions. The semantic transformation entropy component constructs a transition probability matrix by transferring the region-level category distribution along the time axis through frame-by-frame semantic annotation or regional category distribution. The entropy rate is calculated to measure the unpredictability of category switching, and the entry and exit event flow is maintained at the object level to improve sensitivity to occlusion and re-emergence. The visual complexity entropy component starts from texture and edge statistics, introducing indicators such as Laplacian energy and histogram of directional gradients. After obtaining the probability distribution using soft binning, the entropy is calculated. Scene illumination changes are mitigated through adaptive brightness correction and contrast layer normalization.

[0148] The interaction term is used to capture the sources of uncertainty driven by both motion and scene. It can be approximated by mutual information to measure the correlation between the optical flow amplitude distribution and the semantic transformation distribution, or it can use a symmetric combination of normalized covariance or two-way Kullback-Leibler divergence. To ensure numerical stability, the interaction term is temperature-scaled and Laplace-smoothed for both distributions before computation, and non-negativity constraints are imposed on the results. Within windows with large camera motion or global brightness abrupt changes, the interaction term is decayed with robust weights to avoid misjudgment.

[0149] The initial entropy values ​​are combined following a unified aggregation framework. Spatial motion entropy components, directional motion entropy components, temporal motion entropy components, feature distribution entropy components, semantic transformation entropy components, visual complexity entropy components, and interaction terms are mapped to the same scale and then weighted summation or weighted geometric aggregation is performed. The weights can be fixed or adaptively generated by window-level statistics to suit different content formats. To avoid the dominance of a single component, quantile normalization and outlier suppression are used before aggregation. The initial entropy values ​​obtained from aggregation need to be comparable across different videos, different durations, and different resolutions. Therefore, dual calibration at the dataset level and batch level is introduced. The former establishes a global scale through the mean and variance of the entire database or quantile intervals, while the latter fine-tunes the data through robust statistics of the current batch. The two are superimposed in the form of residuals.

[0150] Nonlinear transformation maps the initial entropy value to a bounded interval and improves resolution for medium-to-high complexity regions. The mapping can employ hyperbolic tangent or logistic functions, and the initial entropy value is standardized before mapping to ensure the transformed video entropy factor lies within the expected range and avoids excessive amplification of abnormally high values. Mapping temperature and translation parameters can be determined on the validation set using batch statistics or grid search. In addition to scalar form, the mapping output can also retain component dimensions to form a vector-type video entropy factor, allowing for differentiated scheduling based on different dimensions during subsequent training. To ensure usability during training, the entire implementation uses differentiable soft operators as much as possible, Gaussian kernel smoothing for the histogram distribution, lower-bound approximation for mutual information, and a smoothing threshold function instead of max-min pruning, thus enabling the video entropy factor to both express complexity and participate in end-to-end optimization.

[0151] This embodiment constructs motion entropy and scene entropy into probability distributions across multiple dimensions (space, direction, time, and semantics) and calculates the entropy measure in a differentiable form. It then introduces the interaction term between motion field and scene changes and performs scale-consistent aggregation and bounded mapping. This video entropy factor stably characterizes the temporal information density and content change intensity within a window. The factor remains comparable across different resolutions, durations, and content categories, and is robust to camera shake, lighting fluctuations, and short-term noise. During training, high-value segments are selected and assigned higher reward weights, focusing policy updates on real cross-temporal dependencies rather than single-frame shortcuts, thereby improving the learning efficiency and convergence stability of temporal understanding tasks.

[0152] In one embodiment, step S20 above includes:

[0153] S201, obtain the video entropy factor for the current training phase;

[0154] S202, determine the reward scaling factor based on the video entropy factor;

[0155] S203, Multiply the reward scaling factor by the base reward strength to generate the adjusted reward strength;

[0156] S204, determine the sample priority sequence based on the adjusted reward intensity;

[0157] S205, Select a subset of training data according to the sample priority sequence;

[0158] S206, Organize the subset of training data to form a training process.

[0159] In this embodiment, the goal is to transform the video entropy factor into an executable training scheduling signal, enabling the training process to focus on high-value segments over time and obtain stable reward-driven performance. First, the video entropy factor for the current training phase is obtained. The temporal granularity can be a window, segment, or entire segment, originating from the upstream entropy calculation link and entering the scheduling module in tensor or scalar form. To avoid cross-batch drift, distribution calibration is performed in the current training phase: intra-batch mean-variance standardization or quantile alignment is applied to the video entropy factor to obtain a centralized entropy index; a soft-threshold smoothing function is applied to extremely high values ​​to prevent subsequent scaling distortion. If necessary, a moving average cache is enabled, calculating an exponentially weighted average of the statistics from the most recent K batches to reduce oscillations caused by phase switching.

[0160] When determining the reward scaling factor based on the video entropy factor, a monotonically differentiable scaling mapping is used. Implementation paths can utilize temperature-controlled hyperbolic tangent, logical sigmoid, or piecewise linear-smooth splicing functions. Let the normalized entropy be E, then the reward scaling factor S can be constructed as S = 1 + α·tanh(E / τ), where α is the scaling magnitude and τ is the temperature. To ensure gradient penetration, all thresholding and cropping use smoothing substitutions (e.g., SoftPlus and SoftClip). The scaling factor needs to have lower and upper bounds (e.g., [Smin, Smax]) to ensure the reward space is controlled; and intra-group normalization is introduced to ensure that the average scaling of different videos within the same batch is 1, thereby maintaining a constant overall reward budget. To suppress occasional noise, first-order or second-order time regularization can be applied to S to limit scaling changes between adjacent iterations and avoid frequent priority inversions.

[0161] When multiplying the reward scaling factor by the baseline reward strength to generate the adjusted reward strength, the source and form of the baseline reward strength must first be clarified. The baseline reward strength R_base can come from environmental scores, rule scores, or offline teacher signals. To ensure consistency in additivity or multiplicativity across different sources, R_base is first mapped to a common dimension, and the influence of heavy-tailed distributions is eliminated through linear-quantile calibration or logarithmic domain processing. Subsequently, R_adj = S·R_base is executed, and numerical stabilization is applied to R_adj: a lower bound is set for excessively small values ​​to prevent the training signal from disappearing; excessively large values ​​are smoothly pruned to prevent gradient explosion; and a task weight matrix is ​​enabled for cross-task scenarios to maintain fairness of R_adj in multi-task mixed batches. To reduce variance, the reward can be centered in conjunction with the observed baseline b (R_adj←R_adj-b), but the relative priority order remains unchanged.

[0162] When determining the sample priority sequence based on the adjusted reward intensity, a rigorous sorting and scoring process is introduced. For each candidate sample, a priority score P is calculated. This can be simply taken as P = R_adj, or a multi-factor score P = w1·R_adj + w2·Coverage + w3·Novelty can be formed by combining difficulty, coverage, and novelty. Coverage can be measured by the feature cluster coverage rate, and novelty can be represented by the distance from the embedding space to the centroid of historical samples. To reduce sample congestion within the same video, a redundancy removal constraint is added: if adjacent segments are highly similar, only those with higher priority are retained. A stable sorting method is used, and hierarchical random perturbation is applied when scores are close to avoid training getting bogged down in the repeated use of a few samples. To ensure exposure for long-tail categories and rare shots, binning sampling is employed: first, the samples are divided into bins according to P, then proportionally sampled from the higher bins, while the lower bins retain the minimum quota, forming a sample priority sequence that balances efficiency and diversity.

[0163] When selecting a subset of training data based on sample priority sequence, the quota is dynamically calculated by considering batch size, memory budget, and latency target. In implementation, a priority queue or Top-k filtering is used to fill the first sample in the current batch or several subsequent batches. To improve throughput, multi-view strategies such as data augmentation and timing jitter can be enabled for high-priority samples to ensure diverse presentation of information from the same source. Simultaneously, the cumulative usage count and most recent usage time of each sample are recorded, and a maximum reuse count and cooldown interval are set to avoid oversampling leading to overfitting. For cross-device consistency, the sample priority sequence is synchronized globally via All-reduce in a distributed environment, ensuring that all training nodes use the same selection results. These high-priority items are prefetched in the data loader to reduce I / O wait times.

[0164] When organizing training data subsets into a training pipeline, the subsets are aligned by time, and their resolution and duration are standardized to generate batch structures for subsequent models. Time alignment can fix the frame window and allow for a small amount of overlap, while resolution standardization is achieved through short-side alignment and padding. To fully utilize the adjusted reward intensity, the R_adj tensor is stored within the batch as the weight input for subsequent loss weighting or reinforcement learning objectives, allowing the training pipeline to directly access this weight. For multi-task or multi-modal scenarios, the training pipeline includes task labels and modality masks, facilitating the execution of different decoding heads within the same batch. To ensure stable training and repeatable experiments, the training pipeline defines deterministic random seeds and queue snapshots; when resuming training, the priority sequence position and the state of consumed samples can be accurately restored. The structured batch object output throughout the organization phase includes: sample index, video frame tensor, timestamp, R_adj tensor, task label, and augmentation parameters, used for downstream forward propagation and policy updates.

[0165] This embodiment maps the video entropy factor to a bounded and smooth reward scaling factor and synthesizes it through robust multiplication with the baseline reward intensity under a unified dimension. This numerically constrains the training signal and statistically achieves explicit biasing of segments with high temporal complexity. Combined with a sample priority sequence constructed based on the adjusted reward intensity and scheduling mechanisms such as redundancy removal, bucketing, and cooling, the training data subset achieves a controllable balance between difficulty and coverage. The final training process carries per-sample weights and structured batch information, which reduces variance and improves sensitivity to cross-temporal dependencies in subsequent optimization stages, and achieves more stable convergence with fewer iterations.

[0166] In one embodiment, step S30 above includes:

[0167] S301, Extract video segment sequences from the training process;

[0168] S302, Arrange the video segment sequence in the original time order to generate ordered frame content;

[0169] S303, Randomly shuffle the temporal order of the video segment sequence to generate shuffled frame content;

[0170] S304, determine the first content hash value of the ordered frame content and the second content hash value of the shuffled frame content;

[0171] S305, compare the first content hash value with the second content hash value, and generate a consistency verification result;

[0172] S306, when the consistency verification result is that the content is consistent, the ordered frame content is input into the first inference branch of the training model to generate the ordered frame answer result, and the shuffled frame content is input into the second inference branch of the training model to generate the shuffled frame answer result.

[0173] In this embodiment, when extracting video segment sequences from the training process, continuous windows are preferentially established based on timestamps and frame counts. The window length and step size are fixed during the data loading phase to ensure consistent frame counts across windows within the same batch. To eliminate sampling jitter, a monotonicity check is first performed on the timestamps. If missing frames are found, they are resampled to a uniform frame rate using nearest neighbor interpolation or by resampling dropped frames. The video segment sequences are cached in tensor form, containing a frame pixel matrix, time index, and source identifier, for subsequent input construction.

[0174] When generating ordered frame content from a sequence of video clips arranged in their original chronological order, the sequences are strictly rearranged in ascending order based on their time indices. No transformations that alter pixel content are performed. Resolution alignment and color normalization are completed only in a common preprocessing channel and are fully shared across both types of input. To maintain temporal consistency, empty frames or mirroring can be used to pad sequence boundaries, ensuring consistent sequence lengths. After generating the ordered frame content, frame-level checksums are cached as a benchmark for subsequent consistency checks.

[0175] When generating scrambled frame content by randomly shuffling the temporal order of video clips, only the temporal index is permuted, without altering the individual frame pixel matrix or metadata. The permutation vector is generated by a pseudo-random engine capable of reproducible experiments, with the random seed jointly determined by the batch number and sample index, ensuring that permutations across different batches are independent and traceable. To avoid the accidental occurrence of "approximately ordered" results, a minimum misalignment distance threshold can be set to constrain the permutation results, ensuring that the new position of each frame deviates from the original position by at least several steps. After generating the scrambled frame content, a frame-level checksum is also output for subsequent comparison.

[0176] When determining the first content hash value of ordered frame content and the second content hash value of shuffled frame content, a robust verification mechanism is adopted while preserving the content. In implementation, a perceptual hash (such as pHash or dHash) can be calculated for each frame, and then the frame hashes are aggregated in chronological order using rolling hashing or Merkle to obtain a sequence-level hash. To improve robustness to slight decoding noise, frame hashes are pre-fixed with grayscale and scale normalization, without any data augmentation operations. The first content hash value comes from the aggregation result of ordered frame content, and the second content hash value comes from the aggregation result of shuffled frame content; both reflect only the frame content and are unrelated to the time index.

[0177] When generating a consistency verification result by comparing the hash values ​​of the first and second content, the hash vectors are first measured using Hamming distance or L2 distance. A distance below a threshold is considered consistent; a distance above the threshold triggers a reconstruction of the frame content or removes the sample from the next batch. To reduce false positives, the threshold is adaptively adjusted in conjunction with statistical measures such as video brightness variance, compression ratio, and frame rate. If necessary, a secondary check is added using the CRC of sampled frame pixels to improve the reliability of the consistency verification result. The consistency verification result is represented by a Boolean flag and a distance value for easier subsequent decision-making.

[0178] When the consistency verification result indicates content consistency, the first inference branch of the model is trained with ordered frame content to generate ordered frame answer results, while the second inference branch of the model is trained with shuffled frame content to generate shuffled frame answer results. The first and second inference branches share the underlying encoder weights and vocabulary; the branches are distinguished only by the temporal processing path or attention mask configuration, ensuring equivalence in computational resources and expressive power between the two paths. During inference, both types of inputs use the same preprocessing, temperature, and decoding strategies (such as greedy or bundle search) to avoid introducing additional bias due to decoding differences. To ensure answer comparability, the output format is uniformly set to text sequences and confidence distributions, with time overhead and memory usage balanced through parallel batch processing. If the two branches are executed in parallel on hardware, they must be completed on the same device family to avoid cross-device numerical drift affecting answer consistency. All intermediate tensors and logs are recorded dual-path by sample ID to ensure operability for subsequent comparisons and backtracking.

[0179] This embodiment uses a mechanism combining time index rearrangement and perceptual hash aggregation to align ordered and shuffled frame content at the pixel content level and strictly distinguish them at the time order level. The only difference in the source of the response results converges to the difference between the preservation and destruction of time-series information. Under the constraint of consistency verification results, any potential content drift samples are eliminated or reconstructed, so that the subsequent two-branch inference only reflects the difference in time-dependent capabilities. Two inference branches with shared parameters and equivalent configuration output ordered and shuffled frame response results, providing a standardized, reproducible, and traceable comparison baseline. This provides stable and low-biased data support for subsequent time-series performance gain calculations, thereby reducing interference from non-time factors and improving the resolution of cross-time inference capability assessment.

[0180] In one embodiment, step S40 above includes:

[0181] S401, obtain the ordered frame response result corresponding to the ordered frame content and the scrambled frame response result corresponding to the scrambled frame content;

[0182] S402, extract the confidence feature vector of the ordered frame response result and the confidence feature vector of the shuffled frame response result;

[0183] S403, determine the difference measure between the confidence feature vector of the ordered frame response result and the confidence feature vector of the shuffled frame response result;

[0184] S404, Input the difference metric value into a differentiable mapping function to generate the original gain value;

[0185] S405, Perform range constraint processing on the original gain value to generate timing performance gain.

[0186] In this embodiment, when obtaining ordered frame and shuffled frame answer results, the two outputs are first formatted and aligned to ensure consistency in text sequence, punctuation, capitalization, whitespace, units, and numerical expressions. If the task is a closed-set question-and-answer or multiple-choice decision, the log probability vector and normalized probability distribution of each candidate option are directly obtained. If the task is a generative output, the forward process is replayed with the same decoding configuration to collect the label-by-label log-likelihood. If necessary, a discriminant scorer or answer determiner is used to map the text answer to a discrete label space, forming a comparable multidimensional probability description. Both outputs retain the original logits and the softened distribution, and the temperature coefficient and length mask are recorded as inputs for subsequent extraction of confidence feature vectors.

[0187] When extracting the confidence feature vectors of ordered frame responses and shuffled frame responses, the feature dimensions and statistical methods are unified under the constraint of sample-level comparability. For closed-set scenarios, the candidate probability distributions are directly concatenated into a vector, with additional derived quantities such as the maximum class probability, the second largest class probability, probability gap, distribution entropy, Gini impurity, Top-k cumulative probability, and temperature-calibrated confidence. For generative scenarios, the sequence mean log-likelihood, length-normalized log-likelihood, mean and quantiles of label-by-label entropy, fragment consistency score, answer regularization matching probability, and numerical extraction consistency probability are used as components, while the answer posterior calibration score is incorporated into the feature dimensions. To suppress length effects and low-frequency fluctuations, a mask is applied to the label-by-label statistics, and temperature scaling and label smoothing are used, with a very small positive number added to avoid zero probability and numerical instability. After obtaining the two equal-dimensional confidence feature vectors, z-score or quantile scaling is performed to ensure consistent scale across batches.

[0188] When determining the difference metric between two vectors, the aligned vector pairs are used to calculate the metric, supporting differentiable metrics such as Euclidean distance, weighted cosine distance, symmetric KL divergence, Jensen-Shannon divergence, and Earth's distance. When the task contains multiple sub-objectives or multiple answer slots, the metric is first calculated within the slot and then aggregated according to the task weights to obtain a single difference metric. To enhance robustness, an isomorphic affine transformation can be applied to the two vectors before measurement to eliminate scaling differences; Huber truncation or Winsor transformation is used to suppress extreme effects on outlier samples; when the two answers are completely identical and the log-likelihood difference comes only from numerical noise, a minimum difference threshold is enabled to avoid spurious gains.

[0189] When inputting the difference metric into a differentiable mapping function to generate the original gain value, a family of functions that are monotonic and differentiable with respect to the input across the entire domain is selected to ensure trainability and gradient stability. Temperature-scaled forms of tanh, softsign, sigmoid-linear mixtures, or piecewise smooth soft-threshold functions can be used to map the difference metric to a non-negative interval or a pre-defined symmetric interval. To obtain smooth gradients and controllable sensitivity, a temperature coefficient is introduced to adjust the curvature, a slope coefficient is introduced to control local gain, and a soft threshold is set near zero to suppress the activation of gradients by small noise. Normalization is performed before mapping, for example, by standardizing the difference metric using the intra-batch mean and standard deviation, or by performing steady-state calibration using historical sliding statistics. The output after mapping serves as the original gain value, while retaining Jacobian information for backpropagation.

[0190] When performing range constraints on the original gain values ​​to generate temporal performance gains, a differentiable upper and lower bound constraint strategy is employed to avoid gradient truncation. A smooth pruning function can be used to limit the output to a set interval, or affine recalibration can be used to linearly map the original gain values ​​to the target interval, followed by the addition of temperature and task weights to complete the final calibration. When samples contain abnormally long sequences, extremely low confidence levels, or decoding failures, a mask is applied to place the corresponding gain at a minimum value to ensure training stability. After constraint completion, the temporal performance gain is output, and the difference metric, original gain value, and constrained gain are recorded in the sample metadata for easy backtracking and statistical analysis. The entire process runs under the same random seed and numerical precision environment to avoid systematic shifts introduced by implementation differences.

[0191] This embodiment extracts confidence feature vectors from ordered and shuffled frame responses under a unified standard, and uses differentiability to form a difference metric. Then, a temporal performance gain is obtained through a monotonically differentiable mapping function and smoothing range constraints. The training process yields a scalar signal that is sensitive to cross-temporal dependencies and has stable gradients. This signal originates from the difference in responses with the same content at different temporal intervals, suppressing interference from content factors and decoding strategies, and highlighting the substantial contribution of temporal structure to response quality. Differentiable mapping and smoothing constraints prevent gradient explosion and vanishing, allowing the gain to continue to play a role in policy updates, thereby reducing the model's tendency to rely on single-frame shortcuts and driving parameters towards convergence in a direction that truly utilizes temporal cues.

[0192] In one embodiment, step S50 above includes:

[0193] S501, Input the timing performance gain into the time correction function to generate a time correction term;

[0194] S502, Multiply the adjusted reward intensity by the time correction term to generate a composite reward value;

[0195] S503, Perform a differentiability verification on the synthesized reward value;

[0196] S504, when the repeatability verification passes, the synthetic reward value is used as a reward signal.

[0197] In this embodiment, when the time-series performance gain is input into the time correction function, the gain scalar or gain vector is first scaled and numerically robustned. Mean removal and variance normalization, or standardization using sliding statistics, are then performed to ensure that the gains of different batches and time periods are within a comparable range. Subsequently, a time correction function that is monotonic and differentiable across the entire domain is selected for mapping. The function form can be a temperature-scaled tanh form, a softsign function, a smoothed piecewise linear mapping, or a hybrid mapping of sigmoid and linear terms. Curvature and sensitivity are controlled by temperature parameters, local amplification is controlled by slope coefficients, and small-amplitude noise close to zero is suppressed by soft threshold parameters. To ensure differentiability and avoid saturation caused by extreme inputs, the gain is first limited to a reasonable pre-defined range before mapping is performed. Simultaneously, a Jacobian or an intermediate quantity recoverable by an automatic differentiation system is output in the forward propagation to obtain a continuous gradient during backpropagation. The resulting time correction term maintains the same batch and sample dimensions as the gain, facilitating subsequent element-wise or sample-wise broadcasting operations with the adjusted reward intensity.

[0198] When multiplying the adjusted reward strength with the time correction term to generate the composite reward value, alignment is performed according to the granularity of the training task. If the reward strength is a sample-level scalar and the time correction term is a sequence-level vector, a label-by-label broadcast is used, and a weighted average is calculated in the time dimension to obtain the sample-level composite value. If both are sample-level scalars, scalar multiplication is performed directly. If there are multiple sub-objectives or multiple answer slots, multiplication is performed within the slots first, and then aggregation is performed according to preset weights to obtain a single composite value. To suppress the influence of numerical instability and outliers, differentiable smooth pruning and scale recalibration are added before and after multiplication, and a soft clip function is used instead of hard truncation to prevent gradient interruption. In a mixed-precision training environment, a high-precision accumulator is used to store intermediate results of multiplication and aggregation to avoid the accumulation of rounding errors. When the reward strength or time correction term contains mask bits, it is ensured that the mask is passed throughout the multiplication and aggregation paths to avoid invalid positions contributing to the composite value.

[0199] When performing differentiability verification on the synthetic reward value, three dimensions are checked. First, static verification is performed to examine whether all operations in the mapping and multiplication paths are differentiable, confirming the absence of discrete branches, hard thresholds, or non-differentiable conditional gating. Second, numerical verification is performed to evaluate the range and distribution of the gradient norm, monitoring for gradient explosion, gradient vanishing, non-numerical, or infinite values. If anomalies are found, the process reverts to the most recent stable temperature and slope configuration and re-normalizes the input. Finally, consistency verification is performed by comparing the analytical gradient and the automatic differential gradient within tolerance using one forward and two backward passes. Gradient statistics are also sampled and recorded for subsequent adaptive parameter tuning. The verification process is executed using tensor batch processing, maintaining the same parallel granularity as the training backbone and avoiding the introduction of additional graph structure splits.

[0200] Once the differentiability verification is successful, the synthesized reward value is output as the reward signal to the policy optimization process. To improve stability and controllability, a differentiable interval remapping can be added before the output to ensure the reward signal falls within a specified range, while retaining a slight mean drift to match the policy update expectation. Optionally, a safety term weighted with a very small coefficient can be superimposed so that the reward still carries a weak but consistent training drive on extremely degraded samples. The output is simultaneously written to the training log and sample metadata, recording the gain, correction term, synthesized value, and its gradient statistics, facilitating the tracking of drift over time and closed-loop adjustment of hyperparameters such as temperature, slope, soft threshold, and remapping interval.

[0201] This embodiment performs monotonically differentiable time-corrected mapping on the temporal performance gain and modulates the adjusted reward intensity using a multiplicative method to form a synthetic reward value that retains the basic reward scale while highlighting the time-dependent contribution. Differentiability verification and smoothing constraints are performed before output, ensuring that the reward signal provides a continuous, stable, and directional gradient during policy updates. This suppresses training oscillations and gradient anomalies caused by discrete thresholds or scale mismatches, improves sensitivity and discriminability across temporal cues, concentrates parameter updates on sample regions that enhance temporal understanding, reduces reliance on single-frame shortcuts, and achieves higher convergence efficiency and more reliable generalization performance with the same training budget.

[0202] In one embodiment, step S60 above includes:

[0203] S601, determine the policy gradient update amount based on the reward signal;

[0204] S602, adjust the current policy parameters using the policy gradient update amount, and generate updated policy parameters;

[0205] S603, Load the updated strategy parameters into the model inference engine;

[0206] S604, Extract video segment data from the training process;

[0207] S605, The video segment data is processed by the model inference engine to generate semantic understanding results;

[0208] S606, the semantic understanding result is encapsulated into a training output data packet.

[0209] In this embodiment, when determining the policy gradient update amount based on the reward signal, the objective function increment is constructed using the sample-by-sample reward signal within the training batch as the weight driving term. First, the reward signal is debiased and scaled. Weights with zero mean and unit variance are obtained by subtracting the mean from the standard deviation within the batch, thus reducing drift between different batches. For abnormally large reward signals, a differentiable soft-pruning function is used for smoothing constraints to avoid uncontrollable gradient peaks during backpropagation. To improve variance utilization efficiency, a baseline term is introduced. The baseline can be provided by the value estimation head, moving average, or prior heuristic. The update amount is calculated as the difference between the reward signal and the baseline, ensuring the update direction is consistent with the relative advantage. The normalized advantage weights are multiplied element-by-element by the log-likelihood gradient of the policy output and summed along the batch dimension to obtain the policy gradient update amount. To avoid gradient vanishing or exploding, norm pruning is performed on this update amount, and the first and second moment statistics are recorded for adaptive step size calculation in the downstream update rule. The computation process is completed within the same computation graph, maintaining consistent device, accuracy, and parallel partitioning with forward inference, and avoiding additional graph switching overhead.

[0210] When adjusting the current policy parameters using the policy gradient update, the updated policy parameters are generated by first reading the tensor view of the current policy parameters along with their momentum and second-order statistical cache, and then numerically tuning them using a hyperparameter set consisting of the learning rate, weight decay coefficient, and gradient scaling factor. The update amount and learning rate are scaled parameter by parameter and then superimposed onto the current policy parameters. If a momentum and second-order adaptive update rule is used, the momentum and second-order terms must be updated first, and then the effective step size and bias correction term for this round are calculated and used for the calculation and superposition of parameter increments. For parameters participating in regularization, a decoupled weight decay term is added to reduce the risk of overfitting. For parameters in the embedding matrix or normalization layer that are not expected to be strongly regularized, a mask is used to ensure that they only accept gradient-driven input and do not have decay components superimposed. After the update, parameter validity is checked, including numerical domain checks, non-numerical and infinite value cleanup, and optional parameter quantization-aware correction, to ensure the stability of subsequent inference.

[0211] When loading the updated policy parameters into the model inference engine, parameter broadcasting and shard reconstruction are performed according to the parallel partitioning strategy of the training cluster. In data-parallel scenarios, parameters are synchronized across processes; in tensor-parallel or pipeline-parallel scenarios, parameters are partitioned by dimension and written to the corresponding device memory according to the stage topology, maintaining a unified checkpoint version number to avoid cross-device version inconsistencies. If mixed precision is used, the primary precision parameters are first written to main memory, and then a copy of the computational precision required for inference is generated through a precision converter, while binding the scaling factor and loss scale to ensure the numerical stability of the forward computation. After completing the hot loading of parameters, the inference engine performs an empty forward pass to cache and optimize the key operators, so that subsequent data processing is not affected by initialization jitter.

[0212] When extracting video segment data from the training process, the previously constructed sample priority sequence and training data subset are read. This subset is then segmented into video segment sequences according to batch and duration boundaries. To ensure consistency with the timing test phase, the timestamp continuity and frame content integrity of the segments are verified, and necessary masks, time position encodings, and segment identifiers are appended to the segment metadata. If the training process includes data augmentation strategies, only augmentations that do not change the semantic temporal relationship are allowed at this stage, such as brightness scale fine-tuning or mild noise injection, to avoid target offset caused by changes in temporal order or content. All segment data is converted to a tensor format acceptable to the engine, and frame size, channel order, and memory alignment are standardized. In-batch indexes are also built for result backfilling.

[0213] When processing video segment data through the model inference engine to generate semantic understanding results, forward propagation is performed according to the updated policy parameters. Video frames first enter the visual encoding path to extract spatial representations, then encode cross-frame dependencies through a temporal modeling unit, and subsequently fuse with language-side context or task cues within a multimodal alignment unit to obtain a semantic sequence representation. The generation end outputs response text, action distribution, or intermediate judgment labels based on the training objective, while retaining intermediate quantities such as log-likelihood, attention weights, and temperature coefficients for upstream gain calculation or downstream observability analysis. The inference process follows a hybrid batch and streaming strategy. In long segment scenarios, the sequence is split into sliding windows and overlapped to ensure that the contextual continuity across windows is not disrupted. The generated semantic understanding results are temporarily stored in an indexable structure, with sample identifiers, timestamps, and version numbers to ensure traceability of the association with rewards and gains.

[0214] When encapsulating semantic understanding results into training output data packets, a structured carrier is created for each batch, containing the semantic understanding result primary key, text or label output, log-likelihood component, optional intermediate attention summary, and parameter version number and gain snapshot associated with the current update. Text output exceeding a set length is safely truncated, and the truncation position is recorded to avoid misleading subsequent evaluations. For tasks with multiple answer slots, packets are packaged according to slot order and a mask to ensure consistency in the order of downstream evaluation and replay reproduction. After generation, the data packets are written to a persistent buffer and event log, forming a unified exit for subsequent evaluation, replay, comparison, and diagnosis. When needed, they can be directly used in the next loop for computational paths related to answer result comparison and temporal gain estimation, maintaining the consistency and repeatability of the entire closed loop.

[0215] Example Explanation: In a video intelligent analysis system within the healthcare business domain, a comprehensive entropy-aware reinforcement training mechanism can be employed to enhance the model's temporal reasoning capabilities. The system first collects video training samples from a health management platform. These samples may originate from daily health and wellness scenarios for the elderly, home rehabilitation exercise monitoring, or remote health guidance processes for individuals with chronic diseases. To characterize the motion complexity and scene changes within the video content, the system extracts optical flow feature vectors between consecutive frames and calculates spatial motion entropy, directional motion entropy, and temporal motion entropy to characterize the amplitude, directionality, and rhythmic stability of body movements. Simultaneously, environmental change information is extracted from the video's scene features, calculating feature distribution entropy, semantic transformation entropy, and visual complexity entropy to reflect changes in object position adjustments, the use of rehabilitation aids, and background interference factors in home scenarios. These motion and scene entropy components are combined after calculating the interaction term to form an initial entropy value, which is then used to obtain a video entropy factor through nonlinear mapping. A higher video entropy factor indicates a higher requirement for the model's temporal understanding of the video sample, thus receiving greater attention and weight during the training process.

[0216] In organizing the training process, the system first calculates the reward scaling factor based on the video entropy factor generated at the current stage, and multiplies it by the baseline reward intensity to obtain the adjusted reward intensity. The reward intensity value is used to prioritize training samples, ensuring that high-entropy video segments are selected into the training data subset, guaranteeing that the model training phase encounters more complex videos with higher requirements for temporal inference. The selected training data subset is integrated into a training process, in which ordered frame content and shuffled frame content are constructed. The ordered frame content retains the original temporal order of the video segments, while the shuffled frame content randomly changes the temporal arrangement of the segments. The system generates content hash values ​​for both types of inputs and compares them to ensure that no data is missing or content errors occur during shuffling. After the consistency verification passes, the ordered frame content is sent to the first inference branch of the training model to generate ordered frame response results; the shuffled frame content is sent to the second inference branch to generate shuffled frame response results.

[0217] Subsequently, the system compares the two response results, extracts their respective confidence feature vectors, and calculates the difference metric. This difference metric is then mapped to a differentiable function to generate the original gain value, and after applying range constraints, the temporal performance gain is obtained. The magnitude of the temporal performance gain reflects the degree of difference between the model's inference results under time-order preservation and random-order conditions, directly measuring the model's temporal inference level. This gain is further combined with the adjusted reward intensity. First, a correction term is generated using a time correction function, and then multiplied by the reward intensity to obtain the composite reward value. After passing differentiability verification, the composite reward value is used as the reward signal.

[0218] The reward signal is input into the policy optimization process. The system calculates the gradient update amount, which is used to adjust the current policy parameters and generate updated policy parameters. The updated parameters are loaded into the model inference engine, driving the engine to reprocess the video segment data in the training process and generate semantic understanding results. These results include not only predictions of action categories or rehabilitation action completion, but also the fluency of action completion and abnormal action detection results. The semantic understanding results are packaged into training output data packages and enter subsequent data storage and tracking stages. Based on these training outputs, the model further updates its structure and parameters, thereby continuously improving its ability to understand the temporal logic in complex videos through iterative processes.

[0219] The entire training process is executed iteratively, repeating reward intensity adjustment, temporal testing, gain determination, signal generation, policy optimization, process handling, and model updates until the training termination condition is met. Ultimately, the resulting target model can directly handle target tasks in healthcare scenarios, such as analyzing videos of elderly people's daily exercises to generate results on movement accuracy and potential fall risk, or generating reports on movement completion and progress trends in rehabilitation training. Through this continuous reward-driven and entropy-aware mechanism, the model not only improves the robustness of temporal inference but also demonstrates higher generalization ability in diverse home health management scenarios.

[0220] In the fintech field, entropy-aware reinforcement learning training mechanisms can be applied to complex tasks such as transaction monitoring, compliance review, and user behavior risk modeling. The system first acquires a large number of video or temporal training samples, which may originate from transaction visualization replays, ATM operation records, or remote wealth management interaction data. To characterize the complexity contained in these samples, the system extracts optical flow feature vectors from continuous segments and calculates spatial motion entropy, directional motion entropy, and temporal motion entropy to measure mouse trajectories, interface switching speeds, and interaction continuity during transaction operations. Simultaneously, it extracts scene change features and calculates feature distribution entropy, semantic transformation entropy, and visual complexity entropy to reflect fluctuations in transaction page transitions, form filling steps, and interface element loading. Each entropy component and its interaction are combined into an initial entropy value, which is then transformed nonlinearly to generate a video entropy factor to quantify the complexity of task segments and training priority.

[0221] Based on the video entropy factor, the system dynamically adjusts the reward intensity, calculates the reward scaling factor, and combines it with the baseline reward intensity to obtain a new reward value. This reward value is used to establish a priority sequence for training samples, prioritizing operation segments with high complexity and demanding temporal inference requirements, thus forming the training process. During the training process, the system constructs the sample sequence into ordered and shuffled frame contents, generates response results through the model's two inference branches, and uses hash verification to ensure data consistency. Subsequently, the system compares the two types of response results, extracts confidence feature vectors, calculates the difference metric, and obtains the temporal performance gain through differentiable function mapping and range constraints. The temporal performance gain is combined with the adjusted reward intensity and used to generate a synthetic reward value via a time correction function. After verifying differentiability, this value serves as the final reward signal.

[0222] The reward signal is fed into the policy optimization process to calculate gradient updates and update policy parameters. The updated parameters are then loaded into the model inference engine. The inference engine uses the new parameters to process fragmented data in the training process, and the semantic understanding results are packaged into training output data packets for further model updates. The training process is executed cyclically until a termination condition is met, resulting in the target model. The target model is capable of performing complex cross-temporal inference tasks in financial scenarios, such as identifying abnormal patterns in multi-step transactions, discovering potential chains of fraudulent behavior, or assessing the stability and compliance of user interactions. The final model can not only identify static single-point anomalies but also understand potential risk logic across time spans, significantly enhancing the accuracy and robustness of fintech systems in complex transactions and compliance reviews.

[0223] This embodiment constructs and adjusts parameters using a gradient update process driven by reward signals, ensuring that the direction and magnitude of parameter updates closely align with the training value at the sample level. Through parameter hot-loading and graph optimization, the inference phase can reflect the updated behavior in real time. By extracting and normalizing video segment data from the training process and using it for subsequent forward inference, the output semantic understanding results are more stable in the temporal dimension and seamlessly connect with previous gain and reward paths. This achieves a continuous "reward-update-inference-output" chain, reducing the probability of oscillations caused by the decoupling of rewards and parameter updates, improving the verifiability and traceability of each update for subsequent inference quality, shortening the training rounds required to reach the target performance, and obtaining more robust temporal understanding performance under the same computational budget.

[0224] In one embodiment, a training and processing apparatus based on time-series performance gain is provided, which corresponds one-to-one with the training and processing method based on time-series performance gain in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the training and processing device based on time-series performance gain of the present invention. The modules include: entropy factor generation module 10, reward regulation module 20, input construction and inference module 30, time-series gain evaluation module 40, reward signal generation module 50, policy optimization module 60, model update module 70, loop control module 80, and task execution module 90. Detailed descriptions of each functional module are as follows:

[0225] Entropy factor generation module 10 is used to acquire video training samples and generate video entropy factors based on the motion entropy and scene entropy of the video training samples.

[0226] Reward regulation module 20 is used to adjust the reward intensity and organize the training process based on the video entropy factor;

[0227] The input construction and reasoning module 30 is used to construct ordered frame content and shuffled frame content based on the training process, and to generate answer results based on the ordered frame content and shuffled frame content through the training model;

[0228] The timing gain evaluation module 40 is used to determine the timing performance gain based on the comparison of the response results;

[0229] The reward signal generation module 50 is used to generate a reward signal based on the adjusted reward intensity and the timing performance gain.

[0230] The strategy optimization module 60 is used to input the reward signal into the strategy optimization process to update the strategy parameters, and process the training process based on the updated strategy parameters to generate training output.

[0231] Model update module 70 is used to update the training model based on the training output;

[0232] The loop control module 80 is used to repeatedly execute the reward intensity adjustment step, the timing test step, the gain determination step, the signal generation step, the policy update step, the process processing step, and the model update step until the training termination condition is met and the target model is obtained.

[0233] The task execution module 90 is used to process the target task through the target model and generate the target task result.

[0234] In one embodiment, the entropy factor generation module 10 is specifically used for:

[0235] Obtain the video training samples and extract the continuous frame optical flow feature vectors of the video training samples;

[0236] Based on the optical flow feature vector, determine the spatial motion entropy component, the directional motion entropy component, and the temporal motion entropy component;

[0237] Extract scene change features from the video training samples;

[0238] Based on the scene change characteristics, the feature distribution entropy component, semantic transformation entropy component, and visual complexity entropy component are determined.

[0239] Determine the interaction terms between the spatial motion entropy component, directional motion entropy component, and temporal motion entropy component and the feature distribution entropy component, semantic transformation entropy component, and visual complexity entropy component;

[0240] The initial entropy value is generated by combining the spatial motion entropy component, the directional motion entropy component, the temporal motion entropy component, the feature distribution entropy component, the semantic transformation entropy component, the visual complexity entropy component, and the interaction term.

[0241] A nonlinear transformation process is performed on the initial entropy value to generate a video entropy factor.

[0242] In one embodiment, the reward regulation module 20 is specifically used for:

[0243] Obtain the video entropy factor for the current training phase;

[0244] The reward scaling factor is determined based on the video entropy factor.

[0245] The adjusted reward intensity is generated by multiplying the reward scaling factor by the base reward intensity.

[0246] Determine the sample priority sequence based on the adjusted reward intensity;

[0247] Select a subset of training data according to the aforementioned sample priority sequence;

[0248] Organize the aforementioned subset of training data to form a training process.

[0249] In one embodiment, the input construction and reasoning module 30 is specifically used for:

[0250] Extract video segment sequences from the training process;

[0251] The video segment sequence is arranged in its original chronological order to generate ordered frame content;

[0252] Randomly shuffle the temporal order of the video segment sequence to generate shuffled frame content;

[0253] Determine the first content hash value of the ordered frame content and the second content hash value of the shuffled frame content;

[0254] Compare the hash value of the first content with the hash value of the second content to generate a consistency verification result;

[0255] When the consistency verification result is that the content is consistent, the ordered frame content is input into the first inference branch of the training model to generate the ordered frame answer result, and the shuffled frame content is input into the second inference branch of the training model to generate the shuffled frame answer result.

[0256] In one embodiment, the timing gain evaluation module 40 is specifically used for:

[0257] Obtain the ordered frame response results corresponding to the ordered frame content and the shuffled frame response results corresponding to the shuffled frame content;

[0258] Extract the confidence feature vector of the ordered frame response and the confidence feature vector of the shuffled frame response;

[0259] Determine the difference measure between the confidence feature vector of the ordered frame response and the confidence feature vector of the shuffled frame response;

[0260] The difference metric is input into a differentiable mapping function to generate the original gain value;

[0261] The original gain value is subjected to range constraint processing to generate timing performance gain.

[0262] In one embodiment, the reward signal generation module 50 is specifically used for:

[0263] The timing performance gain is input into the time correction function to generate a time correction term;

[0264] Multiply the adjusted reward intensity by the time correction term to generate a composite reward value;

[0265] Perform a differentiability verification on the synthesized reward value;

[0266] When the doubling verification passes, the synthesized reward value is used as a reward signal.

[0267] In one embodiment, the strategy optimization module 60 is specifically used for:

[0268] The policy gradient update amount is determined based on the reward signal;

[0269] The current policy parameters are adjusted using the policy gradient update amount to generate updated policy parameters;

[0270] Load the updated policy parameters into the model inference engine;

[0271] Extract video segment data from the training process;

[0272] The video segment data is processed by the model inference engine to generate semantic understanding results;

[0273] The semantic understanding results are encapsulated into training output data packets.

[0274] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides deterministic and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps of a training and processing method based on timing performance gains.

[0275] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements user-side functions or steps of a training and processing method based on timing performance gains.

[0276] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0277] Obtain video training samples, and generate a video entropy factor based on the motion entropy and scene entropy of the video training samples;

[0278] The reward intensity is adjusted and the training process is organized based on the video entropy factor;

[0279] Based on the training process, ordered frame content and shuffled frame content are constructed, and the answer result is generated based on the ordered frame content and the shuffled frame content through the training model.

[0280] The timing performance gain is determined based on the comparison of the response results.

[0281] A reward signal is generated based on the adjusted reward intensity and the time-series performance gain;

[0282] The reward signal is input into the policy optimization process to update the policy parameters, and the training process is processed based on the updated policy parameters to generate training output;

[0283] Update the training model based on the training output;

[0284] Repeat the reward intensity adjustment step, timing test step, gain determination step, signal generation step, policy update step, process processing step, and model update step until the training termination condition is met and the target model is obtained.

[0285] The target task is processed through the target model to generate the target task result.

[0286] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0287] Obtain video training samples, and generate a video entropy factor based on the motion entropy and scene entropy of the video training samples;

[0288] The reward intensity is adjusted and the training process is organized based on the video entropy factor;

[0289] Based on the training process, ordered frame content and shuffled frame content are constructed, and the answer result is generated based on the ordered frame content and the shuffled frame content through the training model.

[0290] The timing performance gain is determined based on the comparison of the response results.

[0291] A reward signal is generated based on the adjusted reward intensity and the time-series performance gain;

[0292] The reward signal is input into the policy optimization process to update the policy parameters, and the training process is processed based on the updated policy parameters to generate training output;

[0293] Update the training model based on the training output;

[0294] Repeat the reward intensity adjustment step, timing test step, gain determination step, signal generation step, policy update step, process processing step, and model update step until the training termination condition is met and the target model is obtained.

[0295] The target task is processed through the target model to generate the target task result.

[0296] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0297] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0298] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0299] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A training and processing method based on time-series performance gain, characterized in that, Includes the following steps: Obtain video training samples, and generate a video entropy factor based on the motion entropy and scene entropy of the video training samples; The reward intensity is adjusted and the training process is organized based on the video entropy factor; Based on the training process, ordered frame content and shuffled frame content are constructed, and the answer result is generated based on the ordered frame content and the shuffled frame content through the training model. The timing performance gain is determined based on the comparison of the response results. A reward signal is generated based on the adjusted reward intensity and the time-series performance gain; The reward signal is input into the policy optimization process to update the policy parameters, and the training process is processed based on the updated policy parameters to generate training output; Update the training model based on the training output; Repeat the reward intensity adjustment step, timing test step, gain determination step, signal generation step, policy update step, process processing step, and model update step until the training termination condition is met and the target model is obtained. The target task is processed through the target model to generate the target task result.

2. The training and processing method based on time-series performance gain as described in claim 1, characterized in that, Obtain video training samples, and generate a video entropy factor based on the motion entropy and scene entropy of the video training samples, including: Obtain the video training samples and extract the continuous frame optical flow feature vectors of the video training samples; Based on the optical flow feature vector, determine the spatial motion entropy component, the directional motion entropy component, and the temporal motion entropy component; Extract scene change features from the video training samples; Based on the scene change characteristics, the feature distribution entropy component, semantic transformation entropy component, and visual complexity entropy component are determined. Determine the interaction terms between the spatial motion entropy component, directional motion entropy component, and temporal motion entropy component and the feature distribution entropy component, semantic transformation entropy component, and visual complexity entropy component; The initial entropy value is generated by combining the spatial motion entropy component, the directional motion entropy component, the temporal motion entropy component, the feature distribution entropy component, the semantic transformation entropy component, the visual complexity entropy component, and the interaction term. A nonlinear transformation process is performed on the initial entropy value to generate a video entropy factor.

3. The training and processing method based on time-series performance gain as described in claim 1, characterized in that, Adjusting the reward intensity and organizing the training process based on the video entropy factor includes: Obtain the video entropy factor for the current training phase; The reward scaling factor is determined based on the video entropy factor. The adjusted reward intensity is generated by multiplying the reward scaling factor by the base reward intensity. Determine the sample priority sequence based on the adjusted reward intensity; Select a subset of training data according to the aforementioned sample priority sequence; Organize the aforementioned subset of training data to form a training process.

4. The training and processing method based on time-series performance gain as described in claim 1, characterized in that, Based on the aforementioned training process, ordered frame content and shuffled frame content are constructed. A training model is then used to generate response results based on the ordered frame content and the shuffled frame content, including: Extract video segment sequences from the training process; The video segment sequence is arranged in its original chronological order to generate ordered frame content; Randomly shuffle the temporal order of the video segment sequence to generate shuffled frame content; Determine the first content hash value of the ordered frame content and the second content hash value of the shuffled frame content; Compare the hash value of the first content with the hash value of the second content to generate a consistency verification result; When the consistency verification result is that the content is consistent, the ordered frame content is input into the first inference branch of the training model to generate the ordered frame answer result, and the shuffled frame content is input into the second inference branch of the training model to generate the shuffled frame answer result.

5. The training and processing method based on time-series performance gain as described in claim 1, characterized in that, Based on the comparison of the response results, the timing performance gain is determined, including: Obtain the ordered frame response results corresponding to the ordered frame content and the shuffled frame response results corresponding to the shuffled frame content; Extract the confidence feature vector of the ordered frame response and the confidence feature vector of the shuffled frame response; Determine the difference measure between the confidence feature vector of the ordered frame response and the confidence feature vector of the shuffled frame response; The difference metric is input into a differentiable mapping function to generate the original gain value; The original gain value is subjected to range constraint processing to generate timing performance gain.

6. The training and processing method based on time-series performance gain as described in claim 1, characterized in that, Generate a reward signal based on the adjusted reward intensity and the time-series performance gain, including: The timing performance gain is input into the time correction function to generate a time correction term; Multiply the adjusted reward intensity by the time correction term to generate a composite reward value; Perform a differentiability verification on the synthesized reward value; When the doubling verification passes, the synthesized reward value is used as a reward signal.

7. The training and processing method based on time-series performance gain as described in claim 1, characterized in that, The reward signal is input into the policy optimization process to update the policy parameters, and the training process is processed based on the updated policy parameters to generate training output, including: The policy gradient update amount is determined based on the reward signal; The current policy parameters are adjusted using the policy gradient update amount to generate updated policy parameters; Load the updated policy parameters into the model inference engine; Extract video segment data from the training process; The video segment data is processed by the model inference engine to generate semantic understanding results; The semantic understanding results are encapsulated into training output data packets.

8. A training and processing device based on time-series performance gain, characterized in that, The training and processing device based on time-series performance gain includes: The entropy factor generation module is used to acquire video training samples and generate video entropy factors based on the motion entropy and scene entropy of the video training samples. The reward control module is used to adjust the reward intensity and organize the training process based on the video entropy factor. The input construction and reasoning module is used to construct ordered frame content and shuffled frame content based on the training process, and to generate answer results based on the ordered frame content and shuffled frame content through the training model; The timing gain evaluation module is used to determine the timing performance gain based on the comparison of the response results; A reward signal generation module is used to generate a reward signal based on the adjusted reward intensity and the timing performance gain. The strategy optimization module is used to input the reward signal into the strategy optimization process to update the strategy parameters, and to process the training process based on the updated strategy parameters and generate training output. The model update module is used to update the training model based on the training output; The loop control module is used to repeatedly execute the reward intensity adjustment step, the timing test step, the gain determination step, the signal generation step, the policy update step, the process processing step, and the model update step until the training termination condition is met and the target model is obtained. The task execution module is used to process the target task through the target model and generate the target task result.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a timing performance gain-based training and processing program stored in the memory and executable on the processor. When executed by the processor, the timing performance gain-based training and processing program implements the steps of the timing performance gain-based training and processing method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a training and processing program based on timing performance gain, which, when executed by a processor, implements the steps of the training and processing method based on timing performance gain as described in any one of claims 1-7.

Citation Information

Cited By

  • Inference model training method and device, electronic equipment, medium and product

    CN121257757A

  • Training methods, devices, electronic equipment, media, and products for inference models

    CN121257757B

  • Decision model training method, live broadcast decision determination method and device

    CN121585837A