Method and system for continuous playing of short videos based on multi-modal data

By constructing multimodal data and combining it with user characteristics, precise continuous playback instructions are generated, which solves the problem of inaccurate playback caused by differences in relevance of short videos, and realizes precise control of short videos and overall control of emotional characteristics.

CN121603723BActive Publication Date: 2026-03-31SHANGHAI FULEIDE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-27
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing technologies, when short videos are played continuously, the system cannot achieve precise continuous playback control because the previous short video and the next short video have different relevance dimensions.

Method used

Multimodal data is constructed by labeling the visual, audio, and text streams of short videos. Combined with users' historical viewing data and real-time context information, users' viewing preferences and key video segments are determined. Continuous playback instructions are generated based on emotional features and preset playback order. Multidimensional behavioral features are collected during continuous playback events to determine the final operation instructions.

Benefits of technology

It enables precise control over short videos, improves the accuracy of continuous playback commands and operation commands, and ensures full consideration of different dimensions and overall control of emotional characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121603723B_ABST
    Figure CN121603723B_ABST
Patent Text Reader

Abstract

The application discloses a kind of short video continuous playing method and system based on multi-modal data, and the application relates to the technical field of short video, corresponding multi-modal data is constructed according to visual flow, audio stream and text flow, and key video segment is determined in the matching process of hobby watching feature and multi-modal data, corresponding emotional characteristics are determined based on each key video segment and corresponding user viewing behavior, corresponding continuous playing instruction is determined according to emotional characteristics, key video segment corresponding short video and the preset playing order of multiple short videos, the accuracy of continuous playing instruction is improved, different dimensions are fully considered, and the primary operation instruction is determined according to the multi-dimensional behavior characteristics;In the same time dimension, the playing picture of the continuous playing event of the primary operation instruction and short video is matched, to determine the corresponding matching level, the final operation instruction is determined according to the matching level and the viewing history of user, and the accurate regulation of short video is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of short videos, and in particular to a method and system for continuous playback of short videos based on multimodal data. Background Technology

[0002] With the development of technology, short videos are gradually being applied to people's lives. Multiple short videos are combined for continuous playback. When users operate short video software, they will switch between the previous short video and the next short video. However, there are differences between the previous short video and the next short video in terms of relevance. The system designs a preset playback order for multiple short videos and plays them in a single dimension along the preset playback order, which affects the accuracy of continuous playback instructions and makes it impossible to accurately control the short videos. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a method and system for continuous playback of short videos based on multimodal data.

[0004] This invention provides a method for continuous playback of short videos based on multimodal data, including:

[0005] Multiple short videos are labeled, and the corresponding visual stream, audio stream, and text stream are determined based on the recognition of each short video. The corresponding multimodal data is then constructed based on the visual stream, audio stream, and text stream.

[0006] Collect users' historical viewing data, determine users' viewing preferences based on this historical viewing data and real-time context information, and identify key video segments in the process of matching viewing preferences with multimodal data;

[0007] Based on each key video segment and the corresponding user viewing behavior, determine the corresponding emotional characteristics, and determine the corresponding continuous playback instructions based on the emotional characteristics, the short videos corresponding to the key video segments, and the preset playback order of multiple short videos.

[0008] When a continuous playback command is triggered, a corresponding continuous playback event of a short video is constructed, and multi-dimensional behavioral characteristics of the user in the continuous playback event of the short video are collected. The primary operation command is determined based on these multi-dimensional behavioral characteristics.

[0009] Match the initial operation command and the playback screen of the short video's continuous playback event within the same time dimension to determine the corresponding matching level, and determine the final operation command based on the matching level and the user's viewing history.

[0010] This invention provides a continuous playback system for short videos based on multimodal data, which is applied to the aforementioned continuous playback method for short videos based on multimodal data.

[0011] Compared with the prior art, the beneficial effects of the present invention are:

[0012] (1) Label multiple short videos and determine the corresponding visual stream, audio stream and text stream based on the recognition of each short video. Construct corresponding multimodal data based on the visual stream, audio stream and text stream. Collect users' historical viewing data and determine users' viewing preferences based on the historical viewing data and real-time context information. In the process of matching viewing preferences and multimodal data, key video segments are determined. The visual stream, audio stream and text stream of short videos are introduced to further control the multimodal data. The precise control of key video segments is achieved by combining viewing preferences and multimodal data.

[0013] (2) Based on each key video segment and the corresponding user viewing behavior, the corresponding emotional characteristics are determined. Based on the emotional characteristics, the short videos corresponding to the key video segments and the preset playback order of multiple short videos, the corresponding continuous playback instructions are determined. This achieves overall control of the emotional characteristics, the short videos corresponding to the key video segments and the preset playback order of multiple short videos, and improves the accuracy of the continuous playback instructions, thus achieving full consideration of different dimensions.

[0014] (3) Construct the corresponding short video continuous playback event under the trigger of the continuous playback command, and collect the multi-dimensional behavioral features of the user in the continuous playback event of the short video. Determine the primary operation command based on the multi-dimensional behavioral features. Match the primary operation command and the playback screen of the short video continuous playback event in the same time dimension to determine the corresponding matching level. Determine the final operation command based on the matching level and the user's viewing history. Recognize the primary operation command and combine the matching level and the user's viewing history to trigger the optimization of the primary operation command, so as to improve the accuracy of the final operation command and realize the precise control of the short video. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the continuous playback method for short videos based on multimodal data in an embodiment of the present invention.

[0016] Figure 2 This is a flowchart illustrating step S11 in the continuous playback method of short videos based on multimodal data in an embodiment of the present invention.

[0017] Figure 3 This is a flowchart illustrating step S12 in the continuous playback method of short videos based on multimodal data in an embodiment of the present invention.

[0018] Figure 4 This is a flowchart illustrating step S13 in the continuous playback method of short videos based on multimodal data in an embodiment of the present invention.

[0019] Figure 5 This is a flowchart illustrating step S14 in the continuous playback method of short videos based on multimodal data in an embodiment of the present invention.

[0020] Figure 6 This is a flowchart illustrating step S15 of the continuous playback method for short videos based on multimodal data in an embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0022] Please see Figures 1 to 6 A method for continuous playback of short videos based on multimodal data, applied to short video scenarios; the method for continuous playback of short videos based on multimodal data includes:

[0023] Step S11: Label multiple short videos and determine the corresponding visual stream, audio stream and text stream based on the recognition of each short video, and construct the corresponding multimodal data based on the visual stream, audio stream and text stream;

[0024] Step S12: Collect the user's historical viewing data, determine the user's preferred viewing characteristics based on the historical viewing data and real-time context information, and determine key video segments in the process of matching preferred viewing characteristics with multimodal data;

[0025] Step S13: Determine the corresponding emotional features based on each key video segment and the corresponding user viewing behavior, and determine the corresponding continuous playback instruction based on the emotional features, the short videos corresponding to the key video segments, and the preset playback order of multiple short videos.

[0026] Step S14: Construct the corresponding short video continuous playback event under the trigger of the continuous playback command, and collect the user's multi-dimensional behavioral features in the short video continuous playback event, and determine the primary operation command based on the multi-dimensional behavioral features.

[0027] Step S15: Match the playback screen of the primary operation instruction and the continuous playback event of the short video in the same time dimension to determine the corresponding matching level, and determine the final operation instruction based on the matching level and the user's viewing history.

[0028] refer to Figure 2 In step S11, the specific steps are as follows:

[0029] S111: In the short video database, multiple short videos in the database are labeled, and frame-by-frame parsing is performed on the labeled short videos to trigger feature extraction in different dimensions. At this time, in the visual dimension, a visual stream containing HSV color space and keyframe objects is extracted based on a convolutional network; in the auditory dimension, an audio stream containing background music beats and speech fundamental frequency changes is extracted based on an audio processing network; in the semantic dimension, a text stream containing video embedded subtitles and text in keyframe images is extracted based on a natural language processing network.

[0030] S112: Align and map the extracted visual stream, audio stream, and text stream, project them into a preset semantic space, calculate the interaction weights between modalities through an attention mechanism in the preset semantic space, and deeply fuse them to construct unique multimodal data for each short video.

[0031] In the embodiments of this application, the system preprocesses the massive amount of videos in the short video database. This is not just about building an index, but about timestamping and segmenting the videos using a preset metadata protocol. The system uses a sliding window mechanism to parse the marked short videos frame by frame, discretizing the continuous video stream into a computable image sequence, thereby triggering subsequent parallel feature extraction channels. Its core purpose is to synchronize the timeline of the video stream and ensure that visual, auditory, and textual features can be aligned in a unified time dimension.

[0032] For the visual dimension, the system uses convolutional neural networks (CNNs) and their variants (such as ResNet or EfficientNet) as feature extractors; the system not only extracts RGB features, but also converts image frames to the HSV (hue, saturation, brightness) color space; by calculating the HSV histogram distribution within a specific time window, it captures the color tone and emotional atmosphere of the video.

[0033] By using object detection methods (such as YOLO or Faster R-CNN), CNNs identify high-confidence bounding boxes in keyframes and extract salient object features (such as faces, vehicles, animals, etc.). These features not only include the object's category label, but also the object's spatial location in the image and its motion trajectory over time, thereby constructing dynamic visual stream data.

[0034] For the auditory dimension, the system employs audio processing networks (such as VGGish or Wav2Vec) to perform deep feature analysis on the audio stream. It converts the audio signal into a time-frequency spectrum using Short Time Fourier Transform (STFT) and extracts BPM (beats per minute) and beat start points using rhythm detection. This helps determine the tempo and rhythmic patterns of the video, providing key features for identifying "high points" or "soothing sections." Simultaneously, for speech segments, the system uses signal processing techniques to extract the fundamental frequency trajectory. The fundamental frequency variation curve reflects the speaker's emotional fluctuations (such as excitement, doubt, or low pitch). The system smooths and normalizes the fundamental frequency to eliminate individual vocal differences, focusing on the representation of emotional nuances.

[0035] For the semantic dimension, the system uses optical character recognition (OCR) combined with a natural language processing model for semantic understanding; the OCR engine accurately identifies hard subtitles from video frames and combines the timeline to reassemble discrete characters into a continuous text stream; the NLP model performs word segmentation and word vector embedding on the text to capture narrative logic and keywords; in addition to subtitles, the system also detects iconic text in the scene (such as road signs, poster text, and UI elements). These texts serve as strong semantic features of the scene and are used for entity recognition through the NLP network to help determine the specific scene background of the video.

[0036] Specifically, Group A of short videos is a series of continuous story content about "outdoor off-road adventure". The system retrieves Group A of short videos from the database and marks each video with a millisecond-level timestamp. The parser cuts a 3-minute video stream into a sequence of 4,500 frames (assuming 25fps).

[0037] When the CNN was analyzing the key segment from 20 to 30 seconds, it detected a high proportion of highly saturated orange and ochre in the HSV color space (the Hue value in the HSV histogram was concentrated in the range of 30-60), and the system determined that the scene was a "desert or Gobi" environment. At the same time, the keyframe object detector identified "off-road vehicle" (ConfidenceScore>0.9) and "flying dust" (dynamic texture feature) in consecutive frames, and the visual stream data was labeled as "high-speed motion" and "natural terrain".

[0038] Within the same time window, the audio processing network analyzes the background noise; the STFT time-frequency graph shows a strong low-frequency impact signal, and the rhythm detection method calculates a fast-paced drumbeat with a BPM of 120; at the same time, the voice fundamental frequency analysis module extracts the driver's voice fundamental frequency curve, which rises sharply in a short period of time, and the system quantifies it as the emotional feature of "tension / excitement"; the audio stream is finally encoded into acoustic feature vectors of "high dynamic range" and "urgency".

[0039] At the 25-second mark, the OCR module identified the subtitle text "Beware of sand dunes!" in the lower left corner of the screen, and at the same time, it identified the text "edge of no man's land" on the road sign on the right side of the screen. The NLP network converted these texts into high-dimensional word vectors, and through semantic association analysis, it semantically anchored "sand dunes" and "no man's land" with the "desert" feature in the visual stream, and extracted the semantic tags of "exploration" and "danger warning".

[0040] Furthermore, by interpolation, low-frequency sampled text features and high-frequency sampled audio features are mapped to the timestamps of video frames, ensuring that at any given moment, visual vectors, audio vectors, and text vectors have a strict temporal correspondence. The original features of different dimensions are linearly transformed through their respective fully connected layers and uniformly projected to the same dimension (e.g., 512-dimensional or 1024-dimensional).

[0041] The system constructs a high-dimensional latent semantic space, typically achieved through embedding layers in a deep neural network. In this space, visual, auditory, and semantic features with similar meanings should be close to each other. The aligned feature vectors are then input into a multi-head mapping network, which transforms physical features (such as pixel grayscale and sound wave frequency) into abstract semantic representations through nonlinear transformations.

[0042] In the semantic space, not all modalities are equally important at every moment; the system introduces a multi-head attention mechanism or a cross-modal attention mechanism to dynamically evaluate the interaction strength between modalities; the system calculates the attention score of the visual features at the current moment as the query, and the audio and text features as the key and value, which is essentially calculating the mutual information between modalities, that is, to what extent the current scene depends on the current background sound or subtitle; based on the relevance score, the system generates a set of dynamic weight coefficients αt; for example, in intense action segments, the weight of visuals and audio is increased, while the weight of text is decreased; in narration segments, the weight of text dominates.

[0043] The system fuses multimodal features based on the calculated interaction weights; it uses weighted Hofdinger product or fusion units based on gating mechanisms to nonlinearly combine visual, auditory, and text vectors according to attention weights; the fused features are normalized to generate a multimodal data feature containing rich spatiotemporal and semantic information.

[0044] Specifically, at the critical moment of 28 seconds, the visual stream input is the image feature vector of "off-road vehicle taking off", the audio stream input is the spectral features of "engine roar", and the text stream input is the NLP vector of the subtitle "It's flying!" The system uses timestamp alignment to lock these three at t=28s and maps them into three 1024-dimensional standard vectors respectively through a fully connected layer.

[0045] The system inputs these three standard vectors into a preset semantic space. In this abstract space, the visual features of "off-road vehicle taking off" and the auditory features of "engine roaring" gradually converge in the semantic coordinate system. Because in this space, the visual concept of "high-energy motion" and the auditory concept of "high-decibel noise" have a high similarity measure, they are clustered into the semantic region of "stimulation / orgasm".

[0046] In this clip, the visuals showcase the breathtaking moment of the vehicle taking off, accompanied by the powerful engine sound. The mutual information between the visual and audio streams is extremely high (the visuals and sounds are highly synchronized and mutually interpretive), thus assigning a weight of αv=0.5 to the visual stream and αa=0.4 to the audio stream. In contrast, the text stream "It's flying!" is semantically relevant but has lower information entropy, and is assigned a lower weight of αt=0.1. This dynamic allocation avoids subtitle interference in judging the atmosphere of the scene. Based on the above weights (0.5, 0.4, 0.1), the system performs a deep fusion operation. The resulting multimodal feature tensor is no longer a simple stack of features, but a multimodal data with "visual impact" as the main driver, "auditory shock" as an auxiliary driver, and "textual description" as a supplement. This multimodal data accurately represents the core meaning of the 28th second of the short video in Group A: "a highly impactful visual and auditory climax event."

[0047] refer to Figure 3 In step S12, the specific steps are as follows:

[0048] S121: Based on the database of short videos and the user's entry information, determine the user's data space, and determine the user's historical viewing data based on the traversal of this data space. Use the Long Short-Term Memory Network to mine the user's long-term interests and preferences. At the same time, combine the current timestamp, geographical location information and the audiovisual features of the preceding videos to obtain real-time context information. Determine the user's preferred viewing characteristics based on the multi-dimensional fusion of the user's long-term interests and preferences and real-time context information.

[0049] S122: Based on the viewing characteristics of the preference and multimodal data, determine the corresponding similarity combination, and determine a set of highly relevant candidate videos by matching the similarity combination with the database of short videos in a high dimension. Based on the candidate video set and the corresponding spatiotemporal attention mechanism, determine multiple frame-level features of the candidate videos, and combine multiple frame-level features to determine the corresponding key video segments.

[0050] In the embodiments of this application, the system calls the underlying data warehouse based on the user's unique identifier (such as User_ID) and device fingerprint to construct a dedicated multi-dimensional feature space. This space not only includes a structured historical viewing list (Video_ID list), but also semi-structured interaction logs (likes, shares, comments), unstructured search records, and device metadata.

[0051] The system uses a cursor or traversal method to perform a full scan of the above data space; during this process, outliers (such as millisecond-level playback records generated by accidental touches) and invalid data (such as records that failed to load) are removed through data cleaning strategies, and finally statistically significant valid historical viewing sequences are extracted.

[0052] The extracted historical viewing sequence is transformed into a time series vector and input into a Long Short-Term Memory (LSTM) network. LSTM solves the gradient vanishing problem in traditional RNNs when processing long sequences through its unique cell state design. The network decides which old information to discard (such as outdated trending videos) through the forget gate, which new information to update (such as recently followed new topics) through the input gate, and which hidden states to output through the output gate. Through layer-by-layer transmission, LSTM can uncover deep interest anchors that users maintain stable over a long period of time (such as continuous attention to "tech aesthetics" or "extreme sports") and encode them as long-term interest latent vectors.

[0053] The system collects timestamps in real time and parses them into semantic time (such as "weekend evening" or "commuting peak") through natural language processing (NLP). At the same time, it calls location services (LBS) to obtain latitude and longitude coordinates and encodes them into geographic scenes (such as "residential area", "commercial center", or "moving"). The system extracts the audiovisual features of the preceding video, including the visual dominant color tone of the last frame of the previous video, the rhythm (BPM) of the background music, and the volume decay curve. These features constitute the short-term memory background of the user's current audiovisual experience, which is crucial for ensuring the emotional continuity of continuous playback.

[0054] The system employs a gated fusion unit or attention splicing layer to fuse the obtained "long-term interest latent vector" with the obtained "real-time context vector". The fusion process is not a simple superposition, but a dynamic weighting based on the urgency of the current context. For example, when the real-time context shows that the user is in a "high-speed movement" state, the weight of "deep reading" in the long-term interest is suppressed, while the weight of "fast-paced entertainment" is amplified. Finally, a high-dimensional user preference viewing feature vector is output.

[0055] Specifically, the system retrieved the data space of user ID "User_X"; it scanned the records of the past 3 months in a traversal manner, cleaned up accidental touch records with a duration of less than 3 seconds, and extracted a viewing sequence containing 500 valid Video_IDs; the system found that the sequence frequently contained keywords such as "Wrangler", "Tank 300", and "Dakar Rally".

[0056] The feature vectors of the above 500 videos are sequentially input into the LSTM model. At time step t, the model forgets the accidental "makeup tutorial" viewing record from 2 months ago, but strengthens the memory of watching "off-road modification" videos continuously in the past two weeks. The final output long-term interest latent vector strongly points to "mechanical power" and "unexplored outdoor areas" in the semantic space, and numerically shows a high confidence in the "off-road" category (preset: 0.85).

[0057] Based on the context information, the current timestamp is resolved to "Saturday afternoon 16:30" (leisure time); the geolocation LBS signal shows that the user is located on the "ring expressway" (moving state, vehicle speed 80km / h), the previously played short video (Group A, Episode 4) ends with a scene of a vehicle driving into a mud pit, the visual stream feature is "high brightness natural light", and the audio stream feature is "dull engine low frequency sound".

[0058] Because it was a "Saturday afternoon" and "high-speed movement," the weight of the real-time context was increased, and the system determined that the user needed highly arousing content to prevent driving fatigue. The "long-term off-road interest" output by LSTM was combined with the "fatigue prevention needs" to correct long-term preferences—at this time, the user was less inclined to watch "static explanation" off-road videos and more inclined to watch "dynamic stimulation" off-road videos. The generated viewing preference feature vector was defined as: "high dynamic range, strong visual impact, and strong sense of rhythm in outdoor off-road real shots." Based on the calculation of S121, the system locked the user's current "psychological profile," providing a basis for subsequent steps to accurately match exciting segments such as "vehicles climbing hills" or "high-speed cornering" from the short videos in Group A.

[0059] Furthermore, based on the viewing preferences and multimodal data, corresponding similarity combinations are determined. A highly relevant set of candidate videos is then identified by high-dimensional matching of these similarity combinations with the short video database. Multiple frame-level features of the candidate videos are determined based on the candidate video set and the corresponding spatiotemporal attention mechanism. These multiple frame-level features are then combined to identify the corresponding key video segments. This approach considers both the candidate video set and the corresponding spatiotemporal attention mechanism, ensuring the accuracy of the multiple frame-level features of the candidate videos. Simultaneously, the visual, audio, and text streams of the short videos are introduced to further control the multimodal data. Finally, the combination of viewing preferences and multimodal data enables precise control of key video segments.

[0060] At this point, the system compares the user's preferred viewing feature vector generated in S121 with the multimodal data of each short video in the database (output in S112). This comparison is not single-dimensional, but rather calculates cosine similarity or Mahalanobis distance in the visual subspace, auditory subspace, and semantic subspace respectively.

[0061] The weights of similarity combinations are dynamically determined based on the user's preferences. For example, if the user's preferences lean towards visual stimulation, the weight coefficient of visual similarity increases, while the coefficients of audio and text decrease accordingly. The system outputs a comprehensive similarity combination score matrix, which reflects the potential of each video to meet the user's needs in different dimensions.

[0062] Using the Approximate Nearest Neighbor (ANN) search method or a high-dimensional index structure based on KD-Tree, the system searches for clusters in the potential high-dimensional semantic space that are closest to the user's preferred feature vectors. The system sets a dynamic threshold to remove videos with similarity scores below the threshold, as these videos are considered irrelevant to the user's current state or have very low relevance. After filtering, the remaining videos form a highly relevant candidate video set. The videos in this set are highly consistent with the user's long-term interests and immediate context at the macro level (such as theme and style).

[0063] For each video in the candidate video set, the system inputs its original frame sequence into a spatiotemporal attention network (such as 3D-CNN or VideoTransformer); the model calculates the information gain of each frame along the time axis, assigning high weights to frames with drastic visual changes and large motion amplitudes, while ignoring static or repetitive frames; within each frame, the model generates an attention mask, focusing on salient regions in the image (such as main characters and key objects) and suppressing background noise; through the weighted processing of the above mechanism, the system extracts a series of high-confidence frame-level feature vectors that can characterize the core content of the video. These feature vectors are not only image features, but also contain contextual information of audio and text at that moment.

[0064] The system analyzes the distribution of multiple extracted frame-level features in the feature space; through clustering methods (such as DBSCAN or Mean-Shift), it identifies those closely adjacent and continuous frame sequences in the feature space, which usually correspond to a specific event or plot climax in the video; based on the clustering results, the system determines the start and end timestamps of the video segments; if the density of high-weight frame-level features in a video segment exceeds a preset threshold, the segment is marked as a key video segment; the final output is a specific list of highlight segments with time indexes.

[0065] Specifically, assuming that step S121 determines that the user currently needs content with "high dynamic range and strong visual impact"; assuming that step S121 determines that the user currently needs content with "high dynamic range and strong visual impact"; in the visual subspace, the user prefers "high-speed motion", so videos containing a large number of "vehicle motion blur" features have high visual similarity (0.88); in the auditory subspace, the user prefers "strong rhythm", so videos with prominent "high engine speed" features have high audio similarity (0.85); episode X of group A, "Desert Summit", is given a comprehensive score of 0.92 due to its high visual and audio similarity, ranking first in the similarity combination matrix.

[0066] The system performs searches in a high-dimensional semantic space with a threshold of 0.75. Since the score of "Desert Summit" is 0.92, which far exceeds the threshold, it was successfully selected into the set of highly relevant candidate videos. Another video, "Camping Campsite Setup," although semantically relevant, has weak action and a score of only 0.6, so it was removed from the set.

[0067] The system performed a frame-by-frame scan of "Desert Challenge". In the temporal dimension, during the vehicle's static preparation period from 0 to 5 seconds, the attention weight was low. From 15 to 18 seconds, the vehicle began its full-speed climb, causing the image to shake violently and the motion vector to surge, at which point the temporal attention weight reached its peak. In the spatial dimension, from 15 to 18 seconds, the attention mask tightly locked the central area of ​​the image (the off-road vehicle) and the dust kicked up by the wheels, while the background sky was suppressed. The system successfully extracted a series of high-weight frame-level features at 15.0 seconds, 15.5 seconds, 16.0 seconds...17.5 seconds.

[0068] The system discovered that the frame-level features between 15 and 18 seconds exhibited extremely high clustering in the feature space (all features being "high dynamic range, dust, and engine noise"). After calculation, the feature density of this time period far exceeded the threshold. Therefore, the system determined that 15 to 18 seconds of "Desert Summit" was the key video segment. Ultimately, the system would not recommend the entire 3-minute video, but would precisely locate and prepare to play the most exciting 3 seconds of the "summit moment," meeting the user's immediate need for "strong visual impact."

[0069] refer to Figure 4 In step S13, the specific steps are as follows:

[0070] S131: Perform sentiment analysis on key video segments and determine the corresponding sentiment content in the sentiment analysis; at the same time, determine the corresponding user viewing behavior based on the action detection of key video segments, weight and correct the sentiment content and the corresponding user viewing behavior, and construct the corresponding emotional features.

[0071] S132: Collect the preset playback order of multiple short videos. At the same time, align the emotional features with the current playback progress and narrative structure of the short videos in time and space, and perform dynamic deduction in combination with the preset playback order of multiple short videos to generate an emotional smooth curve that can avoid emotional abrupt changes.

[0072] S133: Based on the recognition of the emotion smoothing curve, multiple emotion nodes are identified. Based on the multiple emotion nodes and the corresponding key video segments, a determined emotion change path is constructed. The playback queue of multiple short videos is determined by parsing the emotion change path, and the corresponding continuous playback instructions are presented in the playback queue.

[0073] In the embodiments of this application, the system utilizes a visual emotion analysis network and an audio emotion recognition model to process key video segments in parallel. In the visual dimension, key points of the face in the image (such as eyebrows and corners of the mouth) are detected by the facial expression analysis unit to identify basic emotions (such as happiness and anger). At the same time, scene semantic understanding is used to analyze color psychology features (such as red representing intensity and blue representing calmness). In the auditory dimension, the intonation of speech is extracted through Mel-frequency cepstral coefficients (MFCC), and the emotional atmosphere is determined by combining the rhythm and harmony of the background music.

[0074] The system inputs the sentiment scores of visual, auditory, and textual (subtitles / comments) into a multimodal fusion classifier. This classifier assigns weights to different modalities based on an attention mechanism (e.g., emphasizing visuals in silent videos and audio in music videos). Finally, it outputs the coordinate position of the segment in the sentiment space and maps it to specific sentiment content labels (such as "excitement", "pleasure", "fear", "sadness", etc.) and their corresponding confidence probabilities.

[0075] The system uses optical flow to calculate pixel-level motion vectors between video frames and generate motion amplitude maps. At the same time, it uses spatiotemporal interest point (3D-STIP) detection or a dual-stream network to identify key action units in the video (such as collision, running, explosion, and stillness).

[0076] A "video action-user behavior" mapping model is established. This model is based on psychological cognition and transforms the motion features in the video into the user's expected reaction state. In high-dynamic scenarios, when high-frequency, large-scale pixel motion (such as vigorous movement) is detected, it is mapped to the user behavior feature of "high arousal" or "focus". In low-dynamic scenarios, when the detected motion vector is sparse and flat, it is mapped to the user behavior feature of "relaxation" or "potential churn risk". The system outputs a vector describing the user's behavioral tendency under the stimulation of the video.

[0077] The system dynamically adjusts the fusion weight of emotional content and user viewing behavior based on the context of the current scene (such as viewing time and historical preferences). For example, during the cold start phase of the recommendation system, it relies more on the emotion of the video itself (high emotion weight); while during periods of user fatigue, it relies more on user behavior prediction (high behavior weight).

[0078] The system concatenates or fuses the weighted emotion vector with the behavior vector using tensors. During this process, the system uses gated recurrent units (GRUs) to perform nonlinear transformations on the features, filtering out noise that causes ambiguity (such as screams in horror movies, which are "fear" emotions but bring "excitement" behavior to the audience). Finally, a high-dimensional composite emotion feature is constructed.

[0079] Specifically, the key video clip currently identified is: "An off-road vehicle is charging up a sand dune in the desert and leaping into the air at the top"; Visual analysis: The system detected sand flying in the scene, with 70% of the screen occupied by warm colors (orange-yellow) with extremely high saturation, and the moment the vehicle takes off has a strong visual impact; The visual emotion model determines it as "high arousal" and "high pleasure".

[0080] Auditory analysis: The engine roar in the audio stream has an extremely high frequency, and the background music is fast-paced heavy metal drum beats; the audio emotion model determines it to be "intense" and "powerful"; combining visual and auditory information, the system determines the emotional content of this segment to be "exhilarating and shocking", and its emotional space coordinates are close (Valence: positive, Arousal: extremely high).

[0081] The optical flow method detected that the pixel displacement speed of the center of the image (vehicle) was extremely fast and the direction was upward diving; the system identified the core action as "leaping into the air"; based on the "video action-user behavior" model, such violent leaping into the air can usually strongly stimulate the user's senses, triggering high attention and dopamine secretion; therefore, the user viewing behavior characteristics deduced by the system are "pupil dilation, high excitement, and high concentration of attention".

[0082] Considering that this is the climax of the videos in Group A, the system assigns a slightly higher weight (0.6) to "user viewing behavior" (excitement) and a lower weight (0.4) to "emotional content" (shock), in order to highlight the user's experience rather than simply video tags; the system merges the emotional vector of "excitement and shock" with the behavioral vector of "high excitement"; and constructs the emotional feature of this segment as "high-energy excitement state". This feature not only tells the system that this video is very exciting, but also more accurately predicts that the user will be in a state of high emotion and unwilling to be interrupted while watching, thereby guiding the subsequent playback logic (for example, do not insert advertisements in the middle of this segment, and the next segment should maintain this high emotion or transition smoothly).

[0083] Furthermore, the preset playback order of multiple short videos is collected. At the same time, the emotional features are spatiotemporally aligned with the current playback progress and narrative structure of the short videos, and dynamic deduction is performed in combination with the preset playback order of multiple short videos to generate an emotional smooth curve that can avoid emotional abrupt changes.

[0084] At this point, the system extracts topological sorting information or plot dependencies of multiple short videos from the content management system. This is usually based on the video's metadata tags (such as "Episode 1", "Episode 2") or the narrative logic of the story graph. The system constructs a directed graph, where nodes represent videos and edges represent recommended connections, thus determining the macroscopic order of playback.

[0085] The system reads the playback timestamp of the current short video in real time and maps it to the timeline of the narrative structure. At the same time, it uses scene segmentation to identify the current narrative structure unit (such as "beginning setup", "conflict development", "climax", "ending"). The system injects the emotional features (such as "tension" and "excitement") generated in step S131 into the intersection of the timestamp and the structure unit. Through this process, the system locks the user's current "emotion-narrative" coordinates in the four-dimensional space (time, space, narrative, emotion) to ensure that subsequent deductions are based on the current precise state.

[0086] The system employs a sequence prediction model (such as TransformerDecoder or Hidden Markov Model HMM) to recursively predict the potential emotional state triggered by each key video segment along the preset playback order. The system does not only predict the next video, but also looks forward N steps (e.g., predicting the emotional sequence of the next 3-5 videos). During the inference process, the system considers emotional inertia, that is, the current emotion will affect the emotion at the next moment with a certain decay coefficient.

[0087] In the simulated sequence, the system calculates the emotional gradient between adjacent nodes; if a gradient change is detected that exceeds a preset psychological comfort threshold (e.g., a sudden jump from extreme "sadness" to extreme "celebration", or a sudden change from "intense" to "dead silence"), it is marked as an "emotional inflection point".

[0088] The system uses spline interpolation or weighted moving average to smooth the original emotion prediction sequence containing abrupt changes. Interpolation strategy: If there is an abrupt change between two videos with significantly different emotions, the system will attempt to insert a transitional emotional segment in between (such as inserting "calm" between "tension" and "relaxation"). Pruning strategy: If the emotion of a video seriously conflicts with the overall flow and cannot be transitioned, the system will adjust the playback order, delaying or removing it from the current queue. After the above calculations, the system generates a smooth emotion curve that changes continuously over time, and the derivative (rate of change) of this curve is limited to a range acceptable to human psychology.

[0089] Specifically, the system retrieves the videos from Group A in the following preset order: Episode 1 (Departure) > Episode 2 (Night Walk in the Wasteland) > Episode 3 (Crossing the Desert) > Episode 4 (Discovering the Oasis); the current timestamp shows that the playback progress of Episode 2 is 95% (about to end); the scene recognition method identifies the current scene as "dawn breaking", and the narrative structure is in the state of "the conflict has ended and a new stage is about to begin"; the current emotional characteristic calculated by S131 is "release after repression"; the system aligns this state to the Tend node on the timeline.

[0090] The system extrapolates backward based on the current state; the opening of episode 3, "Crossing the Desert," contains a large number of high-speed driving scenes, predicting the emotional characteristics as "extreme tension and excitement"; the system calculation found that the emotional gradient was too large, jumping directly from the "suppressed release" (low arousal, neutral Valence) at the end of episode 2 to the "extreme tension" (high arousal, high energy) at the beginning of episode 3. This sudden surge of adrenaline caused discomfort (emotional mutation) when the user had just relaxed.

[0091] To eliminate this abrupt change, the system performed a dynamic adjustment. The system reviewed the data from episode 3 and discovered a segment in the middle of the episode depicting a smooth cruise on sand dunes accompanied by soothing music. The system decided to reconstruct the playback logic: instead of starting from the beginning of episode 3, it would directly cut to the "smooth cruise" segment in the middle of episode 3. The generated emotional smoothing curve presented the following pattern: from the "release" at the end of episode 2 > a smooth transition to the "comfort and vastness" (emotional buffer) in the middle of episode 3 > then gradually rising to the subsequent "tension and excitement." Through this adjustment, the system avoided abrupt editing, ensuring that the user's emotions could naturally "awaken" before entering the climax experience, achieving a seamless emotional transition.

[0092] Therefore, multiple emotion nodes are identified based on the recognition of emotion smoothing curves. An emotion change path is constructed based on these emotion nodes and their corresponding key video segments. A playback queue of multiple short videos is determined by parsing this emotion change path, and corresponding continuous playback instructions are presented in the playback queue. This approach considers the overall construction of multiple emotion nodes and their corresponding key video segments, ensuring the accuracy of the emotion change path. At the same time, it achieves overall control over emotion features, the short videos corresponding to key video segments, and the preset playback order of multiple short videos, and improves the accuracy of continuous playback instructions, thus achieving full consideration of different dimensions.

[0093] At this point, the system performs mathematical analysis on the emotion smoothing curve generated by S132, using extreme point detection and curvature analysis to identify key positions on the curve. The system identifies local maxima on the curve as "emotional high points" (such as peaks of excitement or shock), local minima as "emotional low points" (such as valleys of calm or soothing), and high curvature points (positions where curvature changes drastically) as "emotional turning points." Each identified emotion node is assigned a set of attribute vectors, including: timestamp (relative time in the emotion stream), emotion polarity (positive / negative), arousal level, and duration. These nodes constitute milestones for subsequent path planning.

[0094] The system connects multiple identified emotion nodes in a directed manner according to time sequence to construct a directed acyclic graph (DAG) as a preliminary path framework. The system uses the key video segments generated in step S122 as a resource library and matches them with the attributes of emotion nodes based on their multimodal features. The matching method adopts semantic nearest neighbor search to find the video segments with the smallest spatial distance of emotional features.

[0095] The system calculates the transition cost between adjacent nodes; if the emotional jump between two nodes is too large (although it has been smoothed, there is still a certain slope), the system will try to find intermediate transition segments to reduce the transition cost; and construct an emotional change path composed of alternating "emotional nodes" and "key video segments" to ensure that each step on the path has logical coherence.

[0096] The system analyzes the path of emotional changes and extracts the corresponding key video segments in linear order. It checks the continuity between segments (e.g., whether there is a visual conflict between the ending of one segment and the beginning of the next) and makes fine adjustments. Finally, it generates an ordered list containing Video_ID, In-Point, and Out-Point, i.e., a playback queue. Based on the playback queue, the system generates standardized continuous playback instructions, which are structured data packets containing:

[0097] Sequence control parameters: instruct the player to read segments in the queue in sequence;

[0098] Rendering instructions: Generate specific transition effects (such as fade in / out, dissolve, or hard cut) for the connection points between clips to match the smooth transition or strong contrast of emotions;

[0099] Triggering mechanism: A time trigger is embedded in the instruction to ensure that when the current segment reaches the end point, the next segment is loaded in milliseconds.

[0100] Specifically, the system has completed the emotion deduction and is ready to generate the final playlist for the user; the system analyzed the emotion smoothing curve and found three key feature points: Node A: located at the beginning of the curve, the emotion feature is "expectation and accumulation" (low arousal, positive); Node B: located at the middle peak of the curve, the emotion feature is "passion and explosion" (extremely high arousal, highly positive); Node C: located at the end of the curve, the emotion feature is "tranquility and aftertaste" (low arousal, positive).

[0101] Node A Matching: The system searches the video library of Group A and finds a clip in Episode 1 of Group A, "Ready to Go," showing "engine idling and driver adjusting rearview mirrors," whose emotional characteristics perfectly match Node A. Node B Matching: The system finds a clip in Episode 3 of Group A, "Extreme Summit," showing "vehicles leaping over sand dunes," matching the "passionate outburst" of Node B. Node C Matching: The system finds a clip in Episode 4 of Group A, "Viewing the Panoramic Sunset from the Summit," showing "car parked on the mountaintop, panoramic sunset," matching the "tranquil reflection" of Node C. System Construction Path: Node A (Expectation) > [Transition] > Node B (Outburst) > [Transition] > Node C (Tranquility). The system inserts a micro-segment of "vehicle starting and accelerating" between Nodes A and B as a transition.

[0102] The playback queue sequence determined by the system is as follows: Clip_01: Episode 1 from Group A, timecode 00:15-00:25 (building momentum); Clip_02: Episode 3 from Group A, timecode 12:05-12:08 (starting transition); Clip_03: Episode 3 from Group A, timecode 14:20-14:25 (climax); Clip_04: Episode 4 from Group A, timecode 02:10-02:20 (sunset ending).

[0103] The system generates continuous playback instructions: the instructions specify the player to load Clip_01; at the end of Clip_01, a "hard cut" instruction is generated, seamlessly switching to Clip_02 (to match the increased tempo); at the end of Clip_02, a "dissolve" instruction is generated, cutting into Clip_03 (enhancing the climax); at the end of Clip_03, a "long fade-out" instruction is generated, smoothly transitioning to Clip_04 (guiding the calming of emotions); after the user clicks play, they will see a continuous short video with a complete cinematic emotional narrative, from getting ready to take off to the ultimate leap, and finally ending with a magnificent sunset.

[0104] refer to Figure 5 In step S14, the specific steps are as follows:

[0105] S141: When triggered by a continuous playback command, multiple timeline anchor points are determined based on the recognition of the continuous playback command. The corresponding playback content is determined by tracing each timeline anchor point. A continuous playback event of a short video containing timeline anchor points is constructed based on multiple timeline anchor points, the corresponding playback content, and the corresponding content weight.

[0106] S142: Determine user behavior events based on the parsing of continuous playback events of short videos, identify multiple sub-behavioral contents in the behavior events, and determine corresponding multi-dimensional behavior features based on the multiple sub-behavioral contents, user action characteristics and action trajectories;

[0107] S143: Utilize a multimodal behavior perception network to process multidimensional behavioral features and output the corresponding behavioral intent. Based on the behavioral intent, the current playback progress of the short video, and the corresponding contextual content, generate the corresponding primary operation instructions. These primary operation instructions are not instructions to be executed immediately and include video jump, speed up playback, pause, or immediate termination.

[0108] In the embodiments of this application, the system receives a continuous playback instruction data packet and extracts key parameters from it through a syntax parser, including the source video ID, the time range of the segment, and transition effect parameters. The system maps the time offset in the instruction onto the global rendering timeline of the player. Specifically, the system calculates the physical timestamp corresponding to the current logical time and marks the start frame timestamp, end frame timestamp, and intermediate transition trigger times of the segment as key timeline anchor points. These anchor points not only represent time points but also represent the switching boundaries of the playback state (such as the moment of switching from "current video" to "next video").

[0109] For each defined timeline anchor point, the system performs a reverse index in the multimodal database based on its associated metadata (such as Clip_ID and Frame_Index). The system locates the specific storage location and extracts the playback content corresponding to that anchor point, which includes the video stream (YUV / RGB pixel data), the audio stream (PCM encoded data), and the auxiliary subtitle stream. The system preloads and decodes these data streams, converting them into raw frame buffer data that can be directly read by the rendering engine, ensuring that the data is ready when the anchor point is reached, thus eliminating rendering stutters.

[0110] The system calculates the content weight of each anchor point based on the importance of the segment in the overall narrative (such as whether it is an emotional climax), the user's historical preferences (such as liking a specific visual style), and the information density of the segment (such as frame richness and pacing). The weight value determines the priority of system resource allocation and the position of the segment in the user recommendation logic.

[0111] The system binds the timeline anchor point (time dimension), playback content (data dimension), and content weight (logical dimension) into a triplet and encapsulates it into a continuous playback event object. This object maintains a state machine in memory, which describes the complete lifecycle from "to be played" to "playing" and then to "playing ended", and provides a unified data interface for subsequent user interaction analysis.

[0112] Specifically, the current continuous playback command requires playing a carefully edited montage of content; the system parses the continuous playback command and identifies that the command contains a seamless splicing of two video clips: the first clip is the end of episode 1 of group A, "vehicles driving towards the sand dunes", and the second clip is the beginning of episode 2 of group A, "the moment the vehicle takes off".

[0113] On the global timeline, the system sets T0 as the current moment, T1 as the end of segment 1 (set to 00:15:00 in the command), and T2 as the end of segment 2 (00:20:00). The system identifies T0, T1, and T2 as three key timeline anchor points. T1 is specially marked as the "transition anchor point" because the switch from episode 1 to episode 2 will occur at this moment.

[0114] For the T0 to T1 interval, the system traces back to the index of episode 1 of group A based on the anchor point information, and extracts the video frame sequence from 14 minutes 50 seconds to 15 seconds (the scene shows the wheels kicking up dust) and the corresponding engine roar; for the T1 to T2 interval, the system traces back to the index of episode 2 of group A, and extracts the video frame sequence from 0 minutes 05 seconds to 0 minutes 10 seconds (the scene shows the vehicle flying in the air) and the strong rhythmic drumbeats of the background music. This data is preloaded into the video memory and audio buffer.

[0115] The system analyzes the content and finds that the segment from T1 to T2 is the emotional climax, so it is assigned a content weight of 0.9 (the highest level); while the segment from T0 to T1 is a prelude, and is assigned a content weight of 0.6. The system constructs a continuous playback event Eplay, which includes: anchor points [T0, T1, T2], corresponding frame buffer data, and weight labels. This event object is sent to the player core, instructing it to perform a "dissolve" transition at time T1 to ensure a smooth transition from "driving towards the dunes" to "leaping into the air". Throughout the playback process, resources are prioritized to ensure the smoothness of the high-weight "leaping" segment.

[0116] Furthermore, user behavior events are determined based on the analysis of continuous playback events of short videos, and multiple sub-behavioral contents are identified within these behavioral events. Based on these multiple sub-behavioral contents, user action characteristics, and action trajectories, corresponding multi-dimensional behavioral features are determined. This approach takes into account the overall consideration of multiple sub-behavioral contents, user action characteristics, and action trajectories, ensuring the accuracy of the corresponding multi-dimensional behavioral features.

[0117] At this time, the system receives the continuous playback event constructed by S141 and extracts the time axis anchor point as the time reference; the system defines a sliding time window centered on the current anchor point, and synchronously collects multimodal sensor data (camera image, microphone audio, touch screen signal) within this window.

[0118] Using background modeling methods (such as Gaussian mixture models) to filter out static environmental backgrounds and separate foreground dynamic targets (users); at the same time, based on the user's human skeleton detection results, regions of interest (ROIs) are identified, such as the user's hand area, face area, and the posture of the handheld mobile device.

[0119] The system compares the dynamic changes within the ROI with preset behavior trigger thresholds. When the detected dynamic changes (such as pixel displacement or sound wave amplitude) exceed the threshold, the system determines that a user behavior event has been triggered. This event is given a timestamp and aligned with a specific anchor point (such as the climax of the video) in the continuous playback event.

[0120] The system employs temporal action segmentation technology to finely segment triggered user behavior events along the time dimension. The system calculates the amplitude changes of the optical flow field, identifying peaks or troughs in the optical flow amplitude as the boundaries of actions. Within the segmented time segments, the system uses a pose estimation network (such as OpenPose or MediaPipe) to extract the coordinates of key human body points and combines this with a classifier to identify specific action types. The identified content is defined as sub-behavior content, such as "head turning to the left," "index finger tapping the screen," "pupil dilation," and "rapid hand shaking." The system arranges the identified sub-behavior content in chronological order to form an atomic behavior sequence.

[0121] For each sub-behavior, the system extracts its static geometric features (such as joint angles and limb extension range) and dynamic dynamic features (such as angular velocity, acceleration, and kinetic energy); for example, the force of a finger tap is quantified by a screen pressure sensor or the rate of change of the fingertip contact area.

[0122] The system tracks the positional changes of key points (such as the fingertip and nose tip) between consecutive frames to generate spatiotemporal motion trajectories. The system further calculates parameters such as the curvature, displacement vector, and direction cosine of the trajectory to describe the smoothness or abruptness of the motion. The system then splices and normalizes the above motion features (describing the motion itself) and motion trajectories (describing the motion path) to construct a multidimensional behavior feature vector.

[0123] Specifically, the current playback event is currently at the critical moment where "the vehicle skidded on the edge of a sand dune and nearly overturned" (high-weight anchor point); the system detected that the current timeline anchor point corresponds to 12 minutes and 5 seconds of the video, where the tension is extremely high; during this period, the camera captured a drastic change in the user's upper body posture, and the microphone captured a short "tsk!" sound; the system determined that a user behavior event Euser was triggered and strictly aligned it with the "skidding" scene in the video on the timeline.

[0124] The system analyzes the Euser and identifies the following three key sub-behaviors: Sub-behavior A: The diameter of both pupils instantly dilates (lasting 0.2 seconds); Sub-behavior B: The upper body leans back rapidly, adopting a "retreat to avoid" posture (lasting 0.5 seconds); Sub-behavior C: The right thumb hovers in the lower right corner of the screen but does not click (lasting 1.0 second).

[0125] For sub-behavior B (leaning backward): the system extracts the three-dimensional coordinates of key points in the chest cavity and calculates that the angle of the torso leaning backward is 35 degrees, with an angular velocity as high as 120 degrees / second (high explosive force); for sub-behavior C (hovering): the system extracts the fingertip contact area and finds that the pressure value fluctuates slightly but does not reach the click threshold (hesitation state).

[0126] For sub-behavior A: The system calculates the gaze trajectory and finds that the gaze focus point shifts rapidly from the center of the screen to the upper left corner (the direction of vehicle tilt), with a displacement vector of (x:−120,y:−45) pixels; the system constructs a multi-dimensional behavior feature vector Vbehavior=[tilt angle:35∘, angular velocity:120∘ / s, gaze offset:(−120,−45), thumb pressure:0.2N], this set of features accurately depicts the complex state of the user when seeing a thrilling scene: "startled, instinctively trying to avoid, wanting to pause but hesitating".

[0127] Therefore, a multimodal behavior perception network is used to process multidimensional behavioral features and output the corresponding behavioral intent. Based on the behavioral intent, the current playback progress of the short video and the corresponding contextual content, a primary operation instruction corresponding to the package is generated. This primary operation instruction is not an instruction to be executed immediately, but includes video jump, speed up playback, pause or immediate termination.

[0128] At this point, the system inputs the multi-dimensional behavior feature vector generated in step S142 into the trained multimodal behavior perception network (usually using a Transformer-based fusion architecture or a two-stream neural network); at the same time, the network receives the context features of the current video frame.

[0129] The network's internal attention mechanism calculates the mutual attention weights between behavioral features and video content features. For example, if a user performs a "rapid swipe" action and the video content is "boring dialogue," the network will assign a high relevance. If the video is "high-energy action," the network will determine that the user wants to "rewatch exciting moments."

[0130] After nonlinear transformation by the fully connected layer, the network outputs a probability distribution vector of behavioral intent. The system then uses Softmax normalization to select the category with the highest probability as the determined behavioral intent. Common intent categories include: repeated viewing, skipping boring content, reducing / increasing the information acquisition rate (speeding up), and terminating interaction.

[0131] The system reads the current playback progress (timestamp t) of the short video and maps it to the narrative structure model; the system determines the logical position of the current time point (e.g., "plot setup", "climax", or "ending"); the system modifies the intent based on the context; for example, if the detected intent is "pause", but the current playback progress is in the end credits scrolling area, the credibility of the intent is weakened; if it is at a critical moment with high information density, the "pause" intent is confirmed as valid; the system checks whether the newly generated instruction conflicts with the established continuous playback event (S141); if there is a conflict, the system evaluates the priority of the instruction and decides whether to interrupt continuous playback or suspend it.

[0132] Based on the verified behavioral intent, the system constructs a structured primary operation instruction, which includes the operation type and parameters:

[0133] Video jump: Includes the target timestamp or relative offset (such as Seek:+10s or Seek:Key_Frame_ID).

[0134] Playback speed adjustment: Includes playback rate adjustment (e.g., Playback_Rate: 2.0x or 0.5x);

[0135] Pause: Includes state transition signals (such as State_Change: PAUSE);

[0136] Immediate termination: Includes session end signals (such as Session_Action:TERMINATE).

[0137] The instruction is marked as "pending confirmation" or "preloaded" and temporarily stored in the instruction buffer queue. The system sets a short confidence verification window (e.g., 500 milliseconds). If the user's behavior changes during this period (e.g., canceling the gesture), the instruction will be revoked. If there is no change within the window, the instruction will be promoted to the "executable" state and sent to the rendering engine.

[0138] Specifically, the current playback is at the middle of episode 2, "Crossing the Dangerous Situation"; the system input S142 extracted multi-dimensional behavioral features (eye movement shift -120 pixels, right thumb hovering pressure 0.2N, lasting for more than 1.5 seconds); at the same time, the current video content is "a vehicle stuck in the mud, repeatedly attempting to spin in the mud," with high visual repetition and monotonous engine noise in the audio; the multimodal behavior perception network, through a cross-attention mechanism, found that the user's "eye movement shift" and "hovering without clicking" are highly correlated with the current "monotonous repetitive" video content; in the probability distribution output by the network, the confidence level of the "skip the current scene" category reaches 0.92; therefore, the system determines that the user's behavioral intention is "to advance the plot as quickly as possible."

[0139] The system checks the playback progress; the current timestamp is 08:20. According to the narrative structure script, the next 20 seconds are "the vehicle continues to spin idly," while there is a key plot point at 08:40: "the rescue vehicle arrives." The system determines that the interval from 08:20 to 08:40 is a low-value "anxious waiting period," which meets the conditions for skipping. The system confirms that the "skip" intention logic is valid and will not interrupt the user's acquisition of the core plot.

[0140] Based on the above intent and logic, the system generates a basic operation instruction:

[0141] Type: Video redirect;

[0142] Parameter: Target timestamp Ttarget=08:40 (i.e., the anchor point when "rescue vehicle arrives");

[0143] Pre-execution state: The instruction is marked as PENDING and stored in the buffer; the system begins to preload video frames around 08:40, waiting for the user's final confirmation signal (such as releasing a finger or specific micro-expression confirmation); once confirmed, the player will instantly switch from the scene of the mud spinning to the scene of the rescue vehicle appearing, achieving a smooth viewing experience optimization.

[0144] refer to Figure 6 In step S15, the specific steps are as follows:

[0145] S151: Align the primary operation command with the playback screen of the short video's continuous playback event on the timeline, and use the content recognition command to trigger the short video keyframe features, semantic scene nodes, and music rhythm phase corresponding to that moment.

[0146] S152: Determine the corresponding interactive area based on multi-factor matching of keyframe features of short videos, semantic scene nodes, music rhythm phase, and basic operation instructions; determine multiple corresponding interactive features based on the identification of the interactive area; and determine the corresponding matching level based on the mapping relationship table between multiple interactive features and matching levels.

[0147] S153: Collect the user's viewing history and determine the user's historical behavior pattern based on the user's viewing history. Based on the user's historical behavior pattern, the matching level, and the current context information of the short video, evaluate the experience prediction content for executing the primary instruction. Based on the multi-level iteration of the experience prediction content, determine the operation optimization features. Based on the multi-fusion of the operation optimization features and the primary operation instructions, determine the final operation instructions.

[0148] In the embodiments of this application, the system parses the timestamp information (such as Trigger) or relative offset carried in the primary operation command; simultaneously, the system reads the internal clock of the continuous playback event to determine the precise position of the current playback head (PTS, PresentationTimeStamp); the system uses linear interpolation or dynamic time warping to strictly align the command triggering time with the decoding timeline of the video stream; if the command contains an ambiguous time description (such as "current moment"), the system locks it to the system reference clock (SCR) value at the time the command was generated, and calculates the video frame sequence number (FrameSequenceNumber, FSN) corresponding to that moment; the system determines the specific phase of the command in the screen refresh cycle (V-Sync), and judges whether the command takes effect during the field blanking period of the frame or in the middle of the display of the current frame, thereby ensuring that the alignment operation will not cause screen tearing or jitter.

[0149] Based on the aligned time points, the system broadcasts a content recognition instruction, which is sent to each feature extraction submodule in the video processing pipeline to activate their computational kernels. The system simultaneously triggers the visual analysis channel, semantic understanding channel, and audio analysis channel, which ensures the temporal consistency of feature extraction. That is, visual features, semantic features, and audio features all originate from media data within the same time slice, avoiding modal mismatch problems caused by time drift.

[0150] Using convolutional neural networks (CNN) or visual transformers (ViT), the system extracts low-level visual features (such as color histograms, texture features, and edge gradients) and high-level semantic features (such as object detection boxes, human pose keypoints, and saliency heatmaps) from video frames at alignment time. These features constitute a pixel-level description of the image content.

[0151] The system uses temporal action localization or scene segmentation to analyze the position of the current frame in the entire narrative logic; the system determines which semantic scene node the current moment belongs to, such as "opening introduction", "conflict development", "climax", or "ending foreshadowing", which describes the plot structure and narrative function of the video content.

[0152] The audio stream is analyzed in the frequency domain using short-time Fourier transform (STFT) or wavelet transform to calculate the current beat tracking information. The system extracts the musical rhythm phase, that is, which beat point in the musical measure the current moment is (such as strong beat, weak beat, start beat, and end beat), as well as the instantaneous energy (loudness) and spectral centroid of the audio, which reflects the emotional driving force and rhythmic characteristics of the audio stream.

[0153] Specifically, the current continuous playback event is playing episode 3, "The Ultimate Climb"; the user intends to "pause to watch details" during viewing, and step S143 generates a basic operation command; the system reads the command and determines that the trigger time is 14 minutes 22 seconds and 150 milliseconds of the video playback progress (Ttrigger=14:22.150); the system strictly aligns this time with the timeline of the continuous playback event and determines that this time corresponds to the 21534th frame of the 3rd episode video (assuming a frame rate of 30fps); the system locks this frame as the reference frame for the command to take effect.

[0154] After locking the Trigger moment, the system immediately sends a content recognition instruction to the feature extraction module. This instruction requires all analysis channels to "freeze" at frame 21534 and its corresponding audio sampling point, in preparation for deep analysis.

[0155] The system analyzed frame 21534; visual feature extraction showed that the center of the frame was an off-road vehicle in a state of suspension, with all four wheels off the ground, and a huge cloud of dust billowing behind it; the salience heatmap was concentrated on the vehicle's central axis and the high horizon; combined with contextual analysis, the system determined that frame 21534 was at the climax of the plot, and marked the semantic scene node at this moment as "climax explosion stage: the moment the vehicle leaps," which means that this is the most crucial and visually impactful moment in the entire video; the system analyzed the audio at the same time point; the background music was a fast-paced electronic rock piece; the extraction results showed that 14:22.150 was exactly on the last strong beat of a measure, accompanied by a powerful bass drum hit, with the energy value reaching a local peak; the music rhythm phase was displayed as "Beat_Down".

[0156] Furthermore, the corresponding interaction area is determined by multi-factor matching based on keyframe features of short videos, semantic scene nodes, music rhythm phase, and basic operation instructions; multiple interaction features are determined based on the identification of the interaction area; and the corresponding matching level is determined based on the mapping relationship table between multiple interaction features and matching levels. This overall consideration of the mapping relationship table between multiple interaction features and matching levels ensures the accuracy of the corresponding matching level.

[0157] At this point, the system performs multimodal feature fusion on the keyframe features (visual salience, object edges), semantic scene nodes (plot stages, subject recognition), and music rhythm phase (beat strength) extracted in step S151. An attention mechanism is used to assign weights to different modalities. For example, in fast-paced action scenes, "visual salience" and "rhythm phase" are given higher weights.

[0158] Based on the fused feature map, the system computes an interaction heatmap in the computation space; it excludes semantically unsuitable areas for interaction (such as subtitle areas, UI control areas, and background areas blurred by rapid movement); during strong beats, it appropriately expands the radius of the interaction area to match the energy burst of the audio; during weak beats, it shrinks the interaction area to maintain finesse.

[0159] The system uses non-maximum suppression (NMS) or adaptive threshold segmentation to lock one or more continuous interactive regions in the interaction heatmap. This region is the spatial range with the highest visual attention, the strongest semantic importance, and the best fit with the audio rhythm, ensuring that subsequent interaction feature extraction is concentrated in this effective region.

[0160] The system calculates the geometric properties of the interactive area, including: area (degrees of freedom of interaction), aspect ratio (degree of screen adaptation), centroid position (degree of deviation from the visual center), and edge complexity; the system analyzes the rate of change of the interactive area over time; it calculates the motion magnitude and direction consistency of the optical flow of pixels within the area to determine whether the area is in a static, smooth motion, or violently chaotic state; the system evaluates the synchronization rate between the motion of the visual area and the audio rhythm; for example, it detects whether the expansion and contraction of objects within the interactive area are in sync with the music beat.

[0161] The system assembles multiple interaction features (geometric features, dynamic features, consistency features, etc.) into a high-dimensional interaction feature vector and normalizes it. The system inputs this vector into a predefined matching level mapping table, which can be a rule-based knowledge base (such as "high motion + strong rhythm > pause operation does not match") or a classifier model trained by machine learning.

[0162] The system calculates the similarity distance or confidence score between the interaction feature vector and each operation instruction category, and outputs the corresponding matching level; Level 1 (high matching): the content features and operation instructions match perfectly, resulting in an excellent execution experience; Level 2 (medium matching): the content features and operation instructions partially conflict, requiring parameter optimization; Level 3 (low matching): the content features and operation instructions severely conflict, leading to a very poor experience (such as disrupting the flow of the narrative).

[0163] Specifically, assuming that the short videos in Group A are part of the "Hardcore Off-Road Adventure" series, and the current playback is at the key frame of "Off-Road Vehicle Leaping into the Air" in episode 3 "Extreme Summit", S151 has determined that it is currently at the peak of the music and the visual climax, and the basic operation command is "Pause" (the user wants to see the details clearly).

[0164] The system identified the core of the image as the off-road vehicle located in the center of the screen (keyframe feature), with the semantic node being "leaping to a climax" and the music rhythm being "bursting with a strong beat." The system assigned a very high weight to the vehicle as the main subject, excluding the blurry sand and dust moving rapidly in the background. The system locked the interactive area to a rectangle in the center of the screen containing the body of the off-road vehicle (occupying about 60% of the screen area). This area avoided the bottom subtitle bar and the shaking background at the edges, making it the most stable and attention-grabbing area.

[0165] The system detected that within the interactive area, the off-road vehicle was at its highest point suspended in the air, with its vertical speed close to zero. The optical flow amplitude was extremely low at that instant (a "stagnant" state), which was a relatively stable visual moment. However, the audio characteristics showed that the background music's heavy drumbeat had just ended, the energy value was at its peak, and the aftersound was still oscillating. The system identified a strong conflict between "visual static" and "auditory dynamic." Although the image was at a static apex, the auditory experience was in a period of energy release.

[0166] The system inputs the feature vectors of "visual suspension + auditory strong beat + conflict state" into the matching level mapping table. The mapping table rules indicate that in the "auditory strong beat / energy release" phase, executing the "pause" operation will cause the audio signal to be unnaturally truncated, resulting in a serious "auditory disconnect". Although the picture is suitable for pausing, the overall multimedia experience is seriously negatively affected by the audio. The system determines that the matching level of this primary operation instruction (pause) is Level 2 (medium matching) or Level 3 (low matching, depending on the threshold setting). The system judges that directly executing "pause" is not a good experience and suggests that subsequent steps optimize the instruction (such as changing to slow motion).

[0167] Therefore, by collecting users' viewing history and determining their historical behavior patterns, and then evaluating the expected experience of executing primary instructions based on these patterns, the matching level, and the current context of the short video, operation optimization features are determined through multi-level iterations of these predicted experiences. Finally, the operation instructions are determined by combining these optimization features with the primary operation instructions, thus ensuring the accuracy of the final operation instructions. Furthermore, by identifying the primary operation instructions and combining them with the matching level and the user's viewing history, optimization of the primary operation instructions is triggered, thereby improving the accuracy of the final operation instructions and achieving precise control of the short video.

[0168] At this point, the system collects the user's historical viewing history from the database, including viewing duration, completion rate, frequency of actively dragging the progress bar, distribution of pause points, and usage habits of speeding up the playback. The system performs timestamp alignment and noise filtering on this heterogeneous data to eliminate abnormal operations (such as accidental touches).

[0169] By utilizing Sequential Pattern Mining or clustering methods (such as K-Means or DBSCAN), the system analyzes user operational patterns in specific video scenarios. Based on the mining results, the system constructs a vector of user historical behavior patterns, which contains multiple dimensions of preference indicators, such as: detail preference: whether users are accustomed to pausing and rewatching at exciting parts; rhythm sensitivity: users' tolerance for fast-paced editing; immersion threshold: what types of advertisements or interruptions users dislike.

[0170] The system constructs an experience prediction model (usually based on a reward function or decision tree of reinforcement learning). The input vector of the model consists of three parts: the user's historical behavior pattern (from S153); the current matching level (from S152, such as Level2 - low matching); and the current context information of the short video (such as plot tension, visual complexity, and audio energy).

[0171] The system simulates the result after executing a basic operation command; the model calculates the impact of executing the command on the user experience (UX) in the current context and matching level; for example, if the matching level is low and the user is sensitive to rhythm, the estimated experience score is negative (the user experience is impaired).

[0172] The system initiates a multi-level iterative optimization approach (such as gradient descent or genetic optimization). In each iteration, the system attempts to fine-tune the parameters of the instructions (such as adjusting the pause duration, changing the jump step size, and modifying the playback speed multiplier). The system repeatedly calculates the experience score after fine-tuning until the score converges to a local maximum. The system outputs a set of operation optimization features that maximize the experience score (such as: Play_Speed:0.5x, Transition:Smooth_Fade).

[0173] The system performs multiple fusions between the operation optimization features and the original primary operation instructions. This is not just parameter replacement, but a logical fusion based on a gating mechanism. For example, if the optimization feature suggests "do not execute", the original instruction is suppressed; if the optimization feature suggests "modify parameters", the parameter field of the original instruction is overwritten. The system generates a final operation instruction containing the optimized parameters, which is then sent to the player engine.

[0174] Specifically, in the current playback of episode 3, "Extreme Summit," during the moment when the off-road vehicle leaps through the air, S152 has determined that the matching level for executing the "pause" command at the strong beat of the music is Level 3 (low matching) because it would damage the listening experience. The system collected data on the user's viewing of off-road races over the past month; the system found that the user has a significant "slow-motion playback addiction"; in the past 20 leap scenes, the user manually reduced the playback speed to 0.25x to 0.5x 15 times, and rarely used a completely still "pause"; the system determined that the user's historical behavior pattern is "high attention to detail and preference for dynamic slow motion."

[0175] Predicted input: User mode (slow motion preference) + matching level (Level 3, direct pause provides a poor experience) + context (climax + strong music beat); First iteration: The system simulates executing the basic command "pause"; Predicted model output: Very low score, because abrupt music cut-off would disrupt immersion; Second iteration: The system attempts to modify the parameters to "0.5x speed playback"; Predicted model output: The score significantly improves, because the sound is sustained (becoming a deep, slow tempo) and satisfies the user's slow motion preference; Third iteration: The system attempts to modify the parameters to "0.2x speed + smooth transition"; Predicted model output: Highest score, and operation optimization features are identified.

[0176] The system integrates optimized features with the basic command "pause"; the logic determines that although the user's original intention is to "freeze the frame", under the goal of optimizing the experience, "extremely slow playback" can better achieve the deeper intention of "seeing details clearly" and "not destroying the listening experience"; the system generates the final operation command: "Switch the playback speed to 0.2x speed (slow motion mode) in a smooth transition with exponential decay"; the player will not abruptly freeze the picture, but like a movie special effect, it will smoothly enter the bullet time slow motion within 2 seconds, and the background music will slow down and become lower in sync.

[0177] In another embodiment of this application, the continuous playback system for short videos based on multimodal data includes:

[0178] The short video module is used to tag multiple short videos and determine the corresponding visual stream, audio stream, and text stream based on the recognition of each short video, and construct corresponding multimodal data based on the visual stream, audio stream, and text stream;

[0179] The key video segment module is used to collect users' historical viewing data, determine users' viewing preferences based on this historical viewing data and real-time context information, and identify key video segments in the process of matching viewing preferences with multimodal data.

[0180] The continuous playback instruction module is used to determine the corresponding emotional characteristics based on each key video segment and the corresponding user viewing behavior, and to determine the corresponding continuous playback instruction based on the emotional characteristics, the short videos corresponding to the key video segments, and the preset playback order of multiple short videos.

[0181] The primary operation instruction module is used to construct the corresponding continuous playback event of the short video when triggered by the continuous playback instruction, and to collect the multi-dimensional behavioral characteristics of the user in the continuous playback event of the short video, and determine the primary operation instruction based on the multi-dimensional behavioral characteristics.

[0182] The matching module is used to match the playback screen of the initial operation command and the continuous playback event of the short video in the same time dimension to determine the corresponding matching level, and then determine the final operation command based on the matching level and the user's viewing history.

[0183] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1.A method for continuous playing of short videos based on multi-modal data, characterized in that, The method comprises the following steps: Marking a plurality of short videos and determining corresponding visual streams, audio streams and text streams based on the identification of each short video, and constructing corresponding multi-modal data according to the visual streams, audio streams and text streams; Collecting historical viewing data of the user, determining the favorite viewing features of the user based on the historical viewing data and real-time context information, and determining key video segments in the matching process of the favorite viewing features and the multi-modal data; Determining corresponding emotional features based on each key video segment and the corresponding user viewing behavior, and determining corresponding continuous playing instructions according to the emotional features, the short video corresponding to the key video segment and the preset playing order of the plurality of short videos; Constructing a continuous playing event of the corresponding short video under the triggering of the continuous playing instruction, collecting multi-dimensional behavior features of the user in the continuous playing event of the short video, and determining primary operation instructions according to the multi-dimensional behavior features, including: determining a plurality of time axis anchors based on the identification of the continuous playing instruction under the triggering of the continuous playing instruction, determining corresponding playing content according to the tracing of each time axis anchor, and constructing a continuous playing event of a short video containing time axis anchors based on the plurality of time axis anchors, the corresponding playing content and the corresponding content weight; determining the behavior event of the user based on the analysis of the continuous playing event of the short video, and determining a plurality of sub-behavior contents in the behavior event, determining corresponding multi-dimensional behavior features according to the plurality of sub-behavior contents, the action features and the action trajectories of the user; processing the multi-dimensional behavior features by using a multi-modal behavior perception network, and outputting corresponding behavior intentions, generating corresponding primary operation instructions according to the behavior intentions, the playing progress of the current short video and the corresponding context content, which are not immediate execution instructions, including video jumping, fast playing, pausing or immediate termination; Matching the primary operation instruction and the playing picture of the continuous playing event of the short video in the same time dimension to determine the corresponding matching level, and determining the final operation instruction according to the matching level and the viewing history of the user, including: aligning the primary operation instruction and the playing picture of the continuous playing event of the short video on the time axis, and triggering the short video key frame features, semantic scene nodes and music rhythm phases corresponding to the aligned time points by using the content recognition instruction; determining the corresponding interaction area in the picture according to the multi-factor matching of the short video key frame features, semantic scene nodes and music rhythm phases and the primary operation instruction; determining a plurality of interaction features according to the identification of the interaction area, determining the corresponding matching level according to the mapping relationship table of the plurality of interaction features and the matching level; collecting the viewing history of the user, determining the historical behavior mode of the user according to the viewing history of the user, evaluating the experience estimation content of executing the primary instruction based on the historical behavior mode of the user, the matching level and the current context information of the short video, determining the operation optimization features according to the multi-level iteration of the experience estimation content, and determining the final operation instruction according to the multi-fusion of the operation optimization features and the primary operation instruction. 2.The method of claim 1, wherein, The plurality of short videos are marked, and corresponding visual streams, audio streams and text streams are determined based on the identification of each short video, and the corresponding multi-modal data is constructed according to the visual streams, audio streams and text streams, including: In the database of short videos, a plurality of short videos in the database of short videos are marked, and the marked short videos are frame-by-frame parsed to trigger feature extraction in different dimensions. At this time, in the visual dimension, the visual stream containing the HSV color space and the key frame object is extracted based on the convolutional layer network; in the auditory dimension, the audio stream containing the background music rhythm and the speech fundamental frequency change is extracted based on the audio processing network; in the semantic dimension, the text stream containing the video embedded subtitles and the key frame picture text is extracted based on the natural language processing network; The extracted visual stream, audio stream and text stream are aligned and mapped, and are projected into a preset semantic space. In the preset semantic space, the interaction weight between modalities is calculated through an attention mechanism, and the multi-modal data unique to each short video is constructed through deep fusion. 3.The method of claim 1, wherein, The historical viewing data of the user is collected, the favorite viewing feature of the user is determined based on the historical viewing data and real-time context information, and the key video segment is determined in the matching process of the favorite viewing feature and the multi-modal data, including: The data space of the user is determined based on the database of short videos and the user's input information, and the historical viewing data of the user is determined based on the traversal of the data space. The long-term interest preference of the user is mined by using a long short-term memory network. At the same time, the real-time context information is obtained by combining the current timestamp, geographic location information and audio-visual features of the previous video; the favorite viewing feature of the user is determined according to the multi-dimensional fusion of the long-term interest preference of the user and the real-time context information; According to the favorite viewing feature and the multi-modal data, a corresponding similarity combination is determined, and a candidate video set with high correlation is determined along the high-dimensional matching of the similarity combination and the database of short videos. Based on the candidate video set and the corresponding spatio-temporal attention mechanism, a plurality of frame-level features of the candidate video are determined, and the corresponding key video segment is determined by combining the plurality of frame-level features. 4.The method of claim 1, wherein, The emotion feature corresponding to each key video segment and the corresponding user viewing behavior is determined, and the continuous playback instruction corresponding to the emotion feature, the short video corresponding to the key video segment and the preset playback order of the plurality of short videos is determined, including: The key video segment is subjected to sentiment analysis, and the emotional content corresponding to the key video segment is determined in the sentiment analysis. At the same time, the user viewing behavior corresponding to the key video segment is determined based on the action detection of the key video segment. The emotional content and the corresponding user viewing behavior are weighted and corrected, and the corresponding emotion feature is constructed. 5.The method of claim 4, wherein, The emotion feature corresponding to each key video segment and the corresponding user viewing behavior is determined, and the continuous playback instruction corresponding to the emotion feature, the short video corresponding to the key video segment and the preset playback order of the plurality of short videos is determined, including: The preset playing order of the plurality of short videos is collected, and the emotion feature is spatiotemporally aligned with the playing progress and the narrative structure of the current short video, and is dynamically deduced in combination with the preset playing order of the plurality of short videos to generate an emotion smoothing curve capable of avoiding emotional mutation; A plurality of emotion nodes are determined based on the recognition of the emotion smoothing curve, an emotion change path is constructed according to the plurality of emotion nodes and the corresponding key video segments, a playing queue of the plurality of short videos is determined along the analysis of the emotion change path, and corresponding continuous playing instructions are presented in the playing queue. 6.A system for continuous playing of short videos based on multi-modal data, characterized in that, The continuous playing system of the short video based on the multi-modal data is applied to the continuous playing method of the short video based on the multi-modal data as claimed in any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-mode short video tag recommendation method fusing emotional information

    CN115329127A

  • Video generation method and system based on music rhythm and television on-demand method

    CN118301382A