Video processing method and device, storage medium and electronic equipment

Key video frames are obtained through gesture interaction and multi-modal recognition. The processing model is used to generate personalized annotated content, which solves the problem of poor display flexibility of video content and improves the learning efficiency and interactivity of video viewing.

CN120378698APending Publication Date: 2025-07-25HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510713119.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has poor flexibility in displaying the associated information of video content. It is impossible to adjust or add new associated knowledge information in real time according to user personalized needs, and cannot meet the user's needs for the unspoken part.

Method used

Key video frames are obtained through gesture interaction, multi-modal content recognition is performed, and key search information matching the video viewing progress is generated using the target processing model, and annotated content is dynamically displayed in the video screen.

Benefits of technology

It realizes the dynamic display of personalized annotated content based on user needs, improves the learning efficiency and interactive experience of video viewing, and enriches the flexibility and interactivity of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378698A_ABST
    Figure CN120378698A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method and device, a storage medium and electronic equipment. The method comprises the steps that in the playing process of a target video, in response to triggering operation on video frames, key video frames are acquired, and the triggering operation comprises the triggering operation executed in a gesture interaction mode and marking processing executed on visual elements in the video frames; the method comprises the following steps: performing multi-modal content identification on a key video frame to obtain multi-modal search information; multi-modal search information and historical behavior data are input into a target processing model, a set of key search information matched with the video watching progress is obtained, and the historical behavior data comprise operation data executed on the video frame based on the multi-modal search information; and dynamically displaying the annotation content associated with the information based on the group of key search in the video picture based on the video watching progress. The technical problem of poor flexibility in the process of displaying the associated information of the video content is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing, and in particular, to a video processing method, apparatus, storage medium, and electronic device. Background Art

[0002] During the process of user video playback, the user may be confused about some knowledge points or scenes in the video content. For example, when watching historical figures in historical videos, the user may want to know the era, achievements, and other information of the historical figures; when watching educational videos, the user may search for unfamiliar technical terms or concepts.

[0003] To meet the above needs of users, in the related art, a method of marking pre-set key video frames is usually adopted, that is, developers or video content producers manually indicate important moments in the video and pre-associate corresponding information, so that the associated information (which can also be understood as annotation content) is displayed when the video playback progress reaches the pre-marked important moment.

[0004] However, the above method of pre-marking video frames can only meet the user's need for the associated information of the marked video content, cannot adjust or add new associated knowledge information in real time according to the personalized needs of each user, and cannot meet the user's need to understand the associated knowledge of the unmarked part of the video, resulting in a technical problem of poor flexibility in the process of displaying the associated information of the video content. For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present application provide a video processing method, apparatus, storage medium, and electronic device to at least solve the technical problem of poor flexibility in the process of displaying the associated information of video content.

[0006] According to one aspect of the embodiments of the present application, a video processing method is provided, including: during the playback of a target video, in response to a trigger operation on a video frame, obtaining a key video frame, where the trigger operation includes a trigger operation performed by a gesture interaction method and a marking process performed on a visual element in the video frame; obtaining multi-modal search information by performing multi-modal content recognition on the key video frame; obtaining a set of key search information matching the video viewing progress by inputting the multi-modal search information and historical behavior data into a target processing model, where the historical behavior data includes operation data performed on video frames based on multi-modal search information within a historical period; and dynamically displaying annotation content associated with a set of key search information in the video picture based on the video viewing progress.

[0007] Optionally, during the playback of the target video, in response to a triggering operation on a video frame, obtaining a key video frame includes: obtaining trajectory data of a touch gesture performed based on a gesture interaction method; determining a gesture type matching the gesture interaction method based on the trajectory data; performing a corresponding triggering operation on a first part of the video frames in the video frame based on the gesture type; and determining the first part of the video frames as the key video frame.

[0008] Optionally, the above method further includes: in response to performing a corresponding triggering operation on the first part of the video frames, performing a marking process on a first part of the area of the video picture corresponding to the first part of the video frames; and obtaining multimodal search information by performing multimodal content recognition on visual elements within the first part of the area.

[0009] Optionally, during the playback of the target video, in response to a triggering operation on a video frame, obtaining a key video frame includes: during the playback of the target video, in response to a pause operation performed on the target video picture corresponding to each video frame in a second part of the video frames at different times, performing a marking process on a second part of the area in the target video picture; and determining the second part of the video frames as the key video frame.

[0010] Optionally, obtaining a set of key search information matching the video viewing progress by inputting the multimodal search information and historical behavior data into a target processing model includes: obtaining a score for each search information in the multimodal search information by inputting the multimodal search information and historical behavior data into the target processing model, where the higher the score of the search information, the higher the demand for displaying target annotation content associated with the search information during the playback of the target video; and screening out a set of key search information from the multimodal search information based on the scores.

[0011] Optionally, dynamically displaying annotation content associated with a set of key search information in the video picture based on the video viewing progress includes: in the case of first viewing the target video through a first account and obtaining a set of key search information at a first time, in response to at least one key search information in the set of key search information appearing at a second time, displaying at least one annotation content associated with the at least one key search information at the second time based on the video viewing progress; where the second time corresponding to the second moment is later than the first time corresponding to the first moment.

[0012] Optionally, the above method further includes: when the target video is watched for the first time through the first account and a set of key search information is obtained at the first moment, in response to playing back the target video through the first account, locating the target playback progress at which at least one key search information appears in the video picture; based on the target playback progress, determining a third moment, where the third moment corresponds to the same playback time displayed at the same playback progress as the second moment; based on the video viewing progress, displaying at least one annotation content associated with at least one key search frequency information at the third moment.

[0013] Optionally, the above method further includes: when watching the target video through the first account and obtaining a set of key search information, sharing the target video to the second account through the first account; during the process of watching the target video through the second account, dynamically displaying in the video picture the annotation content associated with a set of key search information.

[0014] According to another aspect of the embodiments of the present application, there is also provided a video processing device, including: a first acquisition unit, configured to obtain a key video frame in response to a trigger operation on a video frame during the playback of the target video, where the trigger operation includes a trigger operation performed by a gesture interaction method and a marking process performed on a visual element in the video frame; a first processing unit, configured to obtain multi-modal search information by performing multi-modal content recognition on the key video frame; a second processing unit, configured to input the multi-modal search information and historical behavior data into a target processing model to obtain a set of key search information that matches the video viewing progress, where the historical behavior data includes operation data performed on the video frame based on the multi-modal search information within a historical period; a first display unit, configured to dynamically display in the video picture the annotation content associated with a set of key search information based on the video viewing progress.

[0015] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is used to execute the above video processing method when being run by an electronic device.

[0016] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including a computer program, where the steps of the above method are implemented when the computer program is executed by a processor.

[0017] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the above video processing method through the computer program.

[0018] By adopting the above embodiments provided in the present application, by using a multi-modal triggering method including a gesture interaction method, an associated video frame for expecting to display annotation content associated with video content is obtained. By identifying the multi-modal content in the associated video frame, multi-modal search information is obtained, and a target processing model is used to determine a set of key search information most relevant to the user's needs. Finally, based on the video viewing progress, annotation content associated with each key search information is dynamically displayed in the video picture. In other words, during the playback of the target video, touch marks are added according to the user's needs, and key search information is identified from the marked key video frames, so as to display annotation content that meets the user's personalized needs in the video picture, improving the learning efficiency during video viewing, enriching the interactive experience of video viewing, and achieving the technical effect of improving the flexibility of displaying annotation content of video content. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application.

[0020] Figure 1 It is a schematic diagram of an application scenario of an optional video processing method according to an embodiment of the present application;

[0021] Figure 2 It is a flowchart of an optional video processing method according to an embodiment of the present application;

[0022] Figure 3 It is an overall flowchart of an optional video processing method according to an embodiment of the present application;

[0023] Figure 4 It is an example 1 of an optional display of annotation content according to an embodiment of the present application;

[0024] Figure 5 It is an example 2 of an optional display of annotation content according to an embodiment of the present application;

[0025] Figure 6 It is an example 3 of an optional display of annotation content according to an embodiment of the present application;

[0026] Figure 7 It is a schematic structural diagram of an optional video processing device according to an embodiment of the present application;

[0027] Figure 8 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0029] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0030] The technical solutions in the embodiments of this application will comply with legal regulations during the implementation process. When operating according to the technical solutions in the embodiments, the data used will not involve user privacy. While ensuring that the operation process is compliant and legal, the security of the data is guaranteed. In addition, when the above embodiments of this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant regulations and standards of the relevant country or region.

[0031] According to one aspect of the embodiments of this application, a video processing method is provided. As an alternative implementation, the above video processing method can be but is not limited to being applied to an application scenario such as Figure 1 shown. In the Figure 1 shown application scenario, the target terminal 102 can communicate with the server 106 through the network 104, and the server 106 can perform operations on the database 108, such as write data operations or read data operations. The above target terminal 102 can include but is not limited to a human-computer interaction screen, a processor, and a memory. The above human-computer interaction screen can be used to display finished files, etc. of the target video played on different video platforms on the target terminal 102. The above processor can be used to respond to the above human-computer interaction operations, perform corresponding operations, or generate corresponding instructions and send the generated instructions to the server 106. The above memory is used to store relevant processing data, such as data change messages, target video identifiers, and target video subtitle notifications.

[0032] Optionally, in this embodiment, the above target terminal may be a terminal configured with a target client, and may include, but are not limited to, at least one of the following: mobile phone (such as Android mobile phone, iOS mobile phone, etc.), laptop computer, tablet computer, handheld computer, MID (Mobile Internet Devices), PAD, desktop computer, smart TV, etc. The target client may be a video client, instant messaging client, browser client, education client, etc. The above network may include, but is not limited to: wired network, wireless network, wherein the wired network includes: local area network, metropolitan area network and wide area network, and the wireless network includes: Bluetooth, WIFI and other networks implementing wireless communication. The above server may be a single server, or a server cluster composed of multiple servers, or a cloud server.

[0033] The technical solution of this application can be widely applied to scenarios such as optimizing the interaction experience of users during video viewing and the acquisition efficiency and flexibility of associated indication information. The following are specific examples of the annotation content for displaying video content in several application scenarios:

[0034] (1) Historical education video viewing: When a user views a video related to historical events or historical figures, the technical solution of this application can automatically identify key historical figures and locations in the picture, generate search keywords, and provide relevant historical backgrounds, biographies of figures, and other supplementary materials to help the audience better understand the video content;

[0035] (2) Professional courses and online education: When viewing complex educational videos, especially science, engineering, or medical courses, users may encounter professional terms or concepts that are difficult to understand. The technical solution of this application promotes autonomous learning and improves learning efficiency by automatically identifying professional terms and linking to detailed explanations and relevant examples;

[0036] (3) Sports event replay: When watching a sports game, specific moments of athletes are marked and shared, and at the same time, the system provides motion analysis, rule explanations, and technical guidance for that moment, enhancing the interactivity and educational value of watching the game;

[0037] (4) Technology demonstration and product review: For product demonstrations or technology content, the technical solution of this application can identify and annotate technical features, parameter specifications, or operation steps in the video to help users quickly master information and improve decision-making quality.

[0038] In summary, the technical solution of the present application can be applied to a variety of different scenarios. It can respond to user needs in real time, intelligently extract key search information of video content, give hierarchical search suggestions based on the context, and support scenario-based sharing of video clips, thereby providing users with rich and personalized additional information without interrupting the video viewing process, effectively improving the explorability and interactivity of video content, and promoting the dissemination of professional knowledge and in-depth learning.

[0039] As described in the above embodiments, the conventional method of displaying the associated information of video content is prone to the problem of poor flexibility. In order to solve the above problem, a video processing method is proposed in the embodiments of the present application. Figure 2 is a flow chart of a video processing method according to an embodiment of the present application, and the flow chart includes the following steps S202 to S208.

[0040] It should be noted that the video processing method shown in step S202 to step S208 can be executed by, but not limited to, an electronic device, wherein the electronic device can be, but not limited to, Figure 1 The target terminal or server is shown.

[0041] Step S202, during the target video playback process, in response to a trigger operation on a video frame, obtaining a key video frame, wherein the trigger operation includes a trigger operation performed by gesture interaction and a marking process performed on a visual element in the video frame;

[0042] Step S204, obtaining multimodal search information by performing multimodal content recognition on key video frames;

[0043] Step S206, obtaining a set of key search information matching the video viewing progress by inputting the multimodal search information and the historical behavior data into the target processing model, wherein the historical behavior data includes operation data performed on the video frame based on the multimodal search information within the historical period;

[0044] Step S208: based on the video viewing progress, dynamically display the annotation content associated with a set of key search information in the video screen.

[0045] Trigger operations include but are not limited to specific actions of users interacting with the video content being played, which can be achieved through gesture recognition (such as sliding and circling on the touch screen) or directly marking specific elements in the video frame (such as people, professional terms). For example, a user can use a finger to draw a circle on the screen to circle a person or object of interest in the video screen, so as to use it as a key object for subsequent analysis.

[0046] The above gesture interaction includes, but is not limited to, the way that users interact with video content by performing touch operations on the screen of a mobile device. In the embodiments of this application, it is possible to, but not limited to, capture the user's gesture trajectory through an event listener, and use the Dynamic Time Warping (DTW) algorithm to match with a predefined gesture template to identify the user's intention. For example, a "circle" gesture may mean that the user wants to focus on at least one specific visual element, while a "swipe" gesture may indicate that the user hopes to browse more relevant information.

[0047] The marking process of visual elements includes, but is not limited to, the user's attention to and selection of specific parts in a video frame, which not only includes the recognition and circling of text in the video, but also includes freehand strokes or automatic adsorption circling of people and objects in the image, enabling the user to finely select the focus points and laying a foundation for subsequent information extraction.

[0048] The multi-modal search information includes, but is not limited to, video information in multiple modalities such as text and images obtained by recognizing key video frames. For example, the avatar and clothing of a historical task, a screenshot of a mathematical formula, etc.

[0049] In this embodiment, it is possible to, but not limited to, use the OCR (Optical Character Recognition) recognition method to perform content recognition on key video frames. Among them, OCR is a technology that converts the text content in an image into editable text. It obtains the image by scanning, photographing or other means, and then uses computer algorithms to recognize the characters in the image and convert them into an electronic text format. In addition to using OCR recognition, image recognition technologies such as YOLOv8 object detection and PointNet++ feature encoding can also be used to identify and understand non-text visual elements in video frames. These technologies can help identify people, objects, shapes and structures.

[0050] In this embodiment, the target processing model can, but is not limited to, implement a model for context association based on marked content during video playback and generate hierarchical search suggestions. The target processing model is an advanced information filtering and recommendation system. Based on the user's viewing history and preferences, by analyzing multi-modal search information and historical behavior data, it provides search suggestions that are most relevant to the current video viewing progress. For example, if a user has searched for a certain historical figure more frequently in the past, then relevant materials about that figure will be recommended first.

[0051] For the same video, the historical behavior data contains operation records performed by multiple users on the video content during the historical period, such as search frequency, circled content, and personal notes added by users. These data help the system understand the specific needs of users, so as to provide more personalized search suggestions and content recommendations.

[0052] Finally, the system dynamically presents the associated annotation content on the video screen according to the time point of video playback, ensuring the timeliness and relevance of the information.

[0053] Dynamically displaying the annotation content means providing explanatory information associated with the current video content in a timely manner without affecting video playback, such as character introductions, explanations of technical terms, or historical background materials, usually in the form of a floating sidebar window. This display method not only retains the continuity of video viewing but also increases the convenience of knowledge acquisition.

[0054] It should be noted that the triggering operation performed on the video frames in the above manner can lock the key video frames in the video file and obtain the marking of specific regions or key visual elements in the key video frames, that is, obtain the user-marked regions or visual elements with marks.

[0055] Through multimodal recognition of the above user-marked regions or visual elements with marks, multimodal search information is obtained. The following gives specific implementation methods for recognizing user-marked regions or visual elements with marks, as well as specific examples of how to generate multimodal search information (which can also be understood as key search terms) based on the recognized multimodal search information.

[0056] Example 1

[0057] Regarding the method for recognizing the core visual elements in the user-marked region, the implementation steps of the computer vision integration solution are as follows:

[0058] S11, detect the visual elements in the marked region;

[0059] Input the key video frame marked by the user and identify the element type in the key video frame. For example, if the visual element is identified as a person through OCR, then further detect the key points of the face and generate a semantic mask.

[0060] S12, perform cross-modal data processing on the above semantic mask;

[0061] (1) Spatiotemporal context modeling: mainly use the LSTM (Long Short-Term Memory) network to analyze the feature changes between adjacent video frames, enhancing the recognition robustness in dynamic scenarios.

[0062] Among them, LSTM is a special recurrent neural network architecture used to solve the problem of gradient disappearance or gradient explosion in traditional recurrent neural networks when dealing with long sequence data. LSTM controls the flow of information by introducing a "gate" structure and can effectively learn and remember long-distance dependencies.

[0063] (2) Hardware acceleration: Achieve CUDA-level parallel computing through WebGPU or Metal API to ensure the real-time performance of complex models on mobile devices.

[0064] S13, Hierarchical processing mechanism;

[0065] One is the low-power mode, which can but is not limited to performing fast detection through 720p resolution + MobileNetV3; the other is the high-performance mode, which performs tracking detection through 4K resolution + EfficientDet-D0 combined with optical flow method.

[0066] S14, Memory management scheme: Implement an LRU frame cache pool to limit the number of video frames processed simultaneously, and at the same time achieve zero-copy transfer of CPU-GPU data.

[0067] In the technical solution of the above-mentioned Embodiment 1, the automatic clarification of the selected area and the synchronous startup of multi-modal content recognition can also be achieved through the technical solution in the following Embodiment 2, which will be described in detail below through Embodiment 2.

[0068] Embodiment 2

[0069] In the process of realizing the automatic clarification of the marked area and the synchronous startup of multi-modal content recognition, it is necessary to construct a complete processing flow including image enhancement, OCR, graphic understanding, and cross-modal retrieval, which specifically includes the following steps:

[0070] S21, Dynamic super-resolution enhancement;

[0071] (1) Blur detection

[0072] Adopt joint judgment of frequency domain analysis (FFT energy entropy) and spatial domain analysis (Laplacian gradient variance):

[0073] def calculate_blur_score(image):

[0074] gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

[0075] laplacian_var = cv2.Laplacian(gray, cv2.CV_64F).var()

[0076] fft_energy = np.log(np.mean(np.abs(np.fft.fft2(gray)) ** 2))

[0077] return 0.7*laplacian_var+0.3*(1-fft_energy / 1e5) # Normalize to the range of 0-1

[0078] (2) Conditional enhancement strategy

[0079] if blur_score<0.4:

[0080] sr_image = ESRGAN_model.predict(region)

[0081] else:

[0082] sr_image = cv2.filter2D(region, -1, np.array([[-1, -1, -1], [-1, 9, -1],

[0083] [-1, -1, -1]]))

[0084] S22, Multimodal content recognition - OCR processing flow;

[0085] graph TD

[0086] A[Region extraction] --> B{Language detection}

[0087] B --> |Chinese / Japanese| C[PP-OCRv3]

[0088] B --> |Latin alphabet| D[Tesseract 5.0]

[0089] C & D --> E[Text normalization]

[0090] E --> F[Knowledge graph matching]

[0091] S23, Multimodal content recognition - Key technical points;

[0092] (1) Multilingual mixed detection: Can but not limited to use the FastText pre-trained model to achieve an accuracy of 0.92+;

[0093] (2) Format unification: Implement table / paragraph structure restoration through the layoutparser module of PaddleOCR;

[0094] (3) Knowledge base matching: Build a FAISS vector library for semantic vector retrieval of the recognition results.

[0095] S24, Graphics processing flow;

[0096] graph LR

[0097] A[Feature Extraction] --> B{Shape Classification}

[0098] B --> |Geometric Shapes| C[OpenCV Hough Transform]

[0099] B --> |Complex Objects| D[PointNet++ Feature Encoding]

[0100] C & D --> E[3D Reconstruction]

[0101] E --> F[CLIP Cross-Modal Search]

[0102] S25, Implementation of details.

[0103] (1) Shape Classification: Use pre-trained MobileNetV3 for binary classification (geometric / non-geometric);

[0104] (2) 3D Modeling: Based on Open3D, implement point cloud registration, and support generating.obj models with depth information from a single image;

[0105] (3) Similar Image Search: Use Sentence-BERT to calculate the cosine similarity between image features and database vectors.

[0106] After detecting and recognizing the user-labeled area through the above-mentioned Example 1 and Example 2, multi-modal search information can be generated according to the recognition content in the following ways, but not limited to these.

[0107] Assume that the text information is obtained after recognizing the visual elements in the user-labeled area through the above method. Then, multi-modal search information can be obtained through the extraction of candidate words. Among them, the extraction of candidate words includes the following steps:

[0108] S31, Associated database matching;

[0109] For example, video-related items, people, professional nouns, technical terms, etc.

[0110] S32, Search popularity associated matching;

[0111] For example, matching the historical search terms and popular search terms of the target video, etc.

[0112] S33, Extract nouns in the text string.

[0113] By responding to the user's needs in real time as described in the above embodiments, providing instant, rich, and relevant information queries greatly enhances the interactivity and educational value of video viewing; the search suggestions generated based on the user's historical behavior and interest points provide customized content recommendations for each user, meeting the learning needs at different levels.

[0114] By integrating visual and text recognition, as well as the analysis of user behavior data, a comprehensive content enhancement and interaction mechanism is formed, which is applicable to a wide range of application scenarios, from historical and educational videos to professional tutorials, etc.

[0115] In summary, the technical solution of this application realizes the in-depth understanding and enhancement of video content through multi-modal information recognition, personalized recommendation, and intelligent annotation technology, providing users with a richer, personalized, and interactive viewing experience.

[0116] As an optional example, during the playback of the target video, in response to a trigger operation on a video frame, obtaining a key video frame includes:

[0117] Obtaining trajectory data of a touch gesture performed based on a gesture interaction method;

[0118] Based on the trajectory data, determining a gesture type that matches the gesture interaction method;

[0119] Based on the gesture type, performing a corresponding trigger operation on the first part of the video frames in the video frame;

[0120] Determining the first part of the video frames as the key video frames.

[0121] In this embodiment, the implementation process of obtaining key video frames through gesture interaction is further refined. Specifically, the trajectory data of the gesture trajectory is captured through an event listener, and then the dynamic time warping (DTW) algorithm is used to match a preset gesture template (for example, circle selection, sliding, pinching, etc.). The specific process is as follows:

[0122] Step 1: Obtaining trajectory data of a touch gesture performed based on a gesture interaction method;

[0123] Among them, the trajectory data of the touch gesture refers to the movement trajectory left by the user during the touch operation on the screen, recorded in the form of a series of coordinate points.

[0124] This step is implemented through a custom View and an event listening mechanism. Whenever the user performs a touch operation on the screen, such as drawing a circle, sliding, or clicking, the system will monitor these events and collect the coordinate information of the touch points.

[0125] Specific implementation process: In the video playback interface, the underlying custom View layer continuously monitors touch events (such as touchstart, touchmove, touchend). When it detects that the user starts to touch the screen, the system immediately starts recording the position coordinates of the touch points until the user ends the touch. These coordinate data form the trajectory of the user's gesture, providing a basis for further gesture recognition and video frame processing.

[0126] Step 2: Based on the trajectory data, determine the gesture type that matches the gesture interaction method;

[0127] Among them, the gesture type refers to the specific touch mode performed by the user on the screen, including but not limited to circle selection, sliding, zooming, etc. The system will use the DTW algorithm to compare the gesture templates and recognize the user's intention.

[0128] Specific implementation process: Through preprocessing steps (such as normalization and path simplification), the system converts the trajectory data of the user's gesture into a format that is easy to compare. Then, the dynamic time warping (DTW) algorithm is used to compare the processed trajectory with the preset gesture templates to find the best match.

[0129] This algorithm can effectively identify the dynamic changes of gestures. Even if the speed or amplitude of the user's gesture is slightly different, it can accurately identify, ensuring the robustness of gesture recognition.

[0130] Step 3: Based on the gesture type, perform corresponding trigger operations on the first part of the video frames in the video frame;

[0131] Among them, the first part of the video frames includes but not limited to the video frames of the area circled or marked by the user through gestures. These video frames are the frame segments that need to be focused on processing and analyzing to extract valuable information.

[0132] The above trigger operations can but are not limited to being different according to the different gesture types of the user. For example, if it is recognized that the user has performed a "circle selection" gesture, the system will process the video frames within the circled area and extract text or image features; if it is a "sliding" gesture, it may trigger functions such as video fast forward or playback.

[0133] Specific implementation process: After recognizing the gesture type, the corresponding trigger operation will be executed, such as recognizing the image and text content of the circled video area. This step can but is not limited to using advanced OCR technology and image recognition algorithms. For example, MobileNet-SSD or YOLOv8 is used to detect objects in the image, and PP-OCRv3 is used to recognize text content, ensuring the accuracy and speed of information extraction.

[0134] Step 4: Determine the first part of the video frames as key video frames.

[0135] Among them, the key video frames refer to the video frames that have been circled or marked by the user and are confirmed by the system to contain valuable information. These frames are the basis for subsequent in-depth analysis and information retrieval.

[0136] Specific implementation process: The images and texts recognized within the video frame area marked by the user are converted into core visual elements and search keywords, thereby determining these video frames as key video frames. Subsequently, the key video frames will be sent to subsequent processing stages. For example, in-depth content correlation analysis and search term generation are performed to provide users with immediate and relevant information.

[0137] By adopting flexible gesture interaction technology, personalized exploration of video content by users is achieved, improving the interactivity and practicality of video viewing. Users can use touch gestures to circle or mark the interesting scenes, thereby extracting the information in these video scenes and providing hierarchical search suggestions based on the video viewing progress.

[0138] The above method not only enhances the user experience, making knowledge acquisition more convenient, but also promotes the in-depth understanding and personalized learning of video content, especially showing great potential in fields such as educational videos and historical documentaries. By intelligently matching user operations with video content, each frame of the video can become an entrance for knowledge exploration and sharing.

[0139] As another alternative example, in addition to the triggering operations performed by the above gesture interaction method, the key video frames can also be determined by using the method of manual marking by the user, specifically including:

[0140] In response to performing the corresponding triggering operation on the first part of the video frames, mark processing is performed on the first part of the area of the video screen corresponding to the first part of the video frames;

[0141] Through multi-modal content recognition of the visual elements within the first part of the area, multi-modal search information is obtained.

[0142] The above mark processing refers to the system circling a specific area in the video screen according to the user's gesture, preparing for subsequent content recognition and information search. Users can freely circle or click on the screen, and the system optimizes the clarity of the selected area based on the user's mark, providing high-quality input for text and image recognition.

[0143] Specifically, when the user selects a certain part of the video frame as the first part of the region through gestures during video playback, the system immediately responds and uses image processing techniques (such as sharpening, noise reduction) to improve the picture quality of this region. Subsequently, the OCR text recognition and image recognition processes are initiated. For text recognition, the system uses PP-OCRv3 or similar technologies, which can accurately recognize text information in complex backgrounds, including different fonts, sizes, and languages. In terms of image recognition, advanced algorithms such as PointNet++ or YOLOv8 are used to recognize and understand non-text visual elements such as people, objects, and environments in the image. The advantage of this process is that it can intelligently process any region marked by the user. Even for non-linear freehand selection, the system can automatically recognize and optimize it, ensuring the flexibility and accuracy of information extraction.

[0144] Secondly, the first part of the region of the marked video frame is sent to the multi-modal content recognition module. Here, in addition to traditional text and image feature extraction, the system also uses spatio-temporal context modeling to analyze the feature changes of adjacent video frames through the LSTM network, enhancing the recognition robustness in dynamic scenarios. In addition, the hardware-software collaborative computing framework ensures the real-time performance of complex models on mobile devices. Even when processing at high resolutions (such as 4K), a smooth user experience can be maintained. The obtained multi-modal search information will be further processed to generate multi-modal search information (which can also be understood as key search terms), providing strong support for the next information retrieval.

[0145] In a specific embodiment, when the user selects a part of the video frame as the first part of the region through gestures during video playback, the system immediately responds and uses image processing techniques (such as sharpening, noise reduction) to improve the picture quality of this region.

[0146] Subsequently, the OCR text recognition and image recognition processes are initiated. For text recognition, it can but is not limited to using PP-OCRv3 or similar technologies, which can accurately recognize text information in complex backgrounds, including different fonts, sizes, and languages.

[0147] For image recognition, algorithms such as PointNet++ or YOLOv8 are used to recognize and understand non-text visual elements such as people, objects, and environments in the image. The advantage of this process is that it can intelligently process any region marked by the user. Even for non-linear freehand selection, the system can automatically recognize and optimize it, ensuring the flexibility and accuracy of information extraction.

[0148] The above multi-modal content recognition refers to comprehensively using various technical means such as image recognition, text recognition, and context understanding to comprehensively analyze the first part of the region in the video frame to extract relevant information of visual elements.

[0149] The above multi-modal search information is keywords, phrases, or image descriptions obtained through content recognition within the first part area of the video frame. This information can be used as input for a search engine to help users obtain more in-depth knowledge or background information.

[0150] Overall recognition process: The system first sends the marked first part area of the video frame to the multi-modal content recognition module. In addition to traditional text and image feature extraction, the system also applies spatio-temporal context modeling, analyzing the feature changes of adjacent video frames through an LSTM network to enhance the recognition robustness in dynamic scenarios.

[0151] In addition, the hardware-software co-designed computing framework ensures the real-time performance of complex models on mobile devices. Even when processing at high resolutions (such as 4K), it can maintain a smooth user experience. The obtained multi-modal search information will be further processed to generate multi-modal search information (which can also be understood as key search terms), providing strong support for the next step of information retrieval.

[0152] In this embodiment, by allowing users to freely mark the first part area in the video frame, the system can intelligently perform multi-modal content recognition on it, extract rich information materials, and then convert them into multi-modal search information. This not only makes the learning of video content more targeted and interesting but also provides users with an immediate and in-depth knowledge acquisition channel. Especially when watching educational, historical, or professional tutorial videos, it can significantly improve the learning efficiency and content understanding level.

[0153] The key to implementing the above technology lies in its high customization ability and intelligent processing flow, effectively combining user intentions and video content.

[0154] As an optional example, during the playback of the target video, in response to a trigger operation on the video frame, obtaining key video frames includes:

[0155] During the playback of the target video, in response to a pause operation performed on the target video frame corresponding to each video frame in the second part of the video frames at different times, perform marking processing on the second part area in the target video frame;

[0156] Determine the second part of the video frames as key video frames.

[0157] In this embodiment, key video frames and related video content can be obtained by, but not limited to, taking screenshots, pausing and selecting areas, or clicking. The specific implementation steps are as follows:

[0158] S41, Pause event response;

[0159] Inject a timeupdate event listener into the video player. When it is detected that the playback pauses (either actively paused by the user or automatically paused), freeze the current frame and activate the selection area module.

[0160] S42, extraction of asynchronous frames;

[0161] It can be but is not limited to achieving low-overhead frame capture through the requestVideoFrameCallback API, supporting real-time processing capabilities of 30 frames per second.

[0162] S43, trigger priority management.

[0163] Specifically include: -graph TDA[user operation]-->B{operation type}B-->|touch gesture|C[start dynamic selection area]B-->|screenshot instruction|D[extract current frame + metadata]B-->|pause operation|E[enter interactive annotation mode.

[0164] In addition to the above gesture interaction methods and the method of determining key video frames manually marked by the user, dynamic selection areas can also be made through non-linear selection in the picture (such as free strokes) and intelligent frame selection (such as automatically attracting key visual elements). The specific implementation steps are as follows:

[0165] S51, non-linear selection (free strokes);

[0166] Path smoothing processing: It can be but is not limited to using the Catmull-Rom spline interpolation algorithm to denoise the original touch points and fit the curve;

[0167] Real-time rendering optimization: Use WebGL shaders for path rasterization to achieve a drawing performance of 60 FPS.

[0168] Implement the above non-linear selection through javascrip:

[0169] / / Example of spline interpolation function smoothPath(points, tension = 0.5) { const spline = new CatmullRomCurve3(points); return spline.getPoints(50); / / Generate smooth path points}

[0170] S52, intelligent frame selection (automatic adsorption);

[0171] Key element detection: Deploy a lightweight object detection model (MobileNet-SSD) to identify the main area in the image;

[0172] Bounding box optimization: Based on the detected object center point and confidence, use the gradient descent method to adjust the box position to achieve an automatic adsorption effect.

[0173] S53, Focus Tracking Engine.

[0174] Visual attention model: Integrate the Transformer-based architecture to build a saliency prediction module;

[0175] Multimodal fusion: Weightedly fuse the touch heat map and the visual saliency map to dynamically adjust the selection area weight.

[0176] Implement the above Focus Tracking Engine through the Python language:

[0177] # Pseudo code for saliency calculation def compute_attention(frame): saliency_map = vision_model.predict(preprocess(frame)) return cv2.resize(saliency_map, (frame.shape[1], frame.shape[0]))

[0178] In a specific example, during video playback, the system continuously monitors for pause operations. Once it detects that the user has performed a pause action, it immediately enters the marking processing mode. The user can select any area in the video frame, and the system automatically optimizes the image quality of the marked area, improving the resolution through intelligent algorithms (such as ESRGAN) to ensure that the image clarity is high enough for subsequent text and image content recognition. The advantage of this processing method is that it provides more flexible user control, allowing them to mark and focus on specific frames at different times in the video, rather than being limited to a preset gesture trigger point, thus making content recognition and information acquisition more in line with the actual needs of users.

[0179] After the user finishes marking, the system will automatically convert these marked video frames into key video frames, that is, identify and save those video segments that carry the user's interests and potential knowledge points. These key video frames will then be sent to the multimodal content recognition module for extraction and analysis of text and image features, and then generate targeted key search terms (which can also be understood as multimodal search information). The determination of key video frames not only helps to accurately capture the user's focus in the video, but also through subsequent automated processing, can provide the user with immediate and detailed background information, avoiding the cumbersome process of manual recording and searching, and greatly improving the efficiency and experience of video viewing.

[0180] In this embodiment, by introducing the pause and mark function during video playback, dynamic and refined analysis of video content is achieved. Users can pause the video at any time, mark and focus on specific areas in the video screen, and determine the video frame where the marked area is located as a key video frame, automatically perform content recognition and information generation, and provide users with instant, relevant and in-depth search information.

[0181] Through the above methods, not only the interactivity between users and video content is enhanced, but also users can explore video content in a more personalized way, whether it is the name and year query in historical videos, or the concept analysis in professional tutorials, which greatly enriches the experience of video viewing and learning.

[0182] As an optional example, by inputting the multimodal search information and historical behavior data into the target processing model, a set of key search information matching the video viewing progress is obtained, including:

[0183] By inputting the multimodal search information and the historical behavior data into the target processing model, a score of each search information in the multimodal search information is obtained, wherein the higher the score of the search information is, the higher the demand for displaying the target annotation content associated with the search information during the target video playback;

[0184] Based on the scores, a set of key search information is filtered out from the multimodal search information.

[0185] In this embodiment, it is described how to utilize the user's historical behavior data and multimodal search information to intelligently filter and recommend the information annotations most relevant to the current video screen.

[0186] Specifically, the model encodes multimodal search information based on deep learning (such as the Transformer architecture), while taking into account user preferences (such as frequently searched or annotated content types) to calculate the score of each search information.

[0187] Among them, search information with high scores means that they are more in line with the user's current knowledge needs and interests, and are therefore recommended first. For example, assuming that for the same video, 10 users have marked picture 1 in the second frame while watching the video in the historical period. If the current user circles picture 1 in the second frame through gesture interaction or pause operation, the score of picture 1 (a search information) output by the target processing model is higher.

[0188] Historical behavior data may include, but is not limited to, multiple users performing the same marking processing or triggering operation on at least one frame of the same video during the historical period, or multimodal search information (text, images, etc.) that other users have used while watching the same video during the historical period.

[0189] Suppose that the user has marked a total of 10 video frame images during the viewing of the target video, and there are a total of 20 multimodal search information identified in the 10 video frame images. After scoring by the target processing model, the top N search information with higher scores is retrieved for associated information, where N is a positive integer greater than or equal to 1.

[0190] The purpose of this processing is to avoid retrieving and displaying annotation content based on such search information when the user does not understand the video content and misinterprets the search information as a technical term or a historical figure, thereby improving the processing efficiency of associated information.

[0191] It should be noted that according to the scores output by the target processing model, the system further screens out the multimodal search information with the highest scores to form a set of key search information (for example, the top N keywords). This set of key search information will be used to retrieve relevant professional knowledge information and be displayed to the user in the form of an annotation content floating layer along with the video playback.

[0192] Among them, the screening process ensures the refinement and efficiency of the recommended information, avoids search information overload, and enables the user to focus on the most relevant and interesting knowledge points. At the same time, the dynamically updated list of key search information can be adjusted as the user's behavior changes, maintaining the real-time nature and relevance of the recommended content. In addition, through the intelligent analysis of the above-mentioned target processing model, the personalized sorting and recommendation of search information are realized, greatly improving the pertinence of information and user satisfaction.

[0193] In an optional embodiment, hierarchical search suggestions can also be generated based on the marked content during video playback and the target processing model. In this process, a multimodal knowledge graph and a real-time inference engine need to be constructed. The specific content is as follows:

[0194] S61, in the case where the marked content is the visual elements in the video frames obtained by using the steps in the above embodiment, perform text processing on the marked content and store it in the database;

[0195] S62, construct a context association model (which can also be understood as the target processing model) to associate these marked contents to form a structure with an associated relationship in content;

[0196] S63, multimodal context modeling - spatio-temporal knowledge graph construction;

[0197] A[Video frame] --> B{Entity recognition}

[0198] B --> C[Person / Place / Event]

[0199] B --> D [Timestamp]

[0200] B --> E [Spatial Coordinates]

[0201] C & D & E --> F [Graph Neural Network Encoding]

[0202] F --> G [Neo4j Graph Database]

[0203] S64, Back-end implementation of key technologies;

[0204] # Build a dynamic graph using PyTorch Geometric

[0205] class VideoGraph(torch.nn.Module):

[0206] def __init__(self):

[0207] super().__init__()

[0208] self.node_encoder = GCNConv(256, 128)

[0209] self.time_embed = nn.LSTM(100, 64, batch_first = True)

[0210] def forward(self, nodes, edge_index, timestamps):

[0211] # Time series encoding

[0212] time_emb = self.time_embed(timestamps)[0]

[0213] # Graph convolution processing

[0214] x = self.node_encoder(nodes, edge_index)

[0215] return torch.cat([x, time_emb], dim = -1)

[0216] S65, Implementation of cross-modal attention mechanism.

[0217] (1) Video-text alignment: Use BERT-wwm to extract video frame text features and perform cross-modal attention with speech recognition text;

[0218] (2) Visual Relation Reasoning: Use the DETR detector to extract object relationships and build an object interaction graph.

[0219] Through the above methods, the quality of information recommendation and user personalized experience in the video viewing process are significantly improved. By deeply analyzing and scoring the user's historical behavior data and multimodal search information extracted from the video screen through an intelligent model, the system can accurately capture the user's needs and intelligently recommend the most relevant information annotations.

[0220] As an optional example, the above-mentioned dynamically displaying annotation content associated with a set of key search information in the video screen based on the video viewing progress includes:

[0221] In a case where the target video is watched for the first time through the first account and a set of key search information is obtained at a first moment, in response to at least one key search information in the set of key search information appearing at a second moment, at least one annotation content associated with the at least one key search information is displayed at the second moment based on the video viewing progress;

[0222] The second time corresponding to the second moment is later than the first time corresponding to the first moment.

[0223] When user A watches the target video for the first time through the first account, the system will record every time the user pauses or marks the video screen, as well as the resulting multimodal search information. This search information will be scored and analyzed by the target processing model to extract a set of key search information. This process enables the system to identify and remember the knowledge points that the user is interested in based on the user's specific behavior in the video, providing basic data for subsequent dynamic annotations.

[0224] When the viewing progress reaches the second moment of the target video, if the picture at this moment contains at least one key search information in a set of key search information, the system will immediately identify it. That is, it can intelligently identify the repeated appearance of knowledge points that user A may be interested in, and the relevant annotations can be prepared in advance without the user having to mark or search again.

[0225] After identifying the second moment associated with the key search information, the system will automatically display the annotations related to these key search information on the video interface based on the current video viewing progress. These annotations are presented in the form of floating knowledge cards, which will not interrupt the video playback. Users can browse or hide the annotations with simple operations, ensuring that the user's knowledge exploration in the video is continuous and natural, avoiding the viewing experience that is disrupted by frequent pauses or switching applications.

[0226] It should be noted that at least one annotation content refers to multimodal information content generated based on key search information, including background information of the knowledge point, extended reading materials, previous user comments, etc., which aims to enhance the interactivity and knowledge depth of video viewing.

[0227] For example, if Figure 4 As shown, assuming that user A circles Formula 1 in the first video frame at the 2nd minute and 31st second of Video 1 for the first time through gesture interaction or other triggering methods, if Formula 1 also appears in the second video frame at the 8th minute and 11th second of the viewing progress, the annotation content associated with Formula 1 will be automatically displayed in the second video frame.

[0228] In an optional embodiment, the process of displaying annotation content during video playback may be implemented, but not limited to, through a streaming learning interface system, and specifically includes:

[0229] S71, design a non-interruption information display solution

[0230] (1) Analyze the information needs of users during the learning process and ensure that information display does not interfere with the main learning content, such as video playback;

[0231] (2) Use knowledge cards with transparent or semi-transparent backgrounds and display them floating on the side to ensure that the video screen is always visible.

[0232] S72, construct a hierarchical information display architecture;

[0233] (1) Basic interpretation layer: provides basic explanations of concepts or terms to ensure that the annotation content is simple and easy to understand;

[0234] (2) Professional interpretation level: In-depth analysis and interpretation from a professional perspective, suitable for learners with a certain foundation.

[0235] (3) Reference layer: lists relevant books, papers, etc. for learners to conduct in-depth research.

[0236] S73, realize dynamic tag persistence;

[0237] (1) Generate visual degree markers on the progress bar, and users can click on the markers to quickly locate relevant knowledge points;

[0238] (2) Marking points should support user customization, such as color and shape, to enhance the personalized experience;

[0239] (3) Ensure persistent storage of marking points so that the marking points still exist even if the system is shut down and reopened.

[0240] S74, creating a "check and save" workflow;

[0241] (1) When the user views the knowledge card, the system automatically records the viewing time and content, generating knowledge nodes with spatio-temporal coordinates;

[0242] (2) The knowledge nodes should support user editing, such as adding remarks, associating with other knowledge points, etc.;

[0243] (3) The system automatically saves the knowledge nodes into the user's knowledge base without manual saving.

[0244] S75, intelligently dock with the original knowledge map system;

[0245] (1) Analyze the data structure and interfaces of the original knowledge map system to ensure that the streaming learning interface system can be compatible with it;

[0246] (2) Implement the automatic import and export of knowledge nodes to maintain the consistency and integrity of the knowledge base;

[0247] (3) Provide a visual interface for the user to intuitively see the docking situation between the streaming learning interface system and the original knowledge map system.

[0248] S76, allow adding personalized memory tags.

[0249] (1) The user can add personalized tags to the knowledge nodes, such as importance level, learning status, etc.;

[0250] (2) The tags should support user customization, such as colors, icons, etc., to improve the personalized experience;

[0251] (3) The system should provide a tag filtering function for the user to quickly find the knowledge nodes with specific tags.

[0252] By adopting the above method, the video viewing experience is greatly optimized. Especially for the same knowledge points that the user encounters multiple times in the video, by recording and analyzing the user's first viewing behavior, the system can intelligently present relevant annotation content automatically when the user encounters these knowledge points again, which not only saves the user's time and energy, but also enhances the coherence and depth of learning.

[0253] As another optional example, the above method further includes:

[0254] In the case of first viewing the target video through the first account and obtaining a set of key search information at the first moment, in response to playing back the target video through the first account, locate the target playback progress at which at least one key search information appears in the video picture;

[0255] Based on the target playback progress, determine the third moment, where the third moment corresponds to the same playback time displayed at the same playback progress as the second moment;

[0256] Based on the video viewing progress, at least one annotation content associated with at least one key search frequency information is displayed at a third moment.

[0257] When the user plays back the target video through the first account, the system will automatically call the user's historical data to find the playback progress associated with the key search information generated during the previous viewing process. By analyzing the degree of match between the video frame and the key search information, the system can determine the moment when these key search information appear again in the video screen. In other words, it can automatically locate the video clips that the user is interested in, without the user having to search manually, saving time and improving the efficiency of video playback and personalized experience.

[0258] After determining the target playback progress, the system can calculate the exact time when the video playback reaches that progress, that is, the third moment. The time corresponding to the third moment is matched with the current progress of the video playback to ensure that when the video screen presents the key search information again, this moment can be accurately captured. In other words, through precise time synchronization, the previous key knowledge points can be automatically reproduced during video playback, providing users with a more coherent and natural learning experience.

[0259] When the video playback reaches the third moment, that is, the moment when the key search information reappears, the system will automatically display the previously generated annotation content on the video interface. As described in the above embodiment, these annotation contents are displayed in a non-interrupted manner, for example, through knowledge cards suspended on the side, which will not interfere with the user's viewing of the video. The user can browse or close this information as needed. Therefore, the instant display of the annotation content can also reduce the user's search operations when replaying the video, thereby improving learning efficiency.

[0260] For example, if Figure 5 As shown, assuming that user A circles formula 1 in the first video frame at the 2nd minute and 31 seconds when watching the video 1 for the first time through gesture interaction or other triggering methods, then the screenshot of formula 1 is determined as a key search information, and the annotation content of formula 1 is retrieved.

[0261] When user A plays back video 1 through account 1, the annotation content related to formula 1 will be directly displayed in the first video frame at 2 minutes and 31 seconds and the second video frame at 8 minutes and 11 seconds. This is because the user ID saves the user's related search and marking behavior data. Therefore, when playing back video 1 through the first account, the search content (i.e., the annotation content) of the associated video frame will be directly displayed.

[0262] By integrating the user's historical behavior data and video playback progress, the accurate positioning of key knowledge points and the automatic reproduction of annotation content are achieved. This not only provides users with a more personalized and efficient learning experience, but also enables users to obtain instant, relevant, and continuous information supplementation when playing back videos, greatly enriching the exploratory nature and learning depth of video content.

[0263] As an optional example, the above method further includes:

[0264] In the case of watching a target video through a first account and obtaining a set of key search information, sharing the target video to a second account through the first account;

[0265] During the process of watching the target video through the second account, dynamically display annotation content associated with the set of key search information in the video frame.

[0266] In this embodiment, a sharing mechanism is proposed, allowing users to share videos containing personalized annotations and key search information with other users, while ensuring that the recipients can dynamically receive these annotation messages during the video viewing process, deepening the dissemination method of video content and enhancing the information sharing and communication experience among users.

[0267] As Figure 6 shown, after the video is watched, assuming that user A shares the target video (Video 1) to the second account through the first account user, then when user B watches at 2 minutes and 30 seconds and 8 minutes and 11 seconds, the annotation content of Formula 1 will be displayed in the video frame. That is, during the sharing process, not only the video itself can be shared, but also all the annotations and key search information generated by user A's marked content can be shared with other users.

[0268] The way of displaying the annotation content in the video frame watched by user B is the same as the way of dynamically displaying the annotation in the form of a transparent or floating card immediately on the video interface described in the above embodiment, and it will not affect the normal process of video playback either.

[0269] During the above process, the system will generate a multi-modal information packet containing the above information, which can be shared, but not limited to, by means of an encrypted link, ensuring the security and integrity of the information. After the recipient (the user of the second account) opens the sharing link, the video and the relevant annotations and search results of the first account user will be automatically loaded, effectively promoting the dissemination of associated information and the sharing of knowledge.

[0270] In an optional embodiment, the specific implementation steps of the above scene-based sharing mechanism with dynamic permission control are as follows:

[0271] S81, each time the search behavior and content marking behavior associated with video playback by the user are triggered, the video frames are automatically saved in association;

[0272] S82, support the sharing of relevant behavior records associated by the user;

[0273] S83, the playback record saves the relevant search and marking behavior data according to the user ID;

[0274] Among them, a mapping relationship between the playback record and the user ID is created in advance.

[0275] S84, support the user to playback and display the search content with the video frame associated in a floating layer;

[0276] S85, implementation of dynamic permissions.

[0277] Among them, it is mainly determined according to the sharing options of the user at the time of sharing to give the relevant content and viewing permissions to be protected by sharing. The viewing permissions can be non-encrypted or the permissions obtained by entering an encryption password.

[0278] To understand the above technical solution more clearly, the following combines Figure 3 the overall flowchart shown below to further describe the above video processing method.

[0279] S302, play the target video;

[0280] S304, obtain multi-modal video content;

[0281] It is possible but not limited to obtain key video frames through touch gestures or conventional marking processing methods, that is, determine the video frames with marks as key video frames.

[0282] S306, identify and process the visual elements in the key video frames;

[0283] It is possible but not limited to determine whether the visual elements marked by the user in the key video frames are people, items, professional terms or formulas, etc. through OCR recognition, image recognition technology, etc.

[0284] S308, convert the recognition result into text;

[0285] It should be noted that converting the recognition result into text is only an example and is not limited thereto. For example, when the recognition result is a picture of an item, the recognition result in picture format can also be retained.

[0286] S310, generation of search terms;

[0287] It is possible but not limited to generate multi-modal search information in the manner of steps S31 - S33 in the above embodiments.

[0288] S312. Perform a search operation based on multimodal search information (text, images, etc.) and a third-party platform, and return the search results;

[0289] Among them, the search results include the annotation content associated with the multimodal search information.

[0290] S314. The user edits and saves the search results;

[0291] For example, save in the form of annotations, save in the form of a floating window, etc.

[0292] S316. Continue playing the video and automatically generate search suggestions;

[0293] For example, in the embodiment as Figure 5 shown, when the user replays Video 1, the pre-saved search suggestions (such as search keywords) are automatically obtained at the corresponding video frames.

[0294] S318. The annotation content obtained by the search is saved along with the video frames;

[0295] Among them, by establishing the mapping relationship between the user ID and the search behavior and annotation content, a video file carrying search suggestions and annotation content under this user ID is generated.

[0296] S320. Scenario-based sharing and playback.

[0297] For example, if User A shares the target video with User B, then during the process of User B watching the target video, the annotation content associated with the multimodal search information marked by User A will be automatically displayed in the corresponding video frames.

[0298] Adopting the above method, by using a multimodal triggering method including a gesture interaction method, the associated video frames expected to display the annotation content associated with the video content are obtained. By identifying the multimodal content in the associated video frames, multimodal search information is obtained, and using the target processing model, a set of key search information most relevant to the user's needs is determined. Finally, based on the video viewing progress, the annotation content associated with each key search information is dynamically displayed in the video picture. In other words, by adding touch marks according to the user's needs during the playback of the target video and identifying the key search information from the marked key video frames, the annotation content that meets the user's personalized needs is displayed in the video picture, improving the learning efficiency during video viewing, enriching the interactive experience of video viewing, and achieving the technical effect of improving the flexibility of displaying the annotation content of video content.

[0299] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0300] According to another aspect of the embodiments of the present application, there is also provided a Figure 7 video processing device as shown, and the device includes:

[0301] A first acquisition unit 702, configured to acquire a key video frame in response to a trigger operation on a video frame during the playback of a target video, where the trigger operation includes a trigger operation performed by a gesture interaction method and a marking process performed on a visual element in the video frame;

[0302] A first processing unit 704, configured to obtain multimodal search information by performing multimodal content recognition on the key video frame;

[0303] A second processing unit 706, configured to input the multimodal search information and historical behavior data into a target processing model to obtain a set of key search information that matches the video viewing progress, where the historical behavior data includes operation data performed on video frames based on multimodal search information within a historical period;

[0304] A first display unit 708, configured to dynamically display annotation content associated with a set of key search information in the video picture based on the video viewing progress.

[0305] Optionally, the foregoing first acquisition unit 702 includes:

[0306] A first acquisition module, configured to acquire trajectory data of a touch gesture performed by a gesture interaction method;

[0307] A first processing module, configured to determine a gesture type that matches the gesture interaction method based on the trajectory data;

[0308] A second processing module, configured to perform a corresponding trigger operation on a first part of the video frames in the video frame based on the gesture type;

[0309] A third processing module, configured to determine the first part of the video frames as the key video frames.

[0310] Optionally, the foregoing device further includes:

[0311] A marking unit, configured to perform marking processing on a first partial area of a video picture corresponding to a first partial video frame in response to a corresponding triggering operation performed on the first partial video frame;

[0312] An identification unit, configured to obtain multimodal search information by performing multimodal content identification on visual elements within the first partial area.

[0313] Optionally, the above-mentioned first obtaining unit 702 includes:

[0314] A pause module, configured to perform marking processing on a second partial area in a target video picture in response to a pause operation performed on each video frame in a second partial video frame at different moments during the playback of the target video;

[0315] A fourth processing module, configured to determine the second partial video frame as a key video frame.

[0316] Optionally, the above-mentioned second processing unit 706 includes:

[0317] A fourth processing module, configured to input the multimodal search information and historical behavior data into a target processing model to obtain a score for each search information in the multimodal search information, where the higher the score of the search information, the higher the demand for displaying target annotation content associated with the search information during the playback of the target video;

[0318] A screening module, configured to screen out a set of key search information from the multimodal search information based on the scores.

[0319] Optionally, the above-mentioned first display unit 708 includes:

[0320] A first display module, configured to, in the case of first viewing the target video through a first account and obtaining a set of key search information at a first moment, in response to the occurrence of at least one key search information in the set of key search information at a second moment, display at least one annotation content associated with the at least one key search information based on the video viewing progress; where the second time corresponding to the second moment is later than the first time corresponding to the first moment.

[0321] Optionally, the above-mentioned device further includes:

[0322] A third processing unit, configured to, in the case of first viewing the target video through a first account and obtaining a set of key search information at a first moment, in response to playing back the target video through the first account, locate the target playback progress at which at least one key search information appears in the video picture;

[0323] A fourth processing unit, configured to determine a third moment based on a target playback progress, where the third moment corresponds to the same playback time as the second moment at the same playback progress;

[0324] A second display unit, configured to display at least one annotation content associated with at least one key search frequency information at the third moment based on a video viewing progress.

[0325] Optionally, the above device further includes:

[0326] A sharing unit, configured to share the target video with a second account through the first account when viewing the target video through the first account and obtaining a set of key search information;

[0327] A third display unit, configured to dynamically display annotation content associated with a set of key search information in a video picture during the process of viewing the target video through the second account.

[0328] It should be noted that the embodiments of the video processing device here can refer to the embodiments of the above video processing method and will not be elaborated here.

[0329] According to another aspect of the embodiments of the present application, an electronic device for implementing the above video processing method is further provided. The electronic device may be Figure 1 the target terminal or server shown in the figure. This embodiment takes the electronic device as the target terminal as an example for illustration. As Figure 8 shown, the electronic device includes a memory 802 and a processor 804. A computer program is stored in the memory 802, and the processor 804 is configured to execute the steps in any one of the above method embodiments through the computer program.

[0330] Optionally, in this embodiment, the above electronic device may be at least one network device among multiple network devices in a computer network.

[0331] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:

[0332] S1. During the playback of the target video, in response to a trigger operation on a video frame, obtain a key video frame, where the trigger operation includes a trigger operation performed through a gesture interaction method and a marking process performed on a visual element in the video frame;

[0333] S2. Obtain multi-modal search information by performing multi-modal content recognition on the key video frame;

[0334] S3, obtaining a set of key search information matching the video viewing progress by inputting the multimodal search information and the historical behavior data into the target processing model, wherein the historical behavior data includes operation data performed on the video frame based on the multimodal search information within the historical period;

[0335] S4, based on the video viewing progress, dynamically displaying annotation content associated with a set of key search information in the video screen.

[0336] Alternatively, a person skilled in the art may understand that: Figure 8 The structure shown is for illustration only. Figure 8 The electronic device and the electronic equipment described above are not limited in structure. Figure 8 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 8 Different configurations are shown.

[0337] Among them, the memory 802 can be used to store software programs and modules, such as program instructions / modules corresponding to the video processing method and device in the embodiments of the present application. The processor 804 executes various functional applications and data processing by running the software programs and modules stored in the memory 802, that is, realizing the above-mentioned video processing method. The memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 802 may further include a memory remotely located relative to the processor 804, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 802 can be specifically used, but is not limited to, for storing key video frames, multimodal search information, and historical behavior data, etc. As an example, such as Figure 8 As shown, the memory 802 may include, but is not limited to, the first acquisition unit 702, the first processing unit 704, the second processing unit 706, and the first display unit 708 in the video processing device. In addition, it may also include, but is not limited to, other module units in the video processing device, which will not be repeated in this example.

[0338] Optionally, the above-mentioned transmission device 806 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 806 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one example, the transmission device 806 is a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0339] In addition, the above-mentioned electronic device further includes: a display 808, which is used to display the video picture of the target video; and a connection bus 810, which is used to connect each module component in the above-mentioned electronic device.

[0340] In other embodiments, the above-mentioned target terminal or server may be a node in a distributed system. Among them, the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes in the form of network communication. Among them, the nodes can form a point-to-point network, and any form of computing device, such as an electronic device such as a server or a target terminal, can become a node in the blockchain system by joining the point-to-point network.

[0341] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the video processing method provided in various optional implementation manners in the above-mentioned server verification processing and other aspects. Among them, the computer program is set to execute the steps in any one of the above method embodiments when running.

[0342] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be set to store a computer program for executing the following steps:

[0343] S1, during the playback of the target video, in response to a trigger operation on the video frame, obtain a key video frame, where the trigger operation includes a trigger operation performed by a gesture interaction method and a marking process performed on a visual element in the video frame;

[0344] S2, through multi-modal content recognition of the key video frame, obtain multi-modal search information;

[0345] S3, obtaining a set of key search information matching the video viewing progress by inputting the multimodal search information and the historical behavior data into the target processing model, wherein the historical behavior data includes operation data performed on the video frame based on the multimodal search information within the historical period;

[0346] S4, based on the video viewing progress, dynamically displaying annotation content associated with a set of key search information in the video screen.

[0347] Optionally, in the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0348] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the target terminal through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0349] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0350] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods of each embodiment of the present application.

[0351] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0352] In several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0353] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0354] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0355] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A video processing method, characterized in that, include: During the target video playback process, in response to a trigger operation on a video frame, a key video frame is acquired, wherein the trigger operation includes a trigger operation performed by gesture interaction and a marking process performed on a visual element in the video frame; Obtaining multimodal search information by performing multimodal content recognition on the key video frame; A set of key search information matching the video viewing progress is obtained by inputting the multimodal search information and historical behavior data into a target processing model, wherein the historical behavior data includes operation data performed on the video frame based on the multimodal search information within a historical period; Based on the video viewing progress, annotation content associated with the set of key search information is dynamically displayed in the video screen.

2. The method according to claim 1, wherein The step of obtaining a key video frame in response to a trigger operation on a video frame during the target video playback process includes: Acquire trajectory data of a touch gesture executed based on the gesture interaction mode; Based on the trajectory data, determining a gesture type that matches the gesture interaction mode; Based on the gesture type, performing a corresponding trigger operation on a first portion of video frames in the video frames; The first portion of video frames is determined as the key video frames.

3. The method according to claim 2, wherein The method further comprises: In response to executing a corresponding trigger operation on the first portion of video frames, performing a marking process on a first portion of the video screen corresponding to the first portion of video frames; The multimodal search information is obtained by performing multimodal content recognition on the visual elements in the first partial area.

4. The method according to claim 1, characterized in that, The step of obtaining a key video frame in response to a trigger operation on a video frame during the target video playback process includes: During the target video playback process, in response to a pause operation performed on the target video screen corresponding to each video frame in the second part of the video frames at different times, a marking process is performed on the second part of the area in the target video screen; The second portion of video frames is determined as the key video frames.

5. The method according to claim 1, characterized in that, The multimodal search information and historical behavior data are input into the target processing model to obtain a set of key search information matching the video viewing progress, including: By inputting the multimodal search information and the historical behavior data into a target processing model, a score of each search information in the multimodal search information is obtained, wherein the higher the score of the search information, the higher the demand for displaying the target annotation content associated with the search information during the target video playback; Based on the score, the group of key search information is filtered out from the multimodal search information.

6. The method according to claim 1, wherein The dynamically displaying annotation content associated with the set of key search information in the video screen based on the video viewing progress includes: In a case where the target video is watched for the first time through the first account and the set of key search information is obtained at a first moment, in response to at least one key search information in the set of key search information appearing at a second moment, at least one annotation content associated with the at least one key search information is displayed at the second moment based on the video viewing progress; Among them, the second time corresponding to the second moment is later than the first time corresponding to the first moment.

7. The method according to claim 6, wherein The method further includes: In the case of first watching the target video through the first account and obtaining the set of key search information at the first moment, in response to playing back the target video through the first account, locating the target playback progress at which at least one of the key search information appears in the video picture; Based on the target playback progress, determining a third moment, where the third moment and the second moment correspond to the same playback time displayed at the same playback progress; Based on the video viewing progress, displaying at the third moment the at least one annotation content associated with the at least one key search frequency information.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: In the case of watching the target video through the first account and obtaining the set of key search information, sharing the target video to the second account through the first account; During the process of watching the target video through the second account, dynamically displaying in the video picture the annotation content associated with the set of key search information.

9. A video processing device, characterized in that, Including: A first acquisition unit, configured to obtain a key video frame in response to a trigger operation on a video frame during the playback of a target video, where the trigger operation includes a trigger operation performed by a gesture interaction method and a marking process performed on a visual element in the video frame; A first processing unit, configured to obtain multi-modal search information by performing multi-modal content recognition on the key video frame; A second processing unit, configured to obtain a set of key search information that matches the video viewing progress by inputting the multi-modal search information and historical behavior data into a target processing model, where the historical behavior data includes operation data performed on the video frame based on the multi-modal search information within a historical period; A first display unit, configured to dynamically display in the video picture the annotation content associated with the set of key search information based on the video viewing progress.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where the program, when executed by a terminal device or a computer, performs the method described in any one of claims 1 to 8.

11. An electronic device, comprising a memory and a processor, characterized in that, A computer program is stored in the memory, and the processor is configured to execute the method described in any one of claims 1 to 8 through the computer program.