Method and device for detecting interaction event in recorded video
By detecting interactive events in the screen-recording video, and using the object detection model and similarity comparison to determine important frames, the problem of excessively long recording video time and dispersed knowledge points is solved, the effect of extracting key content of the video is achieved, and the training efficiency is improved.
Patent Information
- Application Number
- CN202510121483.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
AI Technical Summary
Due to the long time of recording videos and the dispersed knowledge points, learning efficiency is ineffective, and a method is needed to extract key content from the video.
By detecting interactive events in the recorded video, the object detection model is used to detect whether the image frame contains a mouse pointer, and important frames are determined through similarity comparison.
Quickly and at low cost to identify interactive events in videos, assist in extracting key segments from videos, improve training efficiency and strengthen training results.
Smart Images

Figure CN119942419A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of machine learning, and more particularly, to methods and devices for detecting interactive events in recorded videos. Background Art
[0002] Teaching videos play an increasingly important role in modern corporate personnel training. They not only provide standardized training content, but also allow learners to learn flexibly according to their own pace and schedule. One type of teaching video is a screen recording video in which the instructor uses his own electronic device (such as a computer) to demonstrate and explain the operation, and then records the operation interface.
[0003] Common problems with this type of screen recording video include long videos, scattered knowledge points, etc., which lead to low learning efficiency when learners watch directly. Therefore, a method is needed to analyze the video content to extract the key content. Summary of the invention
[0004] One or more embodiments of the present specification describe a method and apparatus for detecting interactive events in a recorded video, and extracting key content from the video by detecting interactive events that occur in the recorded video.
[0005] In a first aspect, a method for detecting an interactive event in a recorded video is provided, comprising:
[0006] Acquire a first image frame and a second image frame extracted from the screen recording video; the first image frame is located before the second image frame in time sequence;
[0007] Inputting the first image frame and the second image frame into the target detection model respectively to detect whether the mouse pointer is included therein;
[0008] When the detection results are all yes, obtaining a first sub-image and a second sub-image output by the target detection model containing the mouse pointer;
[0009] When the first similarity between the first sub-image and the second sub-image is less than a preset first threshold, the first image frame is determined as an important frame containing a mouse interaction event.
[0010] In some possible implementations, the time interval between the first image frame and the second image frame is less than a preset second threshold.
[0011] In some possible implementations, the target detection model is trained by the following steps:
[0012] Obtain a training set comprising a plurality of screenshots, each of which contains a mouse pointer and has bounding box coordinates indicating the location of the mouse;
[0013] According to the training set, a target detection model is trained.
[0014] In some possible implementations, the type of the mouse pointer includes at least one of the following: an arrow pointer and a hand-shaped pointer.
[0015] In some possible implementations, obtaining a first sub-image and a second sub-image output by the target detection model and containing a mouse pointer includes:
[0016] Obtain first bounding box coordinates and second bounding box coordinates output by the target detection model for the first image frame and the second image frame respectively;
[0017] According to the first bounding box coordinates, a first sub-image is obtained by intercepting from the first image frame;
[0018] A second sub-image is captured from the second image frame according to the second bounding box coordinates.
[0019] In some possible implementations, the first similarity is determined by the following steps:
[0020] Performing image encoding on the first sub-image and the second sub-image respectively to obtain a first embedded representation and a second embedded representation;
[0021] A first similarity is determined according to a similarity between the first embedded representation and the second embedded representation.
[0022] In some possible implementations, performing image encoding on the first sub-image and the second sub-image respectively to obtain a first embedded representation and a second embedded representation includes:
[0023] The first sub-image and the second sub-image are respectively input into the image coding model to obtain corresponding first embedded representation and second embedded representation.
[0024] In some possible implementations, the mouse interaction event includes a mouse click event.
[0025] In a second aspect, a device for detecting an interactive event in a recorded video is provided, comprising:
[0026] A first acquisition unit is configured to acquire a first image frame and a second image frame extracted from the screen recording video; the first image frame is located before the second image frame in time sequence;
[0027] The target detection unit is configured to input the first image frame and the second image frame into a target detection model respectively to detect whether a mouse pointer is included therein;
[0028] A second acquisition unit is configured to, when the detection results are all yes, acquire a first sub-image and a second sub-image output by the target detection model containing a mouse pointer;
[0029] The determination unit is configured to determine the first image frame as an important frame containing a mouse interaction event when a first similarity between the first sub-image and the second sub-image is less than a preset first threshold.
[0030] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.
[0031] In a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.
[0032] The method and device for detecting interactive events in recorded videos proposed in the embodiments of this specification, the method extracts two image frames in the recorded video, performs target detection on the two image frames respectively, and detects whether the image of the mouse pointer is contained therein. When both image frames contain a mouse pointer, the sub-image parts of the two image frames respectively containing the mouse pointer are obtained from the results of the target detection. Next, the two sub-images are compared for similarity. When the similarity of the two sub-images is lower than a preset threshold, it means that the background image part near the mouse pointer between the two frames has undergone a large change (for example, the screen jumps to a new page, etc.), then it is considered that an interactive event with the electronic device has occurred in these two image frames, and it is recorded as an important frame. By using the method proposed in the embodiments of this specification, interactive events in videos can be quickly and inexpensively identified to assist in extracting key clips in the video, improve the efficiency of personnel training and enhance the training effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0034] Figure 1 A schematic diagram showing an implementation scenario of a method for detecting interactive events in a recorded video according to an embodiment;
[0035] Figure 2 A flow chart showing a method for detecting interactive events in recorded video according to one embodiment;
[0036] Figure 3 A flow chart showing important frame detection for multiple image frame pairs according to one embodiment is shown;
[0037] Figure 4 A schematic block diagram of an apparatus for detecting interactive events in a recorded video according to one embodiment is shown. DETAILED DESCRIPTION
[0038] The solution provided in this specification is described below in conjunction with the accompanying drawings.
[0039] As mentioned above, screen recordings used as teaching videos have common problems such as long videos and scattered knowledge points. Specifically, the instructor may record too much content in the same screen recording, and the clips related to teaching and operation demonstrations are scattered and have no obvious chapter markings. In addition, when the instructor is demonstrating the operation, the page may freeze, the lecture may be reviewed, the speed of speech may be too slow, and the teaching video may contain more clips without knowledge points. If the trainees watch such screen recordings directly, it is easy to lead to low efficiency and difficulty in capturing key points. At the same time, it may also affect the accurate transmission of knowledge points, resulting in the omission or misunderstanding of important information, which will increase the company's training costs as a whole and affect the overall business level of the staff.
[0040] For example, during the training phase for customer service staff of online e-commerce companies, a large number of key knowledge points need to be summarized to train and teach customer service staff. The original teaching screen recording video data is long, and most of the clips do not contain knowledge points, so the efficiency of direct viewing and learning is low.
[0041] Based on the above analysis, an embodiment of this specification proposes a method for detecting interactive events in recorded videos, which analyzes and compares the image frames in the video to find image frames containing interactive events, called important frames. After collecting these important frames, it is possible to assist in finding important video clips in the original recorded video to improve the efficiency of viewing the video.
[0042] An interaction event may be an interaction between a user using an external device (such as a mouse) and the content displayed in the display area of the electronic device, such as clicking a button with a mouse, scrolling the interface with a mouse wheel, etc. For example, a customer service staff clicks a send button when answering a customer's question; or clicks an order link or a submit modification button when modifying an order for a customer.
[0043] Figure 1 A schematic diagram of an implementation scenario of a method for detecting interactive events in a recorded video according to an embodiment is shown. Figure 1In the example, the screen recording video can be a video obtained by recording the display area of any electronic device. In some specific scenarios, due to the functional limitations of the video recording software and other reasons, the screen recording video only contains video content and audio content, but does not contain information indicating interactive events, such as information indicating mouse interactive events (for example, when the mouse is clicked, an animation with a flashing effect will be generated around the mouse pointer). Therefore, it is necessary to analyze the video content to find out the interactive events.
[0044] First, a pair of image frames is extracted from the screen recording video. The image frame that comes earlier in time can be recorded as image frame 1, and the image frame that comes later in time can be recorded as image frame 2, that is, image frame 1 is located before image frame 2 in time sequence. The time interval between the pair of image frames can be a smaller time interval, for example, the time interval is less than the preset second threshold, so that the result of the important frame generated finally is more accurate.
[0045] Then, image frame 1 and image frame 2 are respectively input into the pre-trained object detection model to detect whether they contain images of mouse pointers. Mouse pointers can include various types, such as conventional arrow-shaped pointers, hand-shaped pointers when placed on clickable buttons, etc. When both image frames contain mouse pointers, the sub-image parts of image frame 1 and image frame 2 containing mouse pointers are extracted from the output information of the object detection model, which can be correspondingly recorded as sub-images. Figure 1 Kazuko Figure 2 .
[0046] Next, Figure 1 Kazuko Figure 2 A similarity comparison is performed. When the similarity between the two is less than a preset first threshold, it indicates that the background information near the mouse pointer has changed significantly between image frame 1 and image frame 2. At this time, it can be considered that an important event has occurred, such as a user clicking a button or a link to jump to a new page. Therefore, at least one of the image frames in the pair can be identified as an important frame and output. The important frame can be used to locate the key position in the original recorded video.
[0047] The above describes the steps of performing important frame judgment on a pair of image frames. It is understandable that multiple pairs of image frames can also be extracted from the recorded video, and the above steps are used to perform important frame judgment based on each pair of image frames to obtain one or more important frames. Then, based on these important frames, key video clips in the recorded video can be extracted more accurately.
[0048] The specific implementation steps of the above method for detecting interactive events in recorded video are described below in conjunction with specific embodiments.
[0049] Figure 2 A flowchart of a method for detecting interactive events in a recorded video according to an embodiment is shown. The execution subject of the method can be any platform or server or device cluster with computing and processing capabilities. Figure 2 As shown, the method at least includes: step 202, obtaining a first image frame and a second image frame extracted from a screen recording video; the first image frame is located before the second image frame in time sequence; step 204, inputting the first image frame and the second image frame into a target detection model respectively to detect whether they contain a mouse pointer; step 206, when the detection results are both yes, obtaining a first sub-image and a second sub-image containing a mouse pointer output by the target detection model; step 208, when a first similarity between the first sub-image and the second sub-image is less than a preset first threshold, determining the first image frame as an important frame containing a mouse interaction event.
[0050] The specific execution process of each of the above steps is described below.
[0051] First, in step 202, a first image frame and a second image frame extracted from the screen recording video are obtained; the first image frame is located before the second image frame in time sequence.
[0052] The screen recording video may be a video obtained by recording the display area of any electronic device. In one embodiment, the screen recording video may include video content and audio content, and does not include additional information indicating interactive events, such as information indicating mouse interactive events. Furthermore, in a more specific embodiment, the screen recording video consists of video content and audio content.
[0053] The first image frame and the second image frame may be image frames extracted from the screen recording video at different time points, wherein the first image frame is located before the second image frame in time sequence.
[0054] In one embodiment, in order to improve the accuracy of video interaction event detection, the time interval between the first image frame and the second image frame can be set to a smaller value. In this way, if the difference between the pair of image frames is still large, it can be considered with a higher confidence that an interaction event exists within the time period of the first image frame and the second image frame.
[0055] Specifically, the time interval between the first image frame and the second image frame is less than a preset second threshold.
[0056] The second threshold can be expressed in terms of time length, such as 0.1 seconds, i.e., the time interval between the first image frame and the second image frame is 0.1 seconds; it can also be expressed in terms of the number of video frames, such as 10 frames, i.e., there are 10 video frames between the first image frame and the second image frame.
[0057] After a pair of image frames are extracted in step 202, in step 204, the first image frame and the second image frame are respectively input into the target detection model to detect whether a mouse pointer is included therein.
[0058] The task of object detection is to find all the objects of interest in the image and determine their categories and locations. The object detection model used in step 204 can be any type of object detection model, such as R-CNN (Region-Convolutional Neural Network) series models, YOLO (You Only Look Once) series models, DETR (Detection Transformer) models, etc., which are not limited here.
[0059] An existing pre-trained target detection model may be used to directly perform target detection on the first image frame and the second image frame, or a target detection model may be obtained by training using a training set.
[0060] In one embodiment, the target detection model is trained through the following steps 2042 to 2044:
[0061] First, in step 2042, a training set including multiple screenshots is obtained, any screenshot including a mouse pointer and having bounding box coordinates indicating the position of the mouse.
[0062] A bounding box is a rectangular box used to locate and select objects in object detection. It completely surrounds the target object with the smallest rectangle. The position of the bounding box is usually represented by the coordinates of the bounding box. For example, it can be represented by the horizontal and vertical coordinates of the upper left corner of the bounding box, and the width and height of the bounding box; or it can be represented by the horizontal and vertical coordinates of the upper left corner of the bounding box, and the horizontal and vertical coordinates of the lower right corner of the bounding box, which are not limited here.
[0063] In the object detection task, the object detection model needs to predict the bounding box position of each object, predict the category of the object in the bounding box, and predict the confidence score.
[0064] Then, in step 2044, a target detection model is trained based on the training set.
[0065] It should be noted that, unlike the targets in conventional target detection tasks, such as faces, animals, etc., these conventional targets often occupy a larger area in the image and are relatively easy to locate and classify. As for the mouse pointer, the image area of the mouse pointer on the screen of the electronic device is much smaller than the conventional targets mentioned above. Therefore, in step 2044, when using the training set to train the target detection model, the learning rate of the gradient descent is set to a smaller value, such as 0.0001, so that the model can be trained for more rounds to improve the final prediction confidence.
[0066] The type of the mouse pointer in step 202 and step 204 includes at least one of the following: an arrow pointer and a hand pointer.
[0067] The arrow pointer is a standard pointer form, and the hand pointer indicates that the user can interact with certain elements such as hyperlinks or buttons. At the same time, the mouse pointer type excludes the text pointer (I-shaped pointer) because it indicates that text can be typed at the cursor position. At this time, it is often possible to use a device such as a keyboard to enter text, and use the mouse for interaction events, so the text pointer should be excluded from the detection type of the target detection model.
[0068] Next, in step 206, when the detection results are all yes, a first sub-image and a second sub-image containing a mouse pointer output by the target detection model are obtained.
[0069] The detection result of a certain image frame may be yes, that is, the target detection model detects the mouse pointer, and the output confidence is greater than a preset third threshold.
[0070] When the target detection results for the first image frame and the second image frame are both yes, that is, the two image frames contain the mouse pointer, then the first sub-image and the second sub-image containing the mouse pointer output by the target detection model are obtained. That is, the image surrounded by the bounding box containing the mouse pointer output by the target detection model in the corresponding image frame.
[0071] In one embodiment, step 206 specifically includes: obtaining first bounding box coordinates and second bounding box coordinates output by the target detection model for the first image frame and the second image frame respectively; based on the first bounding box coordinates, obtaining a first sub-image from the first image frame; based on the second bounding box coordinates, obtaining a second sub-image from the second image frame.
[0072] When at least one of the detection results of the target detection model for the first image frame and the second image frame is negative, that is, at least one of the first image frame and the second image frame does not contain a mouse pointer, steps 206 and 208 are no longer executed, and the first image frame is determined as a non-important frame.
[0073] Finally, in step 208, when the first similarity between the first sub-image and the second sub-image is less than a preset first threshold, the first image frame is determined as an important frame containing a mouse interaction event.
[0074] In one embodiment, the mouse interaction event may include a mouse click event.
[0075] In other embodiments, other mouse interaction events, such as mouse movement events, etc., may be defined according to specific needs, which are not limited here.
[0076] When the first similarity between the first sub-image and the second sub-image is less than a preset first threshold, it means that the background information near the mouse pointer has changed significantly between the first image frame and the second image frame. At this time, it can be considered that an important event has occurred, such as a user clicking a button or a link to jump to a new page. Therefore, at least one of the image frames in the pair, such as the first image frame in the previous time sequence, can be identified as an important frame and output. The important frame can be used to locate the key position in the original recorded video.
[0077] In one embodiment, the first similarity in step 208 is determined by steps 2082 and 2084 as follows:
[0078] In step 2082, image encoding is performed on the first sub-image and the second sub-image respectively to obtain a first embedded representation and a second embedded representation.
[0079] Any method can be used to encode the image and obtain the corresponding embedding. For example, a pre-trained image encoding model can be used, such as the image encoder in the CLIP (Contrastive Language-Image Pre-Training) model, a convolutional neural network, an encoder in the VAE (Variational AutoEncoder) model, etc., without limitation.
[0080] In a more specific embodiment, step 2082 specifically includes:
[0081] The first sub-image and the second sub-image are respectively input into the image coding model to obtain corresponding first embedded representation and second embedded representation.
[0082] After the embedding representations of the two subgraphs are obtained, the similarity between them can be calculated based on the embedding representations. In step 2084, a first similarity is determined based on the similarity between the first embedding representation and the second embedding representation.
[0083] The similarity between the first embedding representation and the second embedding representation can be measured using cosine similarity, Euclidean distance, etc., which is not limited here.
[0084] The above steps describe the process of determining important frames based on a pair of image frames, including the first image frame and the second image frame, and then detecting an interaction event. It can be understood that in other embodiments, multiple pairs of image frames can be extracted from different time positions of the screen recording video, and each pair of image frames can be respectively processed using the following method: Figure 2 The method of each step shown is used to determine important frames and obtain multiple important frames. After arranging these important frames in chronological order, an important frame sequence is obtained. The important video clips in the original screen recording video can be extracted based on the important frame sequence. For example, when the number or density of important frames in a certain time interval exceeds a certain threshold, the video clip corresponding to the time interval can be determined as an important video clip and cut out from the original screen recording video. Furthermore, the important video clip can also be manually checked to confirm whether it contains important information.
[0085] According to the above concept, Figure 3 A flow chart of performing important frame detection on multiple image frame pairs according to one embodiment is shown. Figure 3 In the example, first, a pair of image frame pairs are extracted from the screen recording video, and target detection is performed on them as described in step 204 to determine whether there is a mouse pointer in both of them. When there is no mouse pointer in at least one image frame, another pair of image frame pairs is extracted from the screen recording video for detection. When there is a mouse pointer in both image frames, sub-images containing the mouse pointer in the two image frames are intercepted as described in step 206, and the similarity between them is calculated. When the similarity is lower than a preset threshold, at least one frame in the image frame pair is determined as an important frame, as described in step 208. And another pair of image frame pairs is extracted from the screen recording video for detection. When the similarity is higher than or equal to the preset threshold, another pair of image frame pairs is directly extracted from the screen recording video for detection. Until the traversal of the screen recording video is completed.
[0086] The embodiments of this specification use machine learning models to automatically detect interactive events in screen recordings. By analyzing screen recordings, key knowledge points are automatically identified and extracted, solving the problem of long and inefficient screen recordings in existing training, and significantly improving training efficiency.
[0087] According to another embodiment, a device for detecting interactive events in recorded video is also provided. Figure 4 A schematic block diagram of an apparatus for detecting interactive events in a recorded video according to an embodiment is shown, and the apparatus can be deployed in any device, platform or device cluster with computing and processing capabilities. Figure 4 As shown, the device 400 includes:
[0088] The first acquisition unit 402 is configured to acquire a first image frame and a second image frame extracted from the screen recording video; the first image frame is located before the second image frame in time sequence;
[0089] The target detection unit 404 is configured to input the first image frame and the second image frame into a target detection model respectively to detect whether a mouse pointer is included therein;
[0090] The second acquisition unit 406 is configured to, when the detection results are all yes, acquire the first sub-image and the second sub-image output by the target detection model containing the mouse pointer;
[0091] The determination unit 408 is configured to determine the first image frame as an important frame containing a mouse interaction event when a first similarity between the first sub-image and the second sub-image is less than a preset first threshold.
[0092] According to another aspect of the embodiment, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in any of the above embodiments.
[0093] According to yet another embodiment, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in any one of the above embodiments is implemented.
[0094] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0095] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0096] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0097] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0098] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for detecting an interactive event in a recorded video, comprising: Obtain a first image frame and a second image frame extracted from the screen recording video; The first image frame is located before the second image frame in time sequence; Inputting the first image frame and the second image frame into the target detection model respectively to detect whether the mouse pointer is included therein; When the detection results are all yes, obtaining a first sub-image and a second sub-image output by the target detection model containing the mouse pointer; When the first similarity between the first sub-image and the second sub-image is less than a preset first threshold, the first image frame is determined as an important frame containing a mouse interaction event.
2. The method according to claim 1, wherein: The time interval between the first image frame and the second image frame is less than a preset second threshold.
3. The method according to claim 1, wherein: The target detection model is trained through the following steps: Obtain a training set comprising a plurality of screenshots, each of which contains a mouse pointer and has bounding box coordinates indicating the location of the mouse; According to the training set, a target detection model is trained.
4. The method according to claim 1 or 3, wherein: The type of the mouse pointer includes at least one of the following: an arrow pointer and a hand pointer.
5. The method according to claim 1, wherein: Get the first sub-image and the second sub-image containing the mouse pointer output by the target detection model, including: Obtain first bounding box coordinates and second bounding box coordinates output by the target detection model for the first image frame and the second image frame respectively; According to the first bounding box coordinates, a first sub-image is obtained by intercepting from the first image frame; A second sub-image is captured from the second image frame according to the second bounding box coordinates.
6. The method according to claim 1, wherein: The first similarity is determined by the following steps: Performing image encoding on the first sub-image and the second sub-image respectively to obtain a first embedded representation and a second embedded representation; A first similarity is determined according to a similarity between the first embedded representation and the second embedded representation.
7. The method according to claim 6, wherein: Performing image encoding on the first sub-image and the second sub-image respectively to obtain a first embedded representation and a second embedded representation, including: The first sub-image and the second sub-image are respectively input into the image coding model to obtain corresponding first embedded representation and second embedded representation.
8. The method according to claim 1, wherein: The mouse interaction event includes a mouse click event.
9. A device for detecting interactive events in a recorded video, comprising: A first acquisition unit is configured to acquire a first image frame and a second image frame extracted from the screen recording video; The first image frame is located before the second image frame in time sequence; The target detection unit is configured to input the first image frame and the second image frame into a target detection model respectively to detect whether a mouse pointer is included therein; A second acquisition unit is configured to, when the detection results are all yes, acquire a first sub-image and a second sub-image output by the target detection model containing a mouse pointer; The determination unit is configured to determine the first image frame as an important frame containing a mouse interaction event when a first similarity between the first sub-image and the second sub-image is less than a preset first threshold.
10. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.
11. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
Video processing method and device based on large model and electronic equipment
CN120808240A