A video content tagging processing method, apparatus, and computer-readable storage medium
By identifying and marking video objects during video playback, acquiring operations, and tracking the markings to the video file, the problem of inconvenient video content marking in existing technologies is solved, achieving the effect of fast marking and sharing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NUBIA TECHNOLOGY CO LTD
- Filing Date
- 2022-09-29
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, video content tagging is inconvenient, making it difficult for users to effectively tag key content when sharing videos. This results in incomplete information, difficulty in management, and the need for additional explanations for others to understand.
During video playback, video objects are identified and marked, selected operations and notes are obtained, the marking is tracked until the target object leaves the video content, and the timestamp information is merged into the video file, providing a marked playback option.
It enables users to quickly mark and share key content during video playback, improving the completeness and convenience of information. Users can directly jump to the marked content using timestamps, enhancing the effectiveness of video sharing.
Smart Images

Figure CN115470373B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mobile communications, and more particularly to a video content tagging processing method, device, and computer-readable storage medium. Background Technology
[0002] In existing technologies, with the continuous development of smart terminal devices, users' demands for sharing multimedia files are also increasing. Specifically, currently, when watching videos, users can perform operations such as fast-forwarding, slow-motion, repeating intervals, and pausing. When the video is reopened after playback, there are no changes. Based on this, existing solutions have the following three problems: First, key content seen in the video can only be recorded through screenshots, but screenshots may not capture the desired content, requiring repeated captures; second, the method of saving screenshots lacks the context of the time before and after the screenshot, potentially leading to incomplete information; third, saving both screenshots and videos simultaneously makes managing two separate documents inconvenient; additionally, in existing solutions, when sharing videos with others, a description is required, otherwise the recipient may not understand the intended message.
[0003] Therefore, how to better implement a video content tagging and processing scheme to improve the user experience of video sharing functions has become an urgent technical problem to be solved. Summary of the Invention
[0004] To address the aforementioned technical deficiencies in the prior art, this invention proposes a video content tagging processing method, which includes:
[0005] During video playback, video objects in the video content are identified and marked, and the selection and annotation operations of the marked objects are obtained.
[0006] The video object corresponding to the selected operation and / or the remark operation is taken as the target object;
[0007] Based on the operation time of the selected operation and / or the annotation operation in the video content, the target object is tracked and marked in preset image frames before and after the reference until the target object leaves the video content;
[0008] The timestamp information corresponding to the tracking marker is merged into the video file corresponding to the video content, so that when the video file is played again, a marker playback option corresponding to the timestamp is provided.
[0009] Optionally, the step of identifying and marking video objects in the video content during video playback, and obtaining the selection and annotation operations of the marked objects, includes:
[0010] Obtain the file attributes of the video file corresponding to the video content;
[0011] The content theme of the video content is determined based on the file attributes.
[0012] Optionally, the step of identifying and marking video objects in the video content during video playback, and obtaining the selection and annotation operations of the marked objects, further includes:
[0013] Determine the attention features in the video content based on the content theme;
[0014] During the video playback, an object search is performed on the video content based on the attention features, and the searched objects are used as the video objects.
[0015] Optionally, the step of identifying and marking video objects in the video content during video playback, and obtaining the selection and annotation operations of the marked objects, further includes:
[0016] Determine the corresponding bounding box based on the object attributes of the video object;
[0017] During the video playback, the video objects are marked in real time using the marker boxes.
[0018] Optionally, the step of identifying and marking video objects in the video content during video playback, and obtaining the selection and annotation operations of the marked objects, further includes:
[0019] Obtain the first touch command within the marked box, and execute the selected operation through the first touch command;
[0020] Obtain the second touch command of the marked frame line, and execute the annotation operation through the second touch command.
[0021] Optionally, the step of using the video object corresponding to the selected operation and / or the annotation operation as the target object includes:
[0022] Obtain the target attributes of the target object;
[0023] The object motion range and object outline body range of the target object are determined based on the target attributes.
[0024] Optionally, the step of tracking and marking the target object in preset image frames before and after the selected operation and / or the annotated operation in the video content, based on the operation time of the selected operation and / or the annotated operation, until the target object leaves the video content, includes:
[0025] The preceding and following preset image frames are determined based on the object's motion range, and time-related tracking markers are applied to the target object in the preceding and following preset image frames;
[0026] When the target object extends beyond the main body of the object outline in the image frame, it is determined that the target object has left the video content, and time-related tracking markers are stopped.
[0027] Optionally, the step of fusing the timestamp information corresponding to the tracking marker into the video file corresponding to the video content, so that when the video file is played again, a marker playback option corresponding to the timestamp is provided, includes:
[0028] When the video file is played again, the target object and / or the corresponding annotation content are displayed in a floating position within the video content.
[0029] When the target object and / or the remarks are clicked, the video file is positioned to the video content corresponding to the timestamp and then played.
[0030] The present invention also proposes a video content tagging processing device, the device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the video content tagging processing method as described in any of the preceding claims.
[0031] The present invention also proposes a computer-readable storage medium storing a video content tagging processing program, which, when executed by a processor, implements the steps of the video content tagging processing method as described in any of the preceding claims.
[0032] The video content marking processing method, device, and computer-readable storage medium of the present invention identify and mark video objects in the video content during video playback, and obtain the selected operation and annotation operation of the mark; take the video object corresponding to the selected operation and / or the annotation operation as the target object; use the operation time of the selected operation and / or the annotation operation in the video content as a reference, track and mark the target object in preset image frames before and after the reference until the target object leaves the video content; and fuse the timestamp information corresponding to the tracking mark into the video file corresponding to the video content so that when the video file is played again, a marked playback option corresponding to the timestamp is provided. Attached Figure Description
[0033] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0034] Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal according to the present invention;
[0035] Figure 2 This is a flowchart of the first step of the video content marking and processing method of the present invention;
[0036] Figure 3 This is the second flowchart of the video content marking and processing method of the present invention;
[0037] Figure 4 This is the third flowchart of the video content marking and processing method of the present invention;
[0038] Figure 5 This is the fourth flowchart of the video content marking and processing method of the present invention;
[0039] Figure 6 This is the fifth flowchart of the video content marking and processing method of the present invention;
[0040] Figure 7 This is the sixth flowchart of the video content marking and processing method of the present invention;
[0041] Figure 8 This is the seventh flowchart of the video content marking and processing method of the present invention;
[0042] Figure 9 This is the eighth flowchart of the video content marking and processing method of the present invention;
[0043] Figure 10 This is a schematic diagram of the first video content processing step in the video content marking and processing method of the present invention;
[0044] Figure 11 This is a schematic diagram of the second video content processing step in the video content marking and processing method of the present invention;
[0045] Figure 12 This is a schematic diagram of the third video content processing step in the video content marking processing method of the present invention;
[0046] Figure 13 This is a schematic diagram of the fourth video content processing step in the video content marking and processing method of the present invention. Detailed Implementation
[0047] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0048] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0049] Terminals can be implemented in various forms. For example, the terminals described in this invention may include mobile terminals such as mobile phones, tablets, laptops, handheld computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.
[0050] The following description will use a mobile terminal as an example. Those skilled in the art will understand that, apart from elements specifically designed for mobile purposes, the construction according to embodiments of the present invention can also be applied to fixed-type terminals.
[0051] Please see Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal implementing various embodiments of the present invention. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1 The mobile terminal structure shown does not constitute a limitation on the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0052] The following is combined with Figure 1 A detailed introduction to each component of the mobile terminal:
[0053] The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), and TDD-LTE (Time Division Duplexing-Long Term Evolution).
[0054] WiFi is a short-range wireless transmission technology. Mobile terminals, through the WiFi module 102, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of a mobile terminal and can be omitted as needed without changing the nature of the invention.
[0055] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the mobile terminal 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the mobile terminal 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0056] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage medium) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0057] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0058] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0059] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the mobile terminal. Specifically, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands from processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Specifically, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being limited here.
[0060] Furthermore, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal. The specific implementation is not limited here.
[0061] Interface unit 108 serves as an interface through which at least one external device can connect to mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within mobile terminal 100, or it may be used to transmit data between mobile terminal 100 and the external device.
[0062] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0063] The processor 110 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the mobile terminal, thereby providing overall monitoring of the mobile terminal. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.
[0064] The mobile terminal 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0065] although Figure 1 As not shown, the mobile terminal 100 may also include a Bluetooth module, etc., which will not be described in detail here.
[0066] Based on the above-described mobile terminal hardware structure, various embodiments of the method of the present invention are proposed.
[0067] Figure 2 This is a flowchart of the first embodiment of the video content tagging processing method of the present invention. A video content tagging processing method, the method comprising:
[0068] S1. During video playback, identify and mark video objects in the video content, and obtain the selection operation and annotation operation of the mark;
[0069] S2. Take the video object corresponding to the selected operation and / or the remark operation as the target object;
[0070] S3. Based on the operation time of the selected operation and / or the annotation operation in the video content, track and mark the target object in preset image frames before and after the reference until the target object leaves the video content;
[0071] S4. Merge the timestamp information corresponding to the tracking marker into the video file corresponding to the video content, so that when the video file is played again, a marker playback option corresponding to the timestamp is provided.
[0072] In this embodiment, as Figure 10 As shown, during video playback, users can initiate marking via gestures or buttons; when video playback pauses, the system jumps to the next frame (forward or backward) a certain number of frames based on the user's selection, using the current frame as the reference, to help the user quickly select the target screen; for example... Figure 11 As shown, the system automatically analyzes the current screen, identifies multiple objects within it, and selects the corresponding objects using preset graphic frames; for example... Figure 12 As shown, the user selects one or more objects and completes the input of additional descriptive text; for example... Figure 13 As shown, the system retains the selection box for the target object and displays the entered text next to the object box, while simultaneously deselecting non-target objects. The system identifies whether the target object exists in the video frames before and after the target frame. If it exists, it is considered a continuous frame, and the system continuously adds selection boxes and text to the target object until the object leaves the frame, which is the end of the continuous frame. Finally, the modified frames with added selection boxes and text are saved to the video file. The system records the timestamp of the start of the continuous frame where the user-marked object continues to exist. After the user completes the marking operation, the system automatically saves the video file with added selection boxes and text, as well as the timestamp information. The video file and timestamp information together form a new format video file. The difference between this new format video file and the traditional video file is the addition of timestamp information. During playback, the timestamp information can be parsed first, and then playback can be completed according to traditional video encoding. Users can share the new video file with others, and when others play the video, they can automatically jump to the continuous frame and play it by clicking on the timestamp.
[0073] As can be seen, this embodiment proposes a playback scheme that supports rapid tagging of video content. Users can quickly and conveniently tag content of interest while watching videos and share the tagged videos with others. For example, when watching a game video, if a user finds a particular character outstanding in a team battle, they can quickly tag that character in the battle with the caption "Pro player's play," and share the video with a friend. The friend can then open the video, click on the timestamp, and see the complete process of that character's brilliant performance in the team battle, quickly receiving the friend's sharing intention and improving the effectiveness of message transmission.
[0074] The beneficial effect of this embodiment is that, during video playback, video objects in the video content are identified and marked, and the selected operation and annotation operation of the marked operation are obtained; the video object corresponding to the selected operation and / or the annotation operation is taken as the target object; based on the operation time of the selected operation and / or the annotation operation in the video content, the target object is tracked and marked in preset image frames before and after the benchmark until the target object leaves the video content; the timestamp information corresponding to the tracking mark is fused into the video file corresponding to the video content, so that when the video file is played again, a marked playback option corresponding to the timestamp is provided.
[0075] Figure 3 This is a second flowchart of the video content marking processing method of the present invention. Based on the above embodiment, the step of identifying and marking video objects in the video content during video playback, and obtaining the selection operation and annotation operation of the marking, includes:
[0076] S11. Obtain the file attributes of the video file corresponding to the video content;
[0077] S12. Determine the content theme of the video content based on the file attributes.
[0078] Optionally, in this embodiment, the file attributes include the application that generated the file, thereby determining whether the generated application belongs to video software or game software, etc.
[0079] Optionally, in this embodiment, if it is video software, the content theme is determined based on the title content of the video and the title content in the video content; if it is game software, the content theme is determined based on the game progress, game characters, and task attributes.
[0080] Figure 4This is the third flowchart of the video content marking processing method of the present invention. Based on the above embodiments, the step of identifying and marking video objects in the video content during video playback, and obtaining the selection operation and annotation operation of the marking, further includes:
[0081] S13. Determine the attention features in the video content based on the content theme;
[0082] S14. During the video playback, an object search is performed on the video content based on the attention features, and the searched object is used as the video object.
[0083] Optionally, in this embodiment, as described in the example above, when determining the food-themed video generated by the video software, the features of interest in this embodiment are determined to be one or more of the following: ingredients, cooking utensils, and tableware.
[0084] Optionally, in this embodiment, as described in the example above, when determining the game tower defense themed video generated by the game software, the features of interest in this embodiment are determined to be the tower defense structure and the offensive and defensive activities within the tower defense range.
[0085] Figure 5 This is the fourth flowchart of the video content marking processing method of the present invention. Based on the above embodiments, the step of identifying and marking video objects in the video content during video playback, and obtaining the selection operation and annotation operation of the marking, further includes:
[0086] S15. Determine the corresponding marker box based on the object attributes of the video object;
[0087] S16. During the video playback, the video object is marked in real time using the marker box.
[0088] Optionally, in this embodiment, as described in the example above, the ingredients, cookware, and tableware are distinguished by using one type of marker frame and a different type of marker frame. The form of the marker frame is related to the outline of the object.
[0089] Figure 6 This is the fifth flowchart of the video content marking processing method of the present invention. Based on the above embodiments, the step of identifying and marking video objects in the video content during video playback, and obtaining the selection operation and annotation operation of the marking, further includes:
[0090] S17. Obtain the first touch command inside the mark box, and execute the selection operation through the first touch command;
[0091] S18. Obtain the second touch command of the marked frame line, and execute the annotation operation through the second touch command.
[0092] Optionally, in this embodiment, the first touch instruction inside the mark box and located in the upper half is a single selection instruction, and the selection operation ends after the first touch instruction is executed once.
[0093] Optionally, in this embodiment, the first touch instruction inside the mark box and located in the lower half is used as a selection instruction. After the selection operation is performed once by the first touch instruction, other objects can still be selected.
[0094] Optionally, in this embodiment, the second touch command for the top frame of the marked frame is a single-selection command, and the operation ends after the second touch command is executed once.
[0095] Optionally, in this embodiment, the second touch command for the bottom frame line or side frame line of the marked frame is a selection command. After the annotation operation is performed once by the second touch command, it can still receive annotation operations for other objects.
[0096] Figure 7 This is the sixth flowchart of the video content marking processing method of the present invention. Based on the above embodiments, the step of taking the video object corresponding to the selected operation and / or the annotation operation as the target object includes:
[0097] S21. Obtain the target attributes of the target object;
[0098] S22. Determine the object movement range and object outline body range of the target object based on the target attributes.
[0099] Optionally, in this embodiment, the object's range of motion includes the area where the object is active within the video content.
[0100] Optionally, in this embodiment, the main body of the object outline includes the region formed by the distinguishable features of the object in the video content.
[0101] Figure 8 This is the seventh flowchart of the video content marking processing method of the present invention. Based on the above embodiments, the step of tracking and marking the target object in preset image frames before and after the selected operation and / or the annotation operation in the video content as a reference, until the target object leaves the video content, includes:
[0102] S31. Determine the preceding and following preset image frames based on the object's motion range, and perform time-related tracking marking on the target object in the preceding and following preset image frames;
[0103] S32. When the target object exceeds the main body range of the object outline in the image frame, determine that the target object has left the video content and stop the time-related tracking marker.
[0104] Optionally, in this embodiment, after selecting the target object, a video frame is appropriately acquired based on the change state of the activity area to serve as a supplement to the target object. Then, during the playback, the target object is continuously tracked and marked. When the target object exceeds the main body of the object outline in the image frame, it is determined that the target object has left the video content, and the time-related tracking is stopped.
[0105] Figure 9 This is the eighth flowchart of the video content marking processing method of the present invention. Based on the above embodiments, the step of fusing the timestamp information corresponding to the tracking marker into the video file corresponding to the video content, so that when the video file is played again, a marked playback option corresponding to the timestamp is provided, includes:
[0106] S41. When the video file is played again, the target object and / or the corresponding annotation content are displayed in the video content.
[0107] S42. When the target object and / or the remarks are clicked, the video file is positioned to the video content corresponding to the timestamp and then played.
[0108] Optionally, in this embodiment, when the target object and / or the notes are clicked, a floating box corresponding to the target object and / or the notes is created.
[0109] Optionally, in this embodiment, after locating the video file to the video content corresponding to the timestamp, it is played within the floating frame, while the original video file continues to play. Specifically, the video played within the floating frame is accompanied by sound, while the original video file is played silently.
[0110] Based on the above embodiments, the present invention also proposes a video content tagging processing device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the video content tagging processing method as described in any of the above embodiments.
[0111] It should be noted that the above-described device embodiments and method embodiments belong to the same concept. The specific implementation process can be found in the method embodiments, and the technical features in the method embodiments are also applicable to the device embodiments, which will not be repeated here.
[0112] Based on the above embodiments, the present invention also proposes a computer-readable storage medium storing a video content tagging processing program, which, when executed by a processor, implements the steps of the video content tagging processing method as described in any of the above claims.
[0113] It should be noted that the above-described medium embodiments and method embodiments belong to the same concept. The specific implementation process can be found in the method embodiments, and the technical features in the method embodiments are also applicable to the medium embodiments, which will not be repeated here.
[0114] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0115] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0116] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0117] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method of video content marking processing, characterized by, The method includes: During video playback, video objects within the video content are identified and marked, and selection and annotation operations for the marked objects are obtained. Specifically, the file attributes of the video file corresponding to the video content are obtained; the content theme of the video content is determined based on the file attributes; attention features in the video content are determined based on the content theme; during video playback, an object search is performed on the video content based on the attention features, and the searched objects are used as the video objects; a corresponding marking box is determined based on the object attributes of the video objects; the video objects are marked in real-time using the marking boxes during video playback; a first touch command is obtained within the marking box, and the selection operation is executed using the first touch command; a second touch command is obtained for the marking box outline, and the annotation operation is executed using the second touch command. The video object corresponding to the selected operation and / or the annotation operation is taken as the target object, wherein the target attributes of the target object are obtained; and the object motion range and the main body range of the object outline are determined according to the target attributes. Using the selected operation and / or the annotation operation's operation time in the video content as a reference, the target object is tracked and marked in preset image frames before and after the reference until the target object leaves the video content. The preset image frames before and after the target object are determined based on the object's movement range, and time-related tracking is performed on the target object in these preset image frames. When the target object exceeds the main outline of the object in the image frame, it is determined that the target object has left the video content, and time-related tracking is stopped. The timestamp information corresponding to the tracking marker is merged into the video file corresponding to the video content, so that when the video file is played again, a playback option corresponding to the timestamp is provided. When the video file is played again, the target object and / or the annotation content corresponding to the target object are displayed floating in the video content. When the target object and / or the annotation content is clicked, the video file is positioned to the video content corresponding to the timestamp and then played.
2. A video content marking processing apparatus characterized by comprising: The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the video content tagging processing method as described in claim 1.
3. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a video content tagging processing program, which, when executed by a processor, implements the video content tagging processing method as described in claim 1.