Video processing method and device, equipment and storage medium
By placing display elements on video image frames and setting marker information, the display of automatically following objects is achieved, which solves the inefficiency problem caused by manual frame-by-frame operation in the prior art and improves video processing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-10-12
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies require manual operation frame by frame when adding display elements in video processing, resulting in low processing efficiency.
By placing display elements on the image frames of the video and setting label information for different objects, the display elements automatically follow the selected object in subsequent frames in response to user selection, simplifying the processing steps.
It improves the efficiency of video processing, reduces manual operations, and simplifies the video processing workflow.
Smart Images

Figure CN121888015A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a video processing method, apparatus, device, and storage medium. Background Technology
[0002] In video processing, one method involves adding display elements (such as stickers, emoticons, text, etc.) to the video's image frames to increase its appeal. Of course, display elements can also serve to obscure content.
[0003] In related technologies, adding display elements to video image frames typically requires manual addition frame by frame. Alternatively, one can manually select the area to be followed within an image frame, allowing the display element to follow that area across multiple image frames.
[0004] However, both manually adding display elements frame by frame and manually selecting the area that the display elements follow are cumbersome processes, resulting in low efficiency in video processing. Summary of the Invention
[0005] This application provides a video processing method, apparatus, device, and storage medium, which can improve video processing efficiency. The technical solution provided by this application includes the following aspects.
[0006] According to one aspect of the embodiments of this application, a video processing method is provided, the method comprising the following steps.
[0007] A first display element is placed on the first image frame of the first video, wherein the first image frame is one of a plurality of image frames included in the first video;
[0008] Display at least one object corresponding to each object, wherein the at least one object includes objects that appear in the first image frame that can be followed by the first display element, and different marking information is used to distinguish different objects;
[0009] In response to the operation of selecting a first object among the at least one objects, if the first object appears in a second image frame of the first video, the first display element following the first object is displayed on the second image frame, which is different from the first image frame.
[0010] According to one aspect of the embodiments of this application, a video processing apparatus is provided, the apparatus comprising the following modules.
[0011] A placement module is used to place a first display element on a first image frame of a first video, wherein the first image frame is one of a plurality of image frames included in the first video;
[0012] The display module is used to display the marking information corresponding to at least one object, wherein the at least one object includes objects that appear in the first image frame that can be followed by the first display element, and different marking information is used to distinguish different objects;
[0013] The display module is further configured to, in response to an operation of selecting a first object among the at least one objects, display the first display element following the first object on the second image frame when the first object appears in the second image frame of the first video, wherein the second image frame is different from the first image frame.
[0014] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described video processing method.
[0015] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the above-described video processing method.
[0016] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program, the computer program being loaded and executed by a processor to implement the above-described video processing method.
[0017] The technical solutions provided in this application can bring the following beneficial effects.
[0018] By displaying marker information corresponding to at least one object, different objects in an image frame can be distinguished, allowing the user to select different objects. After selecting the first object, there is no need to manually add the first display element again in the second image frame; instead, the first display element of the first object following the first image frame is directly displayed, simplifying the video processing steps and thus improving video processing efficiency. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a computer system provided in one embodiment of this application;
[0020] Figure 2 This is a schematic diagram of a video processing method provided in one embodiment of this application;
[0021] Figure 3 This is a flowchart of a video processing method provided in one embodiment of this application;
[0022] Figure 4This is a flowchart of a video processing method provided in another embodiment of this application;
[0023] Figure 5 This is a flowchart of a video processing method provided in another embodiment of this application;
[0024] Figure 6 This is a schematic diagram of a video processing method provided in another embodiment of this application;
[0025] Figure 7 This is a block diagram of a video processing apparatus provided in one embodiment of this application;
[0026] Figure 8 This is a block diagram of a video processing apparatus provided in another embodiment of this application;
[0027] Figure 9 This is a structural block diagram of a computer device provided in one embodiment of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0029] Please refer to Figure 1 This illustration shows a schematic diagram of a computer system provided in an exemplary embodiment of this application. The computer system may include: a terminal device 10 and a server 20.
[0030] Terminal device 10 includes, but is not limited to, mobile phones, tablets, smart voice interaction devices, game consoles, wearable devices, multimedia playback devices, PCs (Personal Computers), in-vehicle terminals, smart home appliances, and other electronic devices. A client application for the target application can be installed on terminal device 10. Optionally, the target application can be an application that requires downloading and installation, or it can be an application that can be used instantly; this embodiment of the application does not limit this.
[0031] In this embodiment of the application, the target application is an application for video processing. For example, as shown... Figure 2As shown, a first display element 270 (text "Happy Birthday") is placed on the first image frame 200 of the first video. The first image frame 200 is one of a plurality of image frames included in the first video. At least one object 210 is displayed, each corresponding to a marker information 220. Each object 210 includes objects that the first display element 270 can follow, appearing in the first image frame 200. Different marker information 220 is used to distinguish different objects. In response to the operation of selecting a first object from at least one object, if the first object appears in a second image frame of the first video, a first display element following the first object is displayed on the second image frame. The second image frame is different from the first image frame. Of course, the specific type of the target application is not limited. The target application can be a video processing application, a video editing application, a video playback application, an animation application, a social application, a question-and-answer application, a virtual reality application, an augmented reality application, etc. This application embodiment does not limit the specific category of the target application. In other embodiments, the target application can be considered a separate functional module, such as being implemented as one of the functional modules in a video playback application for implementing video processing. For example, a client running the aforementioned target application is located in terminal device 10.
[0032] Server 20 is used to provide backend services for the client of the target application in terminal device 10. For example, server 20 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but it is not limited to these.
[0033] Terminal device 10 and server 20 can communicate with each other via a network. This network can be a wired network or a wireless network.
[0034] The method provided in this application embodiment can be executed by a computer device in each step. The computer device can be any electronic device capable of data storage and processing. For example, the computer device can be... Figure 1 The terminal device 10 can also be a server 20.
[0035] Please refer to Figure 3This document illustrates a flowchart of a video processing method provided in an embodiment of this application. The execution subject of this method may be the terminal device 10 described above, such as the client of the target application mentioned in the above embodiments, or the server 20 described above. In the following method embodiments, for ease of description, the execution subject of each step is only described as a "computer device". The method may include at least one of the following steps (310-330):
[0036] Step 310: Place a first display element on the first image frame of the first video. The first image frame is one of the multiple image frames included in the first video.
[0037] In some embodiments, the first video is a video to be processed. The first video may be a video uploaded by the user or a video existing in the target application. This application does not limit the duration of the first video, but the first video includes at least two image frames. This application does not limit the type of the first video; it may be a video directly captured using a recording device, a video synthesized frame by frame, or a video generated using an artificial intelligence algorithm.
[0038] In some embodiments, the target application provides video processing functions, such as video editing functions, to the first video. Exemplarily, in processing the first video, the target application first splits the first video into multiple image frames. Exemplarily, the time interval between any two adjacent image frames is the same, such as 0.1 seconds. Exemplarily, each image frame corresponds to the display screen of the first video at one timestamp. Exemplarily, the display screen of the first video at different timestamps is different, that is, different image frames are different.
[0039] In some embodiments, a timeline corresponding to the first video is displayed. This timeline indicates the duration of the first video. In response to selecting a first timestamp in the timeline, image frames of the first video at that first timestamp are displayed.
[0040] In some embodiments, the first image frame is any one of a plurality of image frames in the first video. In some embodiments, the first image frame is the first image frame, an intermediate image frame, or the homepage of the first video. The homepage of the first video is either the first image frame of the first video or the most representative image frame in the first video. The homepage of the first video can be set by the user or the first image frame of the first video can be automatically determined as the homepage.
[0041] In some embodiments, the first display element is an element selected by the user to be placed on top of an image frame. Exemplarily, the type of the first display element includes, but is not limited to, text, images, emoticons, etc. Exemplarily, the first display element can be uploaded by the user or provided by the target application described above. Exemplarily, during the processing of the first video by the target application, multiple display elements are displayed for the user to select and place on the image frames of the first video. Exemplarily, the first display element on the first image frame can be moved, zoomed in, or zoomed out by the user, etc.
[0042] In some embodiments, in response to a click operation on an editing control of the first video, a first image frame and a plurality of display elements that can be placed on the first image frame are displayed. In response to an operation to select a first display element from the plurality of display elements, the first display element is placed at a default position (e.g., the center position) of the first image frame. In response to an editing operation on the placed first display element (after the first display element is placed on the first image frame, the first display element has a corresponding editing bar containing a plurality of controls for editing the first display element, and clicking these controls enables editing of the first display element), the first display element is edited (e.g., modifying the content, display size, display position, etc. of the first display element). Exemplarily, the edited first display element is considered to be the first display element placed on the first image frame in step 310.
[0043] In some embodiments, when processing the first video through the target application described above, the first image frame is displayed directly. Alternatively, the first image frame is displayed in response to an operation on a target timestamp on the timeline. For example, in response to a drag operation on a pointer on the timeline, the timestamp where the pointer is located is considered the selected timestamp, and the image frame corresponding to that timestamp is displayed.
[0044] Step 320: Display the marking information corresponding to at least one object, wherein the at least one object includes objects that appear in the first image frame that can be followed by the first display element, and different marking information is used to distinguish different objects.
[0045] In some embodiments, at least one object includes objects that appear in a first image frame and can be followed by a first display element. Exemplarily, objects that can be followed by a first display element are identified from a plurality of image frames included in the first video. Exemplarily, objects that can be followed by a first display element are identified from a first image frame. Exemplarily, objects that can be followed by a first display element can also be identified from other image frames. In some embodiments, each of the at least one object appears in the first image frame, but there may also be objects that do not appear in the first image frame but appear in other image frames.
[0046] In some embodiments, the object in this application is an object appearing in an image frame. This object may be fully displayed in the image frame or partially displayed, such as only half-displayed due to screen display limitations; this application does not impose this limitation. In some embodiments, the object may be in the form of a person, an animal, a cartoon character, or other forms; this application does not impose this limitation. The object may be displayed in three-dimensional or two-dimensional form; this application does not impose this limitation. In some embodiments, the object is an independent entity, such as a person, an animal, a plant, a building, etc. Conversely, a single hair of a person, a branch of a tree, or the stamen of a flower, which cannot be considered an independent entity, are not considered objects in this application.
[0047] In some embodiments, "display element following object" can be understood as meaning that when the position of the object changes within an image frame, the display element's position within that image frame also changes. For example, if an object moves to the left across multiple image frames, then if that object is the object the display element follows, the display element will also move to the left across those multiple image frames. Similarly, if an object moves to the right across multiple image frames, then if that object is the object the display element follows, the display element will also move to the right across those multiple image frames. In some embodiments, for the object followed by the first display element, the relative positions of the first display element and the object do not change or change only slightly across different image frames. For example, regardless of how the image frames change, as long as the object appears in the image frame, the first display element always appears to the left of the object.
[0048] In some embodiments, at least one object is associated with specific tagging information to distinguish different objects. Exemplarily, the tagging information is information used to mark objects. Exemplarily, the tagging information includes a location tag to indicate the object's display position within an image frame. Exemplarily, the tagging information includes a type tag to indicate the object's type. For example, if three objects appear in the first video—a dog, a cat, and a person—the object's type is directly indicated by text; displaying the text "person," "dog," and "cat" distinguishes the three objects in the first video. Exemplarily, the tagging information includes a feature tag to indicate the object's characteristics. For example, if two objects appear in the first video—a white dog and a black dog—the object's color characteristic is directly indicated by text; displaying the text "white dog" and "black dog" distinguishes the two objects in the first video. For example, if two objects appear in the first video—a large dog and a small dog—the object's volume characteristic is directly indicated by text; displaying the text "large dog" and "small dog" distinguishes the two objects in the first video. Exemplarily, the tagging information includes a name tag to indicate the object's name. For example, if there are two objects in the first video, a dog named Huahua and a dog named Gou Dan, then the names of the objects are directly indicated by text. The two objects in the first video can be distinguished by displaying the text "Huahua" and "Gou Dan".
[0049] Step 330, in response to the operation of selecting a first object among at least one objects, if the first object appears in a second image frame of the first video, a first display element following the first object is displayed on the second image frame, the second image frame being different from the first image frame.
[0050] In some embodiments, the operation of selecting a first object includes the operation of selecting the tag information corresponding to the first object. For example, the operation of selecting the tag information corresponding to the first object includes, but is not limited to, at least one of the following: selecting a location tag corresponding to the first object (e.g., selecting a rectangle within a location tag that frames the puppy), selecting a type tag corresponding to the first object (e.g., selecting the text "person" in a type tag), selecting a feature tag corresponding to the first object (e.g., selecting the text "white dog" in a feature tag), and selecting a name tag corresponding to the first object (e.g., selecting the text "flower" in a name tag).
[0051] In some embodiments, the marker information can be displayed on top of the image frame or next to the image frame. Exemplarily, different marker information can be implemented as different user interface (UI) controls. Exemplarily, each marker is considered a UI control, capable of responding to at least one of a user's click, double-click, long-press, etc., to select the object corresponding to that marker information as the object that the first display element wants to follow.
[0052] UI controls are any visual controls or elements visible on the user interface of an application, such as images, input boxes, text boxes, buttons, labels, etc. Some UI controls respond to user actions, such as a first follow control or a second follow control, used to control the display element to follow the object. The UI controls involved in the embodiments of this application include, but are not limited to: the tag information corresponding to the object, and the first follow control or the second follow control.
[0053] In some embodiments, the second image frame is another image frame among a plurality of image frames that is different from the first image frame. For example, after selecting the first object as the object followed by the first display element, if the first object also appears in the second image frame, the first display element following the first object is automatically displayed on top of the first image frame.
[0054] In some embodiments, the second image frame may be an image frame that precedes the first image frame among a plurality of image frames, or an image frame that follows the first image frame among a plurality of image frames.
[0055] In some embodiments, the first image frame serves only as a reference for the relative position between the first display element and the first object, so that the first display element following the first object is displayed in the second image frame. Whether the first image frame in the final processed first video displays the first display element is not limited in this application.
[0056] In some embodiments, after selecting the first object, it is not necessary to select the video frame to which the first display element applies.
[0057] For example, after selecting a first object, a first display element following the first object is displayed in all image frames containing the first object, including image frames containing the first object before and after the first image frame.
[0058] For example, after selecting the first object, the system responds to an action on the confirmation control, confirming the selection of the first object. Figure 2As shown, after the first object is selected, the operation on the confirmation control 250 is responded to, confirming the selection of the first object.
[0059] For example, after a first object is selected, in response to a click operation on the save control, a first display element is applied in multiple image frames, that is, a first display element following the first object is displayed in multiple image frames in which the first object appears.
[0060] For example, after the first object is selected, in response to a click operation on the save control 260, the saving of the first video processing is confirmed, and the first display element is applied in multiple image frames.
[0061] In other embodiments, after selecting the first object, it is necessary to select the video frame to which the first display element is applied from among multiple image frames.
[0062] For example, the set of video frames to which the selected first display element is applied is considered the first timestamp interval. The first timestamp interval includes multiple image frames, which may be consecutive or discrete, and this application does not limit this. The first timestamp interval may or may not include the first image frame. If the first image frame is not included in the first timestamp interval, the first display element following the first object is not displayed on the first image frame.
[0063] For example, after selecting a first object and a video frame (such as a first timestamp interval) of the selected application, in response to a click operation on the application control, a first display element is applied to multiple image frames included in the first timestamp interval, that is, a first display element following the first object is displayed in multiple image frames in the first timestamp interval where the first object appears.
[0064] The technical solution provided in this application distinguishes different objects in an image frame by displaying marking information corresponding to at least one object, allowing the user to select different objects. After selecting the first object, there is no need to manually add the first display element again in the second image frame; instead, the first display element of the first object following the first object in the second image frame is directly displayed, simplifying the video processing steps and thus improving video processing efficiency.
[0065] Please refer to Figure 4This document illustrates a flowchart of a video processing method provided in another embodiment of this application. The execution subject of this method may be the terminal device 10 described above, such as the client of the target application mentioned in the above embodiments, or the server 20 described above. In the following method embodiments, for ease of description, the execution subject of each step is only described as a "computer device". The method may include at least one of the following steps (410-440):
[0066] Step 410: Place a first display element on the first image frame of the first video. The first image frame is one of the multiple image frames included in the first video.
[0067] In some embodiments, after step 410, step 420 or step 430 may be performed.
[0068] Step 420: On top of the first image frame, at least one location marker is displayed, each of the at least one location marker being used to mark the position of one of the at least one objects in the first image frame.
[0069] For example, position markers are used to indicate the position of an object in the first image frame. For example, for each object in the first image frame, position markers are used to mark each followable object. For example, position markers include marker rectangles, object strokes, etc. For example, each object corresponds to a marker rectangle, which may completely enclose the object, or the center of the marker rectangle may coincide with the center of the object, to indicate the object. For example, each object corresponds to an object stroke, which is used to indicate the edge of the object in the first image frame, and the object stroke is obtained by drawing the edge of the object in the first image frame using a thick line. For example, each position marker can be selected. For example, different selected position markers result in different selected followable objects.
[0070] like Figure 2 As shown, the marking information is a marked rectangle.
[0071] The technical solution provided in this application improves information display efficiency by directly displaying position markers on the first image frame to indicate the position of the object within the first image frame. Furthermore, it simplifies the process of selecting the first object (related technologies require manually selecting an area, while this application allows direct clicking of the marker information), thus improving the efficiency and accuracy of selecting the first object.
[0072] The following section describes the operation of selecting the first object, based on step 420.
[0073] In some embodiments, the operation of selecting a first object includes: selecting a position marker corresponding to the first object.
[0074] In some embodiments, in response to the operation of selecting the position marker corresponding to the first object, the position marker corresponding to the first object and the position markers corresponding to at least one other object besides the first object are displayed separately on the upper layer of the first image frame.
[0075] For example, the operation of selecting the position marker corresponding to the first object is an operation for selecting the position marker corresponding to the first object, and is also an operation for selecting the first object. For example, the operation type of the operation of selecting the position marker corresponding to the first object includes, but is not limited to, clicking, long-pressing, double-clicking, etc., on the position marker corresponding to the first object.
[0076] For example, after selecting the position marker corresponding to the first object, the position marker corresponding to the first object is displayed differently from the position markers corresponding to the other objects among at least one object besides the first object. For example, the position marker corresponding to the first object and the position markers corresponding to the other objects among at least one object are displayed differently by color, such as highlighting the position marker corresponding to the first object in yellow. For example, the position marker corresponding to the first object and the position markers corresponding to the other objects among at least one object are displayed differently by brightness, such as highlighting the position marker corresponding to the first object.
[0077] like Figure 6 As shown, in response to the operation of selecting the position mark 610 corresponding to the first object, the position mark 610 corresponding to the first object and the position marks corresponding to at least one other object besides the first object are displayed separately in the upper layer of the first image frame.
[0078] like Figure 6 As shown, after confirming the selection of the first object, the first display element 630 is controlled to follow the first object 620. Exemplarily, during video processing, a virtual line connecting the first object 620 and the first display element 630 is also displayed; this virtual line indicates the object followed by the first display element 630.
[0079] Furthermore, when the first follow control 650 or the second follow control 640 is clicked again, the first display element can be controlled to unfollow the first object. For example, in response to another operation on the first or second follow control, the tag information corresponding to at least one object is redisplayed. In response to the operation of selecting the tag information of the first object again, following the first object is canceled.
[0080] The technical solution provided in this application clearly informs the user of the selected first object by distinguishing its position marker from the marker information corresponding to other objects, thereby reducing selection errors caused by accidental touches. While enriching the display methods, it also reduces the frequency of video processing errors.
[0081] Step 430: Display object tags corresponding to at least one object, whereby the object tags represent at least one of the following: object name, object thumbnail, object number, and object type.
[0082] For example, referring to the explanation of the above embodiments, the object tag includes at least one of the name tag, type tag, feature tag, etc. For example, the name tag is used to represent the object name, the type tag is used to represent the object type, and the feature tag is used to represent the object features.
[0083] Of course, object labeling can also be object thumbnails and object numbers. For example, at least one object is captured, and key portions of the object are cropped to create thumbnails (e.g., a screenshot of the object's header is used to create an object thumbnail). Displaying different object thumbnails helps the user understand that different object thumbnails correspond to different objects, which is more intuitive. For example, at least one object is captured, and each object is numbered, such as object 1, object 2, and object 3. For example, different object numbers correspond to different objects.
[0084] In some embodiments, unlike location markers, object markers may not be displayed on top of image frames, but rather independently of image frames, such as next to the first image frame, or in a blank space in the editing interface of the first video.
[0085] The technical solution provided in this application improves information display efficiency by displaying object markers for different objects to indicate which objects an object can follow. Furthermore, it simplifies the process of selecting the first object (in related technologies, a region needs to be manually selected, while in this application, the object marker can be selected directly), improving the efficiency and accuracy of selecting the first object.
[0086] The following section describes the operation of selecting the first object, based on step 430.
[0087] In some embodiments, the operation of selecting a first object includes: selecting the object tag corresponding to the first object.
[0088] In some embodiments, in response to the operation of selecting an object tag corresponding to a first object, at least one other object in an object other than the first object is masked on top of the first image frame.
[0089] For example, the operation of selecting the object marker corresponding to the first object is an operation for selecting the object marker corresponding to the first object, and is also an operation for selecting the first object. For example, the operation type of the operation of selecting the object marker corresponding to the first object includes, but is not limited to, clicking, long-pressing, double-clicking, etc., on the object marker corresponding to the first object.
[0090] For example, after selecting the location marker corresponding to the first object, at least one other object in the first image frame is masked on top of the first image frame.
[0091] Here is a brief explanation of masking. Masking refers to obscuring some elements in an image while retaining others. The masking in this application is mainly used to obscure unselected areas while retaining selected areas, such as retaining the area where the first object is displayed, and obscuring other areas besides the first object. Of course, this masking is mainly achieved using a masking layer on top of the image frame, which obscures areas other than the first object. For example, gray or black areas are used to obscure areas other than the first object, while the area containing the first object remains unobscurated.
[0092] The technical solution provided in this application, by masking at least one object other than the first object, clearly informs the user of the selected first object, reducing selection errors caused by accidental touches. While enriching display methods, it can also reduce the frequency of video processing errors.
[0093] As can be seen from step 430 above, the object marker is not directly displayed on the first image frame. Therefore, the object corresponding to the object marker may not appear in the first image frame, but may appear in other image frames. That is, at least one object may include not only the object appearing in the first image frame, but also the object appearing in other image frames.
[0094] Based on this situation, the following describes the positioning and display of image frames combined with object markers.
[0095] In some embodiments, at least one object further includes a second object, which is an object that the first display element can follow and appears in other image frames besides the first image frame.
[0096] For example, the second object is an object that does not appear in the first image frame, but appears in the third image frame.
[0097] For example, the tagging information corresponding to at least one of the displayed objects includes the tagging information corresponding to the second object.
[0098] In some embodiments, in response to the operation of selecting the object tag corresponding to the second object, the first image frame to be displayed is changed to a third image frame, while the first display element is retained.
[0099] For example, the operation of selecting the object marker corresponding to the second object is an operation for selecting the object marker corresponding to the second object, and is also an operation for selecting the second object. For example, the operation type of the operation of selecting the object marker corresponding to the second object includes, but is not limited to, clicking, long-pressing, double-clicking, etc., of the position marker corresponding to the first object.
[0100] For example, after selecting a second object, the image frame in which the second object appears is determined, such as a third image frame, which is different from the first image frame. The third image frame can be an image frame before or after the first image frame.
[0101] In some embodiments, the third image frame is an image frame in which the second object appears among multiple image frames, and the third image frame is different from the first image frame. For example, there are multiple image frames in which the second object appears, and the third image frame is the first of the multiple image frames in which the second object appears, that is, the third image frame is the image frame in which the second object first appears.
[0102] For example, the currently displayed first image frame is changed to a third image frame, and the third image frame is displayed. At the same time, the first display element is retained, and the display position of the first display element remains unchanged, still in the same position as before on the first image frame.
[0103] The technical solution provided in this application, when displaying the marker information corresponding to at least one object, not only displays the marker information corresponding to the object appearing in the current image frame (first image frame), but also displays the marker information corresponding to the object appearing in other image frames. When the marker information corresponding to an object not appearing in the current image frame is selected, the system directly jumps to display the image frame containing the selected object, achieving image frame positioning display based on object selection. This method eliminates the need for the user to select the timestamp of the third image frame before displaying it, directly replacing the display method, simplifying video processing steps and accelerating video processing efficiency. Simultaneously, retaining the display of the first display element avoids repositioning the display element on the third image frame, further reducing video processing costs.
[0104] Step 440, in response to the operation of selecting a first object among at least one objects, if the first object appears in a second image frame of the first video, a first display element following the first object is displayed on the second image frame, the second image frame being different from the first image frame.
[0105] This application provides two methods, steps 420 and 430, to display the marking information, which enriches the information display methods and enhances the diversity and flexibility of object selection.
[0106] Please refer to Figure 5 This document illustrates a flowchart of a video processing method according to another embodiment of this application. The execution subject of this method may be the terminal device 10 described above, such as the client of the target application mentioned in the above embodiments, or the server 20 described above. In the following method embodiments, for ease of description, the execution subject of each step will only be described as a "computer device". The method may include at least one of the following steps (510-530):
[0107] Step 510: Place a first display element on the first image frame of the first video. The first image frame is one of the multiple image frames included in the first video.
[0108] The following steps 512 and 514 describe how to determine at least one of the following objects. Steps 512 and 514 only use the first image frame as an example. When at least one object also includes objects appearing in other image frames, the following subject detection model is also needed to identify objects in other image frames. The specific identification method is detailed in steps 512 and 514 below and will not be repeated here.
[0109] Step 512: Identify the first image frame using the subject detection model to obtain at least one subject in the first image frame and the subject recognition score corresponding to each subject. The subject detection model is a neural network model used to identify subjects in the image frame, and the subject recognition score is used to characterize the confidence level of the subject identified by the subject detection model.
[0110] In some embodiments, the subject detection model is a neural network model for detecting subjects in an image frame. In some embodiments, the input of the subject detection model is a first image frame, and the output is at least one subject and a subject recognition score corresponding to each subject. Exemplarily, the subject recognition score is used to characterize the confidence level of the subject identified by the subject detection model, such as a confidence level ranging from 0 to 1.
[0111] In some embodiments, the input to the subject detection model is a first image frame, and the output is a subject bounding box corresponding to at least one subject, a subject type, and a subject recognition score corresponding to at least one subject. The subject bounding box is used to indicate the subject within the box. In some embodiments, the subject bounding box and the aforementioned marker rectangle can be the same. In some embodiments, subjects with subject recognition scores greater than or equal to a score threshold are retained from the at least one subject as at least one object, and the subject bounding boxes corresponding to these objects are displayed (i.e., the marker information corresponding to at least one object is displayed).
[0112] In some embodiments, the subject detection model is a pre-trained model. A pre-trained model, also known as a foundational model or large model, refers to a deep neural network (DNN) with a large number of parameters, trained on massive amounts of unlabeled data. Utilizing the function approximation capability of the large-parameter DNN, the PTM extracts common features from the data. Through fine-tuning, efficient parameter fine-tuning, prompt-tuning, and other techniques, it is suitable for downstream tasks. Therefore, the pre-trained model can achieve ideal results in small-shot or zero-shot scenarios. PTMs can be categorized according to the data modality they process, including language models, visual models (swin-transformer, ViT (Vision Transformers), V-MOE (Vision Mixture of Experts)), speech models, and multimodal models. Multimodal models refer to models that establish feature representations for two or more data modalities. Pre-trained models are important tools for outputting AI-generated content and can also serve as a general interface connecting multiple specific task models. In the embodiments of this application, the subject detection model, or at least one module or sub-model within the subject detection model, is a pre-trained model.
[0113] Step 514: Retain subjects whose subject identification scores are greater than or equal to a score threshold from at least one subject as at least one object.
[0114] In some embodiments, subjects with a subject identification score greater than or equal to a score threshold are retained as one of at least one object. In some embodiments, subjects with a subject identification score greater than or equal to a score threshold are removed. The score threshold is a preset value.
[0115] The technical solution provided in this application uses a subject detection model to determine at least one object, which not only improves the accuracy of object determination but also increases the efficiency of object determination. This is beneficial for further improving the efficiency of video processing.
[0116] Step 520: Display the marking information corresponding to at least one object, wherein the at least one object includes objects that appear in the first image frame that can be followed by the first display element, and the different marking information is used to distinguish different objects.
[0117] The following explains how to trigger the display of marker information.
[0118] In some embodiments, prior to step 520, an operation on a first follow control on the editing interface of the first video is received. In response to the operation on the first follow control on the editing interface of the first video, marker information corresponding to at least one object is displayed.
[0119] For example, the operation on the first follow control in the editing interface of the first video is used to trigger the first display element to follow, and also to trigger the display of the marker information corresponding to at least one object. For example, the operation types of the operation on the first follow control in the editing interface of the first video include, but are not limited to, clicking, long-pressing, double-clicking, etc. on the first follow control.
[0120] like Figure 2 As shown, in response to an operation of the first follow control 240 on the editing interface for the first video, at least one object's corresponding tag information 220 is displayed.
[0121] In some embodiments, prior to step 520, an operation is received on a second follow control in the edit bar of the first display element. In response to the operation on the second follow control in the edit bar of the first display element, marker information corresponding to at least one object is displayed. The first follow control and the second follow control are different.
[0122] For example, the position of the edit bar of the first display element may be fixed or changeable. For example, when the position of the first display element changes, the position of the edit bar of the first display element remains unchanged. For instance, the edit bar of the first display element is always displayed at a set position on the user interface.
[0123] For example, when the position of the first display element changes, the position of the editing bar of the first display element also changes. For example, the editing bar of the first display element is displayed around the first display element, such as the editing bar surrounding the first display element. For example, the editing bar of the first display element includes multiple controls for editing the first display element, including a second following control. For example, the multiple controls for editing the first display element surround the first display element. When the position of the first display element changes, the positions of the multiple controls for editing the first display element remain centered on the position of the first display element and change synchronously.
[0124] For example, the operation on the second follow control is used to trigger the first display element to follow, and also to trigger the display of the marker information corresponding to at least one object. For example, the operation types on the second follow control include, but are not limited to, clicking, long-pressing, double-clicking, etc. on the second follow control.
[0125] like Figure 2 As shown, in response to an operation on the second follow control 230 in the edit bar of the first display element, at least one object's corresponding tag information 220 is displayed.
[0126] The technical solution provided in this application provides two ways to trigger the display of marker information: one is to directly click the first follow control in the video editing interface, and the other is to click the second follow control in the editing bar of the first display element. Triggering the display of marker information in multiple ways, and further selecting the follow object, facilitates triggering the display element to follow the object, thereby improving video processing efficiency. It also enriches the information display methods.
[0127] Step 530, in response to the operation of selecting a first object among at least one objects, if the first object appears in a second image frame of the first video, a first display element following the first object is displayed on the second image frame, the second image frame being different from the first image frame.
[0128] The following describes how to select or not select the image frame applied to the first display element after selecting the first object.
[0129] In the first case, you need to select an image frame.
[0130] In some embodiments, in response to the operation of selecting a first object, a first timestamp interval is selected from the timestamp intervals of multiple image frames, the first timestamp interval including at least one image frame, and the second image frame being an image frame in the first timestamp interval.
[0131] For example, after selecting the first object, it is also necessary to select the image frames of the application. For example, a first timestamp interval is selected from the timestamp intervals of multiple image frames. For example, a first timestamp interval is selected from the timeline of multiple image frames. For example, the application start image frame and the application end image frame are selected from the multiple image frames, and the timestamp interval between the application start image frame and the application end image frame is used as the first timestamp interval. The application start image frame is the first image frame in the first timestamp interval, and the application end image frame is the last image frame in the first timestamp interval.
[0132] For example, there can be one or more first timestamp intervals. When there are multiple first timestamp intervals, no two of the multiple timestamp intervals overlap. That is, multiple timestamp intervals can be selected simultaneously to apply the first display element.
[0133] In some embodiments, a first display element is controlled to follow a first object appearing in at least one image frame. For example, for multiple image frames within a selected first timestamp interval, when a first object appears, the first display element is controlled to follow that first object.
[0134] The technical solution provided in this application allows users to select the image frame to which the first display element is to be applied, demonstrating the flexibility and diversity of video processing.
[0135] In the second case, there is no need to select an image frame.
[0136] In some embodiments, the first display element is controlled to appear in the first image frame and then in the image frames following the first image frame in the plurality of image frames, and the second image frame is any image frame following the first image frame in the plurality of image frames.
[0137] For example, in multiple image frames starting from the first image frame and continuing into subsequent image frames, such as the second image frame starting from the first image frame, it is determined whether the first object appears in the second image frame. If it appears, the first display element is controlled to follow the first object. If it does not appear, the first display element is not placed in the second image frame.
[0138] For example, when the first display element is controlled to follow the first object in multiple image frames starting from the first image frame, if the processed video is played, the first display element following the first object is displayed in the image frames starting from the first image frame.
[0139] The technical solution provided in this application automatically applies the first display element to image frames starting from the first image frame, reducing operation steps and improving video processing efficiency.
[0140] The following explains how to cancel following.
[0141] In some embodiments, in response to a cancel follow operation for a fourth image frame, the first display element is controlled to stop following the first object appearing in the first video starting from the fourth image frame, which is an image frame following the first image frame among a plurality of image frames.
[0142] For example, the unfollow operation for the fourth image frame includes, after selecting the fourth image frame, performing a second operation on the follow control (including at least one of the first and second follow controls mentioned above). For example, when the operation on the follow control is performed for the first time, following begins; when the operation on the follow control is performed again, following is canceled. For example, when the currently displayed image frame is the first image frame, performing the operation on the follow control is considered as starting following from the first image frame. For example, when the currently displayed image frame is the fourth image frame, performing the operation on the follow control again is considered as starting unfollowing from the fourth image frame.
[0143] For example, the unfollow operation for the fourth image frame includes, after selecting the fourth image frame, an operation on the unfollow control. For example, after starting follow, if the currently displayed image frame is the fourth image frame, performing the operation on the unfollow control is considered as starting unfollowing from the fourth image frame. For example, the unfollow control and the follow control are different controls. For example, the unfollow control is displayed in the editing interface of the first video or in the editing bar of the first display element.
[0144] For example, canceling follow can also be understood as meaning that the first display element following the first object will not appear in any image frame after the fourth image frame, regardless of whether the first object appears or not.
[0145] The technical solution provided in this application allows users to choose to cancel following image frames, which reflects the flexibility of following and enriches the forms of human-computer interaction.
[0146] The following describes the continuous adjustments triggered after the position of the first displayed element is adjusted.
[0147] In some embodiments, in response to an adjustment operation on a first display element displayed on a fifth image frame, the adjusted first display element is displayed, wherein the fifth image frame is a second image frame or another image frame other than the second image frame that displays the first display element.
[0148] For example, the adjustment operation for the first display element displayed on the fifth image frame is an operation used to adjust the first display element displayed on the fifth image frame. For example, the operation types for adjusting the first display element displayed on the fifth image frame include, but are not limited to, clicking, long-pressing, double-clicking, etc., of various adjustment controls in the editing bar of the first display element displayed on the fifth image frame. Different adjustment controls correspond to different adjustment effects, such as adjusting position, adjusting display size, adjusting text content, etc. Of course, a second display element can also be used to replace the first display element, such as the adjusted first display element becoming the second display element.
[0149] It is important to note that if the adjustment of the first display element begins from the fifth image frame, then image frames in which the first display element has already been placed before the fifth image frame will remain unchanged. In other words, adjustments to the first display element in the fifth image frame only affect the first display element placed in subsequent image frames.
[0150] In some embodiments, if a first object appears in the sixth image frame of the first video, an adjusted first display element following the first object is displayed on the sixth image frame, which is an image frame following the fifth image frame among a plurality of image frames.
[0151] For example, in image frames following the fifth image frame, the adjusted first display element is controlled to follow the first object. For example, when playing the processed video, the first display element before adjustment to follow the first object is displayed on image frames before the fifth image frame. Correspondingly, the adjusted first display element to follow the first object is displayed on image frames before the fifth image frame.
[0152] For example, the adjustment operation on the adjusted first display element can also be performed again in subsequent image frames to obtain a second-adjusted first display element, which can then be applied to subsequent image frames.
[0153] The technical solution provided in this application allows adjustment of a first display element applied across multiple image frames, affecting subsequent image frames. Furthermore, it eliminates the need to re-determine the first object. Therefore, it achieves both real-time adjustment of display elements and ensures continuous following (following remains uninterrupted and does not require re-determination), demonstrating the versatility of video processing and improving video processing efficiency.
[0154] The following describes how to determine the display position of the first display element in the second image frame, including at least one of the following steps S1 to S3 (not shown in the figure).
[0155] Step S1: Obtain the position offset of the first display element relative to the first object in the first image frame.
[0156] For example, before step S1, the position of each object in at least one object is obtained in multiple image frames. For example, the position of each object in the first image frame is obtained using a subject position determination model (assuming the at least one object appears in the first image frame). For example, the position of each object in subsequent image frames is obtained using a subject position determination model. In one case, the subject position determination model identifies the position of each object in each subsequent image frame. In this case, the position of each object in each image frame can be directly and accurately obtained. In another case, the subject position determination model identifies the offset value of each object in each subsequent image frame relative to its position in the first image frame. In this case, for the i-th image frame, the position of the object in the i-th image frame can be obtained by superimposing the object's position in the first image frame with the offset value of the object's position in the i-th image frame determined by the subject position determination model, where i is a positive integer greater than 1. In another scenario, the subject position determination model identifies the offset value of each object in each subsequent image frame relative to its position in the previous image frame. In this case, for the i-th image frame, the object's position in the (i-1)-th image frame is obtained by superimposing the object's position in the (i-1)-th image frame with the offset value determined by the subject position determination model. Here, i is a positive integer greater than 1. Both of these latter scenarios can reduce the cost and improve the efficiency of position determination to some extent.
[0157] For example, after determining the position of each of at least one object in multiple image frames, a smoothing process is applied to the position of each object in the multiple image frames to obtain the smoothed position of the object in the multiple image frames. The position in the image frames mentioned in the following embodiments can be understood as the smoothed position. For example, the position of the object in the multiple image frames is smoothed using a smoothing algorithm.
[0158] For example, after determining the position of each object in multiple image frames, the motion trajectory of each object in the multiple image frames is obtained. For example, a smoothing algorithm is used to smooth the motion trajectory to obtain a smooth and natural trajectory result. Based on the smoothed trajectory result, the smoothed position of the object in the multiple image frames is obtained.
[0159] For example, since the first display element is placed on the first image frame, the position offset of the first display element relative to the first object in the first image frame is directly obtained. For example, the position vector pointing from the position of the first object to the position of the first display element in the two-dimensional plane is considered as the position offset. The position offset includes an offset distance and an offset direction, where the offset distance is the distance of the first display element relative to the first object, and the offset direction is the direction of the first display element relative to the first object.
[0160] Step S2: Obtain the position of the first object in the second image frame through the subject position determination model. The subject position determination model is a neural network model used to determine the position of the subject in the image frame.
[0161] In some embodiments, the subject location determination model is a neural network model for determining the position of a subject in an image frame. In some embodiments, the input of the subject location determination model is an image frame, and the output is the position of a first object in the image frame. The subject location determination model in the embodiments of this application, or at least one module or sub-model of the subject location determination model, is a pre-trained model.
[0162] Step S3: Determine the display position of the first display element in the second image frame based on the position offset and the position of the first object in the second image frame.
[0163] For example, the display position of the first display element in the second image frame is determined simultaneously by a position offset and the position of the first object in the second image frame. For example, when either the position offset or the position of the first object in the second image frame changes, the display position of the first display element in the second image frame also changes.
[0164] Furthermore, when the position of the first display element is adjusted by the user, the position offset is re-determined based on the adjusted position. Based on the re-determined position offset and the position of the first object in the second image frame, the display position of the first display element in the second image frame is determined.
[0165] The technical solution provided in this application determines the display position of the first display element in the second image frame by using the position offset on the first image frame and the position of the first object in the second image frame. This helps to improve the accuracy of the display position determination and reflects the video processing effect of the first display element following the first object.
[0166] In some embodiments, the position of the first object in the second image frame includes: the position of the main bounding box of the first object in the second image frame, wherein the main bounding box of the first object is determined by a main object position determination model, and the main bounding box of the first object is used to indicate the position of the first object.
[0167] The following two methods will be used to describe how to determine the display position of the first display element in the second image frame.
[0168] In the first case, the position offset superimposed in each image frame remains constant.
[0169] In some embodiments, the display position of the first display element in the second image frame is obtained by moving the position offset from the position of the main frame of the first object in the second image frame.
[0170] For example, with the position offset unchanged, the position of the main frame of the first object in the second image frame is taken as the starting point, and the position is moved by the distance indicated by the position offset (offset distance) along the direction indicated by the position offset (offset direction) to obtain the display position of the first display element in the second image frame.
[0171] For example, without considering the display status of the first object in the second image frame (display position, display area, etc.), the position offset is directly superimposed on the position of the main body frame of the first object in the second image frame to obtain the display position of the first display element in the second image frame. If the display position of the first display element in the second image frame is still within the second image frame, then the first display element is displayed at the corresponding position. If the display position of the first display element in the second image frame is outside the second image frame, then the first display element is not displayed at the corresponding position.
[0172] In this embodiment, the method of directly superimposing the position offset is beneficial to improving the efficiency of position determination and further improving the efficiency of video processing.
[0173] In the second scenario, the position offset superimposed on each image frame may be adjusted.
[0174] In some embodiments, the position offset is adjusted based on the size of the main body frame of the first object in the second image frame and the size of the main body frame of the first object in the first image frame to obtain the adjusted position offset.
[0175] For example, the dimensions of the main frame include at least one of the length, width, and area of the main frame.
[0176] For example, based on the size of the main bounding box of the first object in the second image frame and the size of the main bounding box of the first object in the first image frame, the change in the size of the main bounding box of the first object can be determined. If the size of the main bounding box of the first object in the second image frame is larger than the size of the main bounding box of the first object in the first image frame, it indicates that the first object appearing in the second image frame has become larger. If the size of the main bounding box of the first object in the second image frame is smaller than the size of the main bounding box of the first object in the first image frame, it indicates that the first object appearing in the second image frame has become smaller.
[0177] For example, if the first object appearing in the second image frame becomes larger, the position offset is increased to obtain the adjusted position offset. For example, if the first object appearing in the second image frame becomes smaller, the position offset is decreased to obtain the adjusted position offset. For example, the adjustment range of the position offset is positively correlated with the change in the size of the first object.
[0178] For example, the adjustment range of the position offset distance is positively correlated with the change in the size of the first object. For example, the offset direction of the position offset does not change.
[0179] In some embodiments, the display position of the first display element in the second image frame is obtained by moving the adjusted position offset, starting from the position of the main frame of the first object in the second image frame.
[0180] For example, starting from the position of the main frame of the first object in the second image frame, the first display element is moved a distance indicated by the adjusted position offset (adjusted offset distance) along the direction indicated by the adjusted position offset (adjusted offset direction) to obtain the display position of the first display element in the second image frame.
[0181] The technical solution provided in this application embodiment flexibly adjusts the position offset based on the size change of the first object, reflecting the diversity and flexibility of determining the position of the first display element, which not only enriches the element display method, but also enriches the video processing effect.
[0182] For example, the ratio of the size of the main body frame of the first object in the second image frame to the size of the main body frame of the first object in the first image frame is used as the first adjustment coefficient.
[0183] For example, the ratio of the area of the main body frame of the first object in the second image frame to the area of the main body frame of the first object in the first image frame is used as the first adjustment coefficient.
[0184] For example, the ratio between the value of the longer side of the main body frame of the first object in the second image frame and the value of the longer side of the main body frame of the first object in the first image frame is used as the first value. For example, the ratio between the value of the wider side of the main body frame of the first object in the second image frame and the value of the wider side of the main body frame of the first object in the first image frame is used as the second value. For example, the product of the first value and the second value is determined as the first adjustment coefficient.
[0185] For example, the position offset is adjusted by a first adjustment factor to obtain the adjusted position offset.
[0186] For example, the product of the first adjustment factor and the offset distance of the position offset is determined as the adjusted offset distance of the position offset. For example, the first adjustment factor does not adjust the offset direction of the position offset.
[0187] The technical solution provided in this application adjusts the position offset based on a first adjustment coefficient determined by the size change, which is beneficial for achieving precise adjustment and thus achieving a better display effect.
[0188] The following describes the display of the first display element in the image frame and the relationship between the object data of the first object in the second image frame.
[0189] In some embodiments, the display of the first display element displayed on the second image frame is related to the object data of the first object in the second image frame.
[0190] For example, the display status of the first display element displayed on the second image frame refers to the display status of the first display element on the second image frame, which includes at least one of the following: rotation angle, area, brightness, filter, etc.
[0191] For example, the object data of the first object in the second image frame refers to the display-related data of the first object on the second image frame, which includes at least one of the following: object orientation, object display area, object brightness, object filter, etc.
[0192] In some embodiments, when the display status of the first display element includes the rotation angle of the first display element and the object data of the first object includes the orientation of the first object in the second image frame, the straight line where the rotated first display element is located is perpendicular to the straight line where the orientation of the first object is located in the second image frame.
[0193] For example, the pose of the first object in the second image frame is analyzed to determine the orientation of the first object. For example, the rotation angle of the first display element refers to the rotation angle relative to the horizontal line on a two-dimensional plane.
[0194] For example, the straight line containing the first display element after rotation based on the rotation angle is perpendicular to the straight line containing the orientation of the first object in the second image frame. For example, when they are perpendicular, it can create a display effect where the first object is directly facing the first display element, thereby enriching the video processing methods.
[0195] In some embodiments, when the display status of the first display element includes the display area of the first display element, and the object data of the first object includes the display area of the first object, the display area of the first display element and the display area of the first object are positively correlated.
[0196] For example, the larger the display area of the first object, the larger the display area of the first display element, that is, the larger the first display element. For example, the smaller the display area of the first object, the smaller the display area of the first display element, that is, the smaller the first display element.
[0197] In some embodiments, in the second image frame, the brightness of the first display element is the same as the brightness of the first object. In some embodiments, in the second image frame, the filter of the first display element is the same as the filter of the first object.
[0198] The technical solution provided in this application determines the display status of the first display element displayed on the second image frame based on the object data of the first object in the second image frame. By closely combining the first display element displayed in the second image frame with the first object, unlike the monotonous and unchanging display elements in related technologies, this application achieves dynamic adjustment of the display element, enriching the video processing methods.
[0199] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0200] Please refer to Figure 7 This diagram illustrates a block diagram of a video processing apparatus according to an embodiment of this application. The apparatus has the function of implementing the video processing method described above; this function can be implemented in hardware or by hardware executing corresponding software. The apparatus can be the computer device described above, or it can be installed within a computer device. For example... Figure 7 As shown, the device 700 may include: a placement module 710 and a display module 720.
[0201] The placement module 710 is used to place a first display element on a first image frame of a first video, wherein the first image frame is one of a plurality of image frames included in the first video.
[0202] Display module 720 is used to display marking information corresponding to at least one object, wherein the at least one object includes objects that appear in the first image frame that can be followed by the first display element, and different marking information is used to distinguish different objects.
[0203] The display module 720 is further configured to, in response to an operation of selecting a first object among the at least one objects, display the first display element following the first object on the second image frame when the first object appears in the second image frame of the first video, wherein the second image frame is different from the first image frame.
[0204] In some embodiments, the display module 720 is configured to display at least one location marker on top of the first image frame, each of the at least one location marker being used to mark the position of one of the at least one objects in the first image frame.
[0205] In some embodiments, the operation of selecting the first object includes: selecting the location marker corresponding to the first object.
[0206] In some embodiments, the display module 720 is configured to, in response to an operation of selecting a position marker corresponding to the first object, distinguish and display the position marker corresponding to the first object and the position markers corresponding to other objects among the at least one object, respectively, on the upper layer of the first image frame.
[0207] In some embodiments, the display module 720 is configured to display object markers corresponding to the at least one object, wherein the object markers represent at least one of object name, object thumbnail, object number, and object type.
[0208] In some embodiments, selecting the first object includes selecting the object tag corresponding to the first object.
[0209] In some embodiments, the display module 720 is configured to, in response to an operation of selecting an object marker corresponding to the first object, mask other objects among the at least one objects besides the first object on top of the first image frame.
[0210] In some embodiments, the at least one object further includes a second object, which is an object that the first display element can follow in other image frames besides the first image frame.
[0211] In some embodiments, the display module 720 is configured to, in response to an operation of selecting an object tag corresponding to the second object, change the first image frame to a third image frame while retaining the display of the first display element; wherein the third image frame is an image frame in which the second object appears among the plurality of image frames, and the third image frame is different from the first image frame.
[0212] In some embodiments, such as Figure 8 As shown, the device 700 also includes a selection module 730 and a control module 740.
[0213] In some embodiments, the selection module 730 is configured to select a first timestamp interval from the timestamp intervals of the plurality of image frames in response to the operation of selecting the first object, wherein the first timestamp interval includes at least one image frame and the second image frame is an image frame in the first timestamp interval.
[0214] The control module 740 is used to control the first display element to follow the first object appearing in the at least one image frame.
[0215] In some embodiments, the control module 740 is further configured to control the first display element to appear in the first object in the plurality of image frames starting from the first image frame, respectively, wherein the second image frame is any one of the plurality of image frames starting from the first image frame.
[0216] In some embodiments, the control module 740 is further configured to, in response to a cancel follow operation for a fourth image frame, control the first display element to stop following the first object appearing in the first video starting from the fourth image frame, the fourth image frame being an image frame following the first image frame among the plurality of image frames.
[0217] In some embodiments, such as Figure 8 As shown, the device 700 also includes an object determination module 750.
[0218] In some embodiments, the object determination module 750 is configured to identify the first image frame through a subject detection model, obtain at least one subject in the first image frame and subject recognition scores corresponding to the at least one subject, wherein the subject detection model is a neural network model for identifying subjects in the image frame, and the subject recognition score is used to characterize the confidence level of the subject identified by the subject detection model; and retain subjects whose subject recognition scores are greater than or equal to a score threshold from the at least one subject as the at least one object.
[0219] In some embodiments, such as Figure 8 As shown, the device 700 also includes a receiving module 760.
[0220] In some embodiments, the receiving module 760 is configured to receive an operation on a first follow control on the editing interface of the first video; or, to receive an operation on a second follow control in the editing bar of the first display element; wherein the first follow control and the second follow control are different.
[0221] In some embodiments, the display module 720 is configured to display an adjusted first display element in response to an adjustment operation on the first display element displayed on a fifth image frame, wherein the fifth image frame is the second image frame or another image frame among the plurality of image frames that displays the first display element besides the second image frame; and, if the first object appears in a sixth image frame of the first video, to display an adjusted first display element following the first object on the sixth image frame, wherein the sixth image frame is an image frame among the plurality of image frames that follows the fifth image frame.
[0222] In some embodiments, the device 700 further includes a position determination module (not shown in the figure).
[0223] In some embodiments, the position determination module is configured to obtain the position offset of the first display element relative to the first object in the first image frame; obtain the position of the first object in the second image frame through a subject position determination model, wherein the subject position determination model is a neural network model for determining the position of a subject in an image frame; and determine the display position of the first display element in the second image frame based on the position offset and the position of the first object in the second image frame.
[0224] In some embodiments, the position of the first object in the second image frame includes: the position of the main bounding box of the first object in the second image frame, wherein the main bounding box of the first object is determined by the main bounding box position determination model, and the main bounding box of the first object is used to indicate the position of the first object.
[0225] In some embodiments, the position determination module is configured to move the position offset starting from the position of the main body frame of the first object in the second image frame to obtain the display position of the first display element in the second image frame; or, based on the size of the main body frame of the first object in the second image frame and the size of the main body frame of the first object in the first image frame, adjust the position offset to obtain the adjusted position offset; and move the adjusted position offset starting from the position of the main body frame of the first object in the second image frame to obtain the display position of the first display element in the second image frame.
[0226] In some embodiments, the position determination module is used to take the ratio of the size of the main body frame of the first object in the second image frame to the size of the main body frame of the first object in the first image frame as a first adjustment coefficient; and adjust the position offset by the first adjustment coefficient to obtain the adjusted position offset.
[0227] In some embodiments, the display status of the first display element displayed on the second image frame is related to the object data of the first object in the second image frame; when the display status of the first display element includes the rotation angle of the first display element and the object data of the first object includes the orientation of the first object in the second image frame, the straight line where the rotated first display element is located is perpendicular to the straight line where the orientation of the first object is located in the second image frame; when the display status of the first display element includes the display area of the first display element and the object data of the first object includes the display area of the first object, the display area of the first display element is positively correlated with the display area of the first object.
[0228] The technical solution provided in this application distinguishes different objects in an image frame by displaying marking information corresponding to at least one object, allowing the user to select different objects. After selecting the first object, there is no need to manually add the first display element again in the second image frame; instead, the first display element of the first object following the first object in the second image frame is directly displayed, simplifying the video processing steps and thus improving video processing efficiency.
[0229] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0230] Please refer to Figure 9 This diagram illustrates a structural block diagram of a computer device 900 provided in one embodiment of this application. The computer device 900 can be any electronic device capable of data calculation, processing, and storage. The computer device 900 can be used to implement the video processing method provided in the above embodiments.
[0231] Typically, computer device 900 includes a processor 901 and a memory 902.
[0232] Processor 901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 901 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 901 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 901 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 901 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0233] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 902 are used to store a computer program configured to be executed by one or more processors to implement the video processing method described above.
[0234] Those skilled in the art will understand that Figure 9 The structure shown does not constitute a limitation on the computer device 900, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0235] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored in the storage medium, and the computer program, when executed by a processor, implements the above-described video processing method. Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0236] In an exemplary embodiment, a computer program product is also provided, comprising a computer program stored in a computer-readable storage medium. A processor of a terminal device reads the computer program from the computer-readable storage medium and executes the computer program, causing the terminal device to perform the aforementioned video processing method.
[0237] It should be noted that the collection and processing of relevant data (including multiple image frames in the first video, the operation of selecting the first object, following, etc.) in this application should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0238] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0239] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method of video processing, the method comprising: The method includes: A first display element is placed on the first image frame of the first video, wherein the first image frame is one of a plurality of image frames included in the first video; Display at least one object corresponding to each object, wherein the at least one object includes objects that appear in the first image frame that can be followed by the first display element, and different marking information is used to distinguish different objects; In response to the operation of selecting a first object among the at least one objects, if the first object appears in a second image frame of the first video, the first display element following the first object is displayed on the second image frame, which is different from the first image frame.
2. The method of claim 1, wherein, The display of the tag information corresponding to at least one object includes: Above the first image frame, at least one location marker is displayed, each of the at least one location markers being used to mark the position of one of the at least one objects in the first image frame.
3. The method of claim 2, wherein, The operation of selecting the first object includes: selecting the position marker corresponding to the first object, and the method further includes: In response to the operation of selecting the position marker corresponding to the first object, the position marker corresponding to the first object and the position markers corresponding to the other objects among the at least one object, excluding the first object, are displayed separately on the upper layer of the first image frame.
4. The method of claim 1, wherein, The display of the tag information corresponding to at least one object includes: Display object tags corresponding to the at least one object, wherein the object tags represent at least one of the following: object name, object thumbnail, object number, and object type.
5. The method of claim 4, wherein, The operation of selecting the first object includes: selecting the object marker corresponding to the first object, and the method further includes: In response to the operation of selecting the object tag corresponding to the first object, other objects among the at least one object, except for the first object, are masked on the upper layer of the first image frame.
6. The method according to claim 4 or 5, characterized in that, The at least one object further includes a second object, which is an object that the first display element can follow in other image frames besides the first image frame; The method further includes: In response to the operation of selecting the object tag corresponding to the second object, the first image frame to be displayed is changed to the third image frame, while the first display element is retained; The third image frame is the image frame in which the second object appears among the plurality of image frames, and the third image frame is different from the first image frame.
7. The method according to any one of claims 1 to 6, characterized in that, Before displaying the first display element following the first object on the second image frame, the method further includes: In response to the operation of selecting the first object, a first timestamp interval is selected from the timestamp intervals of the plurality of image frames, the first timestamp interval including at least one image frame, and the second image frame is an image frame in the first timestamp interval; The first display element is controlled to follow the first object appearing in the at least one image frame.
8. The method according to any one of claims 1 to 6, characterized in that, The method further includes: The first display element is controlled to appear in the first image frame and subsequent image frames in the plurality of image frames, respectively, where the second image frame is any image frame that appears after the first image frame in the plurality of image frames.
9. The method according to claim 8, characterized in that, The method further includes: In response to a cancel follow operation for a fourth image frame, the first display element is controlled to stop following the first object appearing in the first video starting from the fourth image frame, the fourth image frame being an image frame following the first image frame among the plurality of image frames.
10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: The first image frame is identified by a subject detection model, and at least one subject in the first image frame and a subject recognition score corresponding to the at least one subject are obtained. The subject detection model is a neural network model for identifying subjects in an image frame, and the subject recognition score is used to characterize the confidence level of the subject identified by the subject detection model. The subjects whose subject identification scores are greater than or equal to a score threshold are retained from the at least one subject as the at least one object.
11. The method according to any one of claims 1 to 10, characterized in that, Before displaying the tag information corresponding to at least one object, the method further includes: Receive operations on the first follow control in the editing interface of the first video; or, Receive operations on the second follow control in the edit bar of the first displayed element; The first follow control and the second follow control are different.
12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: In response to an adjustment operation on the first display element displayed on the fifth image frame, the adjusted first display element is displayed, wherein the fifth image frame is the second image frame or another image frame among the plurality of image frames other than the second image frame that displays the first display element; If the first object appears in the sixth image frame of the first video, the first display element, adjusted to follow the first object, is displayed on the sixth image frame, which is the image frame following the fifth image frame among the plurality of image frames.
13. The method according to any one of claims 1 to 12, characterized in that, The method further includes: Obtain the position offset of the first display element relative to the first object in the first image frame; The position of the first object in the second image frame is obtained by a subject position determination model, wherein the subject position determination model is a neural network model used to determine the position of the subject in the image frame; The display position of the first display element in the second image frame is determined based on the position offset and the position of the first object in the second image frame.
14. The method according to claim 13, characterized in that, The position of the first object in the second image frame includes: the position of the main bounding box of the first object in the second image frame, wherein the main bounding box of the first object is determined by the main bounding box position determination model, and the main bounding box of the first object is used to indicate the position of the first object; Determining the display position of the first display element in the second image frame based on the position offset and the position of the first object in the second image frame includes: Starting from the position of the main bounding box of the first object in the second image frame, move the position offset to obtain the display position of the first display element in the second image frame; or... Based on the size of the main frame of the first object in the second image frame and the size of the main frame of the first object in the first image frame, the position offset is adjusted to obtain the adjusted position offset; taking the position of the main frame of the first object in the second image frame as the starting point, the adjusted position offset is moved to obtain the display position of the first display element in the second image frame.
15. The method according to claim 14, characterized in that, The step of adjusting the position offset based on the size of the main bounding box of the first object in the second image frame and the size of the main bounding box of the first object in the first image frame to obtain the adjusted position offset includes: The ratio of the size of the main body frame of the first object in the second image frame to the size of the main body frame of the first object in the first image frame is used as the first adjustment coefficient; The position offset is adjusted by the first adjustment coefficient to obtain the adjusted position offset.
16. The method according to any one of claims 1 to 15, characterized in that, The display of the first display element on the second image frame is related to the object data of the first object in the second image frame; In the case where the display status of the first display element includes the rotation angle of the first display element, and the object data of the first object includes the orientation of the first object in the second image frame, the straight line where the rotated first display element is located is perpendicular to the straight line where the orientation of the first object is located in the second image frame; When the display status of the first display element includes the display area of the first display element, and the object data of the first object includes the display area of the first object, the display area of the first display element and the display area of the first object are positively correlated.
17. A video processing apparatus, characterized in that, The device includes: A placement module is used to place a first display element on a first image frame of a first video, wherein the first image frame is one of a plurality of image frames included in the first video; The display module is used to display the marking information corresponding to at least one object, wherein the at least one object includes objects that appear in the first image frame that can be followed by the first display element, and different marking information is used to distinguish different objects; The display module is further configured to, in response to an operation of selecting a first object among the at least one objects, display the first display element following the first object on the second image frame when the first object appears in the second image frame of the first video, wherein the second image frame is different from the first image frame.
18. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the video processing method as described in any one of claims 1 to 16.
19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the video processing method as described in any one of claims 1 to 16.
20. A computer program product, characterized in that, The computer program product includes a computer program that is loaded and executed by a processor to implement the video processing method as described in any one of claims 1 to 16.