Video processing method and apparatus, storage medium and electronic device

By intercepting and processing dynamic pictures and audio data during video playback, the problem of users needing to edit videos to obtain materials is solved, and the material acquisition is simplified and interactive fun is improved.

WO2025156777A1PCT designated stage expired Publication Date: 2025-07-31BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/131367
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-23
Filing Date
2024-11-11
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

In the prior art, users need to download videos from the Internet and edit them to obtain materials, requiring users to have certain video editing capabilities, and the process is poorly fun.

Method used

By inputting a first trigger operation during video playback, the dynamic map and/or audio data of the target video are intercepted as object dynamic data, and processing is performed through the second trigger operation, the material acquisition process is simplified.

Benefits of technology

It simplifies the material editing process, improves the fun and efficiency of user interaction, and allows users to add object dynamic data as material to videos more easily.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024131367_31072025_PF_FP_ABST
    Figure CN2024131367_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a video processing method and apparatus, a storage medium and an electronic device. The method comprises: displaying a video playback page, the video playback page being used for playing video stream data; during the playback of a target video, in response to a first triggering operation on the video playback page, determining a target object, and obtaining object dynamic data corresponding to the target object, the object dynamic data being captured from a plurality of video frames of the target video; and in response to a second triggering operation on the video playback page, performing a processing operation on the object dynamic data.
Need to check novelty before this filing date? Find Prior Art

Description

Video processing method, device, storage medium and electronic equipment

[0001] This application claims priority to the Chinese patent application filed on January 23, 2024, with application number 202410095843.8 and invention name “A video processing method, device, storage medium and electronic device”. The entire contents of that application are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to video processing technology, and more particularly to a video processing method, device, storage medium, and electronic device. Background Art

[0003] With the continuous development of Internet technology, recording life through videos (especially short videos) has become a common way. The process of video generation requires users to have a certain amount of material. Generally, to obtain the material from online videos, users need to download the video from the Internet and edit the downloaded video to obtain the required material.

[0004] Summary of the Invention

[0005] The present disclosure provides a video processing method, device, storage medium and electronic device.

[0006] In a first aspect, an embodiment of the present disclosure provides a video processing method, including:

[0007] Displaying a video playback page, wherein the video playback page is used to play video stream data;

[0008] During playback of a target video, in response to a first triggering operation on the video playback page, determining a target object and obtaining object dynamic data corresponding to the target object, the object dynamic data being captured from a plurality of video frames of the target video;

[0009] In response to a second triggering operation on the video playback page, a processing operation on the object dynamic data is performed.

[0010] In a second aspect, an embodiment of the present disclosure further provides a video processing device, including:

[0011] A page display module is used to display a video playback page, wherein the video playback page is used to play video stream data;

[0012] a data interception module for determining a target object and obtaining object dynamic data corresponding to the target object in response to a first trigger operation on the video playback page during playback of the target video, the object dynamic data being intercepted from multiple video frames of the target video;

[0013] The data processing module is used to execute a processing operation on the object dynamic data in response to a second trigger operation on the video playback page.

[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:

[0015] one or more processors;

[0016] a storage device for storing one or more programs,

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method provided in any embodiment of the present disclosure.

[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to execute the video processing method provided in any embodiment of the present disclosure.

[0019] The technical solution provided by the embodiments of the present disclosure captures the dynamic image and / or audio data of a target video to obtain object dynamic data by inputting a first trigger operation during video playback. Furthermore, by inputting a second trigger operation, processing operations are performed on the object dynamic data. During subsequent video production, the object dynamic data can be added as material to the produced video. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.

[0021] FIG1 is a flow chart of a video processing method provided by an embodiment of the present disclosure;

[0022] FIG2 is a schematic diagram of a video playback page in a cutout mode provided by an embodiment of the present disclosure;

[0023] FIG3 is a schematic diagram of the process of intercepting and processing dynamic data of an object provided by an embodiment of the present disclosure;

[0024] FIG4 is a flow chart of a video processing method provided by an embodiment of the present disclosure;

[0025] FIG5 is a schematic diagram of the process of intercepting and processing dynamic data of an object provided by an embodiment of the present disclosure;

[0026] FIG6 is a flow chart of a video processing method provided by an embodiment of the present disclosure;

[0027] FIG7 is a schematic diagram of processing a video to be edited in a collage mode according to an embodiment of the present disclosure;

[0028] FIG8 is a schematic diagram of processing a video to be edited in a replacement mode provided by an embodiment of the present disclosure;

[0029] FIG9 is a schematic structural diagram of a video processing device provided by an embodiment of the present disclosure;

[0030] FIG10 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0032] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0033] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0034] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0035] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0036] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0037] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0038] Generally, for the video material on the Internet, users need to download the video from the Internet and edit the downloaded video to obtain the required material. The above-mentioned material acquisition method requires users to have certain video editing skills, which is demanding on users and not interesting.

[0039] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0040] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0041] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0042] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0043] Figure 1 is a flow chart of a video processing method provided by an embodiment of the present disclosure. The embodiment of the present disclosure is applicable to the situation where video material is accessed from a video being displayed. The method can be executed by a video processing device, which can be implemented in the form of software and / or hardware. Optionally, it can be implemented by an electronic device, which can be a mobile terminal such as a mobile phone, tablet computer, a PC or a server.

[0044] As shown in FIG1 , the method includes:

[0045] S110 . Display a video playback page, where the video playback page is used to play video stream data.

[0046] S120. During the playback of the target video, in response to a first trigger operation on the video playback page, a target object is determined, and object dynamic data corresponding to the target object is obtained, where the object dynamic data is captured from multiple video frames of the target video.

[0047] S130: In response to a second triggering operation on the video playback page, executing a processing operation on the object dynamic data.

[0048] The electronic device has a video playback application installed. Upon launching the video playback application, a video playback page of the video playback application is displayed, where video stream data is played. The videos played on the video playback page may include, but are not limited to, short videos, regular videos, live videos, etc., and may be determined based on the video types supported by the video playback application.

[0049] On the video playback page, the video being played can be switched according to the user's interactive operations on the video playback page. For example, the video can be switched by sliding operations; or, the video can be switched by triggering operations on tags displayed on the video playback page, where the tags in the video playback page can be video attribute tags such as "Same City" and "Follow", or they can also be type tags such as "Live Broadcast" and "Short Video".

[0050] Play the target video on the video playback page. The target video can be any video currently playing on the video playback page. The target video can include at least one object, which can include but is not limited to people, animals, plants, or any objects that appear in the video.

[0051] During the playback of the target video, a first trigger operation input by the user on the video playback page is detected. When the first trigger operation is detected, the cutout mode is entered in response to the first trigger operation. During the playback of the target video, the selected object (i.e., the target object) in the target video is cutout to obtain the object dynamic data of the target object in the target video. The object dynamic data can be used as new video material for new video editing.

[0052] The first trigger operation is an operation that triggers the cutout mode and can be pre-set according to user needs. The detected input operation is matched with the pre-set first trigger operation. If the match is successful, it is determined that the first trigger operation is detected. The specific form of the first trigger operation is not limited here.

[0053] Optionally, the object dynamic data may be an object dynamic graph formed by capturing target object graphs from multiple video frames including the target video. Based on the above embodiment, the first trigger operation is further used to select the target object. Specifically, the target object is determined based on the position at which the first trigger operation is applied. Exemplarily, the first trigger operation may be a long press operation or a double-click operation, and the target object is determined based on the press position of the long press operation or the click position of the double-click operation. The position at which the first trigger operation is applied may be the coordinate position of the first trigger operation in the video frame.

[0054] In some embodiments, upon detecting a first trigger operation, image segmentation is performed on the first video frame to which the first trigger operation is applied to obtain a plurality of segmented objects. Specifically, the video frame may be input into an image segmentation model to obtain an image segmentation result output by the image segmentation model, wherein the image segmentation result includes at least one object contour in the video frame. A target object is determined based on the location where the first trigger operation was applied and the at least one object contour. Specifically, an object contour including the location where the first trigger operation was applied is determined, and the object to which the object contour belongs is determined as the target object.

[0055] In some embodiments, the first video frame in which the first trigger operation is applied and the location at which the first trigger operation is applied may be input into an image segmentation model to obtain a target object segmentation result output by the image segmentation model. The target object segmentation result may be a mask image of the target object, and the segmented target object is determined by combining the mask image of the target object with the original video. For example, pixels in the mask image that belong to the target object are marked as 1, and pixels that do not belong to the target object are marked as 0. The target object mask image is multiplied by the corresponding value in the original video to obtain the segmented target object.

[0056] After determining the target object, the target object is outlined on the video playback page to highlight the target object. This allows the user to intuitively determine the selected object on the video playback page. If the target object is not selected accurately, the target object can be reselected to avoid incorrect selection of the target object. Optionally, the edge pixels of the target object are enhanced, for example, by increasing the brightness value of the edge pixels of the target object; optionally, the edge pixels of the target object can be set to a set color value.

[0057] When the target object is determined, multiple video frames in the target video are cut out while the target video is being played, to obtain target object images corresponding to the multiple video frames. Accordingly, the contours of the target objects in the multiple video frames are marked during the playback of the target video.

[0058] Specifically, in the cutout mode, a target object is cutout processed for multiple video frames in the target video, and the target object images obtained by cutout in the multiple video frames are sequentially combined to obtain a dynamic image of the target object, i.e., object dynamic data. Specifically, a cutout processing module can be called, and cutout processing can be performed on multiple video frames in the target video based on the cutout processing module, wherein the cutout processing module can include a cutout algorithm, and the cutout algorithm can be a neural network model (i.e., a cutout model) to achieve end-to-end cutout processing, that is, multiple video frames of the target video are input into the above-mentioned cutout model, and the target object image corresponding to each video frame is output.

[0059] In some embodiments, the target object is segmented for the first video frame to which the first trigger operation is applied, and a segmentation result corresponding to the first video frame is obtained, that is, a mask map of the target object. The mask map of the target object and the video frames to be segmented (that is, multiple video frames of the target object) are input into the cutout model to obtain mask maps corresponding to the multiple video frames output by the cutout model. The target object map corresponding to the video frame is determined based on the mask map of the video frame and the original video frame.

[0060] Each target object graph may be configured with a timestamp, and the timestamp of the target object graph is consistent with the timestamp of the corresponding video frame. Multiple target object graphs are sorted and combined according to the timestamps of the target object graphs to obtain a dynamic graph.

[0061] Optionally, the object dynamic data may be an object dynamic graph formed by capturing target object images from multiple video frames of the target video and audio data corresponding to the multiple video frames. Under the cutout model, the cutout processing and audio data capture processing are performed in parallel. The target object is cutout from multiple video frames in the target video to obtain a dynamic graph. Simultaneously, the audio data within the time period of the multiple video frames is captured to obtain audio data. The time range of the dynamic graph and the audio data are consistent. The dynamic graph and the audio data are then merged to obtain the object dynamic data.

[0062] The dynamic image and audio data can be stored in association, and the dynamic image and audio data are independent of each other and can be used as video editing materials. When the user calls the object dynamic data, a selection page of dynamic images, audio data and audio and video data (the confluence of dynamic images and audio data) can be displayed. In response to the selection operation on the selection page, the corresponding dynamic image, audio data and any one of the audio and video data are called. Among them, when audio and video data is selected, it can be obtained by confluence processing of the dynamic image and audio data. By providing options for dynamic images, audio data and audio and video data, it is convenient for users to meet different needs and improve the flexibility of editing the above materials during the video editing process.

[0063] Optionally, the object dynamic data may include audio data corresponding to a plurality of video frames. In the case that the target object is not selected by the first trigger operation, audio data corresponding to a plurality of video frames in the target video are obtained as the object dynamic data.

[0064] In some embodiments, the first trigger operation may include multiple different gesture operations. For example, when the first gesture operation is detected, the target object is determined, and the target object is cut out for multiple video frames in the target video to obtain target object graphs corresponding to the multiple video frames, and the dynamic graph formed by the multiple object dynamic data is determined as the object dynamic data. When the second gesture operation is detected, the audio data corresponding to the multiple video frames in the target video is used as the object dynamic data. When the third gesture operation is detected, the cutout processing and the audio data interception processing are performed in parallel, and the object dynamic graph formed by the target object graph and the audio data are merged to obtain the object dynamic data. The above-mentioned first gesture operation, second gesture operation and third gesture operation can be different gesture operations and are not specifically limited here.

[0065] In the above embodiment, the multiple video frames for intercepting the target object and / or audio data can be partial video frames or all video frames of the target video. For example, when a first trigger operation is input during the playback of the target video, the multiple video frames are the first video frame applied by the first trigger operation to the last video frame of the target video frame, wherein the first video frame applied by the first trigger operation is the initial application moment of the first trigger operation, and the video frame played by the target video. For example, the first trigger operation includes a start operation and an end operation, and the multiple video frames are the video frames between the start operation application moment and the end operation application moment during the playback of the target video. For example, the duration of the multiple video frames is the application duration of the first trigger operation, that is, the multiple video frames are the video frames between the start application moment and the end application moment of the first trigger operation during the playback of the target video. Exemplarily, the first trigger operation is a long press operation, and the multiple video frames are the multiple video frames played by the target video during the application time period of the long press operation.

[0066] Based on the above embodiment, during the capture of object dynamic data, the duration of the object dynamic data is displayed on the video playback page. The duration of the object dynamic data changes in real time as the capture progresses. For example, see Figure 2, which is a schematic diagram of a video playback page in cutout mode provided by an embodiment of the present disclosure. By displaying the duration of the object dynamic data, it is easier for the user to determine whether to terminate the capture process, thereby improving the accuracy of the duration of the object dynamic data.

[0067] After obtaining the object dynamic data, the object dynamic data is processed through the second trigger operation input by the user, wherein the processing of the object dynamic data includes but is not limited to downloading, deleting and applying, etc. In some embodiments, the second trigger operation can be a different gesture operation, and different second trigger operations can correspond to different processing operations. For example, the upward swipe operation on the video playback page corresponds to the deletion processing of the object dynamic data, the downward swipe operation on the video playback page corresponds to the download processing of the object dynamic data, and the right swipe operation on the video playback page corresponds to the application operation of the object dynamic data, that is, entering the video production page, and displaying the object dynamic data on the video production page to realize the production of a new video. The video production page can be a video shooting page, in which the video screen is captured by the camera, and a new video is produced by the object dynamic data and the screen captured by the camera; or, the video production page is a video editing page, in which the video to be edited is displayed, and a new video is produced by the video to be edited and the object dynamic data.

[0068] In some embodiments, when a first trigger operation is detected, the video playback page is switched to a cutout mode, in which the video playback page includes at least one dynamic data processing control, and different dynamic data processing controls correspond to different processing operations. Correspondingly, the second trigger operation is a trigger operation of any of the dynamic data processing controls. Referring to Figure 2, a plurality of dynamic data processing controls are provided at the bottom of the video playback page in Figure 2, including but not limited to download controls, new video shooting controls, collection controls, etc. The type, display form, and display position of the dynamic data processing controls in Figure 2 are only an example. In other embodiments, they can be set according to requirements. A plurality of dynamic data processing controls are provided at the bottom of the video playback page in Figure 2, which are download controls, new video shooting controls, and collection controls, respectively.

[0069] If a download control is triggered, the object's dynamic data is downloaded locally and added to the team's material library for easy access during subsequent video generation. If a favorite control is triggered, the object's dynamic data is added to the collection. If a new video capture control is triggered, the video capture page is switched to and the camera is started. The object's dynamic data is displayed on the video capture page, and a new video is generated based on the object's dynamic data and the camera's captured images.

[0070] In some embodiments, the second trigger operation may be a click operation on the above-mentioned dynamic data processing control, or a drag operation on the object dynamic data, and the release position of the drag operation is the position where the dynamic data processing control is located. For example, refer to Figure 3, which is a schematic diagram of the interception and processing process of the object dynamic data provided by the embodiment of the present disclosure. During the playback of the target video, a long press operation (i.e., the first trigger operation) is applied in the video playback page, and the long press operation is applied to an object in the video frame of the target video, and the target object corresponding to the long press operation is identified. During the playback of the target video, data is intercepted for multiple video frames played during the long press operation application process, that is, the target object cutout processing and the audio data corresponding to the multiple video frames are performed on the above-mentioned multiple video frames to obtain the object dynamic data corresponding to the long press operation application process. The user performs a drag operation (i.e., the second trigger operation) on the video playback page. During the drag operation, a target object diagram of the object's dynamic data (e.g., the target object diagram of the last frame) is displayed at the drag position. Multiple dynamic data processing controls are displayed on the video playback page. When the release position of the drag operation is any dynamic data processing control, the dynamic data processing control is triggered, and the processing operation corresponding to the triggered dynamic data processing control is performed on the object's dynamic data, such as downloading, collecting, or shooting a new video.

[0071] The technical solution of the disclosed embodiment captures the target video's dynamic image and / or audio data by inputting a first trigger operation during video playback to obtain object dynamic data. Furthermore, by inputting a second trigger operation, processing operations are performed on the object dynamic data. During subsequent video production, the object dynamic data can be added as material to the produced video, simplifying the material editing process and enabling interactive editing, thereby enhancing the interactive experience.

[0072] FIG4 is a flow chart of a video processing method provided by an embodiment of the present disclosure. Based on the above embodiment, the processing operation of the object dynamic data is refined. As shown in FIG4, the method includes:

[0073] S210: Display a video playback page, where the video playback page is used to play video stream data.

[0074] S220. During the playback of the target video, in response to a first trigger operation on the video playback page, a target object is determined, and object dynamic data corresponding to the target object is obtained, where the object dynamic data is captured from multiple video frames of the target video.

[0075] S230. In response to a second trigger operation on the video playback page, the video playback page is switched to a video shooting page. During the video shooting process, the object dynamic image is displayed on the video shooting page, and / or the audio data in the object dynamic data is added to the video shooting page to obtain a shot video.

[0076] In this embodiment, the second trigger operation is a trigger operation of a new video capture control. In response to this second trigger operation, a new video is captured based on the object dynamic data captured from the target video. For example, see Figure 5, which is a schematic diagram of the object dynamic data capture and processing process provided by the embodiment of the present disclosure.

[0077] When the video capture page is switched to, the electronic device's camera is activated, and the video captured by the camera is displayed on the video capture page. The video capture page includes a capture control, and when the capture control is triggered, video capture is performed. During the video capture process, the captured image captured by the camera is combined with the object's dynamic data to generate a new captured video.

[0078] The object dynamic data includes object dynamic images and / or audio data. Before the shooting control is triggered, the object dynamic images and the preview screen captured by the camera will be displayed on the video shooting page, and / or the audio data will be played to facilitate preview before video shooting.

[0079] After switching the video playback page to the video shooting page and before shooting a new video, it also includes: displaying an object dynamic image in the video shooting page, and adjusting the display status of the object dynamic image in the video shooting page in response to an adjustment operation on the object dynamic image, wherein the display status includes one or more of display position, display size, and display mode.

[0080] The adjustment operation on the object dynamic image may include a drag operation, and the display position of the object dynamic image is adjusted by dragging the object dynamic image, and the release position of the drag operation is used as the target position of the object dynamic image.

[0081] The adjustment operation on the object dynamic image may include a zooming operation on the object dynamic image, and in response to the zooming operation, the display size of the object dynamic image is adjusted, wherein the zooming operation includes a zooming-in operation and a zooming-out operation.

[0082] The adjustment operation of the object dynamic image may include a mode setting operation, wherein the display mode of the object dynamic image includes loop display and single display, etc. The single display mode is used to control the object dynamic image to be displayed once during the video shooting process, and stop displaying after all target object images of the object dynamic image are displayed; the loop display is used to control the object dynamic image to be displayed in a loop during the video shooting process, and re-display it after all target object images of the object dynamic image are displayed. During the video shooting process, multiple frames of target object images in the object dynamic image are displayed in sequence. When the duration of the object dynamic image is less than the shooting duration, the display mode of the object dynamic image is set by loop display and single display.

[0083] In some embodiments, before shooting a new video, the process may also include calling new object dynamic data from the material library and displaying it on the video shooting page. For example, referring to the right figure in FIG5 , FIG5 shows a material library control provided on the video shooting page. In response to a triggering operation on the material library control, a material library display page is displayed. The material library display page displays multiple stored materials, i.e., object dynamic data. In response to a selection operation on at least one object dynamic data on the material library display page, the selected at least one object dynamic data is displayed on the video shooting page. It is understood that the object dynamic image on the video shooting page can also be deleted.

[0084] In some embodiments, the video capture page also displays audio data information in the object's dynamic data. The audio data information may be the name of a song. As shown in the right figure of FIG5 , the video capture page includes an audio data information display control. Optionally, in response to a triggering operation on the audio data information display control, an audio data display page is displayed. The audio data display page includes selectable audio data information. In response to a selection operation on any audio data information, the audio data information in the video capture page is replaced, as well as the audio data during the video capture process.

[0085] The black solid circle in the right image of Figure 5 represents a capture control. In response to a triggering operation on the capture control, a new video is captured. During the video capture process, the video image captured by the camera is displayed on the video capture page, and the object dynamic image is displayed on the video capture page. The frame of the video image captured by the camera is aligned with the target object image in the object dynamic image, and the frame and the target object image correspond one-to-one, that is, the i-th frame of the frame corresponds to the i-th frame of the target object image. The corresponding frame and target object image are spliced ​​into a new video frame. Multiple new video frames form a captured video image sequence, and the captured video image sequence is merged with the audio data to form the captured video. Optionally, after the image frame and the target object image are spliced, the new video frame can also be re-rendered to unify the light and shadow effects in the new video frame. Specifically, the light and shadow features of the frame can be extracted and the new video frame is re-rendered based on the light and shadow features. For example, the portrait of the person in the right image of Figure 5 can represent the video image captured by the camera, and the simple character can represent a target object image in the object dynamic image.

[0086] For example, if an object dynamic graph includes n frames of target object graphs, and if i>n, in single display mode, the i-th frame is used as a new video frame. In loop display mode, the i-th frame is concatenated with the ik*n frames of the target object graph to form a new video frame, where k is the number of loops of the object dynamic graph, k≥0. In the case of the first display, k=0. The display model of the audio data and the object dynamic graph can be the same, and can be looped or played once.

[0087] The technical solution provided by the disclosed embodiments captures object dynamic data from a target video displayed in a video playback application. After switching to the video capture page, the object dynamic data is displayed on the video capture page as the material for a new video. The object dynamic data and the video footage captured by the camera are combined to generate a new captured video. This simplifies the editing of material and the production of new videos, and enables interactive editing of material, making the interaction more interesting.

[0088] FIG6 is a flow chart of a video processing method provided by an embodiment of the present disclosure. Based on the above embodiment, an application method of object dynamic data and a video production method are provided. Referring to FIG6, the method specifically includes:

[0089] S310: Display a video playback page, where the video playback page is used to play video stream data.

[0090] S320. During the playback of the target video, in response to a first trigger operation on the video playback page, a target object is determined, and object dynamic data corresponding to the target object is obtained, where the object dynamic data is captured from multiple video frames of the target video.

[0091] S330: In response to a second triggering operation on the video playback page, store the object dynamic data.

[0092] S340: Display a video editing page, where the video editing page is used to display the video to be edited. The video editing page includes at least one video editing mode control, and each video editing mode control corresponds to a video editing mode.

[0093] S350. In response to a selection operation on the video editing mode, an edited video in the video editing mode is displayed, where the edited video is obtained by editing the video to be edited in the video editing mode based on the object dynamic data.

[0094] In this embodiment, when object dynamic data is captured in a target video, the object dynamic data can be stored, for example, by downloading or adding it to favorites for easy access during subsequent video production. This can be achieved by triggering a download control or a favorites control on the video playback page (right image in FIG3 ). The object dynamic data can be stored in a material library, which is accessed during video capture, video editing, and other processes.

[0095] In some embodiments, a page switching operation may be input on the video playback page, and in response to the page switching operation, a video editing page may be displayed. The video editing page may include a video selection control, and in response to a triggering operation on the video selection control, a video to be edited is displayed. In response to a selection operation on any video to be edited, the video to be edited is displayed on the video editing page.

[0096] In some embodiments, it may be in a video shooting page such as that in Figure 5, the video shooting page includes an album control, and in response to the triggering operation of the album control, the album display page is limited, the album display page includes captured images or captured videos (the captured videos can be used as videos to be edited), and in response to the selection operation of any captured video, the video editing page is displayed, and the selected captured video is displayed in the video editing page, and the selected captured video can be used as the video to be edited.

[0097] In some embodiments, after exiting the video playback application, in response to a start-up operation of a video editing application, a video editing page may be displayed. The video editing application may call a material library and a video library, and display the selected object dynamic data and the selected captured video on the video editing page.

[0098] In the above embodiment, the video to be edited can be generated based on at least one static image, and the static image is copied by setting the display duration of the static image, and the multiple copied static images are used as video frames in the video to be edited. Referring to the right figure in Figure 5, one or more static images are selected from the album, the display duration of each static image is set, the number of video frames is determined based on the correspondence between the display duration and the number of video frames, the static images are copied based on the above number of video frames, and the video to be edited is obtained by combining multiple video frames. Optionally, the total display duration of at least one static image can be greater than or equal to the duration of the object dynamic data, for example, the total display duration of at least one static image is consistent with the duration of the object dynamic data, and a new edited video is generated by at least one static image and the object dynamic data.

[0099] The video to be edited and the object dynamic data are displayed in the video editing page. The object dynamic data can be the latest generated object dynamic data (i.e., the latest object dynamic data in the material library), or it can be determined by a selection operation. For example, the video editing page can include a material library control, which displays the material library display page in response to a trigger operation (such as a click operation) on the material library control, and selects the object dynamic data on the material library display page. Exemplarily, the latest object dynamic data is displayed in the material library control, and the latest object dynamic data can be directly called by setting a gesture operation on the material library control (for example, it can be a material selection operation, which is different from the trigger operation of entering the material library display page, and exemplarily, it can be a long press operation), without entering the material library display page for selection, thereby simplifying the selection process of the object dynamic data.

[0100] The object dynamic data is used as material to edit the video to be edited to obtain an edited video, wherein the object dynamic data can be edited in multiple ways. In this embodiment, the video editing page includes at least one video editing mode control, which can provide at least one video editing mode. Accordingly, each video editing mode can correspond to a video editing module. When a trigger operation of any video editing mode control is detected, the object dynamic data and the video to be edited are edited by calling the video editing module corresponding to the video editing mode control to obtain an edited video corresponding to the video editing mode. This realizes one-click editing of the video, simplifies the video editing process, and reduces the difficulty of video production.

[0101] Optionally, the video editing mode includes one or more of a collage mode, a fusion mode, and a replacement mode; wherein the collage mode is an editing method for collaging the object dynamic image in the object dynamic data with the video frames in the video to be edited, and the fusion mode is an editing method for fusing the object dynamic image in the object dynamic data with the video frames in the video to be edited, wherein the light and shadow features of the video frames in the edited video obtained in the fusion mode are consistent with the light and shadow features of the original video frames. The replacement mode is an editing method for replacing the image content in the video frames in the video to be edited based on the object dynamic image in the object dynamic data.

[0102] It is understandable that the object dynamic data may also include audio data. In any of the above video editing modes, the video to be edited and the audio data in the object dynamic data may be merged.

[0103] Optionally, the method for generating the edited video in the collage mode includes: collaging the object dynamic graph in the object dynamic data with the corresponding video frame in the video to be edited, and / or, converging the audio data in the object dynamic data with the video to be edited. For example, refer to Figure 7, which is a schematic diagram of the processing of the video to be edited in the collage mode provided by the embodiment of the present disclosure. In response to the triggering operation of the collage mode control in the video editing page, the object dynamic graph in the object dynamic data is collaged with the video to be edited to obtain the edited video. Specifically, the object dynamic graph includes multiple target object graphs, and the multiple target object graphs correspond one-to-one to the multiple video frames of the video to be edited, for example, the i-th video frame corresponds to the i-th target object graph, and the target object graph with the corresponding relationship is collaged with the video frame. The target object graph can be located in the upper layer of the video frame to form a new video frame, and the multiple new video frames form the edited video. It can be understood that the new video frame sequence and the converging data of the audio data form the edited video.

[0104] Optionally, the method of generating the edited video in the fusion mode includes: pasting the object dynamic graph in the corresponding video frame of the video to be edited, and re-rendering the object dynamic graph in the object dynamic data based on the light and shadow characteristics of the video to be edited, and / or, merging the audio data in the object dynamic data with the video to be edited.

[0105] The multiple target object graphs in the object dynamic graph are respectively collaged with the multiple video frames of the video to be edited to form collage video frames, which will not be described in detail here. Since the object dynamic graph is captured from the target video, the light and shadow features of the object dynamic graph may be different from the light and shadow features of the video to be processed, resulting in inconsistent light and shadow features of the collage video frame. After obtaining the collage video frame, the target object graph in the collage video frame can be re-rendered to obtain a fused video frame, and the light and shadow features of the fused video frame are consistent. Specifically, the light and shadow features of the original video frame in the video to be edited are extracted, and the collage video frame is re-rendered based on the light and shadow features. This can be achieved through a pre-set rendering model, and the original video frame and the collage video frame are input into the rendering model. The original video frame is used as prior information, and the collage video frame is re-rendered to obtain a fused video frame. Multiple fused video frames form an edited video. Similarly, the edited video is the confluence data of the fused video frame sequence and the audio data.

[0106] Optionally, the method for generating the edited video in replacement mode includes: replacing the replacement object in the corresponding video frame of the video to be edited with the object dynamic image, re-rendering the object dynamic image in the object dynamic data based on the light and shadow characteristics of the video to be edited, and / or merging the audio data in the object dynamic data with the video to be edited. For example, see Figure 8, which is a schematic diagram of the processing of the video to be edited in replacement mode provided by an embodiment of the present disclosure.

[0107] Before editing a video to be edited, a replacement object is determined in response to a replacement object selection operation on the video editing page. Optionally, when the replacement object is determined, the replacement object is edge-marked to facilitate the user's intuitive viewing of the selected replacement object. The method for determining the replacement object is the same as the method for determining the target object and will not be further described here.

[0108] When a trigger operation for the replacement mode is detected on the video editing page, a replacement object is determined based on the replacement object selection operation, and the replacement object in the video frame of the video to be edited is replaced based on the object dynamic image in the object dynamic data. Specifically, the segmentation result of the replacement object is determined based on the video frame of the video to be edited, which can be, for example, a mask image. An intermediate video frame is generated based on the mask image of the replacement object and the video frame of the video to be edited. The intermediate video frame is a video frame obtained by removing the replacement object from the video frame of the video to be edited. Specifically, the mask image of the replacement object and the video frame of the video to be edited can be input into an image processing model to obtain an intermediate video frame output by the image processing model. The image processing model is pre-trained and based on the function of removing a set object from an image and restoring the image background. The set object is determined based on the input mask image. After the target object image in the object dynamic image and the intermediate video frame are collaged and re-rendered, a replacement video frame is obtained. The multiple replacement video frames form the edited video. It can be understood that the replacement video frame sequence and the confluence data of the audio data form the edited video.

[0109] In the above embodiment, before displaying the edited video in the video editing mode, the process further includes: receiving an editing setting operation, wherein the editing setting operation includes one or more of a position selection operation, a size adjustment operation, and a replacement object selection operation. For example, in collage mode and fusion mode, the position of the object dynamic image can be determined by the position selection operation, and the display size of the object dynamic image can be set by the size adjustment operation. In replacement mode, the replacement object can be determined by the replacement object selection operation.

[0110] The technical solution provided by the disclosed embodiments simplifies the editing of video materials by capturing object dynamic data from the target video being played using touch operations and using this object dynamic data as material for local video editing. Multiple video editing modes are set for the video to be edited. By triggering a video editing mode on the video editing page, one-click editing of the video to be edited and the object dynamic data can be achieved. This simplifies the video editing process, reduces the difficulty and cost of video production, and enhances the interactive fun during video playback.

[0111] FIG9 is a schematic structural diagram of a video processing device provided by an embodiment of the present disclosure. As shown in FIG9 , the device includes: a page display module 410 , a data interception module 420 , and a data processing module 430 .

[0112] A page display module 410 is used to display a video playback page, where the video playback page is used to play video stream data;

[0113] The data interception module 420 is configured to, during playback of a target video, determine a target object in response to a first trigger operation on the video playback page, and obtain object dynamic data corresponding to the target object, the object dynamic data being intercepted from multiple video frames of the target video;

[0114] The data processing module 430 is configured to execute a processing operation on the object dynamic data in response to a second triggering operation on the video playing page.

[0115] The technical solution provided by the disclosed embodiments captures the target video's dynamic image and / or audio data during video playback by inputting a first trigger operation. Furthermore, by inputting a second trigger operation, processing operations are performed on the target video's dynamic data. During subsequent video production, the target video's dynamic data can be added as material to the produced video, simplifying the material editing process and enabling interactive editing, thereby enhancing the interactive experience.

[0116] Based on the above embodiment, optionally, the object dynamic data includes an object dynamic graph formed by a target object graph intercepted from multiple video frames of the target video and / or audio data corresponding to multiple video frames.

[0117] After determining the target object, the method further includes: marking the outline of the target object.

[0118] Based on the above embodiment, optionally, the duration of the object dynamic data is the duration of the first trigger operation;

[0119] The data interception module 420 is further configured to: during the interception process of the object dynamic data, display the duration of the object dynamic data on the video playback page.

[0120] Optionally, the data processing module 430 is further configured to: switch the video playback page to a cutout mode, wherein the video playback page includes at least one dynamic data processing control in the cutout mode, and different dynamic data processing controls correspond to different processing operations;

[0121] The second trigger operation is a trigger operation of any of the dynamic data processing controls.

[0122] Based on the above embodiment, optionally, the data processing module 430 is further configured to:

[0123] The video playback page is switched to the video shooting page. During the video shooting process, the object dynamic image is displayed on the video shooting page, and / or the audio data in the object dynamic data is added to the video shooting page to obtain the shot video.

[0124] Optionally, the data processing module 430 is also used to: after switching the video playback page to the video shooting page, display the object dynamic image in the video shooting page, and in response to the adjustment operation of the object dynamic image, adjust the display status of the object dynamic image in the video shooting page, and the display status includes one or more of display position, display size, and display mode.

[0125] Based on the above embodiment, optionally, the data processing module 430 is further configured to:

[0126] Displaying a video editing page, wherein the video editing page is used to display the video to be edited, and the video editing page includes at least one video editing mode control, each of the video editing mode controls corresponding to a video editing mode;

[0127] In response to a selection operation on the video editing mode, an edited video in the video editing mode is displayed, where the edited video is obtained by editing the video to be edited in the video editing mode based on the object dynamic data.

[0128] Optionally, the video editing mode includes one or more of a collage mode, a fusion mode, and a replacement mode;

[0129] The method for generating the edited video in the collage mode includes: collaging the object dynamic image in the object dynamic data with the corresponding video frame in the video to be edited, and / or merging the audio data in the object dynamic data with the video to be edited;

[0130] The method for generating the edited video in the fusion mode includes: collaging the object dynamic image in the corresponding video frame of the video to be edited, and re-rendering the object dynamic image in the object dynamic data based on the light and shadow characteristics of the video to be edited, and / or merging the audio data in the object dynamic data with the video to be edited;

[0131] The method for generating the edited video in the replacement mode includes: replacing the replacement object in the corresponding video frame of the video to be edited with the re-rendered object dynamic graph, and re-rendering the object dynamic graph in the object dynamic data based on the light and shadow characteristics of the video to be edited, and / or merging the audio data in the object dynamic data with the video to be edited.

[0132] Optionally, the data processing module 430 is further configured to:

[0133] Before displaying the edited video in the video editing mode, an edit setting operation is received, wherein the edit setting operation includes one or more of a position selection operation, a size adjustment operation, and a replacement object selection operation.

[0134] The video processing device provided in the embodiments of the present disclosure can execute the video processing method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.

[0135] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.

[0136] Figure 10 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Referring to Figure 10 below, it shows a schematic diagram of the structure of an electronic device (such as the terminal device or server in Figure 10) 500 suitable for implementing the embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. The electronic device shown in Figure 10 is only an example and should not bring any limitations to the functions and scope of use of the embodiments of the present disclosure.

[0137] As shown in FIG10 , the electronic device 500 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the electronic device 500 are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An edit / output (I / O) interface 505 is also connected to the bus 504.

[0138] Typically, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG10 illustrates the electronic device 500 with various devices, it should be understood that not all of the illustrated devices are required to be implemented or present. More or fewer devices may alternatively be implemented or present.

[0139] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0140] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0141] The electronic device provided by the embodiment of the present disclosure and the video processing method provided by the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0142] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the video processing method provided by the above embodiment is implemented.

[0143] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0144] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.

[0145] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0146] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0147] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: displays a video playback page, which is used to play video stream data; during the playback of the target video, in response to a first trigger operation on the video playback page, determines the target object and obtains object dynamic data corresponding to the target object, wherein the object dynamic data is captured from multiple video frames of the target video; and in response to a second trigger operation on the video playback page, performs a processing operation on the object dynamic data.

[0148] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0149] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0150] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0151] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0152] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0153] According to one or more embodiments of the present disclosure, [Example 1] provides a video processing method, including:

[0154] Displaying a video playback page, wherein the video playback page is used to play video stream data;

[0155] During playback of a target video, in response to a first triggering operation on the video playback page, determining a target object and obtaining object dynamic data corresponding to the target object, the object dynamic data being captured from a plurality of video frames of the target video;

[0156] In response to a second triggering operation on the video playback page, a processing operation on the object dynamic data is performed.

[0157] According to one or more embodiments of the present disclosure, [Example 2] provides the video processing method of Example 1, further comprising:

[0158] The object dynamic data includes an object dynamic graph formed by a target object graph intercepted from multiple video frames of the target video and / or audio data corresponding to the multiple video frames.

[0159] According to one or more embodiments of the present disclosure, [Example 3] provides the video processing method of Example 1, further comprising:

[0160] After the target object is determined, the target object is contour marked.

[0161] According to one or more embodiments of the present disclosure, [Example 4] provides the video processing method of Example 1, further comprising:

[0162] The duration of the object dynamic data is the duration of the first trigger operation;

[0163] The method further includes: during the interception process of the object dynamic data, displaying the duration of the object dynamic data on the video playback page.

[0164] According to one or more embodiments of the present disclosure, [Example 5] provides the video processing method of Example 1, further comprising:

[0165] Switch the video playback page to the cutout mode. In the cutout mode, the video playback page includes at least one dynamic data processing control, and different dynamic data processing controls correspond to different processing operations; the second trigger operation is the trigger operation of any of the dynamic data processing controls.

[0166] According to one or more embodiments of the present disclosure, [Example 6] provides the video processing method of Example 1, further comprising:

[0167] The execution of the processing operation on the object dynamic data includes: switching the video playback page to the video shooting page, displaying the object dynamic image on the video shooting page during the video shooting process, and / or adding the audio data in the object dynamic data to the video shooting page to obtain the shot video.

[0168] According to one or more embodiments of the present disclosure, [Example 7] provides the video processing method of Example 1, further comprising:

[0169] After the video playback page is switched to the video shooting page, an object dynamic image is displayed in the video shooting page. In response to an adjustment operation on the object dynamic image, the display status of the object dynamic image in the video shooting page is adjusted, and the display status includes one or more of display position, display size, and display mode.

[0170] According to one or more embodiments of the present disclosure, [Example 8] provides the video processing method of Example 1, further comprising:

[0171] The method further comprises:

[0172] A video editing page is displayed, wherein the video editing page is used to display the video to be edited, and the video editing page includes at least one video editing mode control, each of the video editing mode controls corresponding to a video editing mode; in response to a selection operation of the video editing mode, an edited video in the video editing mode is displayed, wherein the edited video is obtained by editing the video to be edited in the video editing mode based on the object dynamic data.

[0173] According to one or more embodiments of the present disclosure, [Example 9] provides the video processing method of Example 1, further comprising:

[0174] The video editing mode includes one or more of a collage mode, a fusion mode and a replacement mode;

[0175] The method for generating the edited video in the collage mode includes: collaging the object dynamic image in the object dynamic data with the corresponding video frame in the video to be edited, and / or merging the audio data in the object dynamic data with the video to be edited;

[0176] The method for generating the edited video in the fusion mode includes: collaging the object dynamic image in the corresponding video frame of the video to be edited, and re-rendering the object dynamic image in the object dynamic data based on the light and shadow characteristics of the video to be edited, and / or merging the audio data in the object dynamic data with the video to be edited;

[0177] The method for generating the edited video in the replacement mode includes: replacing the replacement object in the corresponding video frame of the video to be edited with the re-rendered object dynamic graph, and re-rendering the object dynamic graph in the object dynamic data based on the light and shadow characteristics of the video to be edited, and / or merging the audio data in the object dynamic data with the video to be edited.

[0178] According to one or more embodiments of the present disclosure, [Example 10] provides the video processing method of Example 1, further comprising:

[0179] Before displaying the edited video in the video editing mode, an edit setting operation is received, wherein the edit setting operation includes one or more of a position selection operation, a size adjustment operation, and a replacement object selection operation.

[0180] According to one or more embodiments of the present disclosure, [Example 11] provides a video processing device, including:

[0181] A page display module is used to display a video playback page, wherein the video playback page is used to play video stream data;

[0182] a data interception module for determining a target object and obtaining object dynamic data corresponding to the target object in response to a first trigger operation on the video playback page during playback of the target video, the object dynamic data being intercepted from multiple video frames of the target video;

[0183] The data processing module is used to execute a processing operation on the object dynamic data in response to a second trigger operation on the video playback page.

[0184] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0185] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0186] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video processing method, wherein, Including: Display a video playback page for playing video stream data; During the playback of the target video, in response to a first trigger operation on the video playback page, determine a target object and obtain object dynamic data corresponding to the target object, where the object dynamic data is intercepted from multiple video frames of the target video; In response to a second trigger operation on the video playback page, perform a processing operation on the object dynamic data.

2. The method according to claim 1, wherein The object dynamic data includes an object dynamic graph formed by target object graphics intercepted from multiple video frames of the target video and / or audio data corresponding to the multiple video frames.

3. The method according to claim 1, wherein, After determining the target object, it further includes: performing contour marking on the target object.

4. The method according to claim 1, wherein, The duration of the object dynamic data is the duration of the first trigger operation; The method further includes: during the interception of the object dynamic data, displaying the duration of the object dynamic data on the video playback page.

5. The method according to claim 1, wherein The method further includes: Switch the video playback page to a matte painting mode, in which the video playback page includes at least one dynamic data processing control, and different dynamic data processing controls correspond to different processing operations; The second trigger operation is a trigger operation on any one of the dynamic data processing controls.

6. The method according to claim 2, wherein, The performing the processing operation on the object dynamic data includes: Switch the video playback page to a video shooting page. During video shooting, display the object dynamic graph on the video shooting page and / or add the audio data in the object dynamic data to the video shooting page to obtain a shooting video.

7. The method according to claim 6, wherein, After switching the video playback page to the video shooting page, it further includes: Display the object dynamic graph on the video shooting page. In response to an adjustment operation on the object dynamic graph, adjust the display state of the object dynamic graph on the video shooting page, where the display state includes one or more of display position, display size, and display mode.

8. The method according to claim 2, wherein The method further includes: Display a video editing page for displaying a video to be edited. The video editing page includes at least one video editing mode control, and each video editing mode control corresponds to a video editing mode; In response to a selection operation on the video editing mode, display the edited video in the video editing mode, where the edited video is obtained by editing the video to be edited in the video editing mode based on the object dynamic data.

9. The method according to claim 8, wherein, The video editing mode includes one or more of a collage mode, a fusion mode, and a replacement mode; Among them, the generation method of the edited video in the collage mode includes: collaging the object dynamic graph in the object dynamic data with the corresponding video frame in the video to be edited and / or performing a confluence process on the audio data in the object dynamic data and the video to be edited. The generation method of the edited video in the fusion mode includes: pasting the object dynamic graph in the corresponding video frame of the video to be edited, and re-rendering the object dynamic graph in the object dynamic data based on the light and shadow characteristics of the video to be edited, and / or performing a merging process on the audio data in the object dynamic data and the video to be edited; The generation method of the edited video in the replacement mode includes: replacing the replacement object in the corresponding video frame of the video to be edited with the re-rendered object dynamic graph, and re-rendering the object dynamic graph in the object dynamic data based on the light and shadow characteristics of the video to be edited, and / or performing a merging process on the audio data in the object dynamic data and the video to be edited.

10. The method according to claim 8, wherein, Before displaying the edited video in the video editing mode, it further includes: Receiving an editing setting operation, where the editing setting operation includes one or more of a position selection operation, a size adjustment operation, and a replacement object selection operation.

11. A video processing device, wherein, It includes: A page display module for displaying a video playback page, and the video playback page is used to play video stream data; A data intercepting module for determining a target object in response to a first trigger operation on the video playback page during the playback of the target video, and obtaining the object dynamic data corresponding to the target object, where the object dynamic data is intercepted from multiple video frames of the target video; A data processing module for performing a processing operation on the object dynamic data in response to a second trigger operation on the video playback page.

12. An electronic device, wherein, The electronic device includes: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method as described in any one of claims 1-10.

13. A storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the video processing method as described in any one of claims 1-10 when executed by a computer processor.

Citation Information

Patent Citations

  • Video processing method and device

    CN105307051A

  • GIF image generation method and device, server and storage medium

    CN111163358A

  • Method and device for intercepting image in video playing and computing equipment

    CN114173203A

  • Video processing method and device, storage medium and electronic equipment

    CN117880591A

  • Method and apparatus for editing video as well as recording medium for recording computer program for editing video

    JP2002016871A