Video generation method and device, electronic equipment and storage medium

By analyzing the display parameters change trend of the target object in the video frame, determining the scaling coefficient and scaling the target object, the screen shaking problem caused by unbalanced size of the target object during movement is solved, and a smoother and more coordinated video generation effect is achieved.

CN120201241APending Publication Date: 2025-06-24BEIJING BAIDU NETCOM SCI & TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510526725.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

During the movement, the distance between the object and the video acquisition device changes, resulting in uneven size of the target object in the video frame, causing the problem of screen shaking.

Method used

By acquiring the display parameters of the target object in the multiple video frames, analyzing the first change trend of these parameters along the time series, determining the scaling coefficients for each video frame, and scaling the target object to generate a stable video.

Benefits of technology

Reduces screen shaking, ensures that the target object displays more smoothly and coordinated among multiple video frames, and improves the visual effect of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201241A_ABST
    Figure CN120201241A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method, and relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, intelligent editing and the like. The method comprises the following steps: acquiring a plurality of video frames containing a target object, wherein the plurality of video frames indicate spatial position information of the target object along a time sequence; multiple display parameters of the target object are obtained from the multiple video frames, and the display parameters represent the size relation between the target object and the video frames; determining a plurality of scaling coefficients for the plurality of video frames respectively based on a first variation trend of the plurality of display parameters along the time sequence; and performing scaling processing on the target objects in the plurality of video frames based on the plurality of scaling coefficients to generate a target video. The invention further provides a video generation device, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the technical fields of computer vision, intelligent editing, etc., and can be applied to video generation scenarios. More specifically, the present disclosure provides a video generation method, device, electronic device, storage medium, and program product. Background Art

[0002] It is common to shoot the subject's movement, edit the video and share it, for example, people can share their sports moments on social media.

[0003] During the movement, the position of the object is constantly changing, and its distance from the video capture device also changes accordingly. If the object is far away from the video capture device, the object will be displayed smaller in the picture, making it difficult to effectively highlight the main body of the target object. Otherwise, the object will be displayed larger, and the overall composition will be inconsistent. Because the size of the object in the video capture device's picture keeps changing during the movement, if the picture is enlarged or reduced at a fixed ratio during the editing stage, the picture will shake during playback. Summary of the invention

[0004] The present disclosure provides a video generation method, device, electronic device, storage medium and program product.

[0005] According to one aspect of the present disclosure, a video generation method is provided, the method comprising: acquiring multiple video frames containing a target object, the multiple video frames indicating spatial position information of the target object along a time series; acquiring multiple display parameters of the target object from the multiple video frames, wherein the display parameters characterize a size relationship between the target object and the video frames; determining multiple scaling factors respectively used for the multiple video frames based on a first change trend of the multiple display parameters along the time series; and scaling the target objects in the multiple video frames respectively based on the multiple scaling factors to generate a target video.

[0006] According to another aspect of the present disclosure, a video generating device is provided, which includes: a video frame unit, configured to obtain multiple video frames containing a target object, the multiple video frames indicating spatial position information of the target object along a time series; a display parameter unit, configured to obtain multiple display parameters of the target object from the multiple video frames, wherein the display parameters characterize the size relationship between the target object and the video frame; a scaling factor unit, configured to determine multiple scaling factors respectively used for the multiple video frames based on a first change trend of the multiple display parameters along the time series; and a video generating unit, configured to scale the target objects in the multiple video frames respectively based on the multiple scaling factors to generate a target video.

[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.

[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided according to the present disclosure.

[0009] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method provided according to the present disclosure.

[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0012] Figure 1 is a schematic diagram of an exemplary system architecture to which a video generation method and apparatus according to an embodiment of the present disclosure can be applied;

[0013] Figure 2 is a flowchart of a video generation method according to an embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram of a first change trend according to an embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram of scaling based on the display height of a target object according to an embodiment of the present disclosure;

[0016] Figure 5 is a schematic diagram of a second change trend according to an embodiment of the present disclosure;

[0017] Figures 6A to 6C is a schematic diagram of a cropping process according to an embodiment of the present disclosure;

[0018] Figure 7 is a flowchart of a video generation method according to another embodiment of the present disclosure.

[0019] Figure 8 is a block diagram of a video generation apparatus according to an embodiment of the present disclosure; and

[0020] Figure 9 is a block diagram of an electronic device to which a video generation method according to an embodiment of the present disclosure can be applied. Detailed implementation manners

[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0022] Figure 1 is a schematic diagram of an exemplary system architecture to which a video generation method and apparatus according to an embodiment of the present disclosure can be applied. It should be noted that Figure 1 only the examples of the system architecture to which the embodiments of the present disclosure can be applied are shown to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0023] As Figure 1 shown, the system architecture 100 according to this embodiment may include a video capture device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the video capture device 101 and the server 103. The network 102 may include various connection types, such as wired and / or wireless communication links, etc.

[0024] The video capture device 101 may be various devices having a camera and supporting image or video shooting, including but not limited to cameras, smartphones, tablets, drones, etc.

[0025] The server 103 may be a server providing various services, such as a server for processing the video stream captured by the video capture device 101. The server 103 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud computing, network services, middleware services, etc.

[0026] Referring to Figure 1 , taking skiing as an example, the video capture device 101 can be used to shoot a skier. For example, during the skiing process of the skier, his companion uses a smartphone to track and shoot him. For example, through the video recording function of the smartphone, the "record button" can be pressed to shoot an image of the skier.

[0027] Exemplarily, in addition to using a smart phone, one or more cameras can be installed at a ski resort to capture videos of skiers present at the ski resort. For example, a certain skier can wear a dedicated device. After the server 103 detects the dedicated device, it controls the camera to capture a video of the skier and ends the capture after detection fails (such as when the detection distance is exceeded); alternatively, a certain skier can wear clothes with a specific identifier. After the server 103 detects the corresponding specific identifier from the video stream captured by the camera, it starts to control the camera to track and capture the skier until the target cannot be tracked anymore; or, without wearing a dedicated device and clothes with a specific identifier, one or more cameras capture all skiers present at the ski resort, and the server 103 uses a multi-object tracking algorithm to identify and track one or more skiers to obtain a video segment including a certain skier.

[0028] Multi-object tracking is a technology for tracking specific objects in consecutive video frames. It can track one or more objects in multiple video frames continuously captured by a single camera, or can also track one or more objects in multiple video frames captured simultaneously by multiple cameras with overlapping fields of view.

[0029] It should be noted that although the above example is for skiing, the present disclosure is not limited thereto. The video generation method, device, electronic device, storage medium, and program product provided in the embodiments of the present disclosure can be applied to scenarios where due to the change in distance from the video acquisition device 101, some picture objects in multiple video frames are displayed smaller or multiple pictures are displayed unevenly. For example, it can be applied to sports scenarios such as running, marathon competitions, or cycling with people as the target objects; it can be applied to amusement park activity scenarios with people as the target objects, such as roller coaster projects; it can be applied to animal activity scenarios with animals as the target objects; it can be applied to scenarios where non-living objects are the target objects, such as a racing scene of vehicles.

[0030] It should be noted that the video generation method provided in the embodiments of the present disclosure can generally be executed by the server 103. Correspondingly, the video generation device provided in the embodiments of the present disclosure can generally be set in the server 103. The video generation method provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 103 and capable of communicating with the video acquisition device 101 and / or the server 103. Correspondingly, the video generation device provided in the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 103 and capable of communicating with the video acquisition device 101 and / or the server 103.

[0031] In the technical solution of the present application, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, application, etc., all comply with relevant laws, regulations and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] Figure 2 is a flowchart of a video generation method according to an embodiment of the present disclosure. Figure 3 is a schematic diagram of a first change trend according to an embodiment of the present disclosure.

[0033] As Figure 2 shown, the method 200 may include operation S210 to operation S240.

[0034] In operation S210, obtain a plurality of video frames including a target object, and the plurality of video frames indicate the spatial position information of the target object along the time series.

[0035] Exemplarily, videos or a plurality of video frames actively uploaded by users can be obtained, or videos from video acquisition devices can also be obtained. The videos here can be videos with complete content after recording ends, or can also be obtained by loading video streams. For example, through object detection algorithms such as kernel correlation filtering algorithm, YOLO detection algorithm, FasterRCNN, SSD, etc., a plurality of video frames including the target object can be identified and extracted from the video.

[0036] During the movement of the target object, its spatial position changes over time. Through the inter-frame changes of the target object in a plurality of video frames, the spatial position information of the target object along the time series can be obtained. The spatial position information may include the spatial position change trajectory of the target object along the time series.

[0037] In operation S220, obtain a plurality of display parameters of the target object from the plurality of video frames, where the display parameters characterize the size relationship between the target object and the video frames.

[0038] In operation S230, based on the first change trend of the plurality of display parameters along the time series, determine a plurality of scaling factors for the plurality of video frames respectively.

[0039] Exemplarily, as Figure 3As shown, there are corresponding video frames at times t1, t2, t3, and t4. The object detection algorithm can be used to generate detection boxes for the target object in the video frames, which can include detection box identifiers and detection box coordinates. The detection box identifier is used to continuously track the same target object, and the detection box coordinates are used to characterize the position of the target object in the video frame.

[0040] For example, the display parameters of the target object can be extracted from the video frames. The display parameters can include parameters that are positively correlated with the display size of the target object, such as parameters that change with the display size, like the detection box height, face area, torso area, or display ratio. The display size is characterized by a size relationship and can be determined by the ratio of the number of pixels in the target object area to the total number of pixels in the video frame, or by the ratio of the area of the detection box to the total area of the video frame. It can be understood that the size relationship can also be characterized in other ways, which are not specifically limited here.

[0041] Figure 3 In [the figure], the vertical axis represents the measure of the display parameter size, and the horizontal axis represents the time series. Referring to Figure 3 , along the time series of t1, t2, t3, and t4, the first change trend of multiple display parameters can be obtained. The first change trend reflects the change of the size relationship of the target object in the video frame over time.

[0042] For example, an initial scaling coefficient can be assigned to each video frame, and according to the change rate between the display parameters at adjacent times, the initial scaling coefficients of the corresponding two video frames can be adjusted to have basically the same change rate, and finally multiple scaling coefficients are obtained. Or, different historical change trends can be collected in advance and mapped to the corresponding multiple scaling coefficients, and the first change trend in operation S230 is matched with each of the pre-collected historical change trends to use the multiple scaling coefficients of the matched historical change trend as the result of executing operation S230. Or, the first change trend can be smoothed, and then multiple scaling coefficients are obtained based on the smoothed trend.

[0043] In operation S240, the target object in multiple video frames is scaled based on multiple scaling coefficients respectively to generate a target video. For example, the video frames can be scaled, or the target object can be extracted from the video frames for scaling.

[0044] Among them, different video frames can have the same or different scaling coefficients. For example Figure 3 the video frame at time t1 in [the figure] may have a reduction coefficient, the video frame at time t4 may have an enlargement coefficient, or both video frames have reduction coefficients, but they are not the same.

[0045] For example, the multiple scaled video frames are subsequently clipped and then merged in time series to generate a target video.

[0046] Through the embodiments of the present disclosure, compared with the method of scaling according to a fixed ratio, the screen shake is weakened to a certain extent. By using the first change trend to determine multiple scaling factors, the overall display situation of the target object among multiple video frames can be considered, making the scaled target object more smoothly displayed among multiple video frames, and the generated target video can provide a better viewing experience.

[0047] Figure 4 It is a schematic diagram of scaling based on the display height of the target object according to an embodiment of the present disclosure.

[0048] In some embodiments, determining multiple scaling factors respectively for multiple video frames based on the first change trend of multiple display parameters along the time series may include: correcting the multiple display parameters based on the first fitting curve to obtain multiple fitting parameters, where the fitting parameters characterize the corrected size relationship between the target object and the video frame, and the first fitting curve characterizes the first change trend of the smoothed multiple display parameters; determining multiple scaling factors based on the multiple corrected size relationships characterized by the multiple fitting parameters.

[0049] Exemplarily, the first change trend can be fitted (such as polynomial fitting, exponential fitting, or piecewise fitting, etc.) to obtain the first fitting curve, thereby realizing the smoothing process of the first change trend. Among them, the correction process includes obtaining the fitting parameters corresponding to the display parameters of the video frame on the first fitting curve based on the video frame at the same moment.

[0050] Figure 4 The horizontal axis of [diagram] represents the video frames sorted according to the time series. For example, the abscissa "50" refers to the 50th video frame among the multiple video frames obtained. The vertical axis represents the display height of the target object in each video frame, which is characterized by the number of pixels of the target object in the height direction.

[0051] As Figure 4 shown, the first change trend is a curve obtained from the original data of the display height (i.e., display parameter) of the target object in each video frame, and its fluctuations are relatively frequent. Taking the skiing sport as an example, its fluctuations may be related to the terrain of the ski slope in the ski resort. The first change trend as a whole shows a trend of the display height of the target object increasing from small to large. At the end of the first change trend, from Figure 4 the turning point 1 marked, the display height of the target object drops extremely rapidly, which may be caused by the target object gradually approaching the video acquisition device until it leaves the acquisition range.

[0052] Figure 4The first fitting curve can be obtained by performing polynomial fitting on the first trend of change therein. The overall first fitting curve is smoother than the first trend of change. For example, the line segment of the first fitting curve below 150 pixels on the vertical axis reduces the fluctuation of the corresponding part of the first trend of change. Moreover, at the turning point 1 marked by Figure 4 there is no sharp decrease.

[0053] The vertical coordinate on the first fitting curve is the fitting parameter, for example, the corrected display height of the target object. The display height of the target object is proportional to the display ratio. Therefore, the corrected display height of the target object characterizes the corrected size relationship between the target object and the video frame.

[0054] Through the embodiments of the present disclosure, the fluctuations of multiple display parameters in the time series are smoothed by multiple fitting parameters. Based on multiple corrected size relationships, multiple scaling factors are determined, which can make the displayed target object transition naturally between video frames and avoid screen shaking.

[0055] In some embodiments, determining multiple scaling factors based on multiple corrected size relationships characterized by multiple fitting parameters includes: obtaining multiple scaling factors based on the differences between multiple corrected size relationships and a preset size relationship respectively.

[0056] Exemplarily, the preset size relationship is determined according to the expected display effect of the target object. This display effect can be determined according to expert experience or obtained by collecting the scores of different users for statistics. For example, if the human height accounts for 20% of the video frame and has a good display effect, the preset size relationship is 20% (hereinafter represented as 0.2).

[0057] For example, the size of the video frame captured by the video acquisition device is 3840*1920 pixels. Let h represent the fitting parameter, such as h being Figure 4 the display height of the target object after fitting in

[0058]

[0059] As can be seen from the above formula, the difference in this embodiment is the ratio between the preset size relationship and the corrected size relationship, and this ratio is used as the scaling factor.

[0060] Through the embodiments of the present disclosure, the difference between the corrected size relationship and the preset size relationship can reflect the difference from the expected display effect, so that the difference in display can be reduced through scaling processing using the scaling factor.

[0061] It should be noted that the above-described method for determining the difference is only an exemplary illustration that can be implemented in the present disclosure and does not constitute a limitation to the present disclosure. For example, a scaling coefficient can also be obtained based on the difference between the preset size relationship and the corrected size relationship. For instance, the scaling coefficient is assigned according to the difference interval, and the relationship between the difference interval and the assignment can be determined in advance according to expert experience.

[0062] Continuing with the example where the preset size relationship is 0.2, when there is a difference between the corrected size relationship and 0.2, the purpose of the scaling process is to make the size relationship between the scaled target object and the video frame tend to 0.2. However, if the scaling amplitude is too large, it will also affect the visual effect of the target video. Therefore, the scaling coefficient can be restricted. For example, the magnification coefficient is restricted to be between [0.618, 1.618] (only for example). And the size relationship after the scaling process can be restricted. Referring to Figure 4 , the correction result is the height change curve of the target object displayed after the scaling process. At Figure 4 the turning point 2 marked, it means that the displayed height of the target object reaches the maximum value and the size relationship no longer continues to be enlarged.

[0063] Figure 5 is a schematic diagram of the second change trend according to an embodiment of the present disclosure.

[0064] In some embodiments, scaling the target object in multiple video frames based on multiple scaling coefficients to generate a target video may include: scaling multiple video frames based on multiple scaling coefficients to scale the target object in the multiple video frames. Shrinking or enlarging the video frames according to the scaling coefficient, then the target object displayed therein is correspondingly shrunk or enlarged; obtaining multiple candidate centers of the target object from the scaled multiple video frames, where the candidate center represents the center position of the target object in the video frame; determining multiple cropping frames for the scaled multiple video frames respectively based on the second change trend of the multiple candidate centers along the time series; and performing cropping processing on the target object in the scaled multiple video frames respectively based on the multiple cropping frames to generate a target video. For example, editing the cropped video frames and then merging them to obtain the target video.

[0065] For example, after the target object is scaled, its detection box is adaptively scaled. The center position of the target object in the video frame can be the center position of the scaled detection box (such as a rectangular box), or the center position of the region of interest determined based on the target object. The region of interest can be a region determined based on the contour of the target object. For example, the center of gravity position is determined as the candidate center based on the humanoid contour.

[0066] Exemplarily, as Figure 5 shown, candidate centers can be extracted from the video frames at times t1, t2, t3, and t4 respectively, so as to obtain the second change trend.Figure 5 The x-axis and y-axis of the shown coordinate system respectively correspond to the x-axis and y-axis of the coordinate system in the video frame. The second changing trend reflects the spatial position change trajectory of the target object over time.

[0067] For example, the changing trend of the central positions of multiple cropping frames can be substantially consistent with the second changing trend, and there may be a certain deviation between the candidate center in each video frame and the central position of the cropping frame. For example, according to the second changing trend, the central position of the cropping frame can be weighted and adjusted based on the candidate center. If the second changing trend indicates that the target object continuously moves in a certain direction, the central position of the cropping frame can be appropriately offset in that direction, and the offset amount can be determined according to factors such as the moving speed and acceleration. Thus, the cropped result can be smoother and more natural visually, conforming to people's visual perception habits of motion.

[0068] For example, for the size of the cropping frame, it can be determined according to the content of the corresponding video frame and the display requirements. It can be considered to use a cropping frame with a fixed size adapted to the device for playing the target video. For example, if the target video is played on a mobile phone, a 1080P-sized cropping frame is used. It can also dynamically adjust the size of the cropping frame. For example, if the area around the target object is relatively cluttered, the cropping frame can be reduced, and vice versa.

[0069] For example, during the video acquisition process, the spatial position of the target object changes, and accordingly, the central position of the target object continuously changes between video frames. And due to the scaling operation, the central position of the target object further changes, and it may be difficult to highlight the target object as the main body in some video frames.

[0070] Through the embodiments of the present disclosure, by using the spatial position change trajectory of the target object reflected by the second changing trend, determining multiple cropping frames and respectively using them to crop multiple scaled video frames can effectively highlight the main body of the target object, make the transition between adjacent video frames after cropping natural, and to a certain extent avoid situations such as sudden jumps in the picture and sudden changes in the position of the main body of the target object, optimizing the visual effect.

[0071] Figures 6A to 6C It is a schematic diagram of a cropping processing procedure according to an embodiment of the present disclosure.

[0072] In some embodiments, determining multiple cropping frames respectively for multiple scaled video frames based on the second changing trend of multiple candidate centers along the time series may include: correcting multiple candidate centers based on a second fitting curve to obtain multiple fitted centers, the second fitting curve representing the second changing trend of the multiple candidate centers after smoothing processing, and the fitting parameters representing the corrected central position of the target object in the video frame; using the multiple fitted centers as the central positions of the multiple cropping frames respectively to obtain multiple cropping frames.

[0073] Exemplarily, the second variation trend can be fitted (such as polynomial fitting, exponential fitting, piecewise fitting, etc.) to obtain a second fitting curve, thereby realizing the smoothing of the second variation trend. Among them, the calibration process includes obtaining the calibrated center position corresponding to the video frame at the same moment on the second fitting curve, that is, the fitting center.

[0074] Through the embodiments of the present disclosure, multiple fitting centers can smooth the fluctuations of multiple candidate centers in the time series, that is, make the spatial position change trajectory of the target object smoother in the time series. On this basis, the cropping frames of each video frame are determined. The cropping process can improve the coherence of the transition between video frames, make the spatial position change trajectory of the target object have better visual fluency, and simulate the effect of digital camera movement.

[0075] Figure 6A Shows the center position of the cropping frame based on the fitting center. Figure 6B Shows adjusting the size of the cropping frame according to the boundary of the video frame. As Figure 6A Shown, the fitting center is used as the center position of the cropping frame, and the cropping frame is determined accordingly. There is a certain deviation between the fitting center and the candidate center. For example, in order to make the target video have a better playback effect on the mobile phone, a 1080P cropping frame can be generated. After determining the center position and size of the cropping frame, the situation where the cropping frame exceeds the video frame boundary may occur as Figure 6A Shown. In this case, the cropping frame can be moved into the picture of the video frame, or the range of the cropping frame can be reduced as Figure 6B Shown. Figure 6C Shows the cropping result obtained by processing the video frame based on the cropping frame.

[0076] Among them, in order to adapt to the 1080P size, Figure 6B The reduced part in can be generated using AIGC technology or supplemented with special effects. AIGC (Artificial Intelligence Generated Content) is a technology that uses artificial intelligence technology, especially methods such as large pre-trained models, to generate relevant content through learning and pattern recognition of existing data. The core idea of AIGC technology is to use artificial intelligence algorithms to generate content with certain creativity and quality, and it can generate related articles, images, audio, etc. according to the input conditions or guidance.

[0077] Figure 7 Is a flowchart of a video generation method according to another embodiment of the present disclosure.

[0078] As Figure 7As shown, the video generation method 700 of this embodiment includes operations S710 to S730. Taking skiing as an example, for instance, a video acquisition device (such as one or more cameras) is set up beside the snow track. The video acquisition device can acquire videos of one or more skiers in the snow track and push the video stream to the cloud server in real time. The following will further elaborate.

[0079] In operation S710, automated and intelligent video analysis can be achieved based on artificial intelligence technology. In this embodiment, operation S710 can include operations S711 to S715.

[0080] In operation S711, the cloud server pulls the real-time video stream.

[0081] In operation S712, through a multi-object tracking algorithm based on artificial intelligence technology (such as the YOLO detection algorithm), multiple skiers in the video stream are tracked. For example, the video frame sequence and detection box sequence of each skier from entering the frame to leaving the frame are recorded.

[0082] In operation S713, video "stripping" is performed on the selected target object. Video "stripping" means extracting the initial frame sequence of the target object from the video stream according to the tracking record of operation S712. The start time and end time of the extraction are the entry time and exit time of the target object respectively. Among them, the target object can be selected according to the user operation, or each tracked target can be used as the selected target object.

[0083] In operation S714, filters are used to filter multiple initial frame sequences of the target object. In some embodiments, an initial frame sequence containing multiple video frames can be obtained; the initial frame sequence is evaluated based on preset rules to obtain an evaluation result. The preset rules include rules for evaluating at least one of size relationship, video acquisition duration, spatial position information, and target object movement speed information; in the case where the evaluation result is passed, multiple video frames are obtained based on the initial frame sequence.

[0084] Exemplarily, the filter may include a software service capable of invoking and executing preset rules. For example, the rule for evaluating the size relationship may include evaluating whether the average size of the detection boxes in the initial frame sequence is greater than or equal to the detection box threshold, and the size may include at least one of the box height, box width, and box area; the rule for evaluating the video capture duration may include evaluating whether the duration between the entry time and the exit time of the target object is greater than or equal to the duration threshold; the rule for evaluating the spatial position information may include whether the spatial position has changed; the rule for evaluating the moving speed information of the target object may include whether the moving speed is greater than or equal to the speed threshold. It can be understood that if at least one of the following situations occurs: the average size of the detection box is too small, the video capture duration is too short, the target object does not move, or the moving speed is too low, it may be difficult to obtain the target video that meets the expectations, and the evaluation fails, and the initial frame sequence is filtered out.

[0085] Exemplarily, in the case where the evaluation passes, the initial frame sequence is used as multiple video frames of the target object, or some video frames are extracted from the initial frame sequence as multiple video frames of the target object. Among them, extraction may include extraction at intervals and / or evaluation of the video frame quality. Evaluating the video frame quality can be achieved by evaluating one or more of the detection box size, the occlusion situation of the target object, the video frame picture resolution, and the video frame picture contrast of each video frame.

[0086] In operation S715, attribute analysis is performed on the target object. The role of attribute analysis is to extract the object features of the target object for associated storage with its target video.

[0087] For example, object features of the target object are obtained from multiple video frames, and the object features are obtained from at least one of the target object's morphological information, the target object's clothing information, the target object's equipment information, the video capture time, the video capture device information, and the target object's spatial position change trajectory; the object features are stored in association with the target video.

[0088] Among them, the target object's morphological information may include height, body contour, hairstyle, and face information, etc.; the target object's clothing information may include color information and style information, etc.; the target object's equipment information may include the type of equipment carried, such as the type of skateboard; the video capture device information may include the camera identifier, etc.; the target object's spatial position change trajectory may include the spatial position change trajectory of the target object. For example, if the target object triggers the "start skiing" function with a smartphone, the cloud server obtains this instruction and records the spatial position change trajectory of the target object through the positioning information of the smartphone.

[0089] For example, the largest human body frame can be cropped from the frame sequence and position box in the tracking record of the target object, and then analyzed (e.g., analyzed through a neural network algorithm) to identify the clothing color and skateboard type. At the same time, the color, skateboard type, entry time, exit time, and camera number are saved for subsequent quick retrieval of the target video of the target object.

[0090] The video acquisition device can capture multiple skiers on the snow track and extract multiple video frames for each skier as the target object to generate the target video. Among them, multiple target videos may be generated when each skier skis multiple times. Therefore, the cloud may generate a relatively large number of target videos. In the multi-object tracking stage, the identification of the detection box is used to distinguish the tracking target, and it may be difficult to match the specific information of the specific skier. For example, the identification of the detection box is "0001", and it is difficult to represent the specific information of the target object selected by the box. Therefore, in order to accurately push the target video to the target object or facilitate the target object to quickly retrieve its target video, the object features can be associated and stored with the target video.

[0091] In operation S720, the object features of the target object and multiple video frames can be encapsulated into a video production task and pushed to the video production task queue. The video production task queue can be a distributed queue, only for example. Operation S720 can separate video analysis and video production, facilitating horizontal expansion in terms of computing resources.

[0092] In operation S730, several worker processes can be derived. The number of worker processes is determined according to factors such as the number of video production tasks, the resource consumption of video production tasks, and the available computing resources. In operation S730, the worker processes can be used to execute operations S731 to S735.

[0093] In operation S731, the worker process pulls the task. Among them, multiple worker processes can work in parallel in a distributed manner, each taking out the video production task from the video production task queue and executing operations S731 to S735.

[0094] In operation S732, scaling processing is performed. For example, multiple display parameters of the target object are obtained from multiple video frames, and based on the first change trend of the multiple display parameters along the time series, multiple scaling factors for the multiple video frames are determined; then the multiple video frames are scaled based on the multiple scaling factors.

[0095] In some embodiments, the cloud server can detect the size of the detection box of the target object through real-time streaming. In the case where the detected size of the detection box is less than a certain value, the camera can be controlled to change the zoom ratio to increase the size of the detection box of the target object, which is beneficial to providing clear video frames for scaling processing in operation S732.

[0096] In operation S733, digital camera movement is simulated. For example, multiple candidate centers of a target object are obtained from multiple scaled video frames, and based on the second change trend of the multiple candidate centers in a time series, multiple cropping frames for the multiple scaled video frames are determined. Then, after cropping the multiple video frames, a target video is generated to simulate the effect of digital camera movement.

[0097] In some embodiments, the action types of the target object in multiple video frames can be identified respectively, and multiple target video frames that conform to a preset action type are determined; a target video segment is generated based on the multiple target video frames; and the target video segment is added to a preset position in the target video. For example, the target video segment can be added to the target video and played first at the start position.

[0098] For example, the preset action types can include slalom (a small S-shaped trajectory on a ski slope), giant slalom (a large S-shaped trajectory on a ski slope), crossing an obstacle, jumping and spinning (rotating in the air after takeoff and then landing), and somersaulting (controlling a somersault in the air after takeoff and then landing), etc.

[0099] For other scenarios, the preset action types can be passing, dribbling, dunking, overtaking, etc., which are not limited here.

[0100] In operation S734, video encoding is performed. For example, if the video frame format is the YUV format, or image formats such as jpeg or png obtained by converting YUV, after cropping the multiple video frames, they can be merged and encoded to obtain a target video in formats such as MP4 or MKV.

[0101] In operation S735, the target video is saved to cloud storage for users to download.

[0102] Through the embodiments of the present disclosure, a video generation method based on multi-object tracking, video "stripping", and realizing digital camera movement is provided. Taking skiing as an example, by installing a video acquisition device beside the ski slope to collect videos in real time, the cloud server performs multi-object tracking, stripping processing, and digital camera movement on the video stream, and a high-quality target video can be generated for the target object to view and share.

[0103] Figure 8 It is a block diagram of a video generation device according to an embodiment of the present disclosure.

[0104] As Figure 8 shown, the video generation device 800 may include a video frame unit 810, a display parameter unit 820, a scaling factor unit 830, and a video generation unit 840.

[0105] The video frame unit 810 is configured to obtain a plurality of video frames including a target object, and the plurality of video frames indicate the spatial position information of the target object along the time series.

[0106] The display parameter unit 820 is configured to obtain a plurality of display parameters of the target object from the plurality of video frames, where the display parameters characterize the size relationship between the target object and the video frames.

[0107] The scaling factor unit 830 is configured to determine a plurality of scaling factors for the plurality of video frames respectively based on the first change trend of the plurality of display parameters along the time series.

[0108] The video generation unit 840 is configured to perform scaling processing on the target object in the plurality of video frames respectively based on the plurality of scaling factors to generate a target video.

[0109] Exemplarily, the scaling factor unit 830 is further configured to correct the plurality of display parameters based on the first fitting curve to obtain a plurality of fitting parameters, where the fitting parameters characterize the corrected size relationship between the target object and the video frames, and the first fitting curve characterizes the first change trend of the plurality of display parameters after smoothing processing; determine the plurality of scaling factors based on the plurality of corrected size relationships characterized by the plurality of fitting parameters.

[0110] Exemplarily, the scaling factor unit 830 is further configured to obtain the plurality of scaling factors based on the differences between the plurality of corrected size relationships and the preset size relationships respectively.

[0111] Exemplarily, the video generation unit 840 is further configured to perform scaling processing on the plurality of video frames respectively based on the plurality of scaling factors to scale the target object in the plurality of video frames; obtain a plurality of candidate centers of the target object from the scaled plurality of video frames, where the candidate centers represent the center positions of the target object in the video frames; determine a plurality of cropping frames for the scaled plurality of video frames respectively based on the second change trend of the plurality of candidate centers along the time series; perform cropping processing on the target object in the scaled plurality of video frames respectively based on the plurality of cropping frames to generate a target video.

[0112] Exemplarily, the video generation unit 840 is further configured to correct the plurality of candidate centers based on the second fitting curve to obtain a plurality of fitting centers, where the second fitting curve characterizes the second change trend of the plurality of candidate centers after smoothing processing, and the fitting parameters characterize the corrected center positions of the target object in the video frames; use the plurality of fitting centers as the center positions of the plurality of cropping frames respectively to obtain the plurality of cropping frames.

[0113] Exemplarily, the video generation device 800 may further include an associated storage unit configured to obtain object features of a target object from a plurality of video frames, where the object features are obtained based on at least one of target object morphology information, target object clothing information, target object equipment information, video acquisition time, video acquisition device information, and target object spatial position change trajectory; and associate and store the object features with the target video.

[0114] Exemplarily, the video frame unit 810 is further configured to obtain an initial frame sequence including a plurality of video frames; evaluate the initial frame sequence based on a preset rule to obtain an evaluation result, where the preset rule includes a rule for evaluating at least one of size relationship, video acquisition duration, spatial position information, and target object movement speed information; and obtain a plurality of video frames based on the initial frame sequence when the evaluation result passes.

[0115] Exemplarily, the video generation device 800 may further include an identification unit and a segment insertion unit. The identification unit is configured to separately identify the action types of the target object in a plurality of video frames to determine a plurality of target video frames that conform to a preset action type; the video generation unit 840 is further configured to generate a target video segment based on the plurality of target video frames; and the segment insertion unit is configured to add the target video segment to a preset position in the target video.

[0116] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0117] Figure 9 is a block diagram of an electronic device that can apply the video generation method according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0118] As Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 902 or computer programs loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0119] Multiple components in device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0120] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as a text processing method and / or a deployment method of a deep learning framework. For example, in some embodiments, the text processing method and / or the deployment method of the deep learning framework can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the text processing method and / or the deployment method of the deep learning framework described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute a video processing method in any other appropriate manner (e.g., by means of firmware).

[0121] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip (SOC) systems, complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0122] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) monitor or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0125] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0126] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs that run on the respective computers and have a client-server relationship with each other.

[0127] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0128] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A video generation method, comprising: Acquire a plurality of video frames containing a target object, wherein the plurality of video frames indicate spatial position information of the target object along a time series; Acquire a plurality of display parameters of the target object from the plurality of video frames, wherein the display parameters represent a size relationship between the target object and the video frame; Determining a plurality of scaling factors respectively used for the plurality of video frames based on a first change trend of the plurality of display parameters along the time series; The target objects in the multiple video frames are scaled based on the multiple scaling factors to generate target videos.

2. The method according to claim 1, wherein: The determining, based on the first change trend of the plurality of display parameters along the time series, a plurality of scaling factors respectively used for the plurality of video frames comprises: Correcting the plurality of display parameters based on a first fitting curve to obtain a plurality of fitting parameters, wherein the fitting parameters represent a corrected size relationship between the target object and the video frame, and the first fitting curve represents a first variation trend of the plurality of display parameters after smoothing; The plurality of scaling factors are determined based on a plurality of corrected dimensional relationships characterized by the plurality of fitting parameters.

3. The method according to claim 2, wherein: Determining the plurality of scaling factors based on the plurality of corrected size relationships represented by the plurality of fitting parameters comprises: The plurality of scaling factors are obtained based on differences between the plurality of corrected size relationships and the preset size relationships.

4. The method according to claim 1, wherein: The step of respectively scaling the target objects in the multiple video frames based on the multiple scaling factors to generate a target video includes: Performing scaling processing on the multiple video frames respectively based on the multiple scaling coefficients to scale the target objects in the multiple video frames; Acquire multiple candidate centers of the target object from the multiple scaled video frames, where the candidate centers represent the central positions of the target object in the video frames; Determining a plurality of cropping frames respectively used for the plurality of scaled video frames based on a second change trend of the plurality of candidate centers along the time series; The target objects in the scaled multiple video frames are cropped based on the multiple cropping frames to generate the target video.

5. The method according to claim 4, wherein: The determining, based on the second change trend of the plurality of candidate centers along the time series, a plurality of cropping frames respectively used for the plurality of scaled video frames comprises: Correcting the plurality of candidate centers based on a second fitting curve to obtain a plurality of fitting centers, wherein the second fitting curve represents a second change trend of the plurality of candidate centers after smoothing, and the fitting parameter represents a corrected center position of the target object in the video frame; The multiple fitting centers are respectively used as the center positions of the multiple cropping frames to obtain the multiple cropping frames.

6. The method according to claim 1, further comprising: Acquire object features of the target object from the multiple video frames, wherein the object features are obtained according to at least one of target object morphological information, target object clothing information, target object equipment information, video acquisition time, video acquisition device information, and target object spatial position change trajectory; The object feature is associated with the target video and stored.

7. The method according to claim 1, wherein: The acquiring of multiple video frames containing the target object comprises: Acquire an initial frame sequence including the plurality of video frames; Evaluating the initial frame sequence based on preset rules to obtain an evaluation result, wherein the preset rules include a rule for evaluating at least one of the size relationship, the video capture duration, the spatial position information, and the target object movement speed information; When the evaluation result is passed, the plurality of video frames are acquired based on the initial frame sequence.

8. The method according to claim 1, further comprising: Respectively identifying the action types of the target object in the multiple video frames, and determining multiple target video frames that meet the preset action types; generating a target video segment based on the multiple target video frames; The target video segment is added to a preset position in the target video.

9. A video generating device, comprising: A video frame unit, configured to acquire a plurality of video frames containing a target object, wherein the plurality of video frames indicate spatial position information of the target object along a time sequence; A display parameter unit, configured to obtain a plurality of display parameters of the target object from the plurality of video frames, wherein the display parameters represent a size relationship between the target object and the video frame; A scaling factor unit configured to determine a plurality of scaling factors respectively used for the plurality of video frames based on a first variation trend of the plurality of display parameters along the time series; The video generating unit is configured to respectively perform scaling processing on the target objects in the multiple video frames based on the multiple scaling coefficients to generate a target video.

10. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the video generation method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the video generation method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the video generation method according to any one of claims 1 to 8 is implemented.