Short video generation method and system based on virtual-real fusion
By acquiring live data streams to generate virtual images and event tags, this technology solves the problems of high barriers to video production and weak interactivity in existing technologies, enabling low-cost, high-quality personalized short video generation and improving user experience and interactivity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-12
AI Technical Summary
Existing check-in technologies are monotonous and lack personalization. Video production is difficult, hardware costs are high, interactivity is weak, manual labor is inefficient, and high-quality dynamic virtual-real fusion videos cannot be generated, resulting in a poor user experience.
By acquiring live data streams from the shooting location, determining personalized content based on image and audio frames, generating virtual screen information and event tags, filtering video clips, and configuring background music and special effects, fully automated personalized short video generation is achieved.
It enables low-cost, high-quality personalized short video generation, allowing users to deeply participate in content creation through actions, props, and voice. It features low hardware costs, strong interactivity, high immersion, and rich content, making it suitable for large-scale promotion in cultural and tourism scenic spots.
Smart Images

Figure CN121603748B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video technology, and in particular to a method and system for generating short videos that blend virtual and real elements. Background Technology
[0002] With the booming development of the cultural tourism industry, new technologies for checking in have become an important way for tourists to experience and share their experiences. However, existing check-in technologies have obvious limitations:
[0003] (1) Limited format and lack of personalization: Current mainstream technologies focus on generating static photos, such as simple stickers or filters, and cannot generate dynamic, narrative short video content. The user experience is highly homogenized and lacks uniqueness.
[0004] (2) High threshold for video production and high hardware costs: To produce high-quality virtual and real fusion videos, it is usually necessary to rely on professional green screen studios, motion capture equipment and expensive post-editing software and manpower, which are difficult for ordinary tourists and small and medium-sized scenic spots to afford.
[0005] (3) Weak interactivity and passive experience: Most existing technologies are one-way shooting and watching, and users cannot interact with virtual content in real time and in depth (such as driving scene changes through actions and voice), resulting in poor immersion.
[0006] (4) Reliance on manual labor and low efficiency: Even in some high-end experiences, the creation, shooting and editing of videos still rely heavily on professionals, and cannot achieve automated and large-scale personalized content generation, and cannot meet the needs of immediacy. Summary of the Invention
[0007] This application provides a method and system for generating short videos that blend virtual and real elements, with the aim of generating low-cost, high-quality short videos of tourist shooting locations.
[0008] To achieve the above objectives, this application provides the following technical solution:
[0009] A method for generating short videos that blend virtual and real elements, comprising:
[0010] Obtain a live data stream from the shooting location; any live data in the live data stream includes image frames and audio frames;
[0011] For any given live data, determine the corresponding personalized content based on the image frame and the audio frame;
[0012] Based on the human figures shown in the image frames, and combined with the personalized content, corresponding virtual scene information and event tags are generated; the virtual scene information includes virtual scenes from multiple virtual perspectives;
[0013] For each virtual perspective, a corresponding virtual video stream is determined based on the virtual images corresponding to all real-time data throughout the shooting process;
[0014] Based on the event tags, select the corresponding video segments from each virtual video stream, and configure the corresponding background music and special effects for the video segments;
[0015] Personalized short videos are generated based on the obtained video clips.
[0016] A short video generation system that blends virtual and real elements, comprising:
[0017] The shooting data acquisition unit is used to obtain the live data stream of the shooting scene; any live data in the live data stream includes image frames and audio frames;
[0018] A personalized content determination unit is used to determine the corresponding personalized content for any given live data based on the image frame and the audio frame.
[0019] The virtual scene generation unit is used to generate corresponding virtual scene information and event tags based on the human object shown in the image frame and in combination with the personalized content; the virtual scene information includes virtual scenes from multiple virtual perspectives.
[0020] The virtual video determination unit is used to determine the corresponding virtual video stream for each virtual viewpoint based on the virtual images corresponding to all real-time data during the entire shooting process;
[0021] The video segment processing unit allows the user to filter corresponding video segments from each of the virtual video streams based on the event tags, and configure corresponding background music and special effects for the video segments;
[0022] The short video synthesis unit is used to generate personalized short videos based on the obtained video clips.
[0023] An electronic device includes: a processor, a memory, and a bus; the processor and the memory are connected via the bus.
[0024] The memory is used to store the program, and the processor is used to run the program, wherein the program is executed by the processor to perform the virtual-real fusion short video generation method.
[0025] The technical solution provided in this application obtains a live data stream from the shooting location; for any live data, based on image frames and audio frames, it determines corresponding personalized content; based on the human objects shown in the image frames, combined with the personalized content, it generates corresponding virtual screen information and event tags; for each virtual perspective, based on the virtual screens corresponding to all live data throughout the shooting process, it determines the corresponding virtual video stream; according to the event tags, it selects corresponding video segments from each virtual video stream and configures corresponding background music and special effects for the video segments; based on the obtained video segments, it generates personalized short videos. This application generates virtual video streams and event tags based on personalized content, and uses event tags to select corresponding video segments from the virtual video stream to generate personalized short videos, achieving low-cost, high-quality short video generation. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 A first flowchart illustrating a method for generating short videos that blends virtual and real elements, provided in an embodiment of this application;
[0028] Figure 2 A schematic diagram of the second process of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application;
[0029] Figure 3 A schematic diagram of the third process of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application;
[0030] Figure 4 A schematic diagram of the fourth process of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application;
[0031] Figure 5 A schematic diagram of the fifth process of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application;
[0032] Figure 6 A schematic diagram of the sixth process of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application;
[0033] Figure 7 This is a schematic diagram of the architecture of a virtual-real fusion short video generation system provided in an embodiment of this application. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0035] In this application, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0036] like Figure 1 The diagram shown is a first flowchart of a short video generation method that integrates virtual and real elements, according to an embodiment of this application, including the following steps.
[0037] S101: Obtain live data stream from the shooting location.
[0038] In this live data stream, any live data includes both image frames and audio frames.
[0039] In some examples, the live data stream includes a video stream and an audio stream, with the video stream consisting of multiple image frames and the audio stream consisting of multiple audio frames.
[0040] S102: For any live data, determine the corresponding personalized content based on image frames and audio frames.
[0041] Personalized content is used to represent the user's body language and personalized needs during the shooting process.
[0042] In some examples, body language can be determined based on visual information that can be identified from image frames, and personalized needs can be determined based on speech information that can be identified from audio frames.
[0043] Optionally, the process of determining the corresponding personalized content based on image frames and audio frames may include: identifying human object information and prop object information in the image frame; determining the corresponding user intent information based on the user's voice shown in the audio frame; and determining the corresponding personalized content based on the human object information, prop object information, and user intent information.
[0044] In some examples, user body language can be determined based on character and prop information, and user personalized needs can be determined based on user intent information.
[0045] Optionally, the person object information includes the action tags and skeletal data of the person object (specifically, the user). For the process of recognizing person object information, please refer to [link to documentation / reference]. Figure 2 The steps are shown.
[0046] Optionally, the prop object information includes the prop object and its corresponding pose. For the process of recognizing prop object information, please refer to [link / reference needed]. Figure 3 The steps are shown.
[0047] In some examples, the process of determining the corresponding user intent information based on the user's speech shown in the audio frame can be as follows: perform speech recognition on the user's speech shown in the audio frame to obtain the corresponding text information, and then perform keyword recognition on the text information to determine the corresponding user intent information.
[0048] S103: Based on the human figure shown in the image frame, combined with personalized content, generate corresponding virtual screen information and event tags.
[0049] The virtual image information includes virtual images from multiple virtual perspectives.
[0050] In some examples, multiple virtual perspectives include multiple shots from a virtual camera, such as a main camera shot, a close-up shot, a panoramic shot, etc.
[0051] Optionally, the process of generating corresponding virtual image information and event tags based on the person object shown in the image frame, combined with personalized content, can be found in [reference needed]. Figure 4 and Figure 5 The steps are shown.
[0052] In some examples, the human figures shown in the image frame, combined with personalized content and preset virtual resources (used to control the weather, effects, and other attributes of the virtual scene), can be input into a preset rendering engine to generate the corresponding virtual images.
[0053] S104: For each virtual viewpoint, determine the corresponding virtual video stream based on the virtual images corresponding to all real-time data throughout the shooting process.
[0054] Specifically, after the user finishes shooting, the virtual video stream corresponding to any virtual viewpoint is determined based on the virtual images corresponding to all real-time data under any virtual viewpoint.
[0055] In some examples, the audio playback content of the virtual video streams from different virtual perspectives is consistent.
[0056] S105: Based on the event tags, select the corresponding video segments from each virtual video stream, and configure the corresponding background music and effects for the video segments.
[0057] The timestamps of the selected video segments are different, and the video segments can be combined to generate a video stream.
[0058] Optionally, based on event tags, the corresponding video segments are selected from each virtual video stream, and the corresponding background music and effects are configured for each video segment. For details on this process, please refer to [link / reference needed]. Figure 6 The steps are shown.
[0059] S106: Generate personalized short videos based on the obtained video clips.
[0060] The individual video clips can be spliced together in order from morning to night based on their timestamps to generate corresponding personalized short videos.
[0061] In some examples, to ensure that the camera transitions in the personalized short video are aligned with the music beat and to achieve a beat-matching effect, the final video length is determined by the sum of the lengths of all video segments, but the editing rhythm is modulated by a preset beat interval.
[0062] In some examples, personalized short videos can be obtained by encoding the video clip sequence composed of individual video segments, audio tracks, and possible subtitles and effects.
[0063] The virtual-real fusion short video generation method shown in this application embodiment can achieve the following beneficial effects: (1) High personalization: Users can participate deeply in content creation through actions, props, and voice, transforming from passive viewers to active creators, and experiencing something unique; (2) Low hardware cost: The system has low hardware requirements and can be deployed on ordinary cameras and mobile devices, which is conducive to large-scale promotion in cultural and tourism scenic spots; (3) Full automation: It realizes full automation from shooting to finished film, without the need for professional photographers and editors, greatly reducing operating costs and user thresholds; (4) Extremely rich content: The generated dynamic short video integrates a variety of virtual elements and special effects, and the content expressiveness far exceeds that of static photos, making it highly valuable for sharing; (5) Strong interactivity: The multimodal interaction method makes the experience immersive and fun, which can effectively attract and retain tourists.
[0064] It should be noted that the personalization of the short videos obtained based on personalized content is mainly reflected in the following aspects: (1) Scene personalization: Virtual scenes can be switched with one click, and users can be placed in virtual worlds of different historical periods and styles; (2) Prop-driven personalization: Real props are identified and given different virtual effects; (3) Environmental status personalization: Environmental status such as weather, season, day and night can change in real time according to the user's voice command or preset plot; (4) Action trigger personalization: The user's action tags are identified and the virtual scene is triggered to generate corresponding effects; (5) Voice and rhythm personalization: Users can control the content through voice, and the generated video can automatically match the background music to enhance the viewing experience.
[0065] The process described in S101-S106 above obtains corresponding personalized content based on real-time data from the shooting location, generates a virtual video stream and event tags based on the personalized content, and uses the event tags to select corresponding video clips from the virtual video stream to generate personalized short videos, which can achieve low-cost, high-quality short video generation.
[0066] like Figure 2 The diagram shown is a second flowchart of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application, and includes the following steps.
[0067] S201: Obtain the skeletal key points of the human object shown in the current image frame.
[0068] Among them, human pose estimation algorithms (such as the Media Pipe algorithm or the Light weight OpenPose algorithm) can be used to detect skeletal key points in the current image frame and obtain the skeletal key points of the human object shown in the current image frame. The skeletal key points include the coordinates of the main joints of the human body.
[0069] S202: Determine whether the current image frame is a keyframe based on the skeletal keypoints of the previous image frame.
[0070] If the current image frame is a keyframe, then S203 is executed; if the current image frame is not a keyframe, then S204 is executed.
[0071] It should be noted that, based on the skeletal keypoints of the previous image frame, the average motion amplitude M of all skeletal keypoints between the current image frame and the previous image frame can be determined. t Therefore, based on the average motion amplitude M t This determines whether the current image frame is a keyframe. Generally speaking, if the average motion amplitude M... t If the value is below the second preset threshold, it indicates that the user's actions are stabilizing and the posture is clear. The current frame is then determined to be a keyframe, and the computationally intensive global action recognition process is initiated, i.e., step S203 is executed. If the average motion amplitude M...t If the current frame image is not lower than the second preset threshold, it is determined that the current frame image is not a key frame, and a low-overhead local tracking mode is started, i.e., S204 is executed.
[0072] In some examples, the average motion amplitude M t It quantifies the overall rate of change of user actions, specifically the average amplitude of movement M. t The calculation process can be found in formula (1):
[0073] (1).
[0074] In formula (1), K represents the total number of skeletal key points. This represents the coordinates of the k-th skeletal keypoint at time t. This represents the coordinates of the k-th skeletal keypoint at time t-1.
[0075] S203: Based on the preset indexed action library, determine the action label of the current image frame, and based on the skeletal key points of the current image frame, determine the skeletal data of the current image frame.
[0076] The indexed action library includes a product quantization index of multiple sample action labels, and the product quantization index includes a vector of the corresponding codebook.
[0077] Optionally, the process of determining the action label of the current image frame based on a preset indexed action library can be as follows: transform the coordinates of the skeletal key points of the current image frame from the image coordinate system to a coordinate system with the human hip as the origin and height as the normalized scale to obtain the corresponding features; calculate the distance between the features and multiple product quantization indices shown in the preset indexed action library; the indexed action library includes product quantization indices of multiple sample action labels; determine the action label of the current image frame based on the sample action label to which the target product quantization index with the smallest distance belongs.
[0078] It should be noted that if the current image frame is a keyframe, the coordinates of the skeletal keypoints in the current image frame are transformed from the image coordinate system to a coordinate system with the human hip as the origin and height as the normalized scale, forming a normalized feature V. t To eliminate the influence of factors such as the distance between the user and the camera, and the user's height, so that feature V... t It has scale invariance. The feature V t A fast retrieval is performed using the product quantization index shown in the indexed action library to obtain the action label for the current image frame.
[0079] In some examples, the product quantization index divides the high-dimensional feature space into multiple subspaces and builds a codebook for each subspace. During retrieval, only feature V needs to be computed. tThe distance between the prototype vectors of each codebook and the sample action label corresponding to the product quantization index with the smallest distance is obtained by looking up a table. This sample action label is then used as the action label of the current image frame, thereby reducing the processing complexity of action labels and achieving millisecond-level response.
[0080] For example, action tags can be used to represent any type of action such as waving, shooting an arrow, or jumping.
[0081] S204: Based on the action label of the previous keyframe, determine the action label of the current image frame, and based on the skeletal keypoints of the current image frame, determine the skeletal data of the current image frame.
[0082] Among them, optical flow estimation can be performed on the skeletal key points of the current image frame based on the optical flow estimation differential algorithm to obtain the skeletal data of the current image frame.
[0083] It should be noted that the optical flow estimation difference algorithm can adopt the Lucas-Kanade optical flow method. For example, the skeletal keypoints of the previous image frame are used as initial values. The Lucas-Kanade optical flow method is used to iteratively calculate in the neighborhood of the current frame image to solve for the optimal displacement (Δx, Δy) of each skeletal keypoint in the current frame image to obtain skeletal tracking data. The skeletal tracking data obtained based on the Lucas-Kanade optical flow method may have slight jitter. Therefore, prior knowledge such as the length constraint of human bones can be combined to perform micro-optimization on the skeletal tracking data to ensure that the obtained skeletal data of the current image frame is smooth.
[0084] Understandably, the Lucas-Kanade optical flow method is based on the assumption of constant brightness, only needs to process a small window in the image frame, is extremely fast, can effectively reduce computational costs, and achieve lightweight local skeleton tracking.
[0085] In some examples, the quality of the bone tracking data can also be monitored, such as calculating the feature matching degree of image patches around the bone tracking data, or checking the reasonableness of bone length, generating a confidence score C. t If the confidence score C t If the value is below the first preset threshold, it indicates that the Lucas-Kanade optical flow tracking may have failed due to occlusion, rapid movement, or other reasons. The tracking loop is immediately interrupted, and the next image frame is forcibly marked as a key frame to correct the pose.
[0086] It should be noted that, through a carefully designed decision loop, the expensive global recognition computation is used only for necessary key frames, while efficient local tracking is used on the vast majority of non-key frames. This cleverly balances computational accuracy and real-time performance, providing a reliable technical foundation for large-scale personalized interactive applications.
[0087] S205: Based on the motion tags and skeletal data of the current image frame, determine the character object information of the current image frame.
[0088] The processes described in S201-S205 above can extract motion tags and skeletal data of human objects from image frames as personalized content, enabling real-time driving of virtual scenes by user actions.
[0089] like Figure 3 The diagram shown is a third flowchart of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application, and includes the following steps.
[0090] S301: Determine whether the previous image frame contains corresponding prop object information.
[0091] If the previous image frame contains corresponding prop object information, then S302 is executed; if the previous image frame does not contain corresponding prop object information, then S303 is executed.
[0092] It should be noted that the previous image frame contains corresponding prop object information, that is, it is determined that the same prop object exists in both the previous and current image frames. Given that the prop object in the current image frame has been determined, the prop object information in the current image frame can be tracked and obtained by using the pose of the prop object in the previous image frame and combining it with a filter algorithm.
[0093] S302: Based on the prop object information of the previous image frame, combined with the filtering and tracking algorithm, determine the prop object information of the current image frame.
[0094] If the previous image frame contains corresponding prop object information, and the prop object is successfully identified, the tracking mode with less computation can be switched to maintain the continuity of the prop object's pose estimation.
[0095] In some examples, the tracking process of prop object information in the current image frame, based on prop object information from the previous image frame and combined with a filtering tracking algorithm, can be implemented as follows: A Kalman filter is used to track the pose of the prop object. Assuming the prop object is moving at a constant speed, the Kalman filter tracks the pose x of the prop object from the previous image frame. t-1 Predict the pose x of the current image frame t - And error covariance P t - The prediction process can be seen in formulas (2) and (3); in the current image frame, a small number of ORB feature points are re-extracted based on the prediction location and used as the observation value z. t Then calculate the Kalman gain K. t The observed values are then used to correct the predicted values, resulting in the optimal posterior state estimate x. tAnd update the error covariance P, based on the posterior state estimate x. t As the pose of the prop object shown in the current frame image; continuously evaluate the tracking confidence of the prop object information in the current frame image. If the tracking confidence is lower than the third preset threshold, it is determined that the tracking is lost, and immediately switch to S303 for relocalization.
[0096] (2).
[0097] (3).
[0098] In formulas (2) and (3), A represents the state transition matrix, Q represents the process noise covariance, and the process noise covariance represents the uncertainty of the prediction.
[0099] In a possible implementation, the Kalman gain K t Posterior state estimation x t The relationship between the update error covariance P can be found in formulas (4), (5) and (6).
[0100] (4).
[0101] (5).
[0102] (6).
[0103] In formulas (4), (5) and (6), H represents the observation matrix and R represents the observation noise covariance. Formulas (4), (5) and (6) integrate the mechanism of motion model prediction and visual observation, making tracking more robust to occlusion and motion blur.
[0104] It should be noted that if the prop object information of the current image frame cannot be obtained through the filtering tracking algorithm, it is determined that the prop object in the current image frame is different from that in the previous image frame, or that there is no prop object in the current image frame. In this case, S303 can be executed to identify the prop object in the current image frame.
[0105] S303: Perform target detection on the current image frame to obtain the prop region and extract the image block corresponding to the prop region.
[0106] Among them, a target detection algorithm based on a lightweight convolutional neural network can be used to detect targets in the current image frame and quickly locate all potential prop areas in the current image frame.
[0107] In some examples, the models involved in object detection algorithms using lightweight convolutional neural networks include, but are not limited to, YOLO-V5s or SSD-MobileNet.
[0108] S304: Based on image patches as input to the classification model, the corresponding prop type is obtained.
[0109] In this process, the image blocks corresponding to the prop area are input into the classification model for classification. The classification model divides the prop objects into predefined prop types, such as weapons, musical instruments, and magical artifacts, thereby significantly narrowing the search range for accurate recognition of prop objects.
[0110] S305: Determine the prop object corresponding to the current image frame based on the standard prop image corresponding to the prop type.
[0111] For the prop area of the detected prop type, a fine recognition process is initiated to determine its specific identity and initial posture. Specifically, the standard prop image corresponding to the prop type of the prop area is extracted from the corresponding prop image library and used as the prop object corresponding to the prop area.
[0112] S306: Determine the pose of the prop object based on the feature point matching results between the standard prop image and the image block.
[0113] Among them, ORB feature points can be used as feature points, and Hamming distance can be used to match feature points between standard prop images and image blocks. The matching point pairs shown in the feature point matching results are filtered using the random consistency sampling algorithm to obtain the feature point correspondence. Using the feature point correspondence and the solvePnP algorithm, the corresponding six-degree-of-freedom (6DOF) pose is calculated as the pose of the current image frame.
[0114] In some examples, ORB feature points are a type of efficient local feature point that is somewhat invariant to rotation and scale changes.
[0115] In some examples, the random consistency sampling algorithm estimates multiple matrix models by randomly sampling a small number of matching point pairs, counts the number of inliers for each matrix model, selects the target matrix model with the most inliers, and removes all outliers that do not conform to the target matrix model, thereby obtaining a set of accurate feature point correspondences.
[0116] In some examples, the solvePnP algorithm can be used to solve for the camera's rotation matrix R and translation vector t with respect to the prop object, and then calculate the corresponding pose P based on the rotation matrix R and translation vector t. init That is, P init =[R∣t].
[0117] S307: Determine the prop object information of the current image frame based on the prop object and its corresponding position.
[0118] The processes shown in S301-S307 above can extract prop objects and their corresponding poses from image frames as personalized content to achieve in-depth personalized interaction.
[0119] like Figure 4 The diagram shown is a fourth flowchart of a virtual-real fusion short video generation method provided in this application embodiment, which includes the following steps.
[0120] S401: Pre-acquire the shooting mode to be used on the shooting location.
[0121] If the shooting mode used at the shooting location is professional mode, the user will be guided to briefly sample the background before shooting. By calculating the mean and variance of the color in the sampled area, a reference benchmark will be established for the subsequent adaptive keying algorithm to ensure that it can adapt to the lighting characteristics of the specific location.
[0122] S402: If the shooting mode is professional mode, the transparency mask of the image frame is determined based on the chromaticity distance between the pixels in the image frame and the background reference color, combined with the transparency calculation function.
[0123] In the professional mode, the background of the shooting scene is a professional green screen. For this professional mode, the image frame is converted from the RGB color space to the YCbCr color space to separate the brightness and chromaticity of the image frame. Based on each pixel of the image frame in the YCbCr color space, the chromaticity distance between the pixel and the background reference color is calculated using formula (7). To avoid hard edges and retain semi-transparent details, a smooth alpha calculation function is used, combined with two dynamic thresholds (T). low and T high The transparency (alpha) value of each pixel is calculated, and finally, based on the alpha value of each pixel, the transparency mask of the image frame is determined.
[0124] (7).
[0125] In formula (7), the chromaticity distance d p Used to quantify the similarity between pixel p and the background reference color bg. The blue chromaticity component representing pixel p. The blue chromaticity component representing the background base color (bg) The red chromaticity component represents pixel p. This represents the blue chromaticity component of the background base color bg.
[0126] In some examples, the specific expression for the Alpha calculation function can be found in formula (8).
[0127] (8).
[0128] In formula (8), the Alpha calculation function can intelligently handle the shadow and highlight areas on the green screen. Based on the Alpha value of each pixel, the transparency mask of the image frame is determined to undergo a series of post-processing operations to improve quality, including using morphological opening operations to remove small noise, closing operations to fill the holes inside the foreground objects, and adaptive feathering of the foreground edges to make its integration with the virtual background more natural.
[0129] S403: If the shooting mode is normal mode, the image frame is input into the preset semantic segmentation network model to obtain the transparency mask of the image frame.
[0130] The semantic segmentation network model includes, but is not limited to, a neural network model based on an encoder-decoder architecture. Specifically, the encoder is responsible for quickly extracting multi-level features from image frames, while the decoder is responsible for upsampling and fusing the multi-level features to accurately recover the spatial details and edges of the transparency mask.
[0131] In some examples, the encoder-decoder based neural network model is optimized during the training phase by a joint loss function, as shown in Equation (9), where the joint loss function L... total Combined with cross-entropy loss L CE And Dice lost L Dice Specifically, Dice loss L Dice See formula (10) for details.
[0132] (9).
[0133] (10).
[0134] In formulas (9) and (10), L represents the cross-entropy loss CE Preset weights, Represents Dice's loss L Dice Preset weights, Representing the Predicted foreground region for each pixel. Representing the The true foreground region is represented by N pixels, where N represents the total number of pixels. The index representing the pixel.
[0135] It should be noted that the cross-entropy loss L CE To ensure accurate pixel-level classification, and L DiceDirectly optimizing the intersection-union ratio between the predicted foreground region and the real foreground region is particularly beneficial for handling scenes with an imbalance in the number of foreground and background pixels, and can effectively improve the continuity and smoothness of edge segmentation.
[0136] In a possible implementation, the semantic segmentation network model outputs a probability map, where each pixel value represents the probability that it belongs to the foreground. This probability map is then converted into a binary mask by setting a threshold (e.g., 0.5), followed by post-processing such as hole filling and edge smoothing to generate a final transparency mask that can be used for fusion.
[0137] S404: Use transparency masking to extract human figures from the background in an image frame.
[0138] The transparency mask maintains the same resolution as the image frame and precisely defines the transparency of each pixel. It is used to extract the human object in the image frame from the background, laying the foundation for a personalized virtual-real fusion experience.
[0139] S405: Based on the character object and combined with personalized content, generate corresponding virtual screen information and event tags.
[0140] After obtaining the character object, the character object is composited into the virtual scene, and combined with personalized content, the real character generated in the virtual scene is processed to obtain the corresponding virtual screen information.
[0141] The processes described in S401-S405 above can select the corresponding method to obtain the transparency mask of the image frame for different shooting modes, so as to composite the human object in the image frame into the virtual scene.
[0142] like Figure 5 The diagram shown is a fifth flowchart of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application. It includes the following steps.
[0143] S501: Generate a corresponding real-life character in the virtual scene based on the character object shown in the image frame.
[0144] In this process, a person object is extracted from an image frame. This person object is a real photograph, and a corresponding real person character is generated in the virtual scene based on this real photograph.
[0145] It should be noted that the virtual scene can be generated based on a preset virtual resource manager, which contains the virtual resources needed to create the weather and background.
[0146] S502: Based on the character object information shown in the personalized content, drive the real character to perform corresponding actions, and generate corresponding virtual effects in the virtual scene.
[0147] Based on the skeletal data of the character object, the system drives the real character in the virtual scene to perform corresponding actions and triggers the virtual scene to generate virtual effects corresponding to the action tags (such as making the virtual wand glow when the action tag is "casting a spell").
[0148] S503: Based on the prop object information shown in the personalized content, add corresponding virtual props to the real character.
[0149] In this process, virtual props corresponding to the prop objects are added to the real character characters, and the position of the virtual props is adjusted according to the position of the prop objects to achieve the desired effect, such as accurately attaching the virtual props to the real props.
[0150] S504: Based on the user intent information shown in the personalized content, determine the scene type and sound effects of the virtual scene.
[0151] S505: Generates corresponding virtual screen information based on virtual scenes, real characters, virtual special effects, and virtual props.
[0152] Among them, virtual scene information can be generated based on virtual scenes, real characters, virtual special effects and virtual props, combined with light and shadow consistency algorithms.
[0153] S506: Determine the corresponding event tag based on the person / object information and user intent information.
[0154] This involves analyzing the information of the person or object to obtain the corresponding action events, and analyzing the user's intent information to obtain the corresponding voice events. Based on the action events and voice events, corresponding event tags are generated.
[0155] The processes shown in S501-S506 above can utilize character objects and personalized content to render corresponding virtual screen information, and determine corresponding event tags based on personalized content.
[0156] like Figure 6 The diagram shown is a sixth flowchart of a short video generation method that integrates virtual and real elements, provided in an embodiment of this application. It includes the following steps.
[0157] S601: Determine the basic information of multiple interactive events based on the event labels of each image frame.
[0158] The basic information includes the timestamp of the event, the event type, and the event parameters.
[0159] In some examples, basic information about multiple interaction events can be structured into an event metadata stream, which is time-series data and can record basic information about all interaction events.
[0160] In a possible implementation, the basic information of each interactive event can be: {Occurrence timestamp: 5.3s, event type: action, action tag: jump, event parameter: 0.95}; {Occurrence timestamp: 8.1s, event type: voice, user command: sunset, event parameter: turn into evening}; {Occurrence timestamp: 12.5s, event type: prop, prop object: magic wand, event parameter: glow}.
[0161] S602: Based on the timestamp of each interactive event, obtain the set of candidate segments corresponding to each interactive event from each virtual video stream.
[0162] The candidate segment set includes candidate segments from different virtual perspectives within the same time period.
[0163] In some examples, the entire timeline is divided into several candidate segments, with important events (such as action recognition and voice triggering) as key nodes, and each candidate segment revolves around a core event.
[0164] In a possible implementation, the beat of the background music in the virtual video stream can also be analyzed, and the ideal switching point for candidate clip editing should be aligned with or multiple of the beat point as much as possible to enhance the rhythm of the video.
[0165] S603: For each interactive event, based on the preset mapping relationship between the event type and the virtual perspective, candidate segments that match the virtual perspective and the event type are obtained from the video segment set and used as the video segments corresponding to the interactive event.
[0166] Based on the preset mapping relationship between event type and virtual perspective, video clips matching the virtual perspective and event type are obtained from the video clip set, which can select the best performance form for the time period corresponding to each interactive event.
[0167] In some examples, the preset mapping relationship between event types and virtual perspectives can be represented by a rule base, which includes multiple rules used to associate event types with virtual perspectives. For example, if the event type is "action," the corresponding virtual perspective is a close-up shot to exaggerate the dynamism; if the event type is "sunset," the corresponding virtual perspective is a panoramic shot to show the environmental changes; and if the event type is "normal," the corresponding virtual perspective is a main camera shot.
[0168] S604: Based on the event parameters of the interactive event, configure the corresponding background music and special effects for the video clip.
[0169] Among them, the event parameters of interactive events can reflect the content and intensity of video clips, so that appropriate special effects can be assigned according to the event parameters. Specifically, the special effect corresponding to the event type of intense interactive events can be a fast cut, while the special effect corresponding to the event type of lyrical interactive events can be a slow overlay of environmental changes.
[0170] The processes shown in S601-S604 above can determine the basic information of multiple interactive events based on event tags, thereby selecting the corresponding video segments from each virtual video stream and configuring the corresponding background music and special effects.
[0171] like Figure 7 The diagram shown is an architectural schematic of a virtual-real fusion short video generation system provided in an embodiment of this application, including the following units.
[0172] The shooting data acquisition unit 100 is used to acquire the live data stream of the shooting scene; any live data in the live data stream includes image frames and audio frames.
[0173] The personalized content determination unit 200 is used to determine the corresponding personalized content for any given live data based on image frames and audio frames.
[0174] Optionally, the personalized content determination unit 200 is specifically used to: identify the information of human objects and prop objects in the image frame; determine the corresponding user intent information based on the user's voice shown in the audio frame; and determine the corresponding personalized content based on the information of human objects, prop objects, and user intent information.
[0175] Optionally, the personalized content determination unit 200 is specifically used for: obtaining the skeletal key points of the human object shown in the current image frame; determining whether the current image frame is a keyframe based on the skeletal key points of the previous image frame; if the current image frame is a keyframe, determining the action label of the current image frame based on a preset indexed action library, and determining the skeletal data of the current image frame based on the skeletal key points of the current image frame; if the current image frame is not a keyframe, determining the action label of the current image frame based on the action label of the previous keyframe, and determining the skeletal data of the current image frame based on the skeletal key points of the current image frame; and determining the human object information of the current image frame based on the action label and skeletal data of the current image frame.
[0176] Optionally, the personalized content determination unit 200 is specifically used to: transform the coordinates of the skeletal key points of the current image frame from the image coordinate system to a coordinate system with the human hip as the origin and height as the normalized scale to obtain the corresponding features; calculate the distance between the features and multiple product quantization indices shown in the preset indexed action library; the indexed action library includes product quantization indices of multiple sample action labels; and determine the action label of the current image frame based on the sample action label to which the target product quantization index with the smallest distance belongs.
[0177] Optionally, the personalized content determination unit 200 is specifically used for: determining whether there is corresponding prop object information in the previous image frame; if there is no corresponding prop object information in the previous image frame, performing target detection on the current image frame to obtain the prop region and extracting the image patch corresponding to the prop region; using the image patch as input to the classification model to obtain the corresponding prop type; determining the prop object corresponding to the current image frame based on the standard prop image corresponding to the prop type; determining the pose of the prop object based on the feature point matching result between the standard prop image and the image patch; determining the prop object information of the current image frame based on the prop object and the corresponding pose; if there is corresponding prop object information in the previous image frame, determining the prop object information of the current image frame based on the prop object information of the previous image frame and combined with a filtering tracking algorithm.
[0178] The virtual scene generation unit 300 is used to generate corresponding virtual scene information and event tags based on the human object shown in the image frame and combined with personalized content; the virtual scene information includes virtual scenes from multiple virtual perspectives.
[0179] Optionally, the virtual image generation unit 300 is specifically used for: pre-acquiring the shooting mode used at the shooting location; if the shooting mode is professional mode, determining the transparency mask of the image frame based on the chromaticity distance between the pixels in the image frame and the background reference color, combined with the transparency calculation function; if the shooting mode is normal mode, inputting the image frame into a preset semantic segmentation network model to obtain the transparency mask of the image frame; using the transparency mask, extracting the human object in the image frame from the background; and generating corresponding virtual image information and event tags based on the human object and personalized content.
[0180] Optionally, the virtual scene generation unit 300 is specifically used for: adding corresponding virtual props to the real character based on the prop object information shown in the personalized content; determining the scene type and sound effects of the virtual scene based on the user intent information shown in the personalized content; generating corresponding virtual scene information based on the virtual scene, real character, virtual effects and virtual props; and determining corresponding event tags based on the character object information and user intent information.
[0181] The virtual video determination unit 400 is used to determine the corresponding virtual video stream for each virtual viewpoint based on the virtual images corresponding to all real-time data during the entire shooting process.
[0182] The video segment processing unit 500 allows users to filter corresponding video segments from various virtual video streams based on event tags, and configure corresponding background music and special effects for the video segments.
[0183] Optionally, the video segment processing unit 500 is specifically used for: determining basic information of multiple interactive events based on the event tags of each image frame; the basic information includes the occurrence timestamp, event type, and event parameters; obtaining a candidate segment set corresponding to each interactive event from each virtual video stream based on the occurrence timestamp of each interactive event; the candidate segment set includes candidate segments from different virtual perspectives within the same time period; for each interactive event, obtaining candidate segments matching the virtual perspective and event type from the video segment set based on the preset mapping relationship between the event type and the virtual perspective, as the video segment corresponding to the interactive event; and configuring corresponding background music and special effects for the video segment based on the event parameters of the interactive event.
[0184] The short video synthesis unit 600 is used to generate personalized short videos based on the obtained video clips.
[0185] The unit described above acquires corresponding personalized content based on real-time data from the shooting location, generates a virtual video stream and event tags based on the personalized content, and uses the event tags to select corresponding video segments from the virtual video stream to generate personalized short videos, thus achieving low-cost, high-quality short video generation.
[0186] This application also provides a computer-readable storage medium including a stored program, wherein the program executes the virtual-real fusion short video generation method provided in this application.
[0187] This application also provides an electronic device, including a processor, a memory, and a bus. The processor and the memory are connected via the bus. The memory is used to store a program, and the processor is used to run the program. When the program runs, it executes the virtual-real fusion short video generation method provided in this application.
[0188] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0189] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A method for generating short videos that blends virtual and real elements, characterized in that, include: Obtain live data stream from the filming location; Any live data in the live data stream includes image frames and audio frames; For any given live data, identify the information of human objects and prop objects in the image frame; Based on the user's voice shown in the audio frame, determine the corresponding user intent information; Based on the character object information, the prop object information, and the user intent information, the corresponding personalized content is determined; Based on the human figures shown in the image frames, and combined with the personalized content, corresponding virtual scene information and event tags are generated; the virtual scene information includes virtual scenes from multiple virtual perspectives; Multiple virtual perspectives include multiple shooting lenses of a virtual camera, including a main camera lens, a close-up lens, and a panoramic lens; For each virtual perspective, a corresponding virtual video stream is determined based on the virtual images corresponding to all real-time data throughout the shooting process; wherein, the audio playback content of the virtual video streams of each virtual perspective is consistent. Based on the event tags, corresponding video segments are selected from each of the virtual video streams, and corresponding background music and special effects are configured for each video segment; wherein, the timestamps of each selected video segment are different, and the video segments can be combined to generate a video stream; Based on the obtained video clips, generate personalized short videos; The process of generating corresponding virtual screen information and event tags based on the person object shown in the image frame, combined with the personalized content, includes: Based on the human figures shown in the image frames, corresponding real-life human characters are generated in the virtual scene. Based on the character object information shown in the personalized content, the real character is driven to perform corresponding actions, and the virtual scene generates corresponding virtual effects. Based on the prop object information shown in the personalized content, corresponding virtual props are added to the real character. Based on the user intent information shown in the personalized content, the scene type and sound effects of the virtual scene are determined; Based on the virtual scene, the real characters, the virtual effects, and the virtual props, corresponding virtual screen information is generated; Based on the person object information and the user intent information, the corresponding event tag is determined.
2. The method according to claim 1, characterized in that, The process of identifying human object information in the image frame includes: Obtain the skeletal key points of the human object shown in the current image frame; Based on the skeletal key points of the previous image frame, determine whether the current image frame is a key frame; If the current image frame is the keyframe, the action label of the current image frame is determined based on the preset indexed action library, and the skeletal data of the current image frame is determined based on the skeletal key points of the current image frame. If the current image frame is not the keyframe, determine the action label of the current image frame based on the action label of the previous keyframe, and determine the skeletal data of the current image frame based on the skeletal keypoints of the current image frame. Based on the motion tags and skeletal data of the current image frame, the human object information of the current image frame is determined.
3. The method according to claim 2, characterized in that, Based on a pre-defined indexed action library, the action tag of the current image frame is determined, including: The coordinates of the skeletal key points of the current image frame are transformed from the image coordinate system to a coordinate system with the human hip as the origin and height as the normalized scale, in order to obtain the corresponding features. Calculate the distance between the feature and multiple product quantization indices shown in the preset indexed action library; the indexed action library includes product quantization indices of multiple sample action labels; The action label of the current image frame is determined based on the sample action label to which the target product quantization index with the smallest distance belongs.
4. The method according to claim 1, characterized in that, The process of identifying prop object information in the image frame includes: Determine if the previous image frame contains corresponding prop object information; If the previous image frame does not contain corresponding prop object information, target detection is performed on the current image frame to obtain the prop region, and the image block corresponding to the prop region is extracted. The corresponding prop type is obtained by using the image patch as input to the classification model. Based on the standard prop image corresponding to the prop type, determine the prop object corresponding to the current image frame; Based on the feature point matching results between the standard prop image and the image block, the pose corresponding to the prop object is determined; Based on the prop object and its corresponding pose, determine the prop object information of the current image frame; If the previous image frame contains corresponding prop object information, the prop object information of the current image frame is determined based on the prop object information of the previous image frame and combined with a filtering tracking algorithm.
5. The method according to claim 1, characterized in that, Based on the human figure shown in the image frame, and combined with the personalized content, corresponding virtual screen information and event tags are generated, including: Obtain the shooting mode to be used on set in advance; If the shooting mode is professional mode, the transparency mask of the image frame is determined based on the chromaticity distance between the pixels in the image frame and the background reference color, combined with the transparency calculation function. If the shooting mode is normal mode, the image frame is input into a preset semantic segmentation network model to obtain the transparency mask of the image frame; Using the transparency mask, the human figure in the image frame is extracted from the background; Based on the character object and the personalized content, corresponding virtual screen information and event tags are generated.
6. The method according to claim 1, characterized in that, Based on the event tags, corresponding video segments are selected from each virtual video stream, and corresponding background music and effects are configured for the video segments, including: Based on the event tags of each of the image frames, basic information of multiple interactive events is determined; the basic information includes the occurrence timestamp, event type, and event parameters. Based on the timestamps of each of the interactive events, a set of candidate segments corresponding to each interactive event is obtained from each virtual video stream; the set of candidate segments includes candidate segments from different virtual perspectives within the same time period. For each interactive event, based on the preset mapping relationship between event type and virtual perspective, candidate segments that match the virtual perspective and event type are obtained from the video segment set and used as the video segments corresponding to the interactive event; Based on the event parameters of the interactive event, the corresponding background music and special effects are configured for the video clip.
7. A short video generation system that blends virtual and real elements, characterized in that, include: The shooting data acquisition unit is used to obtain the real-time data stream from the shooting location; Any live data in the live data stream includes image frames and audio frames; A personalized content determination unit is used to determine the corresponding personalized content for any given live data based on the image frame and the audio frame. The virtual scene generation unit is used to generate corresponding virtual scene information and event tags based on the human object shown in the image frame and in combination with the personalized content; the virtual scene information includes virtual scenes from multiple virtual perspectives. Multiple virtual perspectives include multiple shooting lenses of a virtual camera, including a main camera lens, a close-up lens, and a panoramic lens; The virtual video determination unit is used to determine the corresponding virtual video stream for each virtual viewpoint based on the virtual images corresponding to all real-time data during the entire shooting process; wherein the audio playback content of the virtual video streams of each virtual viewpoint is consistent. The video segment processing unit allows the user to filter corresponding video segments from various virtual video streams based on the event tags, and configure corresponding background music and special effects for the video segments; wherein the timestamps of the filtered video segments are different, and the video segments can be combined to generate a video stream; The short video synthesis unit is used to generate personalized short videos based on the obtained video clips; The personalized content determination unit is specifically used for: identifying character object information and prop object information in the image frame; determining the corresponding user intent information based on the user's voice shown in the audio frame; and determining the corresponding personalized content based on the character object information, prop object information, and user intent information. The virtual scene generation unit is specifically configured to: generate a corresponding real-life character in a virtual scene based on the character object shown in the image frame; drive the real-life character to perform corresponding actions based on the character object information shown in the personalized content, and generate corresponding virtual effects in the virtual scene; add corresponding virtual props to the real-life character based on the prop object information shown in the personalized content; determine the scene type and sound effects of the virtual scene based on the user intent information shown in the personalized content; generate corresponding virtual scene information based on the virtual scene, the real-life character, the virtual effects, and the virtual props; and determine corresponding event tags based on the character object information and the user intent information.
8. An electronic device, characterized in that, include: Processor, memory, and bus; The processor and the memory are connected via the bus; The memory is used to store the program, and the processor is used to run the program, wherein the program is executed by the processor to perform the short video generation method of virtual-real fusion according to any one of claims 1-6.