Digital human material generation and sharing method and device based on intelligent hardware
Through deep learning and green screen technology, high-quality digital human materials with transparent layers are generated. Combined with user input, explanation videos are generated and abnormal frames are repaired, achieving efficient multi-platform publishing and solving the technical difficulties in the generation and sharing of digital human materials.
Patent Information
- Application Number
- CN202511109959.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-08
AI Technical Summary
In the process of generating digital human materials, there are problems such as inaccurate background removal, insufficient consistency in expressions and movements, difficulty in adapting light and shadow, low video generation efficiency, limited video editing functions, cumbersome and inefficient multi-platform publishing operations, and insufficient multi-tasking capabilities.
Background removal is performed through deep learning portrait segmentation algorithm and green screen chroma keying technology, and digital humans with transparent layers are generated by combining the digital human template library and expression feature points; user input of text and instructions are received, explanation videos are generated and abnormal frames are repaired; videos are edited twice and multi-platform accounts are managed, and video splicing algorithms are used to generate adapted content and publish them simultaneously.
It improves the naturalness and realism of digital human materials, improves the efficiency and quality of video generation, supports multi-task parallel processing, enhances material reuse rate and publishing efficiency, and breaks through the limitations of manual operation.
Smart Images

Figure CN120640079B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and device for generating and sharing digital human materials based on intelligent hardware. Background Art
[0002] Currently, the process of generating digital human material presents numerous technical pain points. Traditional methods often struggle to accurately remove complex backgrounds during digital human creation, resulting in a low degree of integration between the digital human and the subsequent scene. Furthermore, the digital human's expressions and movements lack coherence, making sudden changes in expression and discontinuities in movement more likely to occur, impacting the naturalness and realism of the digital human. Furthermore, ambient lighting significantly impacts the presentation of the digital human, and existing technologies lack effective light and shadow adaptation mechanisms, making it difficult for the digital human to effectively integrate into diverse real-world environments.
[0003] When it comes to generating digital human scene videos, user-entered text, background information, and various commands are difficult to efficiently integrate with the digital human. Most technologies are unable to automatically generate matching digital human action sequences and expression change logic based on the text semantics, requiring extensive manual adjustments. This not only increases user operation costs but also reduces video generation efficiency. Furthermore, the generated videos often contain abnormal video frames such as blur, color cast, and frame skipping. Existing abnormal frame processing technologies are ineffective in repairing these abnormal frames, making it difficult to guarantee the overall quality of the video.
[0004] When it comes to video editing and multi-platform sharing, traditional technologies offer limited secondary editing capabilities for initial videos, making it difficult to meet personalized user needs such as adding text, special effects, picture-in-picture, and adjusting video order, resulting in low material reuse rates. Furthermore, during multi-platform publishing, due to differences in video formats, copywriting styles, and publishing rules across platforms, existing technologies lack unified account management and content adaptation mechanisms, requiring users to perform separate operations for each platform. This is not only cumbersome but also prone to poor compatibility between published content and the platform, impacting content dissemination effectiveness. Furthermore, insufficient multi-tasking parallel processing capabilities also restrict the overall efficiency of digital human material generation and sharing. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] A method for generating and sharing digital human materials based on intelligent hardware is applied to intelligent hardware, comprising: obtaining target video and image information, processing the target video and image information, and generating a digital human with a transparent background layer; receiving user input of text, background information, and workflow instructions including voice, keystrokes, and automation, processing the generated digital human, and generating a digital human scene video that explains the text; processing abnormal video frames in the digital human scene video to generate an initial digital human scene video; performing secondary editing on the initial digital human scene video, and generating an updated digital human scene video based on the user's operating requirements and the preset video selection order; processing the selection order information of multiple videos and the preset publishing rule information and multi-platform account information to generate video concatenation sequence information, platform account configuration information, release copy title information, and release timing planning information; use video splicing algorithm to concatenate videos in sequence, configure platform information with unified account management module, and generate custom copy titles with content adaptation tools to generate concatenation distribution results and platform adaptability evaluation information; based on the concatenation distribution results and platform adaptability evaluation information, verify and adjust the video concatenation content and release parameters, optimize the release parameters for platforms with adaptation exceptions, and generate an optimized concatenation distribution plan; integrate and trigger the optimized concatenation distribution plan, respond to manual, voice or automated workflow instructions, and generate release results that are distributed synchronously to multiple platforms.
[0007] A device for generating and sharing digital human materials based on intelligent hardware, comprising: an acquisition module for acquiring target video and image information, processing the target video and image information, and generating a digital human with a transparent background layer; a processing module for receiving user input of text, background information, and workflow instructions including voice, keystrokes, and automation, processing the generated digital human, and generating a digital human scene video explaining the text; processing abnormal video frames in the digital human scene video to generate an initial digital human scene video; performing secondary editing on the initial digital human scene video, generating an updated digital human scene video based on the user's operating requirements and the preset video selection order; and processing the selection order information and preset release of multiple videos. Rule information and multi-platform account information are extracted and classified to generate video concatenation sequence information, platform account configuration information, release copy title information, and release timing planning information; a video splicing algorithm is used to concatenate videos in sequence, a unified account management module is used to configure platform information, and a content adaptation tool is used to generate custom copy titles, and concatenation distribution results and platform adaptability evaluation information are generated; based on the concatenation distribution results and platform adaptability evaluation information, the video concatenation content and release parameters are verified and adjusted, and release parameters are optimized for platforms with adaptation exceptions to generate an optimized concatenation distribution plan; the optimized concatenation distribution plan is integrated and triggered to respond to manual, voice, or automated workflow instructions to generate release results that are distributed synchronously to multiple platforms.
[0008] Its beneficial effects are: the present invention provides a method for generating and sharing digital human materials based on intelligent hardware. After abnormal frames are repaired and secondary editing, videos are serialized in sequence, multi-platform accounts are configured to generate adapted content, and after verification and optimization, multiple instructions are responded to and distributed synchronously. It supports multi-tasking in parallel, improves material reuse rate and publishing efficiency, and breaks through manual limitations. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 A flowchart of a method for generating and sharing digital human materials based on intelligent hardware provided by an embodiment of the present invention when applied to intelligent hardware;
[0010] Figure 2 A schematic diagram of a module of a digital human material generation and sharing device based on intelligent hardware provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0011] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention. Figure 1 To describe the method for generating and sharing digital human materials based on intelligent hardware according to an exemplary embodiment of the present application.
[0012] In the embodiment of the present application, a method for generating and sharing digital human materials based on intelligent hardware is applied to intelligent hardware, such as Figure 1 As shown:
[0013] S101, obtaining target video and image information, processing the target video and image information, and generating a digital human with a transparent background layer.
[0014] In one embodiment, real-time captured video, uploaded video, or one or more photos are obtained as target video and image information. Users can use the built-in camera of Yunbao smart hardware to shoot a video of themselves explaining a product in real time, upload a previously recorded lecture video from their local album, or upload three personal photos from different angles. These contents together constitute the target video and image information.
[0015] A deep learning-based portrait segmentation algorithm is used to process the target video and image information frame by frame. Through pixel-level semantic segmentation, the model extracts the portrait's outline, facial features, and motion trajectory information. Green screen chroma keying or background separation based on RGB-D depth information is also used to accurately remove green screens or complex backgrounds, ensuring a natural transition between portrait edges. Taking a user-uploaded speech video with a green screen as an example, a deep learning portrait segmentation algorithm based on an improved U-Net architecture is used to process the video frame by frame. This model implements pixel-level semantic segmentation through an encoder-decoder structure. The encoder uses convolutional layers to extract low-level features (such as edges and textures) and high-level features (such as the overall portrait outline) from the video frame. The decoder uses upsampling to restore the feature maps to the original image size and integrates skip connections to fuse feature information from different levels. This allows the model to accurately locate and extract the user's full outline, 48 facial features when the mouth corners are raised (including key coordinates on both sides of the mouth corners and features in the apple cheek area), and the skeletal joint motion trajectories during arm swinging (such as the spatial coordinates of the shoulder, elbow, and wrist joints in each frame).
[0016] At the same time, green screen chroma keying technology is used for background separation. By presetting the HSV threshold range of the green screen color (for example, hue range 45-75°, saturation range 50%-100%, and brightness range 40%-90%), the video frame is converted into color space and threshold segmented to generate a background mask. For the transition area between the green screen and the edge of the character, morphological operations (such as corrosion and dilation) and Gaussian blur processing are introduced to eliminate edge jaggedness, and then the transparency of the edge pixels is optimized through the alpha blending algorithm to make the transition between the edge of the character and the transparent background smooth and natural, ensuring that there is no obvious color cast or residual green screen color blocks on the edge of the portrait after removal.
[0017] Combining the preset body posture templates, clothing templates, and action libraries in the digital human template library, the portrait information after background removal is matched with the template for feature fusion. The template's facial animation is driven based on expression feature points, and the template's body movements are mapped based on the motion trajectory information. This creates an initial digital human with a transparent layer, supporting retrieval and call-up of template libraries by style and scene classification. The digital human template library stores business-style standing posture templates (such as static posture parameters of feet shoulder-width apart, hands hanging naturally or crossed in front of the body), dark suit clothing templates (including texture mapping data for details such as collars, cuffs, and buttons), and a speech action library (such as a collection of dynamic parameters such as the range of hand gesture angles and arm lift height). The specific processing process is as follows:
[0018] First, the SIFT (Scale-Invariant Feature Transform) feature matching algorithm is used to deeply compare the user portrait information after background removal with the business standing posture template: 18 key posture feature points such as the trunk spine key points (such as the coordinates of the connection between the cervical, thoracic and lumbar vertebrae), shoulder endpoints, and hip vertices in the user portrait are extracted, and parameters such as the trunk tilt angle (the deflection angle based on the vertical central axis, accurate to 0.1°) and head orientation (based on the three-dimensional vector direction of the line connecting the nasal root and the mandibular point) are calculated; these parameters are compared with the standard parameters of the business standing posture template (such as a trunk tilt angle of 0°, head facing straight ahead, with an allowable error range of ±2°) to calculate the mean square error (MSE). When the error exceeds the threshold, the user portrait is rotated, translated, and scaled using affine transformation. The matching degree between the corrected portrait posture and the template benchmark must reach more than 95% to ensure that the standing posture meets the template specifications.
[0019] When calling the clothing fusion engine, a CNN-based clothing segmentation model is used to first locate the user's upper body area (the pixel range from the shoulder line to the waistline). Then, UV mapping technology is used to accurately project the texture data of the suit clothing template (a PBR material map with a resolution of 4096×4096 pixels, including diffuse reflection, roughness, and metalness channels) onto the three-dimensional mesh surface of the user's upper body. For the splicing edges between the clothing and the human body contour (such as the collar and neck, cuffs and wrists), a Gaussian blur algorithm is used to feather the edges (blur radius 3-5 pixels) to eliminate hard-edge splicing traces. At the same time, a feature extraction algorithm is used to obtain the user's body parameters such as shoulder width (the distance between the left and right acromion points) and chest circumference (the pixel value of the maximum horizontal circumference of the chest). Based on these parameters, the mesh vertices of the clothing template are non-uniformly scaled (scaling factor = user parameter / template standard parameter) to ensure that the shoulder line and sleeve length of the suit are adapted to the user's body shape, and the fit error is controlled within ±5 pixels.
[0020] Based on the extracted expression feature point data such as the upturned corners of the mouth (change in y-coordinate of the corner feature point ≥ 5 pixels) and the raised apple muscle, the vertex animation system of the template face is driven, the smile expression weight is linearly transitioned from 0 to 0.8 (full weight is 1) through the BlendShape technology, the eye squinting degree is synchronously adjusted (eyelid feature point closure reaches 30%), and a natural smile animation sequence is generated; for the motion trajectory of arm swinging, the mapping relationship between the user's arm joints (shoulder, elbow, wrist) and the template limb skeleton is established through the skeletal binding technology, the joint rotation angle data of the user in each frame is converted into the motion parameters of the template limb through the inverse kinematics (IK) algorithm, so that the swinging amplitude and speed of the template arm are consistent with the user's action (error is controlled within ±3°); finally, the posture, costume, expression and action data are integrated to generate an initial digital human file with a transparent channel (format is.fbx or.glb), and the template library built-in label retrieval system is used to associate and store the template with scene labels such as "business" and "speech", so that the user can trigger fuzzy retrieval by inputting the label keyword, quickly locate and call the template.
[0021] The initial digital human is subjected to inter-frame interpolation processing, the motion trend of the adjacent frames of the human image is predicted through the optical flow estimation algorithm, the transition frame pictures are supplemented, and the time domain filtering algorithm is used to smooth the expression change curve, so as to avoid expression mutation. The Farneback dense optical flow algorithm is used to process adjacent frames (set as the t-th frame and the t+2-th frame), the displacement vectors (including the horizontal direction dx and the vertical direction dy) of each pixel in the arm region of the two frames are calculated, and the optical flow field model is constructed. Based on the model, the motion state of the arm in the intermediate transition frame (the t+1-th frame) is predicted. Specifically, the pixel coordinates (x, y) of the arm region in the t-th frame are linearly transformed according to the optical flow vector (dx, dy) to obtain the predicted pixel position of the t+1-th frame, and then the pixel value is filled through the bilinear interpolation algorithm to generate the transition frame picture. For example, when the swinging angle of the arm changes by 30° from the t-th frame to the t+2-th frame, the supplemented transition frame will decompose the angle change into 15° / frame, and eliminate the motion discontinuity (discontinuity error is controlled to be ≤2°).
[0022] A sliding window mean filter algorithm (with a window size of 5 frames) was used to smooth the expression curve. For example, when an expression changes from smiling to serious, the intensity of the expression in each frame was first quantized into an eigenvalue (smiling set to 1.0, serious set to 0.0). When the eigenvalue of frame t is 0.9 and suddenly drops to 0.2 in frame t+1, the algorithm averages the eigenvalues within the window (frames t-2 to t+2), reallocating the eigenvalues of frame t+1 to 0.6 and frame t+2 to 0.3. This reduces the expression intensity by ≤0.3 per frame, avoiding sudden changes (the amplitude of the sudden change is controlled to ≤0.4 per frame). This process improves the coherence of the initial digital human's movements when waving its arms rapidly, and enhances the smoothness of the expression curve, making it more consistent with natural motion.
[0023] Based on the real scene lighting parameters collected by the ambient light sensing module, the brightness of the highlight area and the contrast of the shadow area of the digital human are dynamically adjusted through HSV color space conversion technology. The super-resolution reconstruction algorithm is used to optimize the facial expression details and the clarity of the portrait edge. Combined with the skin color adaptive adjustment model to match the real scene color temperature, a digital human with a transparent background layer with high integration and excellent naturalness is generated. The ambient light sensing module collects the lighting parameters of the real scene through a light sensor (such as BH1750), and obtains a color temperature value of 3000K (warm light range) and a light intensity of 500lux (medium intensity). Based on this, the following processing is performed: First, the digital human image is converted from RGB color space to HSV space (the conversion formula is ). Furthermore, the 2R-GB term in the formula is the numerator when calculating the red channel's dominance. However, if green or blue is at its maximum, the corresponding offset (+120°, +240°) must be added to cover the full 360° hue circle. The brightness channel (V channel) and hue channel (H channel) are separated. For highlight areas such as the forehead (located using facial landmarks, e.g., the forehead is a rectangular area from the upper edge of the brow bone to the hairline), the pixel values in the V channel are increased by 10% (original V value × 1.1), while the H and S channel values remain unchanged. For the shadowed areas on both sides of the cheeks (from the outer cheekbones to the jawline), a gamma correction algorithm (γ set to 0.9) is used to increase contrast, increasing the difference in dark pixel values by 5% (i.e., the standard deviation of dark pixel values increases from 0.12 to 0.126). After processing, the image is converted back from HSV space to RGB space to complete light and shadow adaptation.
[0024] The digital human facial images were processed using the Enhanced Deep Super-Resolution (EDSR) algorithm. This algorithm extracts detailed features such as facial wrinkles using eight layers of residual blocks. Each residual block consists of two 3×3 convolutional layers and one skip connection. By learning the mapping relationship between low-resolution images and high-resolution images, the facial image resolution is increased from 512×512 to 1024×1024. Specifically, a sub-pixel convolutional layer (Pixel Shuffle) is used to reconstruct high-frequency details at the edges of wrinkles, increasing the gradient value of the wrinkle edge from 20 to 35 (a higher gradient value indicates a sharper edge), improving detail clarity by ≥30%. Skin color samples under real warm light conditions were clustered using a skin color clustering algorithm (K-means clustering, K=5), resulting in the mean RGB value of warm light skin (R=250, G=180, B=150). The digital human's facial skin color pixels are converted using a color mapping matrix (M=[[0.95,0.03,0.02],[0.02,0.90,0.08],[0.01,0.05,0.94]]) so that the Euclidean distance between the RGB value of the digital human's skin color and the mean of the warm light sample is ≤10 (the smaller the distance, the higher the matching degree), ultimately achieving a natural fusion of the skin color and the warm light environment.
[0025] S102, receiving the copy, background information and workflow instructions including voice, keystrokes and automation input by the user, processing them in combination with the generated digital human, and generating a digital human scene video explaining the copy.
[0026] In one embodiment, the system receives text input by the user through a dialog box, background information manually selected or automatically matched based on text keywords, as well as voice commands, button trigger signals, and automated workflow instructions. In the interactive interface of the Yunbao device, the user completes the text input operation through the exclusive dialog box of the digital human avatar, and the input content is "Welcome to learn about the three core functions of the new smart watch." The dialog box supports real-time preview and editing of text and can automatically save input history (for 7 days). In the background selection link, the user manually clicks the "Technology Exhibition Hall" option through the scene classification list in the interface (including "Technology Exhibition Hall", "Business Meeting Room", "Outdoor Scene" and other categories). The scene data corresponding to the background (including 3D models, light and shadow parameters, dynamic element configuration) will be preloaded into the local cache (the cache capacity is capped at 100MB).
[0027] The triggering manner of the task supports three parallel mechanisms: voice instruction triggering: the user issues a "start generating video" voice instruction through the built-in microphone of the cloud treasure device, the device calls the MFCC (Mel Frequency Cepstrum Coefficient) based voice recognition model for real-time analysis, converts the voice signal into a text instruction, and the matching success rate needs to be ≥95%, and the recognition response time is ≤1.5 seconds; entity button triggering: the user presses the customized function button (key stroke 2mm, triggering pressure 50±5g) on the side of the device, the hardware interface notifies the host chip through an interrupt signal, and directly enables the generation task (button triggering response time ≤0.5 seconds);
[0028] Automatic workflow instruction triggering: the user pre-configures the "automatically generate product video at 10 o'clock every day" workflow rule in the task scheduling center of the cloud treasure device, which includes triggering time (accurate to seconds), associated digital human clone ID, default background parameters and other information, and the system automatically calls the generation interface at the set time through the timing task scheduler (based on CRON expression) without human intervention. The three triggering methods will generate a unique task ID (format: UUID) in the task manager of the cloud treasure device, and synchronously record the triggering time, triggering method and associated script and background information, for subsequent task tracking and management.
[0029] A natural language processing model is called to perform semantic analysis on the script, extract core explanation points, and generate digital human action sequences and expression change logic, while the background information is analyzed into scene rendering parameters. When the BERT model is used to perform semantic analysis on the script "Welcome to learn about the three core functions of the new smart watch", the script is first disassembled into "welcome" "learn" "new" "smart watch" "three" "core function" word vectors by a word segmentation tool, and then input into the BERT pre-training model (using a 12-layer Transformer structure with a hidden layer dimension of 768) for context feature extraction. Through the attention mechanism, "smart watch" (entity word weight ratio 35%) and "three core functions" (core phrase weight ratio 40%) are located as key information, and core point labels are generated.
[0030] Based on the core points, the action generation module calls the preset speech action library: for the pointing requirement of "smart watch", a right hand lifting action sequence is generated, with specific parameters of shoulder joint rotation 30°, elbow joint bending 90°, wrist joint internal rotation 15°, hand drop point coordinates corresponding to the left 1 / 3 area of the picture (x=320 pixels, y=540 pixels), and action duration 1.5 seconds; for the introduction of "three core functions", the nodding action logic is configured, and the nodding is triggered once for each function (neck forward 10°, lasting 0.5 seconds), and the action interval is synchronized with the sentence rhythm of the script (interval 2-3 seconds).
[0031] The expression change logic is driven by facial feature points: in the initial smiling state, the corner of the mouth feature points (coordinates (420,680), (580,680)) are offset upward by 8 pixels, and the feature points in the apple muscle area are raised by 5 pixels; when introducing the function, the pupil feature points ((460,520), (540,520)) converge toward the center of the screen (offset ≤ 3 pixels), and the closure degree of the eyelid feature points is maintained at 20% (simulating a state of concentration). The expression parameters are updated in real time with the semantics of the copy (update frequency 30 times / second).
[0032] During the analysis of the "Technology Exhibition Hall" background, the scene parameter extraction module reads the background metadata: the resolution is fixed at 1920×1080 (in line with the standard 16:9 ratio), the lighting parameters are set to an ambient light intensity of 800 lux (warm white light, color temperature 4500K), the main light source direction is 45° from the upper left of the screen, the background blur uses a Gaussian blur algorithm (blur radius 10 pixels), the focus area is the central 2 / 3 of the screen (to ensure the clarity of the digital human body), and dynamic elements are loaded at the same time (such as floating technology icons, which appear randomly 1-2 times per second). All parameters are encapsulated in a JSON-formatted scene rendering configuration file for the animation synthesis engine to call.
[0033] The generated digital human action sequence, expression logic and scene rendering parameters are input into the animation synthesis engine, combined with the created digital human with transparent layer for frame synchronization synthesis, and multi-threaded parallel processing technology is used to support the simultaneous execution of one or more generation tasks. The fusion processing flow of the animation synthesis engine (using the Unity animation system as an example) is as follows: First, the digital human action sequence (frame rate 30fps) is loaded through the Animation component. The keyframe data (including joint rotation angles, displacement coordinates, etc.) for actions such as raising the right hand and nodding is bound to the timeline, so that the actions are executed in the scene according to a preset rhythm (for example, raising the right hand for 1.5 seconds, and nodding for 2-3 seconds). MorphTargets technology is used to drive frame-level facial vertex animation, converting expression logic such as raised mouth corners and focused pupils into displacement parameters of facial mesh vertices (updating more than 500 vertex coordinates per frame), ensuring synchronization between expression and action (time error ≤ 0.03 seconds). Based on the "Technology Exhibition Hall" scene parameters, the Light component is configured with ambient light intensity of 800 lux and warm white light of 4500K color temperature. The main light source is projected at a 45° angle to the upper left. The Camera component is used to set background blur (Gaussian blur radius 10 pixels). As the digital human moves along the action trajectory in the scene, the lighting and shadow effects adapt in real time to the position (for example, the brightness of the highlight area increases by 15% when close to the light source).
[0034] Based on the OpenMP multi-threading framework, the system allocates independent threads (thread IDs Thread_001 and Thread_002) to the two tasks of "Smartwatch Function Introduction" and "Mobile Phone New Product Introduction", respectively. Each thread occupies 2GB of independent memory space and is responsible for the frame rendering calculation of its own task. The shared memory pool (capacity 8GB) is divided into a read-only area (storing digital human templates and basic scene resources) and a read-write area (caching temporary rendering data). The thread access rights are controlled by a mutex mechanism. For example, when Thread_001 reads the digital human texture resources, Thread_002 enters a waiting state to avoid resource competition. The threads transmit synchronization signals through message queues. When a task completes keyframe rendering, it sends a status notification (such as "100th frame synthesis completed") to the other thread, ensuring that the generation progress deviation of the two tasks is controlled within a range of ≤5 frames, ultimately achieving conflict-free parallel execution.
[0035] An inter-frame smoothing algorithm was used to optimize the coherence of the digital human's movements, and alpha channel blending technology was used to achieve a natural superposition of the digital human and the background, generating the initial digital human scene video. To optimize the digital human's arm-waving movements, a third-order Bezier curve interpolation algorithm was used to achieve smooth inter-frame transitions. First, the key frame angles of the arm-waving movement (e.g., a 10° shoulder rotation in frame 10 and a 30° rotation in frame 15) were used as curve control points. The angle parameters of the intermediate frames were calculated using the formula B(t)=P0(1-t)³+3P1(1-t)²t+3P2(1-t)t²+P3t² (t∈[0,1]). This resulted in the angles of frames 11-14 being 14°, 19°, 24°, and 28°, respectively. The angle differences between adjacent frames were 4°, 5°, 5°, and 2°, respectively, all within 5°, ensuring motion coherence (reducing the stuttering rate to below 0.1%).
[0036] The overlay processing of the digital human and the background is implemented based on the Alpha Channel: the edge area (3 pixels wide) of the digital human's transparent layer is divided into a transition zone, and the transparency value from the outer layer to the inner layer increases linearly from 0% to 100% (for example, the transparency of the first pixel is 20%, the second pixel is 50%, and the third pixel is 80%). GPU hybrid rendering technology is used to fuse this layer with the corresponding area of the "Technology Exhibition Hall" background at the pixel level. The fusion formula is: output pixel value = digital human pixel value × transparency + background pixel value × (1-transparency), eliminating hard edges and color casts at the splicing point (edge error is controlled at ΔE≤2).
[0037] The resulting initial digital human scene video uses the H.264 encoding format, with a duration of 60 seconds (corresponding to the length of the presentation), a frame rate of 30fps (each frame contains 307,200 pixels), a bitrate of 8Mbps, and a resolution of 1920×1080 (16:9 ratio). The smoothness of the digital human's movements (assessed by the amount of motion blur pixels) in the video is 40% higher than before optimization, and the integration of the digital human with the background of the technology exhibition hall (assessed by the naturalness of edge transitions) reaches 95 points (out of a maximum of 100), meeting the basic quality requirements for subsequent editing and publication.
[0038] S103: Process the abnormal video frames in the digital human scene video to generate an initial digital human scene video.
[0039] In one implementation, the digital human scene video is scanned frame by frame to extract frame clarity, color consistency, and motion continuity features. Taking the "Smartwatch Function Introduction" digital human scene video (60 seconds, 30fps, 1800 frames in total) as an example, the system uses a frame-by-frame scanning module to extract features from each frame. The specific process is as follows: The sum of the gradient amplitudes of each frame (the result of Sobel operator edge detection) is calculated. Higher values indicate clearer images, thereby obtaining clarity features.
[0040] The deviation between the RGB mean of each frame and the global video RGB mean is extracted. The smaller the deviation, the more consistent the color, thus obtaining the clarity feature. The displacement difference of the digital human joints (such as shoulder joints and wrist joints) in adjacent frames is calculated through the optical flow algorithm. The smaller the difference, the more coherent the movement, thus obtaining the movement continuity feature.
[0041] Quantitatively analyze clarity features, identify blurry and noisy abnormal frames, and generate a clarity anomaly factor. The clarity threshold is set to 80 (sum of gradient amplitudes). A frame with a gradient value below 60 is considered blurry, while a frame with a gradient value above 120 and the presence of isolated pixels is considered noisy. For example, if the gradient value of frame 230 is 52 (blurry), and the gradient value of frame 890 is 135 with 3% noise pixels (noise), the system generates clarity anomaly factors for these two frames (blurry frame factor = 0.8, noisy frame factor = 0.6, with higher values indicating more severe anomalies).
[0042] Quantitative analysis of color consistency features, locate color deviation, exposure abnormal frame, generate color abnormal factor. Calculate the Euclidean distance between the RGB mean value of each frame and the global mean value (R=240, G=230, B=220), and the distance more than 50 is determined as color deviation frame; When the luminance channel (V channel) value is >240 (overexposure) or <30 (underexposure), it is determined as exposure abnormal frame. For example, the RGB mean value of the 512th frame (R=200, G=250, B=180) deviates from the global value by 65 (color deviation), and the V channel value of the 1200th frame is 245 (overexposure), generating color abnormal factor (color deviation frame factor=0.7, overexposure frame factor=0.9). Quantitative analysis of motion continuity features, detect skip frame, motion fault abnormal frame, generate continuity abnormal factor. Set the adjacent frame joint displacement difference threshold to 5 pixels, and when the difference is >15 pixels, it is determined as motion fault; If the timestamp interval between a frame and the previous frame is >40ms (normal interval 33ms), it is determined as skip frame. For example, the wrist joint displacement difference between the 750th frame and the 749th frame is 20 pixels (motion fault), and the timestamp interval of the 1001th frame is 45ms (skip frame), generating continuity abnormal factor (fault factor=0.85, skip frame factor=1.0).
[0043] Based on the video quality standard, the clarity abnormal factor, the color abnormal factor and the continuity abnormal factor are processed, the frame interpolation is used to repair the fuzzy frame, the color correction algorithm is used to adjust the color deviation frame, and the optical flow frame filling technology is used to optimize the skip frame, to generate the abnormal processing result and the corresponding repair weight. Specifically, the fuzzy frame (230th frame), for the problem of detail loss caused by motion blur, a bicubic interpolation algorithm is used for repair. The algorithm calculates the weighted average value of 16 neighboring pixels around each pixel in the frame (the weight is based on the cubic function of the pixel distance), and supplements the texture details of the fuzzy area (such as the button texture of the digital human suit, the equipment outline of the background exhibition hall). The repair weight is set to 0.9 (the weight range is 0-1), indicating that in the parallel processing of multiple abnormal frames, the frame has the highest repair priority, and the gradient amplitude of the processed frame is increased from 52 to 90, reaching the clarity standard. The noise frame (890th frame), for the random isolated noise (such as salt and pepper noise) existing in the frame, a 3x3 window median filter algorithm is used for processing. The algorithm sorts the pixel values in the window by size and replaces the center pixel with the median value, effectively removing isolated noise (noise pixel ratio from 3% to 0), while preserving the facial expression details of the digital human (such as mouth wrinkles). The repair weight is set to 0.6, the priority is medium, and the signal-to-noise ratio (SNR) of the processed frame is increased from 25dB to 38dB.
[0044] The color-shifted frame (frame 512) was corrected using the grayscale world algorithm due to light source color shift, which caused the RGB values to deviate from the global mean. This algorithm assumes that the average values of the R, G, and B channels in the image are equal (all grayscale values). By calculating the ratio of the three channel means (originally R:G:B = 200:250:180), it generates correction coefficients (R=1.2, G=0.92, B=1.22) and scales the RGB values of each pixel. After correction, the Euclidean distance between the RGB mean and the global mean decreased from 65 to 18, and the color anomaly factor decreased from 0.7 to 0.2. A correction weight of 0.7 corresponds to a medium-to-high priority. The overexposed frame (frame 1200) was corrected using the gamma correction algorithm (γ=1.2) to address the luminance channel (V channel) value of 245 (overexposure resulting in loss of detail). The luminance values are nonlinearly compressed using the formula V'=V^γ, reducing the V value of overexposed areas (such as the highlights on the digital human's forehead) from 245 to 200, restoring the gradation of light and dark in the suit collar. The dynamic range of the restored frame is increased by 15%, and a restoration weight of 0.8 corresponds to a higher priority.
[0045] The motion fault frame (frame 750) has a wrist displacement difference of 20 pixels from the previous frame (exceeding the threshold of 15 pixels). Farneback optical flow interpolation is used to generate an interim frame. The algorithm first calculates the optical flow fields of frames 749 and 750, predicts the wrist position of the intermediate interim frame (with an 8-pixel displacement difference), and then uses bilinear interpolation to fill in the pixels of the interim frame, ensuring a smooth transition from "lifting 30°" to "lifting 45°." The motion coherence score (based on the standard deviation of inter-frame displacement) decreases from 18 to 7 after restoration, with a restoration weight of 0.85 corresponding to high priority. The frame skip (frame 1001) has a 45ms timestamp gap (normally 33ms) due to frame loss. A transition frame is inserted using the previous and next frame fusion prediction technique. Based on the digital human's pose in frame 1000 (right hand raised to chest) and the pose in frame 1002 (right hand lowered to waist), a skeletal animation interpolation algorithm is used to generate an intermediate frame (right hand raised to 45° position in front of chest) to complete the motion sequence. After the insertion, the frame interval is restored to 33ms, the continuity anomaly factor is reduced from 1.0 to 0.1, and the repair weight 1.0 is the highest priority to ensure seamless action.
[0046] The anomaly processing results are fused with the restoration weights to generate an initial digital human scene video with smooth imagery and natural colors. A weighted fusion algorithm based on dynamic weight allocation is used to integrate all restored frames. The specific process is as follows: A weight mapping mechanism converts each frame's restoration weight (ranging from 0 to 1) into a fusion coefficient, using the formula: fusion coefficient = restoration weight / sum of all frame restoration weights. This ensures that restored frames with higher weights contribute more to the overall fusion (e.g., a skipped frame restoration weight of 1.0 corresponds to a fusion coefficient of 0.15, and a blurred frame restoration weight of 0.9 corresponds to a fusion coefficient of 0.135). 1,800 frames are segmented using a sliding window (window size 5 frames). The structural similarity index (SSIM) is used to calculate the similarity between adjacent frames. When the SSIM is less than 0.85, a secondary restoration (readjustment of interpolation parameters or frame infill logic) is triggered to ensure a smooth transition between frames. The Laplace operator is used to calculate frame clarity. The Laplace response value of blurred frames repaired through bicubic interpolation is increased from the original 50 to over 80. Combined with secondary optimization of the non-local means denoising algorithm, the proportion of blurred frames is reduced from 2.3% (41 frames) to 0.5% (9 frames). Noisy frames are completely eliminated through 3×3 median filtering, and the peak signal-to-noise ratio (PSNR) is improved from 28dB to 35dB.
[0047] Full-frame color distribution was verified using K-means clustering (K=8). The Euclidean distance between the RGB mean and the global mean of frames corrected for color deviation using the Gray World algorithm was kept below 30 after principal component analysis (PCA) dimensionality reduction. The proportion of frames with color aberrations decreased from 3.1% (56 frames) to 0.8% (14 frames), and the color histogram matching reached 92%. A dynamic time warping (DTW) algorithm was used to align the joint motion trajectories of adjacent frames. Intermediate frames generated using optical flow interpolation were smoothed using a Kalman filter to less than 8 pixels at key joints such as the wrist and shoulder. Motion discontinuities and frame skipping were completely eliminated, and the cosine similarity of inter-frame motion vectors was ≥0.95, meeting the requirements for motion smoothness in subsequent editing. The resulting initial digital human scene video received an overall quality score of 90 out of 100 according to the MPEG-7 video quality assessment standard, providing a high-quality foundation for secondary editing.
[0048] S104 , performing secondary editing on the initial digital human scene video, generating an updated digital human scene video in combination with the user's operation requirements and the preset video selection order.
[0049] In one implementation, the editing requirements, preset video sequence rules, and user operation instructions of the initial digital human scene video are extracted and classified to generate information on text addition requirements, special effects addition requirements, picture-in-picture effect requirements, and video sequence adjustment information. Taking the initial "Smartwatch Function Introduction" video as an example, the user issues operation instructions through the Yunbao device's editing interface: adding the text "Up to 7 days of battery life" to the 10-15 seconds of the video, adding a "flash light effect" special effect at the 20th second, overlaying the digital human explanation at the 30-35 seconds of the video with the smartwatch product demonstration video in picture-in-picture, and swapping the order of the "Function 2" and "Function 3" explanation segments (25-30 seconds and 35-40 seconds) in the original video.
[0050] The system classifies and processes these requirements, as follows: Requirement information for adding text: timestamp 10-15 seconds, text content "up to 7 days of battery life", bold font, 36px font size, white color; Requirement information for adding special effects: timestamp 20 seconds, special effect type "flashing light effect", duration 2 seconds, brightness parameter 50%; Requirement information for picture-in-picture effect: the main screen is the digital human explanation (30-35 seconds), the sub-screen is the product display video (30% of the size, located in the upper right corner), and the transparency is 80%; Video sequence adjustment information: the original clip order "Function 2 (25-30 seconds) → Function 3 (35-40 seconds)" is adjusted to "Function 3 (35-40 seconds) → Function 2 (25-30 seconds)".
[0051] Based on video editing technology standards, the information on text addition requirements, special effects addition requirements, picture-in-picture effect requirements, and video order adjustment is processed. The alpha algorithm is used to achieve picture-in-picture overlay, editing tools are used to add text and special effects, and time reordering technology is used to adjust the video order. The editing processing results and demand satisfaction evaluation information are generated. Text is added at a specified timestamp using editing tools (such as FFmpeg filters). Anti-aliasing algorithms are used to optimize font edges to ensure that the contrast between text and background is ≥4:1. The processing results show that the text is clearly visible, and the demand satisfaction is 100%. The "flash light effect" module of the built-in special effects library is called, and the effect is achieved through an inter-frame brightness gradient algorithm (brightness fluctuation of ±20% per frame). The duration is precisely matched to 2 seconds, and the demand satisfaction is 95% (because the light effect range is slightly larger than the preset area).
[0052] Alpha blending was used, with the alpha channel values of the main and sub-images set to 1.0 and 0.8, respectively. The layer blending formula (output pixels = main image pixels × 1.0 + sub-image pixels × 0.8) achieved natural overlay without hard edges, achieving 100% compliance. Time reordering (based on video frame index remapping) was used to re-time the two clips, ensuring motion continuity at the transitions (the difference in displacement between joints in adjacent frames was ≤ 5 pixels), achieving 100% compliance.
[0053] Based on the editing results and demand satisfaction assessment information, the content of the initial digital human scene video was optimized and corrected. Secondary adjustments were made to areas that did not meet the requirements, generating the edited and optimized results. To address the 95% satisfaction level for adding special effects, the lighting effect range parameters were adjusted again, reducing the area of effect from the original digital human's entire body to just the hands (to match the "endurance" explanation). This ensured that the lighting effect accurately covered the target area, and the satisfaction level was increased to 100% after the correction.
[0054] A second verification of the text effects revealed that at the 12th second, the text and background brightness were similar. The text color saturation was increased (RGB values adjusted from 255,255,255 to 255,255,200) to enhance contrast and ensure readability. The picture-in-picture overlay and sequence adjustment effects met expectations and no further revisions were required. The final result was an optimized editing result that included text, special effects, picture-in-picture, and sequence adjustment.
[0055] The editing optimization results are integrated and synchronized, and combined with the preset video selection order to generate an updated digital human scene video. The system synchronizes all editing optimization results along the timeline, ensuring frame-level alignment of text, special effects, and picture-in-picture with the video (time error ≤ 0.01 seconds). The user's preset video selection order is "Smartwatch Function Introduction → Usage Scenario Demonstration", so after integration, the currently edited video is automatically linked to the "Usage Scenario Demonstration" video in sequence (reserving transition frames). The final updated video is 65 seconds long and has a frame rate of 30fps. After verification, the text is clear, the special effects are natural, the picture-in-picture is seamless, the clip sequence is logical, and it smoothly connects with the next video in the preset selection order, meeting the publishing requirements.
[0056] S105, extracting and classifying the check sequence information, preset release rule information and multi-platform account information of multiple videos to generate video concatenation sequence information, platform account configuration information, release copy title information, and release timing planning information.
[0057] In one embodiment, the selection order information of multiple videos is extracted and classified to generate video concatenation order information. A user selected three videos on the Yunbao device's video management interface: "Smartwatch Function Introduction" (65 seconds), "Usage Scenario Demonstration" (40 seconds), and "User Review Collection" (30 seconds), with the selection order being 1 → 2 → 3. The system extracts this sequence information and generates a video concatenation order table, which includes each video's unique ID, original duration, selection sequence number, and transition method (a 0.5-second fade-in and fade-out transition effect is added by default). For example: Video 1: ID = 20231001, duration 65 seconds, sequence number 1, transition method "fade-in"; Video 2: ID = 20231002, duration 40 seconds, sequence number 2, transition method "fade-in and fade-out"; Video 3: ID = 20231003, duration 30 seconds, sequence number 3, transition method "fade-out."
[0058] The system extracts and classifies the preset publishing rule information to generate the publishing copy title information. The user-preset publishing rules include the general copy template "[New Product Recommendation] {Product Name}'s {Core Selling Point}, do you get it?" and the title template "{Product Name} Three Major Advantages Analysis + Real Experience". The system combines the video content keywords ("smart watch", "battery life", "waterproof") to perform variable replacement and generate targeted publishing copy and title: Publishing copy: "[New Product Recommendation] Smart watch's super long battery life and IP68 waterproof, do you get it?"; Publishing title: "Smart watch's three major advantages analysis + real experience". It also supports platform-specific copy supplements, such as Platform A adding the topic tag "#smartwear#black technology" and Platform B adding the guide "Click to view detailed functions".
[0059] The system extracts and categorizes account information from multiple platforms to generate platform account configuration information. The system extracts account information for the platforms to be published from the user's preset account library and organizes it by platform type, including account ID, authorization status, publishing permissions, and format requirements: Platform A: Account ID = "Tech Trends Gallery", valid authorization, supports 1080p video with a maximum length of 15 minutes; Platform B: Account ID = "Smart Life Guide", valid authorization, supports vertical screen video, and originality protection is enabled by default; Platform C: Account ID = "SmartWatchReview", valid authorization, requires the addition of English subtitles, and supports 4K resolution. Unauthorized or expired accounts (such as "Kuaishou" account authorization has expired) will be marked as "pending" and trigger a reminder.
[0060] The system generates a release schedule based on the above information. Based on the total video length (65 + 40 + 30 + 1.5 transitions = 136.5 seconds), the platform release interval requirements (e.g., ≥ 5 minutes between Platforms A and B), and the user-preset release time (7:00 PM - 8:00 PM), the system creates a schedule: 7:00 PM: Release to Platform A; 7:06 PM: Release to Platform B; 7:12 PM: Release to Platform C. The system also identifies the pre-processing tasks required for each platform release (e.g., automatically generating English subtitles on Platform C is estimated to take 2 minutes) to ensure that the release is executed as planned.
[0061] S106, using a video splicing algorithm to sequentially connect videos, a unified account management module to configure platform information, and a content adaptation tool to generate custom copy titles, and generate concatenation distribution results and platform adaptability evaluation information.
[0062] In one implementation, the system invokes the FFmpeg video concatenation algorithm based on the video concatenation sequence ("Smartwatch Function Introduction → Usage Scenario Demonstration → User Review Highlights"). The algorithm first performs a consistency check on the encoding format (all H.264) and resolution (1920×1080) of the three videos to ensure compatibility. Then, based on a pre-defined concatenation method, a 0.5-second fade-in / fade-out transition frame is added between Video 1 and Video 2 (using frame interpolation to generate an intermediate fade-in / fade-out transition). A similar transition effect is also added between Video 2 and Video 3. The total concatenation duration is 65 + 0.5 + 40 + 0.5 + 30 = 136 seconds, generating a complete concatenated video file (MP4 format, 10Mbps bitrate). A CRC check is then performed to ensure data integrity.
[0063] After reading the platform account configuration information, the unified account management module automatically completes multi-platform authorization configuration. For account A, "Tech Trends," the access token is refreshed via OAuth 2.0, publishing permissions are validated, and default publishing settings are configured (privacy permissions are set to "Public" and comments are allowed). For account B, "Smart Life Guide," the WeChat Open Platform API is used to bind the account and enable the "Original Protection" and "Appreciation" features. For account C, "SmartWatchReview," the video category is configured as "Technology" via the API, and the "Automatic English Subtitle Generation" feature is enabled. After configuration is complete, an account status report is generated, indicating that authorization is valid for all platforms (status code 200) and there are no configuration errors.
[0064] The content adaptation tool generates customized content based on the title information of the published copy and the characteristics of each platform: Specifically, on Platform A, the title retains "Analysis of the three major advantages of smart watches + real experience", and the copy is supplemented with "Click the link in the lower left corner to go directly to the purchase page #Smart Wear #Black Technology" (in line with the platform's traffic label rules); Platform B: The title is adjusted to "
New Product
[0065] Generate serial distribution results and platform adaptability evaluation information. Specifically, the serial distribution results include the serial video storage path (local path + cloud backup URL), customized copy titles for each platform (JSON format), and publishing task ID (such as "PUB20231001001"); evaluate the adaptation effect through quantitative indicators - Platform A (adaptability 95%, because the topic tags comply with platform rules), Platform B (adaptability 90%, because the title conforms to the WeChat style), Platform C (adaptability 98%, because the English translation is accurate and the keyword optimization is in place), and the adaptability of non-adapted platforms (such as Platform D, due to expired account authorization) is marked as 0%, and an abnormal reminder is triggered.
[0066] S107, based on the concatenation distribution result and the platform adaptability evaluation information, the video concatenation content and publishing parameters are verified and adjusted, the publishing parameters are optimized for the platform with adaptation abnormalities, and an optimized concatenation distribution plan is generated.
[0067] In one implementation, the system performs a comprehensive verification based on the concatenated distribution results (136 seconds of concatenated video and customized copy for each platform) and platform compatibility assessment information (95% for Platform A, 90% for Platform B, and 98% for Platform C). Specifically, video content verification: Inter-frame similarity analysis (SSIM ≥ 0.9) confirms that the concatenated video transitions are natural and without image discontinuities. It also detects a 1.2-second dark frame (V channel value < 200) in the Platform B version, which does not meet the platform brightness standard (recommended ≥ 220). Release parameter verification: The compliance of the copy titles on each platform was checked. It was found that the copy on Platform A, "Click the link in the lower left corner," contained a jump guide, potentially triggering platform throttling (a 5% reduction in compatibility). Furthermore, the English subtitles on Platform C contained three translation errors (e.g., "lifetime" was incorrectly translated as "endurance" instead of "batterylife").
[0068] We optimized publishing parameters for platforms experiencing compatibility issues. Specifically, for Platform B, we optimized the image quality by applying gamma correction (γ=0.8) to boost dark areas and adjusting the V channel value from 190 to 230. A contrast enhancement algorithm (gain 1.1) was also used to enhance the depth of the image. This improved compatibility to 95%. We also adjusted the text on Platform A by changing "Click the link in the lower left corner" to "Click the product label for details," complying with platform compliance requirements and restoring compatibility to 100%. We also corrected subtitles on Platform C by updating subtitles using the Google Translate API and combining industry terminology to ensure accurate translation of keywords like "battery life" and "waterproof." We also increased the subtitle display duration from 2 seconds to 3 seconds, maintaining compatibility at 98%.
[0069] An optimized concatenated distribution plan was generated, including the optimized concatenated videos (stored by platform as Platform A, Platform B, and Platform C versions), a revised collection of text and titles (in JSON format), and an updated release schedule: Platform A: Released at 7:00 PM, with the video brightness adjusted to meet platform standards and the text replaced with a compliant version; Platform B: Released at 7:06 PM, with optimized screen brightness and the title "[New Product] Analysis of the Three Major Advantages of Smartwatches"; Platform C: Released at 7:12 PM, with revised subtitles and the addition of the hashtags "#SmartWatch#LongBatteryLife" to enhance search ranking. Compatibility comparisons before and after optimization for each platform were also noted (e.g., Platform B: 90% to 95%), along with exception handling records, to ensure traceability.
[0070] S108, integrating and triggering the optimized serial distribution plan, responding to manual, voice or automated workflow instructions, and generating publishing results that are distributed synchronously to multiple platforms.
[0071] In one implementation, the optimized concatenated distribution plan is integrated: the system combines the optimized concatenated videos (platform A, platform B, and platform C), the revised title text, and the release schedule into a unified task package, which is stored in a local cache (path: / storage / publish / task_20231001 / ) and synchronized to a cloud backup. The integration includes video files: named by platform (e.g., "smartwatch_douyin.mp4," "smartwatch_wechat.mp4"), with a checksum (MD5 value) to ensure integrity; title text: customized for each platform (e.g., platform A's copy: "Click on the product tag to learn more #smartwear #blacktech"); and release parameters: release time for each platform (7:00 PM, 7:06 PM, 7:12 PM), account ID, and format requirements (e.g., English subtitles are required for platform C). Release is triggered in response to manual, voice, or automated workflow commands.
[0072] The user clicks the "publish immediately" button in the "publish management" interface of the cloud treasure device, and the system reads the task package and starts the publishing process. The user issues the "start publishing smart watch video" instruction through the device microphone, and the voice recognition model (based on MFCC algorithm) parses successfully (matching degree 98%), and automatically calls the publishing interface. According to the preset "publish product video every day at 19:00" workflow (CRON expression: 0019**?), the system automatically triggers the publishing after reaching the specified time.
[0073] The publishing result is generated and distributed synchronously to multiple platforms: platform A: published on time at 19:00, returned a publishing success receipt (video ID: 6987654321), displayed "recommended to local traffic pool", and the real-time playback volume was synchronized to the device background; platform B: completed publishing at 19:06, the system obtained playback volume, like number, etc. (playback volume 200+ within 5 minutes after publishing), and started original protection identification; platform C: published successfully at 19:12, generated a video link (https: / / aabb.com / watch?v=abc123), the automatically added English subtitles were verified without errors, and the platform returned a "comply with community standards" prompt. The publishing result is displayed in the form of a visual report on the device interface, including the publishing status of each platform ("success" "under review"), time consumption (A platform 2.3 seconds, B platform 3.5 seconds, C platform 5.8 seconds) and abnormal prompt (no abnormality), and the result log ( / logs / publish_20231001.log) is archived for subsequent tracing.
[0074] The present application extracts portrait features through deep learning portrait segmentation algorithm, combines green screen removal or RGB-D depth information to separate background; integrates pose, costume templates of digital human template library, through optical flow interpolation frame filling, HSV light adjustment and super resolution reconstruction, generates transparent graph digital human adapted to environment. Receive user script and instructions, generate action and expression logic through BERT model parsing script, frame synchronous synthesis through animation synthesis engine, generate scene video through multi-thread parallel processing, alpha channel mixing realizes natural superposition of digital human and background. Scan and repair abnormal frames (frame interpolation, color correction, optical flow frame filling) frame by frame; support text, special effect addition and picture-in-picture in secondary editing, adjust video order to generate updated version. According to the selected order, concatenate the videos, configure multiple platform accounts, generate adapted scripts; after verification and optimization, respond to multiple instruction synchronous distribution to multiple platforms. Support multi-task parallel processing, meet the secondary editing demand, further improve the material reuse rate. Multi-platform synchronous publishing, improve the adaptation degree, shorten the publishing process time, break through the manual operation limit.
[0075] As Figure 2As shown, a digital human material generation and sharing device based on intelligent hardware includes: an acquisition module for acquiring target video and image information, processing the target video and image information, and generating a digital human with a transparent background layer; a processing module for receiving user input of text, background information, and workflow instructions including voice, keystrokes, and automation, and processing the generated digital human to generate a digital human scene video that explains the text; processing abnormal video frames in the digital human scene video to generate an initial digital human scene video; performing secondary editing on the initial digital human scene video, and generating an updated digital human scene video based on the user's operation requirements and the preset video selection order; and processing the selection order information of multiple videos, the preset release order, and the like. Extract and classify distribution rule information and multi-platform account information to generate video concatenation sequence information, platform account configuration information, release copy title information, and release timing planning information; use video splicing algorithm to concatenate videos in sequence, configure platform information with unified account management module, and generate custom copy titles with content adaptation tools to generate concatenation distribution results and platform adaptability evaluation information; verify and adjust video concatenation content and release parameters based on concatenation distribution results and platform adaptability evaluation information, optimize release parameters for platforms with adaptation exceptions, and generate optimized concatenation distribution plans; integrate and trigger the optimized concatenation distribution plans, respond to manual, voice, or automated workflow instructions, and generate release results that are distributed synchronously to multiple platforms.
[0076] A computing device comprises a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute any one of the methods for generating and sharing digital human materials based on intelligent hardware.
[0077] The methods and / or embodiments in the embodiments of the present application can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above-mentioned functions defined in the method of the present application are performed.
[0078] It is instructive to note that the computer readable medium described herein can be a computer readable signal medium or a computer readable storage medium or any combination thereof. More specific examples of the computer readable storage medium can include, but are not limited to: an electronic connection having one or more wires; a portable computer diskette; a hard disk; a random access memory (RAM); a read-only memory (ROM); an erasable programmable read-only memory (EPROM or Flash memory); an optical fiber; a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination of the foregoing. In the present application, a computer readable medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0079] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0080] It will be apparent to those skilled in the art that the present application is not limited to the specific embodiments described above, but that there are many possible variations and modifications that are also within the spirit and scope of the application. Thus, the application is not to be limited to any specific examples described herein, but rather only by the claims that follow.
Claims
1. A method for generating and sharing digital human materials based on intelligent hardware, applied to intelligent hardware, characterized in that: include: Obtain target video and image information, process the target video and image information, and generate a digital human with a transparent background layer; Receive user input of copy, background information, and workflow instructions including voice, keystrokes, and automation, and process them in combination with the generated digital human to generate a scene video of the digital human explaining the copy; Processing abnormal video frames in the digital human scene video to generate an initial digital human scene video; Perform secondary editing on the initial digital human scene video, and generate an updated digital human scene video based on the user's operational requirements and the preset video selection order. This includes extracting and classifying the editing requirements of the initial digital human scene video, preset video sequence rules, and user operational instructions, and generating information on text addition requirements, special effects addition requirements, picture-in-picture effect requirements, and video sequence adjustment requirements. Based on video editing technical standards, the system processes information on text addition requirements, special effects addition requirements, picture-in-picture effect requirements, and video sequence adjustment requirements. It uses the Alpha algorithm to implement picture-in-picture overlays, editing tools to add text and special effects, and time reordering technology to adjust video sequences, generating editing processing results and demand satisfaction evaluation information. Based on the editing results and demand satisfaction evaluation information, the content of the initial digital human scene video is optimized and corrected, and secondary adjustments are made to the parts that do not meet the requirements to generate editing optimization results; Integrate and synchronize the editing optimization results, combine the preset video selection sequence, and generate the updated digital human scene video; Extract and classify the selection order information, preset publishing rule information, and multi-platform account information of multiple videos to generate video concatenation order information, platform account configuration information, publishing copy title information, and publishing timing planning information; Use video splicing algorithms to sequentially connect videos, unify account management modules to configure platform information, and use content adaptation tools to generate custom copy titles, generating connection distribution results and platform adaptability assessment information. Based on the concatenation distribution results and platform adaptability assessment information, the video concatenation content and publishing parameters are verified and adjusted. The publishing parameters are optimized for platforms with abnormal adaptability, and an optimized concatenation distribution plan is generated. Integrate and trigger optimized serial distribution solutions, respond to manual, voice, or automated workflow instructions, and generate publishing results that are distributed synchronously to multiple platforms.
2. The method for generating and sharing digital human materials based on intelligent hardware according to claim 1, characterized in that: Obtain target video and image information, process the target video and image information, and generate a digital human with a transparent background layer, including: Obtain real-time captured videos, uploaded videos, or one or more photos as target video and image information; A deep learning-based portrait segmentation algorithm model is used to process the target video and image information frame by frame. The portrait outline, facial features, and motion trajectory information are extracted through pixel-level semantic segmentation. Green screen chroma keying technology or background separation methods based on RGB-D depth information are used to accurately remove green screens or complex backgrounds, ensuring a natural transition between portrait edges. Combining the preset body posture templates, clothing templates, and action libraries in the digital human template library, the model performs feature matching and fusion between the portrait information after background removal and the template. The facial animation of the template is driven based on the expression feature points, and the body movements of the template are mapped according to the action trajectory information. This creates an initial digital human with a transparent layer. The model library supports retrieval and call by style and scene classification. Perform inter-frame interpolation on the initial digital human, predict the motion trend of adjacent frames through the optical flow estimation algorithm, supplement the transition frames, and use the time domain filtering algorithm to smooth the expression change curve to avoid sudden changes in expression; Based on the real-scene lighting parameters collected by the ambient light sensing module, the HSV color space conversion technology is used to dynamically adjust the brightness of the digital human's highlight area and the contrast of the shadow area. The super-resolution reconstruction algorithm is used to optimize facial expression details and portrait edge clarity. Combined with the skin color adaptive adjustment model to match the real-scene color temperature, a digital human with a transparent background layer is generated with a high degree of integration with the environment and excellent naturalness.
3. The method for generating and sharing digital human materials based on intelligent hardware according to claim 2, characterized in that: The system receives user input of copywriting, background information, and workflow instructions including voice, keystrokes, and automation, processes them in conjunction with the generated digital human, and generates a scene video of the digital human explaining the copywriting, including: Receive text entered by the user through a dialog box, background information manually selected or automatically matched based on text keywords, as well as voice commands, key trigger signals, and automated workflow instructions; The natural language processing model is used to perform semantic analysis on the text, extract the core explanation points, and generate the digital human action sequence and expression change logic. At the same time, the background information is parsed into scene rendering parameters. The generated digital human action sequence, expression logic and scene rendering parameters are input into the animation synthesis engine, and combined with the created digital human with transparent layers for frame synchronization synthesis. Multi-threaded parallel processing technology is used to support the simultaneous execution of one or more generation tasks. The inter-frame smoothing algorithm is used to optimize the continuity of digital human movements, and the alpha channel blending technology is used to achieve the natural superposition of digital humans and backgrounds to generate the initial digital human scene video.
4. The method for generating and sharing digital human materials based on intelligent hardware according to claim 3, characterized in that: Process the abnormal video frames in the digital human scene video to generate the initial digital human scene video, including: Scan the digital human scene video frame by frame to extract the clarity, color consistency, and movement continuity features of the frame; Quantitatively analyze the clarity features, identify blurry and noisy abnormal frames, and generate clarity abnormality factors; Quantitatively analyze color consistency features, locate color deviation and exposure abnormality frames, and generate color anomaly factors; Quantitatively analyze the motion continuity features, detect frame skipping and motion discontinuity abnormal frames, and generate a continuity anomaly factor; Based on the video quality standard, the clarity anomaly factor, color anomaly factor, and coherence anomaly factor are processed. Frame interpolation is used to repair blurred frames, color correction algorithms are used to adjust color-shifted frames, and optical flow interpolation technology is used to optimize frame skipping. The anomaly processing results and corresponding repair weights are generated. The abnormal processing results and the restoration weights are fused to generate an initial digital human scene video with smooth images and natural colors.
5. A digital human material generation and sharing device based on intelligent hardware, characterized in that: The device comprises: The acquisition module is used to obtain target video and image information, process the target video and image information, and generate a digital human with a transparent background layer; The processing module is used to receive the text, background information and workflow instructions including voice, button and automation input by the user, and process them in combination with the generated digital human to generate a digital human scene video explaining the text; process the abnormal video frames in the digital human scene video to generate an initial digital human scene video; perform secondary editing on the initial digital human scene video, and generate an updated digital human scene video in combination with the user's operation requirements and the preset video check order, including extracting and classifying the editing requirements of the initial digital human scene video, the preset video sequence rules and the user's operation instructions, and generating text adding requirement information, special effect adding requirement information, picture-in-picture effect requirement information, and video sequence adjustment information; based on the video editing technical standards, the text adding requirement information, special effect adding requirement information, picture-in-picture effect requirement information, and video sequence adjustment information are processed, and the alpha algorithm is used to realize picture-in-picture overlay, the editing tool adds text and special effects, and the time reordering technology is used to adjust the video sequence, and the editing processing results and demand satisfaction evaluation information are generated; based on the editing processing results and demand satisfaction The content of the initial digital human scene video is optimized and corrected based on the sufficiency assessment information, and secondary adjustments are made to the parts that do not meet the requirements to generate editing optimization results; the editing optimization results are integrated and synchronized, and combined with the preset video selection order to generate an updated digital human scene video; the selection order information, preset publishing rule information and multi-platform account information of multiple videos are extracted and classified to generate video concatenation sequence information, platform account configuration information, release copy title information, and release timing planning information; the video splicing algorithm is used to concatenate videos in sequence, the unified account management module is used to configure platform information, and the content adaptation tool is used to generate custom copy titles to generate concatenation distribution results and platform adaptability assessment information; based on the concatenation distribution results and platform adaptability assessment information, the video concatenation content and publishing parameters are verified and adjusted, and the publishing parameters are optimized for platforms with adaptation anomalies to generate an optimized concatenation distribution plan; the optimized concatenation distribution plan is integrated and triggered to respond to manual, voice or automated workflow instructions to generate publishing results that are distributed synchronously to multiple platforms.
6. An electronic device, characterized in that: include: a first processor; and a memory for storing executable instructions of the first processor; Wherein, the first processor is configured to execute the digital human material generation and sharing method based on intelligent hardware as described in any one of claims 1 to 4 by executing the executable instructions.
7. A computing device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein: When the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Video generation method and device, electronic equipment, storage medium and product
CN118803173A
Video automatic generation method and device, equipment and medium
CN120050463A