A method and apparatus for generating digital images that are adaptive across multiple terminals and scenarios.
By collecting and extracting video image features of physical reference objects in the power industry, a digital image generation intelligent agent is trained, solving the problems of reusability and realism in digital image generation in the power industry, and realizing efficient and accurate interaction in multiple terminals and scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2026-03-06
AI Technical Summary
Existing digital avatar generation technologies in the power industry suffer from problems such as long development cycles, high labor costs, non-reusable models, insufficient data collection accuracy, large errors in capturing subtle facial expressions and device operation gestures, easy misalignment of lip-sync, large fluctuations in model performance, and lag in terminal display. These issues make it difficult to meet the power industry's requirements for high realism, multi-terminal adaptation, and seamless interaction across multiple scenarios.
By collecting video images of physical reference objects, extracting features in single and related dimensions, training a digital avatar generation agent, and using related dimension features to constrain lip-sync with speech, a customized digital avatar is generated that adapts to different terminals and scenarios.
It has achieved highly realistic digital image generation, standardized and efficient production, seamless adaptation to multiple terminals and scenarios, and precise and intelligent interaction in the power sector, lowering the technical threshold and improving the naturalness and rationality of the interaction process.
Smart Images

Figure CN121120886B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and apparatus for generating digital images that are adaptive across multiple terminals and scenarios. Background Technology
[0002] With the deep application of artificial intelligence technology in the power industry, digital avatars have become a core interactive carrier in scenarios such as power operation monitoring data broadcasting, equipment operation training, and customer service consultation. The industry's demand for high realism, multi-terminal adaptation, and seamless interaction across multiple scenarios for digital avatars is becoming increasingly urgent. Current digital avatar generation technology mainly revolves around data collection, feature extraction, model training, and terminal display. It involves collecting videos of physical reference objects using high-definition camera equipment, extracting features using algorithms such as convolutional neural networks and generative adversarial networks, training and generating models, and then adjusting the digital avatar format using a rendering engine to adapt to different terminals, thus initially realizing the scenario-based application of digital avatars.
[0003] However, existing technologies still have significant shortcomings in vertical scenarios within the power industry. Current solutions require training a separate digital avatar model for each scenario, meaning one model corresponds to one scenario, generating a single type of digital avatar. For example, power operation monitoring scenarios require training a dedicated broadcasting avatar model, while power training scenarios require retraining an equipment operation avatar model. Parameters and features cannot be reused between models, leading to long development cycles, high labor costs, and difficulty in meeting the rapidly expanding needs of power industry scenarios. Furthermore, on the one hand, data acquisition accuracy and scenario generalization capabilities are insufficient, relying heavily on ordinary camera equipment and lacking high-precision capture methods such as depth sensors. This results in large errors in capturing key information such as subtle facial expressions and equipment operation gestures. Moreover, the sample of entity reference objects within the same scenario is limited, making it difficult to adapt to the diverse appearances and operating habits of personnel in the power industry. On the other hand, the synchronization of related dimensional features and the practicality of the model are poor. Lip-sync is prone to misalignment, especially in scenarios such as broadcasting power-related technical terms, resulting in insufficient realism. Additionally, model training parameters rely on manual adjustment, and multi-terminal adaptation lacks differentiated strategies, leading to large fluctuations in model performance and lag in terminal display, failing to meet the core requirements of real-time power operation monitoring and professional training.
[0004] These issues make it difficult for existing digital avatar generation technologies to fully support the digital transformation of the power industry. There is an urgent need for a technical solution that can achieve highly realistic digital avatar generation, standardized and efficient production, seamless adaptation to multiple terminals and scenarios, and precise and intelligent interaction in the power sector. This solution would address the shortcomings of current technologies in terms of accuracy, generalization, and practicality, and promote the large-scale and high-quality application of digital avatars in the power industry. Summary of the Invention
[0005] In view of this, this application provides a method and apparatus for generating digital images that are adaptive to multiple terminals and multiple scenarios, which can realize highly realistic generation of digital images, standardized and efficient production, seamless adaptation to multiple terminals and multiple scenarios, and precise and intelligent interaction in the power field.
[0006] Specifically, this application is implemented through the following technical solution:
[0007] The first aspect of this application provides a method for generating digital images that is adaptive across multiple terminals and scenarios, the method comprising:
[0008] Acquire video images of physical reference objects;
[0009] Extract a first feature of a single dimension and a second feature of an associated dimension from the video image, wherein the second feature includes at least a joint feature of the two associated dimensions;
[0010] The associated dimension includes at least the corresponding features of the lip shape and speech dimension. The speech signal and lip shape change image are extracted from the video image. The speech signal is converted into a phoneme sequence. The phoneme sequence is matched with the lip shape feature to obtain multiple pairs of phoneme sequences and lip shape pairs. Supplementary lip shapes are generated according to the lip shape corresponding to adjacent phonemes. The supplementary lip shapes are used as the transitional lip shapes corresponding to the lip shapes corresponding to adjacent phonemes. The second feature corresponding to the lip shape and speech dimension is adjusted.
[0011] Using the first feature and the shooting scene corresponding to the video image as learning objects, and the second feature as constraint rules, a digital image generation intelligent agent is trained. The input of the digital image generation intelligent agent includes at least scene data, and the output includes at least digital image.
[0012] The scene data to be generated for the digital avatar is input into the digital avatar generation intelligent agent to generate a customized digital avatar.
[0013] The second aspect of this application provides a multi-terminal, multi-scenario adaptive digital image generation device, the device comprising a data acquisition module, an extraction module, a training module, and a generation module;
[0014] The acquisition module is used to acquire video images of physical reference objects;
[0015] The extraction module is used to extract a first feature of a single dimension and a second feature of an associated dimension from the video image, wherein the second feature includes at least a joint feature of the two associated dimensions;
[0016] The associated dimension includes at least the corresponding features of the lip shape and speech dimension. The speech signal and lip shape change image are extracted from the video image. The speech signal is converted into a phoneme sequence. The phoneme sequence is matched with the lip shape feature to obtain multiple pairs of phoneme sequences and lip shape pairs. Supplementary lip shapes are generated according to the lip shape corresponding to adjacent phonemes. The supplementary lip shapes are used as the transitional lip shapes corresponding to the lip shapes corresponding to adjacent phonemes. The second feature corresponding to the lip shape and speech dimension is adjusted.
[0017] The training module is used to train a digital image generation agent with the first feature and the shooting scene corresponding to the video image as learning objects, and the second feature as constraint rules. The input of the digital image generation agent includes at least scene data, and the output includes at least digital image.
[0018] The generation module is used to input the scene data of the digital image to be generated into the digital image generation intelligent agent to generate a customized digital image.
[0019] The multi-terminal, multi-scenario adaptive digital image generation method and apparatus provided in this application achieves automatic generation of a scene-matched digital image with synchronized lip movements and speech through a single model. Specifically, it first acquires real-world appearance and movement data by collecting video images of physical reference objects. Then, it extracts a first feature in a single dimension and a second feature in a related dimension. The second feature includes the relationship between the first feature and its related features; that is, the second feature is a feature pair, which includes features with correlation. The agent is trained using the second feature as a constraint, ensuring that during digital image generation, it accurately reproduces the detailed features of individual parts while avoiding disjointed movements of different parts through the constraint of the related dimension. The entire process from data acquisition to model training guarantees the realism of the digital image. Secondly, the shooting scene corresponding to the video image is used as the learning object of the intelligent agent. The intelligent agent's input is clearly defined as the scene data of the digital image to be generated, and the output is a customized digital image. In this way, the intelligent agent for digital image generation can learn the correspondence between scene and image features. There is no need for human intervention in professional aspects such as appearance design and motion choreography of digital images. Even if the operator does not have professional digital image design skills such as 3D modeling and animation design, they only need to input the basic data of the scene to be generated, and the intelligent agent can automatically match the appearance and motion features corresponding to the scene and finally output a matching customized image. This process completely reduces the technical threshold for digital image generation and avoids the problems of low efficiency and high cost caused by the need for professional designers to customize images for each scene in traditional solutions. At the same time, it ensures that digital images in different scenes comply with power industry standards and effectively meet the differentiated image generation needs in multiple scenarios. In addition, for the lip shape and speech dimensions in the association dimension, a transitional supplementary lip shape is generated based on the aligned phonemes and lip shape, which optimizes the smoothness of lip shape changes at different time points. This effectively solves the problems of lip shape and speech misalignment and abrupt transition in traditional technologies, ensuring that the lip shape and speech of the digital image are highly synchronized in the voice interaction scenario, and improving the naturalness and rationality of the interaction process. Attached Figure Description
[0020] Figure 1 A flowchart of an embodiment of the multi-terminal, multi-scenario adaptive digital image generation method provided in this application;
[0021] Figure 2 This is a schematic diagram of the structure of Embodiment 2 of the multi-terminal, multi-scenario adaptive digital image generation device provided in this application. Detailed Implementation
[0022] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0023] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0025] The following specific embodiments are given to illustrate the technical solution of this application in detail.
[0026] Example 1
[0027] Figure 1 This is a flowchart of an embodiment of the multi-terminal, multi-scenario adaptive digital image generation method provided in this application. Please refer to... Figure 1 The method provided in this embodiment may include:
[0028] S101. Acquire video images of the entity reference object.
[0029] It should be noted that acquiring video images of a physical reference object refers to obtaining a sequence of dynamic images containing the target physical reference object through a camera device. The physical reference object can be a real person, animal, or object used to construct the digital avatar. The acquired video images not only include continuous frames but also reflect the appearance characteristics, movement information, and posture changes of the reference object at different points in time.
[0030] Specifically, acquiring video images of physical reference objects includes:
[0031] (1) Determine multiple scene types and the entity reference objects used in the past under each scene type. Multiple entity reference objects under the same scene type have different shapes.
[0032] It's important to clarify that the application scenarios involved in digital avatar generation are first categorized. Taking the power industry as an example, these can be divided into power operation monitoring scenarios, power training scenarios, power news broadcasting scenarios, and power customer service scenarios. Different scenario types correspond to different industry environments and action requirements. For instance, power operation monitoring scenarios require matching a large monitoring screen background with cool-toned lighting, and actions primarily consist of dignified broadcasting gestures. Power training scenarios require matching an equipment operation platform background with natural lighting, and actions primarily consist of equipment operation-related gestures. Simultaneously, historically used physical reference objects are identified for each type of power scenario, such as the professional instructor image in the operation monitoring scenario and the power maintenance personnel image in the training scenario. To enhance the generalization ability of digital avatars in the power industry, multiple physical reference objects with different appearances need to be selected within the same type of power scenario. For example, maintenance personnel of different genders and years of service can be selected for the training scenario. This allows for the acquisition of diverse appearance characteristics (such as different styles of maintenance uniforms) and action data (such as gestures with different operating habits) that are relevant to the actual power industry, thereby improving the adaptability and realism of the digital avatar generated in various power scenarios.
[0033] (2) Marking points are set on the face and movements of the entity reference object, and the changes of the marking points are tracked.
[0034] It should be noted that markers are highly recognizable (e.g., specific color, shape) markers used to assist computer vision systems in recognizing and tracking the movement of objects. Specifically, they are standardized markers placed at specific locations on a physical reference object that differ significantly from the physical object's clothing material / color and the background environment in terms of visual characteristics (color, reflectivity, brightness).
[0035] By attaching or setting high-contrast markers on key facial regions (such as the corners of the eyes, the corners of the mouth, and the eyebrows) and key movement parts (such as hands, arms, legs, and body joints) of physical reference objects, and then using computer vision technology, such as image feature-based tracking algorithms, continuously monitoring the positional changes of these markers in the video sequence, this method can accurately and meticulously capture the facial expression dynamics of the physical reference objects (such as the upward movement of the corners of the mouth markers when smiling, and the convergence of the eyebrow markers when frowning) and body movement trajectories (such as the movement paths of the hand and arm markers when waving). This provides accurate raw data support for the subsequent extraction of facial and movement features of the digital image, ensuring the realism and fluidity of the digital image's expression and movement.
[0036] (3) In the multiple scene types, video is taken of the entity reference objects after the marker points are set.
[0037] After the marker points were deployed, video recording equipment was used to capture images of physical reference objects with marker points for various previously defined scenario types (taking power operation monitoring, training, news broadcasting, and customer service scenarios in the power industry as examples). Specifically, shooting parameters were adjusted according to the characteristics of different scenarios. For example, in the power operation monitoring scenario, the lens focal length might need to be adjusted to suit the large-screen display requirements, and cool-toned lighting might be set to match the monitoring environment. In the power training scenario, natural light simulation lighting was used to ensure that backgrounds such as equipment operation platforms were clearly presented. By shooting in these diverse scenarios, video images with precise movement information of the physical reference objects with marker points could be obtained in different scenarios. Through the above operations, these captured video images contained rich information such as the appearance and movements of the physical reference objects in actual power industry scenarios, providing a high-quality data source covering multiple scenarios for subsequent digital image generation. This allows the generated digital images to better adapt to various application scenarios in the power industry, improving the practicality and realism of the digital images in the power field.
[0038] Furthermore, when collecting data on the actions of physical reference objects, it is advisable to repeatedly collect each action multiple times and select the best data. It should be noted that environmental noise (such as the hum of power equipment, background personnel movement) and equipment noise (such as camera sensor noise, microphone circuit noise) can be introduced during the data collection process. In addition, a single action of a physical reference object may have non-standard deviations (for example, a power maintenance worker's single action of closing a circuit breaker may deviate from the trajectory due to slight hand tremors, and a presenter's single broadcast may be out of sync with the lip movements due to fluctuations in speech rate). By repeatedly collecting the same action, the sample with the lowest noise and the most standard action can be selected from multiple data sets.
[0039] After acquiring video images of physical reference objects, a series of preprocessing operations are required to make these raw video images better serve the subsequent training of digital image generation agents, so as to obtain clean, orderly, and clearly labeled video frames, audio signals, and related label data.
[0040] Specifically, after acquiring the video image of the entity reference object, the process includes:
[0041] (1) Separate the acquired video images into video frames and synchronously acquired audio signals.
[0042] It's important to note that video images are essentially dynamic sequences composed of continuous video frames. Simultaneously, during filming, audio signals related to the video content (such as explanations of physical references, operational instructions, etc.) are often captured concurrently. Specific audio-video separation techniques can break down a complete video image into independent video frames (static image units) and corresponding audio signals (audio data). This provides a basis for targeted processing of video and audio separately, allowing video frames and audio signals to enter different processing flows.
[0043] (2) The video frames and audio signals are filtered and noise-reduced respectively.
[0044] When acquiring video images and audio signals, some noise is inevitably introduced, such as environmental noise (e.g., the slight hum of operating electrical equipment, external background noise) and electronic noise from the recording equipment itself. Filtering and noise reduction processing utilizes signal processing algorithms. For example, for video frames, image filtering algorithms are used to remove grainy noise and stripe interference from the image; for audio signals, audio noise reduction algorithms are used to eliminate background noise. After such processing, video frames become clearer, and audio signals become purer, effectively improving the accuracy of subsequent feature extraction and other operations.
[0045] (3) Add a timestamp to each video frame and perform frame-level time alignment between the video frame and the corresponding audio signal based on the timestamp.
[0046] It's important to note that timestamps are identifiers used to mark the time when video frames and audio signals are generated. By adding a timestamp to each video frame, the video frame and its corresponding audio signal can be precisely matched in time. For example, when a physical reference object performs a specific action (corresponding to a video frame) and simultaneously emits sound (corresponding to a segment of audio signal) at a certain point in time, the timestamp ensures a one-to-one temporal correspondence. This process allows for precise synchronization between the lip movements and audio in subsequent digital avatars, preventing misalignment and significantly improving the realism of the digital avatar.
[0047] (4) Add descriptive tags to the time-aligned video frames and audio signals.
[0048] It should be noted that descriptive tags can be stored in XML format and are standardized textual descriptions of the content contained in video frames and audio signals. These tags must include at least scene information, data association information, and content semantic information. Scene information is generated based on pre-defined scene classifications and on-site environmental records before video image acquisition, ensuring a strong binding between the tags and the power scene and avoiding confusion regarding data scene attribution. Data association information includes video frame timestamps, audio segment IDs, and entity reference object IDs, generated based on the timestamps at the time of video frame separation, audio segment numbers, and entity reference object identity information, achieving precise binding between video and audio. Content semantic information includes action content, audio-text transcription, and key feature markers, generated based on the action trajectories tracked by marker points in the video frame, the audio-text transcription results, and feature judgments output by the semantic classification model (BERT model). Adding tags assigns clear semantic information to this data, facilitating the subsequent training of digital image generation agents. This allows the agent to better understand the meaning represented by different video frames and audio signals, while also facilitating data classification, retrieval, and management. Furthermore, tags can be used to filter specific data for different scenes, preventing irrelevant data from interfering with model training.
[0049] (5) Store the processed video frames, audio signals and corresponding descriptive tags.
[0050] After completing the preceding operations such as separation, noise reduction, time alignment, and labeling, the processed video frames, audio signals, and corresponding descriptive labels are stored, usually in a dedicated database or data storage system. This pre-processed, high-quality data is retained so that it can be quickly and easily accessed when training the digital image generation agent, providing a sufficient and standardized data source for the agent to learn the generation rules of digital images.
[0051] Preferably, before storing the descriptive tags, the method further includes a step of cross-validating and correcting the tag content. This correction process aims to improve the accuracy of the descriptive tags by leveraging the inherent consistency between different modal data. Specifically, the correction process includes: precisely aligning the data from different modalities on the timeline based on the same timestamp system attached to video frames and audio signals. Subsequently, using data from one modality as a benchmark, tags generated based on data from another modality are validated and corrected. For example, if the initial descriptive tag "closing operation" is generated based on a separated audio signal (e.g., recognizing the command "now execute closing"), the correction process will simultaneously retrieve the original video frames from the same time period and the extracted lip-sync images. By analyzing the lip movements using computer vision algorithms, if a definite lip movement pattern matching the "closing" command is identified, the tag is confirmed to be accurate; if the lip-sync changes are significantly inconsistent with the audio content (e.g., lip closure corresponds to a phoneme that should have an open consonant), the tag is marked as "to be verified" or corrected based on the video content. Conversely, the incorrect labels that may be generated by speech recognition can be corrected based on the operation actions recognized in the video.
[0052] This cross-modal cross-validation effectively filters out labeling errors caused by a single source of information (such as audio noise or visual occlusion), ensuring that the final stored descriptive labels truly and accurately reflect the joint semantic content of video and speech. All the video data, speech signals, and high-precision labels that have undergone unified processing and correction together constitute a high-quality dataset that eliminates random scene noise and acquisition specificity.
[0053] S102. Extract a first feature of a single dimension and a second feature of an associated dimension from the video image, wherein the second feature includes at least a joint feature of the two associated dimensions.
[0054] It's important to clarify that the first feature refers to features extracted from a single body part (such as the face, eyes, mouth, or hands) to describe the independent appearance and movement of that part. Its core purpose is to accurately depict the details of each component of the digital avatar. Specifically, the first feature can include facial contour feature vectors (structured data describing the face shape and relative positions of facial features), mouth shape feature vectors (numerical representations describing the shape of the lips at a specific moment, such as opening and closing, and protrusion), hand joint angle sequences (temporal data describing finger bending, wrist rotation, and other movements), and encoding of eye state features (such as open eyes, closed eyes, and gaze direction). The second feature refers to joint features extracted from multiple dimensions with collaborative relationships (i.e., related dimensions) to describe the interactions and constraints between these dimensions. Its core purpose is to ensure the overall coordination and naturalness of the digital avatar's movements across different parts, which is crucial for improving realism. Specifically, the second feature not only includes simple feature vector concatenation but, more importantly, includes constraint parameters representing the relationships between parts. For example, the second feature can include lip-voice association features, hand-gaze association features, and facial expression collaboration features.
[0055] The associated dimension includes at least the corresponding features of the lip shape and speech dimension. Speech signals and lip shape change images are extracted from the video images. The speech signals are converted into phoneme sequences. The phoneme sequences are matched with lip shape features to obtain multiple pairs of phoneme sequences and lip shape pairs. Supplementary lip shapes are generated according to the lip shape corresponding to adjacent phonemes. The supplementary lip shapes are used as transitional lip shapes corresponding to the lip shape corresponding to adjacent phonemes. The second feature corresponding to the lip shape and speech dimension is adjusted.
[0056] It's important to note that after acquiring and preprocessing the video images, to ensure the subsequent digital avatar generation agent can accurately learn the features of the digital avatar, key features need to be extracted from the video images. These include first features in a single dimension and second features in a related dimension. The related dimension features are particularly crucial; for example, the correspondence between lip movements and speech ensures precise synchronization between the digital avatar's lip movements and speech. Specifically, speech signals and lip movement images are extracted from the video images. The speech signals are converted into phoneme sequences, and these sequences are mapped to lip shape features, resulting in multiple pairs of phoneme sequences and lip shape pairs. Then, supplementary lip movements are generated based on the lip shapes corresponding to adjacent phonemes, thereby adjusting the second features of the lip and speech dimensions and providing precise correlation constraints for the realistic generation of the digital avatar.
[0057] Specifically, extracting a first feature of a single dimension and a second feature of a related dimension from the video image includes:
[0058] (1) Locate the entity reference object in the video image, and split the entity reference object into multiple dimensions.
[0059] It's important to note that in this step, "dimension" refers to a specific body part of the entity being referenced. Video images contain information about multiple parts of the entity, such as the face, eyes, mouth, hands, legs, and body. Dimensional decomposition involves dividing the complete video image into multiple independent dimensions based on these different parts. For example, a video image containing the overall digital image can be decomposed into separate dimensions such as face, eyes, mouth, and hands. By obtaining multiple dimensions, targeted feature extraction can be performed on each part, making the feature extraction more detailed and accurate.
[0060] (2) Traverse each dimension. For each dimension, locate the sub-video frames of the entity reference part region corresponding to the dimension in the video image to form the sub-video image of each dimension.
[0061] Each frame primarily contains the target body part itself, as well as its anatomically connected, related regions necessary for movement of that part. After dimensional decomposition, each dimension is processed individually. For each dimension, target detection techniques from computer vision are used to precisely locate the corresponding entity reference region in the video image, and then a sub-video image containing only that part is extracted. For example, for the mouth dimension, the region where the mouth is located in the video image is located, and a sub-video image of the mouth is extracted.
[0062] It's important to clarify that the localization and extraction here isn't simply rectangular cropping. Instead, it involves precise image segmentation to ensure that each extracted sub-video frame meets specific content constraints. For example, for the hand dimension, the sub-video frame should clearly and completely contain the hand. To achieve continuity and naturalness in the movements, the sub-video frame also needs to include related parts that are anatomically directly connected to the target area and change collaboratively during its movement. For instance, when extracting sub-video images of the hand dimension, in addition to the hand itself, it should also include a portion of the wrist and even the forearm area to fully capture the biomechanical chain of actions such as grasping and waving; when extracting sub-video images of the mouth dimension, it should include the chin and part of the cheek area to accurately reflect the coordination of surrounding muscles during mouth shape changes. Sub-video images extracted in this way focus on the independent features of the target area while preserving its local collaborative information with related parts.
[0063] (3) Based on the semantic expression of the sub-video image and the pixel difference between adjacent frames, determine the key frame, and extract the feature vector of the corresponding part region in the key frame as the first feature.
[0064] The semantic expression of a sub-video image refers to the meaning of the content it represents. For example, in a sub-video image of the mouth, is it a smiling or speaking gesture? The pixel differences between adjacent frames reflect the dynamic changes in the body part. By analyzing the semantic expression of sub-video images (determining whether a frame contains key actions or expressions) and the pixel differences between adjacent frames (large differences indicate significant action changes), keyframes are selected. Then, feature extraction algorithms (such as convolutional neural networks) are used to extract feature vectors for the corresponding body parts within these keyframes. These feature vectors constitute the first feature in a single dimension, which accurately describes the characteristics of a single body part of the digital image.
[0065] When determining keyframes based on the semantic representation of the sub-video images and the pixel differences between adjacent frames, it is particularly necessary to employ a difference calculation and semantic content detection mechanism to extract keyframes for long video data (such as power scene videos with a single scene shooting duration of ≥30 minutes). Specifically, for the sub-video images corresponding to the long video, the pixel differences between adjacent frames are calculated frame by frame according to the time series. Specifically, a grayscale difference algorithm can be used to calculate the average of the absolute difference of the grayscale values of corresponding pixels in two frames. A difference threshold is set, and when the average value is greater than the difference threshold, the frame is determined to be a frame with significant action changes and is initially selected as a candidate set of keyframes. Furthermore, the candidate set of keyframes is further filtered using a semantic content detection algorithm to accurately capture keyframes containing important actions and expressions in the long video.
[0066] (4) Determine multiple combinations of related dimensions. For each combination of related dimensions, based on the intersection of the key frames corresponding to the related dimensions, filter the key frames that simultaneously contain two dimension parts.
[0067] It should be noted that a related dimension combination refers to a combination of two dimensions that are related, such as lip shape and voice, eye shape and mouth. For each such related dimension combination, the intersection of the keyframes corresponding to the two dimensions is found. That is, those frames that are keyframes in both dimensions simultaneously. In this way, keyframes that contain parts of both dimensions are selected. This is done to find keyframes that can reflect the common features of the two related parts, preparing for the extraction of the second feature of the related dimension.
[0068] (5) Simultaneously extract the feature vectors of the two dimensional regions in the keyframe to form the constraint features of the associated dimension, and obtain the second feature.
[0069] From the keyframes that simultaneously contain parts of both dimensions, feature vectors are extracted from the regions of both dimensions. These feature vectors are then combined to form the constraint features of the correlation dimension, which is the second feature of the correlation dimension. The second feature reflects the relationship between the two correlation dimensions. For example, the second features of the lip-sync dimension and the speech dimension can constrain the lip-sync of the digital image to synchronize with the speech, thereby making the generated digital image more coordinated and realistic in the performance of the correlation parts.
[0070] Furthermore, in the process of generating supplementary lip shapes based on the lip shapes corresponding to adjacent phonemes, a dynamic transition algorithm based on key point localization and position interpolation is adopted to ensure the smoothness and naturalness of the lip shape changes. Specifically, firstly, for the standard lip shape (i.e., the basic lip shape) corresponding to each phoneme, multiple predefined key points on its lip contour (e.g., the center points of the upper and lower lips, the left and right corners of the mouth, the cupid's bow, etc.) are located using a facial feature point detection model. The two-dimensional or three-dimensional coordinates of these key points constitute the digital representation of the basic lip shape. Then, for any two adjacent phonemes (denoted as phoneme A and phoneme B), their corresponding basic lip shape key point sets are obtained. The corresponding key points in these two sets (such as the upper lip center point pair A and the upper lip center point pair B) are paired to form a series of dynamic point pairs. Furthermore, within the time interval from phoneme A to phoneme B, for each dynamic point pair, based on the preset number of transition frames, linear interpolation or a smoother cubic spline interpolation algorithm is used to calculate the transition position that the keypoint should be in each intermediate frame. During this process, the motion trajectory of all keypoints is calculated synchronously and independently by the interpolation algorithm. The position of each keypoint in any transition frame is precisely calculated through its start and end positions, meaning that many points follow fixed calculation paths. Finally, the new positions of all keypoints in all transition frames are reconnected to form a continuous lip contour sequence. This sequence is the generated supplementary lip shape located between the lip shapes of phoneme A and phoneme B, ensuring that the lip shape change from A to B is continuous, smooth, and conforms to human oral kinematics.
[0071] S103. Using the first feature and the shooting scene corresponding to the video image as learning objects, and the second feature as constraint rules, train a digital image generation intelligent agent. The input of the digital image generation intelligent agent includes at least scene data, and the output includes at least digital images.
[0072] It should be noted that after obtaining the first feature of a single dimension, the second feature of the associated dimension, and clarifying the shooting scene corresponding to the video image, a digital image generation agent needs to be trained in order for the digital image to generate appropriate content according to different scenes. During training, the first feature and the shooting scene are used as the learning objects of the agent, and the second feature is used as the constraint rule, so that the generated digital image not only meets the scene requirements, but also maintains coordination in the features of each part and the associated features.
[0073] Specifically, the training of the digital image generation agent includes:
[0074] (1) Construct training samples, which include shooting scene data, first features and corresponding second features of each component, and corresponding historical digital images.
[0075] It should be noted that the shooting scene data encompasses information such as scene type (e.g., power operation monitoring, training, etc.) and environmental characteristics (e.g., background, lighting, etc.). This shooting scene data serves as the conditional input to the model when constructing training samples. The model's learning objective (i.e., the desired output) is a digital image instance that closely matches the scene data and has historically proven to be effective in practical applications (e.g., the professional announcer image with the highest user rating and the highest frequency of use in power operation monitoring scenarios).
[0076] The first features of each component and the second features of the associated dimensions are not used as direct inputs to the model. Instead, they serve as refined intermediate supervision signals and target outputs that the feature generation layer within the model needs to fit. This means that the training objective of the agent is to learn a generative mapping from scene data to a set of coordinated and realistic component features (first features) and their constraint relationships (second features). Ultimately, by decoding and combining these generated features, a complete and dynamic digital image is rendered.
[0077] By integrating these input conditions (scene data), intermediate supervision and output targets (first and second features), and the final output target (optimized digital image instances) into a training sample, a multi-layered and strongly guided complete data foundation is provided for the subsequent training of the agent. This sample construction method provides precise guidance for the feature generation process within the model. The model not only needs to output the final image, but its internal feature generation layer is also required to output high-quality feature combinations consistent with the optimized samples. This mechanism, where the optimized results guide feature generation in reverse, enables the model to understand more deeply the intrinsic relationship between the scene and the features of the image components, rather than simply learning a surface mapping. This significantly improves the component accuracy, overall coordination, and scene fit of the generated digital image.
[0078] (2) Enhance the facial expressions and motions of the training samples.
[0079] Facial enhancement and motion enhancement are data augmentation techniques used to enrich and expand the facial expressions and movements of digital avatars in training samples. For example, for a sample that originally shows a smiling expression, an enhanced sample is generated with different amplitudes of smiles, smiles accompanied by slight head movements, etc.; for motion samples, versions with slightly different speeds and amplitudes of movements are generated. This increases the diversity of training samples, enabling the digital avatar generation agent to generate more diverse and natural facial expressions and movements after training.
[0080] (3) Take the scene data as input and construct action-related constraint relationships in combination with the second feature.
[0081] It should be noted that the digital image generation agent is essentially a deep learning-based generative model. Its core function is to act as a generator, outputting a complete and coordinated set of digital image driving parameters based on the input scene data. The core logic of this generation process is that the final digital image is composed of the driving parameters (i.e., the first feature) of its various components, and the coordination and naturalness between these component driving parameters are controlled by the constraint relationship defined by the second feature.
[0082] Specifically, during model training and inference, scene data serves as the conditional input, determining the overall style and context of the digital avatar generation (e.g., in a power training scenario, a professional and rigorous operator avatar needs to be generated). The first feature, as the generation target, is the specific parameter that the model needs to predict to drive the various components of the digital avatar (such as facial muscles and hand joints). The second feature, as a constraint rule, is encoded into the model's loss function or network structure to ensure that the first features of the generated components satisfy physical and semantic synergistic relationships. For example, using the lip-speech second feature, temporal synchronization constraints are constructed in the model to ensure that the generated lip-driving parameters are precisely aligned with the input speech signal in time; using the hand-gaze second feature, kinematic association constraints are constructed in the model to ensure that when the hand-driving parameters manifest as pointing movements, the eye-driving gaze direction parameters can synchronously point to the same target area; using the facial expression coordination second feature, muscle linkage constraints are constructed in the model to ensure that the driving parameters of the generated facial areas such as the corners of the mouth, eyes, and eyebrows conform to the real facial muscle movement patterns.
[0083] By deeply integrating scene data with the second feature, the model, when generating the first feature of each component, not only considers the macro requirements of the scene, but is also strictly constrained by the micro-coordination relationship between the components. This mechanism ensures that the actions of the various components of the final digital image are coherent, unified, and highly consistent with the needs of the scene, thereby fundamentally solving the technical problem of component actions being disconnected and inconsistent with the scene context.
[0084] (4) Based on the scene data, the first feature and the action-related constraints, learn the generative mapping relationship between the scene features and the first feature. The relationship between each first feature is constrained by the relevant second feature, and generate the digital image generation intelligent agent.
[0085] It should be noted that scene data contains scene feature information. The first feature is the feature of each component, and the action-related constraint relationship is the constraint rule between the two. Through specific machine learning algorithms (such as generative adversarial networks, variational autoencoders, etc.), the agent learns the generative mapping relationship between scene features and component features. That is, it learns how to generate a combination of component features that meets the constraints based on different scene features and component features, thus laying the foundation for the subsequent generation of a complete digital image.
[0086] When learning the generative mapping relationship between scene features and component features based on the scene data, the first feature, and action-related constraints, joint feature learning is needed to achieve effective fusion of different modal features and improve the overall performance of the digital image. This involves determining the importance of different modal features through feature selection and weighting methods to improve the feature fusion effect. Specifically, a multimodal joint learning network can be constructed, using the first feature of the video modality, the speech features of the audio modality, and the scene data features of the scene modality as network inputs. A cross-modal attention mechanism is used to learn the correlation between different modal features (such as the temporal correlation between lip-sync features and speech features, and the semantic correlation between action features and scene features), outputting preliminary fused features. Further, an L1 regularization algorithm is used to sparsify the preliminary fused features, removing redundant features. Then, the importance weights of the core features of each modality are calculated based on the random forest algorithm, and the final fused features are obtained through weighted summation. Finally, the final fused features are input into Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs), and combined with action-related constraints, the generative mapping relationship between scene features and component features is learned to ensure that the generated digital image has better overall performance under the synergistic effect of multimodal features.
[0087] Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs) are two model frameworks well-suited for training and generating digital images. GANs excel at generating high-quality, realistic images or video sequences, making them suitable for digital image generation tasks requiring highly complex and nuanced representations, often used in face generation and animated character generation. VAEs, on the other hand, learn the latent representations of data, generating diverse and continuous samples, making them suitable for digital image generation tasks requiring control over the features of generated samples. For even more complex scenarios, deep generative models (such as deep autoencoders) combine the powerful representation learning capabilities of deep learning with the generative capabilities of generative models. They can learn and generate complex data distributions, producing high-quality samples and handling multimodal data inputs.
[0088] In addition, during training, a special loss function needs to be designed to ensure the quality of generation. The loss function consists of three parts: reconstruction loss (using mean squared error to calculate the difference between the generated features and the real features), constraint loss (based on the constraint relationship defined by the second feature, calculating the coherence loss between the features of each component), and adversarial loss (when using a generative adversarial network, the adversarial signal provided by the discriminator is used to improve the realism of the generation).
[0089] Optionally, an unsupervised training method based on reinforcement learning can be used to model the generation process as a sequential decision problem, and to guide the model learning by designing reward and punishment functions. Specifically, this includes reward signal design (giving a positive reward when the generated feature combination is close to the distribution of real data), constraint satisfaction reward (giving an additional reward when the generated component features satisfy the constraint relationship defined by the second feature), and diversity reward (giving appropriate rewards for the diversity of generated features to prevent mode collapse).
[0090] It should be noted that the agent is continuously trained using training samples, allowing it to learn how to combine the features of each component in a reasonable way to generate a complete digital image that meets the requirements of the scene and constraints. After sufficient training, a digital image generation agent is finally generated that can output the corresponding digital image based on the input scene data.
[0091] It's important to note that during training, the quantity and quality of training samples, as well as the number of training iterations and the form of the loss function (determined based on the specific circumstances), also need to be considered. Specifically, automated hyperparameter tuning tools or experimental design methods can be used to determine the optimal combination. Automated hyperparameter tuning tools include Hyperopt, which uses a sequential model optimization algorithm to optimize hyperparameters, combining Bayesian optimization and random search principles to effectively optimize complex parameter spaces. Automated machine learning tools like AutoML, H2O.ai, and TPOT go beyond hyperparameter tuning; they can also automate feature engineering, model selection, and the entire machine learning process. This significantly reduces the errors of manual hyperparameter tuning and improves model performance and efficiency.
[0092] It should be noted that after the model training is completed, there is also a usage process. Specifically, in the model usage phase, the scene data of the digital image to be generated is first converted into scene feature vectors according to a preset format, which serve as the input to the agent. The Transformer decoder inside the agent first parses the scene feature vectors, and combined with the scene-feature mapping relationship learned in the training phase, outputs the first feature and second feature constraints of each component of the target digital image. Subsequently, the agent calls the pre-built digital image component library (which is from the same source as the component library in the training samples to ensure feature-component matching), selects the corresponding component according to the first feature, and adjusts the component parameters according to the second feature constraints. Finally, the components and parameters are passed to the terminal rendering engine through the built-in rendering interface of the agent, and the rendering engine generates the final displayable digital image according to the terminal hardware configuration.
[0093] S104. Input the scene data of the digital image to be generated into the digital image generation intelligent agent to generate a customized digital image.
[0094] After training the digital avatar generation agent, the scene data for which the digital avatar is to be generated can be input into the agent to generate a customized digital avatar. Furthermore, to ensure the generated digital avatar is better adapted to different terminal platforms and presents a more complete scene effect, subsequent operations such as terminal adaptation and scene fusion are required.
[0095] Specifically, the step of inputting the scene data of the digital avatar to be generated into the digital avatar generation intelligent agent to generate a customized digital avatar includes:
[0096] (1) The digital image generation agent predicts the set of components required for the target digital image and the action sequence of each component based on the scene data.
[0097] It should be noted that after receiving scene data (such as power operation monitoring scene data) for the digital image generation intelligent agent, it will predict the set of components (such as facial components, hand components, etc.) required to build the target digital image, as well as the action sequence of each component (i.e., what action each component will perform at what time) based on the correspondence between the scene and the digital image components and actions that it has learned. This provides a clear basis for the subsequent retrieval and combination of components.
[0098] (2) Retrieve the component set from the pre-built digital image component library and determine the component combination relationship at each time point according to the predicted time sequence.
[0099] The pre-built digital avatar component library stores various digital avatar components (such as different facial expression components, different gesture action components, etc.). Based on the component set predicted by the agent, the corresponding components are retrieved from the component library. Then, according to the predicted action sequence, it is determined which components need to be combined at each time point, thereby clarifying the combination relationship of the components.
[0100] (3) Based on the action prediction results output by the digital image-generated intelligent agent, generate the transmission parameters of the component interface at each time point, and drive the component to perform the corresponding action.
[0101] The action prediction results output by the digital image generation agent contain specific information about the component's actions. Based on these results, the transmission parameters of the component interface at each time point are generated (these parameters are used to control the component's action execution, such as the magnitude and speed of the action). Then, these transmission parameters are used to drive the corresponding component to perform the corresponding action, so that the component moves in the expected way.
[0102] (4) The combined actions of the components are spliced together in sequence to form a customized digital image.
[0103] After each component executes its corresponding action according to the passed parameters, the combined actions of these components are sequentially assembled according to the action sequence to form a complete, customized digital image that meets the requirements of the scenario.
[0104] After generating a customized digital avatar, terminal adaptation is required to ensure that the digital avatar can be displayed well on different terminal platforms (such as large screens, mobile phones, etc.).
[0105] It should be noted that the process of inputting the scene data for generating the digital avatar into the digital avatar generating agent includes:
[0106] (1) Obtain the hardware configuration parameters and display specifications of the target terminal platform.
[0107] The hardware configuration parameters (such as processor performance and memory size) and display specifications (such as screen resolution and aspect ratio) of the target terminal platform will affect the presentation of the digital avatar. This information can be obtained through relevant technical means (such as communicating with the terminal platform) to allow for targeted adjustments to the digital avatar later.
[0108] (2) The hardware configuration parameters and display specifications are transmitted to the rendering engine.
[0109] The rendering engine is the core module responsible for rendering and displaying digital images. Passing the obtained hardware configuration parameters and display specification information to the rendering engine allows it to understand the capabilities and requirements of the terminal platform.
[0110] (3) Adjust the size, resolution and format of the digital image output by the digital image generation agent according to the hardware configuration parameters and display specifications.
[0111] Based on the hardware configuration parameters and display specifications of the terminal platform, the generated customized digital avatar is adjusted. For example, if the terminal is a small-screen mobile phone, the size of the digital avatar is reduced, the resolution is adjusted to match the resolution of the mobile phone screen, and the format is also adjusted to a format supported by the mobile phone to ensure that the digital avatar can be displayed clearly and normally on the terminal.
[0112] It should be noted that the customized digital avatars output by the digital avatar generation intelligent agent are essentially the prediction results of component combination logic and action timing parameters (not directly displayable images / animations). When actually using this avatar, it needs to be constructed based on a generalized component library and unified templates. Specifically, a pre-built general component library for digital avatars covering all scenarios in the power industry is constructed. This library contains standardized basic components (such as facial components: facial templates for different face shapes / skin tones; limb components: torso templates in maintenance uniform / workwear styles; action components: power-specific action templates such as closing / inspection). All components are stored in a unified format (such as FBX) and have reserved standardized interfaces (supporting parameterized adjustments to size, color, and action amplitude). Based on the prediction results output by the intelligent agent, the corresponding general components are called from the component library, ensuring consistency and compatibility of components called in different scenarios and on different terminals. Furthermore, a general digital avatar construction template is configured. This template defines component combination rules (such as the connection coordinates between facial and torso components, and the kinematic relationship between motion and limb components) and terminal adaptation benchmarks (such as the scaling ratio of components at different resolutions and encoding parameters under different formats). The called general components are imported into the template, and component combination is automatically completed according to the template rules (such as aligning facial and torso components according to neck coordinates, and binding motion and hand components according to joint parameters). Simultaneously, considering the target terminal's hardware configuration and display specifications, the template's built-in parameter adjustment module synchronously completes component size fine-tuning and resolution adaptation. After component combination and template adaptation are completed, the rendering engine generates visualized content from the combined components according to motion timing parameters. For high-performance terminals, a high-resolution (e.g., 4K) real-time animation stream is directly generated; for mobile terminals (such as maintenance phones), lightweight sequence frames (e.g., sprites, single frame size ≤ 50KB) are generated. The final output is a displayable digital avatar that perfectly matches the terminal hardware and display specifications. The entire process relies on a general component library and a unified template to avoid the repeated development of components due to differences in scenarios or terminals. It enables a single prediction result to be used universally across multiple terminals, while ensuring that the digital images displayed on different terminals maintain consistency in appearance and actions, in line with the standardized interaction requirements of the power industry.
[0113] Furthermore, to improve the efficiency and effectiveness of digital avatar display on the terminal, a cloud-edge collaborative approach is adopted for optimization. It should be noted that after inputting the scene data of the digital avatar to be generated into the digital avatar generation intelligent agent to generate a customized digital avatar, the process also includes:
[0114] (1) Send the component action timing data and rendering request corresponding to the customized digital image to the cloud processing node.
[0115] The timing data of the component actions corresponding to the customized digital image (i.e., the time sequence of the actions of each component) and the rendering request are sent to the cloud processing node. The cloud processing node has strong computing power and can perform more complex processing on this data.
[0116] (2) The cloud processing node calls the part action component corresponding to the time series data from the pre-built component library, performs combination operation, generates a lightweight digital image representation model, and performs pre-calculation of multi-channel rendering data on the digital image representation model.
[0117] The cloud processing node calls the corresponding part action components from its own pre-built component library and performs combined calculations on these components to generate a lightweight digital image representation model (which has a smaller data volume while ensuring a certain effect, making it easier to transmit and process). At the same time, it performs pre-calculation of multi-channel rendering data on this lightweight model to complete part of the rendering work in advance and reduce the burden on the edge processing nodes.
[0118] In this step, lightweight digital image representation models, such as pruning, quantization, and distillation, are used to improve computational efficiency and speed. Parallel computing and GPU acceleration are utilized to speed up lip-sync calculation and animation generation, ensuring that performance requirements can be met even in real-time applications.
[0119] (3) Package the digital image representation model and the pre-calculated rendering data and send them to the edge processing node.
[0120] The cloud processing node packages the generated lightweight digital image representation model and pre-computed rendering data, and then sends it to the edge processing node. The edge processing node is closer to the terminal, which can respond to the terminal's rendering needs more quickly.
[0121] (4) The edge processing node performs calculation and real-time rendering on the received data packets according to the display specification information of the target terminal, and generates a final digital image that matches the target terminal.
[0122] After receiving the data packet, the edge processing node decodes the data packet according to the display specifications of the target terminal, then performs real-time rendering, and finally generates a digital image that matches the target terminal, ensuring that the digital image can be displayed on the terminal in real time and smoothly.
[0123] Furthermore, it should be noted that after generating the customized digital avatar, a scene blending operation is required to place the digital avatar in an environment that better fits the scene. Specifically, the process after generating the customized digital avatar includes:
[0124] (1) Fill the customized digital image with preset background constraints and additional component constraints based on the standardized template.
[0125] It should be noted that the standardized template includes background constraints (such as background type and style) and additional component constraints (such as the type and position of auxiliary components to be added) for different scenarios. Based on the standardized template, fill in these preset constraints for the customized digital image and determine the general requirements of the background and additional components of the digital image.
[0126] (2) Based on the constraints of the template, match and generate an appropriate background scene from the pre-built background resource library.
[0127] The pre-built background resource library stores various background scene resources (such as the large screen background for power operation monitoring, the practical platform background for power training, etc.). According to the background constraints in the template, a suitable background scene is matched from the background resource library and generated to provide a suitable background environment for the digital image.
[0128] (3) Call additional components that match the customized digital image and background scene from the pre-built digital image component library.
[0129] Similarly, the pre-built digital avatar component library also contains various additional components (such as power equipment model components, text description components, etc.). Based on the characteristics of the customized digital avatar and the generated background scene, matching additional components can be called from the component library, which can enrich the display scene of the digital avatar.
[0130] (4) Adjust the display parameters of the background scene and the additional components, and merge the customized digital image, background scene and additional components to generate the final digital image display content containing the complete scene.
[0131] Adjust the display parameters (such as size, position, transparency, etc.) of the background scene and additional components to make them blend well with the customized digital avatar. Then, merge the customized digital avatar, the adjusted background scene and additional components to finally generate digital avatar display content containing the complete scene, making the digital avatar display richer and more realistic.
[0132] The method provided in this embodiment accurately captures the appearance features and action data of entities in the power industry by setting high-contrast markers and tracking changes when acquiring video images of physical reference objects, combined with multi-scene differentiated shooting parameter adjustments. Audio-video separation, filtering and noise reduction, and frame-level time alignment in the preprocessing stage further ensure data quality. Feature extraction involves splitting individual dimensions and related dimensions, especially designing phoneme sequence mapping and transitional lip-syncing generation logic for the lip-sync-speech dimension. This effectively solves the problems of lip-sync and speech misalignment and action disconnect in traditional technologies, making the facial expressions, body movements, and voice interactions of digital avatars closely resemble real-world scenarios. The agent is trained using the first feature and the shooting scene as learning objects and the second feature as constraints. Combining facial expressions and actions enhances sample diversity, enabling the agent to learn the correspondence between scene and image features. Subsequent input of scene data can output customized digital avatars, breaking through the limitations of traditional digital avatars being singular, fixed, and disconnected from scenes. This allows for accurate matching of action and appearance requirements in different scenarios such as power operation monitoring, training, and broadcasting. Meanwhile, this method achieves seamless display and efficient interaction across multiple terminals through terminal adaptation and cloud-edge collaborative optimization: for different terminals such as large screens, mobile phones, and tablets, it acquires hardware and display parameters and adjusts the size, resolution, and format of the digital image; by leveraging cloud-generated lightweight models and pre-calculated rendering data, and real-time edge node rendering, it reduces the terminal's computing burden while ensuring smooth digital image display and controllable latency; in the scene integration stage, it matches backgrounds and additional components based on standardized templates, further integrating the digital image with the power scenario environment, ultimately meeting the core needs of the power industry such as real-time operation and monitoring broadcasts, training demonstrations, and precise customer service interaction, providing high-quality digital image support for the digital transformation of the power industry.
[0133] Example 2
[0134] Corresponding to the aforementioned embodiment of a multi-terminal, multi-scene adaptive digital image generation method, this application also provides an embodiment of a multi-terminal, multi-scene adaptive digital image generation device.
[0135] Figure 2 This is a schematic diagram of the structure of Embodiment 2 of the multi-terminal, multi-scenario adaptive digital image generation device provided in this application. Please refer to... Figure 2 The device provided in this embodiment includes a data acquisition module 210, an extraction module 220, a training module 230, and a generation module 240.
[0136] The acquisition module 210 is used to acquire video images of the physical reference object;
[0137] The extraction module 220 is used to extract a first feature of a single dimension and a second feature of an associated dimension from the video image, wherein the second feature includes at least a joint feature of the two associated dimensions;
[0138] The associated dimension includes at least the corresponding features of the lip shape and speech dimension. The speech signal and lip shape change image are extracted from the video image. The speech signal is converted into a phoneme sequence. The phoneme sequence is matched with the lip shape feature to obtain multiple pairs of phoneme sequences and lip shape pairs. Supplementary lip shapes are generated according to the lip shape corresponding to adjacent phonemes. The supplementary lip shapes are used as the transitional lip shapes corresponding to the lip shapes corresponding to adjacent phonemes. The second feature corresponding to the lip shape and speech dimension is adjusted.
[0139] The training module 230 is used to train a digital image generation agent with the first feature and the shooting scene corresponding to the video image as learning objects and the second feature as constraint rules. The input of the digital image generation agent includes at least scene data and the output includes at least digital image.
[0140] The generation module 240 is used to input the scene data of the digital image to be generated into the digital image generation intelligent agent to generate a customized digital image.
[0141] The apparatus of this embodiment can be used to perform... Figure 1 The steps of the method embodiment shown are similar in principle and process, and will not be repeated here.
[0142] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0143] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0144] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A multi-terminal, multi-scene adaptive digital figure generation method, characterized in that, The method comprises: collecting a video image of an entity reference object; extracting a first feature of a single dimension and a second feature of an associated dimension from the video image, the second feature comprising at least a joint feature of two associated dimensions; the extraction of the first feature of a single dimension and the second feature of an associated dimension from the video image comprises: locating the entity reference object in the video image, dimensionally splitting the entity reference object to obtain multiple dimensions; traversing each dimension, for each dimension, locating a sub-video frame of a part region of the entity reference object corresponding to the dimension in the video image to form a sub-video image of each dimension; determining a key frame based on the semantic expression of the sub-video image and the pixel difference of adjacent frames, and extracting a feature vector of the corresponding part region in the key frame as the first feature; determining a plurality of associated dimension combinations, for each associated dimension combination, filtering key frames containing part regions of two dimensions based on the intersection of key frames corresponding to the two dimensions; simultaneously extracting feature vectors of the two dimension part regions in the key frames to form the constraint feature of the associated dimension, and obtaining the second feature; wherein the associated dimension comprises at least corresponding features of mouth shape and speech dimension, and the method further comprises: extracting a speech signal and a mouth shape change image from the video image, converting the speech signal into a phoneme sequence, corresponding the phoneme sequence with the mouth shape feature to obtain a plurality of pairs of phoneme sequence and mouth shape, generating a supplementary mouth shape according to the mouth shape corresponding to adjacent phoneme pairs, taking the supplementary mouth shape as a transition mouth shape corresponding to the mouth shape corresponding to adjacent phoneme pairs, and adjusting the second feature corresponding to the mouth shape and speech dimension; taking the first feature and the shooting scene corresponding to the video image as learning objects, and taking the second feature as a constraint rule, training a digital avatar generation agent, wherein the input of the digital avatar generation agent at least comprises scene data, and the output at least comprises a digital avatar; inputting scene data of a digital avatar to be generated into the digital avatar generation agent to generate a customized digital avatar.
2. The method of claim 1, wherein, The method comprises: determining a plurality of scene types and a plurality of entity reference objects used in the past in each type of scene, and the plurality of entity reference objects in the same type of scene have different appearances; arranging marker points on the face and action of the entity reference object, and tracking the changes of the marker points; video shooting the entity reference object with arranged marker points in the plurality of scene types.
3. The method of claim 1, wherein, The method comprises: constructing a training sample, wherein the training sample comprises shooting scene data, first features of each component and corresponding second features, and corresponding historical digital avatars; performing expression enhancement and action enhancement on the training sample; taking the scene data as input, combining the second features to construct action-related constraint relationship; learning a generative mapping relationship between scene features and first features based on the scene data, the first features and the action-related constraint relationship, the relationship between each first feature being constrained by the related second features, and generating the digital avatar generation agent.
4. The method of claim 1, wherein, The method comprises: Separate the collected video images into video frames and synchronously collected voice signals; Filter and denoise the video frames and voice signals respectively; Add a timestamp to each video frame and perform frame-level time alignment between the video frames and corresponding voice signals based on the timestamp; Add descriptive labels to the time-aligned video frames and voice signals; Store the processed video frames, voice signals, and corresponding descriptive labels.
5. The method of claim 1, wherein, After the scene data of the digital figure to be generated is input into the digital figure generation agent, the following steps are included: Obtain hardware configuration parameters and display specification information of a target terminal platform; Transfer the hardware configuration parameters and display specification information to a rendering engine; Adjust the size, resolution, and format of the digital figure output by the digital figure generation agent according to the hardware configuration parameters and display specification information.
6. The method of claim 1, wherein, After the scene data of the digital figure to be generated is input into the digital figure generation agent to generate a customized digital figure, the following steps are included: Send component action timing data corresponding to the customized digital figure and a rendering request to a cloud processing node; The cloud processing node calls part action components corresponding to the timing data from a pre-built component library, performs combination operation to generate a lightweight digital figure performance model, and precomputes multi-channel rendering data for the digital figure performance model; Pack the digital figure performance model and the precomputed rendering data and issue them to an edge processing node; The edge processing node performs calculation and real-time rendering on the received data packet according to the display specification information of the target terminal to generate a final digital figure matching the target terminal.
7. The method of claim 1, wherein, After the scene data of the digital figure to be generated is input into the digital figure generation agent to generate a customized digital figure, the following steps are included: The digital figure generation agent predicts a component set required by the target digital figure and the action timing of each component based on the scene data; The component set is called from a pre-built digital figure component library, and the component combination relationship at each time point is determined according to the predicted timing; According to the action prediction result output by the digital figure generation agent, the transfer parameters of the component interface at each time point are generated, and the components are driven to perform corresponding actions; The component combination actions are spliced in sequence to form a customized digital figure.
8. The method of claim 1, wherein, After the customized digital figure is generated, the following steps are included: Fill the customized digital figure with preset background constraints and additional component constraints based on a standardized template; According to the constraints of the template, an adaptive background scene is matched and generated from a pre-built background resource library; Call additional components matching the customized digital figure and the background scene from a pre-built digital figure component library; Adjust the display parameters of the background scene and additional components, and fuse the customized digital figure, background scene, and additional components to generate a final digital figure display content containing a complete scene.
9. A multi-terminal, multi-scene adaptive digital figure generation device, characterized in that, The device includes a collection module, an extraction module, a training module, and a generation module; The collection module is configured to collect video images of entity reference objects; The extraction module is configured to extract single-dimension first features and associated-dimension second features from the video image, the second features including at least joint features of two associated dimensions; the extraction of the single-dimension first features and the associated-dimension second features from the video image includes: locating an entity reference in the video image, dimensionally splitting the entity reference to obtain multiple dimensions; traversing each dimension, for each dimension, locating a sub-video frame of a part region of the entity reference corresponding to the dimension in the video image to form a sub-video image of each dimension; determining a key frame based on semantic expression of the sub-video image and pixel difference of adjacent frames, and extracting a feature vector of the corresponding part region in the key frame as the first feature; determining multiple associated-dimension combinations, for each associated-dimension combination, filtering key frames containing part regions of two dimensions based on an intersection of key frames corresponding to the two dimensions; and extracting feature vectors of the two part regions in the key frames to form constraint features of the associated dimensions, thereby obtaining the second features; The associated dimensions include at least corresponding features of a mouth shape and a speech dimension, a speech signal and a mouth shape change image are extracted from the video image, the speech signal is converted into a phoneme sequence, the phoneme sequence is matched with the mouth shape feature to obtain multiple pairs of phoneme sequences and mouth shape pairs, a supplementary mouth shape is generated according to mouth shapes corresponding to adjacent phoneme pairs, the supplementary mouth shape is used as a transition mouth shape corresponding to the mouth shape corresponding to the adjacent phoneme pair, and the second features corresponding to the mouth shape and the speech dimension are adjusted. The training module is configured to take the first features and a shooting scene corresponding to the video image as learning objects, and take the second features as constraint rules, to train a digital figure generation agent, wherein input of the digital figure generation agent includes at least scene data, and output of the digital figure generation agent includes at least a digital figure. The generation module is configured to input scene data of a digital figure to be generated into the digital figure generation agent to generate a customized digital figure.
Citation Information
Patent Citations
Virtual face generation method
CN113781610A
Generation method and generation equipment of digital human interaction video
CN119402720A