Multi-terminal and multi-scene adaptive digital image generation method and device
By collecting and extracting video image features of physical reference objects in the power industry, a digital image generation intelligent agent is trained, which solves the problem of poor model reusability in the power industry, realizes the generation of digital images with high realism and multi-scene adaptability, and improves the interaction effect and efficiency of the power industry.
Patent Information
- Application Number
- CN202511681639.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-17
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2045-11-17
AI Technical Summary
Existing digital image generation technologies in the power industry suffer from poor model reusability, long development cycles, high costs, and insufficient accuracy and generalization capabilities, making it difficult to meet the requirements of seamless adaptation across multiple terminals and scenarios, as well as the high realism demands of the power industry.
By collecting video images of physical reference objects, extracting features in single and related dimensions, training a digital avatar generation agent, and using related dimension features to constrain lip-sync with speech, a customized digital avatar is generated that can adapt to multiple scenarios and multi-terminal environments.
It has achieved highly realistic generation of digital avatars, standardized production, and seamless adaptation to multiple terminals and scenarios, which has improved the naturalness and professionalism of interaction in the power industry and reduced technical barriers and costs.
Smart Images

Figure CN121120886A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a multi-terminal and multi-scene adaptive digital figure generation method and device. BACKGROUND
[0002] With the deep application of artificial intelligence technology in the power industry, digital figures have become the core interactive carrier in power operation monitoring data reporting, equipment operation training, customer service consultation and other scenarios. The industry has increasingly urgent needs for high fidelity, multi-terminal adaptation and seamless multi-scene interaction of digital figures. Current digital figure generation technology mainly constructs a technical system around data collection, feature extraction, model training and terminal display. High-definition camera equipment is used to collect entity reference video, and algorithms such as convolutional neural networks and generative adversarial networks are used to extract features and train a generation model. Then, relying on a rendering engine, the format of the digital figure is adjusted to adapt to different terminals, and the scene application of digital figures is initially realized.
[0003] However, the existing technology still has obvious shortcomings in the vertical scenarios of the power industry. The current solution requires training a digital figure model for each scenario, i.e., one model for one scenario and one type of digital figure. For example, a dedicated reporting figure model needs to be trained for the power operation monitoring scenario, and a device operation figure model needs to be retrained for the power training scenario. The parameters and features of the models cannot be reused, which not only leads to a long development cycle and high labor costs, but also makes it difficult to meet the needs of rapid expansion of scenarios in the power industry. On this basis, on the one hand, the data collection accuracy and scene generalization ability are insufficient, and ordinary camera equipment is mainly used, lacking high-precision capture means such as depth sensors. This results in large capture errors of key information such as facial micro-expressions and device operation gestures, and the entity reference samples in the same scenario are single, making it difficult to adapt to the diverse appearance and operation habits of personnel in the power industry. On the other hand, the correlation dimension feature synchronization and model practicality are not good. The mouth-speech synchronization is prone to misalignment, especially in scenarios such as power professional term reporting, which lacks fidelity. The model training parameters rely on manual adjustment, and there is a lack of differentiated strategies for multi-terminal adaptation, resulting in large fluctuations in model performance and terminal display lag, which cannot meet the core needs of real-time and professional training in the power operation monitoring.
[0004] These problems make it difficult for existing digital figure generation technology to fully support the digital transformation of the power industry. There is an urgent need for a technical solution that can achieve high-fidelity generation of digital figures, standardized and efficient production, seamless adaptation to multiple terminals and multiple scenarios, and precise and intelligent interaction in the power field, to address the shortcomings of current technology in terms of precision, generalization and practicality, and to promote the large-scale and high-quality application of digital figures in the power industry. SUMMARY
[0005] Therefore, the application provides a multi-terminal and multi-scene adaptive digital image generation method and device, which can realize high-realistic generation of digital images, standardized and efficient production, seamless adaptation to multiple terminals and multiple scenes, and precise and intelligent interaction in the power field.
[0006] Specifically, the application is implemented through the following technical solutions.
[0007] The first aspect of the application provides a multi-terminal and multi-scene adaptive digital image generation method, which comprises the following steps:
[0008] collecting a video image of a real reference object;
[0009] extracting a first feature of a single dimension and a second feature of associated dimensions from the video image, wherein the second feature at least comprises joint features of two associated dimensions;
[0010] The associated dimensions at least include corresponding features of the mouth shape and the voice dimension, a voice signal and a mouth shape change image are extracted from the video image, the voice signal is converted into a phoneme sequence, the phoneme sequence is matched with the mouth shape feature to obtain a plurality of pairs of phoneme sequence and mouth shape, a supplementary mouth shape is generated according to the mouth shape corresponding to adjacent phoneme pairs, the supplementary mouth shape is used as a transition mouth shape corresponding to the mouth shape corresponding to adjacent phoneme pairs, and the second feature corresponding to the mouth shape and the voice dimension is adjusted.
[0011] The first feature and the shooting scene corresponding to the video image are taken as learning objects, and the second feature is taken as a constraint rule, and a digital image generation agent is trained, wherein the input of the digital image generation agent at least includes scene data, and the output at least includes a digital image.
[0012] Scene data of a digital image to be generated is input into the digital image generation agent to generate a customized digital image.
[0013] The second aspect of the application provides a multi-terminal and multi-scene adaptive digital image generation device, which comprises a collection module, an extraction module, a training module and a generation module.
[0014] The collection module is used to collect a video image of a real reference object.
[0015] The extraction module is used to extract a first feature of a single dimension and a second feature of associated dimensions from the video image, wherein the second feature at least comprises joint features of two associated dimensions.
[0016] The association dimension at least includes corresponding features of the mouth shape and the speech dimension, a speech signal and a mouth shape change image are extracted from the video image, the speech signal is converted into a phoneme sequence, the phoneme sequence is corresponded with the mouth shape feature, a plurality of pairs of phoneme sequence and mouth shape are obtained, a supplementary mouth shape is generated according to the mouth shape corresponding to adjacent phonemes, the supplementary mouth shape is taken as a transition mouth shape corresponding to the mouth shape corresponding to adjacent phonemes, and a second feature corresponding to the mouth shape and the speech dimension is adjusted;
[0017] The training module is configured to take the first feature and a shooting scene corresponding to the video image as a learning object, take the second feature as a constraint rule, and train a digital figure generation agent, wherein an input of the digital figure generation agent at least includes scene data, and an output of the digital figure generation agent at least includes a digital figure.
[0018] The generation module is configured to input scene data of a digital figure to be generated into the digital figure generation agent, and generate a customized digital figure.
[0019] The multi-terminal and multi-scene adaptive digital image generation method and device provided by the application can automatically generate a digital image matched with a scene and synchronized with a speech through a model, by inputting information matched with the scene into the model. Specifically, first, real appearance and action data are obtained by collecting video images of entity references, and then a first feature of a single dimension and a second feature of associated dimensions are extracted. The second feature includes the relationship between the first feature and its associated features, that is, the second feature is a feature pair, and the feature pair includes features having a correlation relationship. An intelligent agent is trained with the second feature as a constraint, ensuring that the digital image generation can accurately restore the detailed features of individual parts and avoid disconnection of actions of each part through the constraint of associated dimensions. The whole process from data collection to model training guarantees the fidelity of the digital image. Second, the shooting scene corresponding to the video image is taken as the learning object of the intelligent agent, and the input of the intelligent agent is the scene data of the digital image to be generated, and the output is a customized digital image. In this way, the digital image generation intelligent agent can learn the corresponding relationship between the scene and the image features, without the need for human intervention in professional aspects such as appearance design and action arrangement of the digital image. Even if the operator does not have professional digital image design skills such as 3D modeling and animation design, the intelligent agent can automatically match the appearance features and action features corresponding to the scene by only inputting the basic data of the scene to be generated, and finally output the matched customized image. This process completely reduces the technical threshold of digital image generation, avoids the low efficiency and high cost problem caused by the need for professional designers to customize images for each scene in the traditional scheme, and ensures that the digital image in different scenes meets the specifications of the power industry, effectively meeting the differentiated image generation needs in different scenes. In addition, for the mouth shape and speech dimension in the associated dimension, a transition supplementary mouth shape is generated based on the aligned phonemes and mouth shape, which optimizes the smoothness of the mouth shape change at different time points, effectively solves the problem of misalignment and hard conversion of the mouth shape and speech in traditional technologies, ensures the high synchronization of the mouth shape and speech in the speech interaction scene, and improves the naturalness and rationality of the interaction process. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 The flowchart of the multi-terminal and multi-scene adaptive digital image generation method provided by the application;
[0021] Figure 2 The structure diagram of the multi-terminal and multi-scene adaptive digital image generation device provided by the application. DETAILED DESCRIPTION
[0022] The exemplary embodiments will be described in detail herein with reference to several drawings. The following description is presented in connection with the described embodiments, it is not intended to limit the application to the described embodiments, but rather the application should be limited only by the claims.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0024] It is to be understood that the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0025] The following specific embodiments are given to illustrate the technical solutions of the application in detail.
[0026] Embodiment One
[0027] Figure 1 The flowchart of the multi-terminal and multi-scene adaptive digital image generation method provided by the present application is provided in Embodiment One. Please refer to Figure 1 The method provided by the present embodiment can include:
[0028] S101, collect video images of entity references.
[0029] It should be noted that collecting video images of entity references refers to obtaining a dynamic image sequence containing target entity references through a camera device. The entity references can be real people, animals or objects used to construct digital images. The collected video images not only include continuous frames, but also reflect the appearance characteristics, motion information and posture changes of the references at different time points.
[0030] Specifically, collecting video images of entity references includes:
[0031] (1) Determine a plurality of scene types and entity references used in the past in each type of scene. The shapes of the plurality of entity references in the same type of scene are different.
[0032] It should be noted that first, the application scenarios involved in digital image generation are classified, taking the power industry as an example, which can be specifically divided into power operation monitoring scenarios, power training scenarios, power news broadcast scenarios, power customer service scenarios, etc. Different scenario types correspond to different industry environments and action requirements, such as matching the operation monitoring large screen background and cold color lighting in the power operation monitoring scenario, and the main action is the dignified broadcast gesture; the power training scenario needs to match the device operation table background and natural light, and the main action is the device operation related gesture. At the same time, the historical entity reference objects in each type of power scenario are determined, such as the professional lecturer image in the operation monitoring scenario and the power operation and maintenance personnel image in the training scenario. In order to enhance the generalization ability of digital image in the power industry, multiple entity reference objects with different appearances need to be selected in the same type of power scenario, for example, selecting operation and maintenance personnel of different genders and service lengths in the training scenario. In this way, diversified appearance characteristics (such as different styles of operation and maintenance clothes) and action data (such as different operation habits) that fit the actual power industry can be obtained, thereby improving the adaptability and realism of the digital image generation agent in each power scenario.
[0033] (2) Marking points are arranged on the face and action of the entity reference object, and the changes of the marking points are tracked.
[0034] It should be noted that the marking point refers to a high-recognizability (such as a specific color, shape) identification point used to assist the computer vision system in recognizing and tracking the movement of an object. Specifically, it is a standardized identification object arranged at a specific part of the entity reference object, which has a significant difference in visual features (color, reflectivity, brightness) from the material / color of the entity clothing and the background environment of the scene.
[0035] By pasting or setting high-contrast marking points on the key areas of the face of the entity reference object (such as the corners of the eyes, the corners of the mouth, the eyebrows, etc.) and the key parts of the action (such as the hands, arms, legs, and body joints, etc.), and then using computer vision technology such as image feature-based tracking algorithm, the position changes of these marking points in the video sequence are continuously monitored. In this way, the facial expression dynamics of the entity reference object (such as the upward movement of the mouth corner marking point when smiling, and the gathering of the eyebrow marking point when frowning) and the body action trajectory (such as the movement path of the hand and arm marking points when waving) can be accurately and meticulously captured, providing accurate raw data support for subsequent extraction of facial and action features of the digital image, and ensuring the authenticity and fluency of the digital image in expression and action presentation.
[0036] (3) In the plurality of scenario types, the entity reference object with marking points arranged is videoed.
[0037] After completing the layout of the marker points, video shooting of the entity reference object with the marker points is performed using a camera device for the previously divided multiple scene types (for example, the power operation monitoring, training, news broadcast, customer service, and the like in the power industry). Specifically, the shooting parameters are adjusted according to the characteristics of different scenes. For example, in the power operation monitoring scene, the lens focal length may need to be adjusted to adapt to the large-screen display requirement, and cold-tone light is set to match the operation monitoring environment; in the power training scene, natural light simulation is used to set the light to ensure that the device operation table and the like are clearly presented. Through shooting in these diversified scenes, video images of the entity reference object with precise marker point motion information in different scenes can be obtained. Through the above operation, the video images shot contain rich information such as the appearance and action of the entity reference object in the actual scene of the power industry, providing a high-quality data source covering multiple scenes for subsequent digital image generation, so that the generated digital image can better adapt to various application scenarios in the power industry, and the practicality and fidelity of the digital image in the power field are improved.
[0038] In addition, in terms of collecting the related actions of the entity reference object, repeated collection of each action is considered, and the best data is selected. It should be noted that environmental noise (such as the humming of power equipment, the movement of background personnel), device noise (such as camera sensor noise, microphone circuit noise) can be easily introduced during the collection process, and the entity reference object may have non-standard deviation in a single action (such as the power operation and maintenance personnel's single closing action may deviate due to slight hand shaking, and the lecturer's single broadcast may be out of sync between the mouth shape and the voice due to the fluctuation of the speaking speed), and by repeatedly collecting the same action, the sample with the lowest noise and the most standard action can be selected from multiple data.
[0039] After the video images of the entity reference object are collected, in order to better serve the training of the subsequent digital image generation agent, a series of preprocessing operations need to be performed on the original video images, so as to obtain clean, orderly, and clearly labeled video frames, voice signals, and related label data.
[0040] Specifically, after the video images of the entity reference object are collected, the following operations are included:
[0041] (1) The collected video images are separated into video frames and synchronously collected voice signals.
[0042] It should be noted that the video image is essentially a dynamic sequence composed of continuous video frames, and at the same time, during the shooting process, the voice signal related to the video content (such as the sound of the explanation of the entity reference and the operation instruction) is often synchronously collected. Through a specific audio-video separation technology, the complete video image can be split into independent video frames (static image units) and corresponding voice signals (audio data), which can provide a basis for the respective processing of video and audio, so that the video frames and voice signals can enter different processing flows respectively.
[0043] (2) Filter and denoise the video frames and voice signals respectively.
[0044] When collecting video images and voice signals, some noise will inevitably be introduced, such as environmental noise (such as the slight hum of power equipment operation, external background noise), electronic noise of the shooting device itself, etc. Filter and denoise processing is to use signal processing algorithms, for example, for video frames, image filtering algorithms are used to remove grain noise, stripe interference, etc. in the picture; for voice signals, audio noise reduction algorithms are used to eliminate background noise. After such processing, the video frames will be clearer, and the voice signals will be purer, effectively improving the accuracy of subsequent feature extraction and other operations.
[0045] (3) Add a timestamp to each video frame, and perform frame-level time alignment of the video frames and corresponding voice signals based on the timestamp.
[0046] It should be noted that the timestamp is an identifier used to mark the time when the video frame and voice signal are generated. After adding a timestamp to each video frame, the video frames and corresponding voice signals can be accurately matched in the time dimension according to these timestamps. For example, when the entity reference makes a specific action (corresponding to a video frame) at a certain time point and simultaneously emits a sound (corresponding to a voice signal), the timestamp can ensure that the two are one-to-one corresponding in time. Such processing can enable the subsequent digital image to achieve precise synchronization of the mouth shape and voice, avoiding the misalignment of the mouth shape and voice, and greatly improving the fidelity of the digital image.
[0047] (4) Add descriptive labels to the time-aligned video frames and voice signals.
[0048] It should be noted that the descriptive label can be stored in XML format, which is a standardized textual description of the content contained in the video frames and voice signals, and at least includes scene information, data association information, and content semantic information. The scene information is generated according to the pre-set scene classification and on-site environment record data before video image acquisition, ensuring that the label is strongly bound to the power scene and avoiding confusion of data scene ownership; the data association information includes video frame timestamp, voice segment ID and entity reference ID, which is generated according to the timestamp when the video frame is separated, the voice segment number and the entity reference information, to realize the accurate binding of video and voice; the content semantic information includes action content, voice text transcription and key feature label, which is generated according to the action track tracked by the mark point in the video frame, the text transcription result of the voice signal and the feature judgment output by the semantic classification model (BERT model). The role of adding labels is to give these data clear semantic information, which is convenient for the subsequent training of digital image generation agents, so that the agent can better understand the meaning represented by different video frames and voice signals, and also facilitates the classification, retrieval and management of data. It can also filter specific data of different scenes through labels to avoid irrelevant data interference in model training.
[0049] (5) Store the processed video frames, voice signals and corresponding descriptive labels.
[0050] After completing the previous separation, noise reduction, time alignment and adding labels and other operations, the processed video frames, voice signals and corresponding descriptive labels are stored, usually in a special database or data storage system. These pre-processed high-quality data are retained in order to quickly and conveniently call these data in subsequent training of digital image generation agents, providing sufficient and standardized data sources for the agent to learn the generation rules of digital images.
[0051] Preferably, before storing the descriptive tags, the method further includes a step of cross-validating and correcting the tag content. This correction process aims to improve the accuracy of the descriptive tags by leveraging the inherent consistency between different modal data. Specifically, the correction process includes: precisely aligning the data from different modalities on the timeline based on the same timestamp system attached to video frames and audio signals. Subsequently, using data from one modality as a benchmark, tags generated based on data from another modality are validated and corrected. For example, if the initial descriptive tag "closing operation" is generated based on a separated audio signal (e.g., recognizing the command "now execute closing"), the correction process will simultaneously retrieve the original video frames from the same time period and the extracted lip-sync images. By analyzing the lip movements using computer vision algorithms, if a definite lip movement pattern matching the "closing" command is identified, the tag is confirmed to be accurate; if the lip-sync changes are significantly inconsistent with the audio content (e.g., lip closure corresponds to a phoneme that should have an open consonant), the tag is marked as "to be verified" or corrected based on the video content. Conversely, the incorrect labels that may be generated by speech recognition can be corrected based on the operation actions recognized in the video.
[0052] This cross-modal cross-validation effectively filters out labeling errors caused by a single source of information (such as audio noise or visual occlusion), ensuring that the final stored descriptive labels truly and accurately reflect the joint semantic content of video and speech. All the video data, speech signals, and high-precision labels that have undergone unified processing and correction together constitute a high-quality dataset that eliminates random scene noise and acquisition specificity.
[0053] S102. Extract a first feature of a single dimension and a second feature of an associated dimension from the video image, wherein the second feature includes at least a joint feature of the two associated dimensions.
[0054] It should be noted that the first feature refers to the feature extracted from a single body part dimension (such as face, eyes, mouth, hand, etc.) for describing the appearance and motion state of the part independently, and the core use is to accurately depict the details of each component of the digital image. Specifically, the first feature can include face contour feature vector (structured data describing face shape, relative position of facial features), mouth shape feature vector (numerical representation of the shape of the lips at a certain moment, such as opening and closing, protrusion, etc.), hand joint angle sequence (time sequence data describing finger bending, wrist rotation, etc.), and eye state feature (such as open eyes, closed eyes, gaze direction) coding, etc. The second feature refers to the joint feature extracted from multiple dimensions with cooperative relationship (i.e. associated dimensions) for describing the interaction and constraint relationship between these dimensions, and the core use is to ensure the overall coordination and naturalness of the digital image when different parts move, which is the key to improving the fidelity. Specifically, the form of the second feature not only includes simple feature vector splicing, but more importantly, it includes constraint parameters representing the association relationship, for example, the second feature can include mouth shape-speech association feature, hand-eye association feature, expression coordination feature, etc.
[0055] wherein the corresponding features of the mouth shape and the speech dimension are extracted from the video image, the speech signal and the mouth shape change image are extracted from the video image, the speech signal is converted into a phoneme sequence, the phoneme sequence is corresponded to the mouth shape feature to obtain a plurality of pairs of phoneme sequence and mouth shape pair, a supplementary mouth shape is generated according to the mouth shape corresponding to adjacent phoneme pairs, the supplementary mouth shape is used as a transition mouth shape corresponding to the mouth shape corresponding to the adjacent phoneme pairs, and the second feature corresponding to the mouth shape and the speech dimension is adjusted.
[0056] It should be noted that after the collection and preprocessing of the video image, in order to enable the digital image generation agent trained subsequently to accurately learn the features of the digital image, key features need to be extracted from the video image, including the first feature of a single dimension and the second feature of an associated dimension. The associated dimension feature is particularly important, such as the corresponding features of the mouth shape and the speech dimension, which can ensure the accurate synchronization of the mouth shape and the speech of the digital image. Specifically, the speech signal and the mouth shape change image are extracted from the video image, the speech signal is converted into a phoneme sequence, the phoneme sequence is corresponded to the mouth shape feature to obtain a plurality of pairs of phoneme sequence and mouth shape pair, a supplementary mouth shape for transition is generated according to the mouth shape corresponding to adjacent phoneme pairs, and the second feature of the mouth shape and the speech dimension is adjusted to provide accurate association constraints for the realistic generation of the digital image.
[0057] Specifically, the first feature of a single dimension and the second feature of an associated dimension are extracted from the video image, including:
[0058] (1) Position an entity reference in the video image, and dimensionally split the entity reference to obtain multiple dimensions.
[0059] It should be noted that the dimension in this step refers to a specific body part of the entity reference. The video image contains multiple part information of the entity reference, such as face, eyes, mouth, hands, legs, body, etc. Dimensional splitting is to divide the complete video image into multiple independent dimensions according to these different parts. For example, the video image containing the whole digital image is split into face dimension, eye dimension, mouth dimension, hand dimension, etc. By obtaining multiple dimensions, subsequent feature extraction can be targeted for each part, making feature extraction more detailed and accurate.
[0060] (2) Traverse each dimension, and for each dimension, position the sub-video frame of the entity reference part region corresponding to the dimension in the video image to form a sub-video image of each dimension.
[0061] Each frame of image mainly contains the target part itself and the associated part region directly connected in the anatomical structure and necessary for realizing the action of the part. After completing the dimensional splitting, each dimension is processed one by one. For each dimension, the target detection technology in computer vision is used to accurately position the entity reference part region corresponding to the dimension in the video image, and then a sub-video image containing only the part is extracted from the video image. For example, for the mouth dimension, the region where the mouth is located in the video image is positioned, and the sub-video image of the mouth is extracted.
[0062] It should be further explained that the positioning and extraction here are not simple rectangular clipping, but through fine image segmentation to ensure that each frame of sub-video extracted meets the specific content restriction. For example, for the hand dimension, the sub-video frame should clearly and completely contain the hand. In order to realize the continuity and naturalness of the action, the sub-video frame also needs to contain the associated part directly connected with the target part in the anatomical structure and cooperatively changed in the movement process of the target part. For example, when extracting the sub-video image of the hand dimension, in addition to the hand itself, a part of the wrist and even the forearm region should be included to completely capture the mechanical chain of actions such as grabbing and waving hands; when extracting the sub-video image of the mouth dimension, the chin and part of the cheek region should be included to accurately reflect the linkage of surrounding muscles when the mouth shape changes. The sub-video image extracted in this way focuses on the independent features of the target part and also retains the local cooperative information of the associated part.
[0063] (3) Based on the semantic expression of the sub-video image and the pixel difference between adjacent frames, determine the key frame, and extract the feature vector of the corresponding part region in the key frame as the first feature.
[0064] The semantic expression of the sub-video image refers to the content meaning represented by the sub-video image, such as whether the mouth sub-video image is making a smiling action or a speaking action, and the pixel difference between adjacent frames reflects the dynamic change of the part. By analyzing the semantic expression of the sub-video image (judging whether it is a frame with a key action or expression) and the pixel difference between adjacent frames (a large difference indicates a significant action change), key frames are selected, and then in these key frames, feature extraction algorithms (such as convolutional neural networks) are used to extract feature vectors of the corresponding part regions. These feature vectors constitute the first feature of a single dimension, and the first feature can accurately describe the characteristics of a single part of a digital image.
[0065] When determining the key frames based on the semantic expression of the sub-video image and the pixel difference between adjacent frames, for long video data (such as power scene videos with a single scene shooting time of ≥30 min), it is particularly necessary to use the mechanism of difference calculation and semantic content detection to extract key frames. Specifically, for the sub-video images corresponding to the long video, the pixel difference between adjacent two frames is calculated frame by frame in time sequence. Specifically, the gray value difference algorithm can be used to calculate the average value of the absolute difference of the gray values of the corresponding pixel points of the two frames, and a difference threshold is set. When the average value is greater than the difference threshold, it is determined that the frame is a frame with significant action change, and is preliminarily selected as a key frame candidate set. Further, the key frame candidate set is further screened by a semantic content detection algorithm to accurately capture key frames containing important actions and expressions in the long video.
[0066] (4) Determine a plurality of associated dimension combinations. For each associated dimension combination, based on the intersection of the key frames corresponding to the associated dimensions, screen the key frames containing both dimension parts.
[0067] It should be noted that the associated dimension combination refers to the combination of two dimensions that have an association relationship, such as the mouth shape dimension and the speech dimension, the eye dimension and the mouth dimension, etc. For each such associated dimension combination, find the intersection of the key frames corresponding to the two dimensions, that is, the frames that are key frames in both dimensions at the same time. In this way, key frames containing both dimension parts are screened, and this is done to find key frames that can reflect the common characteristics of the two associated parts, in preparation for extracting the second feature of the associated dimension.
[0068] (5) Extract feature vectors of the two dimension part regions in the key frames to form the constraint feature of the associated dimension, and obtain the second feature.
[0069] In the key frame containing two dimension parts, the feature vectors of the two dimension part regions are extracted respectively, and then the feature vectors are combined to form the constraint feature of the associated dimension, which is the second feature of the associated dimension. The second feature can reflect the relationship between the two associated dimensions, for example, the second feature of the mouth shape dimension and the speech dimension can constrain the synchronization of the digital image's mouth shape and speech, so that the generated digital image is more coordinated and realistic in the performance of the associated parts.
[0070] Further, in the process of generating the supplementary mouth shape according to the mouth shape corresponding to the adjacent phonemes, a dynamic transition algorithm based on key point positioning and position interpolation is adopted to ensure the smoothness and naturalness of the mouth shape change. Specifically, first, for each phoneme corresponding to the standard mouth shape (i.e. the basic mouth shape), a facial feature point detection model is used to locate a plurality of predefined key points (such as the center points of the upper and lower lips, the left and right corner points of the mouth, the lip peak points, etc.) on the lip contour, and the two-dimensional or three-dimensional coordinates of these key points form the digital representation of the basic mouth shape. Then, for any two adjacent phonemes (denoted as phoneme A and phoneme B), the corresponding basic mouth shape key point sets are obtained. The corresponding key points in the two sets (such as the upper lip center point pair A and the upper lip center point pair B) are paired to form a series of dynamic point pairs. Further, within the time interval from phoneme A to phoneme B, for each dynamic point pair, according to the preset number of transition frames, linear interpolation or smoother cubic spline interpolation algorithm is used to calculate the transition position of the key point in each intermediate frame, and the motion trajectories of all key points are calculated synchronously and independently by the interpolation algorithm, and the position of each key point in any transition frame is calculated accurately from its starting and ending positions, i.e. many points are fixed calculation paths. Finally, the new positions of all key points in all transition frames are reconnected to form a continuous lip contour sequence, which is the generated supplementary mouth shape between phoneme A and phoneme B, which ensures that the mouth shape change from A to B is continuous, smooth and consistent with human oral kinematics.
[0071] S103, taking the first feature and the shooting scene corresponding to the video image as learning objects, and the second feature as a constraint rule, training a digital image generation agent, wherein the input of the digital image generation agent at least includes scene data, and the output at least includes a digital image.
[0072] It should be noted that after obtaining the first feature of a single dimension, the second feature of the associated dimension, and determining the shooting scene corresponding to the video image, in order to enable the digital image to generate appropriate content according to different scenes, it is necessary to train the digital image generation agent. During training, the first feature and the shooting scene are taken as the learning objects of the agent, and the second feature is taken as the constraint rule, so that the generated digital image not only meets the scene requirements, but also maintains coordination in the features of each part and the associated features.
[0073] Specifically, the training digital image generation agent comprises:
[0074] (1) constructing a training sample, the training sample comprising shooting scene data, first features of each component and corresponding second features, and corresponding historical digital images.
[0075] It should be noted that the shooting scene data covers the type of scene (such as power operation monitoring, training, etc.), environmental characteristics of the scene (such as background, lighting, etc.), and other information. When constructing the training sample, the shooting scene data is taken as the conditional input of the model. The learning goal of the model (i.e. the expected output) is a digital image instance that is highly matched with the scene data and has been verified in actual application in history as having excellent effect (for example, the professional announcer image with the highest user evaluation and the top N usage frequency in the power operation monitoring scene).
[0076] The first features of each component and the second features of the associated dimension are not used as direct inputs of the model here, but as fine-grained intermediate supervision signals and target outputs required for fitting by the internal feature generation layer of the model, which means that the training goal of the agent is to learn a generative mapping from the scene data to a set of coordinated and realistic component features (first features) and their constraint relationships (second features). Finally, a complete and movable digital image is rendered by decoding and combining the generated features.
[0077] The conditional inputs (scene data), intermediate supervision and output targets (first and second features), and final output targets (preferred digital image instances) are integrated together to construct the training sample, providing a complete data basis with multiple levels and strong guidance for subsequent training of the agent. This sample construction method provides precise guidance for the feature generation process within the model. The model not only needs to output the final image, but also requires the internal feature generation layer to output consistent and high-quality feature combinations with the preferred sample. This mechanism of guiding feature generation from the preferred results enables the model to more deeply understand the internal relationship between the scene and the component features of the image, rather than just learning the surface mapping, thereby significantly improving the component accuracy, overall coordination, and scene fit of the generated digital image.
[0078] (2) Perform expression enhancement and action enhancement on the training samples.
[0079] Expression enhancement and action enhancement are to enrich and expand the expressions and actions of digital avatars in the training samples through data enhancement techniques. For example, for a sample originally showing a smiling expression, enhanced samples with different amplitude smiles, smiling accompanied by slight head movements, etc. are generated; for action samples, versions with slightly changed action speed and amplitude are generated. This can increase the diversity of training samples, allowing digital avatar generation agents to generate more diverse and natural expressions and actions after training.
[0080] (3) Use the scene data as input and combine the second feature to construct action-related constraint relationships.
[0081] It should be noted that the digital avatar generation agent is essentially a deep learning-based generation model, and its core function is to act as a generator to output a complete and coordinated set of digital avatar driving parameters based on the input scene data. The core logic of the generation process is that the final presented digital avatar is composed of the driving parameters of its components (i.e., the first feature), and the coordination and naturalness between these component driving parameters are regulated by the constraint relationships defined by the second feature.
[0082] Specifically, during model training and inference, scene data is used as a conditional input to determine the overall style and context of digital avatar generation (e.g., professional and rigorous operator images need to be generated in the power training scenario). The first feature, as the generation target, is the specific parameter that the model needs to predict to drive the components of the digital avatar (such as facial muscles and hand joints). The second feature, as a constraint rule, is encoded into the model's loss function or network structure to ensure that the first features of the generated components meet the physical and semantic coordination relationship. For example, using the lip shape-speech second feature, a time synchronization constraint is constructed in the model to ensure that the generated lip shape driving parameters are accurately aligned in time with the input speech signal; using the hand-view second feature, a kinematic correlation constraint is constructed in the model to ensure that when the hand driving parameters exhibit pointing actions, the eye driving gaze direction parameters can simultaneously point to the same target area; using the expression coordination second feature, a muscle linkage constraint is constructed in the model to ensure that the driving parameters of the mouth, eye, and eyebrow regions of the generated face conform to the real expression muscle movement rules.
[0083] By fusing the scene data with the second features, the model not only considers the macro requirements of the scene when generating the first features of each component, but is also strictly constrained by the micro coordination relationship between components. This mechanism ensures that the final generated digital figure has coherent, unified, and highly consistent actions of each component with the scene requirements, thereby fundamentally solving the technical problem of disconnection between component actions and scene context.
[0084] (4) Learning a generative mapping relationship between the scene features and the first features based on the scene data, the first features, and the action-related constraint relationship, the relationship between each first feature being constrained by the related second features, and generating the digital figure generation agent.
[0085] It should be noted that the scene data contains feature information of the scene, the first features are features of each component, and the action-related constraint relationship is a constraint rule between the two. Through a specific machine learning algorithm (such as a generative adversarial network, a variational autoencoder, etc.), the agent learns the generative mapping relationship between the scene features and the component features, that is, learns how to generate a component feature combination that meets the constraints according to different scene features and component features, thereby laying the foundation for subsequent generation of complete digital figures.
[0086] When learning the generative mapping relationship between the scene features and the component features based on the scene data, the first features, and the action-related constraint relationship, to achieve effective fusion of different modal features and improve the overall performance of the digital figure, joint feature learning is needed. Through feature selection and weighting methods, the importance of different modal features is determined to improve the effect of feature fusion. Specifically, a multi-modal joint learning network can be constructed, with the first features of the video modality, the speech features of the audio modality, and the scene data features of the scene modality as network inputs. Through a cross-modal attention mechanism, the association between different modal features (such as the temporal association between lip feature and speech feature, and the semantic association between action feature and scene feature) is learned, and preliminary fusion features are output. Further, L1 regularization algorithm is used to sparsify the preliminary fusion features to eliminate redundant features. Then, the importance weights of the core features of each modality are calculated based on the random forest algorithm, and the final fusion features are obtained by weighted summation. Finally, the final fusion features are input into a generative adversarial network (GANs) or a variational autoencoder (VAE), combined with the action-related constraint relationship, to learn the generative mapping relationship between the scene features and the component features, ensuring that the generated digital figure has better overall performance under the synergistic effect of multi-modal features.
[0087] Among them, the generative adversarial network and the variational autoencoder are two model frameworks that are more suitable for digital image training generation. The generative adversarial network has the advantage of being able to generate high-quality realistic images or video sequences, and is suitable for digital image generation tasks that require high complexity and delicate performance, such as face generation, animation character generation, etc. The variational autoencoder has the advantage of being able to learn the latent representation of the data, generate samples with diversity and continuity, and is suitable for digital image generation tasks that require control of the features of the generated samples. If more complex scenarios are considered, deep generative models (such as deep autoencoders) combine the powerful representation learning capabilities of deep learning and the generation capabilities of generative models, can learn and generate complex data distributions, have the ability to learn complex data distributions, and generate high-quality samples, and can handle multi-modal data input.
[0088] In addition, during the training process, a special loss function needs to be designed to ensure the generation quality, which consists of three parts: reconstruction loss (using mean square error to calculate the difference between the generated features and the real features), constraint loss (based on the constraint relationship defined by the second feature, calculating the collaborative consistency loss between the component features), and adversarial loss (when using the generative adversarial network, the adversarial signal provided by the discriminator is used to improve the generation fidelity).
[0089] Optionally, an unsupervised training method based on reinforcement learning can also be used, which models the generation process as a sequential decision problem and guides the model learning through the design of reward and punishment functions. Specifically, it includes reward signal design (giving positive reward when the generated feature combination is close to the real data distribution), constraint satisfaction reward (giving additional reward when the generated component features meet the constraint relationship defined by the second feature), and diversity reward (to prevent mode collapse, appropriate reward is given to the diversity of generated features).
[0090] It should be noted that the agent is continuously trained using training samples, so that the agent learns how to combine the component features in a reasonable way to generate complete digital images that meet the scene and constraint requirements. After sufficient training, a digital image generation agent is finally generated that can output corresponding digital images according to the input scene data.
[0091] It should be noted that the number of training samples and the quality requirements thereof, the number of training iterations and the form of loss function (determined according to actual conditions) also need to be considered during the training process. Specifically, an automatic parameter tuning tool or an experimental design method can be used to determine the best combination. The automatic parameter tuning tool can use Hyperopt, which uses a sequential model optimization algorithm to optimize hyperparameters, combining the ideas of Bayesian optimization and random search, which can effectively optimize complex parameter spaces; automatic machine learning tools such as AutoML, H2O.ai and TPOT not only limit the adjustment of hyperparameters, but also automatically perform feature engineering, model selection and optimize the entire machine learning process. It can significantly reduce the error of manual parameter adjustment and improve the performance and efficiency of the model.
[0092] It should be noted that after the model training is completed, there is also a process of using. Specifically, in the model using stage, first, the scene data of the digital image to be generated is converted into a scene feature vector according to a preset format as an input of the agent; the Transformer decoder inside the agent first analyzes the scene feature vector, combines the scene-feature mapping relationship learned in the training stage, and outputs the first feature and the second feature constraint of each component of the target digital image; then, the agent calls the pre-constructed digital image component library (homologous to the component library in the training sample, ensuring that the features match the components), filters the corresponding component according to the first feature, and adjusts the component parameters according to the second feature constraint; finally, the component and the parameter are transmitted to the terminal rendering engine through the rendering interface built-in the agent, and the rendering engine generates the final digital image that can be displayed according to the terminal hardware configuration.
[0093] S104, inputting scene data of a digital image to be generated into the digital image generation agent to generate a customized digital image.
[0094] After the training of the digital image generation agent is completed, the scene data of the digital image to be generated can be input into the agent to generate a customized digital image. Moreover, in order to make the generated digital image better adapt to different terminal platforms and present a more complete scene effect, subsequent operations such as terminal adaptation and scene fusion are also needed.
[0095] Specifically, the inputting scene data of a digital image to be generated into the digital image generation agent to generate a customized digital image comprises:
[0096] (1) predicting, by the digital image generation agent, a component set required by a target digital image and an action time sequence of each component based on the scene data.
[0097] It should be noted that after receiving the scene data of the digital image to be generated (such as the electric power operation monitoring scene data), the digital image generation agent will predict the component set (such as the face component, the hand component, etc.) required for constructing the target digital image according to the corresponding relationship between the scene and the digital image component and action learned by itself, and the action timing of each component (that is, each component performs what action at what time), which can provide a clear basis for subsequent component retrieval and combination.
[0098] (2) Retrieve the component set from the pre-constructed digital image component library, and determine the component combination relationship at each time point according to the predicted timing.
[0099] The pre-constructed digital image component library stores various components of digital images (such as different facial expression components, different gesture action components, etc.). According to the component set predicted by the agent, the corresponding components are retrieved from the component library, and then according to the predicted action timing, it is determined which components need to be combined together at each time point, so as to determine the combination relationship of the components.
[0100] (3) According to the action prediction result output by the digital image generation agent, generate the transmission parameters of the component interface at each time point, and drive the component to perform the corresponding action.
[0101] The action prediction result output by the digital image generation agent contains specific information of the component action. Based on these results, the transmission parameters of the component interface at each time point are generated (these parameters are used to control the action execution of the component, such as the amplitude and speed of the action), and then the corresponding components are driven to perform the corresponding action using these transmission parameters, so that the components move in the expected way.
[0102] (4) The component combination action is spliced in time sequence to form a customized digital image.
[0103] After each component performs the corresponding action according to the transmission parameters, the combination actions of these components are spliced in sequence according to the action timing, thereby forming a complete and customized digital image that meets the scene requirements
[0104] After generating the customized digital image, in order for the digital image to be well displayed on different terminal platforms (such as large screens, mobile phones, etc.), terminal adaptation operation is required.
[0105] It should be noted that after inputting the scene data of the digital image to be generated into the digital image generation agent, the following steps are included:
[0106] (1) Obtain the hardware configuration parameters and display specification information of the target terminal platform.
[0107] The hardware configuration parameters (such as processor performance, memory size, etc.) and display specification information (such as screen resolution, size ratio, etc.) of the target terminal platform will affect the display effect of the digital image. Through relevant technical means (such as communication with the terminal platform), these information are obtained in order to subsequently adjust the digital image.
[0108] (2) The hardware configuration parameters and display specification information are transmitted to the rendering engine.
[0109] The rendering engine is the core module responsible for rendering and displaying the digital image. By transmitting the obtained hardware configuration parameters and display specification information to the rendering engine, the rendering engine can know the capabilities and requirements of the terminal platform.
[0110] (3) According to the hardware configuration parameters and display specification information, the size, resolution and format of the digital image generated by the agent output are adjusted.
[0111] According to the hardware configuration parameters and display specification information of the terminal platform, the generated customized digital image is adjusted. For example, if the terminal is a small-screen mobile phone, the size of the digital image is adjusted to be smaller, the resolution is adjusted to be suitable for the mobile phone screen, and the format is adjusted to be supported by the mobile phone, so as to ensure that the digital image can be displayed clearly and normally on the terminal.
[0112] It should be noted that the customized digital image output by the digital image generation agent is essentially a prediction result of component combination logic and action timing parameters (not directly displayable images / animation). When the image is actually used, it needs to be based on the generalized component library and unified template to complete the instantiation construction. Specifically, a pre-constructed digital image general component library covering the entire scene of the power industry is provided. The library includes standardized basic components (such as face components: face templates of different face shapes / skin colors; body components: torso templates of operation and maintenance clothes / garment styles; action components: power-specific action templates such as closing and inspection), all of which are stored in a unified format (such as FBX) and have standardized interfaces (supporting parameterized adjustment of size, color, and action amplitude). Based on the prediction result output by the agent, the corresponding general components are called from the component library to ensure consistency and compatibility of the components called in different scenarios and different terminals. Further, a digital image general construction template is configured, which defines component combination rules (such as the connection coordinates of face components and torso components, and the kinematics correlation of action components and body components) and terminal adaptation benchmarks (such as scaling ratios of components under different resolutions, and encoding parameters under different formats). The called general components are imported into the template, and the component combination is automatically completed according to the template rules (such as aligning the face components and torso components according to the neck coordinates, and binding the action components and hand components according to the joint parameters). At the same time, combined with the hardware configuration and display specifications of the target terminal, the component size is fine-tuned and the resolution is adapted through the parameter adjustment module built-in the template. After completing the component combination and template adaptation, the combined components are generated into instantiated content according to the action timing parameters through a rendering engine. For high-performance terminals, high-resolution (such as 4K) real-time animation streams are directly generated; for mobile terminals (such as operation and maintenance mobile phones), lightweight sequence frames (such as sprites, single frame size ≤ 50KB) are generated. Finally, the displayable digital image that completely matches the terminal hardware and display specifications is output. The entire process relies on the general component library and unified template, avoiding repeated development of components due to scene or terminal differences, realizing one-time prediction result and multi-terminal generalization, while ensuring consistency of the digital image displayed on different terminals in appearance and action, meeting the standardized interaction needs of the power industry.
[0113] In addition, in order to improve the efficiency and effect of the digital image displayed on the terminal, a cloud-edge collaborative optimization method is adopted. It should be noted that after the scene data of the digital image to be generated is input into the digital image generation agent to generate a customized digital image, the following steps are further included:
[0114] (1) Sending the component action timing data and rendering request corresponding to the customized digital image to a cloud processing node.
[0115] The component action time sequence data corresponding to the customized digital image (i.e., the time sequence of each component action and the like) and the rendering request are sent to a cloud processing node, and the cloud processing node has strong computing capability and can perform more complex processing on the data.
[0116] The cloud processing node calls a part action component corresponding to the time sequence data from a pre-built component library, performs combination operation, generates a lightweight digital image performance model, and precomputes multi-channel rendering data of the digital image performance model.
[0117] The cloud processing node calls a part action component corresponding to the received time sequence data from a pre-built component library of itself, performs combination operation on the components, generates a lightweight digital image performance model (which has a smaller data size under the premise of ensuring a certain effect, and is convenient for transmission and processing), and precomputes multi-channel rendering data of the lightweight model, thereby completing part of the rendering work in advance and reducing the burden of the edge processing node.
[0118] In this step, the lightweight digital image performance model is used, such as pruning, quantization, and distillation, to improve the calculation efficiency and speed, parallel computing and GPU acceleration are used to speed up the speed of mouth shape calculation and animation generation, and the performance requirements in real-time application are ensured.
[0119] (3) The digital image performance model and the precomputed rendering data are packaged and delivered to the edge processing node.
[0120] The cloud processing node packages the generated lightweight digital image performance model and the precomputed rendering data, and then delivers them to the edge processing node. The edge processing node is close to the terminal and can respond to the rendering requirements of the terminal more quickly.
[0121] (4) The edge processing node performs calculation and real-time rendering on the received data packet according to the display specification information of the target terminal, and generates a final digital image matched with the target terminal.
[0122] After receiving the data packet, the edge processing node performs calculation on the data packet according to the display specification information of the target terminal, and then performs real-time rendering to finally generate a digital image matched with the target terminal, so that the digital image can be displayed on the terminal in real time and smoothly.
[0123] In addition, it should be noted that after generating the customized digital image, scene fusion operation is needed to make the digital image more suitable for the scene. Specifically, after generating the customized digital image, the following operations are performed:
[0124] (1) Fill the preset background constraints and additional component constraints for the customized digital figure based on the standardized template.
[0125] It should be noted that the standardized template contains background constraints (such as the type and style of the background) and additional component constraints (such as the type and position of the auxiliary components to be added) in different scenarios. According to the standardized template, these preset constraints are filled for the customized digital figure to determine the general requirements of the background and additional components of the digital figure.
[0126] (2) According to the constraints of the template, match and generate an adapted background scene from the pre-built background resource library.
[0127] The pre-built background resource library stores various background scene resources (such as large screen backgrounds for power operation monitoring scenes and practical operation table backgrounds for power training scenes). According to the background constraints in the template, the appropriate background scene is matched from the background resource library, and the background scene is generated to provide a suitable background environment for the digital figure.
[0128] (3) Call additional components that match the customized digital figure and background scene from the pre-built digital figure component library.
[0129] Similarly, the pre-built digital figure component library also has various additional components (such as power equipment model components and text explanation components). According to the characteristics of the customized digital figure and the generated background scene, the additional components that match them are called from the component library. These additional components can enrich the display scene of the digital figure.
[0130] (4) Adjust the display parameters of the background scene and additional components, and fuse the customized digital figure, background scene, and additional components to generate the final digital figure display content containing complete scenes.
[0131] Adjust the display parameters (such as size, position, transparency, etc.) of the background scene and additional components so that they can be well integrated with the customized digital figure. Then, fuse the customized digital figure, adjusted background scene, and additional components to finally generate digital figure display content containing complete scenes, making the display of the digital figure more rich and realistic.
[0132] The method provided by the embodiment can accurately capture the appearance features and action data of entities in the power industry by laying high-contrast marker points and tracking changes when collecting entity reference video images, combining with multi-scene differentiated shooting parameter adjustment; the audio and video separation, filtering and noise reduction, and frame-level time alignment in the preprocessing stage further guarantee the data quality; the single dimension and the associated dimension are split during feature extraction, and especially for the mouth shape-speech dimension, the phoneme sequence mapping and transition mouth shape generation logic are designed, effectively solving the problems of mouth shape and speech dislocation and action disconnection in traditional technologies, making the facial expression, body movement and speech interaction of the digital image close to the real scene. Taking the first feature and the shooting scene as the learning object and the second feature as the constraint to train the intelligent agent, and combining the expression and action to enhance the diversity of the rich samples, the intelligent agent can learn the corresponding relationship between the scene and the image features, and can output the customized digital image when the scene data is input, breaking through the limitations of traditional digital images that are single and fixed and disconnected from the scene, and accurately matching the action and appearance needs of different scenes such as power operation monitoring, training, and broadcasting. At the same time, the method realizes seamless display and efficient interaction of multiple terminals through terminal adaptation and cloud-edge collaborative optimization: for different terminals such as large screens, mobile phones, and tablets, the hardware and display parameters are obtained, and the size, resolution, and format of the digital image are adjusted; by means of the cloud to generate a lightweight model and pre-compute rendering data, and the edge node to solve the rendering in real time, the digital image display is smooth and the delay is controllable while reducing the computing burden of the terminal; in the scene fusion link, the background and additional components are matched based on the standardized template, further deepening the fit between the digital image and the power scene environment, ultimately meeting the core needs of real-time broadcasting, training practical demonstration, and customer precise interaction in the power industry, and providing high-quality digital image support for the digital transformation of the power industry.
[0133] Embodiment two
[0134] Corresponding to the foregoing embodiment of the method of generating a digital image self-adapted to multiple terminals and multiple scenes, the present application also provides an embodiment of a device for generating a digital image self-adapted to multiple terminals and multiple scenes.
[0135] Figure 2 A structural schematic diagram of the second embodiment of the device for generating a digital image self-adapted to multiple terminals and multiple scenes provided by the present application is shown in FIG. 2. Figure 2 The device provided by the embodiment includes a collection module 210, an extraction module 220, a training module 230, and a generation module 240.
[0136] The collection module 210 is configured to collect video images of entity references.
[0137] The extraction module 220 is configured to extract first features of a single dimension and second features of an associated dimension from the video images, and the second features at least include joint features of two associated dimensions.
[0138] The association dimension at least includes corresponding features of the mouth shape and the speech dimension, a speech signal and a mouth shape change image are extracted from the video image, the speech signal is converted into a phoneme sequence, the phoneme sequence is corresponded to the mouth shape feature to obtain a plurality of pairs of phoneme sequence and mouth shape pair, a supplementary mouth shape is generated according to the mouth shape corresponding to adjacent phoneme pairs, the supplementary mouth shape is taken as a transition mouth shape corresponding to the mouth shape corresponding to adjacent phoneme pairs, and a second feature corresponding to the mouth shape and the speech dimension is adjusted.
[0139] The training module 230 is configured to take the first feature and a shooting scene corresponding to the video image as learning objects, and take the second feature as a constraint rule, train a digital image generation agent, and take scene data as input of the digital image generation agent and take a digital image as output of the digital image generation agent.
[0140] The generation module 240 is configured to input scene data of a digital image to be generated into the digital image generation agent, and generate a customized digital image.
[0141] The apparatus of the embodiment can be used to execute the steps of the method embodiment, and the specific implementation principle and implementation process are similar, and will not be repeated here. Figure 1 The steps of the method embodiment are similar in terms of specific implementation principles and implementation processes, and will not be repeated here.
[0142] The implementation processes of the functions and roles of the units in the apparatus are specifically described in the implementation processes of the corresponding steps in the above method, and will not be repeated here.
[0143] For the device embodiment, since it basically corresponds to the method embodiment, the related parts are described in the method embodiment. The device embodiment described above is only schematic, and the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present application. Those skilled in the art can understand and implement without creative labor.
[0144] The above is only the preferred embodiment of the present application, and is not used to limit the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A method for generating digital images that is adaptive across multiple terminals and scenarios, characterized in that, The method includes: Acquire video images of physical reference objects; Extract a first feature of a single dimension and a second feature of an associated dimension from the video image, wherein the second feature includes at least a joint feature of the two associated dimensions; The associated dimension includes at least the corresponding features of the lip shape and speech dimension. The speech signal and lip shape change image are extracted from the video image. The speech signal is converted into a phoneme sequence. The phoneme sequence is matched with the lip shape feature to obtain multiple pairs of phoneme sequences and lip shape pairs. Supplementary lip shapes are generated according to the lip shape corresponding to adjacent phonemes. The supplementary lip shapes are used as the transitional lip shapes corresponding to the lip shapes corresponding to adjacent phonemes. The second feature corresponding to the lip shape and speech dimension is adjusted. Using the first feature and the shooting scene corresponding to the video image as learning objects, and the second feature as constraint rules, a digital image generation intelligent agent is trained. The input of the digital image generation intelligent agent includes at least scene data, and the output includes at least digital image. The scene data to be generated for the digital avatar is input into the digital avatar generation intelligent agent to generate a customized digital avatar.
2. The method according to claim 1, characterized in that, The video images of the acquired entity reference object include: Identify multiple scene types and the entity references historically used in each scene type; multiple entity references in the same scene type have different shapes. Marking points are placed on the face and movements of the physical reference object, and changes in the marking points are tracked; In the various scene types, video is captured of the physical reference objects after the marker points are set up.
3. The method according to claim 1, characterized in that, The step of extracting a first feature of a single dimension and a second feature of a related dimension from the video image includes: Locate the entity reference object in the video image, and then perform dimensional splitting on the entity reference object to obtain multiple dimensions; Traverse each dimension, and for each dimension, locate the sub-video frames of the entity reference area corresponding to that dimension in the video image to form the sub-video image for each dimension; Based on the semantic representation of the sub-video image and the pixel differences between adjacent frames, key frames are determined, and feature vectors of corresponding regions are extracted from the key frames as the first feature. Multiple association dimension combinations are determined. For each association dimension combination, based on the intersection of the key frames corresponding to the association dimensions, key frames that simultaneously contain two dimension parts are selected. Feature vectors of the two dimensional regions are extracted simultaneously from the keyframe to form the constraint features of the associated dimension, thus obtaining the second feature.
4. The method according to claim 1, characterized in that, The trained digital image generation intelligent agent includes: Construct training samples, which include shooting scene data, first features and corresponding second features of each component, and corresponding historical digital images; The training samples were enhanced with facial expressions and motion. The scene data is used as input, and action-related constraint relationships are constructed by combining the second feature. Based on the scene data, the first feature, and the action-related constraints, the generative mapping relationship between the scene features and the first feature is learned. The relationship between each first feature is constrained by the relevant second feature, and the digital image generation agent is generated.
5. The method according to claim 1, characterized in that, After acquiring the video image of the entity reference object, the following is included: The acquired video images are separated into video frames and synchronously acquired audio signals; The video frames and audio signals are respectively filtered and noise-reduced; Add a timestamp to each video frame and perform frame-level time alignment between the video frame and the corresponding audio signal based on the timestamp; Add descriptive tags to time-aligned video frames and audio signals; The processed video frames, audio signals, and corresponding descriptive tags are stored.
6. The method according to claim 1, characterized in that, The step of inputting the scene data to be generated into the digital image generation agent includes: Obtain the hardware configuration parameters and display specifications of the target terminal platform; The hardware configuration parameters and display specifications are transmitted to the rendering engine. Based on the hardware configuration parameters and display specifications, the size, resolution, and format of the digital image output by the digital image generation agent are adjusted.
7. The method according to claim 1, characterized in that, After inputting the scene data of the digital avatar to be generated into the digital avatar generation agent to generate a customized digital avatar, the process further includes: Send the component action timing data and rendering requests corresponding to the customized digital image to the cloud processing node; The cloud processing node calls the part action component corresponding to the time series data from the pre-built component library, performs combination calculations, generates a lightweight digital image representation model, and pre-calculates multi-channel rendering data for the digital image representation model. The digital image representation model and the pre-calculated rendering data are packaged and sent to the edge processing node; The edge processing node performs calculations and real-time rendering on the received data packets based on the display specifications of the target terminal, generating a final digital image that matches the target terminal.
8. The method according to claim 1, characterized in that, The step of inputting the scene data of the digital avatar to be generated into the digital avatar generation intelligent agent to generate a customized digital avatar includes: The digital image generation agent predicts the set of components required for the target digital image and the action sequence of each component based on the scene data. The component set is retrieved from the pre-built digital image component library, and the component combination relationship at each time point is determined according to the predicted time sequence; Based on the action prediction results output by the digital image-generating intelligent agent, the transmission parameters of the component interface at each time point are generated, and the component is driven to perform the corresponding action. The combined actions of the components are spliced together in sequence to form a customized digital image.
9. The method according to claim 1, characterized in that, After generating the customized digital avatar, the process includes: The customized digital avatar is filled with preset background constraints and additional component constraints based on a standardized template; Based on the constraints of the template, match and generate suitable background scenes from the pre-built background resource library; Call additional components from a pre-built digital avatar component library that match the customized digital avatar and background scene; The display parameters of the background scene and additional components are adjusted, and the customized digital image, background scene and additional components are merged to generate the final digital image display content containing the complete scene.
10. A multi-terminal, multi-scenario adaptive digital image generation device, characterized in that, The device includes a data acquisition module, an extraction module, a training module, and a generation module; The acquisition module is used to acquire video images of physical reference objects; The extraction module is used to extract a first feature of a single dimension and a second feature of an associated dimension from the video image, wherein the second feature includes at least a joint feature of the two associated dimensions; The associated dimension includes at least the corresponding features of the lip shape and speech dimension. The speech signal and lip shape change image are extracted from the video image. The speech signal is converted into a phoneme sequence. The phoneme sequence is matched with the lip shape feature to obtain multiple pairs of phoneme sequences and lip shape pairs. Supplementary lip shapes are generated according to the lip shape corresponding to adjacent phonemes. The supplementary lip shapes are used as the transitional lip shapes corresponding to the lip shapes corresponding to adjacent phonemes. The second feature corresponding to the lip shape and speech dimension is adjusted. The training module is used to train a digital image generation agent with the first feature and the shooting scene corresponding to the video image as learning objects, and the second feature as constraint rules. The input of the digital image generation agent includes at least scene data, and the output includes at least digital image. The generation module is used to input the scene data of the digital image to be generated into the digital image generation intelligent agent to generate a customized digital image.
Citation Information
Patent Citations
Virtual face generation method
CN113781610A
Generation method and generation equipment of digital human interaction video
CN119402720A
Real-time digital human generation system and method based on low-calculation-amount voice driving
CN119991893A
Method for Audio-Driven Character Lip Sync, Model for Audio-Driven Character Lip Sync and Training Method Therefor
US20240054711A1