Multi-scene sign language synthesis data set generation method, system and device based on 3D digital human and medium
By collecting real sign language eye-tracking data to generate attention weight masks, and combining them with semantic encoders and multi-view consistency constraints, a 3D digital human sign language synthesis dataset with high perceptual realism and multi-view consistency is generated. This solves the problems of insufficient perceptual realism and high computational cost in existing sign language synthesis data, and realizes efficient and large-scale multi-scene sign language data generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-03-13
AI Technical Summary
Existing 3D sign language synthesis systems lack effective perception and dynamic adjustment of the eye-tracking attention distribution of real humans when generating the movement and details of key areas such as hands, face, and mouth shapes. This results in insufficient realism of the synthesized data, inconsistencies between sign language actions and expressions from multiple perspectives, and huge computational overhead.
By collecting real sign language eye-tracking data to generate attention weight masks, using semantic encoders and attention region generation subnetworks to generate 3D digital human foregrounds, and merging them with preset backgrounds, and combining multi-view consistency constraints and data augmentation engines, a multi-scene sign language synthesis dataset is generated.
It improves the perceptual realism and multi-view consistency of synthetic sign language data, reduces computational overhead, enriches the diversity and breadth of the dataset, and enhances the training quality and generalization ability of the sign language recognition model.
Smart Images

Figure CN121661436A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer data processing technology, and in particular to a method, system, device and medium for generating multi-scene sign language synthesis datasets based on 3D digital humans. Background Technology
[0002] Sign language, as the primary means of communication in the deaf community, possesses a unique visual-spatial grammatical structure that carries rich information, making it irreplaceable in promoting information accessibility and building a more inclusive social environment. Traditional sign language datasets primarily rely on live video recordings. However, this approach faces numerous inherent challenges and limitations in practice. On one hand, the filming process is constrained by the physical environment, resulting in a severe lack of scene diversity in the dataset. On the other hand, a single filming perspective limits the model's ability to learn sign language features from multiple angles, thus affecting its robust recognition performance for sign language from different perspectives in practical applications. More seriously, the cost of collecting live video data is high, and the subsequent manual annotation work is extremely time-consuming and labor-intensive, significantly reducing the return on investment and making the construction of large-scale, high-quality datasets exceptionally difficult.
[0003] Against this backdrop, sign language synthesis methods based on 3D digital human technology have emerged and are gradually becoming a potential way to overcome the bottlenecks of traditional data generation. However, existing 3D sign language synthesis systems generally lack effective perception and dynamic adjustment mechanisms for the eye-tracking attention distribution of real humans (especially deaf viewers) when generating motion and details in key areas. This greatly reduces the perceptual realism of the synthesized data. Furthermore, after the 3D model generates sign language sequences, passive rendering using multiple externally placed virtual cameras results in multi-view data that is difficult to effectively support sign language recognition tasks that require feature fusion from multiple angles. It may even introduce erroneous training signals, affecting the model's generalization ability.
[0004] Therefore, existing technologies have significant shortcomings in the synthesis of sign language datasets based on 3D digital human technology. More advanced methods are needed to improve the realism and consistency of synthetic sign language datasets. It is necessary to provide a method for generating multi-scene synthetic sign language datasets based on 3D digital humans to solve the above problems. Summary of the Invention
[0005] This invention provides a method, system, device, and medium for generating multi-scene sign language synthesis datasets based on 3D digital humans, in order to solve technical problems such as insufficient perceptual realism and lack of geometric and semantic consistency in multi-viewpoint sign language synthesis data in the prior art.
[0006] In a first aspect, embodiments of the present invention provide a method for generating a multi-scene sign language synthesis dataset based on a 3D digital human, including: Collect real sign language eye-tracking data to generate attention weight masks; Based on the attention weight mask and sign language semantic input, a 3D digital human foreground is generated through a semantic encoder and an attention region generation subnetwork, and the 3D digital human foreground is fused with a preset 3D background to obtain an initial 3D scene. The 3D digital human in the initial 3D scene is optimized by multi-view consistency constraints; Based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated through a data augmentation engine.
[0007] Optionally, collect real sign language eye-tracking data to generate attention weight masks, including: Collect real sign language eye-tracking data and preprocess the eye-tracking data; The preprocessed eye-tracking data is mapped onto the surface of a preset three-dimensional digital human model to generate a dynamic attention heatmap of the key areas of the digital human model. The attention weight mask is generated based on the dynamic attention heatmap.
[0008] Optionally, based on the attention weight mask and sign language semantic input, a 3D digital human foreground is generated through a semantic encoder and an attention region generation subnetwork. The 3D digital human foreground is then fused with a preset 3D background to obtain an initial 3D scene, including: Based on the sign language semantic input, a semantic latent vector is generated by the semantic encoder; Based on the semantic latent vector and the attention weight mask, a 3D digital human foreground is generated through an attention region generation subnetwork. The 3D digital human foreground is blended with a preset 3D background using a background scene injector to obtain an initial 3D scene.
[0009] Optionally, the step of generating a 3D digital human foreground based on the semantic latent vector and the attention weight mask through an attention region generation subnetwork includes: The attention region generation subnetwork generates an initial 3D mesh model of the 3D digital human based on the semantic latent vector, and refines the initial 3D mesh model based on the attention weight mask; The attention region generation subnetwork generates a 3D digital human rendering texture map based on the semantic latent vector and the attention weight mask; The attention region generation subnetwork generates initial skeletal animation parameters for the 3D digital human based on the semantic latent vector, and optimizes the motion trajectory of key skeletal chains in the initial skeletal animation parameters based on the attention weight mask.
[0010] Optionally, the 3D digital human in the initial 3D scene is optimized through multi-view consistency constraints, including: Generate multi-view images based on at least three pre-set virtual cameras with different perspectives; Calculate the geometric consistency loss and semantic consistency loss between different perspectives; Based on the geometric consistency loss and the semantic consistency loss, the parameters of the 3D digital human foreground are optimized in reverse.
[0011] Optionally, based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated using a data augmentation engine, including: Based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated by changing the digital human's identity, background scene, lighting and weather, and camera parameters through a data augmentation engine.
[0012] Secondly, embodiments of the present invention provide a system for generating a multi-scene sign language synthesis dataset based on a 3D digital human. The system is used to execute the method for generating a multi-scene sign language synthesis dataset based on a 3D digital human according to any embodiment of the present invention, including: The acquisition module is used to collect real sign language eye-tracking data and generate attention weight masks; The fusion module is used to generate a 3D digital human foreground based on the attention weight mask and sign language semantic input, through a semantic encoder and attention region generation subnetwork, and to fuse the 3D digital human foreground with a preset 3D background to obtain an initial 3D scene. An optimization module is used to optimize the 3D digital human in the initial 3D scene through multi-view consistency constraints; The generation module is used to generate a multi-scene sign language synthesis dataset based on the optimized 3D digital human through a data augmentation engine.
[0013] Optionally, the acquisition module is specifically used for: Collect real sign language eye-tracking data and preprocess the eye-tracking data; The preprocessed eye-tracking data is mapped onto the surface of a preset three-dimensional digital human model to generate a dynamic attention heatmap of the key areas of the digital human model. The attention weight mask is generated based on the dynamic attention heatmap.
[0014] Thirdly, embodiments of the present invention provide an electronic device, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to execute the method for generating a multi-scene sign language synthesis dataset based on a 3D digital human according to any embodiment of the present invention.
[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions, which are used to cause a processor to execute the method for generating a multi-scene sign language synthesis dataset based on a 3D digital human as described in any embodiment of the present invention.
[0016] (1) By introducing an attention guidance mechanism based on eye-tracking data of real deaf users, the embodiments of the present invention can prioritize the allocation of computational resources and detail fidelity to key expression areas of sign language (hands, face, mouth), thereby ensuring that the geometric details, texture quality and motion smoothness of these areas reach an extremely high level, solving the problems of key detail distortion and unnatural motion in traditional methods. It significantly improves the perceptual realism and dynamic naturalness of sign language synthesis data.
[0017] (2) This embodiment of the invention constructs cross-view geometric consistency loss and semantic consistency loss, thereby constraining the uniformity of the geometric shape and semantic connotation of sign language actions and expressions under different virtual camera perspectives during the generation of 3D digital humans. This solves the problems of inconsistent action-expression and self-occlusion processing errors in the prior art under multiple perspectives, and provides high-quality training data for multi-view sign language recognition tasks. It achieves a high degree of geometric and semantic consistency in the generation of multi-view sign language data.
[0018] (3) In this embodiment of the invention, only high-precision modeling and rendering (foreground) are performed on high-attention areas, while the background is processed using a low-overhead model. Through depth map synthesis technology, the foreground and any diverse background are seamlessly integrated, achieving efficient large-scale data generation. This effectively avoids the high computational cost and background bias problems caused by the tight coupling between the foreground and background in traditional methods, enhances the generalization ability of the sign language recognition model to complex scenes, significantly reduces computational overhead, and effectively avoids background bias.
[0019] (4) Based on the optimized 3D digital human, the present invention uses a data augmentation engine to quickly generate data variants containing different digital human identities, diverse scenes, changing lighting, various weather effects and different camera postures, which greatly enriches the dimensions and breadth of the dataset, has a powerful ability to generate diverse datasets, and comprehensively improves the quality, efficiency and diversity of the 3D sign language synthesis dataset.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of a method for generating a multi-scene sign language synthesis dataset based on a 3D digital human, provided in Embodiment 1 of the present invention; Figure 2 This is a framework diagram of a multi-scene sign language synthesis dataset generation system based on 3D digital humans, provided in Embodiment 2 of the present invention. Figure 3 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] Application Overview: Existing 3D sign language synthesis systems generally lack an effective perception and dynamic adjustment mechanism for the distribution of eye-tracking attention of real humans (especially deaf viewers) when generating motion and details in key areas such as hands, face, and lip movements. This means that the system cannot intelligently allocate more computational resources and detail fidelity to areas of high viewer focus during the generation process, resulting in distortions or unnatural movements in key details such as the curvature of finger bends, facial muscle movements, and subtle changes in lip shapes, significantly reducing the perceptual realism of the synthesized data. This non-perceptual, uniform rendering strategy may result in insufficient semantic accuracy in the generated sign language, even with an increase in the quantity of synthesized data, due to its deficiency in the core dimension of "realism."
[0026] Furthermore, when faced with the need to generate sign language data from multiple perspectives, existing 3D synthesis systems often exhibit geometric or semantic inconsistencies between hand and facial movements from different virtual camera viewpoints. For example, there may be unreasonable self-occlusion processing errors, or semantic conflicts may arise between facial expressions and hand gestures (e.g., a gesture expresses a question, but the facial expression shows affirmation). Simply passively rendering using multiple externally placed virtual cameras cannot effectively support sign language recognition tasks that require feature fusion from multiple angles, and may even introduce incorrect training signals, affecting the model's generalization ability.
[0027] Furthermore, existing sign language synthesis methods generally face enormous computational challenges when pursuing scene diversity and large-scale data generation. This not only consumes massive amounts of computing resources but also greatly limits the speed and scale of data generation. This high computational cost and background-coupled rendering mode make the rapid, efficient, and large-scale generation of highly realistic, diverse, and background-free sign language data a pressing problem that needs to be solved.
[0028] This invention provides a method for generating multi-scene sign language synthesis datasets based on 3D digital humans, aiming to solve technical problems in existing technologies such as insufficient perceptual realism, lack of multi-view geometric and semantic consistency, huge computational overhead, and susceptibility to background interference. This invention generates an attention weight mask by collecting real sign language eye-tracking data; based on the attention weight mask and sign language semantic input, a 3D digital human foreground is generated through a semantic encoder and an attention region generation subnetwork; the 3D digital human foreground is then fused with a preset 3D background to obtain an initial 3D scene; the 3D digital human in the initial 3D scene is optimized through multi-view consistency constraints; based on the optimized 3D digital human, a data augmentation engine is used to generate a multi-scene sign language synthesis dataset, thereby efficiently and on a large scale generating sign language synthesis datasets with high perceptual realism, multi-view consistency, and diverse scenes. Example 1
[0029] Figure 1 This is a flowchart of a method for generating a multi-scene sign language synthesis dataset based on a 3D digital human, provided in Embodiment 1 of the present invention. This embodiment is applicable to the generation of sign language datasets, and the method can be executed by a multi-scene sign language synthesis dataset generation system based on a 3D digital human. Figure 1 As shown, the method includes: S110. Collect real sign language eye-tracking data and generate an attention weight mask. S120. Based on the attention weight mask and sign language semantic input, generate a 3D digital human foreground using a semantic encoder and an attention region generation subnetwork, and fuse the 3D digital human foreground with a preset 3D background to obtain an initial 3D scene. S130. Optimize the 3D digital human in the initial 3D scene using multi-view consistency constraints. S140. Based on the optimized 3D digital human, generate a multi-scene sign language synthesis dataset using a data augmentation engine.
[0030] In this embodiment, an attention weight mask is generated by collecting real sign language eye-tracking data. Based on the attention weight mask and sign language semantic input, a 3D digital human foreground is generated through a semantic encoder and an attention region generation subnetwork. The 3D digital human foreground is then fused with a preset 3D background to obtain an initial 3D scene. The 3D digital human in the initial 3D scene is optimized through multi-view consistency constraints. Based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated through a data augmentation engine. This enables the efficient and large-scale generation of sign language synthesis datasets with high perceptual realism, multi-view consistency, and diverse scenes, effectively improving the training quality and generalization ability of the automatic sign language recognition model.
[0031] The following will describe each step in detail.
[0032] In step S110, real sign language eye-tracking data is collected, and an attention weight mask is generated, including: Collect real sign language eye-tracking data and preprocess the eye-tracking data; The preprocessed eye-tracking data is mapped onto the surface of a preset three-dimensional digital human model to generate a dynamic attention heatmap of the key areas of the digital human model. The attention weight mask is generated based on the dynamic attention heatmap.
[0033] Specifically, eye-tracking data is collected from real deaf users watching standard sign language video sequences using an eye-tracking device. This eye-tracking data includes, but is not limited to, the two-dimensional coordinate sequence (X, Y) of the fixation point on the screen, the duration of each fixation event, the saccade vector information of the eyeball, and auxiliary physiological indicators such as pupil diameter changes. The eye-tracking device is a high-speed eye tracker with a sampling frequency of 200Hz and a visual angle accuracy of 0.5 degrees. The sampling frequency can be set to 200Hz to capture subtle changes in instantaneous eye movements and ensure that its visual angle accuracy reaches 0.5 degrees or higher, thereby accurately locking the fixation point position. The eye-tracking device also integrates a high-definition video recording function to ensure that the eye-tracking data and the sign language video sequence watched by the subject can achieve millisecond-level time synchronization. The standard sign language video sequence covers commonly used words, phrases, and sentences in Chinese Sign Language (CSL) or American Sign Language (ASL), with each sequence lasting approximately 3 to 5 seconds, designed to simulate real sign language communication scenarios.
[0034] The collected eye-tracking data is preprocessed, including data calibration, noise filtering, and timestamp synchronization.
[0035] Specifically, the preprocessing workflow includes data calibration, noise filtering, and timestamp synchronization. For example, a standard 9-point calibration or 5-point calibration paradigm can be used to ensure that the deviation between the gaze point output by the eye tracker and the actual screen gaze position is controlled within 0.5 degrees of visual angle. A median filter or a Kalman filter-based smoothing algorithm is applied to eliminate transient artifacts in the eye-tracking signal introduced by physiological tremors or device noise. For instance, if a median filter is used, the window size can be set to 5 sampling points; if a Kalman filter is used, the process noise covariance Q and measurement noise covariance R in the Kalman filter can be dynamically adjusted according to the actual data characteristics. By analyzing the cross-correlation between the eye-tracking data stream and the video frame timestamps, the eye-tracking data is accurately aligned with the video frames, ensuring that each gaze point corresponds to precise video content. During the preprocessing stage, it is also possible to identify fixation events and saccade events. For example, the I-VT (Velocity-Threshold) algorithm can be used to identify movements with an eye angular velocity exceeding 30 degrees / second as saccades, and events below this threshold with a duration exceeding 100 milliseconds and a diffusion angle of less than 0.5 degrees as fixations.
[0036] After data preprocessing is completed, the preprocessed eye-tracking data is mapped onto the surface of a preset standard 3D digital human model using the Gaussian kernel density estimation method, generating a dynamic attention heatmap of the key regions of the standard 3D digital human model.
[0037] The standard 3D digital human model is a high-precision parametric model with complete skeletal rigging and facial blendshapes, such as a model generated by Epic Games' MetaHuman Creator or an industry-standard SMPL-X model, whose topology has been predefined.
[0038] Specifically, 2D gaze points can be inversely projected back into 3D space using camera intrinsics and depth information used during rendering to obtain a series of 3D gaze points. These 3D gaze points are then projected onto the surface of the digital human model, for example, through nearest neighbor lookup or texture mapping based on UV coordinates. The core of the Gaussian saliency estimation method lies in treating each gaze point as the center of a Gaussian distribution. By superimposing these Gaussian distributions, a dynamic attention heatmap is generated on the surface of the digital human model. The kernel function is typically a two-dimensional Gaussian distribution, and its mathematical expression can be represented as: in,( , The coordinates of the gaze point (x_i, y_i) are the x- and y-axis components of the distance from a gaze point (x_i, y_i) to the point to be evaluated, where u = x - x_i and v = y - y_i. , ) represents the bandwidth parameter. The correlation coefficient is used. The bandwidth parameter is not a fixed value but is adaptively adjusted based on the gaze point distribution density to ensure the smoothness and accuracy of the heatmap. For example, in densely populated gaze point regions, the bandwidth parameter is appropriately reduced to preserve details; in sparse regions, the bandwidth parameter is increased to ensure a smooth transition. An optional implementation is to use the K-Nearest Neighbor (K-NN) method, setting the bandwidth to the distance from each gaze point to its Kth nearest neighbor.
[0039] Heatmaps are specifically generated for the digital human's hands, face, and mouth. These areas are key carriers of emotional expression. The hand area covers all bones and skin surfaces from the wrist joint to the fingertips; the facial area includes all facial features from the forehead to the jaw and from the left ear to the right ear, especially the eye, nose, and lip areas; the mouth area is precisely focused on the lip line and the surrounding muscle tissue.
[0040] Based on the dynamic attention heatmap, a preset attention threshold is set to identify and dynamically delineate high-attention areas such as hands, face, and mouth shape.
[0041] Specifically, this threshold can be determined based on experimental statistical analysis, for example, set to the top 10% or 20% of the heatmap intensity value distribution, or a fixed intensity value (e.g., a threshold set to 0.6 after normalization to the 0-1 range). By binarizing the heatmap or through connected component analysis, the sets of pixels or vertices in these high-attention regions can be precisely delineated.
[0042] Finally, based on the high attention regions and their weight distribution in the heatmap, a series of frame-by-frame attention weight masks are generated.
[0043] The attention weight mask is a spatially varying real-valued matrix aligned with the topology of the digitized human model's surface. Its values are strictly limited to 0 to 1, representing the priority of detail fidelity and motion smoothness for each region of the digitized human model's surface during the generation process. A weight value of 1 indicates that the region enjoys the highest priority in subsequent generation processes, requiring the highest level of detail fidelity and motion smoothness; while a weight value of 0 indicates the lowest priority, allowing for lower generation precision, thus saving computational resources. This mask is stored as an 8-bit grayscale image (pixel values 0-255, corresponding to weights 0-1) or as a vertex color attribute, and is strictly synchronized with each frame of the digitized human animation sequence.
[0044] In step S120, based on the attention weight mask and sign language semantic input, a 3D digital human foreground is generated through a semantic encoder and an attention region generation subnetwork. The 3D digital human foreground is then fused with a preset 3D background to obtain an initial 3D scene, including: Based on the sign language semantic input, a semantic latent vector is generated by the semantic encoder; Based on the semantic latent vector and the attention weight mask, a 3D digital human foreground is generated through an attention region generation subnetwork. The 3D digital human foreground is blended with a preset 3D background using a background scene injector to obtain an initial 3D scene.
[0045] Specifically, it receives sign language text or a symbol sequence (Gloss sequence) as semantic input, and combines with the attention weight mask generated in Step 1. Through a semantic encoder, an attention region generation sub-network, and a background scene injector, it generates a digital human foreground with high perceptual realism and fuses it with diverse background scenes. Among them, the Gloss sequence is a linguistic representation of sign language. For example, the sign language word "hello" can be represented as [HELLO], and a phrase "How are you recently?" can be represented as [YOU][RECENTLY][HOW].
[0046] Specifically, the semantic encoder is a sequence-to-sequence (Seq2Seq) neural network model based on the Transformer architecture. Its input is the Gloss sequence of sign language, and its output is a compact and semantically rich semantic latent vector. The Transformer encoder contains multiple layers of self-attention mechanisms and feed-forward neural network layers to capture long-range dependencies and temporal context information in the sign language semantic sequence. Exemplarily, this semantic encoder can be configured with 6 encoder layers, each layer containing 8 attention heads, and the hidden dimension is 512. The self-attention mechanism enables it to capture long-range dependencies in the sign language semantic sequence. For example, the correlation between gestures across multiple words, while the feed-forward neural network is responsible for further feature transformation. This encoder can be pre-trained on a corpus containing a large number of sign language Gloss sequences and corresponding real sign language videos to learn the mapping relationship between sign language semantics and visual forms. Its output is a semantic latent vector of a fixed dimension, such as a 512-dimensional floating-point vector, which contains the complete semantic information of the sign language sequence and is sufficient to drive subsequent digital human animation generation.
[0047] The attention region generation sub-network receives the semantic latent vector and the attention weight mask as inputs. Its core function is to generate high-precision 3D meshes, textures, and skeletal animations only for the high-attention regions specified by the attention weight mask (i.e., the hand, face, and mouth regions of the digital human). In this way, computing resources are intelligently allocated to the most critical expression regions, avoiding unnecessary ultra-high-precision processing of the entire digital human model.
[0048] This sub-network uses a deep generative model, such as a high-resolution variational autoencoder (VAE) or a denoising diffusion probabilistic model (DDPM), and its architecture design incorporates a conditional generation mechanism. The specific steps are as follows: The attention region generation sub-network generates an initial 3D mesh model of the 3D digital human according to the semantic latent vector, and refines the initial 3D mesh model according to the attention weight mask.
[0049] Specifically, in the 3D mesh generation stage, the attention region generation subnetwork generates an initial 3D mesh model of the digital human based on semantic latent vectors. This initial 3D mesh model is typically a medium-precision parametric mesh, for example, containing approximately 20,000 vertices and 40,000 faces, possessing a basic geometric structure. The attention weight mask is used as a guiding signal for geometric refinement in the 3D mesh generation stage, prompting the subnetwork to perform adaptive tessellation in regions with high weight values (such as fingertips, lip lines, and eyelids), increasing vertex density and the number of faces. For example, for regions with attention weights of 0.8 or higher, such as fingertips, lip lines, and eyelids, the subnetwork applies a first-order Loop subdivision or Catmull-Clark subdivision algorithm, subdividing each triangular face into four, doubling the vertex density; for regions with weights of 0.9 or higher, a second subdivision is performed, further increasing the vertex density. This achieves centimeter-level or even millimeter-level geometric precision in these critical regions. For example, the curvature of the fingers, especially the subtle deformations at the knuckles, will be represented more accurately and naturally through a high-density grid; the subtle deformations of the lips, eyelids, and eyebrows caused by facial expressions can also be captured in detail.
[0050] The attention region generation subnetwork generates a 3D digital human rendering texture map based on the semantic latent vector and the attention weight mask.
[0051] Specifically, in the texture generation stage, the attention region generation sub-network generates high-resolution physically based rendering (PBR) texture maps for the high-attention regions based on semantic latent vectors and attention weight masks. These texture maps include, but are not limited to, diffuse maps, normal maps, roughness maps, and metallic maps. The diffuse map defines the base color of the surface; the normal map simulates the microscopic geometric details of the surface; the roughness map defines the surface gloss; and the metallic map defines the surface metallic properties.
[0052] The attention weight mask guides texture sampling and detail synthesis during the texture generation stage, ensuring that higher resolution and more biologically realistic texture details are generated in areas with higher weight values. For example, the texture resolution in non-attention areas can be as low as 256x256 pixels, while in high-attention areas (such as the face and hands), it can reach 4096x4096 pixels, or even 8192x8192 pixels, to reveal microscopic features such as pores on facial skin, fingerprint textures on fingers, and fine wrinkles on lips. These high-frequency details can be achieved through procedural generation (e.g., based on Perlin noise or Worley noise), high-resolution scan data fusion, or texture synthesis algorithms (e.g., based on StyleGAN or diffusion models).
[0053] The attention region generation subnetwork generates initial skeletal animation parameters for the 3D digital human based on the semantic latent vector, and optimizes the motion trajectory of key skeletal chains in the initial skeletal animation parameters based on the attention weight mask.
[0054] Specifically, in the skeletal animation generation stage, the attention region generation subnetwork generates joint rotation and displacement parameters for the digital human skeletal rig based on semantic latent vectors. The skeletal rig follows industry standards, such as the MetaHuman skeletal system, and includes hundreds of movable joints. The attention weight mask is further integrated into the inverse kinematics (IK) solver or pose generation model during the skeletal animation generation stage to optimize the motion trajectory of key skeletal chains. Key skeletal chains include 21 joints in the hand, such as the proximal, middle, and distal phalangeal joints of each finger, as well as the metacarpal joints; and 52 blendshape parameters for the face, such as those controlling eyebrow raising, eye blinking, mouth opening and closing, and cheek muscle movement. The optimization process prioritizes ensuring the smoothness, accuracy, and biological realism of movements in high-attention areas, avoiding "rubber hand" or "mask face" effects. For example, the minute trembling frequencies of the fingers (typically in the 5-10 Hz range to enhance realism), the precise synchronization of lip movements and sign language pronunciation (millisecond-level delay control), and the natural facial expressions resulting from eye muscle twitching are all achieved through fine-grained animation control guided by attention weight masks. The skeletal animation generation can employ sequence models based on recurrent neural networks (RNNs, such as LSTM or GRU) or temporal convolutional networks (TCNs). These models can learn and predict the temporal sequence of joint movements, thereby ensuring the smoothness and coherence of the animation in the temporal dimension.
[0055] The background scene injector is responsible for seamlessly blending the high-fidelity digital human foreground (including detailed 3D meshes, PBR textures, and skeletal animation) output by the attention region generation subnetwork with an arbitrary preset 3D environment background. The background scene injector receives the 3D geometric information of the digital human foreground (e.g., vertex coordinates, normals, UV coordinates) and a rendered depth map, as well as a pre-selected 3D environment background model. The background model can be in various forms, such as a low-precision environment geometry (containing thousands of faces), a high dynamic range panorama (HDR Skybox for ambient lighting and reflections), a pre-rendered static scene image, or a procedurally generated background. The blending process includes: First, depth map-based geometric compositing is performed. By comparing the foreground depth map of the digitizer with the background scene depth map, the relative occlusion relationship between the foreground and background is accurately determined, ensuring that the digitizer is correctly positioned within the scene and correctly occluded by objects in the scene, and vice versa. The depth map is typically 16-bit floating-point precision (e.g., using OpenEXR format), providing precise information on the distance of each pixel from the camera to the surface to ensure accurate fusion. In the rendering pipeline, this is typically achieved through Z-buffering or depth testing, ensuring that only the surface pixels closest to the camera are rendered.
[0056] Secondly, lighting and shadow matching is performed. The background scene injector analyzes light probes or environment maps in the 3D environment background model to extract the direction of the main light source, ambient light intensity, and color of the scene. For example, if the background is HDRI, the average ambient light and the direction of the main light source can be estimated from it. Then, the foreground of the digital human is re-rendered and tinted according to these lighting parameters to achieve visual consistency between the foreground and background under lighting conditions. This process involves the PBR rendering equation, which adjusts the material parameters of the digital human (such as diffuse, specular, and subsurface scattering) to respond to the lighting conditions of the scene and generate shadows that conform to the scene's lighting conditions. Shadow generation can use real-time shadow mapping technology, or, for higher quality requirements, ray tracing technology to generate soft shadows.
[0057] Finally, post-processing effects are unified. The background scene injector applies post-processing effects such as depth of field, motion blur, and film grain to achieve visual style and realism uniformity in the fused image. Depth of field, by simulating the optical characteristics of a camera lens, makes the focal area sharp while blurring the foreground and background; its parameters can be adjusted according to focal length, aperture value, and blur radius. Motion blur increases the smoothness and realism of the video by applying a blur effect to moving objects; its intensity can be calculated based on the object's velocity vector and shutter speed. Film grain simulates the texture of traditional film photography by overlaying subtle noise textures. Other common cinematic post-processing effects include color correction, exposure adjustment, vignetting, and chromatic aberration correction.
[0058] This decoupled generation and fusion approach significantly reduces the computational overhead of high-precision modeling and rendering of the entire scene, while allowing for the rapid replacement and generation of various background scenes without regenerating the digital human foreground, greatly improving the diversity and efficiency of data generation.
[0059] In step S130, the 3D digital human in the initial 3D scene is optimized through multi-view consistency constraints, including: Generate multi-view images based on at least three pre-set virtual cameras with different perspectives; Calculate the geometric consistency loss and semantic consistency loss between different perspectives; Based on the geometric consistency loss and the semantic consistency loss, the parameters of the 3D digital human foreground are optimized in reverse.
[0060] Specifically, at least three virtual cameras with different perspectives are pre-set to comprehensively capture the sign language expressions of the digital human. For example, a virtual camera located directly in front of the digital human (0-degree perspective) simulates the perspective of a standard video call; a virtual camera located at a 45-degree angle to the right of the digital human (45-degree perspective) provides lateral geometric information; and a virtual camera located above the digital human at a 45-degree downward angle captures the height information of the gestures in three-dimensional space. All virtual cameras are precisely configured to have the same focal length (e.g., 50mm), aperture (f / 2.8), and sensor size (e.g., 36x24mm full-frame) to ensure consistency in their optical imaging characteristics.
[0061] For each virtual camera viewpoint, this step renders the corresponding 2D RGB video sequence, depth map, and 2D projected image of the attention region. The RGB video sequence captures the visual appearance; the depth map (16-bit floating-point precision) provides the precise distance of each pixel to the camera; and the 2D projected image of the attention region (8-bit grayscale image) indicates the position and intensity of the high-attention region in the 2D image at that viewpoint. These rendering tasks can be achieved by integrating with a professional 3D rendering engine API, such as using Blender's Cycles renderer or Unreal Engine's Path Tracer.
[0062] Calculate the cross-view geometric consistency loss, including: From the 2D RGB images rendered from each viewpoint, 21 keypoints for the hand, 68 keypoints for the face, and 17 keypoints for the body are extracted using a pre-trained keypoint detection model (e.g., MediaPipe Hand / Face / Pose or Open Pose). These keypoints include fingertips, knuckles, corners of the lips, corners of the eyes, tip of the nose, shoulders, elbows, etc. Then, using the known intrinsic parameters (e.g., focal length, principal point, radial distortion coefficient) and extrinsic parameters (e.g., rotation matrix and translation vector) of each virtual camera, these 2D keypoints are accurately back-projected into 3D space via inverse projection (e.g., triangulation or homography matrix decomposition) to form a 3D point cloud.
[0063] The reprojection error or L1 distance between these 3D point clouds is calculated across different viewpoints as a measure of geometric consistency. Specifically, for any pair of viewpoints i and j (where i ≠ j), the 3D points of one viewpoint are... (From the 2D point of view i) The result (obtained by inverse projection) is reprojected onto a 2D plane at another viewpoint j, and its relationship to the 2D points directly observed from viewpoint j is calculated. The Euclidean distance between them. The reprojection errors of key points from all viewpoints are accumulated and averaged to obtain the cross-viewpoint geometric consistency loss. The formula for calculating the geometric consistency loss can be precisely expressed as: in, The total number of key points. Let V be the total number of viewpoints, and let V be the set of viewpoints. Let represent the 2D projection function from 3D space to viewpoint j. Represents a 2D point from viewpoint i. Inverse projection function to 3D space (combined with depth information). Let represent the 2D coordinates of the k-th keypoint in viewpoint i. A geometric consistency loss function is used to force the pose and mesh deformation of the digital human skeleton to remain highly consistent in 3D space across all viewpoints.
[0064] Further, a semantic consistency loss is calculated. This semantic consistency loss aims to ensure the consistency of sign language semantics expressed by the digital human from different perspectives. This loss is achieved by inputting the 2D attention region image rendered from each virtual camera perspective (specifically including close-up hand and face regions) into a pre-trained lightweight gesture classifier and facial expression recognizer, respectively. The gesture classifier is a classification model based on a convolutional neural network (CNN), such as MobileNetV3 or EfficientNet-B0, whose training data contains a large number of images with gesture labels. It can classify sign language gestures into a pre-defined gesture vocabulary, for example, recognizing hand shapes representing specific sign language words such as "apple," "drink water," and "thank you." The facial expression recognizer is a classification model based on a deep residual network (e.g., ResNet), such as ResNet-18, whose training data contains various expression labels (e.g., question, affirmation, neutral, surprise, sadness, happiness). It can recognize the intensity and category of facial expressions. The semantic consistency loss is calculated by comparing the differences between the outputs of the classifier and the recognizer (e.g., predicted class probability distribution vectors or feature embedding vectors) from different perspectives. For example, KL divergence can be used to measure the consistency between two probability distributions. in, It is the probability that viewpoint i is predicted to be category c. This represents the probability that viewpoint j is predicted as category c. Alternatively, cosine similarity loss can be used to measure the consistency between feature embedding vectors. When different perspectives produce inconsistent gesture classification or facial expression recognition results for the same sign language semantic input (e.g., one perspective identifies it as "question" while another identifies it as "affirmation"), the semantic consistency loss will be activated and amplified, thereby forcing the generator to maintain cross-perspective consistency at the semantic level.
[0065] Furthermore, this embodiment introduces an occlusion awareness mechanism. This mechanism utilizes the depth map rendered by each virtual camera to identify self-occluded regions generated by the digital human itself. Self-occlusion refers to a part of the digital human's body obscuring another part (e.g., a finger obscuring the palm, an arm obscuring the face). For keypoints or pixels self-occluded at a certain viewpoint, the occlusion awareness mechanism dynamically reduces the weight of the keypoint or pixel in that region when calculating the geometric consistency loss and semantic consistency loss. For example, when a keypoint is completely occluded at a certain viewpoint (i.e., its depth value is greater than the depth value of the occluding object), its weight in the loss calculation will drop to zero; when partially occluded, the weight decreases linearly according to the occlusion ratio (e.g., the percentage of occluded pixels). For example, if a finger keypoint is occluded by 50% of its area, its weight can drop from 1.0 to 0.5. This mechanism effectively avoids erroneous gradient signals introduced by unavoidable self-occlusion, preventing the optimizer from being misled by incorrect or incomplete observations during optimization, thereby improving the effectiveness and robustness of constraints.
[0066] The cross-view geometric consistency loss, semantic consistency loss, and the combined loss weighted by the occlusion perception mechanism are backpropagated to the semantic encoder, the attention region generation subnetwork, and the background scene injector to iteratively optimize the 3D pose parameters (skeletal joint rotation and displacement), mesh deformation parameters (vertex displacement), and facial expression blending shape parameters of the digital human. The total loss function can be expressed as: in, It is geometric consistency loss. It is a semantic consistency loss. It is a temporal smoothing loss used to ensure the smoothness of animation in the time dimension. , , These are hyperparameters used to balance the weights of various losses. The optimization process can employ stochastic gradient descent (SGD) or the Adam optimizer, such as the AdamW optimizer. The initial learning rate can be set to 1e-4 and dynamically adjusted according to the training phase using strategies such as cosine annealing. Through this feedback mechanism, the actions and expressions of the digital human sign language will achieve high coordination and consistency across multiple perspectives, enhancing the practicality of the generated dataset in multi-view sign language recognition tasks.
[0067] In this embodiment, a multi-view consistency constraint mechanism is introduced to ensure that the geometry and semantics of digital human sign language actions are highly consistent under different virtual camera perspectives, thereby solving the problem of inconsistent multi-view data in traditional methods.
[0068] In step S140, based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated through a data augmentation engine. The sign language synthesis dataset generated based on the optimized 3D digital human includes: multi-view RGB video sequences, corresponding depth map sequences, attention mask sequences, 3D skeletal keypoint data, scene labels and metadata, and eye-tracking attention heatmap data.
[0069] Among them, the multi-view RGB video sequence can refer to a color video sequence rendered from at least three different preset viewpoints (such as front, 45-degree side, and 45-degree top view) that conforms to human vision. The corresponding depth map sequence can refer to an image sequence that is strictly synchronized frame by frame with the RGB video sequence and records the physical distance (depth) information between each pixel and the virtual camera. The attention mask sequence can refer to a sequence that is synchronized with the RGB video sequence and identifies the location and weight intensity of "high attention areas" (such as hands, face, and mouth shapes) on the surface of the digital human model in image form. The 3D skeletal keypoint data can refer to the 3D coordinates and confidence scores of the key joints of the entire digital human body in each frame, recorded in a structured data format (such as JSON / CSV). Scene labels and metadata can refer to a set of text labels stored in a structured data format (such as JSON) that describes the background information and semantic context of the entire sign language sequence. Eye-tracking attention heatmap data can refer to the raw data or visualized image sequence provided as supplementary data that records the distribution of visual attention of real deaf users on the surface of the digital human model when watching corresponding sign language content.
[0070] Specifically, the multi-view RGB video sequences are stored in H.264 or H.265 encoding format, with a frame rate of 30 frames per second and a resolution of 1920x1080 pixels. The bitrate is typically set to 10Mbps to balance file size and visual quality. Each sign language sequence contains at least three video streams from different perspectives, each corresponding to a previously set virtual camera.
[0071] The corresponding depth map sequence is strictly synchronized with the RGB video sequence and stored in 16-bit PNG format, providing precise depth information for each pixel. The depth values (i.e., grayscale values) can be mapped to real-world distances (e.g., 0-65535 corresponds to 0-100 meters) and can be used for geometric analysis in 3D reconstruction, SLAM, or deep learning tasks.
[0072] The attention mask sequence is synchronized with the RGB video sequence and stored in 8-bit grayscale PNG format. The pixel values in the image directly represent the attention weight of that region, ranging from 0 (lowest weight) to 255 (highest weight), clearly identifying the pixel positions and weight strengths of high-attention regions such as hands, faces, and mouth shapes.
[0073] 3D skeletal keypoint data is stored in JSON or CSV format, containing the 3D coordinates and confidence scores of keypoints for the entire digital human in each video frame. The full-body keypoints include 21 joints in the hands, 68 facial landmarks (e.g., from the Face Mesh model), and 17 skeletal points in the body (e.g., from the OpenPose or MediaPipe Pose model). The coordinate system can be either world coordinates or the digital human's local coordinate system, and the confidence scores are estimated either from the output of the keypoint detection model or based on visibility at rendering time.
[0074] Scene tags and metadata are stored in JSON format, containing rich semantic information. For example, scene category (indoor / outdoor), specific scene description (e.g., classroom, subway station, rainy night street, sunny beach), lighting conditions (direct light, diffused light, night, backlight), weather conditions (sunny, cloudy, rainy / snowy, foggy), digital human identity information (gender, age range, body type, skin color, clothing type), sign language (Chinese Sign Language / American Sign Language), and original sign language Gloss sequence, etc.
[0075] Eye-tracking attention heatmap data, stored as an additional supervisory signal or research reference, is stored in the form of an 8-bit grayscale image or a data matrix. It records the attention distribution of deaf users when watching each frame and can be used for attention mechanism design or user behavior analysis in subsequent model training.
[0076] Furthermore, based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated by changing the digital human's identity, background scene, lighting and weather, and camera parameters through a data augmentation engine.
[0077] Specifically, changing the digital human identity: By replacing different 3D digital human models (e.g., selecting from a pre-defined library of 50 digital human models with different genders, ages, body types, skin colors, and clothing styles) and rerunning the generation process, the same sign language sequence can be generated in diverse ways under different digital human identities. This helps improve the sign language recognition model's ability to generalize to different individual characteristics.
[0078] Changing background scenes: By combining the high-fidelity digital human foreground with different preset or procedurally generated 3D environmental backgrounds selected from the scene asset library (e.g., including 100 indoor and outdoor scenes, such as cafes, libraries, parks, city streets, etc.), data of the same sign language sequence in various complex scenes can be quickly generated.
[0079] Varying lighting conditions: By adjusting the parameters of virtual light sources in the scene (e.g., light source position, intensity, color, and type, such as directional light, point light, and area light), data of the same sign language sequence can be generated under different lighting conditions. For example, it can simulate strong direct sunlight at noon on a sunny day, soft diffused light on a cloudy day, artificial light sources in an indoor environment, low light at night, and complex backlit scenes.
[0080] Transforming Weather and Environmental Effects: By overlaying weather effects such as rain, snow, fog, and wind during rendering, and applying corresponding physical interactions and visual effects to the digital human foreground (e.g., raindrops reflecting off skin, clothing swaying in the wind, hair being blown by the wind), more realistic and complex environmental data can be generated. This requires the integration of a physics simulation engine, such as NVIDIA PhysX or Havok Physics.
[0081] Transform camera parameters: By adjusting parameters such as the virtual camera's position (e.g., translation ±2 meters on the XYZ axes), rotation (e.g., pitch, yaw, roll angle changes ±30 degrees), focal length (e.g., from 35mm to 85mm), and aperture (f / 1.8 to f / 11), data on the same sign language sequence under different shooting angles, distances, and lens movements (e.g., translation, zoom, panning, tilting) can be generated to simulate diverse shooting conditions.
[0082] Through the aforementioned ability to generate large-scale and diverse datasets, this invention can provide unprecedented high-quality and rich data support for the training of sign language automatic recognition models, significantly improving the robustness and generalization performance of the models, and making them more stable in complex real-world scenarios. Example 2
[0083] Figure 2 This is a framework diagram of a multi-scene sign language synthesis dataset generation system based on a 3D digital human, provided in Embodiment 2 of the present invention. The system is used to execute the multi-scene sign language synthesis dataset generation method based on a 3D digital human, as described in any one of the embodiments of the present invention. Figure 2 As shown, the system includes: The acquisition module 210 is used to acquire real sign language eye-tracking data and generate attention weight masks; The fusion module 220 is used to generate a 3D digital human foreground based on the attention weight mask and sign language semantic input, through a semantic encoder and attention region generation subnetwork, and to fuse the 3D digital human foreground with a preset 3D background to obtain an initial 3D scene. Optimization module 230 is used to optimize the 3D digital human in the initial 3D scene through multi-view consistency constraints; The generation module 240 is used to generate a multi-scene sign language synthesis dataset based on the optimized 3D digital human through a data augmentation engine.
[0084] Optionally, the acquisition module 210 is specifically used for: Collect real sign language eye-tracking data and preprocess the eye-tracking data; The preprocessed eye-tracking data is mapped onto the surface of a preset three-dimensional digital human model to generate a dynamic attention heatmap of the key areas of the digital human model. The attention weight mask is generated based on the dynamic attention heatmap.
[0085] Optional, the fusion module 220 is specifically used for: Based on the sign language semantic input, a semantic latent vector is generated by the semantic encoder; Based on the semantic latent vector and the attention weight mask, a 3D digital human foreground is generated through an attention region generation subnetwork. The 3D digital human foreground is blended with a preset 3D background using a background scene injector to obtain an initial 3D scene.
[0086] Optionally, the fusion module 220 is also used for: The attention region generation subnetwork generates an initial 3D mesh model of the 3D digital human based on the semantic latent vector, and refines the initial 3D mesh model based on the attention weight mask; The attention region generation subnetwork generates a 3D digital human rendering texture map based on the semantic latent vector and the attention weight mask; The attention region generation subnetwork generates initial skeletal animation parameters for the 3D digital human based on the semantic latent vector, and optimizes the motion trajectory of key skeletal chains in the initial skeletal animation parameters based on the attention weight mask.
[0087] Optional, optimization module 230, used for: Generate multi-view images based on at least three pre-set virtual cameras with different perspectives; Calculate the geometric consistency loss and semantic consistency loss between different perspectives; Based on the geometric consistency loss and the semantic consistency loss, the parameters of the 3D digital human foreground are optimized in reverse.
[0088] Optionally, module 240 is generated, specifically for: Based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated by changing the digital human's identity, background scene, lighting and weather, and camera parameters through a data augmentation engine.
[0089] The multi-scene sign language synthesis dataset generation system based on 3D digital human provided in this embodiment of the invention can execute the multi-scene sign language synthesis dataset generation method based on 3D digital human provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method. Example 3
[0090] Figure 3 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0091] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor 11, and the computer program is executed by the at least one processor 11 to enable the at least one processor 11 to perform the method provided by the present invention.
[0092] The processor 11 can perform various appropriate actions and processes based on a computer program stored in the read-only memory (ROM) 12 or a computer program loaded from the storage unit 18 into the random access memory (RAM) 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0093] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0094] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a method for generating a multi-scene sign language synthesis dataset based on a 3D digital human.
[0095] In some embodiments, a method for generating a multi-scene sign language synthesis dataset based on a 3D digital human can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for generating a multi-scene sign language synthesis dataset based on a 3D digital human described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform a method for generating a multi-scene sign language synthesis dataset based on a 3D digital human by any other suitable means (e.g., by means of firmware).
[0096] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0097] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0098] In the context of this invention, a computer-readable storage medium stores computer instructions that, when executed by a processor, implement the method for generating a multi-scene sign language synthesis dataset based on a 3D digital human provided by this invention. The computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0099] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0100] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0101] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0102] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0103] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for generating a multi-scene sign language synthesis dataset based on 3D digital humans, characterized in that, include: Collect real sign language eye-tracking data to generate attention weight masks; Based on the attention weight mask and sign language semantic input, a 3D digital human foreground is generated through a semantic encoder and an attention region generation subnetwork, and the 3D digital human foreground is fused with a preset 3D background to obtain an initial 3D scene. The 3D digital human in the initial 3D scene is optimized by multi-view consistency constraints; Based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated through a data augmentation engine.
2. The method according to claim 1, characterized in that, Collect real sign language eye-tracking data to generate attention weight masks, including: Collect real sign language eye-tracking data and preprocess the eye-tracking data; The preprocessed eye-tracking data is mapped onto the surface of a preset three-dimensional digital human model to generate a dynamic attention heatmap of the key areas of the digital human model. The attention weight mask is generated based on the dynamic attention heatmap.
3. The method according to claim 1, characterized in that, Based on the attention weight mask and sign language semantic input, a 3D digital human foreground is generated through a semantic encoder and an attention region generation subnetwork. The 3D digital human foreground is then fused with a preset 3D background to obtain an initial 3D scene, including: Based on the sign language semantic input, a semantic latent vector is generated by the semantic encoder; Based on the semantic latent vector and the attention weight mask, a 3D digital human foreground is generated through an attention region generation subnetwork. The 3D digital human foreground is blended with a preset 3D background using a background scene injector to obtain an initial 3D scene.
4. The method according to claim 3, characterized in that, The step of generating a 3D digital human foreground based on the semantic latent vector and the attention weight mask through an attention region generation subnetwork includes: The attention region generation subnetwork generates an initial 3D mesh model of the 3D digital human based on the semantic latent vector, and refines the initial 3D mesh model based on the attention weight mask; The attention region generation subnetwork generates a 3D digital human rendering texture map based on the semantic latent vector and the attention weight mask; The attention region generation subnetwork generates initial skeletal animation parameters for the 3D digital human based on the semantic latent vector, and optimizes the motion trajectory of key skeletal chains in the initial skeletal animation parameters based on the attention weight mask.
5. The method according to claim 1, characterized in that, The 3D digital human in the initial 3D scene is optimized through multi-view consistency constraints, including: Generate multi-view images based on at least three pre-set virtual cameras with different perspectives; Calculate the geometric consistency loss and semantic consistency loss between different perspectives; Based on the geometric consistency loss and the semantic consistency loss, the parameters of the 3D digital human foreground are optimized in reverse.
6. The method according to claim 1, characterized in that, Based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated through a data augmentation engine, including: Based on the optimized 3D digital human, a multi-scene sign language synthesis dataset is generated by changing the digital human's identity, background scene, lighting and weather, and camera parameters through a data augmentation engine.
7. A system for generating multi-scene sign language synthesis datasets based on 3D digital humans, characterized in that, The system is used to execute the method for generating a multi-scene sign language synthesis dataset based on a 3D digital human, as described in any one of claims 1-6, including: The acquisition module is used to collect real sign language eye-tracking data and generate attention weight masks; The fusion module is used to generate a 3D digital human foreground based on the attention weight mask and sign language semantic input, through a semantic encoder and attention region generation subnetwork, and to fuse the 3D digital human foreground with a preset 3D background to obtain an initial 3D scene. An optimization module is used to optimize the 3D digital human in the initial 3D scene through multi-view consistency constraints; The generation module is used to generate a multi-scene sign language synthesis dataset based on the optimized 3D digital human through a data augmentation engine.
8. The system according to claim 7, characterized in that, The acquisition module is specifically used for: Collect real sign language eye-tracking data and preprocess the eye-tracking data; The preprocessed eye-tracking data is mapped onto the surface of a preset three-dimensional digital human model to generate a dynamic attention heatmap of the key areas of the digital human model. The attention weight mask is generated based on the dynamic attention heatmap.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method for generating a multi-scene sign language synthesis dataset based on a 3D digital human, as described in any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the method for generating a multi-scene sign language synthesis dataset based on a 3D digital human, as described in any one of claims 1-6.