Sign language broadcast video generation method and system and storage medium
By performing speech recognition and multimodal emotion recognition on the video content, the action parameters of virtual images are determined, and the problem of insufficient sign language broadcast expression in the prior art is solved, rich and accurate sign language video generation is achieved, and the viewing experience of hearing-impaired people is improved.
Patent Information
- Application Number
- CN202510525665.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the posture library of digital human sign language broadcasting is constructed based on real human movement data, resulting in insufficient richness and accuracy of sign language expression and poor picture presentation effect.
By obtaining the audio information and video information of the video to be processed, performing speech recognition and multi-modal emotion recognition, determining expression parameters, lip shaped parameters and body movement parameters, driving virtual images to express sign language, and synchronizing with beat signals and subtitles to generate sign language broadcast videos.
It realizes efficient generation of sign language broadcast videos, improves the richness and accuracy of sign language expressions, and enhances the understanding and viewing experience of video content by hearing-impaired people.
Smart Images

Figure CN120075544A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sign language video, and more specifically, to a method, a system and a storage medium for generating a sign language broadcast video. Background Art
[0002] Currently, for the language translation of video programs for the hearing-impaired population, on-site translation or post-production by sign language interpreters is mainly adopted. Sign language interpreters convert audio content into sign language information through gestures, movements and facial expressions, etc., to meet the understanding needs of the hearing-impaired population. In the prior art, digital human technology has been introduced into the field of sign language translation, but there are still the following drawbacks in the existing digital human sign language broadcast: a) The existing sign language gesture library is mainly constructed based on real human motion data, which is limited by the range of human activities, resulting in insufficient richness and accuracy of digital human sign language expression. b) The picture presentation effect is not good.
[0003] Therefore, there is an urgent need for a method for generating a sign language broadcast video that can achieve efficient information transmission. Summary of the Invention
[0004] In view of the above problems, the purpose of the present invention is to provide a method, a system and a storage medium for generating a sign language broadcast video to solve at least one problem existing in the prior art.
[0005] According to one aspect of the present invention, there is provided a method for generating a sign language broadcast video, which is applied to an electronic device and includes: Obtain the audio information and video information of the video to be processed; wherein, the audio information includes speech information and non-speech information; Perform speech recognition on the speech information to obtain the text content corresponding to the speech information; perform multi-modal emotion recognition on the text content, non-speech information and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category; Determine the expression parameters, lip-sync parameters and limb movement parameters based on the sign language text, the emotion category and the probability distribution of the emotion category of the video to be processed; wherein, the sign language text is obtained by extracting the semantic summary of the text content; Drive the expression, lip-sync and limb movements of the virtual image respectively based on the expression parameters, the lip-sync parameters and the limb movement parameters, and generate a sign language broadcast video after synchronizing with the pre-obtained rich video information; wherein, the rich video information includes synchronized beat signals and subtitles; the beat signals are obtained by extracting beats from the non-speech information; the subtitles are generated according to the text content.
[0006] In addition, an alternative technical solution is a method for performing multi-modal emotion recognition on the text content, non-speech information, and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category, including: Extract features from the text content, non-speech information, and video information respectively to obtain audio features, video features, and text features; Perform dynamic alignment and cross-modal complementary information extraction on the audio features through audio-video cross-attention and audio-text cross-attention respectively, use the Hadamard product for cross-modal feature fusion, and perform normalization processing; obtain new audio features; perform dynamic alignment and cross-modal complementary information extraction on the video features through video-audio cross-attention and video-text cross-attention respectively, use the Hadamard product for cross-modal feature fusion, and perform normalization processing; obtain new video features; Perform dynamic alignment and cross-modal complementary information extraction on the text features, the new audio features, and the new video features through text-audio cross-attention and text-video cross-attention respectively, use the Hadamard product for cross-modal feature fusion, and perform normalization processing; obtain new text features; Concatenate the new text features, the new audio features, and the new video features to obtain multi-modal features; Input the multi-modal features into a preset classification network to obtain the emotion category of the video to be processed and the probability distribution of the emotion category.
[0007] In addition, an alternative technical solution is to use the Hadamard product for cross-modal feature fusion, which is implemented through the following formula;
[0008] Where, is the fusion result; is the cross-modal association weight of the audio-text cross-attention feature; is the cross-modal association weight of the audio-video cross-attention feature; is the Hadamard product;
[0009] Where, , , are the audio feature, text feature, and video feature respectively, and W is a learnable parameter; W Q , W K , W VThey are the weight matrices of the query, key, and value respectively; It is to perform square root scaling on the feature dimension d.
[0010] In addition, an optional technical solution is that the video rich information further includes an emotion disk generated based on the emotion category of the video to be processed and the probability distribution of the emotion category. The emotion disk represents the emotion probability with the sector area. The sector angles corresponding to various emotion categories are the same, and the radius of the sector is proportional to the probability of the corresponding emotion category. Positive emotions and negative emotions are respectively distributed on the left and right sides of the emotion disk, where negative emotions are on the left and positive emotions are on the right.
[0011] In addition, an optional technical solution is that the display interface of the sign language broadcast video adopts a picture-in-picture layout including a main screen and a sub-window. Among them, the main screen is used to display the sign language virtual human and the video rich information, and the sub-window is used to play the original program content.
[0012] In addition, an optional technical solution is that the method for driving the expression of a virtual image based on the expression parameters includes matching the expression parameters with the preset facial key point data to establish a mapping relationship; the facial key point data is obtained through the Mediapipe algorithm; According to the mapping relationship, converting the facial key point data into drive parameters recognizable by a preset virtual image facial model, and updating the facial expression of the virtual image in real time based on the drive parameters through a programming interface.
[0013] In addition, an optional technical solution is that when performing multi-modal emotion recognition on the text content, non-speech information, and video information, the text content further includes the background knowledge of the video to be processed.
[0014] On the other hand, the present invention also provides a sign language broadcast video generation system, which performs image generation by using the sign language broadcast video generation method as described above; the system includes: An information acquisition unit for acquiring the audio information and video information of the video to be processed; wherein, the audio information includes speech information and non-speech information; An emotion recognition unit for performing speech recognition on the speech information to obtain the text content corresponding to the speech information; performing multi-modal emotion recognition on the text content, non-speech information, and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category; A video generation unit is configured to determine expression parameters, lip movement parameters, and limb movement parameters based on the sign language text of the video to be processed, the emotion category, and the probability distribution of the emotion category; drive the expression, lip movement, and limb movement of a virtual avatar based on the expression parameters, the lip movement parameters, and the limb movement parameters respectively, and generate a sign language broadcast video after synchronizing with pre-acquired rich video information; wherein, the sign language text is obtained by performing semantic summary extraction on the text content; the rich video information includes synchronized beat signals and subtitles; the beat signals are obtained by performing beat extraction on the non-speech information; and the subtitles are generated according to the text content.
[0015] The above-mentioned sign language broadcast video generation method, system, and storage medium obtain the audio information and video information of the video to be processed; perform speech recognition on the speech information to obtain the text content corresponding to the speech information; perform multi-modal emotion recognition on the text content, non-speech information, and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category; determine expression parameters, lip movement parameters, and limb movement parameters based on the emotion category of the video to be processed, the probability distribution of the emotion category, and the sign language text; drive the expression, lip movement, and limb movement of a virtual avatar based on the expression parameters, the lip movement parameters, and the limb movement parameters respectively, and generate a sign language broadcast video after synchronizing with pre-acquired rich video information. By integrating speech information, non-speech information, and video information, the present invention realizes the comprehensive understanding and emotion recognition of video content, and then generates a sign language broadcast video based on the emotion recognition result. The virtual avatar in the sign language broadcast video can not only perform sign language expression, but also adjust its expression and limb movement according to the emotion recognition result, making the emotion conveyance more delicate and real. At the same time, the present invention synchronously generates rich information such as beat signals and subtitles with the sign language broadcast video, enriching the video content and improving the integrity of information. The present invention significantly enhances the understanding and viewing experience of the program information by hearing-impaired persons, enabling them to obtain the content in the video more comprehensively and accurately.
[0016] To achieve the above and related purposes, one or more aspects of the present invention include features that will be described in detail later. The following description and the accompanying drawings illustrate certain exemplary aspects of the present invention in detail. However, these aspects merely indicate some of the various ways in which the principles of the present invention can be used. In addition, the present invention is intended to include all these aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] By referring to the following description in conjunction with the accompanying drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become more apparent and easier to understand. In the drawings: Figure 1 It is a flowchart of a sign language broadcast video generation method according to an embodiment of the present invention; Figure 2 Schematic diagram of the principle of a sign language broadcast video generation method according to an embodiment of the present invention; Figure 3 Schematic diagram of the principle of emotion recognition of a sign language broadcast video generation method according to an embodiment of the present invention; Figure 4-1 Schematic diagram of the display interface of a sign language broadcast video according to an embodiment of the present invention Figure 1 ; Figure 4-2 Schematic diagram of the display interface of a sign language broadcast video according to an embodiment of the present invention Figure 2 ; Figure 5 Module schematic diagram of a sign language broadcast video generation system provided by an embodiment of the present invention; Figure 6 Internal structure schematic diagram of an electronic device for implementing a sign language broadcast video generation method provided by an embodiment of the present invention.
[0018] In all the drawings, the same reference numerals indicate similar or corresponding features or functions. Detailed implementation manners
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0020] The technical solutions in the embodiments of the present application will be clearly and elaborately described below with reference to the accompanying drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, "and / or" in the text is only an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations.
[0021] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. Additionally, in the description of the embodiments of the present application, "a plurality" means two or more than two.
[0022] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc., which appear in different places in this specification, are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0023] To describe in detail the sign language broadcast video generation method, system, and storage medium of the present invention, the following will describe in detail the specific embodiments of the present invention with reference to the accompanying drawings.
[0024] Figure 1 The flowchart of the sign language broadcast video generation method according to an embodiment of the present invention is shown.
[0025] As Figure 1 shown, the sign language broadcast video generation method provided in this embodiment is applied to an electronic device and mainly includes the following steps: S110: Obtain the audio information and video information of the video to be processed; wherein, the audio information includes speech information and non-speech information; S120: Perform speech recognition on the speech information to obtain the text content corresponding to the speech information; perform multi-modal emotion recognition on the text content, non-speech information, and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category. S130: Determine the expression parameters, mouth shape parameters, and limb movement parameters based on the sign language script, emotion category, and probability distribution of the emotion category of the video to be processed; wherein, the sign language script is obtained by extracting the semantic summary of the text content. S140: Drive the expression, mouth shape, and limb movements of the virtual image respectively based on the expression parameters, mouth shape parameters, and limb movement parameters, and generate a sign language broadcast video after synchronizing with the pre-obtained rich video information; wherein, the rich video information includes synchronized beat signals and subtitles; the beat signals are obtained by extracting the beats from the non-speech information; the subtitles are generated according to the text content.
[0026] Figure 2 The schematic diagram of the principle of the sign language broadcast video generation method according to an embodiment of the present invention; as Figure 2As shown in the figure, in order to provide a sign language translation solution with rich information, accurate expression and vividness for the hearing-impaired, the present invention is realized through the fusion and processing of multi-modal information, combined with technologies such as emotion recognition, sign language copywriting generation and intelligent driving of virtual digital humans. Specifically, in the first stage, the video program information to be processed is input, and the information may include audio information composed of non-speech information and speech information, video information composed of picture information, and background knowledge for providing auxiliary emotion recognition and other processing. Then, the audio information is processed, and the text content is generated through speech recognition for subsequent semantic understanding, sign language copywriting generation and subtitles in the picture; at the same time, the beat extraction of non-speech information such as background music is output to the metronome as a part of the rich information. In addition, the text information in the picture is recognized and extracted as a part of the rich information to assist information expression. In the second stage, multi-modal emotion recognition of the speech information, non-speech information and video information in the video is carried out, and a multi-modal feature fusion method based on a hierarchical sequence is adopted, and a preset classification network is used to predict the emotion category and its probability distribution, and the emotion category (such as happy, sad, etc.) and the probability value of each emotion category are output. In the third stage, the semantic understanding of the text generated by speech recognition is carried out, and the content summary is extracted to meet the speed requirements of sign language broadcasting; the rare words in the summary are replaced with the words in the "Common Word List of National General Sign Language" to ensure the accuracy and universality of sign language expression; and according to the expression habits of sign language, the word order of the text is adjusted to form a sign language broadcasting copywriting. In the fourth stage, according to the sign language broadcasting copywriting, the expression parameters, mouth shape parameters and limb movement parameters of the corresponding virtual digital human are determined; combined with the emotion category and probability distribution according to the emotion recognition result, the expression, mouth shape and limb movement parameters of the digital human are further refined and adjusted to make the actions and expressions of the digital human more in line with the content and emotion of sign language broadcasting.
[0027] S110: Obtain the audio information and video information of the video to be processed; wherein, the audio information includes speech information and non-speech information; the video information is the picture information.
[0028] Specifically, the audio information is composed of non-speech information (such as music, environmental sounds, noise, etc.) and speech information, and the video information is the picture information (such as video pictures, text, etc.). The video to be processed can be various types of video content such as movies, TV dramas, variety shows, news reports, online videos, teaching videos, advertising videos, sports event videos, documentaries, etc., and no specific limitation is made here.
[0029] S120: Perform speech recognition on the speech information to obtain the text content corresponding to the speech information; perform multi-modal emotion recognition on the text content, non-speech information, background knowledge information and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category.
[0030] In a specific embodiment, when performing multi-modal emotion recognition on the text content, non-speech information, and video information, the text content further includes the background knowledge of the video to be processed. The background knowledge may include plot introduction, character relationships, cultural background, historical background, shooting location, director or production team style, audience feedback, etc. For example, if the video to be processed is a movie, the background knowledge may include the plot introduction of the movie, the character relationships of the main characters, the cultural background reflected by the movie, etc. These information can help the system better understand the emotional content in the video, thereby improving the accuracy and depth of emotion recognition.
[0031] In the specific implementation process, the text content corresponding to the speech information obtained by performing speech recognition on the speech information can be implemented by the whishper algorithm.
[0032] The method for performing multi-modal emotion recognition on the text content, non-speech information, and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category includes: S121, respectively extracting features from the text content, non-speech information, and video information to obtain audio features, video features, and text features; S122, respectively performing dynamic alignment and cross-modal complementary information extraction on the audio features through audio-video cross-attention and audio-text cross-attention, using the Hadamard product for cross-modal feature fusion, and performing normalization processing to obtain new audio features; respectively performing dynamic alignment and cross-modal complementary information extraction on the video features through video-audio cross-attention and video-text cross-attention, using the Hadamard product for cross-modal feature fusion, and performing normalization processing to obtain new video features; S123, respectively performing dynamic alignment and cross-modal complementary information extraction on the text features, the new audio features, and the new video features through text-audio cross-attention and text-video cross-attention, using the Hadamard product for cross-modal feature fusion, and performing normalization processing to obtain new text features; S124, splicing the new text features, the new audio features, and the new video features to obtain multi-modal features; S125, inputting the multi-modal features into a preset classification network to obtain the emotion category of the video to be processed and the probability distribution of the emotion category.
[0033] Figure 3 It is a schematic diagram of the principle of emotion recognition for the sign language broadcast video generation method according to an embodiment of the present invention. As Figure 3As shown in the figure, the process of the multi-modal feature fusion emotion recognition algorithm based on the hierarchical sequence is as follows: First, input the video program information, including audio information and picture information. Then, feature extraction is carried out: extract the text feature sequence from the text content; extract the acoustic features from the non-speech information to obtain the speech feature sequence; extract the visual features (such as human expressions, actions, etc.) from the picture information to obtain the visual feature sequence. Next, the audio features are respectively subjected to audio-video cross-attention and audio-text cross-attention for dynamic alignment and cross-modal complementary information extraction, and the Hadamard product is used for cross-modal feature fusion and normalized processing; new audio features are obtained; the video features are respectively subjected to video-audio cross-attention and video-text cross-attention for dynamic alignment and cross-modal complementary information extraction, and the Hadamard product is used for cross-modal feature fusion and normalized processing; new video features are obtained; the text features, the new audio features and the new video features are respectively subjected to text-audio cross-attention and text-video cross-attention for dynamic alignment and cross-modal complementary information extraction, and the Hadamard product is used for cross-modal feature fusion and normalized processing; new text features are obtained; the new text features, the new audio features and the new video features are concatenated to obtain the multi-modal fusion features. And use a classification network (such as a linear layer) to process the fusion features and predict the corresponding emotion categories and probabilities. Finally, output the emotion category (such as happiness, sadness, etc.) expressed by the video program information and its corresponding probability, and the emotion disk shows various emotions with different sector areas according to these probabilities to assist the hearing-impaired to better understand the program emotion. The present invention can effectively fuse multi-modal information, improve the accuracy and interpretability of emotion recognition, and provide richer information for the emotion drive of digital humans.
[0034] The following takes the process of respectively performing dynamic alignment and cross-modal complementary information extraction on the audio features through audio-video cross-attention and audio-text cross-attention, using the Hadamard product for cross-modal feature fusion, and obtaining new audio features through normalization as an example for specific description.
[0035] The cross-modal feature fusion of the audio-text cross-attention feature and the audio-video cross-attention feature using the Hadamard product can be realized through the following formula:
[0036] Among them, is the fusion result; is the cross-modal correlation weight of the audio-text cross-attention feature; is the cross-modal correlation weight of the audio-video cross-attention feature; is the Hadamard product (element-wise multiplication), which is used to fuse the feature information of different modalities and enhance the complementarity of features.
[0037] The cross-modal association weights of the audio-text cross-attention features and the cross-modal association weights of the audio-video cross-attention features are obtained through the following formulas respectively:
[0038]
[0039] where 、 、 are audio features, text features, and video features respectively, and W is a learnable parameter. Among them W Q 、 W K 、 W V are the weight matrices of Query, Key, and Value respectively. That is, for both attentions, a is used as Q, and t or v is used as K and V; softmax is used to normalize the attention scores into a probability distribution to highlight important feature correspondence relationships. is the square root scaling of the feature dimension d to stabilize the gradient and attention distribution.
[0040] After cross-modal feature fusion using the Hadamard product, the fused features are normalized;
[0041] where LayerNorm represents layer normalization, which is used to standardize the features and stabilize the training process. x represents a hyperparameter that is used to control the weight of the original feature in the residual connection to balance the old and new features.
[0042] Finally, the features of the three modalities are fused by concatenation through cross-progressive. Specifically, cross-progressive first fuses the video features and text features, and the audio features and text features respectively to obtain the new audio feature and the new video feature , and then fuses them with the text features to obtain the new text feature ;
[0043] where represents the feature concatenation operation, which is used to connect the features of different modalities together to form the fused feature .
[0044] Finally, a classification network uses a linear layer and an activation function to predict the sentiment category and probability;
[0045]
[0046] Among them, hi ′ is the feature representation of a certain time step or sample in ; the feature representation of a certain time step or sample in W 1 , b 1 are the weight and bias parameters of the first-layer linear layer; RELU is the activation function, introducing non-linearity to enhance the expression ability of the model. W 2 , b 2 are the weight and bias parameters of the second-layer linear layer, used to map the features to the emotion category space. C i are the output probabilities of each emotion, normalized by softmax so that the sum of all probabilities is 1, and output to the emotion disk.
[0047] Determine the final emotion category according to the emotion probability;
[0048] Among them, Argmax means taking the emotion category with the highest probability as the final prediction result; is the emotion label predicted for the sentence u i , used to determine the expression parameters, mouth shape parameters and body movement parameters subsequently.
[0049] This step-by-step process promotes the selective learning of relevant information by different features, effectively captures the unique and complementary information provided by each modality, and improves the overall performance of the emotion recognition system.
[0050] S130: Determine the expression parameters, mouth shape parameters and body movement parameters based on the sign language text, the emotion category and the probability distribution of the emotion category of the video to be processed; among them, the sign language text is obtained by extracting the semantic summary of the text content.
[0051] The method for obtaining the sign language text by extracting the semantic summary of the text content includes S131, performing semantic understanding and summary extraction on the text content to obtain the sign language text content matching the sign language broadcast speed; S132, performing OOV vocabulary replacement and word order adjustment on the obtained sign language text content matching the sign language broadcast speed to obtain the sign language text.
[0052] S140: Drive the expressions, lip movements, and body movements of the virtual avatar based on the expression parameters, lip movement parameters, and body movement parameters respectively, and generate a sign language broadcast video after synchronizing with the pre-acquired rich video information; wherein, the rich video information includes synchronized beat signals and subtitles; the beat signals are obtained by extracting beats from the non-speech information; the subtitles are generated according to the text content.
[0053] Specifically, first, the correspondence between emotions and action expressions needs to be defined, and a mapping table between emotion categories and body movement and expression parameters should be established. For example, the emotion of "happiness" may correspond to body movements such as the virtual character spreading its arms and jumping, while the facial expression shows a smile and narrowed eyes. This mapping table can be based on the observation and summary of natural expressions and actions of humans under different emotions, or it can refer to the research results on emotion expression in related fields such as psychology. For each emotion category, the corresponding expression parameters (such as the degree of upward curvature of the mouth corners, the shape of the eyebrows, the degree of eye opening and closing, etc.) and body movement parameters (such as the amplitude of arm swing, the angle of body tilt, leg movements, etc.) are defined in detail to ensure that the virtual character can accurately represent the typical characteristics under that emotion. Then, a machine learning model is used for mapping learning. A large amount of human expression and action data with emotion annotations is collected, including facial expression images, body movement videos, or motion capture data, etc. Generative models in deep learning, such as generative adversarial networks (GANs) or variational autoencoders (VAEs), are used to learn the complex mapping relationship between emotions and corresponding action expressions. The model can learn the distribution characteristics of human expressions and actions under different emotions, so that it can generate the corresponding virtual character action and expression parameters according to the given emotion category. Through the trained model, when an emotion category is input, the model can output the expression and body movement parameters of the virtual character that match it, realizing an automated mapping process. For the expression control of the virtual character, a driving method based on facial expression parameters can be adopted. The above-obtained expression parameters are applied to the facial model of the virtual character, and different emotion expressions are realized by adjusting the shapes and positions of various key parts of the face (such as eyes, eyebrows, mouth, etc.). Using facial animation techniques, such as bone-based facial animation or methods that directly drive the change of facial vertex positions, the facial expression of the virtual character is changed in real time according to the expression parameters. For example, actions such as smiling and opening the mouth are achieved by controlling the rotation and scaling of the mouth bones, and different emotions are expressed by adjusting the position and angle of the eyebrow bones. To make the expression more natural and smooth, an expression transition technique can be introduced to perform smooth transition processing between different emotion expressions, avoiding sudden changes in expressions and making the expression changes of the virtual character more in line with human visual habits. Finally, body movement control based on motion capture data or procedurally generated actions. If there is a motion capture device, the body movement data of humans under different emotions can be collected first and stored as an action library. When it is necessary to drive the body movements of the virtual character according to the emotion category, the action data that best matches the emotion is retrieved from the action library and then applied to the bone system of the virtual character to drive the virtual character to make corresponding body movements; or rather, body movements are generated through a procedural method. According to the emotion category and body movement parameters, using character animation techniques in computer graphics, such as inverse kinematics (IK) algorithms, physics-based animation simulation, etc., the body movements of the virtual character are generated in real time.For example, for the jumping motion under the emotion of "excitement", by setting the bone joint constraints and dynamic parameters of the virtual avatar and using the IK algorithm, a reasonable limb posture and motion trajectory can be calculated to enable the virtual avatar to complete a natural jumping motion.
[0054] It should be noted that the virtual avatar can be a virtual pet, an anime character, a robot image, a humanized item, or a digital human. In this embodiment, the virtual avatar is a digital human (i.e., a sign language virtual human). The digital human can be designed according to the normal human proportion. The beat signal is obtained by extracting the beat from the non-speech information. The beat extraction can be implemented through a GRU network. Among them, the metronome can be a red light flashing according to the beat.
[0055] In the specific implementation process, the video rich information further includes an emotion disk generated based on the emotion category of the video to be processed and the probability distribution of the emotion category. The emotion disk represents the emotion probability with the sector area. Positive emotions and negative emotions are respectively distributed on the left and right sides of the emotion disk. The radius represents the probability size. That is to say, the sector angles corresponding to various emotion categories are the same, and the radius of the sector is proportional to the probability of the corresponding emotion category. Positive emotions and negative emotions are respectively distributed on the left and right sides of the emotion disk, where the left side is negative emotions and the right side is positive emotions. Specifically, the emotion disk can display multiple emotions through different sector areas, divided into two categories: positive and negative. Positive emotions such as happiness (red), excitement (magenta), and calmness (yellow) are located on the right side, and negative emotions such as sadness (gray), fear (blue), and anger (green) are located on the left side. The sector radius represents the probability of the emotion occurrence. The larger the radius, the larger the area, and the stronger the emotion. For example, when the area of the red region is the largest, it indicates that the emotion of "happiness" dominates, and other emotions may also be accompanied. And the sector with a small area represents a low probability or absence of the corresponding emotion. This design intuitively presents the emotion distribution of the video program and assists the hearing-impaired to understand the program emotion.
[0056] Driven by gestures, according to the sign language text and emotional information, control the gesture actions of the digital human to accurately express the sign language content; through facial driving, use algorithms such as Mediapipe to detect 468 key points on the face of a human actor, match the determined expression parameters with them, and drive the facial expressions of the digital human in real time to make its expressions rich and conform to the emotions and sign language content; through lip movement driving, drive the lip movement of the digital human, a virtual image, according to the needs of emotions and sign language broadcasting; through body driving, drive the body movements of the digital human according to the needs of emotions and sign language broadcasting to enhance the vividness and accuracy of expression. In the fifth stage, integrate rich information content such as subtitles, metronomes, and emotion disks with the digital human video, display the digital human sign language broadcasting video in the display interface of the sign language broadcasting video, and at the same time open a pip window in the blank area of the interface to play the original video program information, which is convenient for hearing-impaired people to better watch and understand the program content.
[0057] In a specific embodiment, the method for determining expression parameters based on the sign language text, the emotion category, and the probability distribution of the emotion category of the video to be processed includes: S141, determining expression parameters based on the emotion category and sign language text of the video to be processed; and adjusting the expression parameters according to the probability distribution of the emotion category; S142, matching the expression parameters with preset facial key point data to establish a mapping relationship; the facial key point data is obtained through the Mediapipe algorithm; wherein, the Mediapipe algorithm can identify 468 key points on the human face. S143, according to the mapping relationship, convert the facial key point data into drive parameters recognizable by a preset virtual image facial model, and update the facial expressions of the virtual image in real time based on the drive parameters through a programming interface.
[0058] In addition, an optional technical solution is that the display interface of the sign language broadcasting video adopts a pip layout including a main screen and a sub-window. Figure 4-1 and Figure 4-2 The display interface of the sign language broadcasting video is described; wherein, Figure 4-1 is a schematic diagram of the display interface of the sign language broadcasting video according to an embodiment of the present invention Figure 1 ; Figure 4-2 is a schematic diagram of the display interface of the sign language broadcasting video according to an embodiment of the present invention Figure 2 . Figure 4-1 is the horizontal version effect of the display interface of the sign language broadcasting video; Figure 4-2It is the vertical effect of the display interface of the sign language broadcast video. The main screen is used to display the sign language virtual human and video rich information, and the sub-window is used to play the original program content. Currently, the sign language translation programs usually adopt a relatively small screen ratio, resulting in the difficulty of clearly presenting the expressions and action details of the sign language interpreters, which affects the integrity and accuracy of information transmission. In the specific implementation process, in order to convey more information, the head and hands of the digital human can be slightly larger to facilitate seeing the gestures, lip shapes, and expressions clearly.
[0059] The sign language broadcast video generation method of the present invention integrates voice information, non-voice information, and video information to achieve a comprehensive understanding and emotion recognition of the video content, and then generates a sign language broadcast video based on the emotion recognition result. The virtual image in the sign language broadcast video can not only perform sign language expression, but also adjust the expressions and body movements according to the emotion recognition result, making the emotion transmission more delicate and real. At the same time, the present invention synchronously generates rich information such as beat signals and subtitles with the sign language broadcast video, enriching the video content and improving the integrity of information. The present invention significantly enhances the understanding and viewing experience of the hearing-impaired people for the program information, enabling them to obtain the content in the video more comprehensively and accurately.
[0060] As Figure 5 shown, the present invention provides a sign language broadcast video generation system, which uses the sign language broadcast video generation method as described above to generate a sign language broadcast video. According to the implemented functions, the sign language broadcast video generation system 500 may include an information acquisition unit 510, an emotion recognition unit 520, and a video generation unit 530. The unit of the present invention can also be called a module, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0061] In this embodiment, the functions of each module / unit are as follows: Using the sign language broadcast video generation method as described above to generate a video; including: The information acquisition unit 510 is used to acquire the audio information and video information of the video to be processed; wherein, the audio information includes voice information and non-voice information; The emotion recognition unit 520 is used to perform speech recognition on the voice information to obtain the text content corresponding to the voice information; perform multi-modal emotion recognition on the text content, non-voice information, and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category; A video generation unit 530 is configured to determine expression parameters, lip movement parameters, and limb movement parameters based on the sign language text of the video to be processed, the emotion category, and the probability distribution of the emotion category; drive the expression, lip movement, and limb movement of the virtual avatar respectively based on the expression parameters, the lip movement parameters, and the limb movement parameters, and generate a sign language broadcast video after synchronizing with pre-acquired rich video information; wherein, the sign language text is obtained by extracting a semantic summary of the text content; the rich video information includes synchronized beat signals and subtitles; the beat signals are obtained by extracting beats from the non-speech information; and the subtitles are generated according to the text content.
[0062] The sign language broadcast video generation system of the present invention integrates speech information, non-speech information, and video information, realizes a comprehensive understanding and emotion recognition of video content, and then generates a sign language broadcast video based on the emotion recognition result. The virtual avatar in the sign language broadcast video can not only perform sign language expression, but also adjust its expression and limb movements according to the emotion recognition result, making the emotion transmission more delicate and real. At the same time, the present invention synchronously generates rich information such as beat signals and subtitles with the sign language broadcast video, enriches the video content, and improves the integrity of information. The present invention significantly enhances the understanding and viewing experience of the hearing-impaired for program information, enabling them to obtain the content in the video more comprehensively and accurately.
[0063] For more specific implementation manners of the above sign language broadcast video generation system, reference can be made to the description of the embodiments of the sign language broadcast video generation method mentioned above, and details will not be repeated here.
[0064] As Figure 6 shown, the present invention also correspondingly provides an electronic device 1 for the sign language broadcast video generation method.
[0065] The electronic device 1 may include a processor 10, a memory 11, and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 10, such as a sign language broadcast video generation program 12. The memory 11 may also include both an internal storage unit of the sign language broadcast video generation system and an external storage device. The memory 11 can be used not only to store installed application software and various types of data, such as the code of the sign language broadcast video generation program, etc., but also to temporarily store data that has been output or will be output.
[0066] Among them, the memory 11 at least includes one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code of the sign language broadcast video generation program, etc., but also to temporarily store data that has been output or will be output.
[0067] In some embodiments, the processor 10 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules (such as the sign language broadcast video generation program, etc.) stored in the memory 11, and calling the data stored in the memory 11, to perform various functions of the electronic device 1 and process data.
[0068] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.
[0069] Figure 6 Only the electronic device with components is shown. Those skilled in the art can understand that, Figure 6The shown structure does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0070] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management system, so as to implement functions such as charge management, discharge management, and power consumption management through the power management system. The power source may also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0071] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.
[0072] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.
[0073] It should be understood that the embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0074] The sign language broadcast video generation program 12 stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve: obtaining the audio information and video information of the video to be processed; wherein, the audio information includes speech information and non-speech information; performing speech recognition on the speech information to obtain the text content corresponding to the speech information; performing multi-modal emotion recognition on the text content, non-speech information and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category; determining the expression parameters, mouth shape parameters and limb movement parameters based on the sign language text, the emotion category and the probability distribution of the emotion category of the video to be processed; wherein, the sign language text is obtained by performing semantic summary extraction on the text content; driving the expression, mouth shape and limb movements of the virtual image respectively based on the expression parameters, the mouth shape parameters and the limb movement parameters, and generating a sign language broadcast video after synchronizing with the pre-obtained rich video information; wherein, the rich video information includes synchronized beat signals and subtitles; the beat signals are obtained by performing beat extraction on the non-speech information; the subtitles are generated according to the text content.
[0075] Specifically, for the specific implementation method of the above instructions by the processor 10, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here. Further, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or system, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory) that can carry the computer program code.
[0076] An embodiment of the present invention also provides a computer-readable storage medium, which can be non-volatile or volatile. The storage medium stores a computer program, and when the computer program is executed by a processor, it realizes: obtaining audio information and video information of a video to be processed; wherein, the audio information includes speech information and non-speech information; performing speech recognition on the speech information to obtain the text content corresponding to the speech information; performing multi-modal emotion recognition on the text content, non-speech information and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category; determining expression parameters, lip-sync parameters and limb movement parameters based on the sign language text, the emotion category and the probability distribution of the emotion category of the video to be processed; wherein, the sign language text is obtained by performing semantic summary extraction on the text content; driving the expression, lip-sync and limb movements of a virtual avatar respectively based on the expression parameters, the lip-sync parameters and the limb movement parameters, and generating a sign language broadcast video after synchronizing with pre-obtained rich video information; wherein, the rich video information includes synchronized beat signals and subtitles; the beat signals are obtained by performing beat extraction on the non-speech information; the subtitles are generated according to the text content.
[0077] Specifically, the specific implementation method when the computer program is executed by the processor can refer to the description of the relevant steps in the sign language broadcast video generation method in the embodiment, which will not be elaborated here.
[0078] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0079] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0080] In addition, each functional module in various embodiments of the present invention can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional modules.
[0081] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0082] Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.
[0083] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or systems stated in the system claims can also be implemented by one unit or system through software or hardware.
[0084] However, those skilled in the art should understand that various improvements can be made to the sign language broadcast video generation method and the sign language broadcast video generation system proposed for the present invention without departing from the content of the present invention. Therefore, the protection scope of the present invention should be determined by the content of the appended claims.
Claims
1. A method for generating a sign language broadcast video, applied to an electronic device, characterized in that: include: Acquire audio information and video information of the video to be processed; wherein the audio information includes voice information and non-voice information; Performing speech recognition on the speech information to obtain text content corresponding to the speech information; performing multimodal emotion recognition on the text content, non-speech information and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category; Determining expression parameters, lip shape parameters and body movement parameters based on the sign language document of the video to be processed, the emotion category and the probability distribution of the emotion category; wherein the sign language document is obtained by extracting the semantic summary of the text content; Based on the expression parameters, the lip shape parameters and the body movement parameters, the expression, lip shape and body movement of the virtual image are driven respectively, and a sign language broadcast video is generated after synchronization with the pre-acquired video rich information; wherein the video rich information includes a synchronized beat signal and subtitles; the beat signal is obtained by extracting the beat of the non-voice information; and the subtitles are generated according to the text content.
2. The method for generating a sign language broadcast video according to claim 1, characterized in that: The method of performing multimodal emotion recognition on the text content, non-voice information and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category includes: Extracting features of the text content, non-voice information and video information respectively to obtain audio features, video features and text features; The audio features are subjected to dynamic alignment and modal complementary information extraction through audio-video cross attention and audio-text cross attention, cross-modal feature fusion is performed using Hadamard product, and normalization is performed to obtain new audio features; The video features are respectively subjected to video-audio cross attention and video-text cross attention to perform dynamic alignment and modal complementary information extraction, cross-modal feature fusion is performed using Hadamard product, and normalization is performed to obtain new video features; The text features, the new audio features and the new video features are respectively subjected to text-audio cross attention and text-video cross attention for dynamic alignment and modal complementary information extraction, cross-modal feature fusion is performed using Hadamard product, and normalization is performed to obtain new text features; The new text feature, the new audio feature and the new video feature are concatenated to obtain a multimodal feature; The multimodal features are input into a preset classification network to obtain the emotion category of the video to be processed and the probability distribution of the emotion category.
3. The method for generating a sign language broadcast video according to claim 2, characterized in that: Cross-modal feature fusion is performed using the Hadamard product, which is implemented using the following formula: , in, is the fusion result; is the cross-modal correlation weight of the audio-text cross-attention feature; is the cross-modal correlation weight of the audio-video cross-attention feature; is the Hadamard product; among them, , , in, , , They are audio features, text features, and video features, and W is a learnable parameter; W Q , W K , W V are the weight matrices for query, key, and value respectively; It is the square root scaling of the feature dimension d.
4. The method for generating a sign language broadcast video according to claim 1, characterized in that: The video rich information also includes an emotion disk generated based on the emotion category of the video to be processed and the probability distribution of the emotion category. The emotion disk represents the emotion probability with a fan-shaped area. The fan-shaped angles corresponding to various emotion categories are the same, and the radius of the fan is proportional to the probability of the corresponding emotion category. Positive emotions and negative emotions are distributed on the left and right sides of the emotion disk respectively, wherein the left side is negative emotions and the right side is positive emotions.
5. The method for generating a sign language broadcast video according to claim 1, characterized in that: The display interface of the sign language broadcast video adopts a picture-in-picture layout including a main screen and a sub-window, wherein the main screen is used to display the sign language virtual person and video rich information, and the sub-window is used to play the original program content.
6. The method for generating a sign language broadcast video according to claim 1, characterized in that: The method for driving the expression of the virtual image based on the expression parameter includes: Matching the expression parameters with preset facial key point data to establish a mapping relationship; The facial key point data is obtained through the Mediapipe algorithm; According to the mapping relationship, the facial key point data is converted into driving parameters recognizable by a preset virtual image facial model, and the facial expression of the virtual image is updated in real time based on the driving parameters through a programming interface.
7. The method for generating a sign language broadcast video according to claim 1, characterized in that: When multimodal emotion recognition is performed on the text content, non-voice information and video information, the text content also includes background knowledge of the video to be processed.
8. A sign language broadcast video generation system, using the sign language broadcast video generation method according to any one of claims 1 to 7 to generate images; comprising: An information acquisition unit, used to acquire audio information and video information of the video to be processed; wherein the audio information includes voice information and non-voice information; An emotion recognition unit is used to perform speech recognition on the speech information to obtain text content corresponding to the speech information; perform multimodal emotion recognition on the text content, non-speech information and video information to obtain the emotion category of the video to be processed and the probability distribution of the emotion category; A video generation unit is used to determine expression parameters, lip shape parameters and body movement parameters based on the sign language document of the video to be processed, the emotion category and the probability distribution of the emotion category; based on the expression parameters, the lip shape parameters and the body movement parameters, respectively drive the expression, lip shape and body movement of the virtual image, and generate a sign language broadcast video after synchronization with the pre-acquired video rich information; wherein the sign language document is obtained by extracting the semantic summary of the text content; the video rich information includes a synchronized beat signal and subtitles; the beat signal is obtained by extracting the beat of the non-voice information; and the subtitles are generated according to the text content.
9. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a sign language broadcast video generation program stored in the memory and executable on the processor. When the sign language broadcast video generation program is executed by the processor, the sign language broadcast video generation method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for generating a sign language broadcast video as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Sign language video generation method, sign language video translation method, sign language video customer service method and device and readable medium
CN113835522A
AI digital human interaction method, device and system based on emotion recognition
CN116560513A
Method for driving digital human body expression through voice
CN116880695A
Text association type short video multi-mode emotion recognition method and system
CN117636196A
Pacifying chat accompanying method based on multi-mode and multi-angle emotion perception
CN117725178A