Sign language generation method and device, electronic equipment, storage medium and product

By classifying and adjusting the emotional content of speech audio and combining it with the sign language expression style of hearing-impaired individuals to generate a sequence of sign language action images, the problem of insufficient accuracy and emotional content in existing sign language generation technologies has been solved, thereby improving the accuracy and emotional content of sign language generation and enhancing the comprehension ability of hearing-impaired individuals.

CN119418714BActive Publication Date: 2025-12-05IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411485954.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-12-05
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

Existing sign language generation methods have low accuracy and insufficient emotional expression, failing to capture the emotions and rhythm in the audio, making it difficult for hearing-impaired individuals to understand.

Method used

By classifying the speaker's voice audio for emotion, and combining the voice audio with the emotion feature sequence, the facial and hand movements in the speaker's action image frames are adjusted to generate a speaker's sign language action image sequence that conforms to the sign language expression style of hearing-impaired people. The sign language action generation learning is carried out using pre-recorded sign language action videos of hearing-impaired people.

Benefits of technology

It improves the accuracy and emotionality of sign language generation, making it easier for hearing-impaired people to understand the speaker's sign language gestures and facial expressions, thus enhancing communication effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418714B_ABST
    Figure CN119418714B_ABST
Patent Text Reader

Abstract

The application provides a sign language generation method and device, electronic equipment, a storage medium and a product. The method performs emotion classification on each audio frame in a speech audio of a speaker, and determines an emotion feature sequence corresponding to the speech audio. Based on the speech audio and the emotion feature sequence, the speaker's facial action and hand action in a picture frame of the speaker's action are adjusted, and a picture sequence of the speaker's sign language action corresponding to the speech audio is generated. By using the technical solution of the application, the speaker's facial action and hand action in the picture frame of the speaker's action can be adjusted by combining the speech audio with the emotion feature of the speech audio, so that the speaker's sign language action and facial expression have emotion features, and the emotion degree of sign language generation is improved. In addition, the style of the speaker's facial action and hand action in the picture sequence of the speaker's sign language action is the same as the sign language expression style of a hearing-impaired person, the accuracy of sign language generation is improved, and it is more convenient for the hearing-impaired person to understand.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a sign language generation method and device, electronic equipment, storage medium and product. BACKGROUND

[0002] With the development of current technology, there are more and more applications in translating the speech or text of a speaker into sign language actions to generate corresponding fluent and accurate sign language videos for the hearing-impaired to communicate. The existing sign language generation method is to construct a sign language action database by obtaining the sign language videos recorded by standard sign language teachers or sign language announcers, and to generate the sign language corresponding to the audio by learning the sign language actions in the sign language action database. However, the generated sign language may be different from the hearing-impaired person currently communicating, resulting in low accuracy of sign language generation, and the rhythm of the generated sign language is monotonous and stiff, which cannot reflect the emotion and rhythm in the audio, resulting in low emotion of sign language generation. SUMMARY

[0003] Based on the above needs, the present application proposes a sign language generation method, device, electronic equipment, storage medium and product, which can improve the accuracy and emotion of sign language generation.

[0004] To achieve the above object, the present application proposes the following technical scheme:

[0005] According to a first aspect of an embodiment of the present application, a sign language generation method is provided, comprising:

[0006] performing emotion classification on each audio frame in the speech audio of the speaker to determine the emotion feature sequence corresponding to the speech audio;

[0007] adjusting the speaker facial action and hand action in the speaker action picture frame based on the speech audio and the emotion feature sequence to generate a speaker sign language action picture sequence corresponding to the speech audio;

[0008] wherein the style of the speaker facial action and hand action in the speaker sign language action picture sequence is the same as the sign language expression style of the hearing-impaired person.

[0009] Optionally, the method further comprises:

[0010] performing face and hand mask processing on the speaker action picture frame to obtain a mask picture of the speaker, and performing face recognition on the speaker action picture frame to obtain the facial identity feature of the speaker;

[0011] adjusting, based on the speech audio and the sequence of emotion features, speaker facial actions and hand actions in a picture frame of speaker actions to generate a sequence of picture frames of speaker sign language actions corresponding to the speech audio, comprising:

[0012] generating, based on the speech audio, the sequence of emotion features, the mask picture of the speaker, and the facial identity features of the speaker, speaker facial actions and hand actions in the mask picture of the speaker using a predetermined sign language action generation algorithm to obtain the sequence of picture frames of speaker sign language actions corresponding to the speech audio;

[0013] The sign language action generation algorithm is determined by sign language action generation learning using a pre-recorded sign language action video of a hearing-impaired person.

[0014] Optionally, generating, based on the speech audio, the sequence of emotion features, the mask picture of the speaker, and the facial identity features of the speaker, speaker facial actions and hand actions in the mask picture of the speaker using a predetermined sign language action generation algorithm to obtain the sequence of picture frames of speaker sign language actions corresponding to the speech audio, comprising:

[0015] determining, based on the speech audio and the sequence of emotion features, a rhythm amplitude feature corresponding to the speech audio;

[0016] generating, based on the mask picture of the speaker, the facial identity features of the speaker, and the rhythm amplitude feature corresponding to the speech audio, speaker facial actions and hand actions in the mask picture of the speaker to obtain the sequence of picture frames of speaker sign language actions corresponding to the speech audio.

[0017] Optionally, generating, based on the mask picture of the speaker, the facial identity features of the speaker, and the rhythm amplitude feature corresponding to the speech audio, speaker facial actions and hand actions in the mask picture of the speaker to obtain the sequence of picture frames of speaker sign language actions corresponding to the speech audio, comprising:

[0018] determining, based on the rhythm amplitude feature corresponding to the speech audio, an action optical flow feature corresponding to the speech audio;

[0019] generating, based on the mask picture of the speaker, the facial identity features of the speaker, and the action optical flow feature corresponding to the speech audio, speaker facial actions and hand actions in the mask picture of the speaker to obtain the sequence of picture frames of speaker sign language actions corresponding to the speech audio.

[0020] Optionally, based on the speech audio, the sequence of emotional features, the mask picture of the speaker and the facial identity features of the speaker, a predetermined sign language action generation algorithm is used to generate the facial and hand actions of the speaker in the mask picture of the speaker, to obtain a sequence of sign language action pictures of the speaker corresponding to the speech audio, comprising:

[0021] The speech audio, the sequence of emotional features, the mask picture of the speaker and the facial identity features of the speaker are input into a pre-trained sign language action generation model to obtain a sequence of sign language action pictures of the speaker corresponding to the speech audio.

[0022] The sign language action generation model is obtained based on pre-recorded sign language action videos of hearing-impaired persons.

[0023] Optionally, the sign language action generation model comprises a rhythm amplitude feature encoder, an optical flow feature encoder and a sign language action generator.

[0024] The speech audio, the sequence of emotional features, the mask picture of the speaker and the facial identity features of the speaker are input into a pre-trained sign language action generation model to obtain a sequence of sign language action pictures of the speaker corresponding to the speech audio, comprising:

[0025] The speech audio and the sequence of emotional features are input into the rhythm amplitude feature encoder to obtain rhythm amplitude features corresponding to the speech audio.

[0026] The rhythm amplitude features corresponding to the speech audio are input into the optical flow feature encoder to obtain action optical flow features corresponding to the speech audio.

[0027] The mask picture of the speaker, the facial identity features of the speaker and the action optical flow features corresponding to the speech audio are input into the sign language action generator to obtain a sequence of sign language action pictures of the speaker corresponding to the speech audio.

[0028] Optionally, the training process of the sign language action generation model comprises:

[0029] Obtain sign language action videos corresponding to sample audios of hearing-impaired persons in different emotional states and sample emotional features corresponding to the sample audios.

[0030] Perform face and hand masking processing on video frames in the sign language action videos to obtain sample mask pictures of the hearing-impaired persons, and perform facial recognition on the video frames in the sign language action videos to obtain sample facial identity features of the hearing-impaired persons.

[0031] inputting the sample audio, the sample emotion feature, the sample mask picture of the hearing-impaired person and the sample facial identity feature of the hearing-impaired person into a pre-constructed sign language action generation model to obtain a predicted sign language action picture sequence corresponding to the sample audio;

[0032] performing model parameter adjustment on the sign language action generation model based on a difference between the predicted sign language action picture sequence corresponding to the sample audio and a sample sign language action picture sequence, wherein the sample sign language action picture sequence is obtained by frame extraction on a sign language action video corresponding to the sample audio.

[0033] Optionally, inputting the sample audio, the sample emotion feature, the sample mask picture of the hearing-impaired person and the sample facial identity feature of the hearing-impaired person into a pre-constructed sign language action generation model to obtain a predicted sign language action picture sequence corresponding to the sample audio, comprises:

[0034] inputting the sample audio and the sample emotion feature into a rhythm amplitude feature encoder to obtain a sample rhythm amplitude feature corresponding to the sample audio, and performing training on the rhythm amplitude feature encoder based on a difference between a labeled rhythm amplitude feature and the sample rhythm amplitude feature, wherein the labeled rhythm amplitude feature is a rhythm amplitude feature artificially labeled on the sample audio;

[0035] inputting the sample rhythm amplitude feature corresponding to the sample audio into an optical flow feature encoder to obtain a sample action optical flow feature corresponding to the sample audio;

[0036] inputting the sample action optical flow feature, the sample mask picture of the hearing-impaired person and the sample facial identity feature of the hearing-impaired person into a sign language action generator to obtain a predicted sign language action picture sequence corresponding to the sample audio;

[0037] performing joint training on the optical flow feature encoder and the sign language action generator based on a difference between the sample action optical flow feature and a labeled optical flow feature, and a difference between the predicted sign language action picture sequence and a sample sign language action picture sequence, wherein the labeled optical flow feature is obtained by optical flow feature extraction on a sign language action video corresponding to the sample audio.

[0038] According to a second aspect of the embodiment of the present application, a sign language generation device is provided, comprising:

[0039] an emotion classification module configured to perform emotion classification on each audio frame in the speech audio of the speaker to determine an emotion feature sequence corresponding to the speech audio;

[0040] a sign language generation module, configured to adjust a speaker facial action and a speaker hand action in a speaker action picture frame based on the speech audio and the sequence of emotional features, and generate a sequence of speaker sign language action pictures corresponding to the speech audio;

[0041] In the sequence of speaker sign language action pictures, the speaker facial action and the speaker hand action have the same style as a sign language expression style of a hearing-impaired person.

[0042] According to a third aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor;

[0043] The memory is connected with the processor, and is configured to store a program;

[0044] The processor is configured to realize the sign language generation method by running the program in the memory.

[0045] According to a fourth aspect of the embodiments of the present application, a storage medium is provided, and the storage medium stores a computer program. When the computer program is executed by a processor, the sign language generation method is realized.

[0046] According to a fifth aspect of the embodiments of the present application, a computer program product is provided, including computer program instructions. When the computer program instructions are executed by a processor, the processor realizes the sign language generation method.

[0047] The sign language generation method provided in the present application performs emotional classification on each audio frame in the speech audio of the speaker, determines a sequence of emotional features corresponding to the speech audio, adjusts a speaker facial action and a speaker hand action in a speaker action picture frame based on the speech audio and the sequence of emotional features, and generates a sequence of speaker sign language action pictures corresponding to the speech audio. In the sequence of speaker sign language action pictures, the speaker facial action and the speaker hand action have the same style as a sign language expression style of a hearing-impaired person. By using the technical solution of the present application, the speaker facial action and the speaker hand action in the speaker action picture frame can be adjusted by combining the speech audio and the sequence of emotional features of the speech audio, so that the sign language action and the facial expression of the speaker have emotional features, and the emotional degree of the sign language generation is improved. In addition, the speaker facial action and the speaker hand action in the sequence of speaker sign language action pictures have the same style as the sign language expression style of the hearing-impaired person, the accuracy of the sign language generation is improved, and the understanding of the hearing-impaired person is more convenient. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings in the following description only need to be used to explain the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on the provided drawings.

[0049] FIG. 1 A flowchart of a sign language generation method provided by an embodiment of the present application is shown in the figure.

[0050] FIG. 2 A flowchart of another sign language generation method provided by an embodiment of the present application is shown in the figure.

[0051] FIG. 3 A processing flowchart of determining a speaker sign language action picture sequence provided by an embodiment of the present application is shown in the figure.

[0052] FIG. 4 A processing flowchart of training a sign language action generation model provided by an embodiment of the present application is shown in the figure.

[0053] FIG. 5 A processing flowchart of determining a predicted sign language action picture sequence provided by an embodiment of the present application is shown in the figure.

[0054] FIG. 6 A structural diagram of a sign language generation device provided by an embodiment of the present application is shown in the figure.

[0055] FIG. 7 A structural diagram of an electronic device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0056] The technical solutions of the embodiments of the present application are suitable for the application scenarios of generating actions corresponding to audio, and are specifically used in the application scenarios of sign language action generation. By using the technical solutions of the embodiments of the present application, the speaker's facial action and hand action in the speaker action picture frame can be adjusted by combining the speech audio with the emotional features of the speech audio, so that the speaker's sign language action and facial expression have emotional features, and the emotional degree of sign language generation is improved. In addition, the style of the speaker's facial action and hand action in the speaker sign language action picture sequence is the same as the sign language expression style of the hearing-impaired person, the accuracy of sign language generation is improved, and it is more convenient for the hearing-impaired person to understand.

[0057] With the development of current technology, more and more applications have appeared to translate the speech or text of a speaker into sign language action and generate corresponding fluent and accurate sign language video for communication and exchange of the hearing-impaired.

[0058] The current mainstream scheme mainly has the following characteristics:

[0059] First, in the three-dimensional scheme, a neural network is often used to generate a three-dimensional human skeleton point or various feature points as a target, and finally a three-dimensional animation is rendered through an animation engine to serve as the final form of expression. Low-cost animation rendering models are difficult to accurately drive the rhythm, amplitude, and facial expression. The corresponding artificial cost, equipment cost, and time cost of high-realistic three-dimensional animation are extremely high.

[0060] Second, in the two-dimensional scheme, the generation of sign language actions in two-dimensional animation is mostly based on the recording of sign language videos by standard sign language teachers or sign language announcers to construct a sign language action database. The sign language actions in the sign language action database are learned to generate sign language corresponding to the audio. However, the generated sign language may be different from the sign language of the hearing-impaired person currently communicating, resulting in low accuracy of sign language generation. Moreover, the rhythm of the generated sign language is monotonous and stiff, and cannot reflect the emotion and rhythm in the audio, resulting in low emotional degree of sign language generation.

[0061] Therefore, the technical scheme of the present application can adjust the facial action and hand action of the speaker in the picture frame of the speaker action according to the speech audio and the emotional features of the speech audio, so that the sign language action and facial expression of the speaker have emotional features, and the style of the facial action and hand action of the speaker in the picture sequence of the speaker sign language action is the same as the sign language expression style of the hearing-impaired person, thereby solving the problems of low accuracy and emotional degree of sign language generation in the prior art.

[0062] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0063] Exemplary method

[0064] Referring to FIG. 1 The embodiments of the present application propose a sign language generation method. The method comprises:

[0065] S101, classifying the emotion of each audio frame in the speech audio of the speaker to determine the emotional feature sequence corresponding to the speech audio.

[0066] In order to realize normal communication between a speaker and a hearing-impaired person, it is necessary to generate sign language actions same as the meaning of the speech audio spoken by the speaker, so that the hearing-impaired person can understand the content expressed by the speaker according to the sign language actions same as the meaning of the speech audio. However, the speaker has emotional expression when speaking, and even if the content spoken is the same, different emotions may express different meanings. For example, if the speaker says "there is a cat", if it is said with a happy emotion, the content expressed contains love for the cat, and if it is said with a fear emotion, the content expressed contains fear of the cat. Therefore, in order to improve the expression accuracy of the corresponding sign language actions generated according to the speech audio of the speaker, the emotion in the speech audio spoken by the speaker also needs to be integrated into the sign language actions, so that the hearing-impaired person can understand the content expressed by the speaker more accurately.

[0067] Therefore, the embodiment needs to classify the emotion of each audio frame in the speech audio of the speaker, determine the emotion expressed by each audio frame in the speech audio, and obtain the emotion feature corresponding to each audio frame in the speech audio. Then, the emotion features corresponding to all audio frames are collected in the time sequence of the audio frames in the speech audio to obtain the emotion feature sequence corresponding to the speech audio of the speaker.

[0068] In the embodiment, the emotion of the audio frame is classified. First, the audio frame needs to be preprocessed, such as noise reduction processing, standardization processing, etc., to improve the accuracy of subsequent extraction of the sound feature corresponding to the audio frame. Then, the sound feature of the preprocessed audio frame is extracted, and the sound feature related to the emotion in the audio frame is extracted, such as pitch, speech rate, spectrum, mel spectrum coefficient (MFCC), linear predictive coding (LPC), etc. These features can reflect the emotional information in the audio frame. Finally, the sound feature corresponding to the audio frame is classified, and the emotion feature of the audio frame is determined, wherein the emotion feature of the audio frame can reflect the emotion category of the audio frame. In the embodiment, the emotion classification model can be used to classify the sound feature corresponding to the audio frame. The emotion classification model can be trained using pre-collected sample audio. The emotion features of each sample audio frame in the sample audio are manually labeled as emotion labels of the sample audio frames. Then, the sample audio frame is input into the emotion classification model to predict the predicted emotion feature of the sample audio frame. The difference between the predicted emotion feature of the sample audio frame and the emotion label is minimized to adjust the model parameters of the emotion classification model, thereby completing the training of the emotion classification model. In the embodiment, the emotion classification model preferably uses existing classification networks such as BiLSTM network, AlexNet network, wav2vec2 network, etc.

[0069] S102, based on the speech audio and the emotion feature sequence, adjusting the facial action and the hand action of the speaker in the speaker action picture frame to generate a speaker sign language action picture sequence corresponding to the speech audio.

[0070] In this embodiment, the speaker action picture frame is a front picture frame of the speaker, which can be obtained by frame extraction from the speech video of the speaker, or can be obtained by front shooting of the speaker using an image acquisition device (such as a camera, etc.).

[0071] After obtaining the emotion feature sequence corresponding to the speech audio through the above steps, the hand action of the speaker in the speaker action picture frame is adjusted according to the content expressed by the speech audio and the emotion feature sequence, so that the hand action of the speaker can express the content expressed by the speech audio and the emotion expressed by the emotion feature sequence. In addition, the facial action of the speaker in the speaker action picture frame is adjusted according to the content expressed by the speech audio and the emotion feature sequence, so that the facial action of the speaker can express the emotion in the speech audio. Finally, after adjusting the facial action and the hand action of the speaker in the speaker action picture frame, the speaker sign language action picture corresponding to each audio frame in the speech audio is obtained. All speaker sign language action pictures are collected according to the time sequence of each audio frame in the speech audio, and a speaker sign language action picture sequence corresponding to the speech audio is obtained. The speaker in the speaker sign language action picture sequence can not only show the sign language action with the same meaning as the speech audio, but also show the same emotion in the facial expression corresponding to the sign language action and the facial action of the speaker.

[0072] In addition, in this embodiment, the adjustment of the facial action and the hand action of the speaker in the speaker action picture frame according to the speech audio and the emotion feature sequence is determined by learning the sign language expression style and the facial expression style of the hearing-impaired person, so that the adjustment of the facial action and the hand action of the speaker in the speaker action picture frame conforms to the sign language expression style of the hearing-impaired person, that is, the facial action and the hand action of the speaker in the speaker sign language action picture sequence have the same style as the sign language expression style of the hearing-impaired person. When the hearing-impaired person watches the speaker sign language action picture sequence with the same sign language expression style as his own, it is more convenient for the hearing-impaired person to understand the content expressed by the speaker, and the understanding accuracy of the hearing-impaired person is improved, that is, the accuracy of the generated speaker sign language action picture sequence is higher.

[0073] Further, after the sequence of pictures of sign language actions of the speaker corresponding to the voice audio is generated, all pictures in the sequence of pictures of sign language actions of the speaker are merged into a video in sequence, so as to obtain the video of sign language actions of the speaker corresponding to the voice audio, so that the hearing-impaired person can communicate with the speaker.

[0074] As can be seen from the above, the method for generating sign language actions according to the embodiments of the present application classifies the emotion of each audio frame in the voice audio of the speaker, determines the sequence of emotion features corresponding to the voice audio, adjusts the facial action and the hand action of the speaker in the picture frame of actions of the speaker based on the voice audio and the sequence of emotion features, and generates the sequence of pictures of sign language actions of the speaker corresponding to the voice audio, wherein the style of the facial action and the hand action of the speaker in the sequence of pictures of sign language actions of the speaker is the same as the style of sign language expression of the hearing-impaired person. By using the technical solution of the embodiments, the facial action and the hand action of the speaker in the picture frame of actions of the speaker can be adjusted by combining the voice audio with the sequence of emotion features, so that the sign language actions and the facial expressions of the speaker have emotion features, and the emotion degree of the generated sign language is improved. In addition, the style of the facial action and the hand action of the speaker in the sequence of pictures of sign language actions of the speaker is the same as the style of sign language expression of the hearing-impaired person, so that the accuracy of the generated sign language is improved, and the hearing-impaired person can understand more conveniently.

[0075] As an optional implementation, the embodiments of the present application further provide a method for generating sign language.

[0076] Referring to FIG. 2 As shown in the figure, the method comprises:

[0077] S201, performing mask processing on the facial action and the hand action of the picture frame of actions of the speaker to obtain a mask picture of the speaker, and performing facial recognition on the picture frame of actions of the speaker to obtain facial identity features of the speaker.

[0078] The embodiments need to obtain the picture frame of actions of the speaker, and the manner of obtaining the picture frame of actions of the speaker has been specifically described in the above embodiments, which will not be described herein. Before adjusting the facial action and the hand action of the speaker in the picture frame of actions of the speaker, the embodiments need to perform mask processing on the facial action and the hand action of the speaker in the picture frame of actions of the speaker to obtain a mask picture of the speaker. The mask processing on the facial action and the hand action of the picture frame of actions of the speaker needs to first perform image analysis and detection on the picture frame of actions of the speaker to detect the facial contour and the hand contour of the speaker in the picture frame of actions of the speaker, and then perform mask processing on the facial action and the hand action of the speaker in the picture frame of actions of the speaker according to the facial contour and the hand contour of the speaker.

[0079] Specifically, the face contour of the speaker in the speaker action picture frame can be detected by using an existing face detection algorithm to locate the face region in the speaker action picture frame, and then using an existing edge detection algorithm (such as Canny edge detection, Sobel operator, Prewitt operator, etc.) to extract the face contour of the face region. The hand contour of the speaker in the speaker action picture frame can be detected by using an existing hand detection algorithm (such as shape-based hand detection, deep learning-based hand detection, etc.) to locate the hand region in the speaker action picture frame, and then using an existing edge detection algorithm to extract the hand contour of the hand region. The face and hand of the speaker in the speaker action picture frame can be masked by using the face contour and hand contour of the speaker according to the face contour and hand contour of the speaker.

[0080] The embodiment also needs to perform face recognition on the speaker action picture frame, extract the face features of the speaker in the speaker action picture frame, and use the face features as the face identity features of the speaker. The face recognition of the speaker action picture frame is preferably performed by using a face recognizer, based on image processing technology and machine learning algorithms, to detect and track the face in the speaker action picture frame, position and preprocess the detected face, such as descaling, denoising, normalization, etc., and then perform feature extraction on the positioned and preprocessed face to obtain the face identity features of the speaker. The face recognizer can use existing SphereFace network, ArcFace network, etc. The face identity features of the speaker are identified to facilitate the adjustment of the face action of the speaker in the speaker action picture frame, and the face is still the face of the speaker. If the face in the speaker action picture frame is to be replaced by the face of another person, the picture frame of the other person can be obtained, the face identity features of the other person can be identified from the picture frame of the other person, and the face action of the speaker in the speaker action picture frame can be adjusted according to the face identity features of the other person. The face of the speaker can be replaced by the face of the other person, thereby realizing the function of flexible replacement of the face of the speaker.

[0081] S202, performing emotion classification on each audio frame in the speech audio of the speaker to determine a sequence of emotion features corresponding to the speech audio.

[0082] S203, based on the speech audio, the sequence of emotion features, the masked picture of the speaker, and the face identity features of the speaker, using a pre-determined sign language action generation algorithm to generate the face action and hand action of the speaker in the masked picture of the speaker, and obtaining a sequence of sign language action pictures of the speaker corresponding to the speech audio.

[0083] The embodiment predefines a sign language action generation algorithm, which is determined by sign language action generation learning of a pre-recorded sign language action video of a hearing-impaired person. The hearing-impaired person is currently communicating with the hearing-impaired person, and the sign language action generation algorithm learns the generation mode of the sign language expression style of the hearing-impaired person through the pre-recorded sign language action video of the hearing-impaired person, so that the sign language action picture sequence generated by the sign language action generation algorithm conforms to the sign language expression style of the hearing-impaired person, is more convenient for the hearing-impaired person to understand, and improves the understanding accuracy.

[0084] Specifically, the embodiment needs to generate the facial action and hand action of the speaker in the speaker's mask picture according to the voice audio, the emotion feature sequence, the speaker's mask picture and the speaker's facial identity feature, using the pre-determined sign language action generation algorithm, so as to obtain the speaker's sign language action picture sequence corresponding to the voice audio. The sign language action generation algorithm can analyze the rhythm amplitude of each audio frame in the voice audio according to the voice audio and the emotion feature sequence, then determine the moving information of the face and the hand of each audio frame according to the rhythm amplitude of each audio frame, and finally supplement the moving hand action to the hand position in the speaker's mask picture and the moving facial action to the facial identity feature of the speaker on the face position according to the moving information, so as to obtain the speaker's sign language action picture corresponding to each audio frame. In the above manner, after determining the speaker's sign language action picture corresponding to each audio frame in the voice audio, the speaker's sign language action picture sequence is formed according to the time sequence of each audio frame in the voice audio.

[0085] Further, the step specifically includes the following:

[0086] First, determine the rhythm amplitude feature corresponding to the voice audio based on the voice audio and the emotion feature sequence.

[0087] The embodiment analyzes the rhythm feature and the amplitude feature matched with the emotion in the voice audio according to the voice audio and the emotion feature sequence, wherein the rhythm feature includes the rhythm feature corresponding to each audio frame in the voice audio, i.e., the rhythm feature is a sequence composed of the rhythm feature corresponding to each audio frame, and the amplitude feature includes the amplitude feature corresponding to each audio frame in the voice audio, i.e., the amplitude feature is a sequence composed of the amplitude feature corresponding to each audio frame. The rhythm amplitude feature corresponding to the voice audio in the embodiment includes the rhythm feature corresponding to the voice audio and the amplitude feature corresponding to the voice audio, and the rhythm feature and the amplitude feature can be separated or spliced together.

[0088] In this embodiment, the pre-determined sign language action generation algorithm includes a rhythm amplitude feature determination algorithm. The rhythm amplitude feature determination algorithm is used for speech and emotion analysis on the speech audio and the corresponding emotion feature sequence to determine the rhythm amplitude feature corresponding to the speech audio. The rhythm amplitude feature determination algorithm in this embodiment is obtained by rhythm amplitude feature determination training using sample audio, artificially labeled emotion features of the sample audio, and artificially labeled rhythm amplitude features of the sample audio. That is, based on the sample audio and the artificially labeled emotion features of the sample audio, the rhythm amplitude feature determination algorithm is used to determine the predicted rhythm amplitude feature corresponding to the sample audio. Then, according to the difference between the predicted rhythm amplitude feature and the artificially labeled rhythm amplitude feature of the sample audio, and taking the minimum difference as the target, the rhythm amplitude feature determination algorithm is adjusted and improved.

[0089] Second, based on the mask picture of the speaker, the facial identity feature of the speaker, and the rhythm amplitude feature corresponding to the speech audio, the facial action and the hand action of the speaker in the mask picture of the speaker are generated to obtain the sequence of sign language action pictures of the speaker corresponding to the speech audio.

[0090] This embodiment is based on the pre-obtained mask picture of the speaker, the facial identity feature of the speaker, and the rhythm amplitude feature corresponding to the speech audio. According to the rhythm amplitude feature and the facial identity feature of the speaker, the facial movement information and the hand movement information of the speaker are determined. Then, according to the facial movement information and the hand movement information of the speaker, the facial action and the hand action of the facial position and the hand position in the mask picture of the speaker are determined, and the facial action and the hand action are filled into the mask picture of the speaker to obtain the sequence of sign language action pictures of the speaker corresponding to the speech audio.

[0091] Specifically, this embodiment first determines the action optical flow feature corresponding to the speech audio based on the rhythm amplitude feature corresponding to the speech audio.

[0092] According to the rhythm amplitude feature corresponding to the speech audio, the movement information of the face moving according to the rhythm amplitude feature is determined as the facial action optical flow feature, and the movement information of the hand moving according to the rhythm amplitude feature is determined as the hand action optical flow feature. The set of the facial action optical flow feature and the hand action optical flow feature is taken as the action optical flow feature corresponding to the speech audio. Among them, the facial motion trajectory can be determined according to the facial action optical flow feature, and the hand motion trajectory can be determined according to the hand action optical flow feature.

[0093] In this embodiment, the pre-determined sign language action generation algorithm also includes an optical flow feature determination algorithm. Based on the rhythm amplitude features corresponding to the speech audio, the optical flow feature determination algorithm can determine the action optical flow features corresponding to the speech audio.

[0094] Then, based on the speaker's mask image, the speaker's facial identity features, and the motion optical flow features corresponding to the speech audio, the speaker's facial and hand movements in the mask image are generated, resulting in a sequence of speaker sign language motion images corresponding to the speech audio.

[0095] This embodiment, after determining the motion optical flow features corresponding to the speech audio, can determine the speaker's hand movement trajectory based on the hand motion optical flow features, thereby determining the speaker's hand movements. It also determines the facial movement trajectory based on the facial motion optical flow features, and further determines the speaker's facial movements based on the facial identity features. Then, the speaker's facial and hand movements are filled into the speaker's mask image; that is, the speaker's facial movements are filled into the facial area of ​​the mask image, and the speaker's hand movements are filled into the hand area of ​​the mask image, ultimately obtaining a sequence of speaker sign language gesture images corresponding to the speech audio.

[0096] In this embodiment, the pre-determined sign language action generation algorithm also includes an action image determination algorithm. Based on the speaker's mask image, the speaker's facial identity features, and the action optical flow features corresponding to the speech audio, the action image determination algorithm can generate the speaker's facial and hand movements in the speaker's mask image, thereby obtaining a sequence of speaker sign language action images corresponding to the speech audio.

[0097] Step S202 in this embodiment is the same as step S101 in the above embodiment, and the execution method of step S202 will not be described in detail in this embodiment. In addition, this embodiment does not limit the execution order between step S201 and step S202. Step S201 can be executed first and then step S202 can be executed, or step S202 can be executed first and then step S201 can be executed, or steps S201 and step S202 can be executed simultaneously.

[0098] As an optional implementation, see [link to implementation details]. FIG. 3 As shown in another embodiment of this application, step S203 in the above embodiment is disclosed, which involves generating the speaker's facial and hand gestures in the speaker's mask image based on the speech audio, emotional feature sequence, speaker's mask image, and speaker's facial identity features, using a pre-determined sign language gesture generation algorithm, to obtain the speaker's sign language gesture image sequence corresponding to the speech audio. Specifically, this includes the following steps:

[0099] S301, input the speech audio, the emotion feature sequence, the mask picture of the speaker and the facial identity feature of the speaker into the pre-trained sign language action generation model to obtain the sign language action picture sequence corresponding to the speech audio.

[0100] After obtaining the emotion feature sequence corresponding to the speech audio, the mask picture of the speaker and the facial identity feature of the speaker, the speech audio, the emotion feature sequence, the mask picture of the speaker and the facial identity feature of the speaker are input into the pre-trained sign language action generation model to obtain the sign language action picture sequence corresponding to the speech audio. The sign language action generation model in this embodiment is obtained by sign language action generation training based on the pre-recorded sign language action video of the hearing-impaired person. The hearing-impaired person is the hearing-impaired person currently communicating. By pre-recording the sign language action video of the hearing-impaired person, the sign language action generation model learns the generation method of the sign language expression style of the hearing-impaired person, so that the sign language action picture sequence generated by the sign language action generation model conforms to the sign language expression style of the hearing-impaired person, which is more convenient for the hearing-impaired person to understand and improves the understanding accuracy.

[0101] In this embodiment, the sign language action generation model is a model for generating sign language actions using the sign language action generation algorithm in the above embodiment.

[0102] As an optional implementation, in another embodiment of the present application, it is disclosed that the sign language action generation model includes: a rhythm amplitude feature encoder, an optical flow feature encoder and a sign language action generator. The step S301 in the above embodiment specifically includes the following steps:

[0103] First, input the speech audio and the emotion feature sequence into the rhythm amplitude feature encoder to obtain the rhythm amplitude feature corresponding to the speech audio.

[0104] In this embodiment, the speech audio and the emotion feature sequence corresponding to the speech audio are input into the rhythm amplitude feature encoder in the sign language action generation model. The rhythm amplitude feature encoder can execute the rhythm amplitude feature determination algorithm in the above embodiment to obtain the rhythm amplitude feature corresponding to the speech audio. The rhythm amplitude feature encoder is composed of a convolutional neural network (CNN) and a long short-term memory network (LSTM).

[0105] Second, input the rhythm amplitude feature corresponding to the speech audio into the optical flow feature encoder to obtain the action optical flow feature corresponding to the speech audio.

[0106] The rhythm amplitude feature corresponding to the speech audio is input into the optical flow feature encoder, which can perform the optical flow feature determination algorithm in the above embodiment, so as to obtain the motion optical flow feature corresponding to the speech audio. The optical flow feature encoder is composed of a convolutional neural network (CNN) and a recurrent neural network (RNN).

[0107] Thirdly, the mask picture of the speaker, the facial identity feature of the speaker and the motion optical flow feature corresponding to the speech audio are input into the sign language action generator to obtain the sequence of sign language action pictures corresponding to the speech audio.

[0108] The mask picture of the speaker, the facial identity feature of the speaker and the motion optical flow feature corresponding to the speech audio are input into the sign language action generator, which can perform the action picture determination algorithm in the above embodiment, so as to obtain the sequence of sign language action pictures corresponding to the speech audio. The sign language action generator is composed of a generative adversarial network (GAN).

[0109] As an optional implementation, referring to FIG. 4 As shown in the figure, in another embodiment of the present application, the training process of the sign language action generation model includes the following steps:

[0110] S401, obtaining the sample audio corresponding to the sign language action video of the hearing-impaired person in different emotional states and the sample emotional feature corresponding to the sample audio.

[0111] The hearing-impaired person needs to use the sign language generation method to convert the speech audio of the speaker into sign language actions so that the hearing-impaired person can understand the meaning expressed by the speaker. In order to make the sign language actions converted from the speech audio of the speaker more in line with the expression style of the hearing-impaired person with whom the speaker communicates, the sign language action generation model needs to be trained using the sign language action video of the hearing-impaired person when training the sign language action generation model, so that the sign language action generation model can learn the sign language expression style of the hearing-impaired person. Therefore, in order to train the sign language action generation model, the sign language action video corresponding to the sample audio of the hearing-impaired person in different emotional states needs to be obtained first, and the sample emotional feature corresponding to the sample audio is determined according to the emotional state corresponding to the sample audio.

[0112] For example, for emotional states such as happy, angry, excited, etc., and texts "I ran two kilometers in the park today", "The cat is sleeping quietly", "The store is full of goods", "A few yellow and dry leaves fell from the tree", "He is writing homework", the embodiment can record sign language action videos of the hearing-impaired person for each text according to different emotional states. Five different texts, three different states, and 15 different videos can be recorded. In order to improve the accuracy of the sign language action generation model, the sign language action generation model can be realized by expanding the sign language action videos recorded by the hearing-impaired person. The sample audio is the audio corresponding to the text and the emotional state corresponding to the text.

[0113] The sample emotional features corresponding to the sample audio can be manually annotated emotional features according to the emotional state corresponding to the sample audio, the manually annotated emotional features are taken as the emotional features corresponding to each frame of the sample audio frame in the sample audio, and the emotional features corresponding to each frame of the sample audio frame are combined into the sample emotional features corresponding to the sample audio. The emotional features of each sample audio frame can also be determined by performing emotional classification on each sample audio frame in the sample audio by using an emotional classification network, and the emotional features corresponding to each frame of the sample audio frame are combined into the sample emotional features corresponding to the sample audio.

[0114] S402, perform face and hand masking processing on the video frames in the sign language action video to obtain sample mask pictures of the hearing-impaired person, and perform face recognition on the video frames in the sign language action video to obtain sample face identity features of the hearing-impaired person.

[0115] The embodiment needs to perform frame extraction operation on the sign language action video of the hearing-impaired person to obtain the video frames in the sign language action video, and then perform masking processing on the face and hands of the hearing-impaired person in the video frames to obtain the sample mask pictures of the hearing-impaired person. In the embodiment, the specific way of performing face and hand masking processing on the video frames in the sign language action video is the same as the specific way of performing face and hand masking processing on the speaker action picture frames in the above-mentioned embodiment, and the embodiment will not be specifically described.

[0116] The embodiment needs to perform face recognition on the video frames in the sign language action video obtained by the frame extraction operation to obtain the sample face identity features of the hearing-impaired person. The specific way of performing face recognition on the video frames in the sign language action video is the same as the specific way of performing face recognition on the speaker action picture frames in the above-mentioned embodiment, and the embodiment will not be specifically described.

[0117] S403, input the sample audio, the sample emotional features, the sample mask pictures of the hearing-impaired person and the sample face identity features of the hearing-impaired person into the pre-constructed sign language action generation model to obtain a predicted sign language action picture sequence corresponding to the sample audio.

[0118] The sample audio, the sample emotion feature, the sample mask picture of the hearing-impaired person, and the sample facial identity feature of the hearing-impaired person are input into the pre-constructed sign language action generation model, facial action and hand action are generated by using the sign language action generation model, and a predicted sign language action picture sequence corresponding to the sample audio is obtained.

[0119] S404, based on the difference between the predicted sign language action picture sequence corresponding to the sample audio and the sample sign language action picture sequence, the model parameter adjustment is performed on the sign language action generation model.

[0120] After obtaining the sign language action video corresponding to the sample audio, the frame extraction operation is performed on the sign language action video, all video frames contained in the sign language action video are obtained, and all video frames contained in the sign language action video are composed into a sample sign language action picture sequence according to the time sequence thereof in the sign language action video.

[0121] The model parameter adjustment is performed on the sign language action generation model based on the difference between the predicted sign language action picture sequence corresponding to the sample audio and the sample sign language action picture sequence, and the minimum difference is taken as the target.

[0122] As an optional implementation, referring to FIG. 5 It is disclosed in another embodiment of the present application that the sign language action generation model comprises a rhythm amplitude feature encoder, an optical flow feature encoder, and a sign language action generator. The step S403 in the above embodiment specifically comprises the following steps:

[0123] S501, input the sample audio and the sample emotion feature into the rhythm amplitude feature encoder to obtain a sample rhythm amplitude feature corresponding to the sample audio, and train the rhythm amplitude feature encoder based on the difference between the labeled rhythm amplitude feature and the sample rhythm amplitude feature.

[0124] The sample audio and the corresponding sample emotion feature are input into the rhythm amplitude feature encoder in the sign language action generation model to obtain a sample rhythm amplitude feature corresponding to the sample audio. The sample audio is artificially labeled in advance in the embodiment to determine a labeled rhythm amplitude feature corresponding to the sample audio. The artificial labeling of the sample audio in terms of the rhythm amplitude feature can be separate labeling of the rhythm feature and the amplitude feature of the sample audio. If the sample rhythm amplitude feature corresponding to the sample audio generated by the rhythm amplitude feature encoder is a feature of the rhythm feature and the amplitude feature spliced together, the artificially labeled rhythm feature and amplitude feature also need to be spliced according to the format of the sample rhythm amplitude feature.

[0125] The embodiment needs to train the rhythm amplitude feature encoder based on the difference between the labeled rhythm amplitude feature corresponding to the sample audio and the sample rhythm amplitude feature corresponding to the sample audio, with the minimum difference as the goal. Specifically, a first loss function between the labeled rhythm amplitude feature corresponding to the sample audio and the sample rhythm amplitude feature corresponding to the sample audio can be calculated, and then the parameters of the rhythm amplitude feature encoder are adjusted according to the first loss function, so that the first loss function becomes smaller and smaller, until the first loss function reaches the preset loss function range, and the training of the rhythm amplitude feature encoder is completed.

[0126] S502, input the sample rhythm amplitude feature corresponding to the sample audio into the optical flow feature encoder to obtain the sample action optical flow feature corresponding to the sample audio.

[0127] The embodiment inputs the sample rhythm amplitude feature output by the rhythm amplitude feature encoder in the sign language action generation model into the optical flow feature encoder in the sign language action generation model to obtain the sample action optical flow feature corresponding to the sample audio.

[0128] S503, input the sample action optical flow feature, the sample mask picture of the hearing-impaired person, and the sample facial identity feature of the hearing-impaired person into the sign language action generator to obtain the predicted sign language action picture sequence corresponding to the sample audio.

[0129] The embodiment inputs the sample mask picture of the hearing-impaired person, the sample facial identity feature of the hearing-impaired person, and the sample action optical flow feature output by the optical flow feature encoder into the sign language action generator in the sign language action generation model to obtain the predicted sign language action picture sequence corresponding to the sample audio.

[0130] S504, based on the difference between the sample action optical flow feature and the labeled optical flow feature, and the difference between the predicted sign language action picture sequence and the sample sign language action picture sequence, the optical flow feature encoder and the sign language action generator are jointly trained.

[0131] The embodiment needs to use an existing mature optical flow extractor to extract optical flow features from each video frame in the sample audio corresponding sign language action video, to obtain the optical flow features corresponding to each video frame in the sign language action video. The optical flow features corresponding to all video frames in the sign language action video are combined into a sequence according to the time sequence of the video frames in the sign language action video to obtain the labeled optical flow feature. The existing mature optical flow extractor is composed of a Transformer network, for example, a FLowNet network, a RAFT network, etc.

[0132] The embodiment also needs to perform frame extraction on the obtained sample audio corresponding sign language action video, obtain all video frames contained in the sign language action video, and group all video frames contained in the sign language action video into a sample sign language action picture sequence according to the time sequence thereof in the sign language action video.

[0133] The embodiment is based on the difference between the sample action optical flow feature and the labeled optical flow feature, and the difference between the predicted sign language action picture sequence and the sample sign language action picture sequence, and jointly trains the optical flow feature encoder and the sign language action generator with the goal of minimizing both differences. Specifically, a second loss function between the sample action optical flow feature and the labeled optical flow feature is calculated, and a third loss function between the predicted sign language action picture sequence and the sample sign language action picture sequence is calculated. Then, the optical flow feature encoder and the sign language action generator are adjusted in parameters according to the second loss function and the third loss function, so that the second loss function and the third loss function become smaller and smaller, until the second loss function and the third loss function both reach a preset loss function range, and the joint training of the optical flow feature encoder and the sign language action generator is completed.

[0134] The embodiment trains the sign language action generation model by using the sign language action video recorded by the hearing-impaired person. Compared with the standard sign language teacher or sign language announcer required in the existing two-dimensional scheme to record and shoot the sign language action video, the embodiment reduces the labor cost. Compared with the special equipment required for action capture and collection in the existing three-dimensional scheme, the need for art and design personnel to model, and the need for three-dimensional software to render, the embodiment only needs to use an image collection device to shoot the front two-dimensional video of the hearing-impaired person, and also reduces the labor cost.

[0135] Exemplary apparatus

[0136] Correspondingly, the embodiment of the application also provides a sign language generation device, as shown in FIG. 6 The device comprises:

[0137] The emotion classification module 100 is configured to perform emotion classification on each audio frame in the speech audio of the speaker, and determine a sequence of emotion features corresponding to the speech audio.

[0138] The sign language generation module 110 is configured to adjust the facial action and the hand action of the speaker in the picture frame of the speaker action based on the speech audio and the sequence of emotion features, and generate a sequence of picture frames of sign language action of the speaker corresponding to the speech audio.

[0139] The style of the facial action and the hand action of the speaker in the sequence of picture frames of sign language action of the speaker is the same as the sign language expression style of the hearing-impaired person.

[0140] It can be seen from the above that the sign language generation device provided in the embodiments of the present application can adjust the facial action and the hand action of the speaker in the picture frame of the speaker action by combining the speech audio with the emotional characteristics of the speech audio, so that the sign language action and the facial expression of the speaker have emotional characteristics, and the emotional degree of the sign language generation is improved. In addition, the style of the facial action and the hand action of the speaker in the picture sequence of the sign language action of the speaker is the same as the sign language expression style of the hearing-impaired person, the accuracy of the sign language generation is improved, and the understanding of the hearing-impaired person is more convenient.

[0141] As an optional implementation, in another embodiment of the present application, it is disclosed that the sign language generation device further comprises: a picture frame processing module.

[0142] The picture frame processing module is configured to perform face and hand mask processing on the picture frame of the speaker action to obtain a mask picture of the speaker, and perform face recognition on the picture frame of the speaker action to obtain facial identity features of the speaker.

[0143] The sign language generation module 110 is specifically configured to generate the facial action and the hand action of the speaker in the mask picture of the speaker based on the speech audio, the sequence of emotional characteristics, the mask picture of the speaker, and the facial identity features of the speaker, using a predetermined sign language action generation algorithm, to obtain a picture sequence of the sign language action of the speaker corresponding to the speech audio.

[0144] The sign language action generation algorithm is determined by learning the sign language action generation of the pre-recorded sign language action video of the hearing-impaired person.

[0145] As an optional implementation, in another embodiment of the present application, it is disclosed that the sign language generation module 110 comprises: a rhythm amplitude determination unit and an action picture determination unit.

[0146] The rhythm amplitude determination unit is configured to determine the rhythm amplitude characteristics corresponding to the speech audio based on the speech audio and the sequence of emotional characteristics.

[0147] The action picture determination unit is configured to generate the facial action and the hand action of the speaker in the mask picture of the speaker based on the mask picture of the speaker, the facial identity features of the speaker, and the rhythm amplitude characteristics corresponding to the speech audio, to obtain a picture sequence of the sign language action of the speaker corresponding to the speech audio.

[0148] As an optional implementation, in another embodiment of the present application, it is disclosed that the action picture determination unit is specifically configured to:

[0149] determine the action optical flow characteristics corresponding to the speech audio based on the rhythm amplitude characteristics corresponding to the speech audio.

[0150] The speaker's face action and hand action in the speaker's mask picture are generated based on the speaker's mask picture, the speaker's face identity feature, and the action optical flow feature corresponding to the speech audio, to obtain a sequence of sign language action pictures corresponding to the speech audio.

[0151] As an optional implementation, in another embodiment of the present application, it is disclosed that the sign language generation module 110 comprises a model processing unit.

[0152] The model processing unit is configured to input the speech audio, the sequence of emotional features, the speaker's mask picture, and the speaker's face identity feature into a pre-trained sign language action generation model to obtain a sequence of sign language action pictures corresponding to the speech audio.

[0153] The sign language action generation model is obtained by sign language action generation training based on pre-recorded sign language action videos of hearing-impaired persons.

[0154] As an optional implementation, in another embodiment of the present application, it is disclosed that the sign language action generation model comprises a rhythm amplitude feature encoder, an optical flow feature encoder, and a sign language action generator.

[0155] The model processing unit is specifically configured to:

[0156] input the speech audio and the sequence of emotional features into the rhythm amplitude feature encoder to obtain rhythm amplitude features corresponding to the speech audio;

[0157] input the rhythm amplitude features corresponding to the speech audio into the optical flow feature encoder to obtain action optical flow features corresponding to the speech audio;

[0158] input the speaker's mask picture, the speaker's face identity feature, and the action optical flow features corresponding to the speech audio into the sign language action generator to obtain a sequence of sign language action pictures corresponding to the speech audio.

[0159] As an optional implementation, in another embodiment of the present application, it is disclosed that the sign language generation device further comprises a video acquisition module, a sample picture frame processing module, a sample sign language generation module, and a model parameter adjustment module.

[0160] The video acquisition module is configured to acquire sign language action videos corresponding to sample audios of hearing-impaired persons in different emotional states and sample emotional features corresponding to the sample audios.

[0161] The sample picture frame processing module is configured to perform face and hand masking processing on video frames in the sign language action videos to obtain sample mask pictures of the hearing-impaired persons, and perform face recognition on the video frames in the sign language action videos to obtain sample face identity features of the hearing-impaired persons.

[0162] The sample sign language generation module is configured to input the sample audio, the sample emotion feature, the sample mask picture of the hearing-impaired person, and the sample facial identity feature of the hearing-impaired person into a pre-constructed sign language action generation model to obtain a predicted sign language action picture sequence corresponding to the sample audio.

[0163] The model parameter adjustment module is configured to adjust model parameters of the sign language action generation model based on a difference between the predicted sign language action picture sequence corresponding to the sample audio and the sample sign language action picture sequence obtained by frame extraction on the sign language action video corresponding to the sample audio.

[0164] As an optional implementation, in another embodiment of the present application, a sample sign language generation module is specifically configured to:

[0165] The sample audio and the sample emotion feature are input into a rhythm amplitude feature encoder to obtain a sample rhythm amplitude feature corresponding to the sample audio, and the rhythm amplitude feature encoder is trained based on a difference between a labeled rhythm amplitude feature and the sample rhythm amplitude feature; the labeled rhythm amplitude feature is a rhythm amplitude feature artificially labeled on the sample audio.

[0166] The sample rhythm amplitude feature corresponding to the sample audio is input into an optical flow feature encoder to obtain a sample action optical flow feature corresponding to the sample audio.

[0167] The sample action optical flow feature, the sample mask picture of the hearing-impaired person, and the sample facial identity feature of the hearing-impaired person are input into a sign language action generator to obtain a predicted sign language action picture sequence corresponding to the sample audio.

[0168] The optical flow feature encoder and the sign language action generator are jointly trained based on a difference between the sample action optical flow feature and a labeled optical flow feature, and a difference between the predicted sign language action picture sequence and the sample sign language action picture sequence; the labeled optical flow feature is obtained by optical flow feature extraction on a sign language action video corresponding to the sample audio.

[0169] The sign language generation device provided in the embodiment belongs to the same application concept as the sign language generation method provided in the above-mentioned embodiments of the present application, can execute the sign language generation method provided in any of the above-mentioned embodiments of the present application, and has the corresponding function modules and beneficial effects of executing the sign language generation method. Technical details not described in detail in the embodiment can be referred to the specific processing content of the sign language generation method provided in the above-mentioned embodiments of the present application, which will not be described here.

[0170] Exemplary electronic device

[0171] Another embodiment of the present application further provides an electronic device, which is described with reference to FIG. 7As shown, the device comprises:

[0172] The memory 200 and the processor 210;

[0173] The memory 200 is connected with the processor 210, and is configured to store programs;

[0174] The processor 210 is configured to realize the sign language generation method disclosed in any of the above embodiments by running the programs stored in the memory 200.

[0175] Specifically, the above electronic device can further comprise a bus, a communication interface 220, an input device 230 and an output device 240.

[0176] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected with each other through the bus.

[0177] The bus can include a path for transmitting information between various components of the computer system.

[0178] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or can be an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of programs of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a ready-to-use programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0179] The processor 210 can include a main processor, and can further include a baseband chip, a modem, etc.

[0180] The memory 200 stores programs for executing the technical solutions of the present application, and can also store operating systems and other key services. Specifically, the programs can include program codes, and the program codes include computer operation instructions. More specifically, the memory 200 can include read-only memory (ROM), other types of static storage devices that can store static information and instructions, random access memory (RAM), other types of dynamic storage devices that can store information and instructions, disk storage, flash, etc.

[0181] The input device 230 can include devices that receive data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer or a gravity sensor, etc.

[0182] The output device 240 can include any device that allows output of information such as displays, printers, loudspeakers, etc.

[0183] The communication interface 220 can include any transceiver-like mechanism for communicating with other devices or communication networks such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0184] The processor 210 executes the programs stored in the memory 200 and invokes other devices, which can be used to implement each step of any of the sign language generation methods provided by the embodiments described above.

[0185] Exemplary computer program product and storage medium

[0186] In addition to the above methods and devices, the embodiments of the present application can also be computer program products, which include computer program instructions, which, when executed by a processor, cause the processor to perform the steps of the sign language generation methods according to various embodiments of the present application described in the above "Exemplary Methods" section of the specification.

[0187] The computer program product can be written in any combination of one or more programming languages, including an object-oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0188] In addition, the embodiments of the present application can also be storage media, which store computer programs, and the computer programs are executed by a processor to perform the steps of the sign language generation methods according to various embodiments of the present application described in the above "Exemplary Methods" section of the specification.

[0189] For each of the above method embodiments, in order to simply describe, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the present application is not limited by the order of the described actions, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0190] It should be noted that each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between embodiments can be mutually referred to.

[0191] The steps in the method of each embodiment of the application can be adjusted, combined and deleted in sequence according to actual needs. The technical features described in each embodiment can be replaced or combined.

[0192] The modules and sub-modules in the device and terminal of each embodiment of the application can be combined, divided and deleted according to actual needs.

[0193] In several embodiments provided by the application, it should be understood that the disclosed terminal, device and method can be implemented by other ways. For example, the terminal embodiments described above are merely schematic. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, another division manner can be adopted. For example, a plurality of sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed modules can be indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0194] The modules or sub-modules described as separate components can or can not be physically separated, and the components of the modules or sub-modules can or can not be physical modules or sub-modules, that is, they can be located in one place or distributed on a plurality of network modules or sub-modules. Part or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0195] In addition, each functional module or sub-module in each embodiment of the application can be integrated in one processing module, or each module or sub-module can exist physically, or two or more modules or sub-modules can be integrated in one module. The integrated module or sub-module can be realized in the form of hardware or in the form of software functional module or sub-module.

[0196] Those skilled in the art will further appreciate that the units and algorithm steps of the various examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein in terms of their functionality, their composition, and their manner of operation. Whether such functionality is implemented in hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled persons can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0197] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0198] Finally, it should be noted that the terms "first", "second", and the like, herein do not denote any order, quantity, combination, or importance, but rather are used to distinguish one element from another, and are more especially used for the purpose of identification in claims. Moreover, the terms "comprise", "include" or "contain" or any other variant thereof, are intended to cover non-exclusive inclusions, such that processes, methods, articles, or apparatuses that comprise, include, or contain a list of elements, not only include those elements, but also other elements not expressly listed or otherwise inherent to such processes, methods, articles, or apparatuses. Without further limitation, an element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0199] The above description of disclosed embodiments allows a person skilled in the art to implement or use the application. Numerous modifications to these embodiments will be apparent to those skilled in the art, and general principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Therefore, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of generating a sign language, characterized by, The method comprises: performing emotion classification on each frame of audio frame in speech audio of a speaker, determining a sequence of emotion features corresponding to the speech audio, and performing mask processing on face and hand of a picture frame of speaker action to obtain a mask picture of the speaker, and performing face recognition on the picture frame of speaker action to obtain face identity features of the speaker; based on the speech audio, the sequence of emotion features, the mask picture of the speaker and the face identity features of the speaker, using a pre-determined sign language action generation algorithm, generating speaker face action and hand action in the mask picture of the speaker to obtain a sequence of speaker sign language action pictures corresponding to the speech audio; wherein the sign language action generation algorithm is determined by sign language action generation learning using pre-recorded sign language action videos of hearing-impaired persons, so that the sign language action generation algorithm learns the sign language expression style of the hearing-impaired persons, and the style of speaker face action and hand action in the sequence of speaker sign language action pictures is the same as the sign language expression style of the hearing-impaired persons.

2. The method of claim 1, wherein, based on the speech audio, the sequence of emotion features, the mask picture of the speaker and the face identity features of the speaker, using a pre-determined sign language action generation algorithm, generating speaker face action and hand action in the mask picture of the speaker to obtain a sequence of speaker sign language action pictures corresponding to the speech audio, comprising: determining a rhythm amplitude feature corresponding to the speech audio based on the speech audio and the sequence of emotion features; based on the mask picture of the speaker, the face identity features of the speaker and the rhythm amplitude feature corresponding to the speech audio, generating speaker face action and hand action in the mask picture of the speaker to obtain a sequence of speaker sign language action pictures corresponding to the speech audio.

3. The method of claim 2, wherein, based on the mask picture of the speaker, the face identity features of the speaker and the rhythm amplitude feature corresponding to the speech audio, generating speaker face action and hand action in the mask picture of the speaker to obtain a sequence of speaker sign language action pictures corresponding to the speech audio, comprising: determining an action optical flow feature corresponding to the speech audio based on the rhythm amplitude feature corresponding to the speech audio; based on the mask picture of the speaker, the face identity features of the speaker and the action optical flow feature corresponding to the speech audio, generating speaker face action and hand action in the mask picture of the speaker to obtain a sequence of speaker sign language action pictures corresponding to the speech audio.

4. The method of claim 1, wherein, based on the speech audio, the sequence of emotion features, the mask picture of the speaker and the face identity features of the speaker, using a pre-determined sign language action generation algorithm, generating speaker face action and hand action in the mask picture of the speaker to obtain a sequence of speaker sign language action pictures corresponding to the speech audio, comprising: inputting the speech audio, the sequence of emotion features, the mask picture of the speaker and the face identity features of the speaker into a pre-trained sign language action generation model to obtain a sequence of speaker sign language action pictures corresponding to the speech audio; The sign language action generation model is obtained by sign language action generation training based on pre-recorded sign language action videos of hearing-impaired persons.

5. The method of claim 4, wherein, The sign language action generation model comprises a rhythm amplitude feature encoder, an optical flow feature encoder, and a sign language action generator. The voice audio, the emotion feature sequence, the mask picture of the speaker, and the facial identity feature of the speaker are input into a pre-trained sign language action generation model to obtain a speaker sign language action picture sequence corresponding to the voice audio, which comprises: The voice audio and the emotion feature sequence are input into the rhythm amplitude feature encoder to obtain rhythm amplitude features corresponding to the voice audio. The rhythm amplitude features corresponding to the voice audio are input into the optical flow feature encoder to obtain action optical flow features corresponding to the voice audio. The mask picture of the speaker, the facial identity feature of the speaker, and the action optical flow features corresponding to the voice audio are input into the sign language action generator to obtain the speaker sign language action picture sequence corresponding to the voice audio.

6. The method of claim 4, wherein, The training process of the sign language action generation model comprises: Obtaining sample audio corresponding to sign language action videos of hearing-impaired persons in different emotional states and sample emotion features corresponding to the sample audio. The video frames in the sign language action videos are subjected to face and hand mask processing to obtain sample mask pictures of the hearing-impaired persons, and the video frames in the sign language action videos are subjected to face recognition to obtain sample facial identity features of the hearing-impaired persons. The sample audio, the sample emotion features, the sample mask pictures of the hearing-impaired persons, and the sample facial identity features of the hearing-impaired persons are input into a pre-constructed sign language action generation model to obtain a predicted sign language action picture sequence corresponding to the sample audio. Based on the difference between the predicted sign language action picture sequence corresponding to the sample audio and a sample sign language action picture sequence obtained by frame extraction on the sign language action videos corresponding to the sample audio, the model parameters of the sign language action generation model are adjusted.

7. The method of claim 6, wherein, The sample audio, the sample emotion features, the sample mask pictures of the hearing-impaired persons, and the sample facial identity features of the hearing-impaired persons are input into a pre-constructed sign language action generation model to obtain a predicted sign language action picture sequence corresponding to the sample audio, which comprises: The sample audio and the sample emotion features are input into a rhythm amplitude feature encoder to obtain sample rhythm amplitude features corresponding to the sample audio, and the rhythm amplitude feature encoder is trained based on the difference between the labeled rhythm amplitude features and the sample rhythm amplitude features; wherein the labeled rhythm amplitude features are rhythm amplitude features artificially labeled on the sample audio. The sample rhythm amplitude features corresponding to the sample audio are input into an optical flow feature encoder to obtain sample action optical flow features corresponding to the sample audio. The sample rhythm amplitude features corresponding to the sample audio are input into an optical flow feature encoder to obtain sample action optical flow features corresponding to the sample audio. input the sample action optical flow feature, the sample mask picture of the hearing-impaired person and the sample facial identity feature of the hearing-impaired person into a sign language action generator to obtain a predicted sign language action picture sequence corresponding to the sample audio; based on the difference between the sample action optical flow feature and the labeled optical flow feature, and the difference between the predicted sign language action picture sequence and the sample sign language action picture sequence, the optical flow feature encoder and the sign language action generator are jointly trained; wherein the labeled optical flow feature is obtained by extracting the optical flow feature from the sign language action video corresponding to the sample audio.

8. A sign language generating apparatus characterized by comprising: Comprising: an emotion classification module for classifying the emotion of each audio frame in the speech audio of the speaker, and determining the emotion feature sequence corresponding to the speech audio; a picture frame processing module for performing face and hand mask processing on the action picture frame of the speaker to obtain the mask picture of the speaker, and performing face recognition on the action picture frame of the speaker to obtain the facial identity feature of the speaker; a sign language generation module for generating the speaker facial action and hand action in the mask picture of the speaker based on the speech audio, the emotion feature sequence, the mask picture of the speaker and the facial identity feature of the speaker, using a predetermined sign language action generation algorithm, to obtain a speaker sign language action picture sequence corresponding to the speech audio; wherein the sign language action generation algorithm is determined by learning sign language action generation from pre-recorded sign language action videos of hearing-impaired persons, so that the sign language action generation algorithm learns the sign language expression style of the hearing-impaired persons, and the style of the speaker facial action and hand action in the speaker sign language action picture sequence is the same as the sign language expression style of the hearing-impaired persons.

9. An electronic device, comprising: Comprising: a memory and a processor; the memory is connected with the processor and is used for storing programs; the processor is used for realizing the sign language generation method of any one of claims 1 to 7 by running the programs in the memory.

10. A storage medium, characterized by The storage medium has a computer program stored thereon, and the computer program is executed by a processor to realize the sign language generation method of any one of claims 1 to 7.

11. A computer program product, characterised in that, comprising computer program instructions, which when executed by a processor, cause the processor to realize the sign language generation method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Sign language video generation method, sign language video translation method, sign language video customer service method and device and readable medium

    CN113835522A

  • Sign language synthesis method and device thereof, computer equipment and storage medium

    CN114048757A