Method, apparatus, device, and storage medium for processing multimedia information
By visually identifying and image processing of sign language video information of hearing-impaired people, sign language text sequences and tone characteristics are extracted, and natural language text with tone is generated, which solves the problem of inaccurate sign language translation in the prior art and achieves more accurate communication.
Patent Information
- Application Number
- CN202111550961.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-17
AI Technical Summary
Existing sign language translation software cannot fully or accurately express the intentions of hearing-impaired people, resulting in poor communication between normal people and hearing-impaired people.
By obtaining sign language video information for hearing-impaired people, visual recognition and image processing are performed, sign language text sequences and tone characteristics are extracted, and natural language text with tone is generated to more accurately convey the intention.
Accurate translation of the sign language of hearing-impaired people is achieved, and their intentions can be expressed in a complete and accurate manner, thereby improving the communication efficiency between normal people and hearing-impaired people.
Smart Images

Figure CN114495260B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology, and in particular, to a method, apparatus, device, and storage medium for processing multimedia information. Background Art
[0002] In a real communication scenario, communication between normal people can be barrier-free. For example, both parties to the communication can communicate through means such as text and voice. However, in the scenario of communication between a normal person and a hearing-impaired person, the normal person may not understand the sign language of the hearing-impaired person. Therefore, a sign language translation software can be used to achieve two-way translation between sign language and natural language.
[0003] However, the inventors of the present application found that current sign language translation software can translate the sign language of a hearing-impaired person into a sign language text sequence. However, the sign language text sequence is not sufficient to completely or accurately express what the hearing-impaired person really wants to express. Therefore, it will lead to unsmooth communication between normal people and hearing-impaired people. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, apparatus, device, and storage medium for processing multimedia information, so that the first natural language text with a tone can completely and accurately express what the hearing-impaired person really wants to express, thereby ensuring normal communication and interaction between normal people and hearing-impaired people.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for processing multimedia information, including:
[0006] Obtain original video information, where the original video information includes a sign language picture of a first user;
[0007] Perform visual recognition on the sign language picture to obtain a first sign language text sequence;
[0008] Perform image processing on the sign language picture to obtain first feature information, where the first feature information is used to characterize the first tone of the sign language;
[0009] Generate a first natural language text with the first tone according to the first sign language text sequence and the first feature information.
[0010] In a second aspect, an embodiment of the present disclosure provides a multimedia information processing apparatus, including:
[0011] A first acquisition module, configured to acquire original video information, where the original video information includes a sign language picture of a first user;
[0012] A visual recognition module, configured to perform visual recognition on the sign language picture to obtain a first sign language text sequence;
[0013] An image processing module for performing image processing on the sign language screen to obtain first feature information, where the first feature information is used to characterize the first tone of the sign language;
[0014] A first generation module for generating a first natural language text with the first tone according to the first sign language text sequence and the first feature information.
[0015] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:
[0016] A memory;
[0017] A processor; and
[0018] A computer program;
[0019] Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method described in the first aspect.
[0020] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium for processing multimedia information, on which a computer program is stored, and the computer program is executed by a processor to implement the method described in the first aspect.
[0021] The method, device, equipment and storage medium for processing multimedia information provided by the embodiments of the present disclosure obtain the original video information of a deaf person signing, and perform visual recognition on the sign language screen in the original video information to obtain a first sign language text sequence. In addition, the sign language screen can also be subjected to image processing to obtain first feature information that can characterize the tone of the sign language. Further, according to the first sign language text sequence and the first feature information, a first natural language text with this tone is generated, so that the first natural language text with tone can completely and accurately express what the deaf person really wants to express, thereby ensuring that normal people and deaf people can communicate and interact normally. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure.
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a flowchart of the method for processing multimedia information provided by the embodiments of the present disclosure;
[0025] Figure 2 Schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0026] Figure 3 Schematic diagram of another application scenario provided by an embodiment of the present disclosure;
[0027] Figure 4 Flowchart of a method for processing multimedia information provided by another embodiment of the present disclosure;
[0028] Figure 5 Schematic diagram of a sign language recognition link provided by another embodiment of the present disclosure;
[0029] Figure 6 Flowchart of frame extraction provided by another embodiment of the present disclosure;
[0030] Figure 7 Schematic diagram of face region feature extraction provided by another embodiment of the present disclosure;
[0031] Figure 8 Schematic diagram of a single - stream model provided by another embodiment of the present disclosure;
[0032] Figure 9 Schematic diagram of a two - stream model provided by another embodiment of the present disclosure;
[0033] Figure 10 Schematic diagram of a multi - task model provided by another embodiment of the present disclosure;
[0034] Figure 11 Flowchart of a method for processing multimedia information provided by another embodiment of the present disclosure;
[0035] Figure 12 Schematic diagram of a sign language synthesis link provided by another embodiment of the present disclosure;
[0036] Figure 13 Schematic diagram of a multi - task model in the case of single input provided by another embodiment of the present disclosure;
[0037] Figure 14 Schematic diagram of the structure of a multimedia information processing device provided by an embodiment of the present disclosure;
[0038] Figure 15 Schematic diagram of the structure of an electronic device embodiment provided by an embodiment of the present disclosure. Detailed implementation manners
[0039] In order to more clearly understand the above - mentioned objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0040] Numerous specific details are set forth in the following description in order to provide a thorough understanding of the present disclosure, but the present disclosure may be practiced in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all of the embodiments.
[0041] In the scenario of communication between normal people and hearing-impaired people, normal people may not understand the sign language of hearing-impaired people. Therefore, a sign language translation software can be used to achieve two-way translation between sign language and natural language. However, the inventors of the present application have found that current sign language translation software can translate the sign language of hearing-impaired people into a sign language text sequence. However, the sign language text sequence is not sufficient to completely or accurately express what the hearing-impaired people really want to express. Therefore, it will lead to unsmooth communication between normal people and hearing-impaired people. To address this problem, the embodiments of the present disclosure provide a method for processing multimedia information, which can be applied to unidirectional or bidirectional sign language translation products. Among them, unidirectional sign language translation can be translation from sign language to natural language or from natural language to sign language. Bidirectional sign language translation includes a sign language recognition link and a sign language synthesis link. The sign language recognition link completes the translation from sign language to natural language. For example, it translates the sign language of hearing-impaired people into natural language. The sign language synthesis link completes the translation from natural language to sign language. For example, it translates natural language into sign language.
[0042] Specifically, in the sign language recognition link, the embodiments of the present disclosure can generate a natural language sentence with the correct tone according to the sign language text sequence obtained by visual recognition and the factors available for judging the tone. For example, the input is "apple / crispy / good" and the facial expression feature of doubt, and the output is "Is the apple crispy?".
[0043] In the sign language synthesis link, the embodiments of the present disclosure can generate a sign language text sequence and a tone category according to the natural language sentence and the factors available for judging the tone. For example, the input is "Is the apple crispy?" and the intonation, and the output is "apple / crispy / good" and the category of doubt. Among them, the sign language text sequence can also be called a sign language vocabulary sequence. The natural language sentence can also be called a natural language text. The following introduces this method with specific embodiments.
[0044] Figure 1 It is a flowchart of the method for processing multimedia information provided by the embodiments of the present disclosure. This embodiment is applicable to, for example Figure 2The application scenario shown includes a terminal 21 and a server 22, and information interaction can be carried out between the terminal 21 and the server 22. Among them, the terminal 21 specifically includes, but is not limited to, smart phones, palm computers, tablet computers, wearable devices with displays, desktop computers, laptop computers, all-in-one computers, smart home devices, etc. In some communication scenarios, normal people and hearing-impaired people can use the same terminal, such as the terminal 21, to achieve communication and interaction. In some other communication scenarios, normal people and hearing-impaired people can also communicate remotely. For example, Figure 3 As shown, the terminal 21 is the terminal of a normal person, the terminal 23 is the terminal of a hearing-impaired person, and the terminal 21 and the terminal 23 communicate remotely through the server 22. Specifically, the method for processing multimedia information described in this embodiment can be executed by the terminal 21 or the server 22. Hereinafter, the server 22 will be taken as an example to introduce this method. As Figure 1 shown, the specific steps of this method are as follows:
[0045] S101. Obtain the original video information, where the original video information includes the sign language picture of the first user.
[0046] For example, Figure 2 As shown, the terminal 21 can collect the video information of a hearing-impaired person. The hearing-impaired person can be recorded as the first user, and this video information can be recorded as the original video information, and the original video information includes the sign language picture of the hearing-impaired person. Further, the terminal 21 can send the original video information to the server 22. Or as Figure 3 shown, the terminal 23 can collect the original video information of the hearing-impaired person and send the original video information to the server 22. Thus, the server 22 can obtain the original video information.
[0047] S102. Perform visual recognition on the sign language picture to obtain the first sign language text sequence.
[0048] For example, when the server 22 receives the original video information, it can perform visual recognition on the sign language picture in the original video information. Through visual recognition, the actions in the sign language picture can be recognized, and each action can be recognized as one or more words. The sign language picture can be a continuous series of video frames, so that the sign language picture can include multiple actions, and the words corresponding to each action together form the first sign language text sequence. For example, "grape / again / pick / 10 / sweet / good".
[0049] S103. Perform image processing on the sign language picture to obtain first feature information, where the first feature information is used to characterize the first tone of the sign language.
[0050] For example, the server 22 can also perform image processing on the sign language screen to obtain first feature information. For example, the first feature information can be the expression category, facial key points, or facial features of the hearing-impaired person determined by the server 22 from the sign language screen. It can be understood that when a hearing-impaired person uses sign language to express different tones, it is often accompanied by different facial states. For example, when expressing a questioning tone, there are often actions such as frowning and turning the face sideways. Therefore, the facial state can specifically be facial expressions, the relative positional relationship between facial features, etc. That is to say, there is a certain correspondence between the tone and the facial state. In addition, when the facial state is different, the first feature information obtained by performing image processing on the sign language screen is also different. Therefore, different tones can be represented by different first feature information. For example, in this embodiment, the first feature information can be used to represent the first tone when a hearing-impaired person uses sign language.
[0051] S104. Generate a first natural language text with the first tone according to the first sign language text sequence and the first feature information.
[0052] For example, the first tone when a hearing-impaired person uses sign language is a questioning tone. Further, the server 22 can generate a first natural language text with a questioning tone according to the first sign language text sequence and the first feature information. For example, the first sign language text sequence is "grape / again / pull / 10 / sweet / good" as described above, and the first feature information represents a questioning tone. The first natural language text with a questioning tone generated by the server 22 is "Can I have 10 more bunches of grapes? Are they sweet?".
[0053] It can be understood that this embodiment introduces the method for processing multimedia information by taking the server 22 as an example. In other embodiments, the terminal 21 can also execute this method.
[0054] In some other embodiments, some of the steps in S101 - S104 above can be executed by the server 22, and other steps can be executed by the terminal 21. For example, the terminal 21 can execute S101 - S103. Further, the terminal 21 can send the first sign language text sequence and the first feature information to the server 22, so that the server 22 can execute S104, that is, generate a first natural language text with the first tone according to the first sign language text sequence and the first feature information.
[0055] In addition, if S104 is executed by the server 22, then the server 22 can also send the first natural language text with the first tone generated by it to the terminal 21. The terminal 21 can display the first natural language text on the screen, or the terminal 21 can convert the first natural language text with the first tone into an audio with the first tone and play the audio, so that normal people and hearing-impaired people can communicate and interact normally.
[0056] In an embodiment of the present disclosure, the original video information of a sign language user is obtained, and visual recognition is performed on the sign language images in the original video information to obtain a first sign language text sequence. Additionally, image processing can be performed on the sign language images to obtain first feature information that can characterize the sign language tone. Further, according to the first sign language text sequence and the first feature information, a first natural language text with this tone is generated, enabling the first natural language text with the tone to completely and accurately express what the sign language user truly wants to convey, thereby ensuring normal communication between normal people and sign language users.
[0057] Taking the sign language recognition link as an example, the method for processing multimedia information provided in the embodiments of the present disclosure will be introduced below.
[0058] Figure 4 It is a flowchart of the method for processing multimedia information provided in another embodiment of the present disclosure. The method specifically includes the following steps:
[0059] S401. Obtain the original video information, where the original video information includes the sign language images of a first user.
[0060] Specifically, the implementation manners and specific principles of S401 and S101 are the same, and will not be elaborated here.
[0061] S402. Perform visual recognition on the sign language images to obtain a first sign language text sequence.
[0062] For example, in the Figure 5 shown sign language recognition link, the sign language images can obtain a first sign language text sequence through visual recognition, such as "grape / again / pick / 10 / sweet / good".
[0063] S403. Perform image processing on the sign language images to obtain first feature information, where the first feature information is used to characterize the first tone of the sign language.
[0064] As Figure 5 shown, obtain the tone feature from the sign language images. This tone feature is denoted as the first feature information, and this tone feature is a factor that can be used to judge the tone. Further, "grape / again / pick / 10 / sweet / good" and the first feature information are used as the inputs of the text translation model. The text translation model can integrate the first feature information into the translation process from the sign language text sequence to the natural language, thereby generating a first natural language text with the first tone by the text translation model. For example, "Another 10 bunches of grapes, are they sweet?". In the Figure 5 shown sign language recognition link, "grape / again / pick / 10 / sweet / good" can be used as the source text, and "Another 10 bunches of grapes, are they sweet?" can be used as the target text. Additionally, as Figure 5The text translation model shown may be based on an Encoder-Decoder architecture.
[0065] Optionally, the first feature information includes at least one of the following: the facial features corresponding to the face region in the sign language picture; the expression category of the first user in the sign language picture; the facial key points of the first user in the sign language picture.
[0066] Taking the facial features as an example, image processing is performed on the sign language picture to obtain the first feature information, including: extracting one or more frames of images from the sign language picture at a preset time interval or a preset frame number interval; detecting the face region in each frame of the one or more frames of images; obtaining the facial features from the face region, and using the facial features as the first feature information.
[0067] For example, the camera of the terminal 21 collects the original video information of the hearing-impaired person, and the terminal 21 sends the original video information to the server 22. The sign language picture included in the original video information may be a continuous plurality of video frames. For example, Figure 6 the video frames 1 - video frame N shown here, and one video frame here can be understood as one frame of image. Further, the server 22 may extract one or more frames of images from the video frames 1 - video frame N at a fixed time interval or a fixed frame number interval. For example, after frame extraction, video frames 1, 5,..., n are obtained. Figure 7 The image 71 shown may be one of the video frames 1, 5,..., n. Taking the image 71 as an example below, the generation process of the first feature information is introduced. For example, face localization is performed on the image 71. For example, the face region in the image 71 is detected using OpenCV (a cross-platform computer vision and machine learning software library) or other face detection models. The face region may be as Figure 7 shown as 72. Further, feature extraction is performed on the face region 72. For example, a deep convolutional network such as a Residual Network (ResNet) is used to extract facial features from the face region 72. The facial features may be as Figure 7 shown as the numerical features, and the numerical features may be used as the first feature information. Through Figure 7 it can be seen that one frame of image may correspond to one piece of first feature information. Therefore, as Figure 6Each of the video frames 1, video frame 5, …, video frame n shown can correspond to a first feature information. It can be understood that, in some embodiments, frame extraction may not be performed on the sign language picture, so that each video frame in the sign language picture can correspond to a first feature information. In addition, the first feature information can be not only facial features, but also expression categories, facial key points, etc. Among them, the expression categories can specifically be categories such as happy, angry, sad, and delighted. The facial key points can be the points corresponding to the five facial features in the image.
[0068] S404. Determine the first numerical representation corresponding to the first sign language text sequence and the first spatio-temporal representation corresponding to the first feature information.
[0069] For example Figure 5 As shown, when "grape / again / pick / 10 / sweet / good" and the first feature information are input into the text translation model, the text translation model can first determine the first numerical representation corresponding to "grape / again / pick / 10 / sweet / good" and the first spatio-temporal representation corresponding to the first feature information. For example, one character can correspond to a numerical representation, and the numerical representations corresponding to each character in "grape / again / pick / 10 / sweet / good" can form the first numerical representation.
[0070] In addition, the first feature information input into the text translation model can be the first feature information corresponding to multiple video frames respectively, that is, multiple first feature information. The multiple video frames can be multiple consecutive video frames in the sign language picture, or can be multiple video frames obtained after the above-mentioned frame extraction process. For example Figure 7 As shown, the feature extraction can be performed by a two-dimensional convolutional model. Each of the multiple video frames corresponds to a face region. When the two-dimensional convolutional model receives a face region, it can perform feature extraction on the face region to obtain a first feature information, that is, the video frame and the first feature information are in one-to-one correspondence. In some cases, the first feature information corresponding to multiple video frames respectively, that is, multiple first feature information (for example, expression categories) has temporality, and the two-dimensional convolutional model may lose the temporal correlation of the multiple first feature information. Therefore, the multiple first feature information can be input into a recurrent neural network (RNN) model, and the RNN model generates a feature, which is the first spatio-temporal representation corresponding to the multiple first feature information. The first spatio-temporal representation not only includes the multiple first feature information, but also includes the changes in the multiple first feature information in terms of time sequence.
[0071] S405. Input the first numerical representation and the first spatio-temporal representation into a first preset model, and use the first preset model to generate a first natural language text with the first tone.
[0072] For example Figure 5 The text translation model shown includes an RNN model and a first preset model. When the text translation model determines the first numerical representation corresponding to "grape / again / mention / 10 / sweet / good" and the RNN model outputs the first spatio-temporal representation, the text translation model can further input the first numerical representation and the first spatio-temporal representation into the first preset model, and use the first preset model to generate the first natural language text with the first tone. For example, "Another 10 bunches of grapes. Are they sweet?"
[0073] This embodiment does not limit the specific structure of the first preset model. For example, the first preset model can be a single-stream model, a two-stream model, or a multi-task model. They will be introduced one by one below.
[0074] For example, when the first preset model is a single-stream model, the first preset model includes an encoder, a decoder, and a feature fusion module; using the first preset model to generate the first natural language text with the first tone includes: performing a fusion process on the first numerical representation and the first spatio-temporal representation through the feature fusion module to obtain a fusion result; passing the fusion result through the encoder and the decoder in sequence, and generating the first natural language text with the first tone by the decoder.
[0075] For example Figure 8 The overall process shown can be used as Figure 5 The processing process of the text translation model shown for "grape / again / mention / 10 / sweet / good" and the first feature information. In this case Figure 8 The source text shown is as Figure 5 Shown as "grape / again / mention / 10 / sweet / good". Figure 8 The tone feature shown is the first feature information, the spatio-temporal representation is the first spatio-temporal representation, and the numerical representation is the first numerical representation. Among them, the first numerical representation is the numerical representation of "grape / again / mention / 10 / sweet / good". The process of obtaining the first spatio-temporal representation according to the first feature information can be executed by the RNN model, and the specific process will not be elaborated here. The first preset model includes an encoder, a decoder, and a feature fusion module. Specifically, the first spatio-temporal representation and the first numerical representation are fused through the feature fusion module to obtain a fusion result. Further, the fusion result is passed through the encoder and the decoder in sequence, and the decoder generates the target text, which is the first natural language text with the first tone. For example, "Another 10 bunches of grapes. Are they sweet?"
[0076] When the first preset model is a two-stream model, the encoder includes a first encoder and a second encoder; passing the fusion result through the encoder and the decoder in sequence, and generating a first natural language text with the first tone by the decoder, including: processing the first numerical representation through the first encoder to obtain a first encoding result; processing the first spatio-temporal representation through the second encoder to obtain a second encoding result; inputting the first encoding result and the second encoding result into the decoder, and generating a first natural language text with the first tone by the decoder.
[0077] For example Figure 9 The first preset model shown includes a first encoder, a second encoder, and a decoder. The first numerical representation can be processed through the first encoder to obtain a first encoding result; the first spatio-temporal representation can be processed through the second encoder to obtain a second encoding result. Further, inputting the first encoding result and the second encoding result into the decoder, and generating a target-side text by the decoder, where the target-side text is a first natural language text with the first tone. For example, "Another 10 bunches of grapes, are they sweet?". That is to say, when the first preset model is a two-stream model, the two-stream model independently encodes the first numerical representation and the first spatio-temporal representation and uses them as the input to the decoder.
[0078] When the first preset model is a multi-task model, the first preset model further includes a tone classifier; the tone classifier is used to generate category information of the first tone according to the second encoding result and the first natural language text with the first tone.
[0079] For example Figure 10 As shown, the first preset model further includes a tone classifier. That is to say, on the basis of the two-stream model, adding an output path of the tone classifier can form a multi-task model. The tone classifier can generate a tone category according to the second encoding result and the target-side text, where the target-side text is a first natural language text with the first tone, and the tone category is the category information of the first tone, such as interrogative, declarative, rhetorical, etc.
[0080] It can be understood that, as Figure 8 、 9 、10 shown, the process is the usage stage, i.e., the inference stage, of the first preset model. Before the usage stage, i.e., the inference stage, the first preset model needs to be trained. Taking the first preset model shown in Figure 10 as an example, the training process of the first preset model is introduced.
[0081] During the training process, the required sample data includes sample texts in natural language, sign language videos corresponding to the sample texts, and tone labels corresponding to the sample texts. The tone labels can be the annotation results of the tone of the sample texts by humans or machines. The sample text is denoted as y*, and the tone label is denoted as lable. Further, visual recognition is performed on the sign language video to obtain a sign language text sequence, which can be denoted as x, and x is used as the source text as shown in Figure 10 as shown. In addition, image processing is performed on the sign language video to obtain a tone feature, which is denoted as s, and s is used as the first feature information as shown in Figure 10 as shown. After processing x and s through the first preset model as shown in Figure 10 as shown, a target text and a tone category are generated. Among them, the target text is denoted as y. It can be understood that during the training process, the target text and tone category generated by the first preset model may not be accurate. Therefore, according to the true target text, that is, the sample text y*, and the correct tone label lable, the first preset model can be trained. This training process can be achieved through the following formula (1):
[0082] L = Lc(label|y, x, s) + Lg(y*|x, s) (1)
[0083] where L represents the loss function of the first preset model, Lc represents the loss function of the tone classifier, and Lg represents the loss function of text translation. The meanings of other characters are as described above and will not be elaborated here. That is to say, in the multi-task model, while completing the translation task, a well-trained tone classifier can also be used to predict the tone category. It can be understood that in the sign language recognition link, if a multi-task model is selected, then in the training stage of the multi-task model, the multi-task model outputs the target text and the tone category. In the usage stage of the multi-task model, that is, the inference stage, the multi-task model can also output the target text and the tone category. However, the server 22 only uses the target text among them.
[0084] S406. Generate target audio information with the first tone according to the first natural language text with the first tone.
[0085] For example, after the server 22 obtains the first natural language text with the first tone according to the text translation model, it can also generate target audio information with the first tone according to the first natural language text with the first tone, and send the target audio information to the terminal 21.
[0086] Embodiments of the present disclosure perform image processing on the sign language images of hearing-impaired persons to obtain factors for judging tone, i.e., tone features, and incorporate the tone features into the translation process from sign language text sequences to natural language, so as to be able to generate natural language texts with tone. In real communication scenarios, tone plays an important role. For example, for the same language text content, when expressed in different tones, its meaning is also different, and the same is true in sign language translation. For example, for the same sign language text sequence "Apple / crispy / good", if the tone of the hearing-impaired person when signing is a declarative tone, the sign language translation result is "The apple is very crispy", and if the tone of the hearing-impaired person when signing is an interrogative tone, the sign language translation result is "Is the apple crispy?", Therefore, by identifying the tone of the hearing-impaired person when signing, the translation accuracy from sign language text sequences to natural language can be improved. In addition, the method described in this embodiment does not require analyzing the context of the conversation between normal people and hearing-impaired persons or the environment where both parties are located to infer the tone of the speaker. Thus, it avoids the problems of being unable to identify the tone and perform correct translation for isolated sentences due to the lack of context, and also avoids the problems of high hardware and software requirements and difficulty in implementation due to detecting the environment where both parties are located.
[0087] It can be understood that the above embodiments introduce how to identify the tone of a hearing-impaired person when signing. In some scenarios, the tone of a normal person when speaking can also be further identified, so that the natural language text of a normal person can be translated into a more accurate sign language text sequence. The following will be introduced in combination with specific embodiments.
[0088] Figure 11 The following shows a flowchart of a method for processing multimedia information provided by another embodiment of the present disclosure. Based on the above embodiments, this method may further include the following steps:
[0089] S1101. Obtain the original audio information of the second user.
[0090] For example Figure 2 As shown, in a scenario where a normal person and a hearing-impaired person use the same terminal, such as terminal 21, the normal person is denoted as the second user, and terminal 21 can collect the audio information of the normal person when speaking, and this audio information is denoted as the original audio information. Further, terminal 21 can perform subsequent S1102 - S1104. Alternatively, terminal 21 can send the original audio information to server 22, so that server 22 can obtain the original audio information and perform S1102 - S1104.
[0091] In such as Figure 3In the application scenario shown, the terminal 21 collects the original audio information when a normal person is speaking and performs the subsequent S1102 - S1104. Alternatively, the terminal 21 can send the original audio information to the server 22, and the server 22 performs the subsequent S1102 - S1104. Or, after the terminal 21 sends the original audio information to the server 22, the server 22 can forward the original audio information to the terminal 23, and the terminal 23 performs the subsequent S1102 - S1104. The following takes the server 22 as an example for introduction.
[0092] S1102. Perform speech recognition on the original audio information to obtain a second natural language text.
[0093] For example, when the server 22 receives the original audio information from the terminal 21, it can perform speech recognition on the original audio information to obtain a second natural language text. This second natural language text can be, for example, Figure 12 shown as "Another 10 bunches of grapes. Are they sweet?". Figure 12 Shown is a schematic diagram of the sign language synthesis link.
[0094] S1103. Determine second feature information according to the original audio information and the second natural language text, where the second feature information is used to characterize the second tone of the original audio information.
[0095] For example, the server 22 can determine second feature information according to the original audio information and the second natural language text, and this second feature information is used to characterize the second tone when a normal person is speaking.
[0096] Optionally, the second feature information includes at least one of the following: the numerical feature information of the original audio information; the speech rate and intonation of the second user determined according to the original audio information; the tone classification result corresponding to the second natural language text.
[0097] For example, different intonations can represent different intentions. Therefore, the intonation of the second user can be used as a factor to judge the second user's tone. In addition, when the tone of the second user is different when speaking, the speech rate may also be different. Therefore, the speech rate can also be used as a factor to judge the second user's tone. In addition, the numerical feature information of the original audio information, that is, the speech signal, and the tone classification result corresponding to the second natural language text can also be used as factors to judge the second user's tone.
[0098] For example Figure 12As shown, the server 22 can obtain a tone feature from a second natural language text such as "Another 10 bunches of grapes. Are they sweet?" This tone feature can be the tone classification result corresponding to the second natural language text. Additionally, the server 22 can also obtain a tone feature from the original audio information. This tone feature can be the numerical feature information of the original audio information and / or the speaking speed and intonation of a normal person when speaking determined based on the original audio information.
[0099] S1104. Generate a second sign language text sequence and the category information of the second tone according to the second natural language text and the second feature information.
[0100] As Figure 12 shown, at least one of the tone classification result corresponding to the second natural language text, the numerical feature information of the original audio information, and the speaking speed and intonation of a normal person when speaking determined based on the original audio information can be used as the second feature information. The second feature information and the second natural language text are used as the input of a text translation model, and the text translation model generates a second sign language text sequence and the tone category of a normal person when speaking. The second sign language text sequence is, for example, "grape / again / take / 10 / sweet / good", and the tone category is, for example, interrogative.
[0101] Additionally, in Figure 12 , "Another 10 bunches of grapes. Are they sweet?" can be recorded as the source text, and "grape / again / take / 10 / sweet / good" can be recorded as the target text.
[0102] In this embodiment, speech recognition is performed on the original audio signal of a normal person to obtain a second natural language text. Further, a tone feature is obtained from at least one of the original audio information and the second natural language text. This tone feature can be used to judge the tone of a normal person when speaking. Further, according to the second natural language text and this tone feature, a second sign language text sequence and the tone category of a normal person are generated, thereby improving the translation accuracy from natural language to sign language and reducing the ambiguity in the sign language translation process.
[0103] In some embodiments, for example, after S1104, target video information can be further generated according to the second sign language text sequence and the category information of the second tone. The target video information includes a virtual character, and the category information of the second tone is used to determine at least one of the facial features, expression categories, and facial key points of the virtual character.
[0104] For example, the server 22 can also generate target video information including a virtual character according to the second sign language text sequence and the tone categories when a normal person speaks. Specifically, the server 22 generates an action of the virtual character according to one or more words in the second sign language text sequence, and determines at least one of the facial features, expression categories, and facial key points of the virtual character according to the tone categories when a normal person speaks. For example, when a normal person says different things, the corresponding tone categories will also be different. Therefore, the server 22 can generate virtual character images with different expressions according to different tone categories, so that when the virtual character performs different actions, its expressions may be different. Further, the actions of the virtual character are synthesized to obtain the target video information. In Figure 2 the shown scenario, the server 22 can send the target video information to the terminal 21. When the terminal 21 plays the target video information, the hearing-impaired person can experience the tone when a normal person speaks through the expression of the virtual character. In Figure 3 the shown scenario, the server 22 can send the target video information to the terminal 23, so that in the case of remote communication between the hearing-impaired person and the normal person, the hearing-impaired person can experience the tone when a normal person speaks through the expression of the virtual character.
[0105] Specifically, generating the second sign language text sequence and the category information of the second tone according to the second natural language text and the second feature information includes: determining the second numerical representation corresponding to the second natural language text and the second spatio-temporal representation corresponding to the second feature information; inputting the second numerical representation and the second spatio-temporal representation into a second preset model, and using the second preset model to generate the second sign language text sequence and the category information of the second tone.
[0106] For example Figure 12 as shown, when "Another 10 bunches of grapes, are they sweet?" and the second feature information are input into the text translation model, the text translation model can first determine the second numerical representation corresponding to "Another 10 bunches of grapes, are they sweet?" and the second spatio-temporal representation corresponding to the second feature information. Among them, the calculation process of the second spatio-temporal representation can refer to the calculation process of the above-mentioned first spatio-temporal representation, which will not be elaborated here. Further, the second numerical representation and the second spatio-temporal representation are input into the second preset model, and the second preset model is used to generate the second sign language text sequence and the tone category when a normal person speaks. The second preset model can adopt a multi-task model as Figure 10 shown. In this case, the source text as Figure 10 shown is "Another 10 bunches of grapes, are they sweet?". Figure 10 The shown tone feature is the second feature information, the spatio-temporal representation is the second spatio-temporal representation, and the numerical representation is the second numerical representation.
[0107] Optionally, the second preset model includes a first encoder, a second encoder, a decoder, and a tone classifier; generating the second sign language text sequence and the category information of the second tone by using the second preset model includes: processing the second numerical representation through the first encoder to obtain a third encoding result; processing the second spatio-temporal representation through the second encoder to obtain a fourth encoding result; inputting the third encoding result and the fourth encoding result into the decoder, and generating the second sign language text sequence by the decoder; inputting the fourth encoding result and the second sign language text sequence into the tone classifier, and generating the category information of the second tone by the tone classifier.
[0108] As Figure 10 shown, the second preset model includes a first encoder, a second encoder, a decoder, and a tone classifier. The second numerical representation is processed by the first encoder to obtain a third encoding result, and the second spatio-temporal representation is processed by the second encoder to obtain a fourth encoding result. Further, the third encoding result and the fourth encoding result are input into the decoder, and the decoder generates the second sign language text sequence, i.e., the target-side text. In addition, the fourth encoding result and the second sign language text sequence are input into the tone classifier, and the tone classifier generates the tone category when a normal person speaks.
[0109] It can be seen from the above content that, as Figure 10 shown, the multi-task model can be applied to the sign language recognition link or the sign language synthesis link. When the multi-task model is applied to the sign language synthesis link, the tone feature input into the multi-task model can be at least one of the tone classification result corresponding to the second natural language text, the numerical feature information of the original audio information, the speech rate, and the intonation when a normal person speaks determined according to the original audio information. When the tone feature is the tone classification result corresponding to the second natural language text, the structure of the multi-task model as Figure 10 shown can be changed to the structure as Figure 13 shown. In this case, the input of the multi-task model can be one, i.e., the source-side text. For example, "Another 10 bunches of grapes, are they sweet?" At this time, the tone category output by the tone classifier can be the tone classification result corresponding to the source-side text. In addition, when the tone feature includes the numerical feature information of the original audio information, the speech rate, or the intonation, the numerical feature information of the original audio information, the speech rate, or the intonation can assist the tone classifier in judging the tone category.
[0110] Figure 14 is a schematic structural diagram of a multimedia information processing device provided by an embodiment of the present disclosure. The multimedia information processing device provided by the embodiment of the present disclosure can execute the processing flow provided by the embodiment of the multimedia information processing method, as Figure 14As shown, the multimedia information processing device 140 includes:
[0111] A first acquisition module 141, configured to acquire original video information, where the original video information includes a sign language video of a first user;
[0112] A visual recognition module 142, configured to perform visual recognition on the sign language video to obtain a first sign language text sequence;
[0113] An image processing module 143, configured to perform image processing on the sign language video to obtain first feature information, where the first feature information is used to characterize the first tone of the sign language;
[0114] A first generation module 144, configured to generate a first natural language text with the first tone according to the first sign language text sequence and the first feature information.
[0115] Optionally, the first feature information includes at least one of the following:
[0116] Facial features corresponding to the face region in the sign language video;
[0117] The expression category of the first user in the sign language video;
[0118] Facial key points of the first user in the sign language video.
[0119] Optionally, when the image processing module 143 performs image processing on the sign language video to obtain first feature information, it is specifically configured to:
[0120] Extract one or more frames of images from the sign language video at preset time intervals or preset frame intervals;
[0121] Detect the face region in each frame of the one or more frames of images;
[0122] Obtain facial features from the face region, and use the facial features as the first feature information.
[0123] Optionally, when the first generation module 144 generates a first natural language text with the first tone according to the first sign language text sequence and the first feature information, it is specifically configured to:
[0124] Determine a first numerical representation corresponding to the first sign language text sequence and a first spatio-temporal representation corresponding to the first feature information;
[0125] Input the first numerical representation and the first spatio-temporal representation into a first preset model, and use the first preset model to generate a first natural language text with the first tone.
[0126] In some embodiments, the specific structure of the first preset model is not limited. For example, it can be a single-stream model, a two-stream model, or a multi-task model.
[0127] When the first preset model is a single-stream model, the first preset model includes an encoder, a decoder, and a feature fusion module; in this case, when the first generation module 144 uses the first preset model to generate the first natural language text with the first tone, it specifically is used for:
[0128] Performing a fusion process on the first numerical representation and the first spatio-temporal representation through the feature fusion module to obtain a fusion result;
[0129] Sequentially passing the fusion result through the encoder and the decoder, and the decoder generates the first natural language text with the first tone.
[0130] When the first preset model is a two-stream model, the encoder as described above includes a first encoder and a second encoder; in this case, when the first generation module 144 sequentially passes the fusion result through the encoder and the decoder, and the decoder generates the first natural language text with the first tone, it specifically is used for:
[0131] Processing the first numerical representation through the first encoder to obtain a first encoding result;
[0132] Processing the first spatio-temporal representation through the second encoder to obtain a second encoding result;
[0133] Inputting the first encoding result and the second encoding result into the decoder, and the decoder generates the first natural language text with the first tone.
[0134] When the first preset model is a multi-task model, the first preset model further includes a tone classifier; the tone classifier is used to generate category information of the first tone according to the second encoding result and the first natural language text with the first tone.
[0135] In addition, after the first generation module 144 generates the first natural language text with the first tone according to the first sign language text sequence and the first feature information, it is further used to generate target audio information with the first tone according to the first natural language text with the first tone.
[0136] It can be understood that the first tone as described above is the tone when a hearing-impaired person signs, that is to say, the above content involves how to recognize the tone when a hearing-impaired person signs. In some cases, when a normal person communicates with a hearing-impaired person, the tone of the normal person's speech can also be recognized. In this case, the multimedia information processing device 140 further includes the following modules:
[0137] A second acquisition module 145, configured to acquire the original audio information of a second user;
[0138] A speech recognition module 146, configured to perform speech recognition on the original audio information to obtain a second natural language text;
[0139] A determination module 147, configured to determine second feature information according to the original audio information and the second natural language text, where the second feature information is used to characterize the second tone of the original audio information;
[0140] A second generation module 148, configured to generate a second sign language text sequence and category information of the second tone according to the second natural language text and the second feature information.
[0141] Optionally, the second generation module 148 may further be configured to generate target video information according to the second sign language text sequence and the category information of the second tone, where the target video information includes a virtual character, and the category information of the second tone is used to determine at least one of the facial features, expression categories, and facial key points of the virtual character.
[0142] Optionally, the second feature information includes at least one of the following:
[0143] Numerical feature information of the original audio information;
[0144] The speech rate and intonation of the second user determined according to the original audio information;
[0145] The tone classification result corresponding to the second natural language text.
[0146] Optionally, when the second generation module 148 generates a second sign language text sequence and category information of the second tone according to the second natural language text and the second feature information, it is specifically configured to:
[0147] Determine a second numerical representation corresponding to the second natural language text and a second spatio-temporal representation corresponding to the second feature information;
[0148] Input the second numerical representation and the second spatio-temporal representation into a second preset model, and use the second preset model to generate the second sign language text sequence and the category information of the second tone.
[0149] In some embodiments, the specific structure of the second preset model is restricted. For example, the second preset model can be selected as a multi-task model. In this case, the second preset model includes a first encoder, a second encoder, a decoder, and a tone classifier. When the second generation module 148 uses the second preset model to generate the second sign language text sequence and the category information of the second tone, it specifically is used for:
[0150] Processing the second numerical representation through the first encoder to obtain a third encoding result;
[0151] Processing the second spatio-temporal representation through the second encoder to obtain a fourth encoding result;
[0152] Inputting the third encoding result and the fourth encoding result into the decoder, and the decoder generates the second sign language text sequence;
[0153] Inputting the fourth encoding result and the second sign language text sequence into the tone classifier, and the tone classifier generates the category information of the second tone.
[0154] Figure 14 The multimedia information processing device in the illustrated embodiment can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0155] The internal functions and structures of the multimedia information processing device are described above, and the device can be implemented as an electronic device. Figure 15 It is a schematic structural diagram of an electronic device embodiment provided by an embodiment of the present disclosure. As Figure 15 shown, the electronic device includes a memory 151 and a processor 152.
[0156] The memory 151 is used to store programs. In addition to the above programs, the memory 151 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.
[0157] The memory 151 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0158] A processor 152, coupled to a memory 151, executes a program stored in the memory 151 for:
[0159] Obtain original video information, where the original video information includes the sign language video of a first user;
[0160] Perform visual recognition on the sign language video to obtain a first sign language text sequence;
[0161] Perform image processing on the sign language video to obtain first feature information, where the first feature information is used to characterize the first tone of the sign language;
[0162] Generate a first natural language text with the first tone according to the first sign language text sequence and the first feature information.
[0163] Furthermore, as Figure 15 shown, the electronic device may further include: other components such as a communication component 153, a power supply component 154, an audio component 155, a display 156, etc. Figure 15 Only some components are schematically shown in Figure 15 the figure, which does not mean that the electronic device only includes
[0164] The electronic device described in the embodiments of the present disclosure may be the terminal or the server in the above method embodiments. When the electronic device is a terminal, the electronic device may further include a video acquisition module, where the video acquisition module is used to acquire the original video information of a hearing-impaired person, and the audio component 155 is used to acquire the original audio information of a normal person. The processor 152 may generate a first natural language text with the first tone according to the original video information, or generate a second sign language text sequence and the category information of the second tone according to the original audio information. The specific process is not described in detail here. The communication component 153 may send the original video information or the original audio information to the server, or send the first natural language text with the first tone to the server, or may also send the second sign language text sequence and the category information of the second tone to the server.
[0165] When the electronic device is a server, the communication component 153 may receive the original video information of a hearing-impaired person or the original audio information of a normal person from the terminal. The processor 152 may generate a first natural language text with the first tone according to the original video information, or generate a second sign language text sequence and the category information of the second tone according to the original audio information. The specific process is not described in detail here. In addition, the communication component 153 may also send the first natural language text with the first tone to the terminal, or send the second sign language text sequence and the category information of the second tone to the terminal.
[0166] The communication component 153 is configured to facilitate communication between the electronic device and other devices in a wired or wireless manner. The electronic device can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 153 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 153 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0167] The power supply component 154 provides power for various components of the electronic device. The power supply component 154 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device.
[0168] The audio component 155 is configured to output and / or input audio signals. For example, the audio component 155 includes a microphone (MIC). When the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory 151 or transmitted via the communication component 153. In some embodiments, the audio component 155 further includes a speaker for outputting audio signals.
[0169] The display 156 includes a screen, and the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations.
[0170] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium for processing multimedia information, on which a computer program is stored. The computer program is executed by a processor to implement the multimedia information processing method described in the above embodiments.
[0171] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.
[0172] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for processing multimedia information, wherein, The method includes: Obtaining original video information, where the original video information includes the sign language video of the first user; Performing visual recognition on the sign language video to obtain a first sign language text sequence; Performing image processing on the sign language video to obtain first feature information, where the first feature information is used to characterize the first tone of the sign language; Determining a first numerical representation corresponding to the first sign language text sequence and a first spatio-temporal representation corresponding to the first feature information; Inputting the first numerical representation and the first spatio-temporal representation into a first preset model, and using the first preset model to generate a first natural language text with the first tone and category information of the first tone; Wherein, the first preset model includes a first encoder, a second encoder, a decoder, and a tone classifier; using the first preset model to generate a first natural language text with the first tone includes: Processing the first numerical representation through the first encoder to obtain a first encoding result; Processing the first spatio-temporal representation through the second encoder to obtain a second encoding result; Inputting the first encoding result and the second encoding result into the decoder, and generating a first natural language text with the first tone by the decoder; Inputting the second encoding result and the first natural language text with the first tone into the tone classifier, and generating category information of the first tone by the tone classifier.
2. The method according to claim 1, wherein The first feature information includes at least one of the following: Facial features corresponding to the face region in the sign language video; The expression category of the first user in the sign language video; Facial key points of the first user in the sign language video.
3. The method according to claim 1, wherein The method further includes: Obtaining the original audio information of the second user; Performing speech recognition on the original audio information to obtain a second natural language text; Determining second feature information according to the original audio information and the second natural language text, where the second feature information is used to characterize the second tone of the original audio information; Generating a second sign language text sequence and category information of the second tone according to the second natural language text and the second feature information.
4. The method according to claim 3, wherein The method further includes: Generating target video information according to the second sign language text sequence and the category information of the second tone, where the target video information includes a virtual character, and the category information of the second tone is used to determine at least one of the facial features, expression category, and facial key points of the virtual character.
5. The method according to claim 3, wherein, The second feature information includes at least one of the following: Numerical feature information of the original audio information; The speaking speed and intonation of the second user determined according to the original audio information; The tone classification result corresponding to the second natural language text.
6. The method according to claim 3, wherein, Generating a second sign language text sequence and category information of the second tone according to the second natural language text and the second feature information includes: Determining a second numerical representation corresponding to the second natural language text and a second spatio-temporal representation corresponding to the second feature information; Input the second numerical representation and the second spatio-temporal representation into a second preset model, and use the second preset model to generate the second sign language text sequence and the category information of the second tone.
7. The method according to claim 6, wherein The second preset model includes a first encoder, a second encoder, a decoder, and a tone classifier; Using the second preset model to generate the second sign language text sequence and the category information of the second tone includes: Process the second numerical representation through the first encoder to obtain a third encoding result; Process the second spatio-temporal representation through the second encoder to obtain a fourth encoding result; Input the third encoding result and the fourth encoding result into the decoder, and the decoder generates the second sign language text sequence; Input the fourth encoding result and the second sign language text sequence into the tone classifier, and the tone classifier generates the category information of the second tone.
8. A processing device for multimedia information, wherein, The device includes: A first acquisition module for acquiring original video information, where the original video information includes the sign language picture of a first user; A visual recognition module for performing visual recognition on the sign language picture to obtain a first sign language text sequence; An image processing module for performing image processing on the sign language picture to obtain first feature information, where the first feature information is used to characterize the first tone of the sign language; A first generation module for determining a first numerical representation corresponding to the first sign language text sequence and a first spatio-temporal representation corresponding to the first feature information; inputting the first numerical representation and the first spatio-temporal representation into a first preset model, and using the first preset model to generate a first natural language text with the first tone and the category information of the first tone; The first preset model includes a first encoder, a second encoder, a decoder, and a tone classifier; the first generation module is specifically configured to process the first numerical representation through the first encoder to obtain a first encoding result; process the first spatio-temporal representation through the second encoder to obtain a second encoding result; input the first encoding result and the second encoding result into the decoder, and the decoder generates a first natural language text with the first tone; input the second encoding result and the first natural language text with the first tone into the tone classifier, and the tone classifier generates the category information of the first tone.
9. The apparatus according to claim 8, wherein, The device further includes: A second acquisition module for acquiring the original audio information of a second user; A speech recognition module for performing speech recognition on the original audio information to obtain a second natural language text; A determination module for determining second feature information according to the original audio information and the second natural language text, where the second feature information is used to characterize the second tone of the original audio information; A second generation module for generating a second sign language text sequence and the category information of the second tone according to the second natural language text and the second feature information.
10. An electronic device, wherein, Includes: A memory; A processor; And A computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1-7.
11. A computer-readable storage medium for processing multimedia information, having a computer program stored thereon, wherein, When the computer program is executed by the processor, it implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Interaction method and device, terminal equipment and storage medium
CN110807388A
Weak supervision neural network sign language recognition method based on multi-layer time sequence attention fusion mechanism
CN113537024A
Action generation method and device
CN113792537A