Method and equipment for displaying digital human sign language

By generating and adjusting the display speed of digital human sign language gestures in real time through the device, the problem of hearing-impaired users being unable to obtain sign language translations in real time has been solved, thus improving the television program viewing experience.

CN121967769APending Publication Date: 2026-05-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-10-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Most television programs only display subtitles and do not include sign language, which reduces the viewing experience for users with hearing impairments.

Method used

The device generates digital human sign language gestures in real time, and uses the processing power of other devices to extract human voice and generate a sequence of action parameters. Combined with the audio-visual synchronization control module, the display speed is adjusted to ensure that the sign language translation is synchronized with the audio.

Benefits of technology

It enables real-time sign language translation services for hearing-impaired users, improves user experience, avoids sign language translation delays, and adapts to the playback of different audio data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967769A_ABST
    Figure CN121967769A_ABST
Patent Text Reader

Abstract

The invention provides a method and equipment for displaying digital human sign languages. The method can be applied to a first device, and the method can comprise the steps that first audio data are acquired, and the first audio data comprise audio data of human voice; sending the first audio data to a second device; receiving a first action parameter sequence of the digital human sign language sent by the second equipment; the first action parameter sequence is rendered according to the display speed, a multi-frame image of the digital sign language is obtained, the display speed is related to the completion time of rendering a second action parameter sequence corresponding to second audio data by the first device, and the second audio data is audio data of human voice obtained before the first device obtains the first audio data. In the technical scheme, the device can generate the digital human sign language actions in real time according to the human voice when the user watches a video or a television program, and plays the digital human sign language actions in real time, so as to provide sign language translation service for the hearing impaired user in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technology, and more specifically, to a method and apparatus for displaying digital human sign language. Background Technology

[0002] People with hearing impairments also need to watch television. For them, sign language is more convenient than subtitles. However, most television programs only display subtitles and do not include sign language, thus reducing the viewing experience for people with hearing impairments.

[0003] Therefore, how to display sign language on devices in real time and provide sign language translation services for users with hearing impairments has become a technical problem that needs to be solved. Summary of the Invention

[0004] This application provides a method and device for displaying digital human sign language. In this technical solution, the device can generate and play digital human sign language actions in real time based on human voice while a user is watching a video or television program, aiming to provide real-time sign language translation services for hearing-impaired users.

[0005] A first aspect provides a method for displaying digital human sign language, applied to a first device. The method includes: acquiring first audio data, the first audio data being audio data including human voice that is being played in the first device; sending the first audio data to a second device; receiving a first sequence of motion parameters of digital human sign language sent by the second device, wherein the first sequence of motion parameters is related to the audio data of human voice, and each motion parameter in the first sequence of motion parameters corresponds to a frame of digital human sign language image; rendering the first sequence of motion parameters according to a display speed to obtain multiple frames of digital human sign language images, wherein the display speed is related to the completion time of the first device rendering the second sequence of motion parameters corresponding to the second audio data, and the second audio data is the audio data of human voice acquired by the first device before acquiring the first audio data.

[0006] It should be understood that the specific method of displaying multiple frames of digital sign language images is not limited in the embodiments of this application. For example, the first device may display the multiple frames of images uniformly after rendering them; or, the first device may display the rendered images while rendering the multiple frames of images, that is, rendering and displaying simultaneously.

[0007] Based on the embodiments of this application, for any video or television program audio data played in the first device, the first device can display the corresponding sign language translation in real time, thus expanding the application scope of this technical solution. Furthermore, the first device can utilize the processing power of the second device to separate audio data and generate a first sequence of motion parameters for digital human sign language, thereby providing real-time sign language translation services. When rendering and displaying the digital human sign language image, the first device can determine the display speed of the current rendering of the first motion parameter sequence by combining the time when the rendering of the previous person's audio data is completed, thereby ensuring audio-visual synchronization and avoiding sign language translation delays.

[0008] In some alternative solutions, the first device may also convert the first audio data into text and send the text to the second device, which may then generate a first sequence of action parameters based on the text. Alternatively, the first device may extract human voice audio data from the first audio data, convert the human voice audio data into text, and send the text to the second device, which may then generate a first sequence of action parameters based on the text.

[0009] In this way, the technical solution helps to reduce the size of the data transmitted from the first device to the second device.

[0010] In some implementations, before rendering the first action parameter sequence according to the display speed, the method further includes: determining the display speed of the image displaying digital human sign language based on the completion time and the reception time of the first device receiving the first action parameter sequence.

[0011] Based on the embodiments of this application, the first device can determine the display speed of rendering the first action parameter sequence according to the completion time of the second action parameter sequence corresponding to the rendering of the second audio data and the reception time of the first action parameter sequence. This allows the determined display speed to be related to the rendering speed of the previous human voice audio data, thereby controlling the display speed to achieve audio-visual synchronization.

[0012] In some implementations, the first device determines the display speed according to the following formula:

[0013]

[0014] Where α>0, D1 represents the display speed, FPS min This indicates the preset minimum display speed in FPS. max This indicates the preset maximum display speed, T1 represents the completion time, T2 represents the reception time, and Tm represents the speed according to FPS. min The time taken to render the first action parameter sequence.

[0015] It should be understood that min(a, b) can represent taking the minimum value between a and b. max(a, b) can represent taking the maximum value between a and b.

[0016] In some implementations, when the completion time is the same as the reception time, the display speed of the digital human sign language image is determined based on the completion time and the reception time of the first device receiving the first action parameter sequence, including: determining the first preset display speed as the display speed.

[0017] For example, the first preset display speed is the normal speed of sign language translation. For example, the first preset display speed is 30 frames per second or 60 frames per second, etc., and this application embodiment is not limited to this.

[0018] Based on the embodiments of this application, when the completion time and the reception time are the same, the current sign language translation speed is almost equal to the audio playback speed. The first device can display the sign language images of several people at the normal sign language translation speed, which can meet the user's needs and achieve the effect of audio-visual synchronization.

[0019] In some implementations, when the completion time is after the reception time, the first display speed of the image displaying digital human sign language determined by the first device is greater than the second display speed, wherein the second display speed is the display speed of the image displaying digital human sign language determined when the completion time is before the reception time.

[0020] Based on the embodiments of this application, when the completion time is after the reception time, the display speed determined by the first device is the first display speed, and when the completion time is before the reception time, the display speed determined by the first device is the second display speed, and the first display speed is greater than the second display speed.

[0021] In this way, when the current sign language translation lags behind the speech, the speed of sign language translation can be increased, and when the current sign language translation is not lagging behind the speech, the speed of sign language translation can be appropriately reduced in order to maintain the effect of audio-visual synchronization as much as possible.

[0022] A second aspect provides a method for displaying digital human sign language, applied to a second device, the method comprising: receiving first audio data sent by a first device, the first audio data including audio data of a human voice; generating a first sequence of motion parameters for digital human sign language based on the audio data of the human voice, each motion parameter in the first sequence of motion parameters corresponding to a frame of digital human sign language image; and sending the first sequence of motion parameters to the first device.

[0023] Based on the embodiments of this application, the second device can receive the first audio data sent by the first device, generate a first action parameter sequence of digital human sign language based on the audio data of human voice in the first audio data, and send the first action parameter sequence to the first device.

[0024] In this way, the second device can generate a sequence of motion parameters for digital human sign language based on the audio data of the human voice, thereby reducing the performance requirements on the first device.

[0025] In some implementations, the method further includes extracting audio data of the human voice from the first audio data before generating the first sequence of action parameters for digital human sign language based on the audio data of the human voice.

[0026] Based on the embodiments of this application, by extracting the audio data of human voice from the first audio data, the audio data can be made to include only human voice, filtering out noise, thereby making the first action parameter sequence of digital human sign language more accurate.

[0027] In some implementations, generating a first sequence of action parameters for digital human sign language based on the audio data of a human voice includes: converting the audio data of the human voice into first text; and generating a first sequence of action parameters based on the first text.

[0028] In other examples, the second device can also directly generate a first sequence of action parameters based on the audio data of the human voice, without converting it to text.

[0029] A third aspect provides a method for displaying digital human sign language, applied to a first device. The method includes: acquiring first audio data, which is audio data including human voice being played in the first device; extracting audio data of human voice from the first audio data; generating a first sequence of motion parameters for digital human sign language based on the audio data of human voice, where each motion parameter in the first sequence corresponds to a frame of digital human sign language image; and rendering the first sequence of motion parameters according to a display speed to obtain multiple frames of digital human sign language images, wherein the display speed is related to the completion time of the first device rendering the second sequence of motion parameters corresponding to second audio data, and the second audio data is audio data of human voice acquired by the first device before acquiring the first audio data.

[0030] It should be understood that in the embodiments of this application, the first device can display the multiple frames of images uniformly after rendering them; or, the first device can display the rendered image while rendering the multiple frames of images, that is, rendering and displaying at the same time.

[0031] Based on the embodiments of this application, when rendering and displaying digital human sign language images, the first device can determine the display speed of the first action parameter sequence being rendered in this instance by combining the time when the rendering of the previous person's audio data is completed. This ensures audio-visual synchronization and avoids delays in sign language translation. Furthermore, for any video or television program audio data played on the first device, the first device can display the corresponding sign language translation in real time, expanding the application scope of this technical solution.

[0032] In some embodiments, before rendering the first sequence of motion parameters according to the display speed, the method further includes: determining the display speed of the image displaying digital human sign language based on the completion time and the generation time of the first sequence of motion parameters generated by the first device.

[0033] Based on the embodiments of this application, the first device can determine the display speed of rendering the first action parameter sequence according to the completion time of the second action parameter sequence corresponding to the rendering of the second audio data and the generation time of the first action parameter sequence. This allows the determined display speed to be related to the rendering speed of the previous human voice audio data, thereby controlling the display speed to achieve audio-visual synchronization.

[0034] In some implementations, the first device determines the display speed according to the following formula:

[0035]

[0036] Where α>0, D1 represents the display speed, FPS min This indicates the preset minimum display speed in FPS. max This indicates the preset maximum display speed, T1 represents the completion time, T2 represents the generation time, and Tm represents the speed according to FPS. min The time taken to render the first action parameter sequence.

[0037] In some implementations, when the completion time is the same as the generation time, the display speed of the image displaying digital human sign language is determined based on the completion time and the generation time of the first action parameter sequence generated by the first device, including: determining the first preset display speed as the display speed.

[0038] For example, the first preset display speed is the normal speed of sign language translation. For example, the first preset display speed is 30 frames per second or 60 frames per second, etc., and this application embodiment is not limited to this.

[0039] Based on the embodiments of this application, when the completion time is the same as the generation time, the current sign language translation speed is almost equal to the audio playback speed. The first device can display the sign language images of several people at the normal sign language translation speed, which can meet the user's needs and achieve the effect of audio-visual synchronization.

[0040] In some implementations, when the completion time is after the generation time, the first display speed of the image displaying digital human sign language, determined by the first device, is greater than the second display speed, wherein the second display speed is the display speed of the image displaying digital human sign language, determined when the completion time is before the generation time.

[0041] Based on the embodiments of this application, when the completion time is after the generation time, the display speed determined by the first device is the first display speed, and when the completion time is before the generation time, the display speed determined by the first device is the second display speed, and the first display speed is greater than the second display speed.

[0042] In this way, when the current sign language translation lags behind the speech, the speed of sign language translation can be increased, and when the current sign language translation is not lagging behind the speech, the speed of sign language translation can be appropriately reduced in order to maintain the effect of audio-visual synchronization as much as possible.

[0043] In some implementations, the method further includes extracting audio data of the human voice from the first audio data before generating the first sequence of action parameters for digital human sign language based on the audio data of the human voice.

[0044] In this way, by extracting the human voice audio data from the first audio data, the audio data can be made to include only the human voice, filtering out the noise, thereby making the first action parameter sequence for generating digital human sign language more accurate.

[0045] In some implementations, generating a first sequence of action parameters for digital human sign language based on the audio data of a human voice includes: converting the audio data of the human voice into first text; and generating a first sequence of action parameters based on the first text.

[0046] In other examples, the first device can also directly generate a sequence of first action parameters based on the audio data of the human voice, without converting it to text.

[0047] A fourth aspect provides a first device comprising: one or more processors; one or more memories; said one or more memories storing one or more programs that, when executed by said one or more processors, cause a method for displaying digital human sign language as described in the first or third aspect and any possible implementation thereof to be performed.

[0048] A fifth aspect provides a second device comprising: one or more processors; one or more memories; said one or more memories storing one or more programs that, when executed by said one or more processors, cause a method for displaying digital human sign language as described in the second aspect and any possible implementation thereof to be performed.

[0049] A sixth aspect provides a system for displaying digital human sign language, including a first device as described in the fourth aspect and a second device as described in the fifth aspect.

[0050] A seventh aspect provides a chip including a processor and a communication interface for receiving signals and transmitting the signals to the processor, the processor processing the signals such that a method for displaying digital human sign language as described in the first to third aspects and any possible implementation thereof is executed.

[0051] The eighth aspect provides a readable storage medium storing instructions that, when executed on a device, cause a method for displaying digital human sign language as described in the first to third aspects and any possible implementation thereof to be performed.

[0052] A ninth aspect provides a program product comprising program code that, when run on a device, causes a method for displaying digital human sign language as described in the first to third aspects and any possible implementation thereof to be executed. Attached Figure Description

[0053] Figure 1 This is a schematic diagram of a device displaying digital human sign language, provided in an embodiment of this application.

[0054] Figure 2 This is a schematic block diagram of a device provided in an embodiment of this application.

[0055] Figure 3 This is a schematic flowchart illustrating another device for displaying digital human sign language, provided in an embodiment of this application.

[0056] Figure 4 This is a schematic flowchart illustrating another device for displaying digital human sign language, provided in an embodiment of this application.

[0057] Figure 5 This is a schematic flowchart illustrating a method for displaying digital human sign language provided in an embodiment of this application.

[0058] Figure 6 This is a schematic block diagram of a device provided in an embodiment of this application. Detailed Implementation

[0059] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0060] The methods described in this application embodiment can be applied to smart TVs, smart screens, smartphones, tablets, laptops, personal computers (PCs), wearable devices, foldable devices, Internet of Things (IoT) devices, etc.

[0061] Before introducing the technical solution of this application, the following is a brief introduction to some of the technical terms that may be involved in this application.

[0062] Digital human: A virtual character generated or simulating human characteristics and behaviors through a device, which can be used in scenarios such as games, virtual reality, and smart assistants. In the embodiments of this application, the device (such as a smart TV) can display the digital human and provide sign language translation in real time, providing sign language translation services for hearing-impaired users watching TV.

[0063] People with hearing impairments also need to watch television. For them, sign language is more convenient than subtitles. However, most television programs only display subtitles and do not include sign language, thus reducing the viewing experience for people with hearing impairments.

[0064] In view of this, embodiments of this application provide a method and apparatus for displaying digital human sign language. In this technical solution, the apparatus can generate and play digital human sign language actions in real time based on human voice while a user is watching a video or television program, aiming to provide real-time sign language translation services for hearing-impaired users.

[0065] The following will combine Figures 1 to 4 The technical solutions provided in the embodiments of this application are introduced.

[0066] For example, Figure 1 This is a schematic diagram illustrating a device displaying digital human sign language, provided in an embodiment of this application. See also... Figure 1 In (a), the graphical user interface (GUI) is the display interface 115 of device 100. The display interface 115 is used to display the TV program, subtitles, etc. that device 100 is playing, and the display interface 115 may also include a digital human display area 1151.

[0067] The digital human display area 1151 is used to display a digital human that is used to provide real-time sign language translation.

[0068] It should be understood that the specific location of the digital human display area 1151 within the display interface 115 is not limited in the embodiments of this application. For example, the digital human display area 1151 may be displayed in the lower right corner, upper right corner, lower left corner, etc. of the display interface 115.

[0069] In other examples, the digital human display area 1151 can also be displayed side-by-side with the television program being played, that is, the digital human display area 1151 and the television program are displayed in a split-screen format. For example, the display area of ​​the television program occupies a preset proportion of the screen size (such as three-quarters), and the digital human display area 1151 occupies the remaining proportion (such as one-quarter) of the screen size.

[0070] In some optional examples, device 100 provides a setting option to enable or disable sign language translation. When device 100 detects that the user has enabled this setting option, device 100 may display the digital human display area 1151 on the display interface. When device 100 detects that the user has disabled this setting option, device 100 may not display the digital human display area 1151 on the display interface.

[0071] In this embodiment, device 100 can acquire audio data of a TV program being played, extract audio data of human voice from the audio data, generate a sequence of motion parameters for digital human sign language based on the audio data of human voice, render the sequence of motion parameters to generate multiple frames of images, and display the multiple frames of images in the aforementioned digital human display area 1151 to provide real-time sign language translation.

[0072] In some implementations, device 100 may not be able to generate and render the digital human sign language motion parameter sequence in real time, which may result in slower sign language translation and a delay in the translation process, affecting the user experience. In this case, device 100 can also utilize the processing power of other devices to complete the process of extracting human voice and generating the digital human sign language motion parameter sequence.

[0073] See Figure 1 In (b) of the above, device 100 can also send the acquired audio data to server 200. Server 200 extracts the audio data of the human voice from the audio data, generates a sequence of motion parameters for digital human sign language based on the audio data of the human voice, and sends the sequence of motion parameters for digital human sign language to device 100. Device 100 only needs to render and display the sequence of motion parameters, thereby reducing the performance requirements of device 100.

[0074] It should be understood that the server 200 can be a cloud server or a local server.

[0075] The server 200 can also be a device with strong processing capabilities in the user's home, such as a smart host.

[0076] In this way, with the help of the processing power of other devices, device 100 can display digital human sign language in real time, avoiding delays in sign language translation.

[0077] Based on the embodiments of this application, device 100 can obtain corresponding audio data from the system level and display the corresponding digital human sign language. Thus, device 100 can provide corresponding sign language translation services for any television program played by device 100.

[0078] For example, Figure 2This is a schematic block diagram of a device provided in an embodiment of this application. Figure 2 As shown, the device 100 may include an audio acquisition module 110, an audio separation module 120, a speech recognition module 130, a sign language driving service module 140, an audio-visual synchronization control module 150, a rendering module 160, and a display module 170.

[0079] The audio acquisition module 110 can be used to acquire audio data from the device 100.

[0080] For example, when device 100 is playing a video, the audio acquisition module 110 can be used to acquire the audio data A of the video. This audio data A may contain background music (such as accompaniment sounds, instrument sounds), human voices, animal sounds, etc., and digital sign language translation uses human voices. Therefore, device 100 needs to separate this audio data A to extract the human voice audio data from it.

[0081] The audio separation module 120 can be used to separate the audio data A acquired by the audio acquisition module 110 to extract the audio data of human voices.

[0082] For example, the audio separation module 120 can use signal processing to separate audio data A, such as extracting the audio data of the human voice based on the difference in frequency domain characteristics between the human voice signal in audio data A and other signals.

[0083] For example, the audio separation module 120 can also use a deep learning model (such as a convolutional neural network or a recurrent neural network) to predict the audio data of the human voice from the audio data A. For example, the deep learning model takes the audio data A or the features of the audio data A in the frequency domain, or a combination of the two features as input, and predicts the output audio data of the human voice.

[0084] After extracting the audio data of the human voice, the audio separation module 120 can send the audio data of the human voice to the speech recognition module 130 at preset intervals. Correspondingly, the speech recognition module 130 receives the audio data of the human voice sent by the audio separation module 120. For example, the preset interval can be 30 milliseconds or 50 milliseconds, etc.

[0085] It is understandable that if the audio separation module 120 does not extract the audio data of human voice, such as the above audio data A which does not include human voice, then the audio separation module 120 will not send the audio data to the speech recognition module 130.

[0086] The speech recognition module 130 can be used to convert the audio data of human voices separated by the audio separation module 120 into text content. When the speech recognition module 130 recognizes a sentence of text content, it can send the text content to the sign language driven service module 140.

[0087] It should be understood that a sentence here can be interpreted as the text content between two adjacent punctuation marks.

[0088] For example, if the audio separation module sends the audio content "Good morning, today's weather is cloudy" to the speech recognition module 130, and the speech recognition module 130 recognizes the text content "Good morning", it can send the text content "Good morning" to the sign language driving service module 140. Simultaneously, if the speech recognition module 130 continues to recognize the text content "Today's weather is cloudy", it can send the text content "Today's weather is cloudy" to the sign language driving service module 140.

[0089] The sign language driven service module 140 can be used to convert received text content into a sequence of motion parameters for digital human sign language.

[0090] It should be understood that one action parameter of digital human sign language can correspond to one frame of image of digital human to be displayed later, and one sequence of action parameters can correspond to multiple frames of images. The continuous action composed of the digital human sign language actions in these multiple frames of images can correspond to a sentence of text content.

[0091] In one example, the sign language driven service module 140 can convert words in text content into corresponding sign language words, and one sign language word can correspond to one or more sign language actions, each sign language action corresponding to one action parameter. By smoothly concatenating all the sign language words together, the sign language driven service module 140 can obtain a sequence of action parameters for digital human sign language.

[0092] In other examples, the sign language-driven service module 140 may also include a neural network model that converts text content into a sequence of motion parameters for digital human sign language. For example, the text content can be used as input to the neural network model, and the sequence of motion parameters can be used as output. The sign language-driven service module 140 can then convert text content into a sequence of motion parameters for digital human sign language using this neural network model.

[0093] For example, for the text content "Good morning", the sign language driven service module 140 can generate the corresponding action parameter sequence Q1 based on the text content "Good morning". For the text content "Today's weather is cloudy", the sign language driven service module 140 can generate the corresponding action parameter sequence Q2 based on the text content "Today's weather is cloudy".

[0094] It should be understood that other text content received by the sign language driven service module 140 can be converted into the corresponding digital human sign language action parameter sequence in the above manner.

[0095] Because digital humans need to use hand rotations to convert text into sign language, and the time required for sign language translation of the same text at normal speed is longer than the time required for speech playback at normal speed, a delay in sign language translation is likely to occur. Therefore, the audio-visual synchronization control module 150 can be used to control the speed of sign language translation in order to achieve audio-visual synchronization.

[0096] The audio-visual synchronization control module 150 can be used to control the display speed of multiple frames of images corresponding to the action parameter sequence of digital human sign language played by device 100, in order to achieve matching between the voice of the TV program and the displayed digital human sign language actions, realize audio-visual synchronization, and avoid the phenomenon of sign language translation delay.

[0097] For example, this display speed corresponds to 30 frames per second (FPS), or 30 FPS. This can be understood as device 100 displaying 30 frames of the digital human image per second, or device 100 displaying the digital human image at a display speed of 30 frames per second. It can also be understood as the refresh rate of device 100's display screen.

[0098] The audio-visual synchronization control module 150 can send the sequence of motion parameters of digital human sign language and the display speed of the corresponding multi-frame images to the rendering module 160.

[0099] The rendering module 160 can be used to render the received sequence of digital human sign language motion parameters according to the display speed, obtaining multiple frames of digital human images corresponding to the motion parameter sequence. Then, the rendering module 160 can send the multiple frames of digital human images to the display module 170 for display.

[0100] The display module 170 can be used to display multiple frames of digital human images to provide digital human sign language translation.

[0101] It should be understood that the names of the above modules are merely illustrative. In other examples, the above modules may have other names, which are not limited in the embodiments of this application.

[0102] It should be understood that multiple modules among modules 110 to 170 described above can also be combined into one module. For example, audio acquisition module 110 and audio separation module 120 can be combined into one audio module, which can have the functions of audio acquisition module 110 and audio separation module 120.

[0103] It should also be understood that the audio separation module 120, speech recognition module 130, and sign language driven service module 140 are optional modules. In other examples, the device 100 may also exclude the audio separation module 120, speech recognition module 130, and sign language driven service module 140.

[0104] The following will combine Figure 3-4 This application provides a detailed description of the technical solutions for displaying digital human sign language in its embodiments.

[0105] For example, Figure 3 This is a schematic flowchart illustrating another device for displaying digital human sign language, provided in an embodiment of this application. Figure 3 As shown, the method 300 can be applied to the first device, and the method 300 may include steps 310 to 370.

[0106] It should be understood that the first device can be device 100 mentioned above.

[0107] 310, The first device acquires audio data A1.

[0108] It should be understood that the audio data A1 may be audio data from a television program or video being played on the first device.

[0109] For example, the audio data A1 may contain background music (such as accompaniment, instrument sounds), human voices, animal sounds, etc.

[0110] The audio data A1 can be obtained by the first device from the system level. That is, the first device can obtain the corresponding audio data A1 as long as audio exists, regardless of whether the TV program or video is played.

[0111] In this way, the first device can provide sign language translation services for any TV program or video played on the first device.

[0112] 320, The first device extracts the human voice audio data B1 from the audio data A1.

[0113] For example, since the first audio data A1 may contain other types of sounds besides human voices, the first device needs to extract the human voices from the audio data A1.

[0114] As mentioned above Figure 2 As mentioned above, different types of sounds have different characteristics. After converting audio data A1 into a frequency domain signal, the audio data B1 of the human voice can be extracted based on the characteristic distribution of the audio data.

[0115] 330, the first device converts the audio data B1 into text content C1.

[0116] For example, the first device can convert audio data B1 into text content C1 through speech recognition.

[0117] It should be understood that step 330 is an optional step, and in some examples, step 330 may not be performed.

[0118] 340, the first device generates a sequence of motion parameters for digital human sign language 1 based on the text content C1, wherein each motion parameter corresponds to a frame of digital human sign language image.

[0119] For example, the first device can convert the words in the text content C1 into corresponding sign language words, and one sign language word can correspond to one or more sign language actions, each sign language action corresponding to one action parameter. By smoothly concatenating all the sign language words together, the first device can obtain the action parameter sequence 1 of the digital human sign language.

[0120] For example, the first device may further include a neural network model that converts text content into a sequence of motion parameters for digital human sign language. For instance, the text content C1 can be used as input to the neural network model, and the motion parameter sequence 1 can be used as output. Through this neural network model, the first device can convert the text content C1 into the motion parameter sequence 1 of digital human sign language.

[0121] In this embodiment of the application, each motion parameter in the motion parameter sequence 1 can correspond to a frame of digital human sign language image. For example, the motion parameter sequence 1 can include motion parameter 1, motion parameter 2, ..., motion parameter n. Among them, motion parameter 1 can correspond to image 1, motion parameter 2 can correspond to image 2, ..., motion parameter n can correspond to image n.

[0122] In some implementations, each motion parameter includes a data matrix that characterizes the rotation of each joint of the digital human. For example, if the digital human has m joints, and the rotation of each joint is characterized by quaternions, then the data matrix is ​​an m*4 matrix.

[0123] Action parameter 1 corresponds to image 1, which can be understood as image 1 including the sign language action composed of the rotation of m joints of a digital human, represented by the data matrix in action parameter 1.

[0124] In this way, by displaying images 1 to n, the device can fully display all the sign language actions corresponding to the text content C1, providing the user with a complete sign language translation of the text content C1.

[0125] It should be understood that if step 330 is not performed, in step 340, the first device can generate a sequence of motion parameters for digital human sign language based on the audio data B1.

[0126] For example, the first device may include a neural network model that can directly generate a sequence of motion parameters 1 for the corresponding digital human sign language based on the input audio data B1.

[0127] 350, the first device determines the display speed D1 of the image displaying digital human sign language, wherein the display speed D1 is related to the completion time of the motion parameter sequence 2 corresponding to the audio data E1 rendered by the first device.

[0128] Among them, audio data E1 is the preceding audio data of audio data B1, and audio data E1 is also human voice audio data.

[0129] For example, if the audio data includes the content "Good morning, today's weather is cloudy", then the audio data B1 can be the audio data corresponding to "Today's weather is cloudy", and the audio data E1 can be the audio data corresponding to "Good morning".

[0130] It should be understood that the explanation of the displayed speed D1 can be found in the previous text. Figure 2 The relevant description of display speed.

[0131] In this embodiment of the application, the maximum value of the display speed D1 is a preset value FPS. max The minimum value of the display speed D1 is the preset value in FPS. min For example, the FPS max For 120 or 140, FPS min The specific values ​​of the maximum and minimum display speed D1 are not limited to 30 or 60 in this embodiment of the application.

[0132] By setting the maximum value of display speed D1 in FPS max This avoids sign language gestures being too fast for users to understand. The minimum FPS for display speed D1 is set. min This can avoid the phenomenon of sign language translation being interrupted due to slow sign language movements.

[0133] For example, audio data E1 corresponds to motion parameter sequence 2, and audio data B1 corresponds to motion parameter sequence 1. For instance, the first device first obtains motion parameter sequence 2, renders each motion parameter in motion parameter sequence 2 at a preset speed V1, generates multi-frame images of digital human sign language, and displays the multi-frame images.

[0134] Assuming the first device completes rendering motion parameter sequence 2 at time T1, and generates digital sign language motion parameter sequence 1 based on text content C1 at time T2, if the length of motion parameter sequence 1 (i.e., the number of motion parameters included) is L, then the first device operates at the minimum FPS of the aforementioned display speed D1. minThe rendering of action parameter sequence 1 takes the longest time, and this time Tm can be expressed by the following formula:

[0135]

[0136] The display speed D1 can be expressed by the following formula:

[0137]

[0138] Where α>0, the value of α can be used to control the size of D1. For example, when the current sign language action is delayed compared to the speech, the size of D1 can be adjusted by adjusting the size of α, thereby speeding up the rendering speed of action parameter sequence 1 and reducing the delay of sign language action.

[0139] For example, when the difference between T1 and T2 is large, it means that the first device has generated the next action parameter sequence 1, while the previous action parameter sequence 2 has not yet been rendered. That is, there is a large delay in the sign language action. Therefore, the first device can set a larger value for α to speed up the rendering speed of the action parameter sequence 1, reduce the delay in the sign language action, and maintain the effect of audio-visual synchronization.

[0140] When the difference between T1 and T2 is small, it indicates that the sign language movement delay is small. Therefore, the first device can set a smaller value for α to maintain the effect of audio-visual synchronization.

[0141] When the difference between T1 and T2 is close to 0, it means that there is almost no delay in the sign language action. If the minimum display speed is the normal display speed of the sign language action, then the first device can render the action parameter sequence 1 at the minimum display speed to maintain the effect of audio-visual synchronization.

[0142] When the difference between T1 and T2 is less than 0, it means that there is no delay in the sign language action. For example, the previous sign language action has been displayed and after a certain period of time, the first device receives the current sign language action, indicating that there is an interval between the two sentences. In this case, the first device can render the action parameter sequence 1 at the minimum display speed to maintain the effect of audio-visual synchronization.

[0143] In this way, the rendering speed of the first device for the motion parameter sequence is between the set minimum and maximum values ​​and can be dynamically adjusted, so that the rendering speed of the motion parameter sequence is within a suitable range, achieving the effect of audio-visual synchronization.

[0144] 360, the first device renders motion parameter sequence 1 according to the display speed D1 to obtain multi-frame images of digital human sign language.

[0145] After determining the display speed D1, the first device can render the motion parameter sequence 1 according to the display speed D1 to obtain multi-frame images of digital human sign language.

[0146] For example, if the motion parameter sequence 1 includes L motion parameters, then the multi-frame image is an L-frame image.

[0147] It should be understood that the first device can store the modeling data of the digital human, and during rendering, the first device can simply load the modeling data of the digital human.

[0148] 370, the first device displays multiple frames of digital human sign language images.

[0149] For example, see Figure 1 In (a), the first device can display multiple frames of digital human sign language in the digital human display area 1151 to provide users with real-time sign language translation services.

[0150] Based on the embodiments of this application, the first device can provide sign language translation services for any video or television program played by the first device. Furthermore, the first device can adjust the speed of sign language translation by controlling the speed of rendering the motion parameter sequence to achieve audio-visual synchronization and provide a better user experience.

[0151] It should be understood that the specific method of displaying multiple frames of digital sign language images is not limited in the embodiments of this application. For example, the first device may display the multiple frames of images uniformly after rendering them; or, the first device may display the rendered images while rendering the multiple frames of images, that is, rendering and displaying simultaneously.

[0152] In some cases, Figure 3 In this technical solution, the first device needs to perform real-time speech recognition and generate corresponding action parameter sequences, placing high demands on its performance. The first device may exhibit slow sign language translation speeds, causing delays and impacting the user experience. In such cases, the first device can leverage the processing power of other devices to complete the process of extracting audio data of the human voice and generating the action parameter sequences for digital human sign language. The following section will combine... Figure 4 This technical solution will be introduced.

[0153] For example, Figure 4 This is a schematic flowchart illustrating another device for displaying digital human sign language, provided in an embodiment of this application. Figure 4 As shown, the method 400 can be applied to the first device, and the method 400 may include steps 410 to 480.

[0154] 410, The first device acquires audio data A2.

[0155] It should be understood that step 410 can be referred to in the relevant description of step 310 above, and will not be repeated here for the sake of brevity.

[0156] 420, the first device sends audio data A2 to the second device. Correspondingly, the second device receives audio data A2.

[0157] For example, the second device may be a server, such as a cloud server or a local server, or it may be a smart host with strong processing capabilities.

[0158] 430, The second device extracts the human voice audio data B2 from the audio data A2.

[0159] It should be understood that the description of the second device extracting audio data B2 of human voice from audio data A2 can be referred to the process of the first device extracting audio data B1 of human voice from audio data A1.

[0160] 440. The second device generates a sequence of motion parameters for digital human sign language 3 based on the audio data B2, wherein each motion parameter corresponds to a frame of digital human sign language image.

[0161] For example, the second device can first convert the audio data B2 into text content C2, and generate a sequence of motion parameters for digital human sign language based on the text content C2. This technical solution can be found in the description of step 340 above.

[0162] Alternatively, the second device may also include a neural network model that converts audio data into a sequence of motion parameters for digital human sign language. For example, the audio data B2 can be used as input to the neural network model, and the motion parameter sequence 1 can be used as output. Through this neural network model, the second device can convert the audio data B2 into a sequence of motion parameters for digital human sign language 3.

[0163] 450, the second device sends action parameter sequence 3 to the first device. Correspondingly, the first device receives action parameter sequence 3.

[0164] 460, the first device determines the display speed D2 of the image displaying digital human sign language, wherein the display speed D2 is related to the completion time of the motion parameter sequence 4 corresponding to the audio data E2 rendered by the first device.

[0165] 470, the first device renders motion parameter sequence 3 according to the display speed D2 to obtain multi-frame images of digital human sign language.

[0166] 480, the first device displays multiple frames of digital human sign language images.

[0167] It should be understood that steps 460-480 can be referred to in the relevant description of steps 350-370 above, and will not be repeated here for the sake of brevity.

[0168] Based on the embodiments of this application, the first device can provide sign language translation services for any video or television program played by the first device. Furthermore, the first device can leverage the processing power of other devices to provide sign language translation services in real time, avoiding delays in sign language translation. In addition, the first device can adjust the speed of sign language translation by controlling the speed of rendering the motion parameter sequence, ensuring audio-visual synchronization and providing a better user experience.

[0169] Figure 5 This is a schematic flowchart illustrating a method for displaying digital human sign language provided in an embodiment of this application. Figure 5 As shown, the method 500 may include steps 510 to 560.

[0170] 510, The first device acquires the first audio data, which is the audio data including human voices that is being played on the first device.

[0171] For example, the first audio data can be audio data A1 or audio data A2 mentioned above. The first audio data includes audio data of human voices, and may also include audio data of background music other than human voices.

[0172] 520, the first device sends first audio data to the second device. Correspondingly, the second device receives the first audio data.

[0173] The second device can be a server or a device with strong processing power (such as a smart host). For example, the second device can be the second device in method 400 above.

[0174] The first device can send the acquired first audio data to the second device, which can then process the first audio data. For example, the second device can extract human voice audio data from the first audio data, or convert the human voice audio data into text content, or generate a sequence of motion parameters for digital human sign language based on the human voice audio data or text content.

[0175] In other examples, the first device may also convert the first audio data into text and send the text to the second device, which may then generate a first sequence of action parameters based on the text. Alternatively, the first device may extract human voice audio data from the first audio data, convert the human voice audio data into text, and send the text to the second device, which may then generate a first sequence of action parameters based on the text.

[0176] 530, the second device generates a first sequence of motion parameters for digital human sign language based on the audio data of the human voice, and each motion parameter in the first sequence of motion parameters corresponds to a frame of digital human sign language image.

[0177] In some implementations, the second device can first convert the audio data of the human voice into text content, and then generate a sequence of motion parameters for digital human sign language based on the text content. This technical solution can be found in the description of step 340 above.

[0178] In some embodiments, the second device may further include a neural network model that converts audio data of a human voice into a sequence of motion parameters for digital human sign language. For example, the audio data of the human voice can be used as input to the neural network model, and the first sequence of motion parameters can be used as output of the neural network model. Through this neural network model, the second device can convert the audio data of the human voice into a first sequence of motion parameters for digital human sign language.

[0179] 540, the second device sends the first action parameter sequence to the first device.

[0180] 550, the first device renders the first action parameter sequence according to the display speed to obtain a multi-frame image of digital human sign language, wherein the display speed is related to the completion time of the second action parameter sequence corresponding to the second audio data rendered by the first device, and the second audio data is the audio data of human voice acquired by the first device before acquiring the first audio data.

[0181] It should be understood that after receiving the first sequence of action parameters of digital human sign language, the first device can render the first sequence of action parameters at a certain display speed, thereby ensuring the effect of audio-visual synchronization.

[0182] Understandably, the display speed of the multi-frame image of the digital sign language corresponding to each sentence is dynamically changed. If the current sign language translation is delayed, the display speed of the next one can be increased. If the current sign language translation is not delayed, it can be displayed at the normal speed.

[0183] 560, displays multiple frames of digital sign language images.

[0184] It should be understood that the specific method of displaying multiple frames of digital sign language images is not limited in the embodiments of this application. For example, the first device may display the multiple frames of images uniformly after rendering them; or, the first device may display the rendered images while rendering the multiple frames of images, that is, rendering and displaying simultaneously.

[0185] Based on the embodiments of this application, for any video or television program audio data played in the first device, the first device can display the corresponding sign language translation in real time, thus expanding the application scope of this technical solution. Furthermore, the first device can utilize the processing power of the second device to separate audio data and generate a first sequence of motion parameters for digital human sign language, thereby providing real-time sign language translation services. When rendering and displaying the digital human sign language image, the first device can determine the display speed of the current rendering of the first motion parameter sequence by combining the time when the rendering of the previous person's audio data is completed, thereby ensuring audio-visual synchronization and avoiding sign language translation delays.

[0186] In some implementations, before the first device renders the first sequence of motion parameters according to the display speed, the method 500 further includes:

[0187] The first device determines the display speed of the image displaying digital human sign language based on the completion time and the reception time of the first action parameter sequence received by the first device.

[0188] Based on the embodiments of this application, the first device can determine the display speed of rendering the first action parameter sequence according to the completion time of the second action parameter sequence corresponding to the rendering of the second audio data and the reception time of the first action parameter sequence. This allows the determined display speed to be related to the rendering speed of the previous human voice audio data, thereby controlling the display speed to achieve audio-visual synchronization.

[0189] In some implementations, the first device determines the display speed according to the following formula:

[0190]

[0191] Where α>0, D1 represents the display speed, FPS min This indicates the preset minimum display speed in FPS. max This indicates the preset maximum display speed, T1 represents the completion time, T2 represents the reception time, and Tm represents the speed according to FPS. min The time taken to render the first action parameter sequence.

[0192] For example, when the difference between T1 and T2 is large, it means that the first device has generated the next action parameter sequence 1, while the previous action parameter sequence 2 has not yet been rendered. That is, there is a large delay in the sign language action. Therefore, the first device can set a larger value for α to speed up the rendering speed of the action parameter sequence 1, reduce the delay in the sign language action, and maintain the effect of audio-visual synchronization.

[0193] When the difference between T1 and T2 is small, it indicates that the sign language movement delay is small. Therefore, the first device can set a smaller value for α to maintain the effect of audio-visual synchronization.

[0194] When the difference between T1 and T2 is close to 0, it means that there is almost no delay in the sign language action. If the minimum display speed is the normal display speed of the sign language action, then the first device can render the action parameter sequence 1 at the minimum display speed to maintain the effect of audio-visual synchronization.

[0195] In this way, the display speed of the first device rendering the motion parameter sequence is between the set minimum and maximum values, and can be dynamically adjusted, so that the rendering speed of the motion parameter sequence is within a suitable range, achieving the effect of audio-visual synchronization.

[0196] In some implementations, when the completion time and the reception time are the same, the first device determines the display speed of the image displaying digital human sign language based on the completion time and the reception time of the first device receiving the first action parameter sequence, including:

[0197] The first device determines the first preset display speed as the display speed.

[0198] The first preset display speed can be the normal speed for sign language translation. This application embodiment does not limit the specific value of the first preset display speed. For example, the first preset display speed can be 30 FPS, 60 FPS, etc.

[0199] Based on the embodiments of this application, when the completion time and the reception time are the same, the current sign language translation speed is almost equal to the audio playback speed. The first device can display the sign language images of several people at the normal sign language translation speed, which can meet the user's needs and achieve the effect of audio-visual synchronization.

[0200] In some implementations, when the completion time is after the reception time, the first display speed of the image displaying digital human sign language determined by the first device is greater than the second display speed, wherein the second display speed is the display speed of the image displaying digital human sign language determined when the completion time is before the reception time.

[0201] Based on the embodiments of this application, when the completion time is after the reception time, the display speed determined by the first device is the first display speed, and when the completion time is before the reception time, the display speed determined by the first device is the second display speed, and the first display speed is greater than the second display speed.

[0202] In this way, when the current sign language translation lags behind the speech, the speed of sign language translation can be increased, and when the current sign language translation is not lagging behind the speech, the speed of sign language translation can be appropriately reduced in order to maintain the effect of audio-visual synchronization as much as possible.

[0203] In some implementations, before the second device generates a first sequence of action parameters for digital human sign language, the method 500 further includes:

[0204] The second device extracts the audio data of the human voice from the first audio data.

[0205] In this way, by extracting the human voice audio data from the first audio data, the audio data can be made to include only the human voice, filtering out the noise, thereby making the first action parameter sequence for generating digital human sign language more accurate.

[0206] In some implementations, the second device generates a first sequence of motion parameters for digital human sign language based on the audio data of the human voice, including:

[0207] The second device converts the audio data of the human voice into the first text;

[0208] The second device generates a first action parameter sequence based on the first text.

[0209] In other examples, the second device can also directly generate a first sequence of action parameters based on the audio data of the human voice, without converting it to text.

[0210] Figure 6 This is a schematic block diagram of a device provided in an embodiment of this application. Figure 6 As shown, the device 600 includes one or more processors 610; one or more memories 620; the one or more memories 620 storing one or more instructions that, when executed by one or more processors 610, cause the method of displaying digital human sign language as described in any of the possible implementations above to be performed.

[0211] For example, the device 600 can be the device 100, the first device, the second device, etc. mentioned above. The device 600 can be used to execute the methods 300, 400, 500, etc. mentioned above.

[0212] This application also provides an apparatus including a processor and a communication interface. The communication interface is used to receive signals and transmit the signals to the processor. The processor processes the signals so that the method for displaying digital human sign language as described in any of the possible implementations above is executed.

[0213] The device can be a chip. For example, the chip can be a chip system or a standalone chip.

[0214] This application also provides a readable storage medium (also known as a computer-readable storage medium) storing a program that, when run on a device, causes the device to execute the aforementioned method steps to implement the method for displaying digital human sign language in the above embodiments.

[0215] This application also provides a program product (also known as a computer program product) that, when run on a device, causes the device to perform the aforementioned steps to implement the method for displaying digital human sign language in the above embodiments.

[0216] This application also provides an apparatus including a module for implementing the method of displaying digital human sign language as described in any of the preceding embodiments.

[0217] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store instructions, and when the apparatus is running, the processor may execute the instructions stored in the memory to cause the apparatus to perform the method of displaying digital human sign language in the above method embodiments.

[0218] In this embodiment, the device, readable storage medium, program product or apparatus are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0219] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in hardware, or a combination of software and hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0220] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0221] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0222] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0223] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0224] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for displaying digital human sign language, characterized in that, The method is applied to a first device, and the method includes: Acquire first audio data, which is audio data including human voices that is being played on the first device; Send the first audio data to the second device; The device receives a first sequence of motion parameters for digital human sign language sent by the second device, wherein the first sequence of motion parameters is related to the audio data of the human voice, and each motion parameter in the first sequence of motion parameters corresponds to a frame of digital human sign language image. The first action parameter sequence is rendered according to the display speed to obtain a multi-frame image of the digital human sign language. The display speed is related to the completion time of the second action parameter sequence corresponding to the second audio data rendered by the first device. The second audio data is the audio data of human voice acquired by the first device before acquiring the first audio data.

2. The method according to claim 1, characterized in that, Before rendering the first motion parameter sequence according to the display speed, the method further includes: The display speed of the image displaying digital human sign language is determined based on the completion time and the reception time of the first device receiving the first action parameter sequence.

3. The method according to claim 2, characterized in that, The first device determines the display speed according to the following formula: Where α>0, D1 represents the display speed, FPS min This represents the preset minimum display speed, in FPS. max T1 represents the preset maximum value of the display speed, T2 represents the completion time, T2 represents the reception time, and Tm represents the value according to the FPS. min The time taken to render the first action parameter sequence.

4. The method according to claim 2, characterized in that, When the completion time is the same as the reception time Determining the display speed of the digital human sign language image based on the completion time and the reception time of the first device receiving the first action parameter sequence includes: The first preset display speed is determined as the display speed.

5. The method according to claim 2, characterized in that, When the completion time is after the reception time, the first display speed of the image displaying digital human sign language determined by the first device is greater than the second display speed, wherein the second display speed is the display speed of the image displaying digital human sign language determined when the completion time is before the reception time.

6. A method for displaying digital human sign language, characterized in that, The method is applied to a second device, and the method includes: Receive first audio data sent by a first device, the first audio data including audio data of human voice; A first sequence of motion parameters for digital human sign language is generated based on the audio data of the human voice, wherein each motion parameter in the first sequence of motion parameters corresponds to a frame of digital human sign language image; Send the first action parameter sequence to the first device.

7. The method according to claim 6, characterized in that, Before generating the first sequence of action parameters for digital human sign language based on the audio data of the human voice, the method further includes: The audio data of the human voice is extracted from the first audio data.

8. The method according to claim 6 or 7, characterized in that, The step of generating a first sequence of action parameters for digital human sign language based on the audio data of the human voice includes: Convert the audio data of the human voice into first text; Generate the first action parameter sequence based on the first text.

9. A device, characterized in that, include: One or more processors; One or more memory units; The one or more memories store one or more programs that, when executed by one or more processors, cause the method for displaying digital human sign language as described in any one of claims 1-8 to be performed.

10. A readable storage medium, characterized in that, The readable storage medium stores instructions that, when executed on the device, cause the method for displaying digital human sign language as described in any one of claims 1-8 to be performed.