Data processing device, data processing method, and data processing program

The data processing device interprets sign language into text and generates the hearing-impaired person's voice, addressing the challenge of transmitting sign language content in their own voice, facilitating effective communication.

JP7849396B2Active Publication Date: 2026-04-21SOFTBANK GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-01-10
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies cannot transmit the content of sign language of a person with hearing impairment in the voice of the person with hearing impairment themselves.

Method used

A data processing device that includes an image acquisition unit to capture sign language movements, a sign language interpretation unit to interpret the movements into text using AI, and a voice generation unit to generate the hearing-impaired person's voice corresponding to the text, along with optional speech conversion and emotion estimation units to enhance communication.

Benefits of technology

Enables natural conversation between hearing-impaired and hearing individuals by interpreting sign language into text and generating the hearing-impaired person's voice, supporting effective communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007849396000001
    Figure 0007849396000001
  • Figure 0007849396000002
    Figure 0007849396000002
  • Figure 0007849396000003
    Figure 0007849396000003
Patent Text Reader

Abstract

To support hearing-impaired people so that the hearing-impaired people can communicate the content of sign language using their own voice.SOLUTION: A data processing device 12 comprises: an image acquisition unit 310 that acquires images by capturing sign language movements of a hearing-impaired person; a sign language interpretation unit 312 that interprets the sign language movements of the hearing-impaired person from the captured images using a generative AI 58A for sign language interpretation, and outputs the interpreted sign language movements of the hearing-impaired person as text; and a voice generation unit 314 that generates a voice of the hearing-impaired person corresponding to the text using a generative AI 58B for voice generation, and outputs the generated voice of the hearing-impaired person.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a data processing device, a data processing method, and a data processing program.

Background Art

[0002] Patent Document 1 describes a system including a client device that acquires a sign language video including a person performing sign language and transmits the sign language video, and a server device that receives the transmitted sign language video and provides the client device with a support UI (User Interface) that supports an annotation operation of associating a word represented by the sign language performed in the sign language video with the sign language video. This server device identifies the time range during which the sign language operation in the sign language video is performed, uses a sign language word recognition model stored in advance to recognize the word represented by the sign language operation performed during the time range, and presents the identified time range and the recognized word via the support UI.

Prior Art Documents

Patent Documents

[0003] [[ID=B]]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the prior art, although the content of sign language of a person with hearing impairment can be interpreted into characters, the content of sign language of a person with hearing impairment cannot be transmitted in the voice of the person with hearing impairment himself / herself.

[0005] An object of the present disclosure is to provide a data processing device, a data processing method, and a data processing program that can assist in transmitting the content of sign language of a person with hearing impairment in the voice of the person with hearing impairment himself / herself.

Means for Solving the Problems

[0006] A first aspect of the technology of this disclosure is a data processing device comprising: an image acquisition unit that acquires a captured image of a sign language movement of a hearing-impaired person; a sign language interpretation unit that interprets the sign language movement of the hearing-impaired person from the captured image using a sign language interpretation generation AI and outputs the interpreted sign language movement of the hearing-impaired person as text; and a voice generation unit that generates the voice of the hearing-impaired person corresponding to the text using a voice generation generation AI and outputs the generated voice of the hearing-impaired person.

[0007] A second aspect of the technology of this disclosure further comprises, in the first aspect, a speech conversion unit that uses a speech conversion generation AI to convert the speech of a healthy person into sign language or text and output it.

[0008] A third aspect of the technology of the present disclosure further comprises, in the first aspect, an emotion estimation unit that estimates the emotions of the hearing-impaired person from a captured image of the hearing-impaired person, and the speech generation unit uses the speech generation AI to interpret the sign language movements of the hearing-impaired person and generates speech corresponding to the emotions of the hearing-impaired person estimated by the emotion estimation unit for the characters obtained.

[0009] A fourth aspect of the technology of the present disclosure is, in the first aspect, the sign language interpretation unit interprets the sign language movements of the hearing-impaired person and translates the characters obtained into characters of a predetermined language and outputs them, and the voice generation unit uses the voice generation AI to generate the voice of the hearing-impaired person corresponding to the characters translated into the predetermined language by the sign language interpretation unit.

[0010] A fifth aspect of the technology of this disclosure is that, in the first aspect, the speech generation AI learns the voice of the hearing-impaired person and the voice of a close relative of the hearing-impaired person through machine learning.

[0011] A sixth aspect of the technology of this disclosure is a data processing method comprising: acquiring a photographic image of a deaf person's sign language gesture; interpreting the deaf person's sign language gesture from the photographic image using a sign language interpretation generation AI; outputting the interpreted sign language gesture of the deaf person as text; generating the voice of the deaf person corresponding to the text using a voice generation generation AI; and outputting the generated voice of the deaf person.

[0012] A seventh aspect of the technology of this disclosure is a data processing program that causes a computer to perform the following processes: acquire a photographic image of a deaf person's sign language gestures; interpret the deaf person's sign language gestures from the photographic image using a sign language interpretation generation AI; output the interpreted sign language gestures of the deaf person as text; generate the voice of the deaf person corresponding to the text using a voice generation generation AI; and output the generated voice of the deaf person. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system. [Figure 2] This is a conceptual diagram illustrating an example of the essential functions of a data processing device and a smart device. [Figure 3] This diagram schematically shows the functional configuration of a specific processing unit of a data processing device. [Figure 4] This diagram schematically shows an example of the operation flow of a specific process performed by a data processing device. [Figure 5] This figure shows an example of the configuration of a data processing system related to the specific processing of the embodiment. [Figure 6] This is a block diagram showing an example of the functional configuration of a data processing device according to the embodiment. [Figure 7] This figure shows examples of smart device screens for people with hearing impairments and for people without hearing impairments. [Figure 8] This is a flowchart showing an example of the flow of the sign language motion speech generation process according to the embodiment. [Modes for carrying out the invention]

[0014] Next, an example of an embodiment of a data processing apparatus, a data processing method, and a data processing program according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit), etc.

[0017] In the following embodiments, the signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0021] FIG. 1 shows an example of the configuration of a data processing system 10 according to an embodiment.

[0022] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server. An example of the smart device 14 is a smartphone. In this embodiment, the data processing device 12 is an example of the "data processing device" according to the technology of the present disclosure. Note that the smart device 14 may be a terminal device such as a general-purpose PC (Personal Computer), or may be smart glasses, VR (Virtual Reality) goggles, AR (Augmented Reality) goggles, or the like.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the person 20 by outputting the data in a form perceptible to the person 20 (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs sound according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interface 44 and communication interface 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58. The data generation model 58 is used by the specific processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception and output processing. The storage 50 stores the reception and output program 62. The reception and output program 62 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception and output program 62 from the storage 50 and executes the read reception and output program 62 on the RAM 48. The reception and output processing is realized by the processor 46 operating as a control unit 46A according to the reception and output program 62 executed on the RAM 48.

[0032] Next, we will explain the processing of the specific processing unit 290 when the data processing device 12 performs specific processing to support the communication of the content of the sign language of a hearing-impaired person in the hearing-impaired person's own voice.

[0033] As shown in Figure 3, the specific processing unit 290 includes an input unit 292, a processing unit 294, and an output unit 296.

[0034] The input unit 292 acquires user input received by the smart device 14. Specifically, it acquires at least one of the following data from the user received by the smart device 14: text, voice, or image.

[0035] The processing unit 294 performs specific processing using the data generation model 58. Specifically, it inputs character, voice, and image data entered by the user into the data generation model 58 and obtains the generation result.

[0036] The output unit 296 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A then transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0037] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0038] Next, the operation of the data processing system 10 will be explained.

[0039] An example of the flow of a specific processing method will be explained with reference to Figure 4. Note that the flow of a specific processing method shown in Figure 4 is an example of a "data processing method" related to the technology disclosed herein.

[0040] In step S300, the processing unit 294 determines whether or not a predetermined trigger condition is met.

[0041] If the trigger condition is met in step S300 (step S300; Yes), the data processing system 10 proceeds to step S301. On the other hand, if the trigger condition is not met in step S300 (step S300; No), the data processing system 10 terminates the specific processing.

[0042] In step S301, the processing unit 294 generates a prompt by adding an instruction to the text representing the input to obtain the result of a specific process.

[0043] In step S303, the processing unit 294 inputs the generated prompt to the data generation model 58 and obtains the result of a specific process based on the output of the data generation model 58.

[0044] In step S304, the output unit 296 outputs the result of the specific process to the user terminal and terminates the specific process.

[0045] The following provides supplementary information regarding the specific processing of this embodiment. To achieve the following objectives, the specific processing of this embodiment has the following system configuration and functions.

[0046] As mentioned above, while conventional technology can interpret and transcribe the content of sign language used by deaf individuals, it cannot generate the corresponding audio of the deaf person.

[0047] In contrast, the data processing device 12 according to this embodiment interprets the content of the sign language of a hearing-impaired person using the data generation model 58, and generates and outputs the voice of the hearing-impaired person corresponding to the interpreted characters. This makes it possible to support the communication of the content of the sign language of a hearing-impaired person in the hearing-impaired person's own voice.

[0048] Figure 5 shows an example of the configuration of the data processing system 10 related to the specific processing of this embodiment.

[0049] As shown in Figure 5, the data processing system 10 for specific processing in this embodiment includes a data processing device 12, a smart device 14A for a hearing-impaired person, and a smart device 14B for a hearing-controlled person. These data processing device 12, smart device 14A, and smart device 14B are connected to each other via a network 54.

[0050] In the data processing system 10, images of a hearing-impaired person captured by smart device 14A are displayed on the screen of smart device 14B, and images of a hearing person captured by smart device 14B are displayed on the screen of smart device 14A. The captured images are, for example, videos. In the example shown in Figure 5, video conversation is possible between a hearing-impaired person and a hearing person.

[0051] The processing unit 294 of the data processing device 12 shown in Figure 5 uses a data generation model 58 to interpret the content of sign language spoken by a hearing-impaired person, output it as text, and generate the corresponding voice of the hearing-impaired person. Specifically, it functions as shown in Figure 6.

[0052] Figure 6 is a block diagram showing an example of the functional configuration of the data processing device 12 according to this embodiment.

[0053] As shown in Figure 6, the processing unit 294 of the data processing device 12 according to this embodiment functions as an image acquisition unit 310, a sign language interpretation unit 312, a voice generation unit 314, a voice conversion unit 316, and an emotion estimation unit 318.

[0054] The image acquisition unit 310 acquires captured images (videos) of sign language movements of a hearing-impaired person via the smart device 14A.

[0055] The sign language interpretation unit 312 uses the sign language interpretation generation AI 58A to interpret the sign language movements of a hearing-impaired person from a captured image of the hearing-impaired person, and outputs the interpreted sign language movements as text. The sign language interpretation generation AI 58A is an example of a data generation model 58 that has been pre-machine-trained to associate sign language movements with text. According to the sign language interpretation generation AI 58A, by using the generation AI, a natural context is created, so the sign language movements are expressed as natural text. The sign language interpretation generation AI 58A is stored in storage 32, interprets the sign language movements of a hearing-impaired person, and outputs the interpreted sign language movements as text. The text output from the sign language interpretation generation AI 58A is displayed on the screen of the smart device 14B of the hearing person who is the other party in the video conversation.

[0056] The voice generation unit 314 uses the voice generation AI 58B to generate the voice of a hearing-impaired person corresponding to the characters representing the sign language actions of the hearing-impaired person, and outputs the generated voice of the hearing-impaired person. The voice generation AI 58B is an example of a data generation model 58 that has been pre-trained on the voice of a hearing-impaired person. If sufficient voices of hearing-impaired people are not available, it is advisable to train the model on the voices of close relatives who have similar voices to those of the hearing-impaired person. Here, "close relatives" includes, for example, parents, siblings, grandparents, etc. The voice generation AI 58B is stored in the storage 32 and outputs the voice of a hearing-impaired person corresponding to the characters representing the sign language actions of the hearing-impaired person.

[0057] The speech conversion unit 316 uses the speech conversion generation AI 58C to convert the voice of a healthy person into sign language or text and output it. The speech conversion generation AI 58C is an example of a data generation model 58 that has been pre-trained by associating the voice of a healthy person with sign language actions. The speech conversion generation AI 58C may also be trained by associating the voice of a healthy person with text, or by associating the voice of a healthy person with sign language actions and text. The speech conversion generation AI 58C is stored in the storage 32 and outputs the voice of a healthy person converted into sign language or text.

[0058] Figure 7 shows an example screen of a smart device 14A used by a person with hearing impairment and an example screen of a smart device 14B used by a person without hearing impairment.

[0059] As shown in Figure 7, the screen of the hearing-impaired person's smart device 14A displays video footage of a hearing person, along with text obtained by converting the hearing person's speech, and further displays sign language actions obtained by converting the hearing person's speech. The image of the sign language actions is displayed in a separate window from the image of the hearing person, and for example, a character performing the sign language actions may be displayed. Alternatively, the screen of the smart device 14A may display only text along with the video footage of the hearing person, or only sign language actions.

[0060] Meanwhile, the screen of the hearing-controlled smart device 14B displays video footage of the hearing-impaired person along with text obtained by interpreting the hearing-impaired person's sign language movements, and further outputs the hearing-impaired person's voice corresponding to the displayed text. Alternatively, the screen of the smart device 14B may display only video footage of the hearing-impaired person and output the hearing-impaired person's voice.

[0061] Here, the emotion estimation unit 318 estimates the emotions of a person with hearing impairment from the captured image of that person. The emotion estimation unit 318 estimates an emotion value indicating the emotions of the person with hearing impairment based on the state of the person with hearing impairment that can be recognized from the captured image. For example, the state of the person with hearing impairment that can be recognized from the captured image is input into a pre-trained generating AI to obtain an emotion value indicating the emotions of the person with hearing impairment.

[0062] Specifically, the emotion estimation unit 318 recognizes the facial expression and emotions of a hearing-impaired person from an image of the hearing-impaired person captured by the camera. The emotion estimation unit 318 recognizes the facial expression and emotions of the hearing-impaired person based, for example, on the shape and relative position of the hearing-impaired person's eyes and mouth.

[0063] The emotion value, which indicates the feelings of a person with hearing impairment, is a value that shows whether the user's emotion is positive or negative. For example, if the emotion of a person with hearing impairment is a positive emotion accompanied by pleasure or comfort, such as "joy," "pleasure," "happiness," "security," "excitement," "relief," and "fulfillment," it will show a positive value, and the more positive the emotion, the larger the value. If the emotion of a person with hearing impairment is an unpleasant emotion, such as "anger," "sadness," "discomfort," "anxiety," "grief," "worry," and "emptiness," it will show a negative value, and the more unpleasant the emotion, the larger the absolute value of the negative value. If the emotion of a person with hearing impairment is none of the above ("neutral"), it will show a value of 0.

[0064] The speech generation unit 314 may use the speech generation AI 58B to interpret the sign language movements of the hearing-impaired person and generate speech corresponding to the emotions of the hearing-impaired person estimated by the emotion estimation unit 318. For example, if the emotion value of the hearing-impaired person is positive, the speech of the hearing-impaired person will also be made to have a bright tone, and if the emotion value of the hearing-impaired person is negative, the speech of the hearing-impaired person will also be made to have a dark tone.

[0065] Furthermore, the sign language interpretation unit 312 may translate the characters obtained by interpreting the sign language movements of the hearing-impaired person into characters of a pre-specified language and output them. The "pre-specified language" should preferably be a language that the hearing person, who is the conversation partner, can understand. For example, if the hearing-impaired person's native language is Japanese and the hearing person's native language is Korean, the unit should specify that Japanese be translated into Korean. In this case, the speech generation unit 314 uses the speech generation AI 58B to generate the hearing-impaired person's voice corresponding to the characters translated into the pre-specified language by the sign language interpretation unit 312.

[0066] Next, with reference to Figure 8, the operation of the data processing device 12 according to this embodiment will be described.

[0067] Figure 8 is a flowchart showing an example of the flow of the sign language action speech generation process according to this embodiment.

[0068] First, when the data processing device 12 is instructed to perform sign language action voice generation processing between a hearing-impaired person and a hearing-controlled person, the processing unit 294 starts a specific processing program 56 and executes the following steps.

[0069] In step S401 of Figure 8, the processing unit 294 acquires captured images (videos) of a deaf person's sign language movements via the smart device 14A.

[0070] In step S402, the processing unit 294 interprets the sign language movements of a hearing-impaired person from the captured image acquired in step S401 using the sign language interpretation generation AI 58A.

[0071] In step S403, the processing unit 294 outputs the sign language gestures of the hearing-impaired person, as interpreted in step S402, as text. Here, as an example, as shown in Figure 7 above, it is displayed on the screen of the hearing person's smart device 14B.

[0072] In step S404, the processing unit 294 uses the speech generation AI 58B to generate speech from a person with a hearing impairment that corresponds to the characters representing the sign language actions of a person with a hearing impairment.

[0073] In step S405, the processing unit 294 outputs the voice of the hearing-impaired person generated in step S404. Here, as an example, as shown in Figure 7 above, the output is made from the speaker of the hearing-controlled person's smart device 14B.

[0074] In step S406, the processing unit 294 acquires the voice of a healthy person via the smart device 14B.

[0075] In step S407, the processing unit 294 uses the speech conversion generation AI 58C to convert the speech of a healthy person acquired in step S406 into text and sign language actions.

[0076] In step S408, the processing unit 294 outputs the text and sign language actions converted from the voice of the healthy person in step S407. As an example, as shown in Figure 7 above, these are displayed on the screen of the healthy person's smart device 14A.

[0077] In step S409, the processing unit 294 determines whether the video conversation between the hearing-impaired person and the hearing-controlled person has ended. If it determines that the video conversation has not ended (negative determination), the process returns to step S401 and is repeated. If it determines that the video conversation has ended (positive determination), the series of sign language speech generation processes is terminated.

[0078] Thus, according to this embodiment, it is possible to use a generation AI to interpret the content of sign language used by a hearing-impaired person, output it as text, and generate the corresponding voice of the hearing-impaired person. This enables natural conversation between hearing-impaired and hearing-impaired individuals using the hearing-impaired person's own voice. It can also support hearing-impaired individuals in communicating the content of sign language using their own voice.

[0079] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0080] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing method for the specific process may be used, which includes computer 22 and multiple other computers.

[0081] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0082] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0083] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0084] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0085] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0086] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0087] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0088] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0089] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0090] The following additional information is disclosed regarding the embodiments described above.

[0091] (Note 1) An image acquisition unit that acquires images of sign language movements of hearing-impaired persons, A sign language interpretation unit that uses a sign language interpretation generation AI to interpret the sign language movements of the hearing-impaired person from the captured image and outputs the interpreted sign language movements of the hearing-impaired person as text, A voice generation unit that uses a voice generation AI to generate the voice of the hearing-impaired person corresponding to the characters and outputs the generated voice of the hearing-impaired person, A data processing device equipped with [a specific feature]. (Note 2) It further includes a speech conversion unit that uses a speech conversion generation AI to convert the speech of a healthy person into sign language or text and output it. The data processing device described in Appendix 1. (Note 3) The system further includes an emotion estimation unit that estimates the emotions of the hearing-impaired person from the captured images of the hearing-impaired person. The voice generation unit uses the voice generation AI to interpret the sign language movements of the hearing-impaired person and generates voice corresponding to the emotions of the hearing-impaired person estimated by the emotion estimation unit, based on the characters obtained. A data processing device as described in Appendix 1 or Appendix 2. (Note 4) The sign language interpretation unit interprets the sign language movements of the hearing-impaired person, translates the resulting characters into characters of a pre-specified language, and outputs them. The voice generation unit uses the voice generation AI to generate the voice of the hearing-impaired person corresponding to the characters translated into the pre-specified language by the sign language interpretation unit. A data processing device as described in any one of the appendices 1 to 3. (Note 5) The aforementioned AI for generating speech learns from the voice of the hearing-impaired person and the voice of a close relative of the hearing-impaired person. A data processing device as described in any one of the appendices 1 to 4. (Note 6) We obtained images of sign language movements performed by people with hearing impairments. Using a sign language interpretation generation AI, the sign language movements of the hearing-impaired person are interpreted from the captured image, and the interpreted sign language movements of the hearing-impaired person are output as text. Using a speech generation AI, the system generates the voice of the hearing-impaired person corresponding to the aforementioned characters, and outputs the generated voice of the hearing-impaired person. Data processing method. (Note 7) We obtained images of sign language movements performed by people with hearing impairments. Using a sign language interpretation generation AI, the sign language movements of the hearing-impaired person are interpreted from the captured image, and the interpreted sign language movements of the hearing-impaired person are output as text. A process that includes generating the voice of the hearing-impaired person corresponding to the characters using a speech generation AI, and outputting the generated voice of the hearing-impaired person, A data processing program designed to be executed by a computer. [Explanation of Symbols]

[0092] 10 Data Processing Systems 12 Data Processing Devices 14, 14A, 14B Smart Devices 58 Data Generation Models 58A AI for generating sign language interpretation 58B Generation AI for voice generation 58C Voice Conversion Generation AI 290 Specific Processing Unit 292 Input section 294 Processing Unit 296 Output section 310 Image acquisition unit 312 Sign Language Interpretation Department 314 Voice generation unit 316 Voice Conversion Unit 318 Emotion estimation part

Claims

1. An image acquisition unit that acquires images of sign language movements of hearing-impaired persons, A sign language interpretation unit that uses a sign language interpretation generation AI to interpret the sign language movements of the hearing-impaired person from the captured image and outputs the interpreted sign language movements of the hearing-impaired person as text, A voice generation unit that generates the voice of the hearing-impaired person corresponding to the characters using a voice generation AI, and outputs the generated voice of the hearing-impaired person, A voice conversion unit that uses a voice conversion generation AI to convert the voice of a non-speaking person who is communicating with sign language into sign language or text and output it, Equipped with, The voice conversion unit displays, along with the video of the hearing person, the characters obtained by converting the voice of the hearing person, and the images of sign language gestures obtained by converting the voice of the hearing person on the terminal screen of the hearing-impaired person. The image of the sign language movement is displayed as a character performing the sign language movement in a separate window from the video of the healthy person. Data processing device.

2. The system further includes an emotion estimation unit that estimates the emotions of the hearing-impaired person from the captured images of the hearing-impaired person. The voice generation unit uses the voice generation AI to interpret the sign language movements of the hearing-impaired person and generates voice corresponding to the emotions of the hearing-impaired person estimated by the emotion estimation unit, based on the characters obtained. The data processing device according to claim 1.

3. The sign language interpretation unit interprets the sign language movements of the hearing-impaired person, translates the resulting characters into characters of a pre-specified language, and outputs them. The voice generation unit uses the voice generation AI to generate the voice of the hearing impaired person corresponding to the characters translated into the pre-specified language by the sign language interpretation unit. The data processing device according to claim 1.

4. The aforementioned AI for generating speech learns from the voice of the hearing-impaired person and the voice of a close relative of the hearing-impaired person. The data processing device according to claim 1.

5. We obtained images of sign language movements performed by people with hearing impairments. Using a sign language interpretation generation AI, the sign language movements of the hearing-impaired person are interpreted from the captured image, and the interpreted sign language movements of the hearing-impaired person are output as text. Using a speech generation AI, generate the voice of the hearing-impaired person corresponding to the characters, and output the generated voice of the hearing-impaired person. When using a speech conversion generation AI to convert the voice of a hearing person who is communicating with them into sign language or text and output it, the text obtained by converting the voice of the hearing person and the images of the sign language movements obtained by converting the voice of the hearing person are displayed on the hearing-impaired person's terminal screen along with the video of the hearing person. The process includes displaying the image of the sign language action as a character performing the sign language action in a separate window from the video of the healthy person, The data processing methods performed by computers.

6. We obtained images of sign language movements performed by people with hearing impairments. Using a sign language interpretation generation AI, the sign language movements of the hearing-impaired person are interpreted from the captured image, and the interpreted sign language movements of the hearing-impaired person are output as text. Using a speech generation AI, generate the voice of the hearing-impaired person corresponding to the characters, and output the generated voice of the hearing-impaired person. When using a speech conversion generation AI to convert the voice of a hearing person who is communicating with them into sign language or text and output it, the text obtained by converting the voice of the hearing person and the images of the sign language movements obtained by converting the voice of the hearing person are displayed on the hearing-impaired person's terminal screen along with the video of the hearing person. The process includes displaying the image of the sign language action as a character performing the sign language action in a separate window from the video of the healthy person, A data processing program designed to be executed by a computer.

Citation Information

Patent Citations

  • Voice synthesizing device, method and program

    JP2020160319A

  • System, server device and program

    JP6840365B2

  • System and method for providing call service for the hearing impaired

    KR102487847B1