Data processing device, data processing method, and data processing program

The data processing device interprets sign language into text and generates the hearing-impaired person's voice, addressing the inability of existing systems to convey sign language in the hearing-impaired person's voice, enabling natural communication.

JP2025108283AActive Publication Date: 2025-07-23SOFTBANK GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024002118
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-10
Publication Date
2025-07-23
Estimated Expiration
2044-01-10

AI Technical Summary

Technical Problem

Existing systems can interpret sign language into characters but cannot generate the voice of the hearing-impaired person corresponding to the interpreted characters.

Method used

A data processing device that includes an image acquisition unit to capture sign language motions, a sign language interpretation unit to interpret the motions into text using a generative AI, and a voice generation unit to generate the hearing-impaired person's voice based on the text using a generative AI.

Benefits of technology

Enables the transmission of sign language content in the voice of the hearing-impaired person, facilitating natural conversation by allowing the hearing-impaired person's voice to be generated corresponding to their sign language.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025108283000001_ABST
    Figure 2025108283000001_ABST
Patent Text Reader

Abstract

To support hearing-impaired people so that the hearing-impaired people can communicate the content of sign language using their own voice.SOLUTION: A data processing device 12 comprises: an image acquisition unit 310 that acquires images by capturing sign language movements of a hearing-impaired person; a sign language interpretation unit 312 that interprets the sign language movements of the hearing-impaired person from the captured images using a generative AI 58A for sign language interpretation, and outputs the interpreted sign language movements of the hearing-impaired person as text; and a voice generation unit 314 that generates a voice of the hearing-impaired person corresponding to the text using a generative AI 58B for voice generation, and outputs the generated voice of the hearing-impaired person.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a data processing device, a data processing method, and a data processing program.

Background Art

[0002] Patent Document 1 describes a system including a client device that acquires a sign language video including a person performing sign language and transmits the sign language video, and a server device that receives the transmitted sign language video and provides a support UI (User Interface) for assisting an annotation operation of associating a word represented by the sign language performed in the sign language video with the sign language video to the client device. This server device identifies a time range in which a sign language operation in the sign language video is being performed, uses a sign language word recognition model stored in advance to recognize a word represented by the sign language operation being performed in the time range, and presents the identified time range and the recognized word via the support UI.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the prior art, although the content of the sign language of a hearing-impaired person can be interpreted into characters, the content of the sign language of a hearing-impaired person cannot be transmitted in the voice of the hearing-impaired person himself / herself.

[0005] An object of the present disclosure is to provide a data processing device, a data processing method, and a data processing program that can assist in transmitting the content of the sign language of a hearing-impaired person in the voice of the hearing-impaired person himself / herself.

Means for Solving the Problems

[0006] A first aspect of the technology according to the present disclosure is a data processing device, comprising: an image acquisition unit that acquires a captured image of a sign language motion of a hearing-impaired person; a sign language interpretation unit that interprets the sign language motion of the hearing-impaired person from the captured image using a generative AI for sign language interpretation, and outputs the interpreted sign language motion of the hearing-impaired person as text; and a voice generation unit that generates the voice of the hearing-impaired person corresponding to the text using a generative AI for voice generation, and outputs the generated voice of the hearing-impaired person.

[0007] A second aspect of the technology according to the present disclosure is, in the first aspect, further comprising a voice conversion unit that converts the voice of a healthy person into sign language or text and outputs it using a generative AI for voice conversion.

[0008] A third aspect of the technology according to the present disclosure is, in the first aspect, further comprising an emotion estimation unit that estimates the emotion of the hearing-impaired person from the captured image of the hearing-impaired person, and the voice generation unit uses the generative AI for voice generation to generate a voice corresponding to the emotion of the hearing-impaired person estimated by the emotion estimation unit for the text obtained by interpreting the sign language motion of the hearing-impaired person.

[0009] A fourth aspect of the technology according to the present disclosure is, in the first aspect, the sign language interpretation unit translates the text obtained by interpreting the sign language motion of the hearing-impaired person into text in a pre-specified language and outputs it, and the voice generation unit uses the generative AI for voice generation to generate the voice of the hearing-impaired person corresponding to the text translated into the pre-specified language by the sign language interpretation unit.

[0010] A fifth aspect of the technology according to the present disclosure is, in the first aspect, the generative AI for voice generation performs machine learning on the voice of the hearing-impaired person and the voice of a close relative of the hearing-impaired person.

[0011] A sixth aspect of the technology according to the present disclosure is a data processing method, which includes obtaining a captured image of a sign language motion of a hearing-impaired person, interpreting the sign language motion of the hearing-impaired person from the captured image using a generative AI for sign language interpretation, converting the interpreted sign language motion of the hearing-impaired person into characters and outputting the characters, generating the voice of the hearing-impaired person corresponding to the characters using a generative AI for voice generation, and outputting the generated voice of the hearing-impaired person.

[0012] A seventh aspect of the technology according to the present disclosure is a data processing program, which causes a computer to execute a process including obtaining a captured image of a sign language motion of a hearing-impaired person, interpreting the sign language motion of the hearing-impaired person from the captured image using a generative AI for sign language interpretation, converting the interpreted sign language motion of the hearing-impaired person into characters and outputting the characters, generating the voice of the hearing-impaired person corresponding to the characters using a generative AI for voice generation, and outputting the generated voice of the hearing-impaired person.

Brief Description of Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Modes for Carrying Out the Invention

[0014] Next, an example of an embodiment of a data processing apparatus, a data processing method, and a data processing program according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit), etc.

[0017] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applicable to the communication I / F include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B". That is, "A and / or B" means that it may be only A, only B, or a combination of A and B. Also, in this specification, when expressing three or more matters connected by "and / or", the same concept as "A and / or B" is applied.

[0021] FIG. 1 shows an example of the configuration of the data processing system 10 according to the embodiment.

[0022] As shown in FIG. 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server. An example of the smart device 14 is a smartphone. In this embodiment, the data processing device 12 is an example of the "data processing device" according to the technology of the present disclosure. Note that the smart device 14 may be a general-purpose terminal device such as a PC (Personal Computer), or may be smart glasses, VR (Virtual Reality) goggles, AR (Augmented Reality) goggles, or the like.

[0023] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. Also, the database 24 and the communication I / F 26 are connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the reception device 38, the output device 40, and the camera 42 are connected to the bus 52.

[0025] The reception device 38 includes a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by contact of an indicator (for example, a pen or a finger, etc.) by detecting the contact of the indicator. The microphone 38B receives user input by voice by detecting the voice of the user. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires data indicating the user input.

[0026] The output device 40 includes a display 40A, a speaker 40B, etc., and presents data to the person 20 by outputting the data in an expression form (e.g., voice and / or text) that can be perceived by the person 20. The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, a diaphragm, and a shutter, and an imaging device such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] The communication I / F 44 is connected to the network 54. The communication I / F 44 and the communication I / F 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0028] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in FIG. 2, in the data processing device 12, specific processing is performed by the processor 28. The storage 32 stores a specific processing program 56. The specific processing program 56 is an example of the "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores a data generation model 58. The data generation model 58 is used by the specific processing unit 290.

[0031] In the smart device 14, the reception / output processing is performed by the processor 46. In the storage 50, a reception / output program 62 is stored. The reception / output program 62 is used in combination with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception / output program 62 from the storage 50 and executes the read reception / output program 62 on the RAM 48. The reception / output processing is realized by operating as the control unit 46A according to the reception / output program 62 executed by the processor 46 on the RAM 48.

[0032] Next, the processing of the specific processing unit 290 when the data processing device 12 performs specific processing to assist in transmitting the content of the sign language of a hearing-impaired person in the voice of the hearing-impaired person himself / herself will be described.

[0033] As shown in FIG. 3, the specific processing unit 290 includes an input unit 292, a processing unit 294, and an output unit 296.

[0034] The input unit 292 acquires the user input received by the smart device 14. Specifically, it acquires at least one of the data of characters, voice, and images of the user received by the smart device 14.

[0035] The processing unit 294 performs specific processing using the data generation model 58. Specifically, the data of characters, voice, and images input from the user is input to the data generation model 58 to obtain a generation result.

[0036] The output unit 296 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires the voice indicating the user input for the result of the specific processing. Note that the control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0037] The data generation model 58 is a so-called generative AI (Artificial Intelligence). Examples of the data generation model 58 include generative AIs such as ChatGPT (Internet search <URL: https: / / openai.com / blog / chatgpt>). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is input thereto. The data generation model 58 infers the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization, etc.

[0038] Next, the operation of the data processing system 10 will be described.

[0039] An example of the flow of a specific process will be described with reference to FIG. 4. Note that the flow of the specific process shown in FIG. 4 is an example of the "data processing method" according to the technology of the present disclosure.

[0040] In step S300, the processing unit 294 determines whether or not a predetermined trigger condition is satisfied.

[0041] If the trigger condition is satisfied in step S300 (step S300; Yes), the data processing system 10 proceeds to step S301. On the other hand, if the trigger condition is not satisfied in step S300 (step S300; No), the data processing system 10 ends the specific process.

[0042] In step S301, the processing unit 294 adds an instruction sentence for obtaining the result of the specific process to the text representing the input to generate a prompt.

[0043] In step S303, the processing unit 294 inputs the generated prompt into the data generation model 58, and obtains the result of the specific process based on the output of the data generation model 58.

[0044] In step S304, the output unit 296 outputs the result of the specific process to the user terminal and ends the specific process.

[0045] Hereinafter, the specific process of this embodiment will be supplemented. The specific process of this embodiment has the following system configuration and functions for achieving the following problems.

[0046] As described above, in the prior art, it is possible to interpret the content of the sign language of a hearing-impaired person into characters, but it is not possible to generate the voice of the hearing-impaired person corresponding to the interpreted characters.

[0047] On the other hand, the data processing device 12 according to this embodiment uses the data generation model 58 to interpret the content of the sign language of a hearing-impaired person, and generates and outputs the voice of the hearing-impaired person corresponding to the interpreted characters. Thereby, it is possible to assist in transmitting the content of the sign language of a hearing-impaired person in the voice of the hearing-impaired person himself.

[0048] FIG. 5 is a diagram showing an example of the configuration of the data processing system 10 according to the specific process of this embodiment.

[0049] As shown in FIG. 5, the data processing system 10 according to the specific process of this embodiment includes a data processing device 12, a smart device 14A of a hearing-impaired person, and a smart device 14B of a healthy person. These data processing device 12, smart device 14A, and smart device 14B are communicably connected via a network 54.

[0050] In the data processing system 10, a captured image of a hearing-impaired person taken by the smart device 14A is displayed on the screen of the smart device 14B, and a captured image of a healthy person taken by the smart device 14B is displayed on the screen of the smart device 14A. The captured image is, for example, a video. In the example of FIG. 5, a video call is enabled between the hearing-impaired person and the healthy person.

[0051] The processing unit 294 of the data processing device 12 shown in FIG. 5 uses the data generation model 58 to interpret the content of the sign language of the hearing-impaired person and output characters, and generates the voice of the hearing-impaired person corresponding to the interpreted characters. Specifically, it functions as each part shown in FIG. 6.

[0052] FIG. 6 is a block diagram showing an example of the functional configuration of the data processing device 12 according to the present embodiment.

[0053] As shown in FIG. 6, the processing unit 294 of the data processing device 12 according to the present embodiment functions as an image acquisition unit 310, a sign language interpretation unit 312, a voice generation unit 314, a voice conversion unit 316, and an emotion estimation unit 318.

[0054] The image acquisition unit 310 acquires a captured image (video) of the sign language operation of the hearing-impaired person through the smart device 14A.

[0055] The sign language interpretation unit 312 uses the sign language interpretation generation AI 58A to interpret the sign language operation of the hearing-impaired person from the captured image of the hearing-impaired person, and outputs the interpreted sign language operation of the hearing-impaired person as characters. The sign language interpretation generation AI 58A is an example of the data generation model 58 that has been pre-trained by associating sign language operations with characters. According to the sign language interpretation generation AI 58A, by using the generation AI, a natural context is obtained, so that the sign language operation is expressed as a natural sentence. The sign language interpretation generation AI 58A is stored in the storage 32, interprets the sign language operation of the hearing-impaired person, and outputs the interpreted sign language operation of the hearing-impaired person as characters. The characters output from the sign language interpretation generation AI 58A are output to the screen of the smart device 14B of the healthy person who is the video call partner.

[0056] The voice generation unit 314 uses the voice generation AI 58B for generating a voice of a hearing-impaired person corresponding to the characters representing the sign language actions of the hearing-impaired person, and outputs the generated voice of the hearing-impaired person. The voice generation AI 58B for generating is an example of the data generation model 58 that has previously machine-learned the voice of the hearing-impaired person himself / herself. When the voice of the hearing-impaired person cannot be obtained sufficiently, it is advisable to machine-learn the voice of a close relative who has a voice similar to that of the hearing-impaired person. Here, the "close relative" includes, for example, parents, siblings, grandparents, etc. The voice generation AI 58B for generating is stored in the storage 32 and outputs the voice of the hearing-impaired person corresponding to the characters representing the sign language actions of the hearing-impaired person.

[0057] The voice conversion unit 316 uses the voice conversion AI 58C to convert the voice of a healthy person into sign language or characters and outputs them. The voice conversion AI 58C is an example of the data generation model 58 that has previously machine-learned the association between the voice of a healthy person and sign language actions. Note that the voice conversion AI 58C may machine-learn the association between the voice of a healthy person and characters, or may machine-learn the association between the voice of a healthy person, sign language actions, and characters. The voice conversion AI 58C is stored in the storage 32 and converts the voice of a healthy person into sign language or characters and outputs them.

[0058] FIG. 7 is a diagram showing an example of the screen of the hearing-impaired person's smart device 14A and an example of the screen of the healthy person's smart device 14B.

[0059] As shown in FIG. 7, on the screen of the hearing-impaired person's smart device 14A, characters obtained by converting the voice of a healthy person, as well as a moving image of the healthy person, are displayed. Further, sign language actions obtained by converting the voice of a healthy person are displayed. Note that the image of the sign language actions is displayed in a window different from the image of the healthy person, and for example, a character performing the sign language actions may be displayed. Also, only characters or only sign language actions may be displayed on the screen of the smart device 14A together with the moving image of the healthy person.

[0060] On the screen of the smart device 14B of a healthy person, characters obtained by interpreting the sign language movements of the hearing-impaired person are displayed together with the moving images of the hearing-impaired person, and furthermore, the voice of the hearing-impaired person corresponding to the displayed characters is output. Note that only the moving images of the hearing-impaired person may be displayed on the screen of the smart device 14B, and the voice of the hearing-impaired person may be output.

[0061] Here, the emotion estimation unit 318 estimates the emotion of the hearing-impaired person from the captured image of the hearing-impaired person. The emotion estimation unit 318 estimates an emotion value indicating the emotion of the hearing-impaired person based on the state of the hearing-impaired person recognizable from the captured image. For example, the state of the hearing-impaired person recognizable from the captured image is input into a pre-trained generative AI, and an emotion value indicating the emotion of the hearing-impaired person is obtained.

[0062] Specifically, the emotion estimation unit 318 recognizes the expression and emotion of the hearing-impaired person from the captured image of the hearing-impaired person captured by the camera. The emotion estimation unit 318 recognizes the expression and emotion of the hearing-impaired person based on, for example, the shape and position relationship of the eyes and mouth of the hearing-impaired person.

[0063] The emotion value indicating the emotion of the hearing-impaired person is a value indicating the positive or negative of the user's emotion. For example, if the emotion of the hearing-impaired person is a bright emotion accompanied by pleasure or relief, such as "happy", "joyful", "pleasant", "at ease", "excited", "relieved", and "a sense of fulfillment", it indicates a positive value, and the brighter the emotion, the larger the value. If the emotion of the hearing-impaired person is an emotion that makes them feel uncomfortable, such as "angry", "sad", "unpleasant", "uneasy", "sadness", "worried", and "nihility", it indicates a negative value, and the more uncomfortable the emotion, the larger the absolute value of the negative value. If the emotion of the hearing-impaired person is none of the above ("ordinary"), it indicates a value of 0.

[0064] The voice generation unit 314 may generate a voice corresponding to the emotion of the hearing-impaired person estimated by the emotion estimation unit 318 for the characters obtained by interpreting the sign language movements of the hearing-impaired person using the generation AI 58B for voice generation. For example, if the emotion value of the hearing-impaired person is a positive value, the voice of the hearing-impaired person may also be a voice with a bright tone, and if the emotion value of the hearing-impaired person is a negative value, the voice of the hearing-impaired person may also be a voice with a dark tone.

[0065] In addition, the sign language interpretation unit 312 may translate and output the characters obtained by interpreting the sign language movements of the hearing-impaired person into characters in a pre-specified language. It is desirable that the "pre-specified language" be, for example, a language that can be understood by a normal person who is the conversation partner. For example, if the native language of the hearing-impaired person is Japanese and the native language of the normal person is Korean, it may be specified to translate Japanese into Korean. In this case, the voice generation unit 314 uses the generation AI 58B for voice generation to generate the voice of the hearing-impaired person corresponding to the characters translated into the pre-specified language by the sign language interpretation unit 312.

[0066] Next, with reference to FIG. 8, the operation of the data processing device 12 according to the present embodiment will be described.

[0067] FIG. 8 is a flowchart showing an example of the flow of sign language movement voice generation processing according to the present embodiment.

[0068] First, when the data processing device 12 is instructed to execute sign language movement voice generation processing between a hearing-impaired person and a normal person, the specific processing program 56 is started by the processing unit 294, and the following steps are executed.

[0069] In step S401 of FIG. 8, the processing unit 294 acquires a captured image (video) of the sign language movements of the hearing-impaired person via the smart device 14A.

[0070] In step S402, the processing unit 294 interprets the sign language movements of the hearing-impaired person from the captured image acquired in step S401 using the generation AI 58A for sign language interpretation.

[0071] In step S403, the processing unit 294 converts the sign language actions of the hearing-impaired person interpreted in step S402 into characters and outputs them. Here, as an example, as shown in FIG. 7 described above, it is displayed on the screen of the healthy person's smart device 14B.

[0072] In step S404, the processing unit 294 uses the voice generation AI 58B to generate the voice of the hearing-impaired person corresponding to the characters representing the sign language actions of the hearing-impaired person.

[0073] In step S405, the processing unit 294 outputs the voice of the hearing-impaired person generated in step S404. Here, as an example, as shown in FIG. 7 described above, it is output from the speaker of the healthy person's smart device 14B.

[0074] In step S406, the processing unit 294 acquires the voice uttered by the healthy person via the smart device 14B.

[0075] In step S407, the processing unit 294 uses the voice conversion AI 58C to convert the voice of the healthy person acquired in step S406 into characters and sign language actions.

[0076] In step S408, the processing unit 294 outputs the characters and sign language actions obtained by converting the voice of the healthy person in step S407. As an example, as shown in FIG. 7 described above, it is displayed on the screen of the healthy person's smart device 14A.

[0077] In step S409, the processing unit 294 determines whether or not the video conversation between the hearing-impaired person and the healthy person has ended. If it is determined that the video conversation has not ended (in the case of a negative determination), the process returns to step S401 and the process is repeated. If it is determined that the video conversation has ended (in the case of an affirmative determination), the series of sign language action voice generation processes is terminated.

[0078] According to this embodiment, by using a generative AI, the content of the sign language of a hearing-impaired person can be interpreted to output text, and the voice of the hearing-impaired person corresponding to the interpreted text can be generated. As a result, a natural conversation can be held with a hearing person by the voice of the hearing-impaired person himself / herself. It is possible to support the hearing-impaired person so that the content of the sign language can be conveyed by the voice of the hearing-impaired person himself / herself.

[0079] As described above, the system according to the present disclosure has been mainly described in terms of the functions of the data processing device 12. However, the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program operating on a personal computer or an application operating on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).

[0080] In the above embodiment, an example of a form in which specific processing is performed by one computer 22 has been given. However, the technology of the present disclosure is not limited to this, and distributed processing for specific processing by a plurality of computers including the computer 22 may be performed.

[0081] In the above embodiment, an example of a form in which the specific processing program 56 is stored in the storage 32 has been described. However, the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0082] Alternatively, a storage device such as a server connected to the data processing device 12 via the network 54 may store the specific processing program 56, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0083] Note that it is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32. A part of the specific processing program 56 may be stored.

[0084] As hardware resources for executing the specific processing, various types of processors shown below can be used. As the processor, for example, a general-purpose processor such as a CPU that functions as a hardware resource for executing specific processing by executing software, that is, a program, can be mentioned. Also, as the processor, for example, a dedicated electric circuit that is a processor having a circuit configuration designed specifically for executing specific processing such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit) can be mentioned. A memory is built in or connected to any of these processors, and any of these processors executes specific processing by using the memory.

[0085] The hardware resources for executing the specific processing may be configured by one of these various types of processors, or may be configured by a combination of two or more processors of the same type or different types (for example, a combination of a plurality of FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resources for executing the specific processing may be one processor.

[0086] As an example of a configuration consisting of one processor, first, there is a form in which one processor is configured by a combination of one or more CPUs and software, and this processor functions as a hardware resource for executing specific processing. Second, as represented by a SoC (System-on-a-chip), etc., there is a form in which a processor that realizes the functions of an entire system including a plurality of hardware resources for executing specific processing is used in one IC chip. Thus, the specific processing is realized as a hardware resource using one or more of the above various processors.

[0087] Furthermore, as a more specific hardware structure of these various processors, an electric circuit combining circuit elements such as semiconductor elements can be used. Also, the above specific processing is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be changed within the scope of not departing from the gist.

[0088] The description content and illustrated content shown above are detailed descriptions of the part related to the technology of the present disclosure and are merely examples of the technology of the present disclosure. For example, the description regarding the above configuration, function, action, and effect is an example of the configuration, function, action, and effect of the part related to the technology of the present disclosure. Therefore, it goes without saying that within the scope of not departing from the gist of the technology of the present disclosure, the description content and illustrated content shown above may be deleted of unnecessary parts, new elements may be added, or replacements may be made. Also, for the purpose of avoiding intricacies and facilitating the understanding of the part related to the technology of the present disclosure, in the description content and illustrated content shown above, descriptions regarding common technical knowledge, etc., which do not particularly require explanation for enabling the implementation of the technology of the present disclosure, are omitted.

[0089] All documents, patent applications, and technical standards described in this specification are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually stated to be incorporated by reference.

[0090] Regarding the above embodiments, the following additional remarks are disclosed.

[0091] (Appendix 1) An image acquisition unit that acquires a captured image of a sign language motion of a hearing-impaired person, A sign language interpretation unit that interprets the sign language motion of the hearing-impaired person from the captured image using a generation AI for sign language interpretation, and outputs the interpreted sign language motion of the hearing-impaired person as characters, A voice generation unit that generates the voice of the hearing-impaired person corresponding to the characters using a generation AI for voice generation, and outputs the generated voice of the hearing-impaired person, A data processing device comprising the above. (Appendix 2) The data processing device according to Appendix 1, further comprising a voice conversion unit that converts the voice of a healthy person into sign language or characters and outputs the converted result using a generation AI for voice conversion. The data processing device according to Appendix 1. (Appendix 3) The data processing device according to Appendix 1 or Appendix 2, further comprising an emotion estimation unit that estimates the emotion of the hearing-impaired person from the captured image of the hearing-impaired person. The voice generation unit generates a voice corresponding to the emotion of the hearing-impaired person estimated by the emotion estimation unit for the characters obtained by interpreting the sign language motion of the hearing-impaired person using the generation AI for voice generation. The data processing device according to Appendix 1 or Appendix 2. (Appendix 4) The sign language interpretation unit translates the characters obtained by interpreting the sign language motion of the hearing-impaired person into characters of a pre-specified language and outputs the translated result. The voice generation unit generates the voice of the hearing-impaired person corresponding to the characters translated into the pre-specified language by the sign language interpretation unit using the generation AI for voice generation. The data processing device according to any one of Appendices 1 to 3. (Appendix 5) The generation AI for voice generation performs machine learning on the voice of the hearing-impaired person and the voice of a close relative of the hearing-impaired person. The data processing device according to any one of Appendices 1 to 4. (Appendix 6) Acquire a captured image of a sign language motion of a hearing-impaired person, Interpret the sign language movements of the hearing-impaired person from the captured image using a generative AI for sign language interpretation, convert the interpreted sign language movements of the hearing-impaired person into text and output it, Generate the voice of the hearing-impaired person corresponding to the text using a generative AI for voice generation, and output the generated voice of the hearing-impaired person. Data processing method. (Appendix 7) Obtain a captured image of the sign language movements of a hearing-impaired person, Interpret the sign language movements of the hearing-impaired person from the captured image using a generative AI for sign language interpretation, convert the interpreted sign language movements of the hearing-impaired person into text and output it, including a process of generating the voice of the hearing-impaired person corresponding to the text using a generative AI for voice generation and outputting the generated voice of the hearing-impaired person, A data processing program for causing a computer to execute.

Explanation of Signs

[0092] 10 Data processing system 12 Data processing device 14, 14A, 14B Smart device 58 Data generation model 58A Generative AI for sign language interpretation 58B Generative AI for voice generation 58C Generative AI for voice conversion 290 Specific processing unit 292 Input unit 294 Processing unit 296 Output unit 310 Image acquisition unit 312 Sign language interpretation unit 314 Voice generation unit 316 Voice conversion unit 318 Emotion estimation unit

Claims

1. An image acquisition unit that acquires a captured image of a sign language motion of a hearing-impaired person; A sign language interpretation unit that interprets the sign language motion of the hearing-impaired person from the captured image using a generative AI for sign language interpretation, and outputs the interpreted sign language motion of the hearing-impaired person as text; A voice generation unit that generates the voice of the hearing-impaired person corresponding to the text using a generative AI for voice generation, and outputs the generated voice of the hearing-impaired person; A data processing device comprising the above.

2. The data processing device according to claim 1, further comprising a voice conversion unit that converts the voice of a normal person into sign language or text and outputs it using a generative AI for voice conversion. The data processing device according to claim 1.

3. The data processing device according to claim 1, further comprising an emotion estimation unit that estimates the emotion of the hearing-impaired person from the captured image of the hearing-impaired person; The voice generation unit generates a voice corresponding to the emotion of the hearing-impaired person estimated by the emotion estimation unit for the text obtained by interpreting the sign language motion of the hearing-impaired person using the generative AI for voice generation. The data processing device according to claim 1.

4. The sign language interpretation unit translates and outputs the text obtained by interpreting the sign language motion of the hearing-impaired person into text in a pre-specified language; The voice generation unit generates the voice of the hearing-impaired person corresponding to the text translated into the pre-specified language by the sign language interpretation unit using the generative AI for voice generation. The data processing device according to claim 1.

5. The generative AI for voice generation performs machine learning on the voice of the hearing-impaired person and the voice of a close relative of the hearing-impaired person. The data processing device according to claim 1.

6. Acquire a captured image of a sign language motion of a hearing-impaired person; Interpret the sign language motion of the hearing-impaired person from the captured image using a generative AI for sign language interpretation, and output the interpreted sign language motion of the hearing-impaired person as text; Generate the voice of the hearing-impaired person corresponding to the text using a generative AI for voice generation, and output the generated voice of the hearing-impaired person. A data processing method.

7. Acquire a captured image of a sign language motion of a hearing-impaired person; Interpret the sign language motion of the hearing-impaired person from the captured image using a generative AI for sign language interpretation, and output the interpreted sign language motion of the hearing-impaired person as text; A data processing program for causing a computer to execute a process including generating the voice of the hearing-impaired person corresponding to the text using a generative AI for voice generation and outputting the generated voice of the hearing-impaired person. A data processing program for causing a computer to execute the above.

Citation Information

Patent Citations

  • Voice synthesizing device, method and program

    JP2020160319A

  • System and method for providing call service for the hearing impaired

    KR102487847B1

  • System, server device and program

    JP6840365B2