Sign language information generation device, method, program, and system

The system addresses the issue of biased data collection in sign language learning by suggesting word candidates based on user attributes and interests, increasing motivation through practical application, thus collecting diverse and unbiased data.

JP2025145860APending Publication Date: 2025-10-03RISO KAGAKU CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024046324
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing sign language learning applications fail to collect unbiased and balanced sign language video data due to user selection bias, leading to reduced motivation for data registration and limited data collection from a diverse range of individuals.

Method used

A system that provides sign language registrants with word candidates based on their attributes and interests, allowing them to film and register video data, while offering rewards and generating sentences using learned words to demonstrate conversational abilities.

Benefits of technology

Enhances motivation for sign language data registration by allowing users to see the practical application of learned words, thereby collecting diverse and unbiased data from a wider range of people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025145860000001_ABST
    Figure 2025145860000001_ABST
Patent Text Reader

Abstract

To provide a sign language video information collection device, a method, a program, and a system capable of increasing motivation for registering sign language video data and collecting sign language video data from many people evenly and without bias.SOLUTION: A sign language information generation device includes: a storage unit (15) for storing sign language video information corresponding to a plurality of registered words by a sign language registrant, a sentence generation unit (17) for generating a sentence using part or all of the plurality of registered words for which the sign language video information has been stored by the sign language registrant, and a sign language information output unit (19) for outputting information regarding the sentence generated by the sentence generation unit (17).SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sign language information generation device, method, program, and system for generating information about sign language. [Background technology]

[0002] Conventionally, a system has been proposed that recognizes sign language actions from video data of a sign language and displays the action as character data (see, for example, Non-Patent Document 1).

[0003] In order to build a sign language analysis and recognition system, it is necessary to collect a large amount of data on sign language actions. Patent Document 1 proposes a method for collecting video data of sign language by shooting sign language actions using a terminal device. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2020-126144 [Non-patent literature]

[0005] [Non-Patent Document 1] "SureTalk", SoftBank Corp., Internet

[0006] <URL:https: / / www.suretalk.mb.softbank.jp / function / > Summary of the Invention [Problem to be solved by the invention]

[0007] However, since there are individual differences in sign language movements, it is preferable to collect sign language video data of as many words as possible from as many people as possible. In Patent Document 1, the user selects the sign language they want to register based on their own free will, so it is not possible to collect sign language video data that is not biased and is balanced.

[0008] Therefore, the applicant is considering providing sign language registrants with an application for learning sign language, in which the sign language video data collectors would present the sign language registrants with candidates for words they would like to register, and the sign language registrants would then film and register the sign language video data for those words, thereby collecting sign language video data in an evenly and unbiased manner.

[0009] However, while the above-mentioned applications allow users to learn individual sign language words, they have the problem of not being able to realize the benefits of learning sign language, such as being able to converse in sign language. This may reduce the motivation of sign language registrants to learn, making it difficult to collect sign language video data using the above-mentioned applications.

[0010] In view of the above circumstances, the present invention aims to provide a sign language information generation device, method, program, and system that can increase motivation to register sign language video data and collect sign language video data from a large number of people in an even and unbiased manner. [Means for solving the problem]

[0011] The sign language information generation device of the present invention comprises a sign language video information storage unit in which a sign language registrant stores sign language video information corresponding to multiple registered words, a sentence generation unit that generates sentences using some or all of the multiple registered words for which sign language video information has been stored by the sign language registrant, and a sign language information output unit that outputs information regarding the sentences generated by the sentence generation unit. [Effects of the Invention]

[0012] According to the sign language information generation device of the present invention, a sign language registrant generates a sentence using some or all of a plurality of registered words stored in sign language video information, and outputs information about the generated sentence. This allows the sign language registrant to confirm the effect of their own learning, i.e., to realize the effect of being able to converse in sign language using the registered words they have learned, thereby increasing their motivation to register sign language video information. This makes it possible to collect sign language video data from a wide range of people without bias. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a block diagram showing a schematic configuration of an embodiment of a sign language video information collection system of the present invention. [Figure 2] FIG. 10 is a diagram showing an example of a user information table. [Figure 3] FIG. 10 shows an example of a table of the number of registrations by attribute and a table of the number of registrations by registrant. [Figure 4] FIG. 10 is a diagram showing an example of a registered word information table. [Figure 5] A flowchart for explaining an example of a method for determining registered word candidates. [Figure 6] A flowchart for explaining an example of a method for determining registered word candidates. [Figure 7] A diagram showing an example of a list of registered words before and after sorting based on the range of registered words. [Figure 8] FIG. 1 is a diagram showing a schematic diagram of the relationship between the range of all registered words, the range of insufficient registered words, the range of registered words belonging to user interest information, and the range of registered words of priority attribute information. [Figure 9] FIG. 10 is a diagram showing an example of a category association table. [Figure 10] A specific example of a table showing the number of registrations per registrant [Figure 11] A sequence diagram illustrating the processing flow of the sign language video information collection system shown in FIG. [Figure 12] A sequence diagram illustrating the processing flow of the sign language video information collection system shown in FIG. [Figure 13] A diagram showing an example of a list of registered word candidates. [Figure 14] A diagram showing an example of displaying video data of a sign language model DETAILED DESCRIPTION OF THE INVENTION

[0014] A sign language video information collection system 1 using an embodiment of a sign language information generation device of the present invention will be described in detail below with reference to the drawings. Fig. 1 is a block diagram showing a schematic configuration of the sign language video information collection system 1 of this embodiment.

[0015] The sign language video information collection system 1 of this embodiment is a system that collects video data of sign language. Specifically, the sign language video information collection system 1 of this embodiment is a system that uniformly collects sign language video information of various sign language registrants who differ in age, gender, physique, etc., and also allows sign language registrants to register sign language video information of their own sign language while learning sign language.

[0016] As shown in FIG. 1, a sign language video information collection system 1 of this embodiment includes a sign language video information collection device 10 and a terminal device 20 of a sign language registrant.

[0017] The sign language video information collection device 10 and the terminal device 20 of the sign language registrant are connected via a communication line such as an Internet line or a LAN (Local Area Network) line, and are configured to be able to exchange various information with each other. Note that although only one terminal device 20 of the sign language registrant is shown in Fig. 1, in reality, many terminal devices 20 of sign language registrants are connected to the sign language video information collection device 10 and register sign language video information in the sign language video information collection device 10.

[0018] Each device constituting the sign language video information collection system 1 will be described in more detail below.

[0019] As shown in Figure 1, the sign language video information collection device 10 includes an attribute information acquisition unit 11, a user interest information acquisition unit 12, a registered word candidate determination unit 13, a sign language video information acquisition unit 14, a memory unit 15, a reward information output unit 16, a sentence generation unit 17, a sign language video information generation unit 18, and a sign language information output unit 19.

[0020] The attribute information acquisition unit 11 acquires attribute information of a sign language registrant (hereinafter referred to as user attribute information). The user attribute information is information related to the sign language registrant, and includes, for example, at least one of identification information unique to the sign language registrant, gender, age, hometown, occupational information (including student information), purpose for learning sign language, physique information, and sign language speed information. Differences in this attribute information result in differences in sign language movements. For example, sign language movements differ depending on the individual, and also on gender and age. Furthermore, like dialects, sign language has regional characteristics, and the way of signing varies depending on the region. Furthermore, the way of signing varies depending on the occupation and purpose for learning sign language of the sign language registrant.

[0021] Furthermore, physical information includes, for example, height and weight. Such physical characteristics also affect the movements of sign language. Furthermore, sign language speed information is information related to the speed of sign language, and in this embodiment, the sign language registrant sets and inputs information about their self-evaluated speed using the terminal device 20. For example, the sign language registrant sets and inputs the sign language speed information by selecting one of three levels: "fast," "normal," or "slow."

[0022] The user interest information acquisition unit 12 acquires user interest information indicating a range to which a word that a sign language registrant is interested belongs. Specifically, the user interest information acquisition unit 12 of the present embodiment acquires category information and level information as the user interest information.

[0023] Category information is a range of words that is set in advance based on the similarity between words. Specifically, a "weather" category is set as a range to which words such as "rain," "sunny," and "cloudy" belong. A "greetings" category is set as a range to which words such as "hello," "good evening," and "good morning" belong. A "people / family" category is set as a range to which words such as "father," "mother," "friend," and "lover" belong. Other categories such as "color" and "direction" are also set.

[0024] The level information is a range of words that is set in advance based on the difficulty of expressing the words in sign language. In this embodiment, the level information is set based on the sign language proficiency test level. The sign language proficiency test level is the difficulty level in the sign language proficiency test, which is used to test proficiency in sign language. The sign language proficiency test is a test hosted by the NPO Sign Language Proficiency Certification Association to determine the proficiency level of sign language.

[0025] The Sign Language Proficiency Test has seven levels, from 1 to 7, and each level has a set number of target words. For example, there are about 100 words for level 6 and about 200 words for level 5. Levels 5, 6, and 7 are beginner level, levels 3 and 4 are intermediate level, and levels 1 and 2 are advanced level.

[0026] As the level information, a grade unit or a level unit of beginner level, intermediate level, and advanced level is set in advance.

[0027] The user attribute information acquired by the attribute information acquisition unit 11 and the user interest information acquired by the user interest information acquisition unit 12 are set and input into the terminal device 20 of the sign language registrant, and output from the terminal device 20 to the sign language video information collection device 10.

[0028] The sign language video information collection device 10 associates the input user attribute information and user interest information with each identification information of a sign language registrant and stores them in the storage unit 15. Fig. 2 is a diagram showing an example of a user information table in which the user attribute information and user interest information are associated and stored for each sign language registrant A to X. The user information table shown in Fig. 2 is stored in the storage unit 15.

[0029] The registration word candidate determination unit 13 determines registration word candidates to be requested to the sign language registrant for sign language registration based on the above-mentioned user attribute information and user interest information. Then, the registration word candidate determination unit 13 outputs the determined registration word candidates to the terminal device 20 of the sign language registrant.

[0030] Here, a method for determining registered word candidates in the registered word candidate determination unit 13 will be specifically described below. The basic flow of the method for determining registered word candidates in this embodiment is to first determine insufficiently registered words with a small number of registrations based on user attribute information, and then determine, from among the insufficiently registered words, words that the user is likely to be interested in and that the user has not yet learned as registered word candidates. First, a method for determining insufficiently registered words based on user attribute information will be described.

[0031] The registered word candidate determination unit 13 has an attribute-based registration count table that manages the number of registrations for each attribute information of a predetermined registered word. Fig. 3A is a diagram showing an example of the attribute-based registration count table. The attribute-based registration count table shown in Fig. 3A shows the number of registrations for each of different attributes 1 to N for each registered word a to n.

[0032] Specifically, for example, the number of registered words a by sign language registrants belonging to attribute 1 (male, 0-49 years old) is 60, the number of registered words a by sign language registrants belonging to attribute 2 (male, 50 years old or older) is 40, and the number of registered words a by sign language registrants belonging to the other attributes 3 to N is zero. The total number of registered words a is 100.

[0033] For example, the number of registered words b by sign language registrants belonging to attribute 3 (female, 0-49 years old) is 30, the number of registered words b by sign language registrants belonging to attribute 4 (female, 50 years old or older) is 30, and the number of registered words b by sign language registrants belonging to the other attributes 1, 2-5-N is zero. The total number of registered words b is 60.

[0034] That is, in the case of the attribute-by-registration count table shown in Fig. 3A, it can be seen that for registered word a, the number of registered sign language registrants belonging to attributes 3 to N is insufficient, and for registered word b, the number of registered sign language registrants belonging to attributes 1, 2, and 5 to N is insufficient. Also, for registered words d to n, the number of registrations for all attributes is zero, and it can be seen that the number of registrations for all attributes is insufficient. Note that, in the attribute-by-registration count table shown in Fig. 3A, attributes 1 to N are classified by combinations of age and gender, but this is not limiting, and attributes 1 to N may also be classified taking into consideration combinations of other attribute information described above.

[0035] The attribute-specific registration count table is updated whenever new sign language animation information of a registered word is registered.

[0036] Then, the registered word candidate determination unit 13 determines the missing registered words using the attribute-by-attribute registration count table described above. Specifically, for example, when the attribute information of the sign language registrant acquired by the attribute information acquisition unit 11 is attribute 1, the attribute-by-attribute registration count table shown in Fig. 3A is referenced, and registered words of attribute 1 whose registration count is equal to or less than a predetermined threshold are determined as missing registered words. For example, in the attribute-by-attribute registration count table shown in Fig. 3A, for attribute 1, the registration count of registered words b, d to n is zero. Therefore, registered words b, d to n are determined as missing registered words.

[0037] The registered word candidate determination unit 13 also has a registration number per registrant table that manages the number of registrations of predetermined registered words for each sign language registrant. Fig. 3B is a diagram showing an example of the registration number per registrant table. The registration number per registrant table shown in Fig. 3B shows the number of registrations for each registered word a to n for each different sign language registrant.

[0038] Specifically, for example, the number of registered words a by sign language registrant A is 3, the number of registered words b is 2, the number of registered words c is 1, and the number of registered words d to n is 0. In other words, it can be seen that sign language registrant A has already learned and registered words a to c, but has not yet learned registered words d to n.

[0039] The registered word candidate determination unit 13 first identifies a sign language registrant based on identification information included in the attribute information of the sign language registrant acquired by the attribute information acquisition unit 11. Then, the registered word candidate determination unit 13 refers to the registration count table for each registrant shown in FIG. 3B to confirm the registration count of each registered word of the identified sign language registrant, and determines registered words whose registration count is equal to or less than a predetermined threshold as unlearned registered words. For example, if the identified sign language registrant is sign language registrant A, the registration count of registered words d to n is zero, and therefore these registered words d to n are determined as unlearned registered words. Note that in this embodiment, registered words whose registration count is zero are determined as unlearned registered words, but the threshold is not limited to zero. For example, the threshold may be set to 3 or more, and registered words whose registration count is 3 or more may be determined as learned registered words, and registered words whose registration count is 2 or less may be determined as unlearned registered words.

[0040] Next, the registered word candidate determination unit 13 determines, based on the user interest information, registered words that are likely to interest the user from among the insufficient registered words determined as described above, as registered word candidates.

[0041] First, as a prerequisite, the registered word candidate determination unit 13 stores a registered word information table in which category information, level information, and priority attribute information are set for all registered words set in advance. Fig. 4 is a diagram showing an example of the registered word information table. The category information and level information set for each registered word are the same as the category information and level information included in the above-mentioned user interest information. On the other hand, the priority attribute information set for each registered word is information set in advance by the party collecting sign language video information, and is attribute information of users predicted to be interested in each registered word.

[0042] The registered word candidate determination unit 13 refers to the registered word information table shown in FIG. 4 based on the user interest information specified by the sign language registrant, thereby identifying registered words that the sign language registrant is likely to be interested in and determining them as registered word candidates.

[0043] 5 and 6 are flowcharts for explaining a method for determining registered word candidates based on user interest information. Also, Fig. 8 is a diagram schematically showing the relationship between a range R1 of all registered words preset in registered word candidate determination unit 13, a range R2 of insufficient registered words, ranges R3 and R4 of registered words belonging to user interest information, and a range R5 of registered words of priority attribute information including user attribute information of sign language registrants.

[0044] Range R3 indicates the range of registered words that belong to the category information specified by the sign language registrant among the user interest information, and range R4 indicates the range of registered words that belong to the level information specified by the sign language registrant among the user interest information. Also, in Figure 8, R6 indicates the range of registered words that have already been learned.

[0045] Hereinafter, a detailed description will be given with reference to the flowcharts shown in Figures 5 and 6, the list shown in Figure 7, and the Venn diagram shown in Figure 8. In this embodiment, as described above, there are three types of attributes: category information, level information, and user attribute information, and registered word candidates are determined in the order of priority: category information, level information, and user attribute information. Also, in Figure 8, I to VII, the smaller the number, the higher the priority.

[0046] As shown in the flowchart of Figure 5, the registered word candidate determination unit 13 first acquires the user attribute information and user interest information of the sign language registrant (S10), and based on this, checks for all registered words x whether they are unlearned and insufficiently registered words by the sign language registrant, and creates a list of unlearned and insufficiently registered words (S12).

[0047] Next, the registered word candidate determination unit 13 determines to which range of registered word ranges I to VII in the Venn diagram shown in Fig. 8 the unlearned and insufficiently registered words listed in the list belong (S14). Fig. 6 is a flowchart for explaining the flow of the registered word range determination process.

[0048] As shown in FIG. 6, the registered word candidate determination unit 13 first checks whether the category information of the registered word matches the category information of the user interest information (S20).

[0049] If the category information matches (S20, YES), it is checked whether the level information of the registered word matches the level information of the user interest information (S22). If the level information matches (S22, YES), it is checked whether the user attribute information of the sign language registrant is included in the priority attribute information of the registered word (S24). If the user attribute information of the sign language registrant is included in the priority attribute information of the registered word (S24, YES), it is determined that the registered word belongs to registered word range I (S26).

[0050] In S24, if the user attribute information of the sign language registrant is not included in the priority attribute information of the registered word (S24, NO), the registered word is determined to belong to registered word range II (S28).

[0051] In S22, if it is determined that the level information does not match (S22, NO), it is checked whether the user attribute information of the sign language registrant is included in the priority attribute information of the registered word (S30), and if it is included, it is determined that the registered word belongs to registered word range III (S32), and if it is not included, it is determined that the registered word belongs to registered word range IV (S34).

[0052] Furthermore, in S20, if the category information does not match (S20, NO), it is checked whether the level information matches (S36). If the level information matches (S36, YES), it is checked whether the user attribute information of the sign language registrant is included in the priority attribute information of the registered word (S38), and if the user attribute information of the sign language registrant is included in the priority attribute information of the registered word (S38, YES), it is determined that the registered word belongs to registered word range V (S40), and if not (S38, NO), it is determined that the registered word belongs to registered word range VI (S42).

[0053] In S36, if the level information does not match (S36, NO), it is checked whether the user attribute information of the sign language registrant is included in the priority attribute information of the registered word (S44), and if the user attribute information of the sign language registrant is included in the priority attribute information of the registered word (S44, YES), the registered word is determined to belong to registered word range VII (S46), and if not (S44, NO), the registered word is determined to not belong to any of registered word ranges I to VII.

[0054] Then, the judgment process from S20 to S46 is repeated until there are no unlearned and insufficiently registered words to be subjected to the judgment process (S48, NO), and the process ends when the judgment process for all unlearned and insufficiently registered words has been completed (S48, YES).

[0055] Returning to the flowchart shown in Fig. 5, the registered word candidate determination unit 13 sorts the list of unlearned and insufficiently registered words so that they are arranged in the order of registered word ranges I to VII (S16). Fig. 7A shows the list before sorting, and Fig. 7B shows the list after sorting. The lists in Fig. 7A and Fig. 7B indicate for each registered word whether it is unlearned, whether it is an insufficiently registered word, and the registered word range I to VII to which the registered word belongs.

[0056] Then, the registered word candidate determination unit 13 refers to the sorted list shown in FIG. 7B and determines the registered word candidates in the order of registered word ranges I to VII (S18).

[0057] The registered word candidate determination unit 13 determines registered word candidates as described above, and determines registered words that are insufficient to be registered and that the sign language registrant is interested in as registered word candidates. That is, registered words that belong to the shaded ranges in the Venn diagram shown in Fig. 8 are determined as registered word candidates in the order of registered word ranges I to VII. Note that in this embodiment, registered word candidates are determined using three types of attributes, namely, category information, level information, and user attribute information, but the attributes are not limited to these, and other attributes may be used.

[0058] Furthermore, if, as a result of determining registered word candidates according to the flowcharts shown in Figures 5 and 6, there are zero registered words that can be used as registered word candidates for a given sign language registrant, the range of category information and level information of the user interest information can be expanded, that is, by expanding the range R3 of category information and the range R4 of level information in the Venn diagram shown in Figure 8, the number of registered words that can be used as registered word candidates can be increased, and registered word candidates can be determined again according to the flowcharts shown in Figures 5 and 6.

[0059] As a method for expanding the range R3 of category information, for example, other category information (for example, "greetings") may be added to the category information included in the user interest information (for example, "weather"). As a method for determining the newly added category information, for example, category information that is highly related to the category information included in the user interest information may be determined as the newly added category information.

[0060] 9 is a diagram showing an example of a category relevance table in which the relevance between category information items is preset. The registered word candidate determination unit 13, for example, refers to the category relevance table shown in FIG. 9 and determines the category information to be newly added, giving priority to category information that is highly relevant to the category information of the user interest information. For example, if the category information of the user interest information is "weather," the most relevant "greetings" is determined first as the category information to be newly added. After that, "people / family," "colors," and "directions" are added in this order.

[0061] Then, the registration word candidate determination unit 13 outputs the final registration word candidates determined as described above to the terminal device 20 of the sign language registrant.

[0062] In response to the input of the registered word candidate, the sign language video information acquisition unit 14 acquires sign language video information corresponding to the registered word candidate output from the terminal device 20 of the sign language registrant. In this embodiment, the sign language registrant uses the terminal device 20 to demonstrate the sign language corresponding to the registered word candidate and films the demonstration. The terminal device 20 then extracts feature points of the sign language actions from the filmed sign language video data and outputs information on the feature points as sign language video information to the sign language video information collection device 10. The sign language video information acquisition unit 14 then acquires the sign language video information output from the terminal device 20 as described above.

[0063] The storage unit 15 stores the sign language video information corresponding to the registered word candidates acquired by the sign language video information acquisition unit 14 in association with the registered word candidates. At this time, the registration numbers in the registration number per attribute table and the registration number per registrant table are updated based on the stored registered word candidates and attribute information of the sign language registrant.

[0064] The storage unit 15 also stores sign language video data that serves as a model in which all registered words, including registered words that the sign language registrant has not yet learned, are expressed in sign language. In this embodiment, the storage unit 15 corresponds to the sign language video information storage unit of the present invention.

[0065] When the reward information output unit 16 acquires sign language video information output from the terminal device 20 of the sign language registrant, it outputs reward information to the terminal device 20 of the sign language registrant. In this embodiment, information on pet treats to be used in the pet raising game started on the terminal device 20 of the sign language registrant is output as the reward information, the details of which will be described later.

[0066] The sentence generation unit 17 identifies a plurality of registered words for which sign language video information has been registered by a predetermined sign language registrant from among all registered words, and generates a sentence using the identified plurality of registered words. Specifically, the sentence generation unit 17 refers to the registration count table for each registrant shown in Fig. 3B, identifies registered words with a registration count of 1 or more, and generates a sentence.

[0067] FIG. 10 is a table in which each registered word a to n in the registration count table for each registrant shown in FIG. 3B is replaced with a specific example of a registered word. In the registration count table for each registrant shown in FIG. 10, registered words such as "where," "when," "reason," "location," "explanation," "plan," "please," "until," "is it?", and "end" are preset. For example, for sign language registrant A, the number of registered words for "where" is three, and the number of registered words for "reason," "explanation," "please," "until," "is it?", and "end" is one. According to the registration count table for each registrant shown in FIG. 10, these already registered words are registered words that have already been learned by sign language registrant A. Furthermore, other registered words with a registration count of zero are registered words that have not yet been learned by sign language registrant A.

[0068] The sentence generation unit 17 refers to the registration number table for each registrant shown in FIG. 10, identifies learned registered words, and generates a sentence using some or all of the identified registered words. Specifically, the sentence generation unit 17 generates a sentence by inputting the learned registered words of sign language registrant A to a document generation AI (Artificial Intelligence). If there are many learned registered words, the sentence generation unit 17 selects a number (approximately 5 to 20) that will provide a sentence of an appropriate length for the sign language registrant to demonstrate the sign language, and generates the sentence. Examples of sentences generated by the sentence generation unit 17 include, for example, "How much have you finished?" and "Please explain the reason." The document generation AI may be, for example, a natural language processing AI such as ChatGPT (registered trademark), Microsoft 365 Copilot, or Gemini. However, the document generation AI is not limited to these, and other known technologies may be used to generate sentences.

[0069] Furthermore, in the above example, the sign language registrant generates a sentence using learned registered words, but the sentence may also be generated by including unlearned registered words. Specifically, for example, unlearned registered words related to learned registered words may be identified and included in the sentence. As unlearned registered words related to learned registered words, for example, registered words belonging to the same category information or the same level information as the learned registered word may be identified as unlearned registered words. Alternatively, unlearned registered words related to learned registered words may be identified using a table or the like that pre-sets highly related registered words. Related unlearned registered words are, for example, words that are often used together when "reason" has been learned and "explanation" has not been learned, and such words are highly related registered words.

[0070] In addition, for example, registered words that belong to the category information or level information of the user interest information of a sign language registrant and that have not yet been learned may be identified, and a predetermined number of registered words may be randomly selected from among them to be included in a sentence.

[0071] Alternatively, in this embodiment, sentences are generated mainly from registered words that have been learned by a sign language registrant, but this is not limiting, and sentences may be generated mainly from registered words that have not been learned. For example, the number of unregistered registered words may be greater than the number of learned registered words, and these may be input to the sentence generation AI to generate sentences.

[0072] Next, the sign language video information generation unit 18 generates sign language video data representing the sentence generated by the sentence generation unit 17. Specifically, the sign language video information generation unit 18 reads out model sign language video data representing each registered word that makes up the sentence generated by the sentence generation unit 17 from the storage unit 15, and connects this sign language video data to generate sign language video data representing the sentence.

[0073] The sign language information output unit 19 outputs the text data of each registered word that constitutes the sentence generated by the sentence generation unit 17 and the sign language video data representing the sentence generated by the sign language video information generation unit 18 to the terminal device 20 of the sign language registrant.

[0074] The sign language video information collection device 10 includes a CPU (Central Processing Unit), semiconductor memories such as ROM (Read Only Memory) and RAM (Random Access Memory), storage such as a hard disk, and a communication I / F (Interface).

[0075] An embodiment of a sign language video information collection program including a sign language information generation program of the present invention is installed in the storage of the sign language video information collection device 10. When this sign language video information collection program is started by the CPU, the functions of the above-mentioned parts of the sign language video information collection device 10 are executed.

[0076] In addition, in this embodiment, the functions of each part are performed by executing the sign language video information collection program using a CPU, but some or all of the functions performed by the sign language video information collection program may also be configured using hardware such as an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other electrical circuits.

[0077] Next, the terminal device 20 of the sign language registrant will be described.

[0078] As described above, the terminal device 20 of the sign language registrant is used by the sign language registrant, and is configured as a mobile terminal such as a tablet terminal or a smartphone. However, it is not limited to this, and may be configured as a personal computer.

[0079] 1, the terminal device 20 includes a control unit 21, a display unit 22, a storage unit 23, an input unit 24, and an image capturing unit 25. A sign language learning application is installed in the storage unit 23 of the terminal device 20 of this embodiment.

[0080] The control unit 21 controls the entire terminal device 20. In particular, the control unit 21 starts a sign language learning application installed in the storage unit 23, thereby executing functions such as displaying registered word candidates, accepting selection, taking video of sign language, outputting sign language video information, and playing (displaying) the sign language video information.

[0081] In addition to the sign language learning function, the sign language learning application also has a pet raising game function, in which the player raises a pet by feeding it treats.

[0082] Furthermore, the control unit 21 performs a feature point extraction process to extract feature points of sign language movements from the sign language video data captured by the capture unit 25. For example, the positions of joints and fingertips related to sign language are extracted as feature points. In addition to hand movements, feature points may also be extracted from facial movements, expressions, and movements involving the arms and body. Existing image processing can be used for the feature point extraction process, but feature points may also be extracted using a machine learning model that has previously learned feature points by machine learning.

[0083] Furthermore, when extracting feature points, more accurate data can be collected by performing a coordinate correction process, i.e., a normalization process, taking into account the vertical and horizontal positions, orientation, size, etc., depending on the positional relationship between the sign language registrant and the imaging unit 25.

[0084] Furthermore, the control unit 21 receives sign language video information representing a sentence output from the sign language video information collection device 10, and causes the display unit 22 to display the sign language video based on the sign language video information.

[0085] The display unit 22 displays a list of registered word candidates, displays a sign language video model of a predetermined registered word candidate, and displays a sign language video representing the above-mentioned sentence.

[0086] As described above, the sign language learning application is installed in the storage unit 23, and model sign language video data is also stored in the storage unit 23. The storage unit 23 stores all registered words (candidate registered words) and model sign language video data for the registered words in association with each other.

[0087] The input unit 24 accepts various setting inputs from a sign language registrant.

[0088] The image capturing unit 25 has a CMOS (Complementary Metal Oxide Semiconductor) camera, a CCD (Charge Coupled Devices) camera, an image capturing optical system, etc., and captures the sign language of the sign language registrant. The sign language video data captured by the image capturing unit 25 is stored in the storage unit 23, and then the control unit 21 performs feature point extraction processing.

[0089] The sign language learning application may be installed in the storage unit 23 as in this embodiment, or may be an application provided via a web browser.

[0090] Next, the processing flow of the sign language video information collection system 1 of this embodiment will be described with reference to the flowcharts shown in FIGS.

[0091] First, the sign language registrant starts the sign language learning application using the terminal device 20 (S60). Then, when the sign language registrant uses the sign language learning application for the first time, the user attribute information and user interest information of the sign language registrant are set and input in the terminal device 20 (S62).

[0092] The user attribute information and user interest information set and input on the terminal device 20 are acquired from the terminal device 20 by the attribute information acquisition unit 11 and the user interest information acquisition unit 12 of the sign language video information collection device 10, and are registered in association with the identification information of the sign language registrant (S64). Note that if the sign language registrant has used a sign language learning application in the past and the attribute information of the sign language registrant has already been registered, the attribute information acquisition unit 11 reads out the user attribute information and user interest information of the registered sign language registrant based on the identification information of the sign language registrant set and input on the terminal device 20.

[0093] Then, the registration word candidate determination unit 13 of the sign language video information collection device 10 determines registration word candidates to be requested to be registered in sign language from the sign language registrant based on the attribute information and user interest information of the sign language registrant (S66), as described above, and outputs the registration word candidates to the terminal device 20 (S68).

[0094] The terminal device 20 receives the sign language registration request and the registration word candidates output from the sign language video information collection device 10 (S70). Next, the terminal device 20 displays the received registration word candidates in a list (range L shown in FIG. 13) on the display unit 22, as shown in FIG. 13 (S72). The sign language registrant selects the registration word candidate they wish to learn from the multiple registration word candidates displayed on the display unit 22 of the terminal device 20 (S74).

[0095] When a predetermined registered word candidate is selected by the sign language registrant, the control unit 21 of the terminal device 20 reads out the model sign language video data corresponding to the registered word candidate from the storage unit 23 and displays it on the display unit 22 as shown in Figure 14 (S76).

[0096] The sign language registrant then demonstrates the sign language of the selected registration word candidate while viewing the model sign language video data displayed on the display unit 22, and the act is photographed by the photographing unit 25 (S78). At this time, as shown in Fig. 14, the model sign language video data may be displayed on the main screen M while the photographed data of the sign language registrant is displayed on the thumbnail screen S behind it. Furthermore, by selecting a switching button K, the display on the main screen M and the display on the thumbnail screen S may be switched.

[0097] Then, the control unit 21 performs a feature point extraction process on the sign language video data captured by the imaging unit 25 to extract feature points (S80) and acquires the feature points as sign language video information. The control unit 21 outputs the acquired sign language video information to the sign language video information collection device 10 (S82).

[0098] The sign language video information output from the terminal device 20 is acquired by the sign language video information acquisition unit 14 together with information on the registered word candidates (registered words), and the registered word candidates (registered words) and the sign language video information are linked and stored in the storage unit 15 (S84). At this time, the registration numbers in the registration number per attribute table and the registration number per registrant table are updated based on the stored registered word candidates and user attribute information of the sign language registrant (S86).

[0099] Furthermore, the reward information output unit 16 of the sign language video information collection device 10 outputs information about pet treats to be used in the pet raising game as reward information to the terminal device 20 (S88). The terminal device 20 stores the received information about the treats and makes it available for use when the pet raising game is started.

[0100] Furthermore, in order to check the effectiveness of his / her sign language learning, the sign language registrant uses the terminal device 20 to send a request for output of text sign language information to the sign language video information collection device 10 (S90).

[0101] In response to a request to output sentence sign language information, the sign language video information collection device 10 generates a sentence using registered words that have been learned by the sign language registrant (S92), and generates sign language video data that represents the sentence (S94). The generation of the sentence in S92 is as described above for the sentence generation unit 17, and although a case where learned registered words are used is described, unlearned registered words may also be included.

[0102] Then, the sign language video information collection device 10 outputs to the terminal device 20 of the sign language registrant the text data of each registered word that constitutes the sentence generated in S92 and the sign language video data representing the sentence (S96), and the terminal device 20 displays the input text data of each registered word and generates and displays a sign language video based on the input sign language video data of the sentence (S98). The display in S98 can be realized in the same manner as in S76 described above. Also, as described in S78, the sign language registrant may sign the sentence while looking at the displayed model sign language video data, and film the process.

[0103] In the sign language video information collection system 1 of this embodiment, the sign language video information collection device 10 outputs to the terminal device 20 the text data of each registered word that constitutes a sentence and the sign language video data representing the sentence, but it is also possible to output only one of these to the terminal device 20.

[0104] In addition, in the sign language video information collection system 1 of this embodiment, when the "random" tag shown in FIG. 14 is selected, the registered word candidates determined according to the flowcharts shown in FIG. 5 and FIG. 6 are displayed as a list. However, other tags such as "category," "level," "child," and "female" may be further provided, and different lists of registered word candidates may be displayed for each tag. For example, when the "category" tag is selected, registered words belonging to the category information of the user interest information that are unlearned and insufficiently registered may be displayed as a list. When the "level" tag is selected, registered words belonging to the category information of the user interest information that are unlearned and insufficiently registered may be displayed as a list. Furthermore, when the "child" tag is selected, registered words whose priority attribute information includes the child's age may be selected and displayed as a list. When the "female" tag is selected, registered words whose priority attribute information includes the gender "female" may be selected and displayed as a list.

[0105] According to the sign language video information collection system 1 of the above embodiment, a sentence is generated using a plurality of registered words learned by a sign language registrant, and information about the generated sentence is output to the terminal device 20 of the sign language registrant. This allows the sign language registrant to confirm the effect of their own learning, i.e., to realize the effect of being able to converse in sign language using the registered words they have learned, thereby increasing the motivation of the sign language registrant to register sign language video information. This makes it possible to collect sign language video information from a wide range of people in an unbiased manner.

[0106] In addition, in the sign language video information collection system 1 of the above embodiment, sign language video data representing sentences is generated, and the generated sign language video data is output to and displayed on the terminal device 20 of the sign language registrant, so that the sign language registrant can learn sign language that represents sentences, rather than just registered words.

[0107] Furthermore, in the sign language video information collection system 1 of the above embodiment, the sign language registrant generates sentences including registered words that have not yet been learned, which increases the variety of sentences and allows the sign language registrant to learn the sign language for registered words that have not yet been learned.

[0108] Furthermore, in the sign language video information collection system 1 of the above embodiment, registered words related to learned registered words are used as unlearned registered words, so more appropriate sentences can be created. Learning unlearned registered words related to learned registered words as sentences provides an opportunity to learn new registered words, and also has the effect of increasing opportunities to utilize learned registered words.

[0109] Furthermore, in the sign language video information collection system 1 of the above embodiment, registered words that the sign language registrant is interested in are used as unlearned registered words, which can further increase the sign language registrant's motivation to learn.

[0110] Furthermore, in the sign language video information collection system 1 of the above embodiment, registered word candidates are determined based on the user attribute information of the sign language registrant, so that sign language video information can be collected without bias in the number of registrations based on the attributes of the sign language registrant.

[0111] Furthermore, in the sign language video information collection system 1 of the above embodiment, information on the feature points of sign language actions extracted from sign language videos is obtained as sign language video information, i.e., information on the feature points of sign language actions is output from the terminal device 20 to the sign language video information collection device 10, which reduces the amount of data transmitted and protects the privacy of sign language registrants because the sign language image data itself is not output.

[0112] However, the sign language video data itself may be output as sign language video information from the terminal device 20 to the sign language video information collection device 10, and the feature point extraction process may be performed in the sign language video information collection device 10.

[0113] Furthermore, in the sign language video information collection system 1 of the above embodiment, when sign language video information output from the terminal device 20 of a sign language registrant is acquired, information about treats is output as reward information to the terminal device 20 of the sign language registrant, so that the sign language registrant can learn sign language while having fun by using the treat information in the pet raising game, which motivates them to continue learning sign language. Furthermore, from the perspective of collecting sign language video information, it is possible to encourage the sign language registrant to film sign language and output sign language video information.

[0114] In addition, the success reward information output from the sign language video information collection device 10 to the terminal device 20 is not limited to information about treats used in the pet raising game as in the above embodiment, but may also be information such as points that can be used for shopping on e-commerce sites or miles that can be used by airlines, etc.

[0115] Furthermore, the present invention is not limited to the above-described embodiments, and the components can be modified and embodied in practice without departing from the spirit of the invention. Furthermore, various inventions can be formed by appropriately combining multiple components disclosed in the above-described embodiments. For example, all of the components shown in the embodiments can be appropriately combined. Naturally, various modifications and applications are possible within the scope of the invention.

[0116] The present invention further discloses the following supplementary notes.

[0117] (Appendix 1) The sign language information generation device of the present invention comprises a sign language video information storage unit in which a sign language registrant stores sign language video information corresponding to multiple registered words, a sentence generation unit that generates sentences using some or all of the multiple registered words for which sign language video information has been stored by the sign language registrant, and a sign language information output unit that outputs information related to the sentences generated by the sentence generation unit.

[0118] (Appendix 2) The sign language information generation device described in Appendix 1 is provided with a sign language video information generation unit that generates sign language video information related to the sentence generated by the sentence generation unit, and a sign language information output unit that outputs the sign language video information generated by the sign language video information generation unit to a sign language registrant.

[0119] (Appendix 3) In the sign language information generation device described in Supplementary Note 1 or 2, the sentence generation unit can generate sentences including unregistered registered words for which sign language video information by a sign language registrant is not stored.

[0120] (Appendix 4) In the sign language information generation device described in Supplementary Note 3, the sentence generation unit can identify, as the unregistered registered word, a registered word related to a registered word for which the sign language video information is stored.

[0121] (Appendix 5) In the sign language information generation device described in Supplementary Note 3 or 4, the sentence generation unit can identify, as unregistered registered words, registered words in which the sign language registrant is interested.

[0122] (Appendix 6) The sign language information generation system of the present invention includes a sign language information generation device according to any one of Supplementary Notes 1 to 5, and a terminal device that receives and displays information about a sentence generated by the sign language information generation device.

[0123] (Appendix 7) The sign language information generation method of the present invention involves a sign language registrant storing sign language video information corresponding to a plurality of registered words, generating a sentence using some or all of the plurality of registered words for which the sign language registrant has stored sign language video information, and outputting information about the generated sentence.

[0124] (Appendix 8) The sign language information generation program of the present invention causes a computer to execute the steps of: a step in which a sign language registrant stores sign language video information corresponding to a plurality of registered words; a step in which the sign language registrant generates a sentence using some or all of the plurality of registered words for which sign language video information has been stored; and a step in which information regarding the generated sentence is output. [Explanation of symbols]

[0125] 1. Sign language video information collection system 10 Sign language video information collection device 11 Attribute information acquisition section 12 User interest information acquisition unit 13 Registered word candidate determination unit 14 Sign language video information acquisition unit 15 Storage section 16 Reward information output section 17 Sentence generation section 18 Sign language video information generation unit 19 Sign language information output unit 20 Terminal equipment 21 Control section 22 Display section 23 Memory section 24 Input section 25 Photography Department

Claims

1. a sign language video information storage unit in which a sign language registrant stores sign language video information corresponding to a plurality of registered words; a sentence generation unit that generates a sentence by the sign language registrant using a part or all of a plurality of registered words stored in the sign language video information; A sign language information generation device comprising a sign language information output unit that outputs information about the sentence generated by the sentence generation unit.

2. a sign language video information generation unit that generates sign language video information related to the sentence generated by the sentence generation unit; The sign language information generation device wherein the sign language information output unit outputs the sign language video information generated by the sign language video information generation unit to the sign language registrant.

3. 2. The sign language information generating device according to claim 1, wherein the sentence generating unit generates a sentence including an unregistered registered word for which sign language video information by the sign language registrant is not stored.

4. 4. The sign language information generating device according to claim 3, wherein the sentence generating unit specifies, as the unregistered registered word, a registered word related to a registered word stored in the sign language video information.

5. 4. The sign language information generating device according to claim 3, wherein the sentence generating unit identifies, as the unregistered registered word, a registered word in which the sign language registrant is interested.

6. The sign language information generation device according to claim 1; and a terminal device that receives and displays information about the sentence generated by the sign language information generation device.

7. A sign language registrant stores sign language video information corresponding to a plurality of registered words; A sentence is generated by the sign language registrant using a part or all of the plurality of registered words stored in the sign language video information; A sign language information generation method that outputs information about the generated sentence.

8. A step in which a sign language registrant stores sign language video information corresponding to a plurality of registered words; A step of generating a sentence by the sign language registrant using a part or all of the plurality of registered words stored in the sign language video information; and a step of outputting information about the generated sentence.

Citation Information

Patent Citations

  • System, server device, and program

    JP2020126144A