Electronic apparatus that provide sign language practice functionality and the operating method thereof

An electronic device with a sign language database and real-time feedback mechanism helps users learn accurate sign language gestures by comparing their movements with pre-recorded gestures, addressing the challenge of feedback accuracy in learning sign language.

KR1020260113804APending Publication Date: 2026-07-21TJHE HANCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
TJHE HANCOMM INC
Filing Date
2025-01-14
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Individuals learning sign language face challenges in determining the accuracy of their movements due to the lack of effective feedback, making it difficult to learn accurate sign language gestures.

Method used

An electronic device equipped with a sign language database, a split display unit, an acquisition unit, a similarity calculation unit, and a judgment unit to provide real-time feedback on the accuracy of sign language movements by comparing user performances with pre-recorded gestures.

Benefits of technology

The device supports users in learning accurate sign language movements by providing immediate feedback, enhancing the learning experience and improving gesture accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PAT00006_ABST
    Figure PAT00006_ABST
Patent Text Reader

Abstract

The present invention provides an electronic device that provides a sign language practice function and a method of operation thereof, thereby supporting users to learn accurate sign language movements more effectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an electronic device that provides a sign language practice function and a method of operation thereof. Background Technology

[0002] Sign language is a visual language that conveys meaning through the movements of hands or fingers, and it is an important means of communication used by people who cannot hear or speak due to hearing impairments.

[0003] Recently, with the growing social interest in sign language, the number of non-disabled people, as well as the hearing impaired, wishing to learn sign language is gradually increasing.

[0004] Previously, when people wanted to learn sign language, they had to learn it by watching specific sign language learning materials or sign language videos and following along. However, because it was difficult for people to know whether they were performing the sign language movements correctly, they had difficulty learning accurate sign language movements.

[0005] In this regard, it is necessary to introduce technology that supports users in receiving rapid feedback on the accuracy of the sign language movements they have performed.

[0006] Specifically, research is needed on a technology that supports a user in learning accurate sign language movements more effectively by providing a sign language movement video corresponding to a predetermined practice text, acquiring a video recording of the user performing the sign language movement according to the video, and determining whether the user has performed the sign language movement accurately by measuring the similarity between the sign language movement video and the video recording. The problem to be solved

[0007] The present invention aims to support users in learning accurate sign language movements more effectively by presenting an electronic device that provides a sign language practice function and a method of operation thereof. means of solving the problem

[0008] An electronic device according to an embodiment of the present invention comprises: a sign language database in which a plurality of predetermined practice texts and sign language movement images corresponding to each of the plurality of practice texts are stored; a split display unit that, when a command to start practicing a sign language movement corresponding to a first practice text, which is one of the plurality of practice texts, is applied from a user, extracts the first practice text and a first sign language movement image corresponding to the first practice text from the sign language database, displays the first practice text in a pre-set first area on a screen, and displays the first sign language movement image in a pre-set second area on a screen; an acquisition unit that, when a command to start shooting requesting the user to shoot their own sign language movement is applied, activates a camera mounted on the electronic device to start shooting the user's sign language movement, and when a command to end shooting requesting the user to end shooting is applied, deactivates the camera and ends shooting, thereby acquiring a first captured image of the user's sign language movement; a similarity calculation unit that calculates a similarity between the first sign language movement image and the first captured image; and the calculated similarity It includes a judgment unit that checks whether a preset threshold is exceeded, and if it is confirmed that it is exceeded, displays an accuracy guidance message on the screen indicating that the user's sign language gesture is accurate, and if it is confirmed that it is not exceeded, displays an inaccuracy guidance message on the screen indicating that the user's sign language gesture is inaccurate.

[0009] In addition, a method of operating an electronic device according to an embodiment of the present invention comprises: maintaining a sign language database in which a plurality of predetermined practice texts and sign language action images corresponding to each of the plurality of practice texts are stored; when a command to start practicing a sign language action corresponding to a first practice text, which is one of the plurality of practice texts, is applied from a user, extracting the first practice text and a first sign language action image corresponding to the first practice text from the sign language database, then displaying the first practice text in a pre-set first area on the screen and displaying the first sign language action image in a pre-set second area on the screen; when a command to start shooting requesting to shoot one's own sign language action is applied from the user, activating a camera mounted on the electronic device to start shooting the user's sign language action, and when a command to end shooting requesting to end the shooting is applied from the user, deactivating the camera and ending the shooting, thereby obtaining a first captured image of the user's sign language action; calculating a similarity between the first sign language action image and the first captured image; and the calculated similarity is a dictionary The method includes the step of checking whether the set threshold is exceeded, and if it is confirmed that it is exceeded, displaying an accuracy guidance message on the screen indicating that the user's sign language gesture is accurate, and if it is confirmed that it is not exceeded, displaying an inaccuracy guidance message on the screen indicating that the user's sign language gesture is inaccurate. Effects of the invention

[0010] The present invention provides an electronic device that provides a sign language practice function and a method of operation thereof, thereby supporting users to learn accurate sign language movements more effectively. Brief explanation of the drawing

[0011] FIG. 1 is a drawing illustrating the structure of an electronic device according to an embodiment of the present invention. FIGS. 2 and 3 are drawings for explaining an electronic device according to an embodiment of the present invention. FIG. 4 is a flowchart illustrating a method of operation of an electronic device according to an embodiment of the present invention. Specific details for implementing the invention

[0012] Embodiments according to the present invention will be described in detail below with reference to the accompanying drawings. This description is not intended to limit the present invention to specific embodiments and should be understood to include all modifications, equivalents, and substitutions that fall within the spirit and scope of the present invention. Similar reference numerals have been used for similar components in describing each drawing, and unless otherwise defined, all terms used in this specification, including technical or scientific terms, have the same meaning as generally understood by a person skilled in the art to which the present invention pertains.

[0013] In this document, when a part is described as "including" a component, it means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, in various embodiments of the present invention, each component, functional block, or means may be composed of one or more sub-components, and the electrical, electronic, or mechanical functions performed by each component may be implemented by various known devices or mechanical elements, such as electronic circuits, integrated circuits, and ASICs (Application Specific Integrated Circuits), and may be implemented separately or two or more may be integrated into one.

[0014] Meanwhile, the blocks in the attached block diagram or the steps in the flowchart may be interpreted as computer program instructions that perform designated functions by being loaded into the processor or memory of data-processing equipment, such as general-purpose computers, specialized computers, portable notebook computers, and network computers. Since these computer program instructions may be stored in memory provided in a computer device or in memory readable by a computer, the functions described in the blocks in the block diagram or the steps in the flowchart may be produced as manufactured products containing means of instruction to perform them. Furthermore, each block or each step may represent a module, segment, or part of code containing one or more executable instructions for executing a specific logical function(s). Also, it should be noted that in some alternative embodiments, the functions mentioned in the blocks or steps may be executed in a different order than the prescribed order. For example, two blocks or steps shown in succession may be performed substantially simultaneously or in reverse order, and in some cases, some blocks or steps may be omitted.

[0015] FIG. 1 is a drawing illustrating the structure of an electronic device according to an embodiment of the present invention.

[0016] Referring to FIG. 1, an electronic device (110) according to one embodiment of the present invention includes a sign language database (111), a segmented display unit (112), an acquisition unit (113), a similarity calculation unit (114), and a judgment unit (115).

[0017] The sign language database (111) stores a plurality of pre-specified practice texts and sign language action videos corresponding to each of the plurality of practice texts.

[0018] Here, the aforementioned multiple practice texts may consist of specific short sentences or expressions frequently used in daily life. Additionally, the sign language gesture video corresponding to each of the aforementioned multiple practice texts may consist of a video in which a specific expert, such as a sign language interpreter, expresses each practice text in sign language.

[0019] For example, the sign language database (111) may have information stored as shown in Table 1 below.

[0021] Practice text Sign language movement video Practice Text 1 Sign language gesture video 1 Practice Text 2 Sign language gesture video 2 Practice Text 3 Sign language gesture video 3 ... ...

[0023] When a split display unit (112) receives a command from a user to start practicing a sign language action corresponding to a first practice text, which is one of the plurality of practice texts, it extracts the first practice text and the first sign language action video corresponding to the first practice text from the sign language database (111), displays the first practice text in a preset first area on the screen, and displays the first sign language action video in a preset second area on the screen.

[0024] For example, let us assume that information such as Table 1 is stored in the sign language database (111), as in the example above.

[0025] At this time, when a command to start practicing a sign language action corresponding to a first practice text, 'practice text 1', which is one of the plurality of practice texts, 'practice text 1, practice text 2, practice text 3, ...', is applied to the electronic device (110) of the present invention, the split display unit (112) can extract the first practice text, 'practice text 1', and the first sign language action image, 'sign language action image 1', which corresponds to the first practice text, from the sign language database (111).

[0026] Then, as shown in the figure in FIG. 2, the split display unit (112) can display the first practice text, 'Practice Text 1', in a preset first area (210) on the screen and the first sign language action video, 'Sign Language Action Video 1', in a preset second area (220) on the screen.

[0027] When the first practice text and the first sign language movement video are displayed on the screen in this way, the user can familiarize themselves with the sign language movement by viewing the first sign language movement video displayed on the screen multiple times or by following the sign language movement according to the first sign language movement video.

[0028] Then, the user may apply a shooting start command to the electronic device (110) of the present invention requesting to film their sign language movements, thereby allowing a process to proceed on the electronic device (110) to determine the accuracy of their sign language movements.

[0029] When a shooting start command requesting the user to film their sign language movements is applied to the electronic device (110) in this manner, the acquisition unit (113) activates the camera mounted on the electronic device (110) to start filming the user's sign language movements, and then when a shooting end command requesting the user to stop filming the sign language movements is applied, the camera is deactivated and the filming is stopped, thereby acquiring a first filmed image of the user's sign language movements.

[0030] In this regard, when the above-mentioned shooting start command is applied to the electronic device (110) of the present invention from the above-mentioned user, the acquisition unit (113) can activate the camera mounted on the electronic device (110) to start shooting the user's speech motion.

[0031] Then, the user can apply a shooting termination command to the electronic device (110) of the present invention, requesting to end the shooting of the sign language gesture, after performing the sign language gesture they have learned toward the camera.

[0032] When the shooting termination command is applied to the electronic device (110) from the user, the acquisition unit (113) disables the camera and terminates the shooting, thereby acquiring a first shooting image of the user's handwriting motion.

[0033] When the first captured image is obtained through the acquisition unit (113) in this way, the similarity calculation unit (114) calculates the similarity between the first handwriting motion image and the first captured image.

[0034] In this case, according to one embodiment of the present invention, the similarity calculation unit (114) may include a generation unit (116), a playback section pair configuration unit (117), a section similarity calculation unit (118), and a calculation processing unit (119).

[0035] The generating unit (116) generates k first playback segments by dividing the entire playback segment constituting the first sign language motion video into k playback segments (k is a natural number greater than or equal to 2), and generates k second playback segments by dividing the entire playback segment constituting the first captured video into k playback segments.

[0036] For example, if k is '5', the generating unit (116) can divide the entire playback section constituting the first sign language action video into five equal playback sections. If the playback time of the first sign language action video is '20 seconds', the generating unit (116) can generate five first playback sections, such as 'playback section 1, playback section 2, playback section 3, playback section 4, and playback section 5', by dividing the entire playback section constituting the first sign language action video based on the playback points being '4 seconds, 8 seconds, 12 seconds, and 16 seconds'.

[0037] In this way, the generating unit (116) can also generate five second playback sections, such as 'playback section A, playback section B, playback section C, playback section D, and playback section E', by dividing the entire playback section constituting the first captured image into five playback sections, evenly for the first captured image.

[0038] The playback section pair configuration part (117) is configured with k playback section pairs by sequentially matching the k first playback sections and the k second playback sections one by one.

[0039] For example, as in the example above, if k is set to '5' and the five first playback sections are generated as 'playback section 1, playback section 2, playback section 3, playback section 4, playback section 5' and the five second playback sections are generated as 'playback section A, playback section B, playback section C, playback section D, playback section E', then the playback section pair component (117) can generate a total of five playback section pairs by sequentially matching each playback section one by one as '(playback section 1, playback section A)', '(playback section 2, playback section B)', '(playback section 3, playback section C)', '(playback section 4, playback section D)', and '(playback section 5, playback section E)'.

[0040] The section similarity calculation unit (118) calculates the section similarity for each of the k pairs of playback sections by calculating the section similarity between the two playback sections constituting each pair of playback sections for each of the k pairs of playback sections.

[0041] In this case, according to one embodiment of the present invention, the segment similarity calculation unit (118) may include a sampling unit (120), a frame pair configuration unit (121), an image similarity calculation unit (122), and a segment similarity calculation processing unit (123).

[0042] The sampling unit (120) performs a process to calculate the segment similarity for each of the k pairs of playback segments sequentially, and when it is time to calculate the segment similarity between two playback segments constituting a first pair of playback segments, which is one of the k pairs of playback segments, for one of the two playback segments constituting the first pair of playback segments, n first frames are generated by sampling n frames (n is a natural number greater than or equal to 5) among the plurality of frames constituting the playback segment at equal intervals, and for the other of the two playback segments constituting the first pair of playback segments, n second frames are generated by sampling n frames among the plurality of frames constituting the playback segment at equal intervals.

[0043] The frame pair configuration unit (121) is configured with n frame pairs by sequentially matching the n first frames and the n second frames one by one.

[0044] The image similarity calculation unit (122) calculates the image similarity for each of the n pairs of frames by calculating the image similarity between the two frames constituting each pair of frames for each of the n pairs of frames.

[0045] At this time, according to one embodiment of the present invention, the image similarity calculation unit (122) performs a process for calculating image similarity one by one for each of the n frame pairs sequentially, and when it is the order to calculate image similarity between two frames constituting a first frame pair, which is one of the n frame pairs, for one of the two frames constituting the first frame pair, the frame is divided equally into (txt) first divided regions by dividing the frame into t horizontal (t is a natural number greater than or equal to 10) and t vertical regions, and then for each of the first divided regions, a feature vector corresponding to each of the first divided regions is generated by constructing a 3-dimensional vector having the R value, G value, and B value of the pixel located at the center point of each divided region as components, and for the other frame constituting the first frame pair, the frame is divided equally into t horizontal and t vertical regions to generate (txt) second divided regions, and then the second divided For each of the regions, a feature vector corresponding to each of the second divided regions is generated by constructing a 3-dimensional vector having the R, G, and B values ​​of the pixel located at the center point of each divided region as components, and then a total of (txt) vector similarities are calculated by computing the vector similarity of the feature vectors between divided regions at the same location in the first divided regions and the second divided regions, and then a reference matrix of size (txt) having each of the calculated vector similarities as components is generated, and then the Frobenius Norm of the reference matrix can be calculated as the image similarity between the two frames constituting the first frame pair. Here, vector similarity can be calculated using methods such as Euclidean distance or cosine similarity.

[0046] The section similarity calculation processing unit (123) calculates the overall average of the image similarity for each of the n pairs of frames as the section similarity between the two playback sections constituting the first pair of playback sections.

[0047] Below, the operation of the sampling unit (120), the frame pair configuration unit (121), the image similarity calculation unit (122), and the segment similarity calculation processing unit (123) will be explained in detail with examples.

[0048] First, let k be '5' as in the example above, and assume that the above 5 pairs of playback intervals were generated as '(playback interval 1, playback interval A)', '(playback interval 2, playback interval B)', '(playback interval 3, playback interval C)', '(playback interval 4, playback interval D)', and '(playback interval 5, playback interval E)'.

[0049] Under these conditions, the process of calculating the segment similarity between 'playback segment 1' and 'playback segment A', which are two playback segments constituting the first playback segment pair '(playback segment 1, playback segment A)', one of the five playback segment pairs mentioned above, is explained as follows.

[0050] First, the sampling unit (120) can generate n first frames by sampling n frames at equal intervals from among the plurality of frames constituting 'playback section 1', which is one of the two playback sections constituting the first playback section pair '(playback section 1, playback section A)'. For example, if n is '100', the sampling unit (120) can generate 100 first frames such as 'frame 1, frame 2, frame 3, ...' by sampling 100 frames from among the plurality of frames constituting 'playback section 1' at equal intervals.

[0051] Additionally, the sampling unit (120) can generate 100 second frames such as ‘Frame A, Frame B, Frame C, ...’ by sampling 100 frames among the plurality of frames constituting ‘Frame A, Frame B, Frame C, ...’ at equal intervals for ‘Frame A,’ which is the other of the two playback sections constituting the first playback section pair of ‘(Playback Section 1, Playback Section A)’.

[0052] Then, the frame pair configuration part (121) can sequentially match the 100 first frames, 'Frame 1, Frame 2, Frame 3, ...' and the 100 second frames, 'Frame A, Frame B, Frame C, ...' one by one to form a total of 100 frame pairs in the form of '(Frame 1, Frame A)', '(Frame 2, Frame B)', '(Frame 3, Frame C)', ...

[0053] Then, the image similarity calculation unit (122) can calculate the image similarity for each of the 100 frame pairs by calculating the image similarity between the two frames constituting each frame pair for each of the 100 frame pairs, '(Frame 1, Frame A)', '(Frame 2, Frame B)', '(Frame 3, Frame C)', ...

[0054] That is, the image similarity calculation unit (122) can calculate the image similarity between ‘Frame 1’ and ‘Frame A’, the image similarity between ‘Frame 2’ and ‘Frame B’, the image similarity between ‘Frame 3’ and ‘Frame C’, ... one by one in sequence, thereby calculating the image similarity for each of the 100 pairs of frames, ‘(Frame 1, Frame A)’, ‘(Frame 2, Frame B)’, ‘(Frame 3, Frame C)’, ...

[0055] At this time, assuming that the image similarity calculation unit (122) needs to calculate the image similarity between ‘Frame 1’ and ‘Frame A’, which are two frames constituting the first frame pair ‘(Frame 1, Frame A)’, which is one of the 100 frame pairs ‘(Frame 1, Frame A)’, ‘(Frame 2, Frame B)’, ‘(Frame 3, Frame C)’, ..., the details of how the image similarity calculation unit (122) calculates the image similarity between ‘Frame 1’ and ‘Frame A’ are explained in detail as follows.

[0056] For example, if t is set to '10' and 'Frame 1' and 'Frame A' are as shown in the figures illustrated in reference numerals 310 and 320 of FIG. 3, the image similarity calculation unit (122) can generate 100 first divided regions by equally dividing 'Frame 1' (310), which is one of the two frames constituting the first frame pair '(Frame 1, Frame A)', into 10 horizontal and 10 vertical partial regions as shown in the figure illustrated in reference numeral 311 of FIG. 3. Then, for each of the first divided regions, the image similarity calculation unit (122) can construct a three-dimensional vector having the R value, G value, and B value of the pixel located at the center point of each divided region as components. For example, if 'divided area 1' exists as one of the first divided areas, and the R, G, and B values ​​of the pixel located at the center point of 'divided area 1' are respectively 'R: 200, G: 190, B: 210', the image similarity calculation unit (122) can generate a feature vector corresponding to 'divided area 1' as '[200 190 210]'. In this way, the image similarity calculation unit (122) can generate a feature vector corresponding to each of the first divided areas.

[0057] Additionally, the image similarity calculation unit (122) can generate 100 second divided regions by equally dividing 'Frame A' (320), which is the other of the two frames constituting the first frame pair '(Frame 1, Frame A)', into 10 horizontal and 10 vertical partial regions as shown in the figure illustrated in reference numeral 321 of FIG. 3. Then, for each of the second divided regions, the image similarity calculation unit (122) can construct a three-dimensional vector having the R value, G value, and B value of the pixel located at the center point of each divided region as components. For example, if 'divided area 2' exists as one of the second divided areas, and the R, G, and B values ​​of the pixel located at the center point of 'divided area 2' are respectively 'R: 150, G: 180, B: 220', the image similarity calculation unit (122) can generate a feature vector corresponding to 'divided area 2' as '[150 180 220]'. In this way, the image similarity calculation unit (122) can generate a feature vector corresponding to each of the second divided areas.

[0058] Then, the image similarity calculation unit (122) can calculate the vector similarity of feature vectors between division areas at the same location among the first division areas and the second division areas. For example, if the division area corresponding to the area indicated by reference numeral 312 in FIG. 3 among the first division areas is called 'division area 1', and the division area corresponding to the area indicated by reference numeral 322 in FIG. 3 among the second division areas is called 'division area 2', the image similarity calculation unit (122) can calculate the vector similarity of feature vectors between 'division area 1' and 'division area 2', which are division areas at the same location. That is, the image similarity calculation unit (122) can calculate the vector similarity between '[200 190 210]', a feature vector corresponding to 'division area 1', and '[150 180 220]', a feature vector corresponding to 'division area 2'. In this way, the image similarity calculation unit (122) can calculate all 100 vector similarities by calculating the vector similarity of feature vectors between segmented regions at the same location for the remaining segmented regions as well.

[0059] Then, the image similarity calculation unit (122) can generate a reference matrix of size (10 x 10) having each of the 100 calculated vector similarities as components, and then calculate the Frobenius norm of the reference matrix as the image similarity between the two frames, 'Frame 1' (310) and 'Frame A' (320), which constitute the first frame pair '(Frame 1, Frame A)'.

[0060] Here, the Frobenius norm is a norm representing the size of a matrix, and can be expressed as shown in Equation 1 below.

[0062]

[0064] At this time, in the above mathematical formula 1 ga means Frobenius's game, and represents the element at the i-th row and j-th column of the matrix.

[0065] In this way, the image similarity calculation unit (122) calculates the image similarity for each of the 100 pairs of frames, such as '(Frame 1, Frame A)', '(Frame 2, Frame B)', '(Frame 3, Frame C)', ... as 'L1, L2, L3, ... , L 100 It can be calculated as follows.

[0066] Then the interval similarity calculation processing unit (123) calculates the overall average (V1) of the image similarity for each of the 100 frame pairs, i.e., ' ' can be calculated as the segment similarity between 'playback segment 1' and 'playback segment A', which are two playback segments constituting the first playback segment pair '(playback segment 1, playback segment A)'.

[0067] In this way, the segment similarity calculation unit (118) can calculate the segment similarity for each of the five pairs of playback segments, '(playback segment 1, playback segment A)', '(playback segment 2, playback segment B)', '(playback segment 3, playback segment C)', '(playback segment 4, playback segment D)', and '(playback segment 5, playback segment E)', as 'V1, V2, V3, V4, V5'.

[0068] Then the calculation processing unit (119) calculates the total sum (S) of the segment similarity for each of the five pairs of playback segments, i.e., ' ' can be calculated as the similarity between the first handwriting motion video and the first captured video.

[0069] When the similarity between the first sign language motion video and the first captured video is calculated through the similarity calculation unit (114) in this way, the judgment unit (115) checks whether the calculated similarity exceeds a preset threshold. If it is confirmed that it exceeds the threshold, it displays an accurate guidance message on the screen indicating that the user's sign language motion is accurate, and if it is confirmed that it does not exceed the threshold, it displays an inaccurate guidance message on the screen indicating that the user's sign language motion is inaccurate.

[0070] For example, if it is confirmed that the calculated similarity exceeds the threshold value, the judgment unit (115) reports that the user has performed the sign language action correctly and can display the accuracy guidance message on the screen. Then, the user can see the accuracy guidance message and recognize that the sign language action they performed was accurate.

[0071] However, unlike the example described above, if it is confirmed that the calculated similarity does not exceed the threshold value, the judgment unit (115) may report that the user has failed to perform the sign language action correctly and display the inaccuracy guidance message on the screen. Then, the user can see the inaccuracy guidance message and realize that they have not performed the sign language action correctly.

[0072] Through this, the user can quickly receive feedback on whether the sign language gestures they performed were accurate, thereby enabling them to learn sign language gestures more effectively.

[0073] According to one embodiment of the present invention, the electronic device (110) may further include a query section (124), a division section (125), and a section display section (126).

[0074] When the inaccurate guidance message is displayed on the screen through the judgment unit (115), the inquiry unit (124) displays a question message on the screen asking whether to proceed with learning sign language movements through transcription of the first practice text.

[0075] In this regard, if the judgment unit (115) determines that the user's sign language movement is inaccurate and the inaccurate guidance message is displayed on the screen, the inquiry unit (124) may display a question message on the screen asking whether to proceed with learning the sign language movement through transcription of the first practice text.

[0076] Then, the user can see the query message displayed on the screen and decide whether to proceed with sign language movement learning. If the user wishes to proceed with sign language movement learning, the user can apply a learning start command to the electronic device (110) of the present invention instructing to start sign language movement learning in response to the query message.

[0077] When a learning start command is issued from the user in response to the query message to start learning sign language movements, the division unit (125) divides the entire playback section constituting the first sign language movement video into multiple playback sections corresponding to the total number of characters constituting the first practice text.

[0078] For example, if the total number of characters constituting the first practice text is '10 characters', the dividing unit (125) can divide the entire playback section constituting the first sign language action video into '10' parts, such as 'playback section A, playback section B, playback section C, ...', and divide it into a total of 10 playback sections.

[0079] The section display unit (126) displays the plurality of playback sections sequentially one by one on the second area whenever a character constituting the first practice text is transcribed one by one by the user through a typing input device connected to the electronic device (110).

[0080] For example, let the first practice text be "I like you" and the plurality of playback sections be "playback section A, playback section B, playback section C, ...". In addition, in this embodiment, it is assumed that spaces are also treated as characters.

[0081] At this time, if the character 'Na' is transcribed by the user through a typing input device connected to the electronic device (110), the section display unit (126) can display 'Playback section Ga' on the second area. Next, if the character 'Neun' is transcribed by the user through the typing input device, the section display unit (126) can display 'Playback section Na' on the second area. Next, if the character ' ' (blank) is transcribed by the user through the typing input device, the section display unit (126) can display 'Playback section Da' on the second area.

[0082] In this way, the section display unit (126) can sequentially play and display the plurality of playback sections, ‘playback section A, playback section B, playback section C, ...’, one by one on the second area whenever the characters constituting the first practice text, ‘I like you,’ are transcribed one by one.

[0083] Through this, the user will be able to study the first sign language movement video in detail by dividing it into sections, thereby enabling more accurate sign language movement learning.

[0084] FIG. 4 is a flowchart illustrating a method of operation of an electronic device according to an embodiment of the present invention.

[0085] In step (S410), a sign language database is maintained in which a plurality of pre-specified practice texts and sign language movement videos corresponding to each of the plurality of practice texts are stored.

[0086] In step (S420), when a command to start practicing a sign language action corresponding to a first practice text, which is one of the plurality of practice texts, is issued by the user, the first practice text and a first sign language action video corresponding to the first practice text are extracted from the sign language database, and then the first practice text is displayed in a pre-set first area on the screen and the first sign language action video is displayed in a pre-set second area on the screen.

[0087] In step (S430), when a shooting start command is issued requesting the user to film their sign language movements, the camera mounted on the electronic device is activated to start filming the user's sign language movements, and then when a shooting end command is issued requesting the user to stop filming the sign language movements, the camera is deactivated and filming is stopped, thereby obtaining a first filmed video of the user's sign language movements.

[0088] In step (S440), the similarity between the first speech motion video and the first captured video is calculated.

[0089] In step (S450), it is checked whether the calculated similarity exceeds a preset threshold. If it is confirmed that it exceeds the threshold, an accuracy guidance message indicating that the user's sign language gesture is accurate is displayed on the screen. If it is confirmed that it does not exceed the threshold, an inaccuracy guidance message indicating that the user's sign language gesture is inaccurate is displayed on the screen.

[0090] At this time, according to one embodiment of the present invention, step (S440) may include: generating k first playback segments by dividing the entire playback segment constituting the first sign language motion video into k playback segments (k is a natural number greater than or equal to 2), and generating k second playback segments by dividing the entire playback segment constituting the first captured video into k playback segments; sequentially matching the k first playback segments and the k second playback segments one by one to form k pairs of playback segments; calculating the segment similarity for each of the k pairs of playback segments by calculating the segment similarity between the two playback segments constituting each pair of playback segments for each of the k pairs of playback segments; and calculating the total sum of the segment similarities for each of the k pairs of playback segments as the similarity between the first sign language motion video and the first captured video.

[0091] According to one embodiment of the present invention, the step of calculating the segment similarity comprises performing a process for calculating the segment similarity sequentially for each of the k pairs of playback segments, wherein when the sequence is to calculate the segment similarity between two playback segments constituting a first pair of playback segments, which is one of the k pairs of playback segments, for one of the two playback segments constituting the first pair of playback segments, n first frames are generated by sampling n frames (where n is a natural number greater than or equal to 5) among a plurality of frames constituting the said playback segment at equal intervals, and for the other of the two playback segments constituting the first pair of playback segments, n second frames are generated by sampling n frames among a plurality of frames constituting the said playback segment at equal intervals, the n first frames and the n second frames are sequentially matched one by one to form n pairs of frames, and for each of the n pairs of frames, the image similarity for each of the n pairs of frames is calculated by calculating the image similarity between the two frames constituting each pair of frames. The method may include a step of calculating the overall average of the image similarity for each of the n pairs of frames as the segment similarity between the two playback segments constituting the first pair of playback segments.

[0092] At this time, according to one embodiment of the present invention, the step of calculating the image similarity performs a process for calculating the image similarity one by one for each of the n frame pairs, wherein when it is the order to calculate the image similarity between two frames constituting a first frame pair, which is any one of the n frame pairs, for one of the two frames constituting the first frame pair, the frame is equally divided into (txt) first divided regions by dividing the frame horizontally (t is a natural number greater than or equal to 10) and vertically t, and for each of the first divided regions, a feature vector corresponding to each of the first divided regions is generated by constructing a 3-dimensional vector having the R, G, and B values ​​of the pixel located at the center point of each divided region as components, and for the other frame constituting the first frame pair, the frame is equally divided into (txt) second divided regions by dividing the frame horizontally and t vertically, and for each of the second divided regions For this, by constructing a 3-dimensional vector having the R, G, and B values ​​of the pixel located at the center point of each segmented region as components, a feature vector corresponding to each of the second segmented regions is generated, and then a total of (txt) vector similarities are calculated by computing the vector similarity of the feature vectors between segmented regions at the same location in the first segmented regions and the second segmented regions, and then a reference matrix of size (txt) having each of the calculated vector similarities as components is generated, and then the Frobenius norm of the reference matrix can be calculated as the image similarity between the two frames constituting the first frame pair.

[0093] In addition, according to one embodiment of the present invention, the method of operating the electronic device may further include the step of displaying on the screen, the step of displaying a question message on the screen asking whether to proceed with sign language movement learning through transcription of the first practice text when the inaccurate guidance message is displayed on the screen, the step of dividing the entire playback section constituting the first sign language movement video into a plurality of playback sections according to the total number of characters constituting the first practice text when a learning start command instructing to start sign language movement learning in response to the question message is applied from the user, and the step of sequentially playing and displaying the plurality of playback sections one by one on the second area whenever a character constituting the first practice text is transcribed one by one by the user when the first practice text begins to be transcribed through a typing input device connected to the electronic device.

[0094] Hereinafter, a method of operation of an electronic device according to an embodiment of the present invention has been described with reference to FIG. 4. Since the method of operation of an electronic device according to an embodiment of the present invention may correspond to the configuration of the operation of the electronic device (110) described using FIG. 1 to 3, a more detailed description thereof will be omitted.

[0095] A method of operation of an electronic device according to one embodiment of the present invention can be implemented as a computer program stored in a storage medium for execution through combination with a computer.

[0096] In addition, a method of operation of an electronic device according to an embodiment of the present invention may be implemented in the form of computer program instructions for execution through combination with a computer and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the present invention, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.

[0097] As described above, the present invention has been explained by specific details such as specific components, limited embodiments, and drawings; however, this is provided merely to aid in a more comprehensive understanding of the invention, and the invention is not limited to the above embodiments. A person skilled in the art can make various modifications and variations from this description.

[0098] Therefore, the scope of the present invention is not limited to the described embodiments, and all things equivalent to or having equivalent variations to the claims set forth below, as well as the claims set forth below, shall be considered to fall within the scope of the concept of the present invention. Explanation of the symbols

[0099] 110: Electronic device 111: Sign Language Database 112: Split display 113: Acquisition Department 114: Similarity Calculation Unit 115: Judgment Division 116: Generation section 117: Playback section pair component 118: Interval Similarity Calculation Unit 119: Output processing unit 120: Sampling section 121: Frame pair component 122: Image Similarity Calculator 123: Interval Similarity Calculation Processing Unit 124: Inquiry Section 125: Divided section 126: Section indicator

Claims

Claim 1 In an electronic device, a sign language database storing a plurality of pre-specified practice texts and sign language movement images corresponding to each of the plurality of practice texts; a split display unit that, when a command to start practicing a sign language movement corresponding to a first practice text, which is one of the plurality of practice texts, is applied from a user, extracts the first practice text and a first sign language movement image corresponding to the first practice text from the sign language database, displays the first practice text in a pre-set first area on the screen, and displays the first sign language movement image in a pre-set second area on the screen; an acquisition unit that, when a command to start shooting requesting the user to shoot their sign language movement is applied, activates a camera mounted on the electronic device to start shooting the user's sign language movement, and when a command to end shooting requesting the user to end shooting is applied, deactivates the camera and ends shooting, thereby acquiring a first captured image of the user's sign language movement; and a similarity calculation unit that calculates the similarity between the first sign language movement image and the first captured image. An electronic device comprising a judgment unit that checks whether the similarity calculated above exceeds a preset threshold, and if the check result confirms that it exceeds the threshold, displays an accuracy guidance message on the screen indicating that the user's sign language gesture is accurate, and if the check result confirms that it does not exceed the threshold, displays an inaccuracy guidance message on the screen indicating that the user's sign language gesture is inaccurate. Claim 2 An electronic device according to claim 1, wherein the similarity calculation unit generates k first playback segments by equally dividing the entire playback segment constituting the first sign language motion video into k playback segments (k is a natural number greater than or equal to 2) and generates k second playback segments by equally dividing the entire playback segment constituting the first captured video into k playback segments; a playback segment pair formation unit that sequentially matches the k first playback segments and the k second playback segments one by one to form k playback segment pairs; a segment similarity calculation unit that calculates the segment similarity for each of the k playback segment pairs by calculating the segment similarity between the two playback segments constituting each playback segment pair for each of the k playback segment pairs; and a calculation processing unit that calculates the total sum of the segment similarities for each of the k playback segment pairs as the similarity between the first sign language motion video and the first captured video. Claim 3 In paragraph 2, the segment similarity calculation unit performs a process for calculating segment similarity one by one for each of the k pairs of playback segments sequentially, wherein when the order of calculating segment similarity between two playback segments constituting a first pair of playback segments, which is any one of the k pairs of playback segments, is reached, for one of the two playback segments constituting the first pair of playback segments, a sampling unit generates n first frames by sampling n frames (where n is a natural number greater than or equal to 5) among a plurality of frames constituting the said playback segment at equal intervals, and for the other of the two playback segments constituting the first pair of playback segments, a sampling unit generates n second frames by sampling n frames among a plurality of frames constituting the said playback segment at equal intervals; a frame pair formation unit that sequentially matches the n first frames and the n second frames one by one to form n pairs of frames; and an image that calculates image similarity for each of the n pairs of frames by calculating image similarity between the two frames constituting each pair of frames for each of the n pairs of frames. An electronic device comprising: a similarity calculation unit; and a segment similarity calculation processing unit that calculates the overall average of image similarity for each of the n pairs of frames as the segment similarity between two playback segments constituting the first pair of playback segments. Claim 4 In paragraph 3, the image similarity calculation unit performs a process for calculating image similarity sequentially for each of the n frame pairs, wherein when it is time to calculate image similarity between two frames constituting a first frame pair, which is any one of the n frame pairs, for one of the two frames constituting the first frame pair, the frame is equally divided into (txt) first divided regions by dividing the frame horizontally (t is a natural number greater than or equal to 10) and vertically t, and for each of the first divided regions, a feature vector corresponding to each of the first divided regions is generated by constructing a 3-dimensional vector having the R, G, and B values ​​of the pixel located at the center point of each divided region as components, and for the other frame constituting the first frame pair, the frame is equally divided into (txt) second divided regions by dividing the frame horizontally and t vertically, and for each of the second divided regions, each divided region An electronic device characterized by: generating a feature vector corresponding to each of the second segmented regions by constructing a three-dimensional vector having the R, G, and B values ​​of a pixel located at a center point as components; then calculating a total of (txt) vector similarities by calculating the vector similarity of the feature vectors between segmented regions at the same location in the first segmented regions and the second segmented regions; generating a reference matrix of size (txt) having each of the calculated vector similarities as components; and then calculating the Frobenius Norm of the reference matrix as the image similarity between the two frames constituting the first frame pair. Claim 5 An electronic device according to claim 1, further comprising: a query unit that displays a query message on a screen asking whether to proceed with sign language movement learning through transcription of the first practice text when the inaccurate guidance message is displayed on the screen through the judgment unit; a division unit that divides the entire playback section constituting the first sign language movement video into a plurality of playback sections corresponding to the total number of characters constituting the first practice text when a learning start command instructing to start sign language movement learning in response to the query message is applied from the user; and a section display unit that sequentially plays and displays the plurality of playback sections one by one on the second area whenever a character constituting the first practice text is transcribed one by one by the user through a typing input device connected to the electronic device. Claim 6 A method of operating an electronic device comprises: maintaining a sign language database in which a plurality of pre-specified practice texts and sign language movement images corresponding to each of the plurality of practice texts are stored; when a command to start practicing a sign language movement corresponding to a first practice text, which is one of the plurality of practice texts, is applied from a user, extracting the first practice text and a first sign language movement image corresponding to the first practice text from the sign language database, then displaying the first practice text in a pre-set first area on the screen and displaying the first sign language movement image in a pre-set second area on the screen; when a command to start shooting requesting to shoot one's own sign language movement is applied from the user, activating a camera mounted on the electronic device to start shooting the user's sign language movement, and when a command to end shooting requesting to end the shooting is applied from the user, deactivating the camera and ending the shooting, thereby obtaining a first captured image of the user's sign language movement; and calculating a similarity between the first sign language movement image and the first captured image. A method of operating an electronic device comprising the step of checking whether the similarity calculated above exceeds a preset threshold, and if it is confirmed that it exceeds the threshold, displaying an accuracy guidance message on the screen indicating that the user's sign language gesture is accurate, and if it is confirmed that it does not exceed the threshold, displaying an inaccuracy guidance message on the screen indicating that the user's sign language gesture is inaccurate. Claim 7 In claim 6, the calculating step comprises: generating k first playback segments by equally dividing the entire playback segment constituting the first sign language motion video into k playback segments (k is a natural number greater than or equal to 2), and generating k second playback segments by equally dividing the entire playback segment constituting the first captured video into k playback segments; sequentially matching the k first playback segments and the k second playback segments one by one to form k pairs of playback segments; calculating the segment similarity for each of the k pairs of playback segments by calculating the segment similarity between the two playback segments constituting each pair of playback segments for each of the k pairs of playback segments; and calculating the total sum of the segment similarities for each of the k pairs of playback segments as the similarity between the first sign language motion video and the first captured video. Claim 8 In claim 7, the step of calculating the segment similarity comprises performing a process for calculating the segment similarity sequentially for each of the k pairs of playback segments, wherein when the sequence of calculating the segment similarity between two playback segments constituting a first pair of playback segments, which is one of the k pairs of playback segments, is reached, for one of the two playback segments constituting the first pair of playback segments, n first frames are generated by sampling n frames (where n is a natural number greater than or equal to 5) among the plurality of frames constituting the said playback segment at equal intervals, and for the other of the two playback segments constituting the first pair of playback segments, n second frames are generated by sampling n frames among the plurality of frames constituting the said playback segment at equal intervals; the step of sequentially matching the n first frames and the n second frames one by one to form n pairs of frames; and the step of calculating the image similarity for each of the n pairs of frames by calculating the image similarity between the two frames constituting each pair of frames for each of the n pairs of frames. A method of operation of an electronic device comprising the step of calculating the overall average of image similarity for each of the n pairs of frames as the segment similarity between two playback segments constituting the first pair of playback segments. Claim 9 In claim 8, the step of calculating the image similarity performs a process for calculating image similarity sequentially for each of the n frame pairs, wherein when it is the order to calculate the image similarity between two frames constituting a first frame pair, which is any one of the n frame pairs, for one of the two frames constituting the first frame pair, the frame is equally divided into (txt) first divided regions by dividing the frame horizontally (t is a natural number greater than or equal to 10) and vertically t, and for each of the first divided regions, a feature vector corresponding to each of the first divided regions is generated by constructing a 3-dimensional vector having the R, G, and B values ​​of the pixel located at the center point of each divided region as components, and for the other frame constituting the first frame pair, the frame is equally divided into (txt) second divided regions by dividing the frame horizontally and t vertically, and for each of the second divided regions, each divided region A method of operation of an electronic device characterized by: generating a feature vector corresponding to each of the second divided regions by constructing a three-dimensional vector having the R, G, and B values ​​of a pixel located at a center point as components; then calculating a total of (txt) vector similarities by calculating the vector similarity of the feature vectors between divided regions at the same location in the first divided regions and the second divided regions; then generating a reference matrix of size (txt) having each of the calculated vector similarities as components; and then calculating the Frobenius Norm of the reference matrix as the image similarity between the two frames constituting the first frame pair. Claim 10 A method of operation of an electronic device according to claim 6, further comprising: a step of displaying a question message on a screen inquiring whether to proceed with sign language movement learning through transcription of the first practice text when the inaccurate guidance message is displayed on the screen through the step of displaying on the screen; a step of dividing the entire playback section constituting the first sign language movement video into a plurality of playback sections corresponding to the total number of characters constituting the first practice text when a learning start command instructing to start sign language movement learning in response to the question message is issued from the user; and a step of sequentially playing and displaying the plurality of playback sections one by one on the second area whenever a character constituting the first practice text is transcribed one by one by the user through a typing input device connected to the electronic device. Claim 11 A computer-readable recording medium having a computer program for executing the method of any one of paragraphs 6 through 10 in combination with a computer. Claim 12 A computer program stored on a storage medium for executing the method of any one of paragraphs 6 through 10 through combination with a computer.