A sign language recognition system adapted for sign language rephrasing processing

By designing a sign language recognition system adapted to re-sign language processing and adopting word-level and sentence-level continuous translation modes, the system solves the problem of insufficient processing of re-sign language in daily use by hearing-impaired people, realizes two-way communication between hearing-impaired and hearing people, and improves communication efficiency and user experience.

CN116704548BActive Publication Date: 2026-04-14ZHEJIANG UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-17
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing sign language recognition devices are insufficient in handling complex sign language scenarios in daily use by hearing-impaired individuals, resulting in a reduced user experience and the inability to achieve two-way communication between hearing-impaired and hearing individuals.

Method used

Design a sign language recognition system adapted to sign language re-typing processing, including a host, screen, acquisition components and control unit. It adopts word-level and sentence-level continuous translation modes, recognizes and processes re-typed sign language through re-typing mark actions, provides a re-typing confirmation process, and improves the user experience and communication efficiency of sign language recognition devices.

Benefits of technology

It improves the user experience of sign language recognition devices for the hearing impaired, reduces the frustration of using them, enables two-way communication between the hearing impaired and hearing people, and improves communication efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116704548B_ABST
    Figure CN116704548B_ABST
Patent Text Reader

Abstract

The application relates to a sign language recognition system suitable for sign language retyping processing, which comprises a host computer, at least two screens matched with the host computer, a collecting assembly and an adjusting assembly matched with one of the screens, and a control unit matched with the screens, the collecting assembly and the adjusting assembly, and carries out sign language retyping processing in a word-level sign language continuous translation mode and a sentence-level sign language continuous translation mode. In the word-level sign language continuous translation mode, a whole sentence of content needing to be re-expressed is replaced, and in the sentence-level sign language continuous translation mode, a sentence needing to be re-expressed is divided, a part is reserved, and another part is replaced. The application improves the utilization rate of key point information, conforms to the daily use habits of users, especially hearing-impaired people, and obtains better sign language recognition equipment use experience; reduces the use of abruptness, improves fluency, improves the sign language retyping experience of hearing-impaired people on the sign language recognition equipment, and to a certain extent, solves the inconvenience of hearing-impaired people in communication with normal people.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of image or video recognition or understanding, and in particular to a sign language recognition system based on video data processing and adapted sign language re-processing. Background Technology

[0002] Hearing impairment (dysaudia) refers to organic or functional abnormalities in the various levels of the auditory system's nerve centers responsible for sound transmission, perception, and comprehensive sound analysis, resulting in varying degrees of hearing loss. Its causes mainly include genetic factors, environmental factors, drug and chemical-induced deafness, noise-induced hearing loss, and trauma. Hearing-impaired individuals primarily communicate through writing and sign language. However, due to the low prevalence of sign language, the inability of hearing-impaired individuals to communicate effectively with untrained individuals using sign language remains a pressing issue.

[0003] With increased societal care for the hearing-impaired, the development of image processing technology, and the widespread application of smart terminal devices, more and more assistive devices are available for the hearing-impaired to choose from, improving their communication efficiency.

[0004] However, current sign language recognition devices still have certain limitations. This is mainly reflected in the fact that current devices focus on selecting more suitable sign language data collection methods to improve the accuracy and speed of sign language recognition, while relatively insufficient consideration is given to the practical daily use by hearing-impaired users. Hearing-impaired users may be dissatisfied with or make mistakes in their sign language expression. Most current sign language translation devices process these situations only once; when a hearing-impaired user wants to rephrase a previous statement, all the previous content must be typed out again in sign language. This significantly reduces the user experience and communication efficiency.

[0005] Regarding sign language translation, patent CN115171212A discloses a sign language recognition method, device, equipment, and storage medium. It acquires joint and skeletal pose data of the target sign language movement, further expanding this into joint motion data streams, skeletal pose data streams, and skeletal motion data streams. It also constructs the topological relationships corresponding to the human body's pose structure, inputting all of these data into a sign language graph convolutional neural network for sign language recognition to obtain the predicted sign language vocabulary corresponding to the target sign language movement. This method can improve the recognition accuracy of sign language videos, but its limitation lies in not considering the daily use by hearing-impaired individuals and ignoring scenarios where hearing-impaired users need to repeat sign language.

[0006] Regarding the issue of communication with hearing-impaired individuals, patent CN114721515A discloses an AI-powered gesture recognition bracelet and its gesture recognition method. After collecting the gestures of hearing-impaired individuals through the smart bracelet, the AI ​​chip built into the smart bracelet processes the data, and then the data is played out through a voice player, thus fulfilling the need for one-way communication between hearing-impaired and hearing individuals. However, this method cannot achieve two-way communication between hearing-impaired and hearing individuals, and when hearing-impaired users need to re-express themselves, it is easy for hearing individuals to have difficulty understanding.

[0007] In summary, people with hearing impairments need a sign language approach that facilitates the reinterpretation of information and allows hearing people to understand it. Summary of the Invention

[0008] This invention solves the problems existing in the prior art and provides a sign language recognition system adapted to sign language re-typing processing.

[0009] The technical solution adopted in this invention is a sign language recognition system adapted to sign language re-printing processing. The system includes a host computer and at least two screens in conjunction with the host computer. One of the screens is equipped with a data acquisition component and an adjustment component.

[0010] The screen, acquisition component, and adjustment component are equipped with a control unit. The control unit performs sign language re-translation in two modes: word-level sign language continuous translation mode and sentence-level sign language continuous translation mode. In the word-level sign language continuous translation mode, the entire sentence that needs to be re-expressed is replaced. In the sentence-level sign language continuous translation mode, the sentence that needs to be re-expressed is segmented, retaining one part and replacing the other part that needs to be re-expressed.

[0011] Preferably, the control unit collects and saves the re-signature gesture and other signature gestures through the acquisition component. During the sign language translation process, when the acquisition component captures the re-signature gesture, it enters the re-signature confirmation process. Other signature gestures include cancellation, approval, and selection. Signature gestures include, but are not limited to, facial expressions, body movements, and hand gestures. The capture of facial expressions and body movements is based on relevant deep learning detection models, and the capture of hand gestures is based on key point detection models of human pose estimation networks, thereby distinguishing re-signature gestures from normal sign language gestures.

[0012] Preferably, the re-call confirmation process includes the following steps:

[0013] Step 1.1: In the normal sign language process, when the acquisition component captures the re-signature action, proceed to the next step; otherwise, repeat step 1.1.

[0014] Step 1.2: The control unit pre-selects sentences that have been translated. Pre-selection means that the current sentence is considered to need to be re-typed by default, which helps the sign language interpreter select the sentences that need to be re-typed.

[0015] Step 1.3: The screen displays a countdown. If the user cancels within the countdown, the process returns to Step 1.2. Otherwise, the process restarts after the countdown ends or the user acknowledges within the countdown, until the process returns to the normal sign language flow. The countdown on the screen gives the user time to reconsider whether to restart, preventing accidental re-signing. The countdown time can be set by the hearing-impaired user according to their own situation.

[0016] Preferably, in step 1.2, assisting the sign language interpreter in selecting the phrases that need to be retyped includes the following steps:

[0017] Step 1.2.1: Pre-select the current statement by changing the font color;

[0018] Step 1.2.2: The control unit prompts whether to retype the current statement. If the statement to be retyped is not the current statement or the retyped statement needs to be cancelled, the sign language interpreter cancels the statement and proceeds to the next step. Otherwise, the sign language interpreter acknowledges the statement and proceeds to step 1.3.

[0019] Step 1.2.3: The control unit prompts whether to cancel the re-signing. If yes, the signer makes the cancel action again, exits the re-signing process, and returns to the normal sign language process. Otherwise, the signer selects the preceding statement by selecting an action.

[0020] Step 1.2.4: Repeat step 1.2.2.

[0021] Preferably, in the re-typing process of the word-level sign language continuous translation mode, the user re-typs the sign language, while the text flashes to indicate that it is being processed. At the same time, another screen indicates that the content of the sentence is being modified. After the new sentence is translated, the original sentence is overwritten, the cursor jumps to the end, the sign language re-typing process is exited, and the normal sign language process continues.

[0022] Preferably, in the sentence-level sign language continuous translation mode, during the normal sign language translation process, the control unit obtains the key point coordinate data of each frame of sign language action, which is recorded as the raw data.

[0023] Preferably, the retyping process of the sentence-level sign language continuous translation mode includes the following steps:

[0024] Step 2.1: Calculate the parameter matrix of the original data frame in reverse order, denoted as... This represents the relative distances and relative angles between keypoints in the s-th frame of the original data frame. In the parameter matrix, the upper triangular matrix [j,k] represents the relative distance between keypoints j and k, and the lower triangular matrix [j,k] represents the angle between keypoints j and k, satisfying the following conditions:

[0025]

[0026] Step 2.2: The user re-enters the sign language, obtaining the key point data of each frame of the sign language, which is recorded as a secondary data frame. The parameter matrix of each secondary data frame is then obtained.

[0027] Step 2.3: Determine if each frame is invalid. If it is, clear t and return to step 2.2; otherwise, ... With each Perform matrix subtraction to obtain the staggered matrix PDM. st , Here, t refers to the t-th frame;

[0028] Step 2.4: For each parametric matrix PDM st The contrast coefficient CC is obtained by performing a weighted summation. st ;

[0029] Step 2.5: Define a sensitive parameter n, n≥10; if t is less than the sensitive parameter n, then t=t+1, return to step 2.2, otherwise proceed to the next step; here t refers to the t-th frame;

[0030] Step 2.6: Based on the contrast coefficient CC st Calculate the n consecutive frames in the original data frame that have the smallest sum of contrast coefficients.

[0031]

[0032] Where m represents the total number of original data frames in the current sentence, and the first frame in n frames is denoted as frame k; retain the key point data of the original frames before frame k, discard the key point data of the original frames after frame k, and replace them with secondary data frames;

[0033] Step 2.7: Merge the original data with the secondary data of the secondary data frame, send it to the sign language model for translation, replace the original statement after the new statement is translated, jump the cursor to the end, exit the sign language re-typing process, and continue the normal sign language process.

[0034] Preferably, in step 2.3, the determination of invalid frames includes the following steps:

[0035] Step 2.3.1: Collect the raw data of sign language action frames, send them into the human pose estimation network, extract the human key points involved in any frame read in Step 2.3.1, and obtain the key point matrix, including the nodes of the head, torso, and hands;

[0036] Step 2.3.2: Determine whether the necessary key points are included (key points of the head, torso, and hands are all present), whether the relative positions of the key points (relative positions of the head and torso) are reasonable, and whether the absolute positions of the key points are reasonable. If all of these are true, the data frame is valid; otherwise, it is invalid.

[0037] Preferably, step 2.4 includes the following steps:

[0038] Step 2.4.1: For parameter matrix PM1,

[0039]

[0040] Where m represents the total number of original data frames for the current sentence, and p i,j This represents the parameter array consisting of all parameters in the i-th row and j-th column of PM1, where i ≠ j;

[0041] Calculate the standard deviation of each parameter.

[0042] The standard deviation matrix SDM is obtained.

[0043] Step 2.4.2: Unify the upper and lower triangular matrices in the SDM to obtain the weighted matrix WM. in,

[0044] Step 2.4.3: Based on each weighting coefficient in the weighting matrix, perform a PDM analysis on each non-parallel matrix. st The weighted summation is performed to obtain the corresponding contrast coefficient CC. st ,

[0045] Where k represents the number of key points, i represents the i-th row, j represents the j-th column, and i ≠ j.

[0046] Preferably, distinguishing between the re-signing action and other sign actions includes the following steps:

[0047] Step 3.1: Input each frame of sign language motion image data captured by the acquisition component into the human pose estimation network to obtain the coordinate data of the sign language key points in each frame;

[0048] Step 3.2: Sequentially calculate the frame difference (Frame) of keypoint coordinate data between every two frames. i-1 –Frame i If i ≥ 1, save as Diff i ;

[0049] Step 3.3: When there are at least two frame differences, perform a t-test on two consecutive frame differences; the significance level of the t-test is set to 0.01. The preset gesture for hearing-impaired users is a stationary gesture. Therefore, when the key points of two consecutive frames of the same gesture are subtracted, it can be considered that the difference Diff meets the requirements of a normal distribution.

[0050] Step 3.4: If Diff i With Diff i-1 If the difference is greater than the preset value, then clear i and return to step 3.1; otherwise, proceed to the next step.

[0051] Step 3.5: Calculate the frame difference between the sign language key points in each frame and the preset gesture key points. Calculate the average offset of the head and shoulder key points in two frames. Correct the remaining key points using the average offset. After correction, calculate the offset between each remaining key point and the preset key point.

[0052] Step 3.6: Determine the offset of each key point (! x ,! y If the error rate is within the allowable tolerance range, then clear i and return to step 3.1; otherwise, i = i + 1. f is typically set to 5%, and width and height are matched to the resolution of the sign language images captured by the camera.

[0053] Step 3.7: Determine if i is equal to the sensitive parameter n. If not, return to step 3.1. Otherwise, assume the user has made a re-signing action, pause the current sign language reasoning, and enter the re-signing confirmation process.

[0054] This invention relates to a sign language recognition system adapted for sign language re-typing processing. The system includes a host computer and at least two screens. One of the screens is equipped with a data acquisition component and an adjustment component. A control unit is equipped with the screens, the data acquisition component, and the adjustment component. The control unit performs sign language re-typing processing in two modes: word-level continuous sign language translation and sentence-level continuous sign language translation. In the word-level continuous sign language translation mode, the entire sentence that needs to be re-expressed is replaced. In the sentence-level continuous sign language translation mode, the sentence that needs to be re-expressed is segmented, retaining one part and replacing the other part that needs to be re-expressed.

[0055] The beneficial effects of this invention are as follows:

[0056] (1) Based on key point information, the key point data that originally needed to be discarded in large quantities was utilized and analyzed, thereby improving the utilization rate of key point information.

[0057] (2) By using special processing strategies for key points of sign language action data, this method is made to conform to the daily usage habits of users, especially the hearing-impaired, and the hearing-impaired have a better experience in using sign language recognition devices.

[0058] (3) Different sign language redo symbols can be selected according to the user's communication habits, which reduces the user's sense of stagnation, improves the fluency, and improves the sign language redo experience of hearing-impaired people on sign language recognition devices, thus solving the inconvenience of hearing-impaired people when communicating with hearing people to a certain extent. Attached Figure Description

[0059] Figure 1 This is a schematic block diagram of the sign language recognition system in this invention;

[0060] Figure 2 This is a flowchart of the re-sign language process for the word-level sign language continuous translation mode of the present invention;

[0061] Figure 3 This is a flowchart of the re-sign language process for the sentence-level sign language continuous translation mode of the present invention;

[0062] Figure 4 This is a flowchart illustrating the distinction between the re-marking action and other marking actions in this invention. Detailed Implementation

[0063] The present invention will be further described in detail below with reference to embodiments, but the scope of protection of the present invention is not limited thereto.

[0064] like Figures 1-4 As shown, the present invention relates to a sign language recognition system adapted to sign language re-printing processing. The system includes a host computer and at least two screens in conjunction with the host computer; one of the screens is equipped with a data acquisition component and an adjustment component.

[0065] The screen, acquisition component, and adjustment component are equipped with a control unit. The control unit performs sign language re-translation in two modes: word-level sign language continuous translation mode and sentence-level sign language continuous translation mode. In the word-level sign language continuous translation mode, the entire sentence that needs to be re-expressed is replaced. In the sentence-level sign language continuous translation mode, the sentence that needs to be re-expressed is segmented, retaining one part and replacing the other part that needs to be re-expressed.

[0066] In this invention, the system is presented in the form of a sign language recognition device. The device includes two screens, typically touchscreens, one facing the hearing-impaired user and the other facing the hearing user, used to display interactive information and input text data, effectively enabling two-way communication between the hearing-impaired and hearing users. The acquisition component is typically a wide-angle lens (camera) used to capture the sign language gestures of the hearing-impaired user. The adjustment component includes a button function area located on the outside of the sign language recognition device, used to adjust settings such as brightness and volume. Figure 1 The diagram illustrates one implementation, which integrates an audio module, a sign language module, and a video module within the system module. The video module captures sign language gestures via a camera, and after collaborative processing by the sign language module and the posture estimation module, the output is displayed on two screens (LCD screen 1 and LCD screen 2) to both the hearing-impaired and hearing-sounding individuals via an interactive interface module. The audio module provides audio input via speakers, microphones, etc., and the button function area enables the application of adjustment components.

[0067] In this invention, the hearing-impaired user presets a sensitive parameter n and a sign language re-match countdown sec in the sign language recognition device. The sensitive parameter indicates that n frames of secondary data are adopted for matching calculation. The larger the value of n, the higher the accuracy but the slower the matching speed. The smaller the value of n, the lower the accuracy but the faster the matching speed (n≥10). The countdown sec is used to give the hearing-impaired user to re-determine whether a rematch is needed and to give the hearing-impaired user time to think about the sign language action.

[0068] In the implementation of this invention, the re-typing process includes a re-typing confirmation process and a re-typing process. The re-typing process includes a word-level sign language continuous translation mode and a sentence-level sign language continuous translation mode.

[0069] The control unit collects and saves the re-printing sign action and other sign actions through the acquisition component. During the sign language translation process, when the acquisition component captures the re-printing sign action, it enters the re-printing confirmation process; other sign actions include cancellation, approval, and selection actions.

[0070] In the implementation of this invention, the user needs to preset the re-marking action and other marking actions.

[0071] The re-call confirmation process includes the following steps:

[0072] Step 1.1: In the normal sign language process, when the acquisition component captures the re-signature action, proceed to the next step; otherwise, repeat step 1.1.

[0073] Step 1.2: The control unit pre-selects the translated sentences to assist the sign language interpreter in selecting the sentences that need to be re-typed;

[0074] Step 1.2 is essentially a reconfirmation of whether the user truly needs to retype the statement; in Step 1.2, assisting the sign language interpreter in selecting the statement that needs to be retyped includes the following steps:

[0075] Step 1.2.1: Pre-select the current statement by changing the font color;

[0076] Step 1.2.2: The control unit prompts whether to retype the current statement. If the statement to be retyped is not the current statement or the retyped statement needs to be cancelled, the sign language interpreter cancels the statement and proceeds to the next step. Otherwise, the sign language interpreter acknowledges the statement and proceeds to step 1.3.

[0077] Step 1.2.3: The control unit prompts whether to cancel the re-signing. If yes, the signer makes the cancel action again, exits the re-signing process, and returns to the normal sign language process. Otherwise, the signer selects the preceding statement by selecting an action.

[0078] Step 1.2.4: Repeat step 1.2.2.

[0079] Step 1.3: The screen displays a countdown. If the user cancels within the countdown, return to Step 1.2. Otherwise, after the countdown ends or the user acknowledges within the countdown, start re-entering the process until it is completed and return to the normal sign language process.

[0080] This invention involves distinguishing between re-marking actions and other marking actions, including the following steps:

[0081] Step 3.1: Input each frame of sign language motion image data captured by the acquisition component into the human pose estimation network to obtain the coordinate data of the sign language key points in each frame;

[0082] Step 3.2: Sequentially calculate the frame difference (Frame) of keypoint coordinate data between every two frames. i-1 –Frame i If i ≥ 1, save as Diff i ;

[0083] Step 3.3: When there are at least two frame differences, perform a t-test on two consecutive frame differences;

[0084] Step 3.4: If Diff i With Diff i-1 If the difference between Diff and i is greater than a preset value (significant difference exists), then i is reset to zero and the process returns to step 3.1; otherwise, proceed to the next step. Generally, the preset value here can be set to 0.05, that is, if Diff... i With Diff i-1 The t-test result was less than 0.05, indicating a significant difference between the two.

[0085] Step 3.5: Calculate the frame difference between the sign language key points in each frame and the preset gesture key points. Calculate the average offset of the head and shoulder key points in two frames. Correct the remaining key points using the average offset. After correction, calculate the offset between each remaining key point and the preset key point.

[0086] Step 3.6: Determine the offset of each key point (! x ,! y If the error rate is within the allowable tolerance range, then clear i and return to step 3.1; otherwise, i = i + 1.

[0087] Step 3.7: Determine if i is equal to the sensitive parameter n. If not, return to step 3.1. Otherwise, assume the user has made a re-signing action, pause the current sign language reasoning, and enter the re-signing confirmation process.

[0088] In the re-typing process of the word-level sign language continuous translation mode, the user re-typs the sign language, while another screen displays the message content being modified. After the new message is translated, it completely overwrites the original message, the cursor jumps to the end, and the sign language re-typing process is exited, continuing the normal sign language process.

[0089] In this invention, the word-level sign language continuous translation mode is as follows:

[0090] The sign language recognition device is activated, reading the sensitive parameters 'n' preset by the hearing-impaired user, the sign language redo countdown 'sec', and the sign language redo marker gesture; the hearing-impaired user adjusts their position to ensure their hand movements are fully captured by the camera; sign language recognition and translation proceed normally.

[0091] When the sign language recognition device detects the sign language re-signature action, the normal sign language recognition and translation process is interrupted and the sign language re-signature process begins. The interface for hearing-impaired users defaults to selecting the current sentence and prompts the hearing-impaired user by changing the font appearance of the selected sentence.

[0092] When a hearing-impaired user uses the interface, they are asked whether to retype the current sentence. After several confirmations and a countdown timer, a final confirmation is made. If the countdown ends or the hearing-impaired user makes an approval gesture within the countdown, the sentence in the chat interface will be highlighted with a blinking icon, indicating that the sentence content is being modified. The hearing-impaired user then retypes the sentence in sign language. After the new sentence is translated, it replaces the original sentence, the cursor jumps to the end, and the retyped sign language process is exited, allowing the user to continue with normal sign language input.

[0093] In the sentence-level sign language continuous translation mode, during the normal sign language translation process, the control unit obtains the key point coordinate data of each frame of sign language action, which is recorded as the raw data.

[0094] The retyping process for the sentence-level sign language continuous translation mode includes the following steps:

[0095] Step 2.1: Calculate the parameter matrix of the original data frame in reverse order, denoted as... Characterizes the relative distances and relative angles between keypoints in the s-th frame of the original data frame;

[0096] Step 2.2: The user re-enters the sign language, obtaining the key point data of each frame of the sign language, which is recorded as a secondary data frame. The parameter matrix of each secondary data frame is then obtained.

[0097] Step 2.3: Determine if each frame is invalid. If it is, clear t and return to step 2.2; otherwise, ... With each Perform matrix subtraction to obtain the staggered matrix PDM. st ,

[0098] In step 2.3, the determination of invalid frames includes the following steps:

[0099] Step 2.3.1: Collect the raw data of sign language action frames, feed them into the human pose estimation network, extract the human key points involved in any frame read in Step 2.3.1, and obtain the key point matrix;

[0100] Step 2.3.2: Determine whether the frame contains necessary key points, whether the relative positional relationship between key points is reasonable, and whether the absolute position of key points is reasonable. If all of these are true, the frame is valid; otherwise, it is invalid.

[0101] Step 2.4: For each parametric matrix PDM st The contrast coefficient CC is obtained by performing a weighted summation. st ;

[0102] Step 2.4 includes the following steps:

[0103] Step 2.4.1: For parameter matrix PM1,

[0104]

[0105] Where m represents the total number of original data frames for the current sentence, and p i,j This represents the parameter array consisting of all parameters in the i-th row and j-th column of PM1, where i ≠ j;

[0106] Calculate the standard deviation of each parameter.

[0107] The standard deviation matrix SDM is obtained.

[0108] Step 2.4.2: Unify the upper and lower triangular matrices in the SDM to obtain the weighted matrix WM. in,

[0109] Step 2.4.3: Based on each weighting coefficient in the weighting matrix, perform a PDM analysis on each non-parallel matrix. st The weighted summation is performed to obtain the corresponding contrast coefficient CC. st ,

[0110] Where k represents the number of key points, i represents the i-th row, j represents the j-th column, and i ≠ j.

[0111] Step 2.5: Define the sensitivity parameter n, n≥10; if t is less than the sensitivity parameter n, then t=t+1, return to step 2.2, otherwise proceed to the next step;

[0112] Step 2.6: Based on the contrast coefficient CC st Calculate the n consecutive frames in the original data frame that have the smallest sum of contrast coefficients.

[0113]

[0114] Where m represents the total number of original data frames in the current sentence, and the first frame in n frames is denoted as frame k; retain the key point data of the original frames before frame k, discard the key point data of the original frames after frame k, and replace them with secondary data frames;

[0115] Step 2.7: Merge the original data with the secondary data of the secondary data frame, send it to the sign language model for translation, replace the original statement after the new statement is translated, jump the cursor to the end, exit the sign language re-typing process, and continue the normal sign language process.

[0116] In this invention, the pre-processing steps of the sentence-level sign language continuous translation mode are similar to those of the word-level sign language continuous translation mode. However, the control unit obtains the key point coordinate data of each frame of sign language action, which is recorded as the original data. After starting to re-encode the sign language, the parameter matrix of the original data frame is calculated in reverse order to obtain the key point data of each frame of the re-encoded sign language, which is recorded as the secondary data frame. The parameter matrix of each secondary data frame is then obtained. If it is an invalid frame, t is cleared to zero and the process returns. Otherwise, the difference matrix is ​​calculated, and a weighted summation is performed on each difference matrix to obtain the contrast coefficient CC. st ;

[0117] After obtaining enough data, based on the contrast coefficient CC st Calculate the n frames with the smallest sum of contrast coefficients among the n consecutive frames in the original data frame. The key point data of the original frames before frame k is retained, while the key point data of the original frames after frame k is discarded and replaced with secondary data frames. The original data and secondary data are merged and fed into the sign language model for inference. After the new sentence is translated, the original sentence is replaced, the cursor jumps to the end, the sign language re-typing process is exited, and the normal sign language process continues.

Claims

1. A sign language recognition system adapted to sign language re-typing processing, characterized in that: The system includes a host computer, and at least two screens are provided in conjunction with the host computer; one of the screens is equipped with a data acquisition component and an adjustment component. A control unit is provided in conjunction with the screen, acquisition component, and adjustment component. This control unit performs sign language re-translation processing using both word-level and sentence-level continuous translation modes. The word-level mode replaces the entire sentence requiring re-expression, while the sentence-level mode segments the sentence, retaining one part and replacing the other. The control unit acquires and saves re-translation marker actions and other marker actions through the acquisition component. During sign language translation, when the acquisition component captures a re-translation marker action, a re-translation confirmation process is initiated. Other marker actions include cancellation, approval, and selection. Distinguishing between re-translation marker actions and other marker actions involves the following steps: Step 3.1: Input each frame of sign language motion image data captured by the acquisition component into the human pose estimation network to obtain the coordinate data of the sign language key points in each frame; Step 3.2: Sequentially calculate the frame difference of key point coordinate data between every two frames. Save as ; Step 3.3: When there are at least two frame differences, perform a t-test on two consecutive frame differences; Step 3.4: If and If the difference is greater than the preset value, then i is cleared and the process returns to step 3.1; otherwise, proceed to the next step. Step 3.5: Calculate the frame difference between the sign language key points in each frame and the preset gesture key points. Calculate the average offset of the head and shoulder key points in two frames. Correct the remaining key points using the average offset. After correction, calculate the offset between each remaining key point and the preset key point. Step 3.6: Determine the offset of each key point If the error rate is within the allowable tolerance range, then clear i and return to step 3.1; otherwise, i = i + 1. Step 3.7: Determine if i is equal to the sensitive parameter n. If not, return to step 3.

1. Otherwise, assume the user has made a re-signing action, pause the current sign language reasoning, and enter the re-signing confirmation process.

2. The sign language recognition system adapted for sign language re-typing processing according to claim 1, characterized in that: The re-call confirmation process includes the following steps: Step 1.1: In the normal sign language process, when the acquisition component captures the re-signature action, proceed to the next step; otherwise, repeat step 1.

1. Step 1.2: The control unit pre-selects the translated sentences to assist the sign language interpreter in selecting the sentences that need to be re-typed; Step 1.3: The screen displays a countdown. If the user cancels within the countdown, return to Step 1.

2. Otherwise, after the countdown ends or the user acknowledges within the countdown, start re-entering the process until it is completed and return to the normal sign language process.

3. A sign language recognition system adapted for sign language re-typing processing according to claim 2, characterized in that: In step 1.2, assisting the sign language interpreter in selecting the phrases that need to be retyped includes the following steps: Step 1.2.1: Pre-select the current statement by changing the font color; Step 1.2.2: The control unit prompts whether to retype the current statement. If the statement to be retyped is not the current statement or the retyped statement needs to be cancelled, the sign language interpreter cancels the statement and proceeds to the next step. Otherwise, the sign language interpreter acknowledges the statement and proceeds to step 1.

3. Step 1.2.3: The control unit prompts whether to cancel the re-signing. If yes, the signer makes the cancel action again, exits the re-signing process, and returns to the normal sign language process. Otherwise, the signer selects the preceding statement by selecting an action. Step 1.2.4: Repeat step 1.2.

2.

4. A sign language recognition system adapted for sign language re-typing processing according to claim 2, characterized in that: In the re-typing process of the word-level sign language continuous translation mode, the user re-typs the sign language, while another screen displays the message content being modified. After the new message is translated, it completely overwrites the original message, the cursor jumps to the end, and the sign language re-typing process is exited, continuing the normal sign language process.

5. A sign language recognition system adapted for sign language re-typing processing according to claim 2, characterized in that: In the sentence-level sign language continuous translation mode, during the normal sign language translation process, the control unit obtains the key point coordinate data of each frame of sign language action, which is recorded as the raw data.

6. A sign language recognition system adapted for sign language re-typing processing according to claim 5, characterized in that: The retyping process for the sentence-level sign language continuous translation mode includes the following steps: Step 2.1: Calculate the parameter matrix of the original data frame in reverse order, denoted as... , representing the relative distance and relative angle between each key point in the s-th frame of the original data frame; Step 2.2: The user re-enters the sign language, obtaining the key point data of each frame of the sign language, which is recorded as a secondary data frame. The parameter matrix of each secondary data frame is then obtained. ; Step 2.3: Determine if each frame is invalid. If it is, clear t and return to step 2.2; otherwise, ... With each Perform matrix subtraction to obtain the staggered matrix. , ; Step 2.4: For each parametric matrix The contrast coefficient is obtained by performing a weighted summation. ; Step 2.5: Define the sensitivity parameter n. If t is less than the sensitive parameter n, then t = t + 1, return to step 2.2; otherwise, proceed to the next step. Step 2.6: Based on the contrast coefficient Calculate the n consecutive frames in the original data frame that have the smallest sum of contrast coefficients. , Where m represents the total number of original data frames in the current sentence, and the first frame in n frames is denoted as frame q; retain the key point data of the original frames before frame q, discard the key point data of the original frames after frame q, and replace them with secondary data frames; Step 2.7: Merge the original data with the secondary data of the secondary data frame, send it to the sign language model for translation, replace the original statement after the new statement is translated, jump the cursor to the end, exit the sign language re-typing process, and continue the normal sign language process.

7. A sign language recognition system adapted for sign language re-typing processing according to claim 6, characterized in that: In step 2.3, the determination of invalid frames includes the following steps: Step 2.3.1: Collect the raw data of sign language action frames, feed them into the human pose estimation network, extract the human key points involved in any frame read in Step 2.3.1, and obtain the key point matrix; Step 2.3.2: Determine whether the frame contains necessary key points, whether the relative positional relationship between key points is reasonable, and whether the absolute position of key points is reasonable. If all of these are true, the frame is valid; otherwise, it is invalid.

8. A sign language recognition system adapted for sign language re-typing processing according to claim 6, characterized in that: Step 2.4 includes the following steps: Step 2.4.1: For the parameter matrix , , Where m represents the total number of original data frames for the current sentence. Indicates in The parameter array consisting of all parameters in the i-th row and j-th column. ; Calculate the standard deviation of each parameter. , Obtain the standard deviation matrix , ; Step 2.4.2: For The weighted matrix is ​​obtained by unifying the upper and lower triangular matrices respectively. , ,in, ; Step 2.4.3: Based on each weighting coefficient in the weighting matrix, for each staggered matrix... Perform a weighted summation to obtain the corresponding contrast coefficient. , Where k represents the number of key points, i represents the i-th row, and j represents the j-th column. .

Citation Information

Patent Citations

  • Gesture recognition bracelet for artificial intelligence and gesture recognition method thereof

    CN114721515A

  • Sign language recognition method and device, equipment and storage medium

    CN115171212A

  • Sign language interpreting device and method

    CN108877408A

  • Sign language translation confirming device

    JP1994337628A