Sign language recognition method and device, electronic equipment and storage medium
By detecting, classifying, labeling, and connecting key points of sign language videos, and using limb-color mapping relationships for sign language recognition, the problems of high cost and low accuracy of sign language recognition devices are solved, achieving high-accuracy and low-cost sign language recognition.
Patent Information
- Application Number
- CN202211635635.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-19
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2042-12-19
AI Technical Summary
Existing sign language recognition devices are expensive and inaccurate, mainly because they treat all information in an RGB image equally without considering the relationships between the information.
By detecting, classifying, labeling, and connecting pose key points in each frame of the video to be recognized, using the limb-color mapping relationship for labeling, extracting and explicitly constructing the association relationship of pose key points, sign language recognition is performed.
It improves the accuracy of sign language recognition and reduces equipment costs. It is suitable for ordinary imaging devices and can be adapted to the recognition of different sign language movements.
Smart Images

Figure CN116259102B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of sign language translation, and in particular to a sign language recognition method and device, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of computer vision technology, the application scenarios of sign language recognition are becoming more and more extensive. Sign language recognition is to translate the collected sign language video into text or voice.
[0003] At present, human posture action information is mainly obtained based on an RGBD depth camera for sign language recognition. However, it has a high requirement for imaging hardware, resulting in a high cost of sign language recognition. For sign language recognition of an RGB video, the prior art treats all information represented by an RGB image equally, which does not consider the mutual relationship of the information in the image, and thus the accuracy of sign language recognition is not high. SUMMARY
[0004] The present application provides a sign language recognition method, device, electronic device and storage medium to solve the defect of low accuracy of sign language recognition in the prior art and realize high-accuracy sign language recognition.
[0005] The present application provides a sign language recognition method, comprising:
[0006] Performing posture key point detection on each frame of image in a to-be-recognized video to obtain a posture key point map corresponding to each frame of image, wherein any posture key point map comprises a plurality of posture key points;
[0007] Classifying and marking the plurality of posture key points of each posture key point map to obtain a posture map corresponding to each posture key point map;
[0008] Performing sign language recognition on each posture map to obtain a sign language recognition result of the to-be-recognized video.
[0009] According to the sign language recognition method provided by the present application, the plurality of posture key points of each posture key point map are classified and marked to obtain a posture map corresponding to each posture key point map, which comprises:
[0010] Based on a preset key point classification rule, the plurality of posture key points of each posture key point map are classified to obtain a key point classification result corresponding to each posture key point map, wherein the preset key point classification rule is used to represent the posture key points included in each limb of a human body, and the limbs are limbs related to sign language actions;
[0011] Based on the key point classification result, the plurality of posture key points of each posture key point graph are connected to obtain a connection graph corresponding to each posture key point graph.
[0012] Based on the key point classification result, the limb connection part of each connection graph is marked to obtain a posture graph corresponding to each connection graph.
[0013] According to the hand language recognition method provided by the application, the limb connection part of each connection graph is marked based on the key point classification result to obtain a posture graph corresponding to each connection graph, which comprises:
[0014] Based on the limb-color mapping relationship and the key point classification result, the limb connection part of each connection graph is color marked to obtain a posture graph corresponding to each connection graph.
[0015] The limb-color mapping relationship is used to represent the mapping relationship between the limbs and different colors.
[0016] According to the hand language recognition method provided by the application, the hand language recognition result comprises a sentence-level recognition result and / or a word-level recognition result.
[0017] The sentence-level recognition result is obtained by decoding a sentence-level encoding vector, and the sentence-level encoding vector is determined based on the context information of each posture graph.
[0018] The word-level recognition result is obtained by decoding a word-level encoding vector of each posture graph.
[0019] According to the hand language recognition method provided by the application, the sentence-level recognition result is obtained based on the following steps:
[0020] Each posture graph is word-level encoded to obtain a word-level encoding vector of each posture graph.
[0021] Based on the context information of each word-level encoding vector, each word-level encoding vector is sentence-level encoded to obtain a sentence-level encoding vector.
[0022] The sentence-level encoding vector is decoded to obtain the sentence-level recognition result.
[0023] According to the hand language recognition method provided by the application, the sentence-level encoding vector is decoded to obtain the sentence-level recognition result, which comprises:
[0024] If the current decoding turn is not the first decoding turn, decoding the sentence-level encoding vector based on the word recognition result of the previous decoding turn to obtain the word recognition result of the current decoding turn;
[0025] updating the previous decoding turn as the current decoding turn until the current decoding turn is the last decoding turn;
[0026] determining the sentence-level recognition result based on the word recognition result of each decoding turn;
[0027] If the current decoding turn is the first decoding turn, decoding the sentence-level encoding vector to obtain the word recognition result of the current decoding turn;
[0028] taking the current decoding turn as the previous decoding turn.
[0029] According to the hand language recognition method provided by the application, any frame image in the video to be recognized is subjected to posture key point detection based on the following steps:
[0030] If the any frame image is not the first frame image, the any frame image is subjected to aggregation processing with the previous frame posture image to obtain aggregated image data;
[0031] The aggregated image data is subjected to posture key point detection to obtain a posture key point image corresponding to the any frame image;
[0032] The previous frame posture image is updated as a posture image corresponding to the posture key point image corresponding to the any frame image until the any frame image is the last frame image;
[0033] If the any frame image is the first frame image, the any frame image is subjected to posture key point detection to obtain a posture key point image corresponding to the any frame image.
[0034] The application further provides a hand language recognition device, comprising:
[0035] The detection module is configured to perform posture key point detection on each frame image in the video to be recognized to obtain a posture key point image corresponding to each frame image, and any posture key point image comprises a plurality of posture key points;
[0036] The classification module is configured to classify and mark the plurality of posture key points of each posture key point image to obtain a posture image corresponding to each posture key point image;
[0037] The recognition module is configured to perform hand language recognition on each posture image to obtain a hand language recognition result of the video to be recognized.
[0038] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the sign language recognition method according to any one of the above when executing the program.
[0039] The present application also provides a non-transitory computer readable storage medium, which stores a computer program, wherein the computer program is executable on a processor to implement the sign language recognition method according to any one of the above.
[0040] The present application provides a sign language recognition method, device, electronic device and storage medium, which detects posture key points of each frame image in a video to be recognized to obtain posture key point maps corresponding to each frame image, thereby removing redundant information in the video to be recognized and extracting posture information related to sign language actions, so as to improve the accuracy of sign language recognition; classifies and labels a plurality of posture key points of each posture key point map to obtain posture maps corresponding to each posture key point map, thereby distinguishing each posture key point from each other, grouping posture key points explicitly, and constructing grouped posture key points explicitly, so that when each posture map is subjected to sign language recognition, the association between each part of the human body represented by the grouped posture key points and the action characteristics of the sign language word can be recognized, thereby further improving the accuracy of sign language recognition. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0042] Figure 1 One of the flowcharts of the sign language recognition method provided by the present application;
[0043] Figure 2 The layout diagram of the posture key points provided by the present application;
[0044] Figure 3 The second flowchart of the sign language recognition method provided by the present application;
[0045] Figure 4 The third flowchart of the sign language recognition method provided by the present application;
[0046] Figure 5 The structural diagram of the sign language recognition device provided by the present application;
[0047] Figure 6 The structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0048] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the protection scope of the present application.
[0049] Sign language is an important language for deaf people to communicate with each other, which mainly relies on hand shape, hand position, movement, facial expression and body movement to express specific meaning to convey information.
[0050] In the traditional way, sign language translation talents are trained to ensure the communication between hearing-impaired people and the communication between hearing-impaired people and hearing people. However, there is a large gap in sign language translation talents, and most places still lack effective sign language translation services. Therefore, with the rapid development of computer vision technology, the application scenarios of sign language recognition are becoming more and more extensive, so as to realize the translation between sign language and text and the translation between sign language and voice by means of deep learning technology. Sign language recognition is to translate the collected sign language video into text or voice to realize the barrier-free communication and exchange between hearing-impaired people and hearing people.
[0051] At present, most of the human posture action information is obtained based on RGBD depth camera for sign language recognition. However, it has high requirements for imaging hardware, so that the overall sign language recognition device cannot be light and low in price, resulting in high cost of sign language recognition. In addition, the realization of sign language recognition based on wearable devices requires users to wear necessary action information capture devices such as wristbands, gloves and the like, which limits the use scene and makes the overall device heavy and high in price, which is not suitable for popularization and application. For sign language recognition of RGB video, the prior art treats each information represented by the RGB image equally, which leads to the fact that the mutual relationship of each information in the image is not considered in the sign language recognition process, and thus the accuracy of sign language recognition is not high.
[0052] Even if the key points of the sign language video are detected, the prior art is to splice each key point into a sequence and then input the sign language recognition model for sign language recognition, which is easily affected by the specific value fluctuation of the key points. When the key points are slightly dithered, the sign language recognition model will output completely different sign language recognition results, while the meaning of the sign language action does not change significantly. In addition, different key points can only be treated equally, while different key points have mutual relationships, and the mutual relationships have a strong correlation with the expression of sign language words. Therefore, the accuracy of sign language recognition in the prior art is not high.
[0053] To solve the above problems, the present application provides the following embodiments. Figure 1 As shown in a flowchart of a sign language recognition method provided by the present application, Figure 1 The sign language recognition method comprises the following steps.
[0054] In step 110, posture key point detection is performed on each frame image in the video to be recognized to obtain posture key point maps corresponding to the frame images.
[0055] Here, the video to be recognized is a video to be recognized for sign language, which includes a sign language action video, that is, the video to be recognized includes multiple frames of sign language images to form a sign language action video based on the multiple frames of sign language images, in other words, the video to be recognized includes multiple frames of images of continuous sign language actions.
[0056] The video to be recognized can be collected by a low-cost RGB imaging device, such as a CMOS (Complementary Metal Oxide Semiconductor) camera, or by other imaging devices, which are not limited in the embodiments of the present application.
[0057] Here, any frame image in the video to be recognized represents posture information of sign language actions, which can include but is not limited to at least one of the following: hand shape, hand position, limb shape, limb position, facial expression, facial position, etc. Based on this, the information represented by each frame image can include but is not limited to hand action, facial action, limb action, etc.
[0058] In an embodiment, each frame image can be an RGB image, which can be collected by a low-cost RGB imaging device. Of course, each frame image can also be an image of other forms. Based on this, the embodiments of the present application do not need to rely on specific hardware, but only need a common imaging device to complete sign language recognition.
[0059] Here, the posture key point map can be used to represent the spatial position of each posture key point, that is, to cover the spatial coordinates of each posture key point. In an embodiment, the posture key point map is a heat map, and the peak position of the heat map is the spatial position of the corresponding posture key point.
[0060] Here, the posture key point is a key point related to sign language actions, that is, a key point in the limb related to sign language actions, which plays a key role in the representation of sign language. The positions or locations of the multiple posture key points in the human body can be set according to actual needs.
[0061] In an embodiment, as shown in Figure 2As shown, the plurality of pose keypoints can include, but are not limited to, at least one of the following: a set of hand thumb keypoints, a set of hand index finger keypoints, a set of hand middle finger keypoints, a set of hand ring finger keypoints, a set of hand pinky finger keypoints, a set of head-shoulder keypoints, a set of body trunk keypoints, a set of hand trunk keypoints, and the like. The set of hand thumb keypoints is illustrated by way of example with a set of thumb keypoints of one hand, which can include, but are not limited to, at least one of the following: a first metacarpal keypoint 14, a second metacarpal keypoint 15, a first thumb joint keypoint 16, a second thumb joint keypoint 17, a third thumb joint keypoint 18, and the like. The set of hand index finger keypoints is illustrated by way of example with a set of index finger keypoints of one hand, which can include, but are not limited to, at least one of the following: the first metacarpal keypoint 14, a first index finger joint keypoint 19, a second index finger joint keypoint 20, a third index finger joint keypoint 21, a fourth index finger joint keypoint 22, and the like. The set of hand middle finger keypoints is illustrated by way of example with a set of middle finger keypoints of one hand, which can include, but are not limited to, at least one of the following: the first metacarpal keypoint 14, a first middle finger joint keypoint 23, a second middle finger joint keypoint 24, a third middle finger joint keypoint 25, a fourth middle finger joint keypoint 26, and the like. The set of hand ring finger keypoints is illustrated by way of example with a set of ring finger keypoints of one hand, which can include, but are not limited to, at least one of the following: the first metacarpal keypoint 14, a first ring finger joint keypoint 27, a second ring finger joint keypoint 28, a third ring finger joint keypoint 29, a fourth ring finger joint keypoint 30, and the like. The set of hand pinky finger keypoints is illustrated by way of example with a set of pinky finger keypoints of one hand, which can include, but are not limited to, at least one of the following: the first metacarpal keypoint 14, a first pinky finger joint keypoint 31, a second pinky finger joint keypoint 32, a third pinky finger joint keypoint 33, a fourth pinky finger joint keypoint 34, and the like. The set of head-shoulder keypoints can include, but are not limited to, at least one of the following: a right eye keypoint 1, a left eye keypoint 2, a nose keypoint 3, a right ear keypoint 4, a left ear keypoint 5, a right shoulder keypoint 6, a left shoulder keypoint 7, and the like. The set of body trunk keypoints can include, but are not limited to, at least one of the following: the right shoulder keypoint 6, the left shoulder keypoint 7, a right hip joint keypoint 12, a left hip joint keypoint 13, and the like. The set of hand trunk keypoints can include, but are not limited to, at least one of the following: the right shoulder keypoint 6, a right elbow joint keypoint 8, a right wrist joint keypoint 10, the left shoulder keypoint 7, a left elbow joint keypoint 9, a left wrist joint keypoint 11, and the like.
[0062] Specifically, based on the posture key point detection model, posture key points of each frame image in the to-be-identified video are detected frame by frame to obtain each posture key point map. It can be understood that the to-be-identified video usually only includes one action subject, and based on this, the above-mentioned each posture key point map of the action subject can be obtained, that is, the posture key point detection model is used to detect the posture key points of the action subject in the to-be-identified video. The specific structure and training method of the posture key point detection model can be set according to actual needs, and the embodiments of the present application do not make specific limitations thereto.
[0063] In step 120, the plurality of posture key points of each of the posture key point maps are classified and labeled to obtain a posture map corresponding to each of the posture key point maps.
[0064] It should be noted that considering that the whole process of sign language action is jointly constituted by the motion trajectories of each part of the human body, and each part is indispensable, and most importantly, each part has different weights on the expressed semantics. For example, when expressing the semantics related to "politeness", the main semantic related to the action is bowing, and the details of the hand joints are not important; for example, when expressing the semantics related to "thank you", the main semantic related to the action is the whole hand posture and the thumb action, and other joint details are not important. Therefore, the plurality of posture key points need to be classified and labeled, so that different vocabularies are strongly associated with the motion trajectories of the corresponding body parts when subsequent sign language recognition is performed, thereby improving the accuracy of sign language recognition.
[0065] Specifically, any of the pose keypoint maps is classified and labeled based on the following steps: based on the mutual relationship of the plurality of pose key points of the any of the pose keypoint maps, the plurality of pose key points are classified to obtain a key point classification result corresponding to the any of the pose keypoint maps; based on the key point classification result, the plurality of pose key points are labeled to obtain a pose map corresponding to the any of the pose keypoint maps. More specifically, based on a preset key point classification rule, the plurality of pose key points of the any of the pose keypoint maps are classified to obtain a key point classification result corresponding to the any of the pose keypoint maps, and the preset key point classification rule is determined based on the mutual relationship of the plurality of pose key points. For example, the right shoulder key point has a connection relationship with the right elbow joint key point, the right elbow joint key point has a connection relationship with the right wrist joint key point, and the right shoulder key point, the right elbow joint key point and the right wrist joint key point all belong to the key points of the right hand trunk, then the right shoulder key point, the right elbow joint key point and the right wrist joint key point are classified into the same class, i.e. the right hand trunk class. For another example, the left shoulder key point has a connection relationship with the left elbow joint key point, the left elbow joint key point has a connection relationship with the left wrist joint key point, and the left shoulder key point, the left elbow joint key point and the left wrist joint key point all belong to the key points of the left hand trunk, then the left shoulder key point, the left elbow joint key point and the left wrist joint key point are classified into the same class, i.e. the left hand trunk class. Wherein, the right hand trunk class and the left hand trunk class can be the same class, i.e. the double hand trunk class, or can be different classes.
[0066] Here, the pose map is used to represent the pose information and state of the action subject person in the current frame. The pose information can include but is not limited to at least one of the following: hand shape, hand position, limb shape, limb position, facial expression, facial position, etc. Based on this, the information represented by each pose map can include but is not limited to: hand action, facial action, limb action, etc., that is, by processing each frame image of the to-be-identified video frame by frame, the pose stream data formed by each pose map can be obtained to represent a series of sign language action state information of the sign language expresser based on the pose stream data, that is, the pose stream data records the action direction, hand shape and other key information of the expresser, thereby converting the to-be-identified video into pose stream data, significantly reducing the redundancy information of the to-be-identified video, and further improving the accuracy of sign language recognition. In other words, the pose stream data records the changes of the spatial positions of each joint and limb over time, and can fully represent the motion state information.
[0067] It can be understood that the multiple gesture key points are classified and marked, that is, the gesture key point map is rendered, so that the gesture map explicitly constructs the gesture key points, and each gesture key point is distinguished from each other, so that when subsequent sign language recognition is performed, different vocabularies are strongly associated with the motion trajectories of the corresponding body parts, and the accuracy of sign language recognition is improved. At the same time, when subsequent sign language recognition is performed, the information of each part of the body after grouping of the gesture key points can be treated with different weights on the basis of the classified and marked gesture key points; and when each gesture key point is classified and marked, the accurate sign language recognition result can still be recognized when the coordinate values of each gesture key point are small-range jittered, so that the gesture key points have a good tolerance for jitter.
[0068] Step 130, performing sign language recognition on each gesture map to obtain a sign language recognition result of the video to be recognized.
[0069] Here, the sign language recognition result can include but is not limited to a sentence-level recognition result and a word-level recognition result, etc. The sentence-level recognition result is a continuous sentence text result represented by the video to be recognized; further, the continuous sentence text result can be a continuous sentence text under a normal hearing person's order, which takes into account that the order of sign language can be different from that of a normal hearing person. The word-level recognition result is a sign language word text result represented by each segment video in the video to be recognized.
[0070] Specifically, based on a sign language recognition model, each gesture map is encoded to obtain each encoding vector, and each encoding vector is decoded to obtain a sign language recognition result. The specific structure of the sign language recognition model is not specifically limited here.
[0071] The sign language recognition method provided by the embodiment of the present application performs gesture key point detection on each frame image in the video to be recognized to obtain a gesture key point map corresponding to each frame image, thereby removing redundant information in the video to be recognized and extracting gesture information related to sign language actions, so as to improve the accuracy of sign language recognition; multiple gesture key points of each gesture key point map are classified and marked to obtain a gesture map corresponding to each gesture key point map, so as to distinguish each gesture key point from each other, explicitly group the gesture key points, and explicitly construct the grouped gesture key points, so that when each gesture map is subjected to sign language recognition, the association relationship between each part of the body represented by the grouped gesture key points and the motion characteristics of the sign language words can be recognized, thereby further improving the accuracy of sign language recognition.
[0072] Based on the above embodiment, Figure 3 The flowchart of the sign language recognition method provided by the present application is shown in Figure 2, which shows that the above step 120 includes: Figure 3 As shown in Figure 2, the above step 120 includes:
[0073] Step 121: Based on the preset key point classification rules, classify the multiple posture key points of each posture key point map to obtain the key point classification results corresponding to each posture key point map. The preset key point classification rules are used to characterize the posture key points included in each limb of the human body. Each limb is a limb related to sign language movements.
[0074] It should be noted that the entire process of sign language gestures is composed of the movement trajectories of various limbs and trunks of the human body, and each limb and trunk is indispensable. Most importantly, each limb and trunk has a different weight in expressing semantics. Therefore, it is necessary to classify multiple postural key points to establish a strong association between different sign language words and the corresponding limb and trunk movement trajectories during subsequent sign language recognition, thereby improving the accuracy of sign language recognition.
[0075] For example, such as Figure 2 As shown, each limb trunk may include, but is not limited to: head and shoulders, main body trunk, main hands, thumb, index finger, middle finger, ring finger, little finger, etc. The head and shoulders may include, but is not limited to, at least one of the following: right eye key point 1, left eye key point 2, nose key point 3, right ear key point 4, left ear key point 5, right shoulder key point 6, left shoulder key point 7, etc. The main body trunk may include, but is not limited to, at least one of the following: right shoulder key point 6, left shoulder key point 7, right hip joint key point 12, left hip joint key point 13, etc. The main hands may include, but is not limited to, at least one of the following: right shoulder key point 6, right elbow joint key point 8, right wrist joint key point 10, left shoulder key point 7, left elbow joint key point 9, left wrist joint key point 11, etc. The thumb may include, but is not limited to, at least one of the following: first palmar root key point 14, second palmar root key point 15, first thumb joint key point 16, second thumb joint key point 17, third thumb joint key point 18, etc. The index finger may include, but is not limited to, at least one of the following: key point 14 of the first palm base, key point 19 of the first index finger joint, key point 20 of the second index finger joint, key point 21 of the third index finger joint, key point 22 of the fourth index finger joint, etc. The middle finger may include, but is not limited to, at least one of the following: key point 14 of the first palm base, key point 23 of the first middle finger joint, key point 24 of the second middle finger joint, key point 25 of the third middle finger joint, key point 26 of the fourth middle finger joint, etc. The ring finger may include, but is not limited to, at least one of the following: key point 14 of the first palm base, key point 27 of the first ring finger joint, key point 28 of the second ring finger joint, key point 29 of the third ring finger joint, key point 30 of the fourth ring finger joint, etc. The little finger may include, but is not limited to, at least one of the following: key point 14 of the first palm base, key point 31 of the first little finger joint, key point 32 of the second little finger joint, key point 33 of the third little finger joint, key point 34 of the fourth little finger joint, etc. Based on the above, preset key point classification rules can be set according to actual needs.
[0076] Step 122: Based on the classification results of each key point, connect the multiple attitude key points of each attitude key point map to obtain the connection map corresponding to each attitude key point map.
[0077] Specifically, based on the classification results of each keypoint, multiple keypoints in each pose keypoint map are grouped and connected to obtain the connection map corresponding to each pose keypoint map. Each connection map includes at least one limb connection part.
[0078] For example, such as Figure 2 As shown, assuming the classification results in groups for head and shoulder key points, body trunk key points, hands trunk key points, thumb trunk key points, index finger trunk key points, middle finger trunk key points, ring finger trunk key points, and little finger trunk key points; based on this, the right eye key point 1, left eye key point 2, and nose key point 3 in the head and shoulder key point group are connected sequentially; the right eye key point 1, right ear key point 4, and right shoulder key point 6 in the head and shoulder key point group are connected sequentially; and the left eye key point 1, right ear key point 4, and right shoulder key point 6 in the head and shoulder key point group are connected sequentially. Connect the key points 2 for the eye, 5 for the left ear, and 7 for the left shoulder in sequence; connect the key points 6 for the right shoulder, 7 for the left shoulder, 13 for the left hip joint, and 12 for the right hip joint in sequence from the group of key points on the main body; connect the key points 6 for the right shoulder, 8 for the right elbow joint, and 10 for the right wrist joint in sequence from the group of key points on the main body of both hands; connect the key points 7 for the left shoulder, 9 for the left elbow joint, and 11 for the left wrist joint in sequence from the group of key points on the main body of both hands; connect the key point on the first palm base in the group of key points on the thumb. Connect key points 14, 15, 16, 17, and 18 of the first thumb joint in sequence; connect key points 14, 19, 20, 21, and 22 of the first palmar root joint in sequence; connect key points 14, 19, 20, 21, and 22 of the first index finger joint in sequence; connect key points 14, 23, 24, and 25 of the first palmar root joint in sequence; connect key points 16, 17, and 28 of the second thumb joint in sequence; connect key points 18, 19, 20, 21, and 22 of the third index finger joint in sequence; connect key points 14, 23, 24, and 25 of the first palmar root joint in sequence; connect key points 19, 20, 21, and 22 of the second index finger joint in sequence; connect key points 14, 23, 24, and 25 of the second palmar root joint in sequence; connect key points 16, 20, 21, and 22 of the third index finger joint in sequence; connect key points 19, 20, 21, and 22 of the third palmar root ... Key points 25 and 26 of the fourth middle finger joint are connected sequentially; key points 14 of the first palm base, 27 of the first ring finger joint, 28 of the second ring finger joint, 29 of the third ring finger joint, and 30 of the fourth ring finger joint are connected sequentially in the key point group of the ring finger; key points 14 of the first palm base, 31 of the first little finger joint, 32 of the second little finger joint, 33 of the third little finger joint, and 34 of the fourth little finger joint are connected sequentially in the key point group of the little finger.
[0079] It can be understood that, based on the key point classification results, the plurality of posture key points are connected respectively to obtain each limb connection part. Even if the individual posture key point coordinate values in one limb connection part are small range jitter, since the overall limb connection part has little influence, an accurate sign language recognition result can still be recognized, and the jitter of the posture key points has good tolerance.
[0080] At step 123, based on the key point classification results, the limb connection parts of each connection graph are marked to obtain the posture graph corresponding to each connection graph.
[0081] Specifically, any connection graph is marked based on the following manner: based on the key point classification result corresponding to the any connection graph, the marking manner of each limb connection part is determined, and each limb connection part of the any connection graph is marked based on each marking manner to obtain the posture graph corresponding to the any connection graph. The specific marking manner can be set according to actual needs, as long as the subsequent sign language recognition can distinguish the information of each limb connection part.
[0082] The marking manners corresponding to different limb connection parts are different, for example, the head and shoulder correspond to a first marking manner, the body trunk corresponds to a second marking manner, the double-hand trunk corresponds to a third marking manner (the left-hand trunk and the right-hand trunk can also correspond to different marking manners), the thumb corresponds to a fourth marking manner (the left-hand thumb and the right-hand thumb can also correspond to different marking manners), the index finger corresponds to a fifth marking manner (the left-hand index finger and the right-hand index finger can also correspond to different marking manners), the middle finger corresponds to a sixth marking manner (the left-hand middle finger and the right-hand middle finger can also correspond to different marking manners), the ring finger corresponds to a seventh marking manner (the left-hand ring finger and the right-hand ring finger can also correspond to different marking manners), and the little finger corresponds to an eighth marking manner (the left-hand little finger and the right-hand little finger can also correspond to different marking manners).
[0083] It can be understood that, based on the key point classification results, the limb connection parts of each connection graph are marked to explicitly construct each limb connection part, so that when the sign language recognition is performed on each posture graph, the association relationship between each limb part and the motion characteristics of the sign language word can be recognized, thereby improving the accuracy of sign language recognition.
[0084] The sign language recognition method provided by the embodiment of the present application classifies the plurality of posture key points of each posture key point graph according to limb trunks, thereby connecting the plurality of posture key points of each posture key point graph based on the classification results of each key point, obtaining each limb trunk connection part of each connection graph, thereby distinguishing each posture key point from each other, thereby improving the accuracy of sign language recognition; then, based on the classification results of each key point, the limb trunk connection part of each connection graph is marked to explicitly construct each limb trunk connection part, so that when each posture graph is recognized, the association between each limb trunk part and the motion characteristics of the sign language word can be recognized, thereby further improving the accuracy of sign language recognition.
[0085] Based on any of the above embodiments, the method comprises:
[0086] Based on the limb trunk-color mapping relationship and the classification results of each key point, the limb trunk connection part of each connection graph is color marked to obtain the posture graph corresponding to each connection graph.
[0087] The limb trunk-color mapping relationship is used to represent the mapping relationship between the plurality of limb trunks and different colors.
[0088] Specifically, any connection graph is marked based on the following manner: based on the key point classification results corresponding to the any connection graph and the limb trunk-color mapping relationship, the marking color of each limb trunk connection part is determined, and each limb trunk connection part of the any connection graph is marked based on the marking color to obtain the posture graph corresponding to the any connection graph.
[0089] The marking colors of different limb trunk connection parts are different. Specifically, the marking color can be set according to actual needs, as long as the subsequent sign language recognition can distinguish the information of each limb trunk connection part, that is, the color difference between different limb trunk connection parts is large, and the specific color is not limited.
[0090] The sign language recognition method provided by the embodiment of the present application color marks the limb trunk connection part of each connection graph based on the limb trunk-color mapping relationship and the classification results of each key point to obtain the posture graph corresponding to each connection graph, thereby providing support for the specific marking manner, marking the limb trunk connection part of each connection graph, explicitly constructing each limb trunk connection part, and when each posture graph is recognized, the association between each limb trunk part and the motion characteristics of the sign language word can be recognized, thereby further improving the accuracy of sign language recognition.
[0091] Based on any of the above embodiments, the sign language recognition result comprises a sentence-level recognition result and / or a word-level recognition result.
[0092] The sentence-level recognition result is obtained by decoding a sentence-level encoding vector, and the sentence-level encoding vector is determined based on context information of each gesture graph.
[0093] The word-level recognition result is obtained by decoding a word-level encoding vector of each gesture graph.
[0094] Here, the sentence-level recognition result is a continuous sentence text result represented by the video to be recognized. Further, the continuous sentence text result can be a continuous sentence text in a normal hearing person's order, considering that the order of the sign language can be different from the order of the normal hearing person.
[0095] Here, the sentence-level encoding vector is a feature representation of the video to be recognized.
[0096] In an embodiment, the sentence-level encoding vector can be decoded by an attention mechanism to obtain the sentence-level recognition result, thereby further improving the accuracy of sign language recognition. Further, a multi-head attention mechanism can be used to strengthen the feature extraction capability, thereby further improving the accuracy of sign language recognition. For example, the decoder for decoding the sentence-level encoding vector can be composed of multiple transformer layers.
[0097] In an embodiment, considering that the order of the sign language can be different from the order of the normal hearing person, the word recognition result obtained in the previous decoding round is also considered when recognizing each sign language word, so as to finally obtain the sentence-level recognition result corresponding to the normal hearing person's order. Based on this, the sentence-level recognition result is decoded based on the following steps: if the current decoding round is not the first decoding round, the word recognition result of the previous decoding round is used to decode the sentence-level encoding vector to obtain the word recognition result of the current decoding round; the previous decoding round is updated to the current decoding round until the current decoding round is the last decoding round; and the word recognition result of each decoding round is used to determine the sentence-level recognition result. If the current decoding round is the first decoding round, the sentence-level encoding vector is decoded to obtain the word recognition result of the current decoding round; and the current decoding round is used as the previous decoding round. The specific execution process can be referred to the following embodiments, which will not be described here.
[0098] Here, the word-level recognition result is a sign language word text result represented by each segment video in the video to be recognized.
[0099] For example, the word-level encoding vector is encoded based on the isolated word encoder. Based on this, one word-level encoding vector is encoded based on multiple frames of pose graphs, that is, one isolated word corresponds to multiple frames of pose graphs; in other words, the isolated word encoder is composed of a 3D convolution block, so that multiple frames of images can be encoded. Of course, the key pose graph corresponding to the current isolated word can be first determined from the multiple frames of pose graphs, so that the key pose graph is encoded to obtain the word-level encoding vector, and the key pose graph is the pose graph that best represents the current isolated word.
[0100] In an embodiment, the word-level encoding vector can be decoded by an attention mechanism to obtain a word-level recognition result, so as to further improve the accuracy of sign language recognition. Further, a multi-head attention mechanism can be used to strengthen the feature extraction capability, so as to further improve the accuracy of sign language recognition. For example, the decoder for decoding the word-level encoding vector can be composed of multiple transformer layers.
[0101] The sign language recognition method provided by the embodiment of the present application not only can recognize a sentence-level recognition result, but also can recognize a word-level recognition result, thereby improving the richness of sign language recognition.
[0102] Based on any of the above embodiments, Figure 4 As shown in FIG. 3, the sentence-level recognition result is recognized based on the following steps: Figure 4
[0103] Step 410: word-level encoding is performed on each of the pose graphs to obtain a word-level encoding vector of each of the pose graphs.
[0104] Specifically, the word-level encoding is performed on each of the pose graphs based on the isolated word encoder to obtain a word-level encoding vector of each of the pose graphs, that is, to obtain a high-level semantic feature. It can be understood that one word-level encoding vector is encoded based on multiple frames of pose graphs, that is, the number of word-level encoding vectors is less than or equal to the number of pose graphs.
[0105] In an embodiment, the isolated word encoder is composed of a 3D convolution block, so that multiple frames of images can be encoded. In another embodiment, the key pose graph corresponding to the current isolated word can be first determined from the multiple frames of pose graphs, so that the key pose graph is encoded to obtain the word-level encoding vector, and the key pose graph is the pose graph that best represents the current isolated word.
[0106] Step 420: sentence-level encoding is performed on each of the word-level encoding vectors based on context information of each of the word-level encoding vectors to obtain a sentence-level encoding vector.
[0107] Specifically, each word-level encoding vector is input into a sentence encoder to obtain a sentence-level encoding vector output by the sentence encoder. The sentence encoder is configured to extract context information of each word-level encoding vector, and encode each word-level encoding vector based on the context information. The specific structure of the sentence encoder can be set according to actual needs, for example, the sentence encoder can be composed of multiple transformer layers.
[0108] At step 430, the sentence-level encoding vector is decoded to obtain the sentence-level recognition result.
[0109] Specifically, the sentence-level encoding vector is decoded based on a decoder to obtain the sentence-level recognition result. The specific structure of the decoder can be set according to actual needs, for example, the decoder can be composed of multiple transformer layers.
[0110] In an embodiment, the sentence-level encoding vector can be decoded by an attention mechanism to obtain the sentence-level recognition result, thereby further improving the accuracy of sign language recognition. Further, the sentence-level encoding vector can be decoded by a multi-head attention mechanism to obtain the sentence-level recognition result, thereby further improving the accuracy of sign language recognition.
[0111] The sign language recognition method provided by the embodiment of the present application first encodes each posture graph at a word level to obtain a word-level encoding vector of each posture graph, and then encodes each word-level encoding vector at a sentence level based on context information of each word-level encoding vector to obtain a sentence-level encoding vector, that is, a word-level encoding vector at a higher level is first extracted, and then the word-level encoding vector is encoded to obtain a sentence-level encoding vector at a next level, thereby improving the representation ability of the sentence-level encoding vector through layer-by-layer encoding, and further improving the accuracy of sign language recognition.
[0112] Based on any of the above embodiments, in the method, considering that the word order of sign language can be different from that of normal listeners, based on this, the step 430 includes:
[0113] If the current decoding round is not the first decoding round, the sentence-level encoding vector is decoded based on the word recognition result of the previous decoding round to obtain the word recognition result of the current decoding round;
[0114] The previous decoding round is updated to the current decoding round until the current decoding round is the last decoding round;
[0115] Based on the word recognition result of each decoding round, the sentence-level recognition result is determined;
[0116] If the current decoding round is the first decoding round, the sentence-level encoding vector is decoded to obtain a word recognition result of the current decoding round;
[0117] The current decoding round is taken as the previous decoding round.
[0118] It should be noted that the decoding rounds are performed in sequence, so that when the decoding of the current decoding round is performed, the word recognition result of the previous decoding round is also referred to, to ensure that the word recognition result of the previous decoding round is associated with the word recognition result of the current decoding round, to avoid difficulty in understanding due to the word order problem, to finally obtain a sentence-level recognition result corresponding to a normal word order, and to further improve the accuracy of sign language recognition.
[0119] For ease of understanding, it is assumed that the to-be-recognized video includes a 'friend' sign language video, a 'today' sign language video, a 'evening' sign language video, a 'go home' sign language video, and a 'eat' sign language video. If the word recognition result obtained by the first decoding round is 'friend', the sentence-level encoding vector is decoded based on the word recognition result of 'friend' to obtain the word recognition result of the second decoding round as 'today'; the sentence-level encoding vector is decoded based on the word recognition result of 'today' to obtain the word recognition result of the third decoding round as 'evening'; the sentence-level encoding vector is decoded based on the word recognition result of 'evening' to obtain the word recognition result of the fourth decoding round as 'go home'; the sentence-level encoding vector is decoded based on the word recognition result of 'go home' to obtain the word recognition result of the fifth decoding round as 'eat'; finally, the word recognition results 'friend', 'today', 'evening', 'go home', and 'eat' are spliced in the order of the decoding rounds to obtain the sentence-level recognition result 'friend today evening go home eat?'.
[0120] The sign language recognition method provided in the embodiments of the present application ensures that the word recognition result of the previous decoding round is associated with the word recognition result of the current decoding round in the decoding process of each decoding round, avoids difficulty in understanding of the finally obtained sentence-level recognition result due to the word order problem, finally obtains a sentence-level recognition result corresponding to a normal word order, and further improves the accuracy of sign language recognition.
[0121] Based on any of the above embodiments, any frame image in the to-be-recognized video is subjected to pose key point detection based on the following steps:
[0122] If the any frame image is not the first frame image, the any frame image is subjected to aggregation processing with the previous frame pose image to obtain aggregated image data.
[0123] perform pose key point detection on the aggregated image data to obtain a pose key point map corresponding to the any frame image;
[0124] update the last pose map as a pose map corresponding to the pose key point map corresponding to the any frame image until the any frame image is a last frame image;
[0125] If the any frame image is a first frame image, perform pose key point detection on the any frame image to obtain a pose key point map corresponding to the any frame image.
[0126] Here, the aggregated image data is obtained by concatenating the any frame image and the last pose map in the channel dimension. If the any frame image is an RGB image, the aggregated image data is obtained by concatenating the any frame image and the last pose map in the three channel dimensions of RGB.
[0127] Specifically, if the any frame image is not the first frame image, perform pose key point detection on the aggregated image data based on the pose key point detection model to obtain the pose key point map corresponding to the any frame image. If the any frame image is the first frame image, perform pose key point detection on the any frame image based on the pose key point detection model to obtain the pose key point map corresponding to the any frame image.
[0128] It should be noted that each frame image is sequentially subjected to pose key point detection, so that when the pose key point detection is performed on the current image frame, the pose map corresponding to the pose key point map of the last frame is also introduced to guide the pose key point detection of the current image frame, thereby improving the accuracy of the pose key point detection.
[0129] The sign language recognition method provided by the embodiment of the present application introduces the pose map corresponding to the pose key point map of the last frame in the pose key point detection process of each frame image to guide the pose key point detection of the current image frame, thereby improving the accuracy of the pose key point detection and further improving the accuracy of the sign language recognition.
[0130] In actual application, based on the above embodiments, the redundancy information in the video to be recognized is greatly reduced, and factors affecting sign language recognition such as crowd occlusion and background interference can be effectively dealt with, and the sign language recognition model can be applied to complex gesture recognition, thereby improving the accuracy of sign language recognition and improving the robustness of the sign language recognition model.
[0131] The sign language recognition device provided by the present application is described below, and the sign language recognition device described below can be correspondingly referred to the sign language recognition method described above.
[0132] Figure 5 The structure diagram of the sign language recognition device provided by the present application is as follows,Figure 5 The sign language recognition device shown in the embodiment comprises:
[0133] The detection module 510 is configured to perform pose key point detection on each frame image in the video to be recognized to obtain a pose key point map corresponding to each frame image, and any pose key point map comprises a plurality of pose key points.
[0134] The classification module 520 is configured to perform classification labeling on the plurality of pose key points of each pose key point map to obtain a pose map corresponding to each pose key point map.
[0135] The recognition module 530 is configured to perform sign language recognition on each pose map to obtain a sign language recognition result of the video to be recognized.
[0136] The sign language recognition device provided in the embodiment performs pose key point detection on each frame image in the video to be recognized to obtain a pose key point map corresponding to each frame image, thereby removing redundant information in the video to be recognized and extracting pose information related to sign language actions, so as to improve the accuracy of sign language recognition; the plurality of pose key points of each pose key point map are classified and labeled to obtain a pose map corresponding to each pose key point map, so as to distinguish each pose key point from each other, explicitly group the pose key points, and explicitly construct the grouped pose key points, so that when each pose map is subjected to sign language recognition, the association between each part of the human body represented by the grouped pose key points and the action characteristics of the sign language word can be recognized, thereby further improving the accuracy of sign language recognition.
[0137] According to any one of the above embodiments, the classification module 520 comprises:
[0138] The key point classification unit is configured to classify the plurality of pose key points of each pose key point map based on a preset key point classification rule to obtain a key point classification result corresponding to each pose key point map, and the preset key point classification rule is used to represent pose key points included in each limb of the human body, and the limbs are limbs related to sign language actions.
[0139] The key point connection unit is configured to connect the plurality of pose key points of each pose key point map based on each key point classification result to obtain a connection map corresponding to each pose key point map.
[0140] The limb labeling unit is configured to label a limb connection part of each connection map based on each key point classification result to obtain a pose map corresponding to each connection map.
[0141] According to any one of the above embodiments, the limb labeling unit is further configured to:
[0142] The limb trunk-color mapping relationship is used to represent the mapping relationship between the limbs and different colors.
[0143] The limb trunk-color mapping relationship is used to represent the mapping relationship between the limbs and different colors.
[0144] According to any one of the above embodiments, the sign language recognition result includes a sentence-level recognition result and / or a word-level recognition result.
[0145] The sentence-level recognition result is obtained by decoding a sentence-level encoding vector, and the sentence-level encoding vector is determined based on context information of the gesture graphs.
[0146] The word-level recognition result is obtained by decoding a word-level encoding vector of each gesture graph.
[0147] According to any one of the above embodiments, the recognition module 530 includes:
[0148] a word encoding unit configured to perform word-level encoding on the gesture graphs to obtain word-level encoding vectors of the gesture graphs;
[0149] a sentence encoding unit configured to perform sentence-level encoding on the word-level encoding vectors based on context information of the word-level encoding vectors to obtain a sentence-level encoding vector;
[0150] a vector decoding unit configured to decode the sentence-level encoding vector to obtain the sentence-level recognition result.
[0151] According to any one of the above embodiments, the vector decoding unit is further configured to:
[0152] if the current decoding turn is not the first decoding turn, decode the sentence-level encoding vector based on a word recognition result of a previous decoding turn to obtain a word recognition result of the current decoding turn;
[0153] update the previous decoding turn to the current decoding turn until the current decoding turn is the last decoding turn;
[0154] determine the sentence-level recognition result based on the word recognition results of the decoding turns;
[0155] if the current decoding turn is the first decoding turn, decode the sentence-level encoding vector to obtain a word recognition result of the current decoding turn;
[0156] take the current decoding turn as the previous decoding turn.
[0157] Based on any of the above embodiments, the detection module 510 comprises:
[0158] an image aggregation unit configured to, if the any frame image is not a first frame image, aggregate the any frame image with a previous frame pose map to obtain aggregated image data;
[0159] an image detection unit configured to perform pose key point detection on the aggregated image data to obtain a pose key point map corresponding to the any frame image;
[0160] an image updating unit configured to update the previous frame pose map to a pose map corresponding to a pose key point map corresponding to the any frame image until the any frame image is a last frame image;
[0161] The image detection unit is further configured to, if the any frame image is a first frame image, perform pose key point detection on the any frame image to obtain a pose key point map corresponding to the any frame image.
[0162] Figure 6 An example of a schematic diagram of a physical structure of an electronic device is shown in FIG. 6. Figure 6 As shown in FIG. 6, the electronic device can include a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 can communicate with each other through the communications bus 640. The processor 610 can invoke a logical instruction in the memory 630 to execute a sign language recognition method, which includes: performing pose key point detection on each frame image in a video to be recognized to obtain a pose key point map corresponding to each frame image, wherein any pose key point map includes a plurality of pose key points; classifying and labeling the plurality of pose key points of each pose key point map to obtain a pose map corresponding to each pose key point map; and performing sign language recognition on each pose map to obtain a sign language recognition result of the video to be recognized.
[0163] In addition, the logic instructions in the memory 630 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0164] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to perform the sign language recognition method provided by the above-mentioned methods. The method comprises: performing pose key point detection on each frame image in a to-be-recognized video to obtain a pose key point map corresponding to each frame image, any pose key point map comprising a plurality of pose key points; performing classification labeling on the plurality of pose key points of each pose key point map to obtain a pose map corresponding to each pose key point map; and performing sign language recognition on each pose map to obtain a sign language recognition result of the to-be-recognized video.
[0165] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to perform the sign language recognition method provided by the above-mentioned methods. The method comprises: performing pose key point detection on each frame image in a to-be-recognized video to obtain a pose key point map corresponding to each frame image, any pose key point map comprising a plurality of pose key points; performing classification labeling on the plurality of pose key points of each pose key point map to obtain a pose map corresponding to each pose key point map; and performing sign language recognition on each pose map to obtain a sign language recognition result of the to-be-recognized video.
[0166] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components explicitly shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0167] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0168] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for sign language recognition, characterized in that, The method comprises the following steps: performing posture key point detection on each frame image in the to-be-identified video to obtain a posture key point map corresponding to each frame image, wherein any posture key point map comprises a plurality of posture key points; performing classification labeling on the plurality of posture key points of each posture key point map to obtain a posture map corresponding to each posture key point map; performing sign language recognition on each posture map to obtain a sign language recognition result of the to-be-identified video; the step of performing classification labeling on the plurality of posture key points of each posture key point map to obtain a posture map corresponding to each posture key point map comprises: performing classification on the plurality of posture key points of each posture key point map based on a preset key point classification rule to obtain a key point classification result corresponding to each posture key point map, wherein the preset key point classification rule is used to represent posture key points included in each limb of a human body, and the limbs are limbs related to sign language actions; connecting the plurality of posture key points of each posture key point map based on each key point classification result to obtain a connection map corresponding to each posture key point map; labeling a limb connection part of each connection map based on each key point classification result to obtain a posture map corresponding to each connection map; any connection map is labeled based on the following manner: determining a labeling manner of each limb connection part based on a key point classification result corresponding to the connection map; respectively labeling each limb connection part of the connection map based on each labeling manner to obtain a posture map corresponding to the connection map.
2. The sign language recognition method according to claim 1, characterized in that, the step of labeling a limb connection part of each connection map based on each key point classification result to obtain a posture map corresponding to each connection map comprises: color labeling a limb connection part of each connection map based on a limb-color mapping relationship and each key point classification result to obtain a posture map corresponding to each connection map; wherein the limb-color mapping relationship is used to represent a mapping relationship between the limbs and different colors.
3. The sign language recognition method of claim 1, wherein, the sign language recognition result comprises a sentence-level recognition result and / or a word-level recognition result; the sentence-level recognition result is obtained by decoding a sentence-level encoding vector, and the sentence-level encoding vector is determined based on context information of each posture map; the word-level recognition result is obtained by decoding a word-level encoding vector of each posture map.
4. The sign language recognition method according to claim 3, characterized in that, the sentence-level recognition result is obtained based on the following steps: performing word-level encoding on each posture map to obtain a word-level encoding vector of each posture map; performing sentence-level encoding on each word-level encoding vector based on context information of each word-level encoding vector to obtain a sentence-level encoding vector; decoding the sentence-level encoding vector to obtain the sentence-level recognition result.
5. The sign language recognition method according to claim 4, characterized in that, the step of decoding the sentence-level encoding vector to obtain the sentence-level recognition result comprises: if a current decoding round is not a first decoding round, decoding the sentence-level encoding vector based on a word recognition result of a previous decoding round to obtain a word recognition result of the current decoding round; updating the previous decoding round as the current decoding round until the current decoding round is the last decoding round; determining the sentence-level recognition result based on the word recognition result of each decoding round; if the current decoding round is the first decoding round, decoding the sentence-level encoding vector to obtain the word recognition result of the current decoding round; updating the current decoding round as the previous decoding round.
6. The sign language recognition method of claim 1, wherein, any frame image in the to-be-recognized video is subjected to gesture key point detection based on the following steps: if the any frame image is not the first frame image, performing aggregation processing on the any frame image and the previous frame gesture image to obtain aggregated image data; performing gesture key point detection on the aggregated image data to obtain a gesture key point image corresponding to the any frame image; updating the previous frame gesture image as a gesture image corresponding to the gesture key point image corresponding to the any frame image until the any frame image is the last frame image; if the any frame image is the first frame image, performing gesture key point detection on the any frame image to obtain a gesture key point image corresponding to the any frame image.
7. A sign language recognition apparatus characterized by comprising: comprise: a detection module configured to perform gesture key point detection on each frame image in a to-be-recognized video to obtain a gesture key point image corresponding to the frame image, and any gesture key point image comprising a plurality of gesture key points; a classification module configured to classify the plurality of gesture key points of each gesture key point image to obtain a gesture image corresponding to the gesture key point image; an identification module configured to perform sign language recognition on each gesture image to obtain a sign language recognition result of the to-be-recognized video; the classification module comprises: a key point classification unit configured to classify the plurality of gesture key points of each gesture key point image based on a preset key point classification rule to obtain a key point classification result corresponding to the gesture key point image, the preset key point classification rule being used to represent gesture key points included in each limb of a human body, the limbs being limbs related to sign language actions; a key point connection unit configured to connect the plurality of gesture key points of each gesture key point image based on each key point classification result to obtain a connection graph corresponding to the gesture key point image; a limb marking unit configured to mark a limb connection part of each connection graph based on each key point classification result to obtain a gesture image corresponding to the connection graph; any connection graph is marked based on the following manner: determining a marking manner of each limb connection part based on the key point classification result corresponding to the connection graph; respectively marking each limb connection part of the connection graph based on each marking manner to obtain a gesture image corresponding to the connection graph.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the sign language recognition method of any one of claims 1-6. 9.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the sign language recognition method of any one of claims 1-6.
Citation Information
Patent Citations
Posture recognition method and device, electronic equipment and computer readable medium
CN114495173A
Real-time video stream sign language recognition method based on human body key points
CN115457654A