Sign language recognition method and device, electronic equipment and readable storage medium

By extracting palm images and key points of the human skeleton from sign language images, and combining hand shape encoding models and attention mechanisms, the problem of inaccurate sign language features in existing sign language recognition schemes is solved, achieving higher recognition accuracy.

CN117115918BActive Publication Date: 2026-03-20VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing pure RGB schemes and skeletal keypoint schemes are greatly affected by lighting conditions and background noise in sign language recognition, resulting in inaccurate sign language feature extraction and low accuracy of sign language recognition results.

Method used

By extracting palm images and human skeletal key points from sign language images, the hand shape feature information and position feature information of the palm are determined respectively, and the sign language information is determined based on these feature information. The sign language recognition is then performed by combining the trained hand shape encoding model and attention mechanism.

Benefits of technology

It improves the accuracy of sign language feature extraction and enhances the accuracy of sign language recognition results, and can eliminate the differences in hand shape and position between left-handed and right-handed sign language without increasing the amount of computation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117115918B_ABST
    Figure CN117115918B_ABST
Patent Text Reader

Abstract

The application discloses a sign language recognition method and device, an electronic device and a readable storage medium, and belongs to the technical field of sign language recognition. The sign language recognition method comprises the following steps: acquiring a sign language image; extracting a palm image and human skeleton key points in the sign language image; determining hand shape feature information of the palm according to the palm image; determining position feature information of the palm according to the human skeleton key points; and determining sign language information corresponding to the sign language image according to the hand shape feature information and the position feature information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of sign language recognition, and particularly relates to a sign language recognition method and device, an electronic device and a readable storage medium. BACKGROUND

[0002] At present, sign language recognition schemes mainly include a pure RGB scheme and a skeleton key point scheme. However, the pure RGB scheme is greatly affected by light conditions and background noise, has a high requirement for signal bandwidth, and needs many training resources, and the training cost is very high. The skeleton key point scheme has the problem of inaccurate or missing hand skeleton joints, thereby causing the loss of sign language information captured by the model. Based on this, when sign language recognition is performed by using the above two sign language recognition schemes, the extracted sign language features are not accurate enough, thereby reducing the accuracy of the sign language recognition result. SUMMARY

[0003] Embodiments of the application provide a sign language recognition method and device, an electronic device and a readable storage medium, which can improve the accuracy of sign language feature extraction and sign language recognition result.

[0004] In a first aspect, the embodiments of the application provide a sign language recognition method, which includes: acquiring a sign language image; extracting a palm image and a human skeleton key point in the sign language image; determining hand shape feature information of the palm according to the palm image; determining position feature information of the palm according to the human skeleton key point; and determining sign language information corresponding to the sign language image according to the hand shape feature information and the position feature information.

[0005] In a second aspect, the embodiments of the application provide a sign language recognition device, which includes: an acquisition unit configured to acquire a sign language image; a processing unit configured to extract a palm image and a human skeleton key point in the sign language image; the processing unit is further configured to determine hand shape feature information of the palm according to the palm image; the processing unit is further configured to determine position feature information of the palm according to the human skeleton key point; and the processing unit is further configured to determine sign language information corresponding to the sign language image according to the hand shape feature information and the position feature information.

[0006] In a third aspect, the embodiments of the application provide an electronic device, which includes a processor and a memory. The memory stores programs or instructions that can be run on the processor. When the programs or instructions are executed by the processor, the steps of the sign language recognition method of the first aspect are implemented.

[0007] In a fourth aspect, the embodiments of the application provide a readable storage medium, which stores programs or instructions. When the programs or instructions are executed by the processor, the steps of the sign language recognition method of the first aspect are implemented.

[0008] In a fifth aspect, an embodiment of the present application provides a chip, which comprises a processor and a communication interface, the communication interface and the processor being coupled, and the processor being configured to run programs or instructions to implement the steps of the sign language recognition method according to the first aspect.

[0009] In a sixth aspect, an embodiment of the present application provides a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the sign language recognition method according to the first aspect.

[0010] In the sign language recognition method provided by the embodiment of the present application, in the process of sign language recognition, the palm image and the human body skeleton key points are extracted respectively, then the hand shape feature information of the palm is determined based on the palm image, the position feature information of the palm is determined based on the human body skeleton key points, and the sign language information is determined based on the hand shape feature information and the position feature information. In this way, in the process of sign language recognition, not only the hand region feature, i.e., the palm hand shape feature, can be accurately captured based on the palm image, but also the body posture feature can be accurately captured based on the human body skeleton key points, and the accurate palm position feature is obtained based on the body posture feature, which improves the accuracy of extracting the sign language feature, and thus improves the accuracy of the sign language recognition result. BRIEF DESCRIPTION OF DRAWINGS

[0011] Figure 1 A flowchart of the sign language recognition method provided by the embodiment of the present application is shown in the figure;

[0012] Figure 2 One of the principle diagrams of the sign language recognition method provided by the embodiment of the present application is shown in the figure;

[0013] Figure 3 Another of the principle diagrams of the sign language recognition method provided by the embodiment of the present application is shown in the figure;

[0014] Figure 4 A third of the principle diagrams of the sign language recognition method provided by the embodiment of the present application is shown in the figure;

[0015] Figure 5 A fourth of the principle diagrams of the sign language recognition method provided by the embodiment of the present application is shown in the figure;

[0016] Figure 6 A fifth of the principle diagrams of the sign language recognition method provided by the embodiment of the present application is shown in the figure;

[0017] Figure 7 A sixth of the principle diagrams of the sign language recognition method provided by the embodiment of the present application is shown in the figure;

[0018] Figure 8 A structural block diagram of the sign language recognition device provided by the embodiment of the present application is shown in the figure;

[0019] Figure 9A structural block diagram of an electronic device provided by an embodiment of the present application is shown in the figure.

[0020] Figure 10 A hardware structure schematic diagram of an electronic device provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.

[0022] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than that illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category, not limited to the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.

[0023] The sign language recognition method provided by the embodiments of the present application will be described in detail below in combination with the drawings, specific embodiments and application scenarios.

[0024] As shown in the figure, the present application provides a sign language recognition method, which can include the following S102 to S110: Figure 1

[0025] S102: Obtain a sign language image.

[0026] The sign language recognition method proposed in the embodiments of the present application is executed by an electronic device, which can be specifically a smart phone, a tablet computer, a notebook computer, a smart watch, and other smart electronic devices, which are not specifically limited here.

[0027] Specifically, the sign language image can be a video frame extracted from a sign language video.

[0028] Further, the sign language image includes a character image of a person signing a sign language, and the character in the character image has independent hands or the character in the character image has overlapping hands, which are not specifically limited here.

[0029] S104: Extract a palm image and a human skeleton key point in the sign language image.

[0030] ​The one frame of sign language image includes two palm images, and the two palm images correspond to left and right hands of the character image in the sign language image respectively.

[0031] Specifically, in a case where the left and right hands of the character image in the sign language image are independent of each other, the palm images of the left and right hands of the character image in the sign language image are extracted by the target detection model. In a case where the left and right hands of the character image in the sign language image overlap each other, that is, in a case where the sign language image includes crossed hands, the crossed hand image of the character image in the sign language image is extracted by the target detection model, and the crossed hand image is taken as the palm image of the left and right hands of the character image in the sign language image respectively.

[0032] In actual application, the target detection model can be a YoLo model, and the specific type of the target detection model can be selected by a person skilled in the art according to actual conditions, which is not limited here.

[0033] Further, the human body skeleton key points can be posture key points of the character image in the sign language image, so as to reduce the response time and improve the feature extraction efficiency. Specifically, the posture key points of the character image in the sign language image, such as wrist joint points, eyes, nose, mouth, ears, shoulders and elbow joint points, are extracted by a human body skeleton key point detection algorithm.

[0034] In actual application, the human body skeleton key point detection algorithm can be a mediapipe algorithm, and the specific type of the human body skeleton key point detection algorithm can be selected by a person skilled in the art according to actual conditions, which is not limited here.

[0035] S106: determining hand shape feature information of the palm according to the palm image.

[0036] The hand shape feature information is used to describe the palm posture.

[0037] Specifically, in the sign language recognition method provided in the embodiment of the present application, after the palm images of the left and right hands of the character image in the sign language image are extracted, the hand shape feature information of the left hand palm of the character image is obtained by analyzing and processing the palm image of the left hand of the character image, and the hand shape feature information of the right hand palm of the character image is obtained by analyzing and processing the palm image of the right hand of the character image.

[0038] S108: determining position feature information of the palm according to the human body skeleton key points.

[0039] The position feature information of the palm is relative position information of the palm relative to other human body skeleton key points.

[0040] Specifically, the human body skeleton key points include wrist joint points, eyes, nose, mouth, ears, shoulders, and elbow joint points, etc. On this basis, after the human body skeleton key points of the character image in the sign language image are extracted, the position feature information of the left hand palm can be determined by the relative position information between the key points of eyes, nose, mouth, ears, shoulders, and elbow joint points relative to the left hand wrist joint point; and the position feature information of the right hand palm can be determined by the relative position information between the key points of eyes, nose, mouth, ears, shoulders, and elbow joint points relative to the right hand wrist joint point.

[0041] S110: determining the sign language information corresponding to the sign language image according to the hand shape feature information and the position feature information.

[0042] Specifically, in the sign language recognition method provided in the embodiment of the present application, after the hand shape feature information and the position feature information of the left and right hand palms of the character image in the sign language image are obtained, the hand shape feature information and the position feature information of the left hand palm of the character image are fused to obtain a first palm feature, and the hand shape feature information and the position feature information of the right hand palm of the character image are fused to obtain a second palm feature. On this basis, the sign language feature information is obtained by analyzing and fusing the first palm feature and the second palm feature, and then the corresponding sign language information is obtained by classifying and recognizing the sign language feature information.

[0043] The above-mentioned sign language recognition method provided in the embodiment of the present application extracts the palm image and the human body skeleton key points respectively in the process of sign language recognition, then determines the hand shape feature information of the palm based on the palm image, determines the position feature information of the palm based on the human body skeleton key points, and determines the sign language information based on the hand shape feature information and the position feature information. In this way, in the process of sign language recognition, not only the hand region feature, i.e. the palm hand shape feature, can be accurately captured based on the palm image, but also the body posture feature can be accurately captured based on the human body skeleton key points, and the accurate palm position feature is obtained based on the body posture feature, which improves the accuracy of extracting the sign language feature, thereby improving the accuracy of the sign language recognition result.

[0044] In the embodiment of the present application, the palm image includes a first image and a second image, the first image corresponds to a first palm, the second image corresponds to a second palm, the hand shape feature information includes hand shape feature information of the first palm and hand shape feature information of the second palm, and on this basis, the above-mentioned S106 can specifically include the following S106a to S106c:

[0045] S106a: performing symmetry processing on the first image to obtain a third image.

[0046] The first image corresponds to the first palm, which can be the left hand palm or the right hand palm of the character image in the sign language image, and is not specifically limited here.

[0047] Further, the first image can only include the image of the first palm of the character image in the sign language image, or the first image can include the image of the crossed hands of the character image in the sign language image, and is not specifically limited here.

[0048] Specifically, in the sign language recognition method provided in the embodiment of the present application, after the palm image of the left and right hands of the character image in the sign language image is extracted, the palm image corresponding to the first palm, i.e., the first image, is symmetrically processed to obtain a third image. In this way, after the hand shape feature is extracted based on the third image, the hand shape features extracted for left-handed sign language and right-handed sign language can be unified with each other, thereby eliminating the difference in hand shape between left-handed sign language and right-handed sign language, and ensuring the accuracy of hand shape feature extraction without increasing the amount of hand shape feature extraction.

[0049] The left-handed sign language is a sign language mode in which the left hand is mainly used and the right hand is used as an auxiliary, and the right-handed sign language is a sign language mode in which the right hand is mainly used and the left hand is used as an auxiliary. For the same sign language, when the sign language actions of the left and right hands are inconsistent, the palm hand shapes extracted based on the left-handed sign language and the right-handed sign language are in opposite states.

[0050] S106b: encoding the third image based on the trained hand shape coding model to obtain hand shape feature information of the first palm.

[0051] Specifically, in the sign language recognition method provided in the embodiment of the present application, after the palm image corresponding to the first palm, i.e., the first image, is symmetrically processed to obtain a third image, the third image is encoded based on the trained hand shape coding model to obtain hand shape feature information of the first palm.

[0052] The hand shape coding model can be a BackBone model, such as a ResNet model and an EfficientNet model, and the specific type of the hand shape coding model can be selected by a person skilled in the art according to actual conditions, and is not specifically limited here.

[0053] S106c: encoding the second image based on the trained hand shape coding model to obtain hand shape feature information of the second palm.

[0054] The second image corresponds to the second palm, which can be the left hand palm or the right hand palm of the character image in the sign language image, and is not specifically limited here.

[0055] Further, the second image can only include a single image of the second palm of the character image in the sign language image, or the second image can include an image of crossed hands of the character image in the sign language image, which is not specifically limited here.

[0056] Specifically, in the sign language recognition method provided in the embodiments of the present application, after the palm image of the left hand and the right hand of the character image in the sign language image is extracted, the second image corresponding to the second palm is encoded based on the trained hand shape coding model to obtain the hand shape feature information of the second palm.

[0057] In this way, in the process of determining the hand shape feature information of the palm according to the palm image, for the palm image of the first palm, the hand shape feature information of the first palm is determined based on the trained hand shape coding model after the symmetric processing of the palm image, and for the palm image of the second palm, the hand shape feature information of the second palm is directly determined based on the trained hand shape coding model. In this way, the difference in hand shape between left-handed sign language and right-handed sign language can be eliminated, and the hand shape feature information of the first palm and the second palm obtained can be unified with each other regardless of left-handed sign language or right-handed sign language, so that the accuracy of hand shape feature extraction can be ensured without increasing the amount of hand shape feature extraction and the amount of model calculation.

[0058] In the above embodiments provided by the present application, the palm image includes a first image and a second image, the first image corresponds to a first palm, the second image corresponds to a second palm, the hand shape feature information includes hand shape feature information of the first palm and hand shape feature information of the second palm, in the process of determining the hand shape feature information of the palm according to the palm image, the first image is symmetrically processed to obtain a third image; the third image is encoded based on the trained hand shape coding model to obtain the hand shape feature information of the first palm; and the second image is encoded based on the trained hand shape coding model to obtain the hand shape feature information of the second palm. In this way, the difference in hand shape between left-handed sign language and right-handed sign language can be eliminated, and the hand shape feature information of the first palm and the second palm obtained can be unified with each other regardless of left-handed sign language or right-handed sign language, so that the accuracy of hand shape feature extraction can be ensured without increasing the amount of hand shape feature extraction and the amount of model calculation.

[0059] In the embodiments of the present application, before the above S106b, the above sign language recognition method can further include the following S112 to S120:

[0060] S112: Obtain a sign language sample image.

[0061] Specifically, the sign language sample image can be a video frame extracted from a sign language video.

[0062] Further, the sign language sample image includes a person image in which a person is signing a sign language, and the hands of the person in the person image are in a state of being independent of each other or in a state of being overlapped with each other, and the present application is not limited in this regard.

[0063] S114: In a case where the palms in the sign language sample image are overlapped with each other, the sign language sample image is encoded based on the hand shape coding model, and the hand shape coding model is iteratively trained according to the encoding result of the sign language sample image and the first loss function.

[0064] Specifically, in the sign language recognition method provided in the present application, before the third image is encoded based on the hand shape coding model, a sign language sample image is obtained, and the hand shape coding model is trained according to the sign language sample image. Specifically, in a case where the palms in the sign language sample image are overlapped with each other, the sign language sample image is encoded based on the hand shape coding model, and then the encoding result of the sign language sample image is classified and learned based on the first loss function, so as to train the hand shape coding model, thereby obtaining the trained hand shape coding model.

[0065] As shown in the formula (1), the hand shape coding model includes a hand shape encoder, and the first loss function can be a cross-entropy loss function. Based on a category fully connected layer, the hand shape encoder can accurately learn the hand shape code of the crossed hand shape through the cross-entropy loss function. Figure 2

[0066] S116: In a case where the palms in the sign language sample image are independent of each other, a first recognition frame is determined according to the palm skeleton key points in the sign language sample image, and a second recognition frame is determined according to the palm image in the sign language sample image.

[0067] Specifically, the palm skeleton key points can include wrist joint points and finger joint points.

[0068] Further, the first recognition frame is a minimum frame that covers the palm skeleton key points of a single palm.

[0069] Further, the second recognition frame is a recognition frame of the palm image of a single palm when the palm image in the sign language sample image is extracted by a target detection model such as a YoLo model.

[0070] Specifically, in the sign language recognition method provided in the present application, after the sign language sample image is obtained, in a case where the palms in the sign language sample image are independent of each other, for each single palm in the sign language sample image, a minimum frame that covers the palm skeleton key points of each single palm, i.e., a first recognition frame, is determined, and a second recognition frame of the palm image of the single palm when the palm image of the single palm is extracted is determined.

[0071] ​S118: Determine an intersection over union of the first bounding box and the second bounding box according to size information of the first bounding box and the second bounding box.

[0072] The size information can be area information of the first bounding box and the second bounding box.

[0073] Further, the intersection over union of the first bounding box and the second bounding box can be a ratio value of an intersection area of the first bounding box and the second bounding box to a union area of the first bounding box and the second bounding box.

[0074] Specifically, in the sign language recognition method provided in the embodiments of the present application, in the case that the palms in the sign language sample image are independent of each other, for each single palm in the sign language sample image, after obtaining the first bounding box and the second bounding box of the palm, the intersection over union of the first bounding box and the second bounding box is determined according to area information of the first bounding box and the second bounding box.

[0075] In actual application, the intersection over union of the first bounding box and the second bounding box can be calculated by the following formula (1):

[0076] IOU = (S (B1∩B2) / S (B1 ∪ B2) ), (1)

[0077] wherein, IOU represents the intersection over union of the first bounding box and the second bounding box, B1 represents the first bounding box, B2 represents the second bounding box, B1∩B2 represents the intersection of the first bounding box and the second bounding box, B1∪B2 represents the union of the first bounding box and the second bounding box, S (B1∩B2) ∩ represents the intersection area of the first bounding box and the second bounding box, S (B1 ∪ B2) represents the union area of the first bounding box and the second bounding box.

[0078] S120: In the case that the intersection over union is greater than a first threshold value, iteratively train the hand shape coding model according to the coordinate information of the palm skeleton key points in the sign language sample image and the second loss function, so as to obtain the trained hand shape coding model.

[0079] The first threshold value can be 0.6, 0.7, 0.8, etc., and the specific value of the first threshold value can be set by those skilled in the art according to actual conditions, which is not limited herein.

[0080] Specifically, in the sign language recognition method provided in this application embodiment, when the palms in the sign language sample images are independent of each other, for each individual palm in the sign language sample image, after determining the intersection-union ratio (IU) of the first and second recognition boxes corresponding to that palm, if the IU is greater than a first threshold, the key points of the palm's skeletal structure are used as supervision signals to perform regression fitting on the coordinate information of the key points of the palm's skeletal structure. Based on this, the fitting results of the coordinate information of the key points of the palm's skeletal structure are trained using a second loss function to train the hand shape encoding model, thereby obtaining the trained hand shape encoding model.

[0081] Among them, such as Figure 2 As shown, the above hand shape encoding model includes a hand shape encoder. The second loss function can specifically be the mean squared error loss function. Based on the fully connected layer of skeleton point categories, the mean squared error loss function helps the hand shape encoder accurately learn the hand shape encoding of a single hand.

[0082] The embodiments provided in this application, before encoding the third image based on the trained hand shape encoding model, acquire sign language sample images; when the palms in the sign language sample images overlap, encode the sign language sample images based on the hand shape encoding model, and iteratively train the hand shape encoding model according to the encoding results of the sign language sample images and a first loss function, thereby obtaining a trained hand shape encoding model; when the palms in the sign language sample images are independent, determine a first bounding box based on the key points of the hand skeleton in the sign language sample images, and determine a second bounding box based on the palm images in the sign language sample images; determine the intersection-union ratio (IUR) of the first and second bounding boxes based on their size information; when the IUR is greater than a first threshold, iteratively train the hand shape encoding model according to the coordinate information of the key points of the hand skeleton in the sign language sample images and a second loss function, thereby obtaining a trained hand shape encoding model. Thus, by training the hand shape encoding model based on sign language sample images before encoding the third image based on the trained hand shape encoding model, the accuracy of the hand shape encoding model in learning single hand shapes and crossed hand shapes is improved.

[0083] In this embodiment of the application, the key points of the human skeleton include a first key point, a second key point, and multiple third key points. The first key point corresponds to the first palm, the second key point corresponds to the second palm, and the multiple third key points correspond to the head and arms. The positional feature information includes the positional feature information of the first palm and the positional feature information of the second palm. Based on this, the above-mentioned S108 may specifically include the following S108a to S108e:

[0084] S108a: Establish a first coordinate system with the first key point as the origin, and establish a second coordinate system with the second key point as the origin.

[0085] The first key point corresponds to the first palm, and specifically can be a wrist joint point of the first palm, a palm center point of the first palm, etc., which are not limited herein.

[0086] Further, the second key point corresponds to the second palm, and specifically can be a wrist joint point of the second palm, a palm center point of the second palm, etc., which are not limited herein.

[0087] Specifically, in the sign language recognition method provided in the embodiments of the present application, after the human body skeleton key points in the sign language image are extracted, a first coordinate system is established with the first key point corresponding to the first palm as the origin, and a second coordinate system is established with the second key point corresponding to the second palm as the origin.

[0088] S108b: In the first coordinate system, first coordinate information of each third key point relative to the first key point is determined.

[0089] The plurality of third key points correspond to the head and the arm.

[0090] In actual application, the third key points specifically can include skeleton key points of the eyes, the nose, the mouth, the ears, the shoulders, and the elbows, etc., which are not limited herein.

[0091] Specifically, after the first coordinate system is established with the first key point corresponding to the first palm as the origin, in the first coordinate system, the relative coordinates of the third key points such as the skeleton key points of the eyes, the nose, the mouth, the ears, the shoulders, and the elbows, etc., relative to the first key point are calculated, to obtain the first coordinate information.

[0092] S108c: In the second coordinate system, second coordinate information of each third key point relative to the second key point is determined.

[0093] Specifically, after the second coordinate system is established with the second key point corresponding to the second palm as the origin, in the second coordinate system, the relative coordinates of the third key points such as the skeleton key points of the eyes, the nose, the mouth, the ears, the shoulders, and the elbows, etc., relative to the second key point are calculated, to obtain the second coordinate information.

[0094] S108d: The first coordinate information is symmetrically processed to obtain third coordinate information, and the position feature information of the first palm is determined according to the third coordinate information.

[0095] Specifically, after obtaining the first coordinate information by calculating the relative coordinates of each third key point relative to the first key point in the first coordinate system, the first coordinate information is symmetrically processed to obtain third coordinate information, and then the first palm position feature information is represented by the third coordinate information. In this way, the palm position difference between left-handed sign language and right-handed sign language can be eliminated, and the palm position features of left-handed sign language and right-handed sign language can be unified with each other, thereby ensuring the accuracy of the palm position feature determination.

[0096] S108e: determining the second palm position feature information according to the second coordinate information.

[0097] Specifically, after obtaining the second coordinate information by calculating the relative coordinates of each third key relative to the second key point in the second coordinate system, the second palm position feature information is directly represented by the second coordinate information.

[0098] That is, in the sign language recognition method provided in the embodiments of the present application, as shown in Figure 3 After the human body skeleton key points in the sign language image are extracted, for the first palm, the first key point P1 corresponding to the first palm, such as the wrist joint of the first palm, is taken as the origin, and the relative coordinates of the third key points, i.e., the human body skeleton key points other than the first key point P1 and the second key point P2, relative to the first key point P1 are calculated to obtain the first coordinate information, and then the first coordinate information is symmetrically processed to obtain the third coordinate information, and the first palm position feature information is represented by the third coordinate information. For the second palm, the second key point P2 corresponding to the second palm, such as the wrist joint of the second palm, is taken as the origin, and the relative coordinates of the third key points, i.e., the human body skeleton key points other than the first key point P1 and the second key point P2, relative to the second key point P2 are calculated to obtain the second coordinate information, and the second palm position feature information is directly represented by the second coordinate information. In this way, the palm position difference between left-handed sign language and right-handed sign language can be eliminated, and the palm position features of left-handed sign language and right-handed sign language can be unified with each other, thereby ensuring the accuracy of the palm position feature extraction without increasing the model calculation amount.

[0099] In the above embodiment provided by the present application, the human body skeleton key points include a first key point, a second key point and a plurality of third key points, the first key point corresponds to the first palm, the second key point corresponds to the second palm, and the plurality of third key points correspond to the head and the arm; the position feature information includes position feature information of the first palm and position feature information of the second palm; in the process of determining the position feature information of the palm according to the human body skeleton key points, a first coordinate system is established with the first key point as the origin, and a second coordinate system is established with the second key point as the origin; in the first coordinate system, first coordinate information of each third key point relative to the first key point is determined; in the second coordinate system, second coordinate information of each third key point relative to the second key point is determined; the first coordinate information is symmetrically processed to obtain third coordinate information, and the position feature information of the first palm is determined according to the third coordinate information; and the position feature information of the second palm is determined according to the second coordinate information. In this way, the palm position difference between left-handed sign language and right-handed sign language can be eliminated, and the position feature information of the first palm and the second palm obtained by the left-handed sign language and the right-handed sign language can be unified, so that the accuracy of palm position feature extraction can be ensured without increasing the model calculation amount.

[0100] In the embodiment of the present application, the palm includes a first palm and a second palm, and on this basis, S110 can specifically include S110a-S110d as follows:

[0101] S110a: determining a first palm feature according to the hand shape feature information and the position feature information of the first palm, and determining a second palm feature according to the hand shape feature information and the position feature information of the second palm.

[0102] Specifically, in the sign language recognition method provided by the embodiment of the present application, after obtaining the hand shape feature information and the position feature information of the first palm and the second palm respectively, the hand shape feature information and the position feature information of the first palm and the second palm are respectively subjected to layer normalization standardization processing. On this basis, the hand shape feature information and the position feature information of the processed first palm are spliced and fused to obtain the first palm feature, and the hand shape feature information and the position feature information of the processed second palm are spliced and fused to obtain the second palm feature.

[0103] On this basis, as shown in Figure 4 , for each frame of sign language image, a pair of first palm feature and second palm feature can be obtained, and in the subsequent process of processing the first palm feature and the second palm feature by the conversion model, the first palm feature and the second palm feature can be used as a step feature of the conversion model, that is, the first palm feature and the second palm feature can be used as an input feature of the conversion model. In this way, the number of input features input into the conversion model is twice the number of sign language image frames.

[0104] S110b: performing attention learning on the first palm feature and the second palm feature based on an attention mechanism to obtain a first attention feature and a second attention feature.

[0105] The conversion model includes a self-attention learning module.

[0106] Specifically, in the sign language recognition method provided in the embodiments of the present application, after obtaining the first palm feature and the second palm feature of each frame of sign language image, the first palm feature and the second palm feature of each frame of sign language image are input into the conversion model according to the sign language time sequence, and the self-attention learning module in the conversion model is used to perform attention learning on the first palm feature and the second palm feature of each frame of sign language image based on the attention mechanism, so as to obtain the first attention feature corresponding to the first palm and the second attention feature corresponding to the second palm.

[0107] S110c: performing convolution processing on the first attention feature and the second attention feature to obtain sign language feature information.

[0108] The conversion model further includes a convolution module.

[0109] Specifically, in the sign language recognition method provided in the embodiments of the present application, after the self-attention learning module in the conversion model is used to perform attention learning on the first palm feature and the second palm feature of each frame of sign language image to obtain the first attention feature and the second attention feature corresponding to each frame of sign language image, the first attention feature and the second attention feature of each frame of sign language image are input into the convolution module of the conversion model according to the sign language time sequence, and the convolution module is used to perform convolution processing on the first attention feature and the second attention feature of each frame of sign language image to fuse the first attention feature and the second attention feature of the same frame of sign language image, so as to obtain the sign language feature information of each frame of sign language image.

[0110] S110d: adding first classification feature information to the sign language feature information, performing attention learning on the sign language feature information to obtain second classification feature information, and determining sign language information according to the second classification feature information.

[0111] The conversion model further includes a conversion module.

[0112] Specifically, in the sign language recognition method provided in the embodiments of the present application, after obtaining the sign language feature information of each frame of sign language image, the sign language feature information of each frame of sign language image is arranged according to the sign language time sequence to obtain a sign language feature sequence, and first classification feature information is added before the sign language feature sequence. On this basis, the sign language feature sequence to which the first classification feature information is added is input into the conversion module of the conversion model, and based on the self-attention learning mechanism, the sign language feature sequence to which the first classification feature information is added is subjected to self-attention learning, so that the features of the two hands of the same frame of sign language image can be further fused to obtain a sign language feature sequence containing second classification feature information. Further, the second classification feature information in the sign language feature sequence output by the conversion module is extracted, the second classification feature information is classified by a classifier, and the corresponding sign language information is determined according to the classification result.

[0113] Exemplarily, as shown in Figure 5 , based on the multiple frames of sign language images in the sign language video, a first palm feature sequence f1 1 f2 1 f3 1 …f i 1 and a second palm feature sequence f1 2 f2 2 f3 2 …f i 2 are obtained. Wherein, f1 1 represents the first palm feature in the 1st frame of sign language image, f i 1 represents the first palm feature in the i-th frame of sign language image, f1 2 represents the second palm feature in the 1st frame of sign language image, f i 2 represents the second palm feature in the i-th frame of sign language image. On this basis, according to the sign language time sequence, the first palm feature sequence f1 1 f2 1 f3 1 …f i 1 and the second palm feature sequence f1 2 f2 2 f3 2 …f i 2 are input into the conversion model, i.e. f1 1 f1 2 f2 1 f2 2 f3 1 f3 2 …f i 1 f i 2The input conversion model, through a self-attention learning module in the conversion model, learns f1 1 f1 2 f2 1 f2 2 f3 1 f3 2 …f i 1 f i 2 Attention learning is performed to obtain attention features t1 1 t1 2 t2 1 t2 2 t3 1 t3 2 …t i 1 t i 2 , wherein t1 1 t2 1 t3 1 …t i 1 is a first attention feature, t1 2 t2 2 t3 2 …t i 2 is a second attention feature, t1 1 represents the first attention feature in the first frame of sign language image, t i 1 represents the first attention feature in the i-th frame of sign language image, t1 2 represents the second attention feature in the first frame of sign language image, t i 2 represents the second attention feature in the i-th frame of sign language image.

[0114] Further, t1 1 t1 2 t2 1 t2 2 t3 1 t3 2 …t i 1 t i 2 The convolution module of the input conversion model, through the convolution module, learns t1 1 t1 2 t2 1 t2 2 t3 1 t3 2 …t i 1 t i 2Perform convolution processing and output t1t2t3…t i Where t1 represents the sign language feature information of the first frame of the sign language image, t i This represents the sign language feature information of the i-th frame sign language image.

[0115] Furthermore, in the sign language feature sequence t1t2t3…t i Add the first classification feature information t before CLS The sign language feature sequence t is obtained. CLS t1t2t3…t i Furthermore, the sign language feature sequence t CLS t1t2t3…t i The input transformation model's transformation module, based on a self-attention learning mechanism, uses sign language feature sequences t CLS t1t2t3…t i Self-attention learning is performed to obtain the sign language feature sequence t' CLS t1t2t3…t i Furthermore, the sign language feature sequence t' is extracted. CLS t1t2t3…t i The second classification feature information t' CLS , using a classifier to classify t' CLS Classify the information and determine the corresponding sign language information based on the classification results.

[0116] It should be noted that for right-handed sign language, the hand feature sequence of the self-attention learning module in the input conversion model is f1. 1 f1 2 f2 1 f2 2 f3 1 f3 2 …f i 1 f i 2 After self-attention learning, the attention feature sequence output by the self-attention learning module is t1. 1 t1 2 t2 1 t2 2 t3 1 t3 2 …t i 1 t i 2For the left-handed hand language, since the palm image and palm position feature of the first palm are symmetrically processed in the application, the first palm feature in the left-handed hand language is the same as the second palm feature in the right-handed hand language, and the second palm feature in the left-handed hand language is the same as the first palm feature in the right-handed hand language. That is, for the left-handed hand language, the palm feature sequence of the self-attention learning module of the input conversion model is f1 2 f1 1 f2 2 f2 1 f3 2 f3 1 … i 2 f i 1 After self-attention learning, the attention feature sequence output by the self-attention learning module is t1 2 t1 1 t2 2 t2 1 t3 2 t3 1 … i 2 t i 1 On this basis, when the attention feature sequence t1 1 t1 2 t2 1 t2 2 t3 1 t3 2 … i 1 t i 2 of the right-handed hand language and the attention feature sequence t1 2 t1 1 t2 2 t2 1 t3 2 t3 1 … i 2 t i 1 of the left-handed hand language are subjected to convolution processing, since the convolution does not consider the spatial relationship, there is Y(t i 1 , t i 2 ) = Y(t i 2 , t i 1 ), that is, the hand language feature information corresponding to the right-handed hand language and the left-handed hand language is t1t2t3…t i .

[0117] That is, the convolution result Y(t i 1 , t i 2 ) of the attention feature sequence of the right-handed sign language is the same as the convolution result Y(t i 2 , t i 1 ) of the attention feature sequence of the left-handed sign language. That is, the corresponding sign language feature information of the right-handed sign language and the left-handed sign language is the same. In this way, in the subsequent process of sign language recognition based on the sign language feature information, the sign language information difference between the left-handed sign language and the right-handed sign language can be eliminated, so that the sign language information of the left-handed sign language and the right-handed sign language can be unified with each other, thereby accurately recognizing the left-handed sign language and the right-handed sign language without increasing the model calculation amount, and ensuring the accuracy of the sign language recognition result.

[0118] In the above embodiments provided by the present application, the palm includes a first palm and a second palm. In the process of determining the sign language information corresponding to the sign language image according to the hand shape feature information and the position feature information, the first palm feature is determined according to the hand shape feature information and the position feature information of the first palm, and the second palm feature is determined according to the hand shape feature information and the position feature information of the second palm. The attention learning is performed on the first palm feature and the second palm feature based on the attention mechanism, to obtain the first attention feature and the second attention feature. The convolution processing is performed on the first attention feature and the second attention feature, to obtain the sign language feature information. The attention learning is performed on the sign language feature information after adding the first classification feature information to the sign language feature information, to obtain the second classification feature information, and the sign language information is determined according to the second classification feature information. In this way, the palm feature is determined based on the hand shape feature information and the position feature information of the palm, the sign language feature information of each frame of sign language image is obtained by fusing the palm features of the left hand and the right hand, and the accuracy of the sign language recognition result is improved when the sign language recognition is performed based on the sign language feature information.

[0119] In the embodiments of the present application, the step of performing the attention learning on the first palm feature and the second palm feature can specifically include the following S122 and S124:

[0120] S122: determining the attention weight matrix according to the first mask.

[0121] In the attention weight matrix, the attention weight between the first palm feature and the second palm feature corresponding to different sign language images is zero.

[0122] Further, in the process of attention learning of the first palm feature and the second palm feature, the first palm feature and the second palm feature of each frame of sign language image are input into the conversion model as a step feature according to the sign language time sequence, and therefore, the number of steps of the conversion model is twice the number of frames of sign language images. On this basis, the first mask can be specifically a 2Tx2T matrix, where T is the number of frames of sign language images in the sign language video.

[0123] In actual application, the element value in the first mask can be calculated by the following formula (2):

[0124] mask[i, j] = (i + j) % 2, (2)

[0125] Wherein, mask[i, j] represents the element value of the i-th row and the j-th column in the first mask, i and j represent the order of inputting the step feature into the conversion model, and (i + j) % 2 represents the remainder operation on (i + j) / 2, that is, (i + j) % 2 represents the remainder value after (i + j) is divided by 2.

[0126] On this basis, since the first palm feature and the second palm feature of each frame of sign language image are input into the conversion model alternately according to the sign language time sequence, if the palm feature of the i-th input conversion model and the palm feature of the j-th input conversion model correspond to the same palm, i + j is odd, (i + j) % 2 = 1; if the palm feature of the i-th input conversion model and the palm feature of the j-th input conversion model correspond to different palms, i + j is even, (i + j) % 2 = 0. That is, the element value distribution of the first mask can be specifically as shown in the following table. Figure 6

[0127] Further, the attention weight matrix is the relevance weight matrix between each input feature of the input conversion model, i.e. the palm feature and other input features.

[0128] In actual application, the attention weight between each input feature of the input conversion model and other input features can be determined by the following formula (3):

[0129] A ij = softmax((QK T ) / ((d k ) 1 / 2 )-MASK ij ), (3)

[0130] Wherein, A ij represents the attention weight between the i-th input feature and the j-th input feature of the input conversion model, Q and K are input features, T in K T represents transposition, and (QK T ​) / ((d k ) 1 / 2 )Matrix element value is generally small, MASK ij = mask[i, j] x 10 8 .

[0131] On this basis, if the palm feature of the i-th input conversion model and the palm feature of the j-th input conversion model correspond to the same palm, mask[i, j] = 0, MASK ij = 0, at this time, A ij = softmax((QK T ) / ((d k ) 1 / 2 ));If the palm feature of the i-th input conversion model and the palm feature of the j-th input conversion model correspond to different palms, mask[i, j] = 1, MASK ij = 10 8 , at this time, (QK T ) / ((d k ) 1 / 2 )-10 8 approaches negative infinity, A ij = softmax((QK T ) / ((d k ) 1 / 2 )-10 8 ) = 0. In this way, in the process of attention learning of the first palm feature and the second palm feature, the attention weight between the first palm feature and the second palm feature is zero, so that the conversion model can only perform self-attention learning on the palm features of the same palm.

[0132] S124: According to the attention weight matrix, performing self-attention learning on the first palm feature and the second palm feature.

[0133] In the embodiments of the present application, the first palm feature and the second palm feature can be specifically self-attention learned based on the following formula (4):

[0134] A(Q, K, V) = softmax((QK T ) / ((d k ) 1 / 2 )-MASK ij )V, (4)

[0135] Wherein, Q, K and V are input features, (QK T ) / ((d k ) 1 / 2 )Matrix element value is generally small, MASK ij = mask[i, j] x 10 8 .

[0136] In the above embodiments provided by the present application, in the process of attention learning of the first palm feature and the second palm feature, the attention weight matrix is determined according to the first mask, wherein the attention weight between the first palm feature and the second palm feature corresponding to different sign language images in the attention weight matrix is zero; and the first palm feature and the second palm feature are self-attention learned according to the attention weight matrix. In this way, in the process of attention learning of the first palm feature and the second palm feature, the attention weight between the first palm feature and the second palm feature is zero, so that the conversion model can only self-attention learn the palm features of the same palm, thereby ensuring the accuracy of the first attention feature and the second attention feature.

[0137] In summary, the sign language recognition method provided by the embodiments of the present application, as shown in Figure 7 The sign language video is extracted, and for each frame of sign language image, the human body skeleton key points, the first palm image and the second palm image in the sign language image are extracted. Further, the first palm image is symmetrically processed to obtain a third palm image, the third palm image is encoded to obtain the hand shape feature information of the first palm, and the second palm image is encoded to obtain the hand shape feature information of the second palm. Further, based on the human body skeleton key points, the position feature information of the first palm and the second palm is determined, wherein the position feature information of the first palm needs to be symmetrically processed. Further, after the hand shape feature information and the position feature information of the first palm of the same frame of sign language image are standardized, the hand shape feature information and the position feature information of the first palm of the same frame of sign language image are fused to obtain the first palm feature of each frame of sign language image, and after the hand shape feature information and the position feature information of the second palm of the same frame of sign language image are standardized, the hand shape feature information and the position feature information of the second palm of the same frame of sign language image are fused to obtain the second palm feature of each frame of sign language image.

[0138] Further, the first palm feature and the second palm feature of the plurality of hand sign images in the hand sign video are alternately input into the conversion model according to the hand sign time sequence, the first palm feature and the second palm feature of the plurality of hand sign images are analyzed and recognized by the conversion model, and corresponding hand sign information is obtained. Specifically, the conversion model includes a self-attention learning module, a convolution module, and a conversion module. The self-attention learning module is used for self-attention learning of the first palm feature and the second palm feature of the plurality of hand sign images, to obtain first attention features and second attention features. Then, the convolution module is used for convolution processing of the first attention features and the second attention features, to fuse the left-hand features and the right-hand features of the same frame of hand sign image, to obtain hand sign feature information of each frame of image, and to obtain a hand sign feature sequence of the plurality of hand sign images. Further, first classification feature information is added before the hand sign feature sequence, the conversion module is used for self-attention learning of the hand sign feature sequence, to obtain a hand sign feature sequence containing second classification feature information. Further, the second classification feature information in the hand sign feature sequence is extracted, the second classification feature information is classified by a classifier, and corresponding hand sign information is determined according to a classification result.

[0139] The hand sign recognition method provided in the embodiments of the present application can be executed by a hand sign recognition device. The hand sign recognition device is taken as an example to execute the hand sign recognition method to illustrate the hand sign recognition device provided in the embodiments of the present application.

[0140] As shown in Figure 8 The hand sign recognition device 800 provided in the embodiments of the present application can include the following acquisition unit 802 and processing unit 804.

[0141] The acquisition unit 802 is configured to acquire a hand sign image.

[0142] The processing unit 804 is configured to extract a palm image and a human skeleton key point in the hand sign image.

[0143] The processing unit 804 is further configured to determine hand shape feature information of the palm according to the palm image.

[0144] The processing unit 804 is further configured to determine position feature information of the palm according to the human skeleton key point.

[0145] The processing unit 804 is further configured to determine hand sign information corresponding to the hand sign image according to the hand shape feature information and the position feature information.

[0146] The hand language recognition device 800 provided in the embodiments of the present application extracts a palm image and human body skeleton key points respectively in the process of hand language recognition, then determines hand shape feature information of the palm based on the palm image, determines position feature information of the palm based on the human body skeleton key points, and determines hand language information based on the hand shape feature information and the position feature information. In this way, in the process of hand language recognition, not only can the hand shape feature of the palm region be accurately captured based on the palm image, but also the body posture feature can be accurately captured based on the human body skeleton key points, and the accurate palm position feature is obtained based on the body posture feature, which improves the accuracy of extracting the hand language feature, thereby improving the accuracy of the hand language recognition result.

[0147] In the embodiments of the present application, the palm image includes a first image and a second image, the first image corresponds to a first palm, and the second image corresponds to a second palm. The hand shape feature information includes hand shape feature information of the first palm and hand shape feature information of the second palm. The processing unit 804 is specifically configured to: perform symmetry processing on the first image to obtain a third image; encode the third image based on the trained hand shape coding model to obtain the hand shape feature information of the first palm; and encode the second image based on the trained hand shape coding model to obtain the hand shape feature information of the second palm.

[0148] In the embodiments of the present application, the palm image includes a first image and a second image, the first image corresponds to a first palm, and the second image corresponds to a second palm. The hand shape feature information includes hand shape feature information of the first palm and hand shape feature information of the second palm. In the process of determining the hand shape feature information of the palm based on the palm image, the first image is symmetrically processed to obtain a third image. The third image is encoded based on the trained hand shape coding model to obtain the hand shape feature information of the first palm. The second image is encoded based on the trained hand shape coding model to obtain the hand shape feature information of the second palm. In this way, the difference in hand shape between left-handed and right-handed hand languages can be eliminated. The hand shape feature information of the first palm and the second palm obtained for left-handed and right-handed hand languages can be unified with each other, so that the accuracy of hand shape feature extraction can be ensured without increasing the amount of hand shape feature extraction and the amount of model calculation.

[0149] In the embodiments of the present application, before the third image is encoded based on the trained hand shape coding model, the obtaining unit 802 is further configured to obtain a sign language sample image; the processing unit 804 is further configured to, in a case where the palms in the sign language sample image overlap each other, encode the sign language sample image based on the hand shape coding model, and iteratively train the hand shape coding model according to an encoding result of the sign language sample image and a first loss function; in a case where the palms in the sign language sample image are independent of each other, determine a first bounding box according to palm skeletal points in the sign language sample image, and determine a second bounding box according to a palm image in the sign language sample image; determine an intersection over union of the first bounding box and the second bounding box according to size information of the first bounding box and the second bounding box; in a case where the intersection over union is greater than a first threshold, iteratively train the hand shape coding model according to coordinate information of the palm skeletal points in the sign language sample image and a second loss function, thereby obtaining the trained hand shape coding model.

[0150] The above embodiments provided in the present application obtain a sign language sample image before the third image is encoded based on the trained hand shape coding model; in a case where the palms in the sign language sample image overlap each other, encode the sign language sample image based on the hand shape coding model, and iteratively train the hand shape coding model according to an encoding result of the sign language sample image and a first loss function, thereby obtaining the trained hand shape coding model; in a case where the palms in the sign language sample image are independent of each other, determine a first bounding box according to palm skeletal points in the sign language sample image, and determine a second bounding box according to a palm image in the sign language sample image; determine an intersection over union of the first bounding box and the second bounding box according to size information of the first bounding box and the second bounding box; in a case where the intersection over union is greater than a first threshold, iteratively train the hand shape coding model according to coordinate information of the palm skeletal points in the sign language sample image and a second loss function. In this way, the hand shape coding model is trained based on the sign language sample image before the third image is encoded based on the trained hand shape coding model, thereby improving the accuracy of the hand shape coding model in learning hand shape coding of single hand shape and cross hand shape.

[0151] In the embodiment of the present application, the human body skeleton key points include a first key point, a second key point and a plurality of third key points, the first key point corresponds to the first palm, the second key point corresponds to the second palm, and the plurality of third key points correspond to the head and the arm; the position feature information includes the position feature information of the first palm and the position feature information of the second palm; the processing unit 804 is specifically configured to: establish a first coordinate system with the first key point as the origin, and establish a second coordinate system with the second key point as the origin; in the first coordinate system, determine the first coordinate information of each third key point relative to the first key point; in the second coordinate system, determine the second coordinate information of each third key point relative to the second key point; perform symmetry processing on the first coordinate information to obtain third coordinate information, and determine the position feature information of the first palm according to the third coordinate information; and determine the position feature information of the second palm according to the second coordinate information.

[0152] In the above embodiment provided by the present application, the human body skeleton key points include a first key point, a second key point and a plurality of third key points, the first key point corresponds to the first palm, the second key point corresponds to the second palm, and the plurality of third key points correspond to the head and the arm; the position feature information includes the position feature information of the first palm and the position feature information of the second palm; in the process of determining the position feature information of the palm according to the human body skeleton key points, a first coordinate system is established with the first key point as the origin, and a second coordinate system is established with the second key point as the origin; in the first coordinate system, the first coordinate information of each third key point relative to the first key point is determined; in the second coordinate system, the second coordinate information of each third key point relative to the second key point is determined; symmetry processing is performed on the first coordinate information to obtain third coordinate information, and the position feature information of the first palm is determined according to the third coordinate information; and the position feature information of the second palm is determined according to the second coordinate information. In this way, the palm position difference between left-handed sign language and right-handed sign language can be eliminated, and the palm position feature information of the first palm and the second palm obtained by the left-handed sign language and the right-handed sign language can be unified, so that the accuracy of palm position feature extraction can be ensured without increasing the model calculation amount.

[0153] In the embodiment of the present application, the palm includes a first palm and a second palm, and the processing unit 804 is specifically configured to: determine a first palm feature according to the hand shape feature information and the position feature information of the first palm, and determine a second palm feature according to the hand shape feature information and the position feature information of the second palm; perform attention learning on the first palm feature and the second palm feature based on an attention mechanism to obtain a first attention feature and a second attention feature; perform convolution processing on the first attention feature and the second attention feature to obtain sign language feature information; add first classification feature information to the sign language feature information, perform attention learning on the sign language feature information to obtain second classification feature information, and determine sign language information according to the second classification feature information.

[0154] In the above embodiments provided in the application, the palm includes a first palm and a second palm. In the process of determining sign language information corresponding to the sign language image according to the hand shape feature information and the position feature information, first palm features are determined according to the hand shape feature information and the position feature information of the first palm, and second palm features are determined according to the hand shape feature information and the position feature information of the second palm. Attention learning is performed on the first palm features and the second palm features based on an attention mechanism to obtain first attention features and second attention features. The sign language feature information is obtained by performing convolution processing on the first attention features and the second attention features. The second classification feature information is obtained by performing attention learning on the sign language feature information after adding the first classification feature information to the sign language feature information, and the sign language information is determined according to the second classification feature information. In this way, the palm features are determined based on the hand shape feature information and the position feature information of the palm, the sign language feature information of each frame of sign language image is obtained by fusing the palm features of the left and right hands, and the accuracy of the sign language recognition result is improved when sign language recognition is subsequently performed based on the sign language feature information.

[0155] In the embodiments of the application, the processing unit 804 is specifically configured to determine an attention weight matrix according to the first mask, wherein the attention weight between the first palm features and the second palm features corresponding to different sign language images in the attention weight matrix is zero, and perform self-attention learning on the first palm features and the second palm features according to the attention weight matrix.

[0156] In the above embodiments provided in the application, in the process of performing attention learning on the first palm features and the second palm features, an attention weight matrix is determined according to the first mask, wherein the attention weight between the first palm features and the second palm features corresponding to different sign language images in the attention weight matrix is zero, and self-attention learning is performed on the first palm features and the second palm features according to the attention weight matrix. In this way, in the process of performing attention learning on the first palm features and the second palm features, the attention weight between the first palm features and the second palm features is zero, so that the conversion model can only perform self-attention learning on the palm features of the same palm, and the accuracy of the first attention features and the second attention features is ensured.

[0157] The sign language recognition apparatus 800 in the embodiments of the present application can be an electronic device, or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices other than the terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), and the like, and can also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, and the like, and the embodiments of the present application are not limited in this regard.

[0158] The sign language recognition apparatus 800 in the embodiments of the present application can be an apparatus having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, and the embodiments of the present application are not limited in this regard.

[0159] The sign language recognition apparatus 800 provided in the embodiments of the present application can implement the method embodiments Figure 1 , and each process of the method embodiments is not repeated here.

[0160] Optionally, as shown in Figure 9 , the embodiments of the present application further provide an electronic device 900, which includes a processor 902 and a memory 904, and the memory 904 has a program or instructions stored thereon, which can be run on the processor 902. When the program or instructions are executed by the processor 902, each step of the above-mentioned sign language recognition method embodiments is implemented, and the same technical effects are achieved, and each process is not repeated here.

[0161] It should be noted that the electronic device in the embodiments of the present application includes the above-mentioned mobile electronic device and non-mobile electronic device.

[0162] Figure 10 To implement the hardware structure of an electronic device according to an embodiment of the present application.

[0163] The electronic device 1000 includes, but is not limited to, a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, and a processor 1010, etc.

[0164] Those skilled in the art can understand that the electronic device 1000 can also include a power supply (such as a battery) for supplying power to each component, and the power supply can be logically connected to the processor 1010 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 10 The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than the figure, or combine certain components, or different component arrangements, which are not described here.

[0165] The processor 1010 is configured to acquire a sign language image.

[0166] The processor 1010 is further configured to extract a palm image and a human body skeleton key point in the sign language image.

[0167] The processor 1010 is further configured to determine hand shape feature information of the palm according to the palm image.

[0168] The processor 1010 is further configured to determine position feature information of the palm according to the human body skeleton key point.

[0169] The processor 1010 is further configured to determine sign language information corresponding to the sign language image according to the hand shape feature information and the position feature information.

[0170] In the embodiments of the present application, in the process of sign language recognition, the palm image and the human body skeleton key point are extracted respectively, and then the hand shape feature information of the palm is determined based on the palm image, the position feature information of the palm is determined based on the human body skeleton key point, and the sign language information is determined based on the hand shape feature information and the position feature information. In this way, in the process of sign language recognition, not only the hand region feature, i.e. the palm hand shape feature, can be accurately captured based on the palm image, but also the body posture feature can be accurately captured based on the human body skeleton key point, and the accurate palm position feature is obtained based on the body posture feature, which improves the accuracy of extracting sign language features, and thus improves the accuracy of the sign language recognition result.

[0171] Optionally, the palm image includes a first image and a second image, the first image corresponds to a first palm, and the second image corresponds to a second palm, the hand shape feature information includes hand shape feature information of the first palm and hand shape feature information of the second palm, and the processor 1010 is specifically configured to: perform symmetry processing on the first image to obtain a third image; encode the third image based on the trained hand shape coding model to obtain the hand shape feature information of the first palm; and encode the second image based on the trained hand shape coding model to obtain the hand shape feature information of the second palm.

[0172] In the above embodiments provided by the present application, the palm image includes a first image and a second image, the first image corresponds to a first palm, and the second image corresponds to a second palm, the hand shape feature information includes hand shape feature information of the first palm and hand shape feature information of the second palm, in the process of determining the hand shape feature information of the palm according to the palm image, symmetry processing is performed on the first image to obtain a third image; the third image is encoded based on the trained hand shape coding model to obtain the hand shape feature information of the first palm; and the second image is encoded based on the trained hand shape coding model to obtain the hand shape feature information of the second palm. In this way, the difference in hand shape between left-handed sign language and right-handed sign language can be eliminated, and the hand shape feature information of the first palm and the second palm obtained for both left-handed sign language and right-handed sign language can be unified, so that the accuracy of hand shape feature extraction can be ensured without increasing the amount of hand shape feature extraction and the amount of model calculation.

[0173] Optionally, before encoding the third image based on the trained hand shape coding model, the processor 1010 is further configured to: obtain a sign language sample image; in a case where the palms in the sign language sample image overlap each other, encode the sign language sample image based on the hand shape coding model, and iteratively train the hand shape coding model according to an encoding result of the sign language sample image and a first loss function; in a case where the palms in the sign language sample image are independent of each other, determine a first recognition box according to palm skeleton key points in the sign language sample image, and determine a second recognition box according to palm images in the sign language sample image; determine an intersection-over-union of the first recognition box and the second recognition box according to size information of the first recognition box and the second recognition box; and in a case where the intersection-over-union is greater than a first threshold, iteratively train the hand shape coding model according to coordinate information of the palm skeleton key points in the sign language sample image and a second loss function.

[0174] The above embodiments provided in the application, before the third image is encoded based on the trained hand shape coding model, a sign language sample image is obtained; in the case that the palms in the sign language sample image overlap each other, the sign language sample image is encoded based on the hand shape coding model, and the hand shape coding model is iteratively trained according to the encoding result of the sign language sample image and a first loss function; in the case that the palms in the sign language sample image are independent of each other, a first recognition box is determined according to the palm skeletal key points in the sign language sample image, and a second recognition box is determined according to the palm image in the sign language sample image; the size information of the first recognition box and the second recognition box is determined to determine the intersection over union of the first recognition box and the second recognition box; in the case that the intersection over union is greater than a first threshold, the hand shape coding model is iteratively trained according to the coordinate information of the palm skeletal key points in the sign language sample image and a second loss function. In this way, before the third image is encoded based on the trained hand shape coding model, the hand shape coding model is trained based on the sign language sample image, and the accuracy of the hand shape coding model in learning single hand shape and cross hand shape is improved.

[0175] Optionally, the human body skeletal key points include a first key point, a second key point, and a plurality of third key points, the first key point corresponds to a first palm, the second key point corresponds to a second palm, and the plurality of third key points correspond to a head and arms, the position feature information includes position feature information of the first palm and position feature information of the second palm, and the processor 1010 is specifically configured to: establish a first coordinate system with the first key point as an origin, and establish a second coordinate system with the second key point as an origin; in the first coordinate system, determine first coordinate information of each third key point relative to the first key point; in the second coordinate system, determine second coordinate information of each third key point relative to the second key point; perform symmetry processing on the first coordinate information to obtain third coordinate information, and determine the position feature information of the first palm according to the third coordinate information; and determine the position feature information of the second palm according to the second coordinate information.

[0176] In the above embodiments provided in the application, the human body skeleton key points include a first key point, a second key point and a plurality of third key points, the first key point corresponds to a first palm, the second key point corresponds to a second palm, and the plurality of third key points correspond to a head and arms, the position feature information includes position feature information of the first palm and position feature information of the second palm, in the process of determining the position feature information of the palms according to the human body skeleton key points, a first coordinate system is established with the first key point as the origin, and a second coordinate system is established with the second key point as the origin; in the first coordinate system, first coordinate information of each third key point relative to the first key point is determined; in the second coordinate system, second coordinate information of each third key point relative to the second key point is determined; the first coordinate information is symmetrically processed to obtain third coordinate information, and the position feature information of the first palm is determined according to the third coordinate information; and the position feature information of the second palm is determined according to the second coordinate information. In this way, the palm position difference between left-handed sign language and right-handed sign language can be eliminated, and the palm position feature information of the first palm and the second palm obtained for left-handed sign language and right-handed sign language can be unified, so that the accuracy of palm position feature extraction can be ensured without increasing the model calculation amount.

[0177] Optionally, the palms include a first palm and a second palm, and the processor 1010 is specifically configured to: determine a first palm feature according to the hand shape feature information and the position feature information of the first palm, and determine a second palm feature according to the hand shape feature information and the position feature information of the second palm; perform attention learning on the first palm feature and the second palm feature based on an attention mechanism to obtain a first attention feature and a second attention feature; perform convolution processing on the first attention feature and the second attention feature to obtain sign language feature information; perform attention learning on the sign language feature information after adding first classification feature information to the sign language feature information to obtain second classification feature information, and determine sign language information according to the second classification feature information.

[0178] In the above embodiments provided in the application, the palm includes a first palm and a second palm, in the process of determining sign language information corresponding to the sign language image according to the hand shape feature information and the position feature information, a first palm feature is determined according to the hand shape feature information and the position feature information of the first palm, and a second palm feature is determined according to the hand shape feature information and the position feature information of the second palm; the first palm feature and the second palm feature are subjected to attention learning based on an attention mechanism, to obtain a first attention feature and a second attention feature; the first attention feature and the second attention feature are subjected to convolution processing, to obtain sign language feature information; after adding first classification feature information to the sign language feature information, the sign language feature information is subjected to attention learning, to obtain second classification feature information, and the sign language information is determined according to the second classification feature information. In this way, the palm feature is determined based on the hand shape feature information and the position feature information of the palm, the sign language feature information of each frame of sign language image is obtained by fusing the palm features of the left and right hands, and the accuracy of the sign language recognition result is improved when sign language recognition is subsequently performed based on the sign language feature information.

[0179] Optionally, the processor 1010 is specific for determining an attention weight matrix according to the first mask, wherein the attention weight between the first palm feature and the second palm feature corresponding to different sign language images in the attention weight matrix is zero; and performing self-attention learning on the first palm feature and the second palm feature according to the attention weight matrix.

[0180] In the above embodiments provided in the application, in the process of performing attention learning on the first palm feature and the second palm feature, an attention weight matrix is determined according to the first mask, wherein the attention weight between the first palm feature and the second palm feature corresponding to different sign language images in the attention weight matrix is zero; and the first palm feature and the second palm feature are subjected to self-attention learning according to the attention weight matrix. In this way, in the process of performing attention learning on the first palm feature and the second palm feature, the attention weight between the first palm feature and the second palm feature is zero, so that the conversion model can only perform self-attention learning on the palm features of the same palm, and the accuracy of the first attention feature and the second attention feature is ensured.

[0181] It should be understood that in the embodiments of the present application, the input unit 1004 can include a graphics processor (GPU) 10041 and a microphone 10042. The graphics processor 10041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 1006 can include a display panel 10061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes at least one of a touch panel 10071 and other input devices 10072. The touch panel 10071 is also referred to as a touch screen. The touch panel 10071 can include two parts of a touch detection device and a touch controller. The other input devices 10072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which will not be described here.

[0182] The memory 1009 can be used to store software programs and various data. The memory 1009 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), etc. In addition, the memory 1009 can include a volatile memory or a non-volatile memory, or the memory 1009 can include both volatile and non-volatile memories. The non-volatile memory can be a Read-Only Memory (ROM), a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a Random Access Memory (RAM), a Static RAM (SRAM), a Dynamic RAM (DRAM), a Synchronous DRAM (SDRAM), a Double Data Rate SDRAM (DDR SDRAM), an Enhanced SDRAM (ESDRAM), a Synch link DRAM (SLDRAM), and a Direct Rambus RAM (DRRAM). The memory 1009 in the embodiments of the present application includes but is not limited to these and any other suitable types of memories.

[0183] The processor 1010 can include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 1010.

[0184] The embodiments of the present application also provide a readable storage medium, and the readable storage medium stores programs or instructions, the programs or instructions are executed by a processor to realize various processes of the above-mentioned sign language recognition method embodiments and achieve the same technical effects. To avoid repetition, details are not described here.

[0185] The processor is a processor in the electronic device in the above embodiments. The readable storage medium includes a computer readable storage medium, such as a computer readable memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0186] The chip provided in the embodiments of the present application includes a processor and a communication interface. The communication interface is coupled with the processor. The processor is configured to execute programs or instructions to implement various processes of the above-mentioned sign language recognition method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0187] It should be understood that the chip mentioned in the embodiments of the present application can also be referred to as a system-level chip, a system chip, a chip system, or a system-on-chip, etc.

[0188] The embodiments of the present application provide a computer program product stored in a storage medium. The program product is executed by at least one processor to implement various processes of the above-mentioned sign language recognition method embodiments and achieve the same technical effects. To avoid repetition, details are not described herein.

[0189] It should be noted that in this document, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of another identical element in the process, method, article or device that includes the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to the order of performing the functions as shown or discussed, but can also include performing the functions in a substantially simultaneous manner or in reverse order, for example, the described method can be performed in an order different from that described, and various steps can be added, omitted or combined. In addition, the features described with reference to certain examples can be combined in other examples.

[0190] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer software product, which is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk), and includes a plurality of instructions for making a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) execute the methods of various embodiments of the present application.

[0191] The embodiments of the present application are described above with reference to the accompanying drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative, but not restrictive, and a person of ordinary skill in the art can make many forms under the inspiration of the present application without departing from the purpose of the present application and the scope protected by the claims.

Claims

1. A sign language recognition method, characterized in that, The sign language recognition method includes: Acquire sign language images; Extract the palm image and human skeleton key points from the sign language image. The human skeleton key points include a first key point, a second key point and multiple third key points. The first key point corresponds to the first palm, the second key point corresponds to the second palm, and the multiple third key points correspond to the head and arms. The hand shape feature information of the palm is determined based on the palm image; The positional features of the hand are determined based on the key points of the human skeleton. Based on the hand shape feature information and the position feature information, the sign language information corresponding to the sign language image is determined; The palm image includes a first image and a second image, the first image corresponding to the first palm and the second image corresponding to the second palm. The hand shape feature information includes hand shape feature information of the first palm and hand shape feature information of the second palm. Determining the hand shape feature information of the palm based on the palm image includes: The first image is symmetrically processed to obtain the third image; The third image is encoded based on the trained hand shape encoding model to obtain the hand shape feature information of the first palm; The second image is encoded based on the trained hand shape encoding model to obtain the hand shape feature information of the second palm.

2. The sign language recognition method according to claim 1, characterized in that, Before encoding the third image based on the trained hand shape encoding model, the sign language recognition method further includes: Obtain sign language sample images; When the palms of the sign language sample images overlap, the sign language sample images are encoded based on the hand shape encoding model, and the hand shape encoding model is iteratively trained according to the encoding results of the sign language sample images and the first loss function to obtain the trained hand shape encoding model; When the palms in the sign language sample images are independent of each other, a first recognition box is determined based on the key points of the palm bones in the sign language sample images, and a second recognition box is determined based on the palm images in the sign language sample images; Based on the size information of the first recognition frame and the second recognition frame, determine the intersection-union ratio of the first recognition frame and the second recognition frame; If the intersection-union ratio is greater than the first threshold, the hand shape encoding model is iteratively trained based on the coordinate information of the key points of the hand bones in the sign language sample image and the second loss function to obtain the trained hand shape encoding model.

3. The sign language recognition method according to claim 1, characterized in that, The positional feature information includes the positional feature information of the first palm and the positional feature information of the second palm. Determining the positional feature information of the palm based on the key points of the human skeleton includes: A first coordinate system is established with the first key point as the origin, and a second coordinate system is established with the second key point as the origin; In the first coordinate system, determine the first coordinate information of each of the third key points relative to the first key point; In the second coordinate system, determine the second coordinate information of each of the third key points relative to the second key point; The first coordinate information is symmetrically processed to obtain the third coordinate information, and the positional feature information of the first palm is determined based on the third coordinate information. The positional features of the second hand are determined based on the second coordinate information.

4. The sign language recognition method according to claim 1, characterized in that, The hand includes a first hand and a second hand. Determining the sign language information corresponding to the sign language image based on the hand shape feature information and the position feature information includes: The features of the first palm are determined based on the hand shape and position features of the first palm, and the features of the second palm are determined based on the hand shape and position features of the second palm. Based on the attention mechanism, attention learning is performed on the first palm feature and the second palm feature to obtain the first attention feature and the second attention feature; The first attention feature and the second attention feature are convolved to obtain sign language feature information; After adding first classification feature information to the sign language feature information, attention learning is performed on the sign language feature information to obtain second classification feature information, and the sign language information is determined based on the second classification feature information.

5. The sign language recognition method according to claim 4, characterized in that, The attention learning of the first palm features and the second palm features includes: An attention weight matrix is ​​determined based on a first mask, wherein the attention weight between the first palm feature and the second palm feature corresponding to different sign language images is zero in the attention weight matrix. Self-attention learning is performed on the first palm features and the second palm features based on the attention weight matrix.

6. A sign language recognition device, characterized in that, The sign language recognition device includes: Acquisition unit, used to acquire sign language images; The processing unit is used to extract the palm image and human skeleton key points from the sign language image. The human skeleton key points include a first key point, a second key point and a plurality of third key points. The first key point corresponds to the first palm, the second key point corresponds to the second palm, and the plurality of third key points correspond to the head and arms. The processing unit is also used to determine the hand shape feature information of the palm based on the palm image; The processing unit is also used to determine the positional feature information of the palm based on the key points of the human skeleton; The processing unit is further configured to determine the sign language information corresponding to the sign language image based on the hand shape feature information and the position feature information; The palm image includes a first image and a second image, the first image corresponding to the first palm and the second image corresponding to the second palm. The hand shape feature information includes hand shape feature information of the first palm and hand shape feature information of the second palm. The processing unit is specifically used for: The first image is symmetrically processed to obtain the third image; The third image is encoded based on the trained hand shape encoding model to obtain the hand shape feature information of the first palm; The second image is encoded based on the trained hand shape encoding model to obtain the hand shape feature information of the second palm.

7. The sign language recognition device according to claim 6, characterized in that, Before encoding the third image based on the trained hand-shaped encoding model, the acquisition unit is further configured to: Obtain sign language sample images; The processing unit is also used for: When the palms of the sign language sample images overlap, the sign language sample images are encoded based on the hand shape encoding model, and the hand shape encoding model is iteratively trained according to the encoding results of the sign language sample images and the first loss function to obtain the trained hand shape encoding model; When the palms in the sign language sample images are independent of each other, a first recognition box is determined based on the key points of the palm bones in the sign language sample images, and a second recognition box is determined based on the palm images in the sign language sample images; Based on the size information of the first recognition frame and the second recognition frame, determine the intersection-union ratio of the first recognition frame and the second recognition frame; If the intersection-union ratio is greater than the first threshold, the hand shape encoding model is iteratively trained based on the coordinate information of the key points of the hand bones in the sign language sample image and the second loss function to obtain the trained hand shape encoding model.

8. The sign language recognition device according to claim 6, characterized in that, The positional feature information includes the positional feature information of the first palm and the positional feature information of the second palm, and the processing unit is specifically used for: A first coordinate system is established with the first key point as the origin, and a second coordinate system is established with the second key point as the origin; In the first coordinate system, determine the first coordinate information of each of the third key points relative to the first key point; In the second coordinate system, determine the second coordinate information of each of the third key points relative to the second key point; The first coordinate information is symmetrically processed to obtain the third coordinate information, and the positional feature information of the first palm is determined based on the third coordinate information. The positional features of the second hand are determined based on the second coordinate information.

9. The sign language recognition device according to claim 6, characterized in that, The palm includes a first palm and a second palm, and the processing unit is specifically used for: The features of the first palm are determined based on the hand shape and position features of the first palm, and the features of the second palm are determined based on the hand shape and position features of the second palm. Based on the attention mechanism, attention learning is performed on the first palm feature and the second palm feature to obtain the first attention feature and the second attention feature; The first attention feature and the second attention feature are convolved to obtain sign language feature information; After adding first classification feature information to the sign language feature information, attention learning is performed on the sign language feature information to obtain second classification feature information, and the sign language information is determined based on the second classification feature information.

10. The sign language recognition device according to claim 9, characterized in that, The processing unit is specifically used for: An attention weight matrix is ​​determined based on a first mask, wherein the attention weight between the first palm feature and the second palm feature corresponding to different sign language images is zero in the attention weight matrix. Self-attention learning is performed on the first palm features and the second palm features based on the attention weight matrix.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the steps of the sign language recognition method as described in any one of claims 1 to 5.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the sign language recognition method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Continuous sign language recognition method based on multiple feature points

    CN113780059A

  • Sign language recognition method and device

    CN114821761A