Sign language recognition method and device, electronic equipment and storage medium

By using low-dimensional linguistic features instead of high-dimensional skeletal point features in electronic devices to recognize sign language movements, the problem of low recognition efficiency in existing technologies is solved, and more efficient sign language recognition is achieved.

CN115661927BActive Publication Date: 2026-04-07VIVO MOBILE COMM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, electronic devices are less efficient at recognizing sign language, mainly because the dimensionality of the skeletal points output by the skeletal point feature template is high, which requires feature retrieval from a large number of vocabulary templates.

Method used

By acquiring the hand feature information of the target user in the video frames of the target video, M video segments are determined, and sign language recognition is performed on the M video segments. Low-dimensional linguistic features are used to replace high-dimensional skeletal point features, and N video segments are further filtered to obtain N video segments.

Benefits of technology

It improves the efficiency of electronic devices in recognizing sign language gestures, reduces the amount of computation, and increases the recognition speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115661927B_ABST
    Figure CN115661927B_ABST
Patent Text Reader

Abstract

This application discloses a sign language recognition method, device, electronic device, and storage medium, belonging to the field of artificial intelligence technology. The method includes: acquiring hand feature information of a target user from video frames of a target video; determining M video segments from the target video based on the hand feature information, each video segment containing at least one consecutive video frame, and each video segment including video content corresponding to similar sign language actions, where M is an integer greater than 1; performing sign language recognition on N video segments from the M video segments to obtain the target user's sign language information, where each of the N video segments contains video content corresponding to the target sign language action, and N is an integer less than or equal to M.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a sign language recognition method, device, electronic device, and storage medium. Background Technology

[0002] Currently, users can use electronic devices to translate sign language, allowing them to communicate with other users based on the translated sign language. In existing technology, electronic devices can input images or videos containing sign language into a skeletal point feature template, thereby recognizing the meaning of the sign language in the image or video through the skeletal point feature template.

[0003] However, in the above methods, because the skeletal point feature template outputs a high dimensionality of skeletal points (i.e., the skeletal point feature template outputs a lot of skeletal point feature information), electronic devices need to perform feature retrieval in a large number of vocabulary templates in order to achieve a relatively accurate result. As a result, the efficiency of electronic devices in recognizing sign language is low. Summary of the Invention

[0004] The purpose of this application is to provide a sign language recognition method, device, electronic device, and storage medium that can solve the problem of low efficiency in sign language recognition by sub-devices.

[0005] In a first aspect, embodiments of this application provide a sign language recognition method, which includes: acquiring hand feature information of a target user in video frames of a target video; determining M video segments from the target video based on the hand feature information, each video segment containing at least one consecutive video frame, and each video segment including video content corresponding to similar sign language actions, where M is an integer greater than 1; performing sign language recognition on N video segments from the M video segments to obtain the sign language information of the target user, where each of the N video segments contains video content corresponding to the target sign language action, and N is an integer less than or equal to M.

[0006] Secondly, embodiments of this application provide a sign language recognition device, which includes an acquisition module, a determination module, and a processing module. The acquisition module is used to acquire hand feature information of a target user from video frames of a target video. The determination module is used to determine M video segments from the target video based on the hand feature information. Each video segment contains at least one consecutive video frame, and each video segment includes video content corresponding to similar sign language actions, where M is an integer greater than 1. The processing module is used to perform sign language recognition on N video segments out of the M video segments to obtain the sign language information of the target user. Each of the N video segments contains video content corresponding to the target sign language action, where N is an integer less than or equal to M.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions, when executed by the processor, implementing the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0011] In this embodiment, the electronic device can acquire the hand feature information of the target user in video frames of the target video, and then determine M video segments from the target video based on the hand feature information. It then performs sign language recognition on N video frames out of the M video frames to obtain the target user's sign language information. In this solution, the electronic device can use the target user's hand feature information instead of skeletal point feature information. That is, the electronic device can abstract complex and high-dimensional skeletal point features into low-dimensional linguistic features, thereby determining M video segments containing sign language actions from the target video based on the low-dimensional linguistic features. Furthermore, the electronic device can further filter the M video segments to obtain N video segments. Since the electronic device uses low-dimensional linguistic features, it can obtain N video segments containing the target user's sign language actions with less computation, thus improving the efficiency of the electronic device in recognizing sign language actions. Attached Figure Description

[0012] Figure 1 This is a flowchart of a sign language recognition method provided in an embodiment of this application;

[0013] Figure 2 This is a schematic diagram of the structure of a sign language recognition device provided in an embodiment of this application;

[0014] Figure 3 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;

[0015] Figure 4 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0018] The sign language recognition method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0019] Currently, with the development of communication technology, electronic devices are becoming increasingly functional. For example, electronic devices can recognize videos containing sign language gestures to obtain sign language gesture information. In the prior art, (1) electronic devices can randomly cut the video by sliding window to obtain continuous video frames containing sign language actions, and then perform sign language action recognition on the cut video frames to obtain sign language action information; (2) electronic devices can perform sign language action recognition on the video according to the skeletal point feature template to obtain sign language action information; (3) electronic devices can predict the segmentation points that should be segmented in the video according to the convolutional neural network, and then obtain sign language action information through the segmented video frames; however, for the above scheme (1), the electronic device cuts out a large number of video segments by brute force cutting method, and the electronic device is prone to false recall when performing sign language action recognition through multiple video segments; for the above scheme (2), the feature information output by the skeletal point feature template has many dimensions, and after obtaining the feature information, the electronic device needs a large number of vocabulary templates for feature retrieval, resulting in low efficiency of the electronic device in obtaining sign language action information; for the above scheme (3), the electronic device needs to perform a large amount of training in the early stage when predicting the segmentation points of the video through the convolutional neural network in order to achieve the preset effect.

[0020] In this embodiment, the electronic device can acquire the hand feature information of the target user in video frames of the target video. Then, based on this hand feature information, it determines M video segments from the target video and performs sign language recognition on N video frames out of the M video frames to obtain the target user's sign language information. In this solution, the electronic device can use the target user's hand feature information instead of skeletal point feature information. That is, the electronic device can abstract complex and high-dimensional skeletal point features into low-dimensional linguistic features. Based on these low-dimensional linguistic features, it can determine M video segments containing sign language actions from the target video. Furthermore, the electronic device can further filter these M video segments to obtain N video segments. Since the electronic device uses low-dimensional linguistic features, it can obtain N video segments containing the target user's sign language actions with less computation, thus improving the efficiency of the electronic device in recognizing sign language actions.

[0021] The execution subject of the sign language recognition method provided in this application embodiment can be a sign language recognition device, which can be an electronic device or a functional module in an electronic device. The following uses an electronic device as an example to illustrate the technical solution provided in this application embodiment.

[0022] This application provides a sign language recognition method. Figure 1 A flowchart of a sign language recognition method provided in an embodiment of this application is shown. Figure 1 As shown, the sign language recognition method provided in this application embodiment may include the following steps 201 to 203.

[0023] Step 201: The electronic device acquires the hand feature information of the target user in the video frame of the target video.

[0024] Optionally, in this embodiment of the application, the above-mentioned hand feature information includes at least one of the following features: target mark, palm orientation, wrist orientation, and hand shape. The target mark is used to characterize whether the target user performs sign language actions.

[0025] It is understood that the target video mentioned above is a video containing the sign language gestures of the target user.

[0026] Optionally, in this embodiment of the application, the target video may be downloaded by the user through a browser application on the electronic device; or obtained by the user through a video application on the electronic device.

[0027] It is understandable that the target users mentioned above are the users who perform sign language gestures in the target video.

[0028] Optionally, in this embodiment of the application, if the target video contains multiple users, the electronic device can select the target user from the multiple users and obtain the hand feature information of the target user.

[0029] For example, an electronic device can input each frame of the target video into a skeletal point feature extraction model (e.g., the MediaPipe model) to obtain skeletal point information of multiple users in each frame. The electronic device can then select a target user from multiple users based on the skeletal point information and obtain the target user's hand feature information through the target user's skeletal point information.

[0030] Optionally, in this embodiment of the application, the above-mentioned skeletal point information includes body skeletal point information and hand skeletal point information.

[0031] For example, an electronic device can select a target user from multiple users through body skeletal point information, and then obtain the target user's hand feature information through the target user's hand skeletal point information.

[0032] Optionally, in this application, after obtaining the skeletal point information of the target user, the electronic device can save the skeletal point information of the target user through the skeletal point sequence.

[0033] In this embodiment of the application, after obtaining the skeletal point information of the user in the target video, the electronic device can determine whether the user is performing sign language actions based on the skeletal points of the user in the target video, and then add target markers to the video frames containing sign language actions.

[0034] For example, taking the left hand as an example, the electronic device can use the following formula (1) to calculate whether the angle formed by the three points of the shoulder (represented by 11), elbow (represented by 13), and wrist (represented by 15) is less than 150 degrees, thereby determining whether the user is making a sign language gesture. The specific formula is as follows:

[0035] angle = tan -1 ((y 15 -y 13 ) / (x 15 -x 13 ))-tan -1 ((y 11 -y 13 ) / (x 11 -x 13 )) (Formula 1)

[0036] Where angle is the angle measure, y 15 Information on wrist bone points on the Y-axis, y 13 Information on elbow bone points on the Y-axis, y 11 Information on the shoulder bone points on the Y-axis.

[0037] Optionally, in this embodiment of the application, the target marker can be directly added to the video frame containing sign language actions, or added and stored in the target storage space in a key-value pair storage manner.

[0038] Optionally, in this embodiment of the application, when both of the target user's hands are performing sign language gestures, the electronic device can use a target marker to indicate that both of the target user's hands are performing sign language gestures; or, if one of the target user's hands is performing sign language gestures while the other hand is not, the electronic device can use different target markers to mark the target user's hands.

[0039] Optionally, in the embodiments of this application, the target mark may include at least one of the following: numerical identifiers, letter identifiers, and special symbol identifiers, etc.

[0040] For example, an electronic device can input video frames that have been marked with targets into a YOLO target detection model to obtain the hand shapes of the target user's hands.

[0041] Optionally, in the embodiments of this application, step 201 can be implemented by steps 201a to 201e as described below.

[0042] Step 201a: The electronic device acquires the horizontal slope of the target user's hand, the vertical slope of the hand, the variance of the horizontal line of the hand, and the variance of the vertical line of the hand.

[0043] In this embodiment of the application, the electronic device can perform linear regression calculation on the horizontal bone points of the target user's hand to obtain the horizontal line of the target user's hand, that is, to obtain the slope of the horizontal line of the target user's hand; and perform linear regression calculation on the vertical bone points of the target user's hand to obtain the vertical line of the target user's hand, that is, to obtain the slope of the horizontal line of the target user's hand. Then, the variance of the horizontal line of the hand corresponding to the palm horizontal line is obtained according to a preset formula.

[0044] Optionally, in the embodiments of this application, the above step 201a can be implemented by the following steps 201a1 and 201a2.

[0045] Step 201a1: The electronic device determines the slope of the horizontal line and the slope of the vertical line of the hand based on the hand shape and skeletal point information.

[0046] Optionally, in this embodiment of the application, the above-mentioned variance of the horizontal line of the hand may include the variance of the horizontal line of the hand corresponding to the X-axis and the variance of the horizontal line of the hand corresponding to the Y-axis.

[0047] For example, taking a single video frame as an example, suppose the first bone point in the hand skeleton points corresponds to bone points 5, 9, 13, and 17. The electronic device can perform linear regression on the bone point information corresponding to 5, 9, 13, and 17, and then obtain the horizontal lines of the skeleton corresponding to 5, 9, 13, and 17. It can be understood that after obtaining the horizontal lines of the hand, the electronic device obtains the slope of the horizontal lines of the hand through the corresponding relationship between the horizontal lines of the hand and the slope of the horizontal lines of the hand.

[0048] Optionally, in this embodiment of the application, the above-mentioned variance of the hand vertical line may include the variance of the hand vertical line corresponding to the X-axis and the variance of the hand vertical line corresponding to the Y-axis.

[0049] For example, taking a single video frame as an example, assuming that the second bone point in the hand skeleton points corresponds to bone points 0 and 13, the electronic device can perform a linear regression operation on the bone point information corresponding to 0 and 13, and then obtain the bone vertical line corresponding to 0 and 13. It can be understood that after obtaining the hand vertical line, the electronic device obtains the slope of the hand vertical line through the corresponding relationship between the hand vertical line and the slope of the hand vertical line.

[0050] Step 201a2: The electronic device determines the variance of the horizontal line of the hand based on the slope of the horizontal line, and determines the variance of the vertical line of the hand based on the slope of the vertical line.

[0051] For example, after obtaining the slope of the horizontal line of the hand, the electronic device can obtain the variance of the horizontal line of the hand using the following formulas (2) to (5), specifically:

[0052]

[0053]

[0054] in, Var represents the average value of the horizontal target bone points corresponding to the X-axis. x This represents the variance of the horizontal target skeleton points corresponding to the X-axis.

[0055]

[0056]

[0057] in, Var represents the average value of the horizontal target bone points corresponding to the Y-axis. y This represents the variance of the horizontal target skeleton points corresponding to the Y-axis.

[0058] It should be noted that the above embodiments are illustrated using a single video frame. The slope and variance of the horizontal hand line can be obtained using the above method for at least one video frame of the target video.

[0059] For example, after obtaining the slope of the vertical line of the hand, the electronic device can obtain the variance of the vertical line of the hand using the following formulas (6) to (9), specifically:

[0060]

[0061]

[0062] in, Var represents the average value of the longitudinal target bone points corresponding to the X-axis. x This represents the variance of the longitudinal target skeleton points corresponding to the X-axis.

[0063]

[0064]

[0065] in, Var represents the average value of the longitudinal target bone points corresponding to the Y-axis. y This represents the variance of the longitudinal target skeleton points corresponding to the Y-axis.

[0066] It should be noted that the above embodiments are illustrated using a single video frame. The slope and variance of the hand's vertical line can be obtained using the above method for at least one video frame of the target video.

[0067] Step 201b: The electronic device determines the parallelism of the palm horizontal lines based on the slope and variance of the palm horizontal lines.

[0068] Optionally, in this embodiment of the application, the above-mentioned parallel state of the palm horizontal line can be represented by a three-bit binary number, where each bit indicates whether it is parallel to the corresponding coordinate axis.

[0069] For example, 100 indicates parallel to the Z-axis, 010 indicates parallel to the Y-axis, 001 indicates parallel to the X-axis, and 101 indicates parallel to either the X-axis or the Z-axis.

[0070] Optionally, in the embodiments of this application, step 201b above can be implemented by steps 301 and 302 below.

[0071] Step 301: The electronic device determines the target horizontal reference parallel axis corresponding to the horizontal line of the hand based on the slope and variance of the horizontal line of the hand.

[0072] In this embodiment, the electronic device can determine the target horizontal reference parallel axis corresponding to the horizontal line of the hand based on whether the slope of the horizontal line of the hand matches a preset slope parallel threshold and whether the variance of the horizontal line of the hand matches a lower variance threshold and an upper variance threshold.

[0073] For example, if the slope of the horizontal line of the hand is less than a preset slope parallel threshold, the electronic device considers that the horizontal line of the hand may be parallel to the X-axis or Z-axis, and then judges the variance of the horizontal line of the hand. If the variance of the horizontal line of the hand is less than the lower limit threshold of variance, the horizontal line of the hand is close to a point on the xy plane, i.e., case 1, the horizontal line of the hand is parallel to the Z-axis, and the parallel state is set to 100; if the variance of the horizontal line of the hand is greater than the upper limit threshold, i.e., case 2, it is parallel to the X-axis, and the parallel state is set to 001; if the variance of the horizontal line of the hand is between the lower limit threshold of variance and the upper limit threshold of variance, i.e., case 3, the horizontal line of the hand is between the X-axis and the Z-axis, and the parallel state is set to 101.

[0074] Step 302: The electronic device determines the parallel state of the palm's horizontal line based on the target's horizontal reference parallel axis.

[0075] In this embodiment of the application, the electronic device can determine the parallel state of the palm's horizontal line based on the target lateral reference parallel axis of the hand's horizontal line.

[0076] In this embodiment, the electronic device can determine the palm orientation and wrist orientation of the target user by the parallel state of the palm's horizontal line, and then recognize the target user's sign language movements by palm orientation and wrist orientation. That is, the electronic device can recognize the target user's sign language movements by using low-dimensional linguistic features, which reduces the amount of computation required for the electronic device to recognize the target user's sign language movements by using high-dimensional skeletal point information, and improves the efficiency of the electronic device in recognizing the target user's sign language movements.

[0077] Step 201c: The electronic device determines the parallelism of the palm's vertical line based on the slope and variance of the vertical line of the hand.

[0078] Optionally, in this embodiment of the application, the above-mentioned parallel state of the palm vertical line can be represented by a three-bit binary number, where each bit indicates whether it is parallel to the corresponding coordinate axis.

[0079] For example, 100 indicates parallel to the Z-axis, 010 indicates parallel to the Y-axis, 001 indicates parallel to the X-axis, and 101 indicates parallel to either the X-axis or the Z-axis.

[0080] Optionally, in the embodiments of this application, step 201c can be implemented by steps 401 and 402 as described below.

[0081] Step 401: The electronic device determines the target longitudinal reference parallel axis corresponding to the longitudinal line of the hand based on the slope and variance of the longitudinal line of the hand.

[0082] In this embodiment, the electronic device can determine the target longitudinal reference parallel axis corresponding to the hand longitudinal line based on whether the slope of the hand longitudinal line matches a preset slope perpendicular threshold and whether the variance of the hand longitudinal line matches a lower variance threshold and an upper variance threshold.

[0083] For example, if the slope of the vertical line of the hand is greater than the vertical slope threshold, the electronic device considers that the vertical line of the hand may be parallel to the Y-axis or Z-axis, and then makes the same judgment on the variance of the vertical line of the hand. If the variance of the vertical line of the hand is less than the lower limit threshold, i.e., case 1, parallel to the Z-axis, and the parallel state is set to 100; if the variance of the vertical line of the hand is greater than the upper limit threshold, i.e., case 2, parallel to the Y-axis, and the parallel state is set to 010; if the variance of the vertical line of the hand is between the lower limit threshold and the upper limit threshold, i.e., case 3, the parallel state is between the Y-axis and the Z-axis, and the parallel state is set to 110.

[0084] For example, if the slope of the vertical line or the slope of the horizontal line of the hand is between the parallel threshold and the vertical threshold, and if both the variance of the horizontal line and the variance of the vertical line are less than the lower threshold, it belongs to case 1, which is parallel to the Y-axis, and the parallel state is set to 100; if only the variance of the vertical line of the hand is less than the lower threshold, it belongs to case 2, which is parallel between the X-axis and the Z-axis, and the parallel state is set to 101; if the variance of the horizontal line of the hand is less than the lower threshold, it belongs to case 3, which is parallel between the Y-axis and the Z-axis, and the parallel state is set to 110; if both the variance of the horizontal line and the variance of the vertical line of the hand are greater than the lower threshold, it belongs to case 4, which is parallel between the X-axis and the Y-axis, and the parallel state is set to 011.

[0085] Step 402: The electronic device determines the parallel state of the palm's longitudinal line based on the target's longitudinal reference parallel axis.

[0086] In this embodiment of the application, the electronic device can determine the parallel state of the palm's longitudinal line based on the target longitudinal reference parallel axis of the hand's longitudinal line.

[0087] In this embodiment, the electronic device can determine the palm orientation and wrist orientation of the target user by the parallel state of the palm's vertical lines, and then recognize the target user's sign language movements by palm orientation and wrist orientation. That is, the electronic device can recognize the target user's sign language movements by using low-dimensional linguistic features, which reduces the amount of computation required for the electronic device to recognize the target user's sign language movements by using high-dimensional skeletal point information, and improves the efficiency of the electronic device in recognizing the target user's sign language movements.

[0088] Step 201d: The electronic device determines the orientation of the target user's hand based on the skeletal point information of the hand.

[0089] Optionally, in this embodiment of the application, the electronic device can determine the direction of the target user's hand by using the vertex bone point information of the horizontal line of the hand (i.e., the bone points 5 and 17 mentioned above) or by using the vertex bone point information of the vertical line of the hand (i.e., the bone points 0 and 133 mentioned above).

[0090] For example, based on the vertex skeletal point information of the aforementioned hand horizontal line, the electronic device can calculate the direction of the hand horizontal line using the XYZ coordinates of skeletal points 5 and 17, specifically as follows:

[0091] If X_17>X_5: the direction of the horizontal line of the hand is left; if X_17<=X_5: the direction of the horizontal line of the hand is right; if Y_17>Y_5: the direction of the horizontal line of the hand is up; if Y_17<=Y_5: the direction of the horizontal line of the hand is down; if Z_17>Z_5: the direction of the horizontal line of the hand is backward; if Z_17<=Z_5: the direction of the horizontal line of the hand is forward.

[0092] For example, based on the vertex skeletal point information of the hand's vertical line, the electronic device can calculate the direction of the hand's vertical line using the XYZ coordinates of skeletal points 0 and 13, specifically as follows:

[0093] If X_13>X_0: the direction of the vertical line of the hand is left; if X_13<=X_0: the direction of the vertical line of the hand is right; if Y_13>Y_0: the direction of the vertical line of the hand is up; if Y_13<=Y_0: the direction of the vertical line of the hand is down; if Z_13>Z_0: the direction of the vertical line of the hand is backward; if Z_13<=Z_0: the direction of the vertical line of the hand is forward.

[0094] Step 201e: The electronic device determines the palm orientation and wrist orientation of the target user's hand based on the parallel state and orientation of the target user's hand.

[0095] In this embodiment of the application, the aforementioned hand feature information includes: palm orientation and wrist orientation, and the parallel state of the target user's hand includes either the horizontal parallel state of the palm or the vertical parallel state of the palm.

[0096] In this embodiment of the application, the electronic device can determine the palm orientation and wrist orientation of the target user's hand based on the parallel state of the target user's hand and the correspondence between the direction of the target user's hand and the preset combination number.

[0097] Optionally, in the embodiments of this application, step 201e can be implemented by step 501 as described below.

[0098] Step 501: Based on the correspondence, the electronic device determines the first hand feature corresponding to the parallel state of the target user's hand and the direction of the target user's hand.

[0099] In the implementation of this application, the above correspondence includes the parallel state of the target user's hand and the mapping relationship between the direction of the target user's hand and the hand feature information, and the first hand feature is the feature determined based on the correspondence in the hand feature information.

[0100] For example, taking the right hand as an example, if the first digit of the horizontal line of the hand is 1 (parallel to the X-axis), then the first digit of the combination number is A. If the horizontal line of the hand is to the right, then the second digit of the combination number is x, and to the left it is y. When the third digit of the vertical line of the palm is 1 (parallel to the Z-axis), if the vertical line of the palm is forward, then the third digit of the combination number is 1, and to the back it is 4. Specifically, this can be found in Table 1 below.

[0101]

[0102]

[0103]

[0104] Table 1

[0105] Optionally, in this embodiment of the application, since the left and right hands are mirror images, after calculating the combination number with reference to the calculation formula of the right hand, the XY in the middle of the combination number can be interchanged to obtain the first hand feature information of the left hand.

[0106] Step 202: The electronic device determines M video segments from the target video based on the hand feature information.

[0107] In this embodiment of the application, each of the above M video segments contains at least one consecutive video frame, and each video segment includes video content corresponding to similar sign language actions, where M is an integer greater than 1.

[0108] Optionally, in the embodiments of this application, step 202 above can be specifically implemented by steps 202a to 202c below.

[0109] Step 202a: The electronic device uses hand feature information to identify at least one video frame in the target video to obtain sign language movement difference information between every two video frames in at least one video frame.

[0110] For example, taking two video frames as an example, the electronic device can first compare the target markers in the two video frames. If they are different, the difference is marked as dissimilar. If the target markers in the two video frames are the same, the palm orientation and wrist orientation in the two video frames are then compared. When the palm orientation and wrist orientation in video frame 1 are exactly the same as those in video frame 2, the difference is marked as the same. If at least one of the orientations is the same, the difference is marked as similar; otherwise, they are dissimilar. When the palm orientation and wrist orientation of the terminal are the same in both video frames, the hand shape of the target user in the two video frames is then compared. If the hand shapes are the same, the difference is marked as equal; otherwise, they are similar.

[0111] Optionally, in this embodiment of the application, the electronic device can merge the difference marks of the left and right hands according to the priority order of dissimilar > similar > same, and take the one with the highest priority as the difference mark of the frame image.

[0112] Step 202b: The electronic device determines I target video frames from at least one video frame based on the difference category of the sign language movement difference information.

[0113] In this embodiment of the application, I target video frames are video frames containing similar sign language actions, where I is an integer greater than 1.

[0114] In this embodiment of the application, each time the electronic device reads a new video frame, the electronic device can choose to add the new video frame to the video segment or create a new video segment based on the difference marker between the video frame and the previous video frame.

[0115] For example, using windows to represent video segments, when an electronic device reads in the first video frame, it can create a new window (hereinafter referred to as the first window) based on that video frame. Then, it continues to read in the second video frame and determines the difference identifier (e.g., equal, similar, dissimilar) between the second video frame and the first video frame. If the difference identifier between the first and second video frames is equal, the electronic device can add the second video frame to the first window. If the difference identifier between the first and second video frames is similar, the electronic device can add the second video frame to the first window, create a new second window, and use the second video frame as the first video frame in the second window. If the difference identifier between the first and second video frames is dissimilar, the electronic device can create a new third window and use the second video frame as the first video frame in the second window. In this way, M video segments are determined in at least one video frame.

[0116] Step 202c: The electronic device obtains M video segments based on I target video frames.

[0117] In this embodiment of the application, the electronic device can obtain I target video frames with similar sign language actions based on the difference identifier between every two video frames, and then obtain M video segments from the I target video frames according to the original timing information of the target video.

[0118] In this embodiment of the application, the electronic device can determine M video segments with similar sign language actions from at least one video frame based on the difference identifier of the video frame. In this way, the electronic device can perform sign language recognition on the segmented M video segments, thereby improving the efficiency of the electronic device in recognizing sign language actions.

[0119] Optionally, in this embodiment of the application, after step 202 above, the sign language recognition method provided in this embodiment of the application further includes steps 601 to 603.

[0120] Step 601: The electronic device determines L video segments from the M video segments based on the first value of each video segment and the length of each video segment.

[0121] In this embodiment of the application, the first value is the average value of the difference program corresponding to each video segment.

[0122] For example, the electronic device can obtain the average value of the difference program corresponding to each video segment using the following formula (10):

[0123] Score avg =Score / (End-Start+1) (Formula 10)

[0124] Among them, Score avg The average difference score for each video segment is denoted as Score, the total difference score for each video segment is denoted as Score, the difference score for the last video frame in the video segment is denoted as End, and the difference score for the first video frame in the video segment is denoted as Start.

[0125] For example, the electronic device can obtain the length of each video segment using the following formula (11):

[0126] Length=End-Start+1 (Formula 11)

[0127] Where Length is the length of the video segment corresponding to each video segment, End is the index of the last frame of the video segment, and Start is the index of the first frame of the video segment.

[0128] Optionally, in this embodiment of the application, the electronic device can select a video segment that meets a preset length from M video segments according to the video segment length corresponding to each video segment, and then sort the video segments that meet the preset length according to a first value.

[0129] Step 602: The electronic device sorts the L video segments according to the first value and calculates the similarity value between the target video segment in the sorted L video segments and the other video segments in the L video segments.

[0130] For example, taking two windows as an example, the electronic device can calculate the similarity value between the target video segment in L video segments and the other video segments in L video segments using the following formulas (12) to (18); the formulas are as follows:

[0131] Start = max(Start) A Start B ) (Formula 12)

[0132] Where Start is the start window. A For the first video frame of window A, Start B This is the first video frame of window B.

[0133] End = min(End) A End B ) (Formula 13)

[0134] Where Ended is the end window, End A For the last video frame of window A, End B This is the last video frame of window B.

[0135] Inter=End-Start+1 (Formula 14)

[0136] Where Inter is the intersection of the start window and the end window.

[0137] Score A =Inter / (Length) A +Length B -Inter) (Formula 15)

[0138] Among them, Score A The intersection score of the start and end windows, Length A Length is the length of the starting window. B The window length at which the window ends.

[0139] Score B =|Start A +End A -Start B -End B | / 2 / Length A (Formula 16)

[0140] Among them, Score B The score represents the overlap between the center points of the start and end windows.

[0141] Score C =|Length A -Length B | / Length A(Formula 17)

[0142] Among them, Score C This is the length of the Score window.

[0143] Score = Score A -M*Score B -N*Score C (Formula 18)

[0144] Step 603: The electronic device determines N video segments from the sorted L video segments based on similarity values.

[0145] In this embodiment of the application, if the score calculated according to the above formula is greater than the preset threshold, then window B is removed from the filter window until the filter window no longer contains any windows.

[0146] In this embodiment, the electronic device can further filter the M video segments to obtain N video segments, thereby reducing the number of video segments that the electronic device can recognize and improving the efficiency of the electronic device in recognizing sign language actions.

[0147] Step 203: The electronic device performs sign language recognition on N video segments out of M video segments to obtain the sign language information of the target user.

[0148] In this embodiment of the application, each of the above N video segments contains video content corresponding to the target sign language action, where N is an integer less than or equal to M.

[0149] This application provides a sign language recognition method. An electronic device can acquire hand feature information of a target user from video frames in a target video. Then, based on this hand feature information, it determines M video segments from the target video and performs sign language recognition on N video frames out of the M video frames to obtain the target user's sign language information. In this solution, the electronic device can use the target user's hand feature information to replace skeletal point feature information. That is, the electronic device can abstract complex and high-dimensional skeletal point features into low-dimensional linguistic features. Based on these low-dimensional linguistic features, it can determine M video segments containing sign language actions from the target video. Furthermore, the electronic device can further filter the M video segments to obtain N video segments. Since the electronic device uses low-dimensional linguistic features, it can obtain N video segments containing the target user's sign language actions with less computation, thus improving the efficiency of the electronic device in recognizing sign language actions.

[0150] It should be noted that the sign language recognition method provided in this application can be executed by a sign language recognition device, an electronic device, or a functional module or entity within an electronic device. This application uses a sign language recognition device executing the sign language recognition method as an example to illustrate the sign language recognition device provided in this application.

[0151] Figure 2 A schematic diagram of a possible structure of the sign language recognition device involved in an embodiment of this application is shown. For example... Figure 2 As shown, the sign language recognition device 70 may include: an acquisition module 71, a determination module 72, and a processing module 73.

[0152] The system includes three modules: an acquisition module 71, which acquires hand feature information of the target user from video frames of the target video; a determination module 72, which determines M video segments from the target video based on the hand feature information, where each video segment contains at least one consecutive video frame and includes video content corresponding to similar sign language actions, and M is an integer greater than 1; and a processing module 73, which performs sign language recognition on N video segments out of the M video segments to obtain the target user's sign language information, where each of the N video segments contains video content corresponding to the target sign language actions, and N is an integer less than or equal to M.

[0153] In one possible implementation, the acquisition module 71 is specifically used to acquire the horizontal slope, vertical slope, variance of the horizontal line, and variance of the vertical line of the target user's hand; determine the parallel state of the palm's horizontal line based on the horizontal slope and variance of the horizontal line; determine the parallel state of the palm's vertical line based on the vertical slope and variance of the vertical line; determine the orientation of the target user's hand based on the skeletal point information of the hand; and determine the palm orientation and wrist orientation of the target user based on the parallel state and orientation of the target user's hand; wherein the hand feature information includes: palm orientation and wrist orientation, and the parallel state of the target user's hand includes either the horizontal parallel state or the vertical parallel state.

[0154] In one possible implementation, the acquisition module 71 is specifically used to determine the slope of the horizontal line and the slope of the vertical line of the hand based on the hand shape and skeletal point information; and to determine the variance of the horizontal line of the hand based on the slope of the horizontal line of the hand, and to determine the variance of the vertical line of the hand based on the slope of the vertical line of the hand.

[0155] In one possible implementation, the acquisition module 71 is specifically used to determine the target horizontal reference parallel axis corresponding to the horizontal line of the hand based on the slope of the horizontal line and the variance of the horizontal line of the hand; and to determine the parallel state of the horizontal line of the palm based on the target horizontal reference parallel axis; and to determine the target vertical reference parallel axis corresponding to the vertical line of the hand based on the slope of the vertical line of the hand and the variance of the vertical line of the hand; and to determine the parallel state of the vertical line of the palm based on the target vertical reference parallel axis.

[0156] In one possible implementation, the acquisition module 71 is specifically used to determine a first hand feature corresponding to the parallel state and direction of the target user's hand based on the correspondence relationship; wherein, the correspondence relationship includes the mapping relationship between the parallel state and direction of the target user's hand and the hand feature information, and the first hand feature is a feature determined based on the correspondence relationship in the hand feature information.

[0157] In one possible implementation, the determining module 72 is specifically used to identify at least one video frame in the target video using hand feature information to obtain sign language movement difference information between every two video frames in the at least one video frame; determine I target video frames from the at least one video frame according to the difference category of the sign language movement difference information, where I is an integer greater than 1, and the I target video frames are video frames containing similar sign language movements; and obtain M video segments based on the I target video frames.

[0158] In one possible implementation, the determining module 72 is further configured to, after determining M video segments from the target video based on hand feature information, determine L video segments from the M video segments based on a first value and the length of each video segment, where the first value is the average of the difference programs corresponding to each video segment. The processing module 73 is further configured to, sort the L video segments according to the first value, and calculate the similarity value between the target video segment and the other video segments in the sorted L video segments. The determining module 72 is further configured to, based on the similarity value, determine N video segments from the sorted L video segments.

[0159] This application provides a sign language recognition device. The device can replace skeletal feature information with the hand feature information of the target user. That is, the sign language recognition device can abstract complex and high-dimensional skeletal feature information into low-dimensional linguistic features. Based on these low-dimensional linguistic features, it can determine M video segments containing sign language actions from the target video. Furthermore, the device can further filter these M video segments to obtain N video segments. Since the sign language recognition device uses low-dimensional linguistic features, it can obtain N video segments containing the target user's sign language actions with less computation, thus improving the efficiency of electronic devices in recognizing sign language actions.

[0160] The sign language recognition device in this application embodiment can be a device, or a component, integrated circuit, or chip in an electronic device. The device can be a mobile electronic device or a non-mobile electronic device. For example, a mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0161] The sign language recognition device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0162] The sign language recognition device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0163] Optionally, such as Figure 3As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described sign language recognition method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0164] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0165] Figure 4 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0166] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0167] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0168] The processor 110 is configured to acquire hand feature information of the target user in the video frames of the target video; and based on the hand feature information, determine M video segments from the target video, each video segment containing at least one consecutive video frame, each video segment including video content corresponding to similar sign language actions, where M is an integer greater than 1; and perform sign language recognition on N video segments among the M video segments to obtain the sign language information of the target user, where each of the N video segments contains video content corresponding to the target sign language actions, where N is an integer less than or equal to M.

[0169] This application provides an electronic device that can replace skeletal point feature information with hand feature information of the target user. That is, the electronic device can abstract complex and high-dimensional skeletal point features into low-dimensional linguistic features, and then determine M video segments containing sign language actions from the target video based on the low-dimensional linguistic features. Moreover, the electronic device can further filter the M video segments to obtain N video segments. Since the electronic device uses low-dimensional linguistic features, it can obtain N video segments containing the target user's sign language actions with less computation, thus improving the efficiency of the electronic device in recognizing sign language actions.

[0170] Optionally, in this embodiment of the application, the processor 110 is specifically used to acquire the horizontal slope, vertical slope, variance of the horizontal line, and variance of the vertical line of the target user's hand; determine the parallel state of the palm's horizontal line based on the horizontal slope and variance of the horizontal line; determine the parallel state of the palm's vertical line based on the vertical slope and variance of the vertical line; determine the direction of the target user's hand based on the skeletal point information of the hand; and determine the palm orientation and wrist orientation of the target user's hand based on the parallel state and orientation of the target user's hand; wherein, the hand feature information includes: palm orientation and wrist orientation, and the parallel state of the target user's hand includes either the parallel state of the horizontal line or the parallel state of the vertical line.

[0171] Optionally, in this embodiment of the application, the processor 110 is specifically used to determine the slope of the horizontal line and the slope of the vertical line of the hand based on the hand shape and skeletal point information; to determine the variance of the horizontal line of the hand based on the slope of the horizontal line of the hand; and to determine the variance of the vertical line of the hand based on the slope of the vertical line of the hand.

[0172] Optionally, in this embodiment of the application, the processor 110 is specifically configured to determine the target lateral reference parallel axis corresponding to the hand horizontal line based on the slope of the hand horizontal line and the variance of the hand horizontal line; and determine the parallel state of the palm horizontal line based on the target lateral reference parallel axis; and determine the target longitudinal reference parallel axis corresponding to the hand longitudinal line based on the slope of the hand longitudinal line and the variance of the hand longitudinal line; and determine the parallel state of the palm longitudinal line based on the target longitudinal reference parallel axis.

[0173] Optionally, in this embodiment of the application, the processor 110 is specifically used to determine a first hand feature corresponding to the parallel state of the target user's hand and the direction of the target user's hand based on the correspondence relationship; wherein, the correspondence relationship includes the mapping relationship between the parallel state of the target user's hand and the direction of the target user's hand and the hand feature information, and the first hand feature is a feature determined based on the correspondence relationship in the hand feature information.

[0174] Optionally, in this embodiment of the application, the processor 110 is specifically configured to perform recognition processing on at least one video frame in the target video using hand feature information to obtain sign language action difference information between every two video frames in the at least one video frame; determine I target video frames from the at least one video frame according to the difference category of the sign language action difference information, wherein the I target video frames are video frames containing similar sign language actions, and I is an integer greater than 1; and obtain M video segments based on the I target video frames.

[0175] Optionally, in this embodiment of the application, the processor 110 is further configured to: determine M video segments from the target video based on hand feature information; determine L video segments from the M video segments based on a first value of each video segment and the length of each video segment, wherein the first value is the average value of the difference program corresponding to each video segment; sort the L video segments according to the size of the first value; calculate the similarity value between the target video segment in the sorted L video segments and the other video segments in the L video segments; and determine N video segments from the sorted L video segments based on the similarity value.

[0176] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0177] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.

[0178] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0179] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0180] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0181] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0182] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0183] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0184] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0185] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described sign language recognition method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0186] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0187] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0188] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A sign language recognition method, characterized in that, The method includes: Obtain the hand feature information of the target user in the video frames of the target video; Based on the hand feature information, M video segments are determined from the target video. Each video segment contains at least one consecutive video frame and includes video content with similar sign language movements. M is an integer greater than 1. Sign language recognition is performed on N video segments out of the M video segments to obtain the sign language information of the target user. Each of the N video segments contains video content corresponding to the target sign language action, where N is an integer less than or equal to M. The step of obtaining the hand feature information of the target user in at least one video frame of the target video includes: Obtain the slope of the horizontal line, the slope of the vertical line, the variance of the horizontal line, and the variance of the vertical line of the hand for the target user. The parallel state of the palm horizontal lines is determined based on the slope of the horizontal lines and the variance of the horizontal lines. Based on the slope of the vertical line of the hand and the variance of the vertical line of the hand, the parallel state of the vertical line of the palm is determined; Based on the skeletal point information of the hand, the orientation of the target user's hand is determined; Based on the parallel state of the target user's hand and the direction of the target user's hand, determine the palm orientation and wrist orientation of the target user's hand; The hand feature information includes: palm orientation and wrist orientation, and the parallel state of the target user's hand includes either the horizontal parallel state of the palm or the vertical parallel state of the palm.

2. The method according to claim 1, characterized in that, The acquisition of the horizontal line slope, vertical line slope, horizontal line variance, and vertical line variance of the target user's hand includes: Based on the hand shape and the skeletal point information, determine the slope of the horizontal line and the slope of the vertical line of the hand; The variance of the horizontal line of the hand is determined based on the slope of the horizontal line, and the variance of the vertical line of the hand is determined based on the slope of the vertical line.

3. The method according to claim 1, characterized in that, The step of determining the parallelism of the palm's horizontal lines based on the slope of the horizontal lines and the variance of the horizontal lines includes: Based on the slope of the horizontal hand line and the variance of the horizontal hand line, determine the target lateral reference parallel axis corresponding to the horizontal hand line; The parallel state of the palm's horizontal line is determined based on the target's lateral reference parallel axis; Determining the parallelism of the palm's vertical lines based on the slope of the hand's vertical lines and the variance of the hand's vertical lines includes: Based on the slope of the hand longitudinal line and the variance of the hand longitudinal line, determine the target longitudinal reference parallel axis corresponding to the hand longitudinal line; Based on the target longitudinal reference parallel axis, the parallel state of the palm's longitudinal line is determined.

4. The method according to claim 1, characterized in that, The step of obtaining the palm orientation and wrist orientation of the target user's hand based on the parallel state and orientation of the target user's hand includes: Based on the correspondence, a first hand feature corresponding to the parallel state of the target user's hand and the direction of the target user's hand is determined; The correspondence includes the mapping relationship between the parallel state of the target user's hand and the direction of the target user's hand and the hand feature information, wherein the first hand feature is a feature determined based on the correspondence in the hand feature information.

5. The method according to claim 1, characterized in that, The step of determining M video segments from the target video based on the hand feature information includes: Using the hand feature information, at least one video frame in the target video is identified to obtain sign language movement difference information between every two video frames in the at least one video frame; Based on the difference category of the sign language movement difference information, I target video frames are determined from the at least one video frame, where the I target video frames are video frames containing similar sign language movements, and I is an integer greater than 1; Based on the I target video frames, M video segments are obtained.

6. The method according to claim 1, characterized in that, After determining M video segments from the target video based on the hand feature information, the method further includes: Based on the first value of each video segment in the M video segments and the video segment length of each video segment, L video segments are determined from the M video segments, where the first value is the average value of the difference program corresponding to each video segment; According to the first value, the L video segments are sorted, and the similarity value between the target video segment in the sorted L video segments and the other video segments in the L video segments is calculated. Based on the similarity value, N video segments are determined from the sorted L video segments.

7. A sign language recognition device, characterized in that, The sign language recognition device includes: an acquisition module, a determination module, and a processing module; The acquisition module is used to acquire hand feature information of the target user in the video frame of the target video; The determining module is used to determine M video segments from the target video based on the hand feature information. Each video segment contains at least one consecutive video frame and includes video content with similar sign language actions. M is an integer greater than 1. The processing module is used to perform sign language recognition on N video segments out of the M video segments to obtain the sign language information of the target user. Each of the N video segments contains video content corresponding to the target sign language action, and N is an integer less than or equal to M. The acquisition module is specifically used for The system acquires the horizontal and vertical slopes of the target user's hand, as well as the variances of the horizontal and vertical lines. Based on the horizontal slopes and variances, it determines the parallelism of the palm's horizontal lines; and based on the vertical slopes and variances, it determines the parallelism of the palm's vertical lines. It also determines the orientation of the target user's hand based on the skeletal point information of the hand; and based on the parallelism of the target user's hand and the orientation of the target user's hand, it determines the palm orientation and wrist orientation of the target user. The hand feature information includes: palm orientation and wrist orientation, and the parallel state of the target user's hand includes either the horizontal parallel state of the palm or the vertical parallel state of the palm.

8. The apparatus according to claim 7, characterized in that, The acquisition module is specifically used to determine the slope of the horizontal line and the slope of the vertical line of the hand based on the hand shape and the skeletal point information; and to determine the variance of the horizontal line of the hand based on the slope of the horizontal line of the hand, and to determine the variance of the vertical line of the hand based on the slope of the vertical line of the hand.

9. The apparatus according to claim 7, characterized in that, The acquisition module is specifically used to determine the target lateral reference parallel axis corresponding to the hand horizontal line based on the slope of the hand horizontal line and the variance of the hand horizontal line; and to determine the parallel state of the palm horizontal line based on the target lateral reference parallel axis. And based on the slope of the hand longitudinal line and the variance of the hand longitudinal line, determine the target longitudinal reference parallel axis corresponding to the hand longitudinal line; and based on the target longitudinal reference parallel axis, determine the parallel state of the palm longitudinal line.

10. The apparatus according to claim 7, characterized in that, The acquisition module is specifically used to determine a first hand feature corresponding to the parallel state of the target user's hand and the direction of the target user's hand based on a correspondence relationship; wherein, the correspondence relationship includes the mapping relationship between the parallel state of the target user's hand and the direction of the target user's hand and hand feature information, and the first hand feature is a feature determined based on the correspondence relationship in the hand feature information.

11. The apparatus according to claim 7, characterized in that, The determining module is specifically configured to: identify at least one video frame in the target video using the hand feature information to obtain sign language movement difference information between every two video frames in the at least one video frame; determine I target video frames from the at least one video frame based on the difference category of the sign language movement difference information, wherein the I target video frames are video frames containing similar sign language movements, and I is an integer greater than 1; and obtain M video segments based on the I target video frames.

12. The apparatus according to claim 7, characterized in that, The determining module is further configured to determine M video segments from the target video based on the hand feature information, and then determine L video segments from the M video segments based on the first value of each video segment and the video segment length of each video segment, wherein the first value is the average value of the difference program corresponding to each video segment; The processing module is further configured to sort the L video segments according to the first value, and calculate the similarity value between the target video segment in the sorted L video segments and the other video segments in the L video segments respectively; The determining module is further configured to determine N video segments from the sorted L video segments based on the similarity values.

13. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the sign language recognition method as described in any one of claims 1 to 6.

14. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the sign language recognition method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Artificial intelligence-based sign language recognition method and system

    CN114708648A