A vehicle-mounted sign language recognition method, a vehicle-mounted terminal and a vehicle
By extracting keyframe images from video information and recognizing human posture and gestures through the in-vehicle terminal, and using a combined semantic model to output speech or text, the problem of inaccurate sign language recognition is solved, and the safety of in-vehicle interaction is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-28
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, sensor gloves cannot combine hand gestures with movements of other parts of the body for semantic recognition, resulting in inaccurate sign language recognition.
By capturing keyframe images of video information of people inside the vehicle through the vehicle terminal, dividing the human body parts into regions, extracting posture feature points and hand joint feature points, and using a pre-trained combined semantic model to comprehensively recognize gestures and body postures, the system outputs voice or text information.
It achieves more accurate sign language recognition, improving the safety of interactions during driving.
Smart Images

Figure CN115546842B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a vehicle-mounted sign language recognition method, a vehicle-mounted terminal, and a vehicle. Background Technology
[0002] In existing technologies, since drivers cannot see sign language while driving, it is necessary to convert the sign language of other passengers in the vehicle into speech and other information output. When performing sign language recognition inside the vehicle, sensors in various finger parts of an external sensing glove are generally used to convert the user's finger movements into electrical signals, and the semantic information of sign language is obtained by analyzing the results of the electrical signals. Since the sensing glove can only recognize the user's hand gestures, it cannot combine the movements of other parts of the user's body to perform semantic recognition.
[0003] It is evident that existing technologies often encounter the problem of being unable to combine gestures with movements of other parts of the body for semantic recognition when performing sign language recognition, resulting in insufficient accuracy in recognizing users' sign language. Summary of the Invention
[0004] This invention proposes a vehicle-mounted sign language recognition method, a vehicle-mounted terminal, and a vehicle. After capturing key frame images from video information of people inside the vehicle through the vehicle-mounted terminal, human posture recognition results and gesture recognition results are obtained. Then, the human posture recognition results and the gesture recognition results are input into a trained combined semantic model to calculate semantic results, thereby achieving more accurate sign language recognition.
[0005] This invention provides a vehicle-mounted sign language recognition method, including:
[0006] The system acquires video information of people inside the vehicle through an in-vehicle terminal and extracts key frame images from the video information of people inside the vehicle.
[0007] The keyframe image is divided into at least two regions according to different parts of the human body, and the human pose feature points contained in each region constitute a pose dataset.
[0008] The Gini index is calculated for all pose datasets to obtain the output action that corresponds one-to-one with each pose dataset. All output actions constitute the human pose recognition result.
[0009] The palm region is identified from the keyframe image, and palm joint feature points are obtained;
[0010] The gesture with the highest similarity score to the palm joint feature point is searched from the preset gesture database and used as the gesture recognition result;
[0011] The human posture recognition result and the gesture recognition result are input into a pre-trained combined semantic model to calculate the semantic result.
[0012] Furthermore, after inputting the human pose recognition result and the gesture recognition result into the pre-trained combined semantic model to calculate the semantic result, the method further includes: comparing the semantic result with the semantic result obtained from the next keyframe image; if there are differences in the information keywords between the semantic results, then the semantic result obtained from the next keyframe image is corrected; wherein, the information keywords of the semantic result include the subject and predicate of the statement.
[0013] Furthermore, the Gini index is calculated for all pose datasets to obtain the output action corresponding to each pose dataset. All output actions constitute the human pose recognition result. Specifically, the Gini index is calculated for all pose datasets to obtain the probability of at least one action corresponding to the same pose dataset. The action with the highest probability is selected as the output action. All output actions corresponding to each pose dataset constitute the human pose recognition result.
[0014] Furthermore, the human posture recognition result and the gesture recognition result are input into a pre-trained combined semantic model to calculate the semantic result. Specifically, the human posture recognition result and the gesture recognition result are input into a pre-trained combined semantic model to calculate the similarity with different output results, and the output result with the highest similarity is selected as the semantic result.
[0015] Furthermore, video information of people inside the vehicle is obtained through the vehicle terminal, and key frame images are extracted from the video information of people inside the vehicle. Specifically, the video information of people inside the vehicle is obtained through the vehicle terminal; the inter-frame difference is obtained by performing a difference operation on two adjacent frames of the video information of people inside the vehicle; if the inter-frame difference is greater than a preset threshold, the next frame of the two adjacent frames is used as the key frame image.
[0016] Furthermore, after obtaining the semantic results, voice information is output through the voice device configured in the vehicle terminal; or, text information is output through the display device configured in the vehicle terminal.
[0017] This invention also provides an in-vehicle terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the above-described in-vehicle sign language recognition method.
[0018] This invention also provides a vehicle equipped with the aforementioned vehicle-mounted terminal.
[0019] The present invention has the following beneficial effects:
[0020] In this embodiment of the invention, keyframe images are extracted from the acquired video information of people inside the vehicle via an in-vehicle terminal. The keyframe images are divided into regions according to human body parts, and the human posture feature points contained in each region constitute a posture dataset. The Gini index is calculated on the posture dataset to obtain the human posture recognition result. The palm region is identified from the keyframe images and the palm joint feature points are obtained. Then, the gesture with the highest similarity score to the palm joint feature points is searched from a preset gesture database and used as the gesture recognition result. The human posture recognition result and the gesture recognition result are input into a pre-trained combined semantic model to calculate the semantic result, and the corresponding voice information or text information is output through the in-vehicle terminal. In this embodiment of the invention, the human posture recognition results and gesture recognition results obtained separately are input into a pre-trained combined semantic model to calculate the similarity to different output results. The output result with the highest similarity is selected as the semantic result. Thus, when performing in-vehicle sign language recognition, the gestures and body postures of the recognition object, including limb movements and facial movements, can be comprehensively considered. This results in more accurate in-vehicle sign language recognition, and the recognition results are output in the form of voice or text, thereby improving the safety of interaction between drivers and passengers during driving. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating an embodiment of the vehicle-mounted sign language recognition method provided by the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] See Figure 1 This is a flowchart illustrating an embodiment of the vehicle-mounted sign language recognition method provided by the present invention. The method includes steps S1 to S6, as detailed below:
[0024] S1, acquire video information of people inside the vehicle through the vehicle terminal, and extract key frame images from the video information of people inside the vehicle;
[0025] Furthermore, video information of people inside the vehicle is obtained through the vehicle terminal; a difference operation is performed on two adjacent frames of the video information of people inside the vehicle to obtain the inter-frame difference value. If the inter-frame difference value is greater than a preset threshold, the next frame of the two adjacent frames is used as the key frame image.
[0026] Specifically, one embodiment is as follows: the vehicle terminal includes all electronic devices inside the vehicle, wherein a monocular camera is installed in front of the passenger seat, the right side of the rear seat, and the left side of the rear seat. The camera records video information of the people in the vehicle and transmits it to the central control system of the vehicle terminal. The central control system performs a difference operation on two adjacent frames of the video information of the people in the vehicle to obtain the inter-frame difference. If the inter-frame difference is greater than a preset threshold, it means that there is a dynamic change caused by the movement of people in the video during the time period represented by the two frames. Compared with the previous frame image, which is a static image, the subsequent frame image records the movement of people and is selected as the key frame image.
[0027] S2, the keyframe image is divided into at least two regions according to different parts of the human body, and the human pose feature points contained in each region constitute a pose dataset.
[0028] Specifically, the keyframe image is divided into at least two regions according to different parts of the human body in the keyframe image, such as the face and limbs. The number of regions is determined by the number of human body parts captured in each keyframe image. All human pose feature points contained in each region are treated as a pose dataset. One embodiment of the human pose feature points is the feature points of facial expressions or the feature points of limb joints extracted using the Mediapipe algorithm.
[0029] S3, calculate the Gini index for all pose datasets to obtain the output action that corresponds to each pose dataset. All output actions constitute the human pose recognition result.
[0030] Furthermore, the Gini index is calculated for all pose datasets to obtain the probability of at least one action corresponding to the same pose dataset, and the action with the highest probability is selected as the output action; all output actions that correspond one-to-one with each pose dataset constitute the human pose recognition result.
[0031] Specifically, the Gini index is calculated for all pose datasets using the following formula, where each pose dataset corresponds to K actions, and the probability of a pose dataset belonging to the k-th action is p. k The probability of each action corresponding to each pose dataset is calculated using the following formula. The action with the highest probability is selected as the output action. The human pose recognition result is composed of the output actions that correspond one-to-one with each pose dataset.
[0032]
[0033] S4, Identify the palm region from the keyframe image and obtain palm joint feature points;
[0034] Specifically, one particular embodiment is as follows: after identifying the palm region from the keyframe image, the SSD algorithm (Single Shot MultiBox Detector algorithm) is used to locate key points of the palm joints, thereby obtaining the palm joint feature points.
[0035] S5. Search the preset gesture database for the gesture action with the highest similarity score to the palm joint feature point, and use it as the gesture recognition result.
[0036] S6. Input the human posture recognition result and the gesture recognition result into the pre-trained combined semantic model to calculate the semantic result.
[0037] Furthermore, the human posture recognition result and the gesture recognition result are input into a pre-trained combined semantic model to calculate the similarity to different output results, and the output result with the highest similarity is selected as the semantic result.
[0038] Specifically, the combined semantic model is trained using a semi-supervised learning tri-training method.
[0039] Furthermore, after comparing the semantic result with the semantic result obtained from the next keyframe image, if there are differences in the information keywords between the semantic results, the semantic result obtained from the next keyframe image is corrected; wherein, the information keywords of the semantic result include the subject and predicate of the statement.
[0040] Furthermore, after obtaining the semantic results, voice information is output through the voice device configured in the vehicle terminal; or, text information is output through the display device configured in the vehicle terminal.
[0041] Accordingly, this embodiment of the invention also provides an in-vehicle terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the above-described in-vehicle sign language recognition method.
[0042] In addition, this invention also provides a vehicle equipped with the above-mentioned vehicle-mounted terminal.
[0043] Implementing the embodiments of the present invention has the following beneficial effects:
[0044] In this embodiment of the invention, keyframe images are extracted from the acquired video information of people inside the vehicle via an in-vehicle terminal. The keyframe images are divided into regions according to human body parts, and the human posture feature points contained in each region constitute a posture dataset. The Gini index is calculated on the posture dataset to obtain the human posture recognition result. The palm region is identified from the keyframe images and the palm joint feature points are obtained. Then, the gesture with the highest similarity score to the palm joint feature points is searched from a preset gesture database and used as the gesture recognition result. The human posture recognition result and the gesture recognition result are input into a pre-trained combined semantic model to calculate the semantic result, and the corresponding voice information or text information is output through the in-vehicle terminal. In this embodiment of the invention, the human posture recognition results and gesture recognition results obtained separately are input into a pre-trained combined semantic model to calculate the similarity to different output results. The output result with the highest similarity is selected as the semantic result. Thus, when performing in-vehicle sign language recognition, the gestures and body postures of the recognition object, including limb movements and facial movements, can be comprehensively considered. This results in more accurate in-vehicle sign language recognition, and the recognition results are output in the form of voice or text, thereby improving the safety of interaction between drivers and passengers during driving.
[0045] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary hardware platforms, and of course, it can also be implemented entirely by hardware. Based on this understanding, all or part of the technical solution of the present invention that contributes to the background art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0046] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A vehicle-mounted sign language recognition method, characterized in that, include: The video information of the people inside the vehicle is obtained through the vehicle terminal, and key frame images are extracted from the video information of the people inside the vehicle. The keyframe image is divided into at least two regions according to different parts of the human body, and the human pose feature points contained in each region constitute a pose dataset. The Gini index is calculated for all pose datasets to obtain the output action that corresponds one-to-one with each pose dataset. All output actions constitute the human pose recognition result. The palm region is identified from the keyframe image, and palm joint feature points are obtained; The gesture with the highest similarity score to the palm joint feature point is searched from the preset gesture database and used as the gesture recognition result; The human posture recognition result and the gesture recognition result are input into a pre-trained combined semantic model to calculate the semantic result; Specifically, calculating the Gini index for all pose datasets yields output actions corresponding to each pose dataset, and all output actions constitute the human pose recognition result. For all pose datasets, calculate the Gini index according to the following formula to obtain the probability of at least one action corresponding to the same pose dataset, and select the action with the highest probability as the output action; all output actions that correspond one-to-one with each pose dataset constitute the human pose recognition result. Each pose dataset corresponds to K actions, and the probability that a pose dataset belongs to the k-th action is: .
2. The vehicle-mounted sign language recognition method as described in claim 1, characterized in that, After inputting the human pose recognition result and the gesture recognition result into a pre-trained combined semantic model to calculate the semantic result, the method further includes: After comparing the semantic result with the semantic result obtained from the next keyframe image, if there are differences in the information keywords between the semantic results, the semantic result obtained from the next keyframe image is corrected; wherein, the information keywords of the semantic result include the subject and predicate of the statement.
3. The vehicle-mounted sign language recognition method as described in claim 1, characterized in that, The step of inputting the human posture recognition result and the gesture recognition result into a pre-trained combined semantic model to calculate the semantic result is as follows: The human pose recognition result and the gesture recognition result are input into a pre-trained combined semantic model to calculate the similarity with different output results, and the output result with the highest similarity is selected as the semantic result.
4. The vehicle-mounted sign language recognition method as described in claim 1, characterized in that, The step of acquiring video information of occupants inside the vehicle via an in-vehicle terminal and extracting keyframe images from the video information of occupants inside the vehicle specifically involves: Obtain video information of people inside the vehicle through the vehicle-mounted terminal; The inter-frame difference is obtained by performing a differential operation on two adjacent frames of the video information of the people inside the vehicle. If the inter-frame difference is greater than a preset threshold, the next frame of the two adjacent frames is used as the key frame image.
5. The vehicle-mounted sign language recognition method as described in claim 1, characterized in that, Also includes: After obtaining the semantic results, the voice information is output through the voice device configured in the vehicle terminal; Alternatively, text information can be output through the display device configured in the vehicle terminal.
6. A vehicle-mounted terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the vehicle-mounted sign language recognition method according to any one of claims 1 to 5.
7. A vehicle, characterized in that, The vehicle is equipped with the vehicle-mounted terminal as described in claim 6.
Citation Information
Patent Citations
Gesture language teaching method, device and system based on gesture action generation and recognition
CN114842547A
Environmental semantic understanding-based body movement recognition method, apparatus, device, and storage medium
WO2021114892A1