Vehicle-mounted sign language recognition model training method and intelligent vehicle control method and device

By training the on-board sign language recognition model, using video stream data and key point label data to generate sign language semantic text sequences, the problem of insufficient sign language recognition accuracy for deaf and dumb people is solved, and more efficient vehicle-machine command interaction is achieved.

CN120564262AActive Publication Date: 2025-08-29CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510659306.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-29
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing intelligent vehicle control system fails to effectively consider the use of people with language barriers, especially those with deaf and mute people in non-main driving positions, which leads to low accuracy of vehicle-machine command interaction.

Method used

By obtaining the training sample data set, including original video stream data, hand key point label data, shoulder hand key point label data and sign language semantic text sequence label data, the vehicle sign language recognition model is trained, and the key point detection sub-model and sign language recognition sub-model are used to extract feature and capture gesture change trajectory, generate sign language semantic text sequence data, and control the execution of instructions of intelligent vehicles.

Benefits of technology

It improves the recognition accuracy of complex sign language, improves the accuracy of vehicle-machine command interaction, and realizes sign language recognition for deaf and dumb people in non-main driving positions in the cockpit of the intelligent vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120564262A_ABST
    Figure CN120564262A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle-mounted sign language recognition model training method and an intelligent vehicle control method and device, relates to the technical field of intelligent control, and aims to solve the technical problem that the recognition precision of complex sign languages is low due to the fact that tiny dynamic features of fingers are difficult to capture through a sensor in the prior art. The sign language semantic text sequence data is obtained by inputting the video data stream into the vehicle-mounted sign language recognition model, the control instruction is generated based on the sign language semantic text sequence data, and the intelligent vehicle is controlled to execute the control instruction, so that the precision of complex sign language recognition is improved, and the accuracy of vehicle instruction interaction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent control technology, and in particular to a vehicle-mounted sign language recognition model training method, an intelligent vehicle control method and device. Background Art

[0002] With the progress and development in the field of intelligent control technology, the functions of smart cars are becoming more and more abundant. Thanks to the development of machine learning and deep learning, intelligent vehicle control technology based on speech recognition is becoming more and more mature and accurate. However, the existing intelligent vehicle control system is obviously not considerate enough for people with language barriers and other problems, and does not consider the use problems of people with language barriers and other problems. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide an in-vehicle sign language recognition model training method, an intelligent vehicle control method and device to improve the accuracy of complex sign language recognition.

[0004] In a first aspect, the present invention provides a method for training an in-vehicle sign language recognition model, comprising:

[0005] Acquire a training sample data set; wherein the training sample data set includes a plurality of training sample data; each training sample data includes original video stream data and hand key point label data, shoulder-hand key point label data, and sign language semantic text sequence label data corresponding to the sign language action in the original video stream data;

[0006] Based on the training sample data set, a training operation is iteratively performed on the in-vehicle sign language recognition model; wherein the in-vehicle sign language recognition model includes a key point detection sub-model and a sign language recognition sub-model with skip connections; the training operation includes:

[0007] Select target training sample data from the training sample data set;

[0008] The original video stream data in the target training sample data is input into the in-vehicle sign language recognition model, so that the in-vehicle sign language recognition model performs feature extraction and key point detection on multiple frames of original video frame images in the original video stream data through the key point detection sub-model to obtain hand key point feature data and shoulder-hand key point feature data, and the hand key point feature data and shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of original video frame images are captured by the sign language recognition sub-model through gesture change trajectory capture and sign language action semantic analysis to obtain sign language semantic text sequence prediction data;

[0009] Based on the sign language semantic text sequence prediction data, hand key point feature data and shoulder-hand key point feature data, as well as the sign language semantic text sequence label data, hand key point label data and shoulder-hand key point label data in the target training sample data, the current loss value is calculated, and the first weight parameter of the key point detection sub-model and the second weight parameter of the sign language recognition sub-model are updated based on the current loss value.

[0010] Optionally, calculating the current loss value based on the sign language semantic text sequence prediction data, the hand key point feature data, and the shoulder-hand key point feature data, as well as the sign language semantic text sequence label data, the hand key point label data, and the shoulder-hand key point label data in the target training sample data includes:

[0011] Based on the hand key point feature data and the hand key point label data in the target training sample data, the target loss function is used to calculate the hand loss value;

[0012] Based on the shoulder-hand key point feature data and the shoulder-hand key point label data in the target training sample data, the target loss function is used to calculate the shoulder-hand loss value;

[0013] Based on the hand loss value and the shoulder-hand loss value, a feature loss value is obtained;

[0014] Based on the sign language semantic text sequence prediction data and sign language semantic text sequence label data, the connection temporal classification loss function is used to calculate the sequence loss value;

[0015] Calculate the current loss value based on the feature loss value and sequence loss value.

[0016] Optionally, the key point detection sub-model includes a first local feature extraction module, a first multi-scale feature extraction module, a first bidirectional feature extraction module, a multi-scale feature fusion module, and a key point detection module connected in sequence; the key point detection sub-model performs feature extraction and key point detection on the original video stream data to obtain hand key point feature data and shoulder-hand key point feature data, including:

[0017] Preprocessing and feature extraction are performed on the original video stream data by a first local feature extraction module to obtain a first key point feature image;

[0018] Performing multi-scale feature extraction on features of the first key point feature image using a first multi-scale feature extraction module to obtain a second key point feature image;

[0019] Performing temporal feature extraction on features of the second key point feature image using the first bidirectional feature extraction module to obtain a third key point feature image;

[0020] The third key point feature image is extracted and fused through a multi-scale feature fusion module to obtain a feature fusion image;

[0021] The key point detection module is used to perform key point detection on the feature fusion image to obtain the hand key point feature data and the shoulder-hand key point feature data.

[0022] Optionally, the first local feature extraction module includes a first 3D convolution layer; the first multi-scale feature extraction module includes a plurality of first multi-scale convolution units connected in sequence, the first multi-scale convolution unit includes a first hole convolution layer, a first convolution layer, a second convolution layer connected in parallel, and a first splicing layer connected to the first hole convolution layer, the first convolution layer and the second convolution layer respectively, and a first compression-excitation attention layer jump-connected to the first splicing layer; the first bidirectional feature extraction module includes two first bidirectional long short-term memory layers connected in sequence; the multi-scale feature fusion module includes a multi-branch convolution layer; the key point detection module includes two first fully connected layers; the key point detection sub-model is used to perform feature extraction and key point detection on the original video stream data to obtain hand key point feature data and shoulder-hand key point feature data, including:

[0023] Preprocessing and feature extraction are performed on the original video stream data through the first 3D convolution layer in the first local feature extraction module to obtain a first key point feature image;

[0024] Performing multi-scale feature extraction on features of the first key point feature image using first multi-scale convolution units in the plurality of first multi-scale feature extraction modules to obtain a second key point feature image;

[0025] Performing temporal feature extraction on features of the second key point feature image using two first bidirectional long short-term memory layers in the first bidirectional feature extraction module to obtain a third key point feature image;

[0026] The third key point feature image is extracted and fused through the multi-branch convolution layer in the multi-scale feature fusion module to obtain a feature fusion image;

[0027] The first fully connected layer in the key point detection module is used to perform key point detection on the feature fusion image to obtain the hand key point feature data and the shoulder-hand key point feature data.

[0028] Optionally, the key point detection sub-model also includes a random multi-frame feature extraction module; the random multi-frame feature extraction module is connected to the first local feature extraction module, and is used to group and extract the original video stream data to obtain multiple frames of original video frame images.

[0029] Optionally, the sign language recognition sub-model includes a second local feature extraction module, a second multi-scale feature extraction module, a temporal feature extraction module, a second bidirectional feature extraction module, and a sign language action semantic parsing module connected in sequence; the sign language recognition sub-model performs gesture change trajectory capture and sign language action semantic parsing on the original video stream data and the hand key point feature data and the shoulder-hand key point feature data corresponding to the multiple frames of original video frame images to obtain sign language semantic text sequence prediction data, including:

[0030] The second local feature extraction module captures the hand key point feature data and the shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of original video frame images to perform gesture change trajectory capture, thereby obtaining a first gesture feature image;

[0031] Performing multi-scale feature extraction on the features of the gesture feature image by a second multi-scale feature extraction module to obtain a second gesture feature image;

[0032] Extracting time sequence information from the second gesture feature image using a time sequence feature extraction module to obtain a third gesture feature image;

[0033] Performing temporal feature extraction on the third gesture feature image through the second bidirectional feature extraction module to obtain a fused image;

[0034] The sign language action semantic parsing module is used to perform sign language action semantic parsing on the fused image to obtain sign language semantic text sequence prediction data.

[0035] Optionally, the second local feature extraction module includes a second 3D convolution layer; the second multi-scale feature extraction module includes a plurality of second multi-scale convolution units connected in sequence, the second multi-scale convolution unit includes a second hole convolution layer, a third convolution layer, and a fourth convolution layer connected in parallel, and a second splicing layer connected to the second hole convolution layer, the third convolution layer and the fourth convolution layer respectively, and a second compression incentive attention layer jump-connected to the second splicing layer; the temporal feature extraction module includes a time-space convolution layer; the second bidirectional feature extraction module includes a second bidirectional long short-term memory layer; the sign language action semantic parsing module includes a second fully connected layer; the hand key point feature data and the shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of original video frame images are captured by the sign language recognition sub-model through gesture change trajectory capture and sign language action semantic parsing to obtain sign language semantic text sequence prediction data, including:

[0036] The second 3D convolution layer in the second local feature extraction module captures the hand key point feature data and the shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of original video frame images to obtain a first gesture feature image;

[0037] Performing multi-scale feature extraction on features of the gesture feature image using a second multi-scale convolution unit in a second multi-scale feature extraction module to obtain a second gesture feature image;

[0038] Extracting temporal information from the second gesture feature image through the time-space convolution layer in the temporal feature extraction module to obtain a third gesture feature image;

[0039] Performing temporal feature extraction on the third gesture feature image using the second bidirectional long short-term memory layer in the second bidirectional feature extraction module to obtain a fused image;

[0040] The second fully connected layer in the sign language action semantic parsing module performs sign language action semantic parsing on the fused image to obtain sign language semantic text sequence prediction data.

[0041] In a second aspect, the present invention provides an intelligent vehicle control method, comprising:

[0042] Obtain video stream data inside the smart vehicle cabin;

[0043] Inputting the video stream data into an in-vehicle sign language recognition model to obtain sign language semantic text sequence data corresponding to the video stream data; wherein the in-vehicle sign language recognition model is trained using the above-mentioned in-vehicle sign language recognition model training method;

[0044] Based on the sign language semantic text sequence data, control instructions are generated and the intelligent vehicle is controlled to execute the control instructions.

[0045] In a third aspect, the present invention provides an in-vehicle sign language recognition model training device, comprising:

[0046] A data acquisition module is configured to acquire a training sample data set; wherein the training sample data set includes a plurality of training sample data; each training sample data includes original video stream data and hand key point label data, shoulder-hand key point label data, and sign language semantic text sequence label data corresponding to the sign language action in the original video stream data;

[0047] The model training module is used to iteratively perform training operations on the in-vehicle sign language recognition model based on the training sample data set; wherein the in-vehicle sign language recognition model includes a key point detection sub-model and a sign language recognition sub-model with jump connections; the training operation includes: selecting target training sample data from the training sample data set; inputting the original video stream data in the target training sample data into the in-vehicle sign language recognition model, so that the in-vehicle sign language recognition model performs feature extraction and key point detection on multiple frames of original video frame images in the original video stream data through the key point detection sub-model to obtain hand key point feature data and shoulder-hand key point feature data, and The sign language recognition sub-model captures the hand key point feature data and the shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of original video frame images, and performs semantic analysis of sign language movements to obtain sign language semantic text sequence prediction data; based on the sign language semantic text sequence prediction data, the hand key point feature data and the shoulder-hand key point feature data, and the sign language semantic text sequence label data, the hand key point label data and the shoulder-hand key point label data in the target training sample data, the current loss value is calculated, and the first weight parameter of the key point detection sub-model and the second weight parameter of the sign language recognition sub-model are updated based on the current loss value.

[0048] In a fourth aspect, the present invention provides an intelligent vehicle control device, comprising:

[0049] An acquisition module, used to acquire video stream data in the cockpit of an intelligent vehicle;

[0050] A recognition module, configured to input the video stream data into an in-vehicle sign language recognition model to obtain sign language semantic text sequence data corresponding to the video stream data; wherein the in-vehicle sign language recognition model is trained using the in-vehicle sign language recognition model training method described above;

[0051] The control module is used to generate control instructions based on sign language semantic text sequence data and control the intelligent vehicle to execute the control instructions.

[0052] In a fifth aspect, the present invention also provides an electronic device comprising: a processor and a memory; the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the above-mentioned in-vehicle sign language recognition model training method or the above-mentioned intelligent vehicle control method is implemented.

[0053] In a sixth aspect, the present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the above-mentioned vehicle-mounted sign language recognition model training method or the above-mentioned intelligent vehicle control method.

[0054] The embodiments of the present invention provide a vehicle-mounted sign language recognition model training method, an intelligent vehicle control method and a device. By inputting a video data stream into the vehicle-mounted sign language recognition model to obtain sign language semantic text sequence data, control instructions are generated based on the sign language semantic text sequence data, and the intelligent vehicle is controlled to execute the control instructions, thereby improving the accuracy of complex sign language recognition and improving the accuracy of vehicle-computer command interaction.

[0055] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0057] Figure 1 A flow chart of a vehicle-mounted sign language recognition model training method provided by an embodiment of the present invention is shown;

[0058] Figure 2 A schematic diagram showing a flow chart of a method for obtaining a training sample data set provided by an embodiment of the present invention;

[0059] Figure 3 A schematic diagram of the hand key point feature data structure provided by an embodiment of the present invention is shown;

[0060] Figure 4 A schematic diagram of the shoulder-hand key point feature data structure provided by an embodiment of the present invention is shown;

[0061] Figure 5 A schematic structural diagram of an in-vehicle sign language recognition model provided by an embodiment of the present invention is shown;

[0062] Figure 6 A schematic diagram of the key point detection sub-model structure provided by an embodiment of the present invention is shown;

[0063] Figure 7 A schematic diagram of a feature dimension expansion structure provided by an embodiment of the present invention is shown;

[0064] Figure 8 A schematic flow chart illustrating a method for calculating the loss of an in-vehicle sign language recognition model provided by an embodiment of the present invention is shown;

[0065] Figure 9 A schematic diagram showing a flow chart of an intelligent vehicle control method provided by an embodiment of the present invention;

[0066] Figure 10 A schematic structural diagram of a vehicle-mounted sign language recognition model training device provided by an embodiment of the present invention is shown;

[0067] Figure 11 A schematic structural diagram of an intelligent vehicle control device provided by an embodiment of the present invention is shown;

[0068] Figure 12 A schematic structural diagram of an electronic device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0069] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.

[0070] At present, there are mainly two methods in the technology related to sign language recognition: the first is to obtain the motion data of sign language by wearing sensors on the hands (such as accelerometers, gyroscopes, etc.), and use deep learning models to recognize the motion data of sign language. This method can accurately capture the motion trajectory and deformation of gestures, and is not affected by ambient light sources, nor is it easily affected by individual differences such as age, body shape and other factors. However, this method also has certain defects, such as wearing affects the user experience, the equipment cost is high, and it is difficult to capture the subtle dynamic features of fingers, resulting in low recognition accuracy for complex sign language; the second is based on deep learning. A sign language recognition method based on Xi's experience is proposed. This method uses a camera to capture a video sequence of an individual's sign language movements and uses a deep learning model to detect and recognize sign language. This solution has strong feature extraction and pattern recognition capabilities and can automatically learn high-level semantic features of sign language from large amounts of data. However, this method also has certain defects: it is easily affected by factors such as ambient lighting changes, occlusion, and background clutter, has high requirements for hardware equipment, and the recognition accuracy is easily reduced in complex scenarios. In addition, the model training process requires reliance on a large amount of labeled data. The lack of labeled data resources limits the improvement of its generalization performance.

[0071] The above two methods have not yet been applied to the scenario of sign language recognition for deaf and mute people in non-driver positions in the cockpit of smart vehicles. This technology has not been effectively explored and applied in the cockpit scenario, and related development is almost blank.

[0072] To this end, in this application, an on-board sign language recognition model is trained by using original video stream data and hand key point label data, shoulder-hand key point label data and sign language semantic text sequence label data corresponding to sign language movements in the original video stream data, and text data is obtained by inputting video frame images into the on-board sign language recognition model. Control instructions are generated based on the text data, and the intelligent vehicle is controlled to execute control instruction actions, thereby improving the accuracy of complex sign language recognition and improving the accuracy of vehicle-computer command interaction, thereby filling the missing technology for sign language recognition for deaf and mute people in non-driving positions in the cockpit of intelligent vehicles.

[0073] After introducing the application scenarios and design concepts of the present invention, the technical solutions provided by the present invention are described in detail below.

[0074] The embodiment of the present invention provides a method for training a vehicle-mounted sign language recognition model. Figure 1 As shown, an embodiment of the present invention provides a method for training an in-vehicle sign language recognition model, which includes at least the following steps:

[0075] Step 110: Obtain a training sample data set; wherein the training sample data set includes multiple training sample data; each training sample data includes original video stream data and hand key point label data, shoulder-hand key point label data and sign language semantic text sequence label data corresponding to the sign language action in the original video stream data.

[0076] In the embodiment of the present application, when obtaining the training sample data set, the following methods may be used but are not limited to:

[0077] The occupancy monitoring system (OMS) in the car cabin collects RGB-IR (RGB-Infrared) raw video stream data in the car cabin, and performs frame segmentation processing on the raw video stream data to obtain multiple frames of raw video frame images;

[0078] The key point data and sign language data in the original video frame image are annotated manually or automatically by annotating tools to obtain hand key point label data, shoulder-hand key point label data, and sign language semantic text sequence label data;

[0079] Based on the original video stream data, hand key point label data, shoulder-hand key point label data and sign language semantic text sequence label data, a training sample data set is obtained.

[0080] Specifically, by performing key point tasks and sign language tasks on the RGB-IR original video stream data, a training sample data set is obtained. The specific process is as follows: Figure 2 As shown:

[0081] First, for the key point task, the RGB-IR video data is frame-cut and each frame of the video image is annotated to obtain the hand key point label data and the shoulder-hand key point label data. The hand key point label data consists of 2×21 key points, and the shoulder-hand key point label data consists of 7 key points. Specifically, the preset annotation model is first used to automatically annotate the human hand and arm shoulder in each frame of the video image. Then, the key point detection and annotation tool is used for manual verification to obtain the hand key point label data GT and the arm shoulder key point label data GT, as shown in the figure. Figure 3 As shown in , each key point in the hand key point label data has a label, starting from 0 and ending at 41, where the first 21 points represent the left hand, the last 21 points represent the right hand, and the combination of multiple points represents different fingers; Figure 4 As shown, there are 7 shoulder-hand key point label data from left to right, namely left wrist, left elbow, left shoulder, neck, right shoulder, right elbow and right wrist, and each key point includes (x, y) coordinates.

[0082] For the sign language task, the videos belonging to the same sign language content in the RGB-IR original video stream data are segmented to obtain sign language video images, and the sign language video images are annotated to obtain sign language semantic text sequence label data GT, wherein the sign language video images are annotated by professionals, and the specific annotation format is [segmented video path, "Today's weather in Shanghai is cloudy to rainy", "Shanghai / Today / Weather / Cloudy / Change / Light / Rain"], wherein "segmented video path" represents the path of the video data and is in mp4 format, "Today's weather in Shanghai is cloudy to rainy" represents the content of the video sign language, and "Shanghai / Today / Weather / Cloudy / Change / Light / Rain." represents the sign language expression.

[0083] Furthermore, the label data GT is processed. Specifically, for the hand key point label data GT, the coordinate information [x, y, index] of each label needs to be converted into the [tx, ty, index] format, where tx = x / w, w is the width of the sign language video image, ty = y / h, h is the height of the sign language video image, and the index is [0-41]. The index of the hand key point label data GT is combined and sorted by combining the 42 labels together to form a one-dimensional data group [1, 42, 2], which is recorded as the hand key point label data GT-KEY-Hand.

[0084] For the shoulder-hand key point label data GT, the coordinate information of each label [x, y, index] needs to be converted into the [tx, ty, index] format, where tx = x / w, w is the width, ty = y / h, h is the height, and the index is [0-6]. The index of the shoulder-hand key point label data GT is combined and sorted by combining the 7 labels together to form a one-dimensional data group [1, 7, 2], which is recorded as the shoulder-hand key point label data GT-KEY-JS;

[0085] For the sign language semantic text sequence label data GT, for example, the original text label is: "Shanghai / Today / Weather / Cloudy / Change / Light / Rain / ", first, based on all the original text labels, a word dictionary table is formed, and each word or phrase corresponds to a digital ID. The text with the original text label separated by " / " is converted into ["Shanghai", "Today", "Weather", "Cloudy", "Change", "Light", "Rain"], and the converted sentence and the word dictionary are matched, that is, converted into [23, 564, 12, 234, 54, 170, 102], and the sign language semantic text sequence label data is recorded as GT-CTC.

[0086] Since there are two image formats in the RGB-IR original video stream data used, one is a normal RGB color image and the other is a grayscale image at night, in the embodiment of the present application, the RGB color image and the grayscale image are normalized to obtain an image with a single channel value between [0-1].

[0087] Specifically, in an embodiment of the present application, the cv2.imread function is used to read the image data in the original video stream. First, the image data in the original video stream is converted from BGR to RGB using the cv2.cvtColor function to obtain the original dimension number [B, C, H, W], where B represents the batch dimension, C represents the channel dimension, C=3, H represents the height dimension, and W represents the width dimension. Then, the original dimension number [B, 1, H, W] of the grayscale image and the original dimension number [B, C, H, W] of the RGB image are feature spliced ​​to obtain the original feature dimension number [B, 4, H, W]. Finally, the spliced ​​original features are normalized, for example, RGB-I The original value of R is [230, 100, 20, 100]. After normalization, it is divided by 255 to [0.0784, 0.0824, 0.0980, 0.0824]. Then, the average value [0.485, 0.456, 0.406, 0.456] is subtracted in the RGB-IR order and divided by the standard deviation [0.229, 0.224, 0.225, 0.224] to finally obtain [-1.7754, -1.6681, -1.3687, -1.6681]. This ensures that all channel values ​​after processing have zero mean and unit variance. The processed feature data is input into the in-vehicle sign language recognition model to accelerate model convergence and more stable training.

[0088] Step 120: Select target training sample data from the training sample data set.

[0089] Step 130: Input the original video stream data in the target training sample data into the on-board sign language recognition model, so that the on-board sign language recognition model performs feature extraction and key point detection on the multiple frames of original video frame images in the original video stream data through the key point detection sub-model to obtain hand key point feature data and shoulder-hand key point feature data, and performs gesture change trajectory capture and sign language action semantic analysis on the original video stream data and the hand key point feature data and shoulder-hand key point feature data corresponding to the multiple frames of original video frame images through the sign language recognition sub-model to obtain sign language semantic text sequence prediction data.

[0090] Step 140: Calculate the current loss value based on the sign language semantic text sequence prediction data, the hand key point feature data, the shoulder-hand key point feature data, and the sign language semantic text sequence label data, the hand key point label data, and the shoulder-hand key point label data in the target training sample data, and update the first weight parameter of the key point detection sub-model and the second weight parameter of the sign language recognition sub-model based on the current loss value.

[0091] Step 150, determine whether the iterative training termination condition is met; if so, execute step 160; if not, return to step 120; wherein, the iterative training termination condition is that the number of iterations is not less than the number threshold, or the first loss value is not higher than the first loss value threshold.

[0092] Step 160: Obtain an in-vehicle sign language recognition model based on the parameters of the key point detection sub-model and the parameters of the sign language recognition sub-model updated during the last iterative training operation.

[0093] In the embodiments of this application, Figure 5 As shown, the on-board sign language recognition model includes a key point detection sub-model and a sign language recognition sub-model. The original video stream data is input into the on-board sign language recognition model, so that the on-board sign language recognition model performs feature extraction and key point detection on multiple frames of original video frame images in the original video stream data through the key point detection sub-model to obtain hand key point feature data and shoulder-hand key point feature data. The sign language recognition sub-model performs gesture change trajectory capture and sign language action semantic analysis on the original video stream data and the hand key point feature data and shoulder-hand key point feature data corresponding to the multiple frames of original video frame images to obtain sign language semantic text sequence prediction data; based on the sign language semantic text sequence prediction data, the hand key point feature data and the shoulder-hand key point feature data and the sign language semantic text sequence label data GT-CTC, the hand key point label data GT-KEY-Hand and the shoulder-hand key point label data GT-KEY-JS in the target training sample data, the current loss value is calculated, and the first weight parameter of the key point detection sub-model and the second weight parameter of the sign language recognition sub-model are updated based on the current loss value.

[0094] Furthermore, the backbone of the key point detection sub-model is based on the traditional residual network ResNet. It introduces the multi-scale feature extraction module InceptionBlock to achieve multi-scale feature fusion through parallel convolution kernels of different sizes. It introduces the bidirectional feature extraction module BiLSTM to extract higher-level temporal features, enhance sensitivity to detail changes and accurately capture forward and reverse temporal dependencies. It introduces the multi-scale feature fusion module Multi-CNN to further refine feature information, including the trend, direction and amplitude of the movement. It introduces two fully connected layers FC to predict 42 key points of the hand key point feature data and 7 key points of the shoulder-hand key point feature data respectively.

[0095] The backbone of the sign language recognition sub-model is based on the traditional residual network ResNet. The temporal feature extraction module TCNN is introduced to enhance the extraction and compression of time series features. The bidirectional feature extraction module BiLSTM is introduced to keep the feature dimension of the extracted features unchanged. A fully connected layer is introduced to obtain text data.

[0096] The loss calculation of the in-vehicle sign language recognition model is based on the hand key point feature data and the hand label data GT-KEY-Hand to obtain the hand loss value L2-hand-loss; based on the shoulder-hand key point feature data and the shoulder-hand label data GT-KEY-JS to obtain the shoulder-hand loss value L2-JS-loss; based on the text data and the sign language label data GT-CTC to obtain the sign language loss value CTC-loss; based on the hand loss value L2-hand-loss and the shoulder-hand loss value L2-JS-loss, the weight of the key point detection sub-model is updated, and based on the sign language loss value CTC-loss, the weight of the sign language recognition sub-model is updated.

[0097] In an optional embodiment, the key point detection submodel includes a first local feature extraction module, a first multi-scale feature extraction module, a first bidirectional feature extraction module, a multi-scale feature fusion module and a key point detection module connected in sequence; the key point detection submodel performs feature extraction and key point detection on the original video stream data to obtain hand key point feature data and shoulder-hand key point feature data, including: preprocessing and feature extraction of the original video stream data by the first local feature extraction module to obtain a first key point feature image; performing multi-scale feature extraction on the features of the first key point feature image by the first multi-scale feature extraction module to obtain a second key point feature image; performing temporal feature extraction on the features of the second key point feature image by the first bidirectional feature extraction module to obtain a third key point feature image; extracting and fusion processing on the third key point feature image by the multi-scale feature fusion module to obtain a feature fusion image; performing key point detection on the feature fusion image by the key point detection module to obtain hand key point feature data and shoulder-hand key point feature data.

[0098] Furthermore, the first local feature extraction module includes a first 3D convolution layer; the first multi-scale feature extraction module includes a plurality of first multi-scale convolution units connected in sequence, the first multi-scale convolution unit includes a first hole convolution layer, a first convolution layer, a second convolution layer connected in parallel, and a first splicing layer connected to the first hole convolution layer, the first convolution layer and the second convolution layer respectively, and a first compression-excitation attention layer jump-connected to the first splicing layer; the first bidirectional feature extraction module includes two first bidirectional long-short-term memory layers connected in sequence; the multi-scale feature fusion module includes a multi-branch convolution layer; the key point detection module includes two first fully connected layers; the key point detection sub-model is used to perform feature extraction and key point detection on the original video stream data to obtain hand key point feature data and shoulder-hand key point feature data, including: through The first 3D convolution layer in the first local feature extraction module preprocesses and extracts features from the original video stream data to obtain a first key point feature image; the first multi-scale convolution units in multiple first multi-scale feature extraction modules perform multi-scale feature extraction on the features of the first key point feature image to obtain a second key point feature image; the two first bidirectional long short-term memory layers in the first bidirectional feature extraction module perform temporal feature extraction on the features of the second key point feature image to obtain a third key point feature image; the multi-branch convolution layer in the multi-scale feature fusion module extracts and fuses the third key point feature image to obtain a feature fusion image; the first fully connected layer in the key point detection module performs key point detection on the feature fusion image to obtain hand key point feature data and shoulder-hand key point feature data.

[0099] In the embodiments of this application, Figure 6 As shown in the figure, the key point detection sub-model includes the first 3D convolution layer Conv3d in the first local feature extraction module, four first multi-scale convolution units InceptionBlock in the first multi-scale feature extraction module, two first bidirectional long short-term memory layers BiLSTM in the first bidirectional feature extraction module, a multi-branch convolution layer Multi-CNN in the multi-scale feature fusion module, and two first fully connected layers FC in the key point detection module. The number of convolution kernels in the 3D convolution layer Conv3d is 64, and the convolution kernel depth is 4. The four multi-scale feature extraction layers improve the ability to extract fine-grained features. The two first bidirectional long short-term memory layers BiLSTM extract higher-level temporal features to enhance sensitivity to detailed changes and accurately capture the forward and reverse temporal dependencies to improve the ability to model dynamic changes. The multi-branch convolution layer Multi-CNN further refines feature information, mainly including the trend, direction, and amplitude of the movement. Two fully connected layers predict 42 key points of the hand and 7 key points of the shoulder and hand, respectively.

[0100] In order to highlight key detail information, an attention mechanism SE module is embedded in the multi-scale feature extraction module. Specifically, the multi-scale feature extraction module also includes a first compressed excitation attention layer, which adaptively adjusts the feature weights of different channels or spatial positions, especially the fine movement features of fingers.

[0101] In order to perform subsequent feature fusion more efficiently, reduce the computational cost of the in-vehicle sign language recognition model and reduce feature redundancy, in an embodiment of the present application, the key point detection submodel also includes a random multi-frame feature extraction module; the random multi-frame feature extraction module is connected to the first local feature extraction module, and is used to group and extract the original video stream data to obtain multiple frames of original video frame images.

[0102] Specifically, a random multi-frame feature extraction module is introduced into the in-vehicle sign language recognition model, and the original video stream data is input into the random multi-frame feature extraction module, so that the random multi-frame feature extraction module extracts k frames as the input of the key point detection sub-model; wherein, the original video stream data is evenly divided into K groups, and one frame is extracted from each of the K groups to form k frames, and the k frames are not frames of fixed time, wherein K and k are both integers.

[0103] In an optional embodiment, the sign language recognition submodel includes a second local feature extraction module, a second multi-scale feature extraction module, a temporal feature extraction module, a second bidirectional feature extraction module and a sign language action semantic parsing module connected in sequence; the sign language recognition submodel performs gesture change trajectory capture and sign language action semantic parsing on the original video stream data and the hand key point feature data and the shoulder-hand key point feature data corresponding to the multiple frames of original video frame images, including: performing gesture change trajectory capture on the original video stream data and the hand key point feature data and the shoulder-hand key point feature data corresponding to the multiple frames of original video frame images by the second local feature extraction module to obtain a first gesture feature image; performing multi-scale feature extraction on the features of the gesture feature image by the second multi-scale feature extraction module to obtain a second gesture feature image; extracting temporal information from the second gesture feature image by the temporal feature extraction module to obtain a third gesture feature image; performing temporal feature extraction on the third gesture feature image by the second bidirectional feature extraction module to obtain a fused image; and performing sign language action semantic parsing on the fused image by the sign language action semantic parsing module to obtain sign language semantic text sequence prediction data.

[0104] Furthermore, the second local feature extraction module includes a second 3D convolution layer; the second multi-scale feature extraction module includes a plurality of second multi-scale convolution units connected in sequence, the second multi-scale convolution unit includes a second hole convolution layer, a third convolution layer, and a fourth convolution layer connected in parallel, and a second splicing layer connected to the second hole convolution layer, the third convolution layer and the fourth convolution layer respectively, and a second compression-excitation attention layer jump-connected to the second splicing layer; the temporal feature extraction module includes a time-space convolution layer; the second bidirectional feature extraction module includes a second bidirectional long-short-term memory layer; the sign language action semantic parsing module includes a second fully connected layer; the hand key point feature data and the shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of original video frame images are captured by the sign language recognition sub-model to perform gesture change trajectory capture and sign language action semantic parsing to obtain sign language semantic text sequence prediction data , including: capturing the hand key point feature data and shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of original video frame images through the second 3D convolution layer in the second local feature extraction module to obtain the first gesture feature image; performing multi-scale feature extraction on the features of the gesture feature image through the second multi-scale convolution unit in the second multi-scale feature extraction module to obtain the second gesture feature image; extracting the timing information of the second gesture feature image through the time-space convolution layer in the temporal feature extraction module to obtain the third gesture feature image; performing temporal feature extraction on the third gesture feature image through the second bidirectional long short-term memory layer in the second bidirectional feature extraction module to obtain the fused image; performing sign language action semantic analysis on the fused image through the second fully connected layer in the sign language action semantic analysis module to obtain the sign language semantic text sequence prediction data.

[0105] To ensure the integrity of sign language actions, in the embodiments of this application, one input of the sign language recognition sub-model is the original video stream data, whose feature dimension is [B, maxLen, C, w, h], where B represents the batch size, maxLen represents the maximum value of the short video frame data in the current batch, which is a dynamic value, C represents the number of channels, usually 3, w represents the width, and h represents the height. After preprocessing, it is sampled to [B, maxLen, C, 224, 224]. The preprocessed original video is processed by the residual network of the sign language recognition sub-model to obtain a residual feature dimension of [xLen, 1, 1024], where xLen is the dynamic value of the feature length, 1 is the single position after spatial dimension compression, and 1024 is the number of feature channels. The residual feature dimension is denoted as fea1. Another input of the sign language recognition sub-model is the hand key point feature data and shoulder-hand key point feature data output by the key point detection sub-model, whose feature dimension is [xmLen, 1, 1024], and the feature dimension is denoted as fea0. Since xmLen < xLen and the two features cannot be directly added, fea0 is first extended in the first dimension to xLen. The specific operation is as Figure 7 shown. Assume xmLen = 3 and xLen = 6. The first dimension xmLen = 3 is extended to xLen = 6. The extension is a uniform extension to ensure the generality of the features. The extended fea0 feature is denoted as fea0n. Secondly, fea1 and fea0n are added. The dimension after addition remains unchanged, and the feature dimension after addition is [xLen, 1, 1024], denoted as the feature fea. Then, the feature dimension after addition is input into the temporal feature extraction module to extract and compress the time series features, and important temporal information is refined through layer-by-layer convolution and pooling operations. Then, the image output by the temporal feature extraction module is input into the second bidirectional feature extraction module BiLSTM to keep the number of feature dimensions unchanged. Finally, the image feature dimension output by the second bidirectional feature extraction module BiLSTM is input into the second fully connected layer FC to obtain the sign language semantic text sequence prediction data. Among them, the feature dimension number output by the second fully connected layer FC is [xLen, dictNum], and dictNum is the size of the sign language corresponding Chinese character dictionary table.

[0106] In an optional embodiment, the current loss value is calculated based on the sign language semantic text sequence prediction data, the hand key point feature data and the shoulder-hand key point feature data, as well as the sign language semantic text sequence label data, the hand key point label data and the shoulder-hand key point label data in the target training sample data, including: calculating the hand loss value using the target loss function based on the hand key point feature data and the hand key point label data in the target training sample data; calculating the shoulder-hand loss value using the target loss function based on the shoulder-hand key point feature data and the shoulder-hand key point label data in the target training sample data; obtaining the feature loss value based on the hand loss value and the shoulder-hand loss value; calculating the sequence loss value using the connection temporal classification loss function based on the sign language semantic text sequence prediction data and the sign language semantic text sequence label data; and calculating the current loss value based on the feature loss value and the sequence loss value.

[0107] In an optional embodiment, the current loss value is calculated based on the sign language semantic text sequence prediction data, hand key point feature data and shoulder-hand key point feature data, as well as the sign language semantic text sequence label data, hand key point label data and shoulder-hand key point label data in the target training sample data, and also includes: calculating the sequence loss value using a connection temporal classification loss function based on the sign language semantic text sequence prediction data and the sign language semantic text sequence label data; determining the gradient value corresponding to the sign language recognition sub-model based on the sequence loss value and the sign language semantic text sequence prediction data; updating the gradient value based on the optimization function, and inputting the updated gradient value into the sign language recognition sub-model to perform back propagation to update the second weight parameter of the sign language recognition sub-model.

[0108] In an embodiment of the present application, the key point detection submodel uses two losses, namely L1-hand-loss and L1-JS-loss, to calculate the loss value between the predicted key point coordinates and the standard coordinates. The sign language recognition submodel uses the loss value CTC-loss, which is used for sequence-to-sequence tasks to solve the loss function that solves the problem of misalignment between the input length and the output length. It efficiently calculates the alignment path through dynamic planning and automatically learns the optimal input-to-output mapping during the training process. The total loss loss of the L1-hand-loss loss value, the L1-JS-loss loss value and the CTC-loss loss value are multiplied by the adaptive weight parameters and added together to jointly optimize the parameters of the on-board sign language recognition model.

[0109] Specifically, if Figure 8As shown, in the present application, the optimization of the in-vehicle sign language recognition model through loss calculation mainly includes two branches. One branch is to input the original video stream data into the key point detection sub-model to obtain hand key point feature data and shoulder-hand key point feature data, and obtain the first loss value, namely result 1, based on the hand key point feature data and shoulder-hand key point feature data and the hand key label data and the shoulder-hand key label data; the other branch is to input the original video stream data, the hand key point feature data and the shoulder-hand key point feature data into the sign language recognition sub-model to obtain sign language semantic text sequence prediction data, and obtain the second loss value, namely result 2, based on the sign language semantic text sequence prediction data and the sign language semantic text sequence label data, and perform back propagation and gradient optimization based on the second loss value to update the first weight parameter of the sign language recognition sub-model and the second weight parameter of the sign language recognition sub-model until the loss value reaches the preset threshold, thereby obtaining the weight parameters of the in-vehicle sign language recognition model.

[0110] In an embodiment of the present application, by analyzing the sign language recognition judgment factors, combining the overall information of the user's upper body, the key point information of the hands and the key point information of the arms, the dependent features of the entire in-vehicle sign language recognition model are formed to perform sign language translation accurately in all directions; and the in-vehicle sign language recognition model is obtained by modeling and learning based on time-frequency data and deep learning models. Compared with the single modal information modeled based on sensor data, the in-vehicle sign language recognition model can include more information, making the in-vehicle sign language recognition model generalized and stable.

[0111] Figure 9 A flow chart of an intelligent vehicle control method provided by an embodiment of the present invention is shown as follows: Figure 9 As shown, the intelligent vehicle control method includes at least the following steps:

[0112] Step 210: Acquire video stream data in the smart vehicle cabin.

[0113] Step 220: Input the video stream data into the in-vehicle sign language recognition model to obtain sign language semantic text sequence data corresponding to the video stream data; wherein the in-vehicle sign language recognition model is trained using the above-mentioned in-vehicle sign language recognition model training method.

[0114] Step 230: Generate a control instruction based on the sign language semantic text sequence data, and control the intelligent vehicle to execute the control instruction.

[0115] In an embodiment of the present application, the sign language information in the video stream data is translated by an on-board sign language recognition model, and intelligent vehicle control instructions are generated based on the sign language translation information, so that disabled people (non-drivers) can control the air conditioning, audio and other equipment of the intelligent vehicle through sign language; and the video stream data is input into the on-board sign language recognition model to obtain sign language semantic text sequence data, thereby abandoning sensor equipment such as hand wear and using sensors already arranged in the intelligent vehicle to obtain video image information to reduce costs.

[0116] Figure 10 A structural diagram of a vehicle-mounted sign language recognition model training device provided by an embodiment of the present invention is shown in FIG. Figure 10 As shown, the model training device at least includes:

[0117] The data acquisition module 310 is configured to acquire a training sample data set, wherein the training sample data set includes a plurality of training sample data; each training sample data includes original video stream data and hand key point label data, shoulder-hand key point label data, and sign language semantic text sequence label data corresponding to the sign language action in the original video stream data;

[0118] The model training module 320 is used to iteratively perform training operations on the in-vehicle sign language recognition model based on the training sample data set; wherein the in-vehicle sign language recognition model includes a key point detection sub-model and a sign language recognition sub-model with jump connections; the training operation includes: selecting target training sample data from the training sample data set; inputting the original video stream data in the target training sample data into the in-vehicle sign language recognition model, so that the in-vehicle sign language recognition model performs feature extraction and key point detection on multiple frames of original video frame images in the original video stream data through the key point detection sub-model to obtain hand key point feature data and shoulder-hand key point feature data, and The sign language recognition sub-model captures the hand key point feature data and the shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of original video frame images, and performs semantic analysis of sign language movements to obtain sign language semantic text sequence prediction data; based on the sign language semantic text sequence prediction data, the hand key point feature data and the shoulder-hand key point feature data, and the sign language semantic text sequence label data, the hand key point label data and the shoulder-hand key point label data in the target training sample data, the current loss value is calculated, and the first weight parameter of the key point detection sub-model and the second weight parameter of the sign language recognition sub-model are updated based on the current loss value.

[0119] Figure 11 The present invention provides a schematic diagram of the structure of an intelligent vehicle control device. Figure 11 As shown, the control device at least includes:

[0120] An acquisition module 410 is used to acquire video stream data in the cockpit of the smart vehicle;

[0121] The recognition module 420 is configured to input the video stream data into an in-vehicle sign language recognition model to obtain sign language semantic text sequence data corresponding to the video stream data; wherein the in-vehicle sign language recognition model is trained using the in-vehicle sign language recognition model training method described above;

[0122] The control module 430 is used to generate control instructions based on the sign language semantic text sequence data and control the intelligent vehicle to execute the control instructions.

[0123] Based on the above embodiments, an embodiment of the present invention provides an electronic device, referring to Figure 12 As shown, the electronic device 600 provided in the embodiment of the present invention includes at least a controller 601, a memory 602, and a computer program stored in the memory 602 and executable on the controller 601. When the controller 601 executes the computer program, the above-mentioned vehicle-mounted sign language recognition model training method or intelligent vehicle control method provided in the embodiment of the present invention is implemented.

[0124] The electronic device 600 provided in this embodiment of the present invention may further include a bus 603 connecting different components (including the controller 601 and the memory 602). The bus 603 represents one or more of several types of bus structures, including a memory bus, a peripheral bus, a local bus, and the like.

[0125] The memory 602 may include a readable storage medium in the form of a volatile memory, such as a random access memory (RAM) 6021 and / or a cache memory 6022, and may further include a read-only memory (ROM) 6023. The memory 602 may also include a program tool 6025 having a set (at least one) of program modules 6024. The program modules 6024 include, but are not limited to, an operating subsystem, one or more application programs, other program modules, and program data. Each of these examples or some combination thereof may include the implementation of a network environment.

[0126] The controller 601 may be a processing element or a collective term for multiple processing elements. For example, the controller 601 may be a central processing unit (CPU) or one or more integrated circuits configured to implement the above-mentioned in-vehicle sign language recognition model training method or intelligent vehicle control method provided in an embodiment of the present invention. Specifically, the controller 601 may be a general-purpose controller, including but not limited to a CPU, an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.

[0127] The electronic device 600 can communicate with one or more external devices 604 (e.g., keyboard, remote control, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600 (e.g., mobile phone, computer, etc.), and / or communicate with a device that enables the electronic device 600 to communicate with one or more other electronic devices 600 (e.g., router, modem, etc.). Such communication can be performed through an input / output (I / O) interface 605. In addition, the electronic device 600 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 606. Figure 6 As shown, the network adapter 606 communicates with other modules of the electronic device 600 via the bus 603. Figure 6 Not shown, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to microcode, device drivers, redundant controllers, external disk drive arrays, disk arrays (Redundant Arrays of Independent Disks, RAID) subsystems, tape drives, and data backup storage subsystems.

[0128] It should be noted that Figure 12 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0129] The following describes the computer-readable storage medium provided in an embodiment of the present invention. The computer-readable storage medium provided in an embodiment of the present invention stores computer instructions that, when executed by a controller, implement the aforementioned in-vehicle sign language recognition model training method or intelligent vehicle control method provided in an embodiment of the present invention. In a specific implementation, the computer instructions may be built into or installed in the controller. Thus, the controller can implement the aforementioned in-vehicle sign language recognition model training method or intelligent vehicle control method provided in an embodiment of the present invention by executing the built-in or installed computer instructions.

[0130] In addition, the in-vehicle sign language recognition model training method or intelligent vehicle control method provided in the embodiments of the present invention can also be implemented as a computer program product, which includes program code. When the program code is executed on a controller, it implements the in-vehicle sign language recognition model training method or intelligent vehicle control method provided in the embodiments of the present invention.

[0131] The computer program product provided by the embodiments of the present invention may adopt one or more computer-readable storage media, and the computer-readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any suitable combination of the above. Specifically, more specific examples of computer-readable storage media (a non-exhaustive list) include an electrical connection with one or more wires, a portable disk, a hard disk, RAM, ROM, Erasable Programmable Read Only Memory (EPROM), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.

[0132] The computer program product provided in the embodiments of the present invention may utilize a CD-ROM and include program code. It may also be run on electronic devices such as an in-vehicle sign language recognition model training device or an intelligent vehicle control device. However, the computer program product provided in the embodiments of the present invention is not limited thereto. In the embodiments of the present invention, a computer-readable storage medium may be any tangible medium that contains or stores program code, and the program code may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0133] It should be noted that although several units or subunits of the device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more units described above may be embodied in a single unit. Conversely, the features and functions of a single unit described above may be further divided and embodied by multiple units.

[0134] Furthermore, although the operations of the method of the present invention are described in a particular order in the accompanying drawings, this does not require or imply that these operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0135] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0136] Obviously, those skilled in the art may make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if such changes and modifications of the embodiments of the present invention fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A method for training an in-vehicle sign language recognition model, characterized in that: include: Obtain a training sample data set; The training sample data set includes a plurality of training sample data; each of the training sample data includes original video stream data and hand key point label data, shoulder-hand key point label data and sign language semantic text sequence label data corresponding to the sign language action in the original video stream data; Iteratively performing a training operation on an in-vehicle sign language recognition model based on the training sample data set; wherein the in-vehicle sign language recognition model includes a key point detection sub-model and a sign language recognition sub-model with skip connections; the training operation includes: Selecting target training sample data from the training sample data set; The original video stream data in the target training sample data is input into the in-vehicle sign language recognition model, so that the in-vehicle sign language recognition model performs feature extraction and key point detection on multiple frames of original video frame images in the original video stream data through the key point detection sub-model to obtain hand key point feature data and shoulder-hand key point feature data, and performs gesture change trajectory capture and sign language action semantic analysis on the original video stream data and the hand key point feature data and the shoulder-hand key point feature data corresponding to the multiple frames of the original video frame images through the sign language recognition sub-model to obtain sign language semantic text sequence prediction data; Based on the sign language semantic text sequence prediction data, the hand key point feature data and the shoulder-hand key point feature data, as well as the sign language semantic text sequence label data, the hand key point label data and the shoulder-hand key point label data in the target training sample data, a current loss value is calculated, and based on the current loss value, a first weight parameter of the key point detection sub-model and a second weight parameter of the sign language recognition sub-model are updated.

2. The in-vehicle sign language recognition model training method according to claim 1, characterized in that: Calculating a current loss value based on the sign language semantic text sequence prediction data, the hand key point feature data, the shoulder-hand key point feature data, and the sign language semantic text sequence label data, the hand key point label data, and the shoulder-hand key point label data in the target training sample data, including: Calculating a hand loss value using a target loss function based on the hand key point feature data and the hand key point label data in the target training sample data; Calculating a shoulder-hand loss value using a target loss function based on the shoulder-hand key point feature data and the shoulder-hand key point label data in the target training sample data; Obtaining a feature loss value based on the hand loss value and the shoulder-hand loss value; Calculating a sequence loss value using a connection temporal classification loss function based on the sign language semantic text sequence prediction data and the sign language semantic text sequence label data; The current loss value is calculated based on the feature loss value and the sequence loss value.

3. The in-vehicle sign language recognition model training method according to claim 1, characterized in that: The key point detection sub-model includes a first local feature extraction module, a first multi-scale feature extraction module, a first bidirectional feature extraction module, a multi-scale feature fusion module and a key point detection module connected in sequence; The key point detection sub-model is used to perform feature extraction and key point detection on the original video stream data to obtain hand key point feature data and shoulder-hand key point feature data, including: Preprocessing and feature extraction are performed on the original video stream data by the first local feature extraction module to obtain a first key point feature image; Performing multi-scale feature extraction on features of the first key point feature image by the first multi-scale feature extraction module to obtain a second key point feature image; Performing temporal feature extraction on features of the second key point feature image using the first bidirectional feature extraction module to obtain a third key point feature image; Extracting and fusing the third key point feature image through the multi-scale feature fusion module to obtain a feature fusion image; The key point detection module performs key point detection on the feature fusion image to obtain hand key point feature data and shoulder-hand key point feature data.

4. The in-vehicle sign language recognition model training method according to claim 3, characterized in that: The first local feature extraction module includes a first 3D convolution layer; the first multi-scale feature extraction module includes a plurality of first multi-scale convolution units connected in sequence, the first multi-scale convolution unit includes a first hole convolution layer, a first convolution layer, a second convolution layer connected in parallel, and a first splicing layer connected to the first hole convolution layer, the first convolution layer and the second convolution layer respectively, and a first compression-excitation attention layer jump-connected to the first splicing layer; the first bidirectional feature extraction module includes two first bidirectional long short-term memory layers connected in sequence; the multi-scale feature fusion module includes a multi-branch convolution layer; the key point detection module includes two first fully connected layers; The key point detection sub-model is used to perform feature extraction and key point detection on the original video stream data to obtain hand key point feature data and shoulder-hand key point feature data, including: Preprocessing and extracting features from the original video stream data using the first 3D convolutional layer in the first local feature extraction module to obtain a first key point feature image; Performing multi-scale feature extraction on features of the first key point feature image by using the first multi-scale convolution units in the plurality of first multi-scale feature extraction modules to obtain a second key point feature image; Performing temporal feature extraction on features of the second key point feature image using the two first bidirectional long short-term memory layers in the first bidirectional feature extraction module to obtain a third key point feature image; Extracting and fusing the third key point feature image through the multi-branch convolutional layer in the multi-scale feature fusion module to obtain a feature fusion image; The first fully connected layer in the key point detection module performs key point detection on the feature fusion image to obtain hand key point feature data and shoulder-hand key point feature data.

5. The in-vehicle sign language recognition model training method according to claim 3 or 4, characterized in that: The key point detection sub-model also includes a random multi-frame feature extraction module; the random multi-frame feature extraction module is connected to the first local feature extraction module, and is used to group and extract the original video stream data to obtain multiple frames of original video frame images.

6. The in-vehicle sign language recognition model training method according to claim 1, characterized in that: The sign language recognition submodel includes a second local feature extraction module, a second multi-scale feature extraction module, a temporal feature extraction module, a second bidirectional feature extraction module, and a sign language action semantic parsing module connected in sequence; the sign language recognition submodel performs gesture change trajectory capture and sign language action semantic parsing on the original video stream data and the hand key point feature data and the shoulder-hand key point feature data corresponding to multiple frames of the original video frame images to obtain sign language semantic text sequence prediction data, including: Capturing gesture change trajectories on the original video stream data and the hand key point feature data and the shoulder-hand key point feature data corresponding to multiple frames of the original video frame images by a second local feature extraction module to obtain a first gesture feature image; Performing multi-scale feature extraction on features of the gesture feature image by the second multi-scale feature extraction module to obtain a second gesture feature image; Extracting time sequence information from the second gesture feature image using the time sequence feature extraction module to obtain a third gesture feature image; Performing temporal feature extraction on the third gesture feature image by the second bidirectional feature extraction module to obtain a fused image; The sign language action semantic parsing module performs sign language action semantic parsing on the fused image to obtain sign language semantic text sequence prediction data.

7. The in-vehicle sign language recognition model training method according to claim 6, characterized in that: The second local feature extraction module includes a second 3D convolution layer; the second multi-scale feature extraction module includes a plurality of second multi-scale convolution units connected in sequence, the second multi-scale convolution unit includes a second hole convolution layer, a third convolution layer, and a fourth convolution layer connected in parallel, and a second splicing layer connected to the second hole convolution layer, the third convolution layer and the fourth convolution layer respectively, and a second compression-excitation attention layer jump-connected to the second splicing layer; the temporal feature extraction module includes a time-space convolution layer; the second bidirectional feature extraction module includes a second bidirectional long-short-term memory layer; the sign language action semantic parsing module includes a second fully connected layer; the hand key point feature data and the shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of the original video frame images are captured by the sign language recognition sub-model through gesture change trajectory capture and sign language action semantic parsing to obtain sign language semantic text sequence prediction data, including: The second 3D convolutional layer in the second local feature extraction module captures the hand key point feature data and the shoulder-hand key point feature data corresponding to the original video stream data and multiple frames of the original video frame images to obtain a first gesture feature image; performing multi-scale feature extraction on features of the gesture feature image by the second multi-scale convolution unit in the second multi-scale feature extraction module to obtain a second gesture feature image; Extracting time sequence information from the second gesture feature image through the time-space convolution layer in the time sequence feature extraction module to obtain a third gesture feature image; performing temporal feature extraction on the third gesture feature image using the second bidirectional long short-term memory layer in the second bidirectional feature extraction module to obtain a fused image; The second fully connected layer in the sign language action semantic parsing module performs sign language action semantic parsing on the fused image to obtain sign language semantic text sequence prediction data.

8. An intelligent vehicle control method, characterized in that: include: Obtain video stream data inside the smart vehicle cabin; Inputting the video stream data into an in-vehicle sign language recognition model to obtain sign language semantic text sequence data corresponding to the video stream data; wherein the in-vehicle sign language recognition model is trained using the in-vehicle sign language recognition model training method according to any one of claims 1 to 7; Based on the sign language semantic text sequence data, a control instruction is generated, and the smart vehicle is controlled to execute the control instruction.

9. A vehicle-mounted sign language recognition model training device, characterized in that: include: A data acquisition module is used to obtain a training sample data set; The training sample data set includes a plurality of training sample data; each of the training sample data includes original video stream data and hand key point label data, shoulder-hand key point label data and sign language semantic text sequence label data corresponding to the sign language action in the original video stream data; The model training module is used to iteratively perform training operations on the in-vehicle sign language recognition model based on the training sample data set; wherein the in-vehicle sign language recognition model includes a key point detection sub-model and a sign language recognition sub-model with jump connections; the training operation includes: selecting target training sample data from the training sample data set; inputting the original video stream data in the target training sample data into the in-vehicle sign language recognition model, so that the in-vehicle sign language recognition model performs feature extraction and key point detection on multiple frames of original video frame images in the original video stream data through the key point detection sub-model to obtain hand key point feature data and shoulder-hand key point feature data, and The sub-model captures the gesture change trajectory and performs sign language action semantic analysis on the original video stream data and the hand key point feature data and the shoulder-hand key point feature data corresponding to multiple frames of the original video frame images to obtain sign language semantic text sequence prediction data; based on the sign language semantic text sequence prediction data, the hand key point feature data and the shoulder-hand key point feature data and the sign language semantic text sequence label data, the hand key point label data and the shoulder-hand key point label data in the target training sample data, the current loss value is calculated, and the first weight parameter of the key point detection sub-model and the second weight parameter of the sign language recognition sub-model are updated based on the current loss value.

10. An intelligent vehicle control device, characterized in that: include: An acquisition module, used to acquire video stream data in the cockpit of an intelligent vehicle; a recognition module, configured to input the video stream data into an in-vehicle sign language recognition model to obtain sign language semantic text sequence data corresponding to the video stream data; wherein the in-vehicle sign language recognition model is trained using the in-vehicle sign language recognition model training method according to any one of claims 1 to 7; A control module is used to generate a control instruction based on the sign language semantic text sequence data, and control the intelligent vehicle to execute the control instruction.

11. An electronic device, characterized in that: include: processor and memory; The memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements the in-vehicle sign language recognition model training method according to any one of claims 1 to 7 or the intelligent vehicle control method according to claim 8.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the in-vehicle sign language recognition model training method according to any one of claims 1 to 7 or the intelligent vehicle control method according to claim 8.

Citation Information

Patent Citations

  • Sign language recognition and translation method

    CN112487951A

  • Sign language recognition method and device based on neural network model, equipment and medium

    CN114067362A

  • Region positioning method and device, storage medium and electronic equipment

    CN114972713A

  • Sign language recognition model training method, dynamic sign language recognition method and dynamic sign language recognition device

    CN118609208A

  • Vehicle-machine interaction method and device, vehicle, electronic equipment and storage medium

    CN119428730A