Continuous sign language recognition word segmentation method and device

By using parallel multi-scale visual feature extraction and cross-modal alignment constraints, the problem of imprecise characterization of the temporal length of sign language actions in existing technologies has been solved, thereby improving the accuracy of fine word segmentation and recognition of sign language actions.

CN116665304BActive Publication Date: 2026-04-21TIANJIN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TIANJIN UNIVERSITY OF TECHNOLOGY
Filing Date
2023-06-11
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately characterize sign language movements of arbitrary temporal lengths, and the probability labels for similar sign language movements are inaccurate, making it difficult to effectively capture sign language movements of various temporal lengths.

Method used

A parallel multi-scale visual feature extraction model is adopted. By extracting text features and segmenting visual features at multiple scales, combined with cross-modal alignment constraints and the CTC objective function, a continuous sign language recognition and word segmentation system is trained to finely characterize sign language actions of each time length.

Benefits of technology

It achieves fine-grained word segmentation of sign language actions, improves the accuracy and generalization ability of sign language recognition, and enhances the word segmentation performance of sign language recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116665304B_ABST
    Figure CN116665304B_ABST
Patent Text Reader

Abstract

This invention provides a sign language recognition and segmentation method and apparatus, relating to the technical field of artificial intelligence, and applied to a continuous sign language recognition and segmentation system. The continuous sign language recognition and segmentation system includes a text extraction model and a parallel multi-scale visual feature extraction model, specifically comprising the following steps: inputting a continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset; determining sign language recognition data videos using the continuous sign language recognition dataset, inputting the sign language recognition data videos into the parallel multi-scale visual feature extraction model to segment the sign language recognition data videos according to different time spans, and extracting multi-scale sign language visual features; training the continuous sign language recognition and segmentation system using the sign language word text features and the multi-scale sign language visual features. This application allows for precise characterization of sign language actions of each temporal length, achieving refined sign language recognition and segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human intelligence, and in particular to a method and apparatus for continuous sign language recognition and word segmentation. Background Technology

[0002] Sign language is a visual language and a primary means of daily communication for the hearing impaired. It uses hand and other body movements, including gestures and their trajectories, facial and oral expressions, and head and body movements, to collaboratively convey information. However, sign language has a different grammatical structure and expression than natural spoken language, making effective communication between the hearing impaired and hearing people difficult in daily life. As a core research area for using artificial intelligence to promote barrier-free communication between the hearing impaired and hearing people, continuous sign language recognition (CSLR) utilizes computer vision and natural language processing technologies to sequentially identify multiple sign language words in a sign language video.

[0003] To effectively capture sign language movements, models need to build an effective temporal receptive field to extract temporal features. Existing technologies employ the following methods: 1) Combining 2D convolutional neural networks (CNNs) with long short-term memory (LSTM) networks, or combining 3D convolutional neural networks (CNNs) with dilated convolutional models, more effectively increases the temporal receptive field of the visual feature extraction network, focusing only on the extraction of long-term visual information in sign language. 2) Since most sign language movements are not performed over long periods, combining 2D convolutional neural networks (CNNs) with temporal convolutional neural networks (1D-TCNs) builds a short-term temporal receptive field, focusing only on the extraction of visual information from the majority of short-term sign language movements. 3) To capture sign language movements more comprehensively, many methods combine 2D convolutional neural networks (CNNs) with temporal convolutional neural networks (1D-TCNs) and long short-term memory (LSTM) networks. This aims to build long- and short-term temporal receptive fields, enabling simultaneous extraction of visual information from both long-term and short-term sign language movements. 4) CTC (Concurrent Transcriptional Transformation) is used to maximize the sum of probabilities of all feasible alignment paths between video frames and sign language words in a sentence, thereby obtaining a probability label for each video frame. This allows for end-to-end training of the model in a fully supervised manner.

[0004] Although current methods extract visual information from sign language movements by combining long and short temporal receptive fields, the constructed temporal receptive fields are fixed. This limits the extraction results to these two types of receptive fields and fails to finely characterize sign language movements of every temporal length. Therefore, they face the problem of effectively capturing sign language movements of arbitrary temporal lengths. Furthermore, since most sign language movements have similar appearances and trajectories, training the model solely using CTC also faces the problem of inaccurate probability labels for similar sign language movements, making it difficult to effectively capture sign language movements of multiple temporal lengths. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a continuous sign language recognition and word segmentation method and apparatus to finely characterize sign language actions of each temporal length and to finely segment sign language actions.

[0006] This application provides a continuous sign language recognition and word segmentation method, applied to a continuous sign language recognition and word segmentation system. The continuous sign language recognition and word segmentation system includes a text extraction model and a parallel multi-scale visual feature extraction model, specifically including the following steps:

[0007] The continuous sign language recognition dataset is input into a text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset;

[0008] The sign language recognition data video is determined using a continuous sign language recognition dataset. The sign language recognition data video is then input into the parallel multi-scale visual feature extraction model to segment the sign language recognition data video according to different time spans and extract multi-scale sign language visual features.

[0009] The continuous sign language recognition and segmentation system is trained using the text features of the sign language words and the visual features of the multi-scale sign language.

[0010] The technical solution of this application can precisely depict sign language movements of each temporal length and perform fine word segmentation of sign language movements.

[0011] One possible approach is that the text extraction model includes a text feature extraction sub-model and a mapping sub-model.

[0012] The steps of inputting the continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset include:

[0013] The continuous sign language recognition dataset is input into the text feature extraction sub-model to extract continuous sign language text features.

[0014] The continuous sign language text features are input into the mapping sub-model, the continuous sign language text features are dimensionally transformed, and the sign language word text features are output.

[0015] One possible approach is to utilize a continuous sign language recognition dataset to determine sign language recognition data videos, input these videos into the parallel multi-scale visual feature extraction model, and segment the video according to different time spans to extract multi-scale sign language visual features.

[0016] The parallel multi-scale visual feature extraction model includes ResNet18G. sp (.;θ sp Parallel multi-scale temporal network G hpt (.;θ hpt ) and video frame sequence network G se (.;θ se ).

[0017] One possible approach is that the parallel multi-scale temporal network comprises an H-layer parallel temporal network;

[0018] The operation of the H-layer parallel temporal network is as follows:

[0019]

[0020] —The r-th one-dimensional dilated convolutional layer in the H-th layer of the parallel temporal network;

[0021] —Input features of the Hth layer parallel temporal network, f in =f Sp This represents the input features of the first PT network structure;

[0022] —Output characteristics of the Hth layer parallel temporal network;

[0023] —Output features of a single one-dimensional dilated convolutional layer in the Hth layer of the parallel temporal network;

[0024] * — Convolution operation;

[0025] W r ∈ d×3 ,b r ∈ d —Refers to the weights and biases of the dilated convolutional layer.

[0026] d — Feature dimension;

[0027] R 1×1 —A 1D convolutional layer with a kernel of 1;

[0028] BN—Batch Normalization Layer;

[0029] ReLU — ReLU activation function;

[0030] —Multi-scale sign language temporal visual features.

[0031] One possible approach is that the video frame sequence network G se (.;θ se It consists of BI-GRU units and a fully connected layer, with the following specific structure:

[0032] f GRU =G se (f HPT );

[0033] —The output features of BI-GRU represent the visual features of sign language that integrate multi-scale temporal and sequence information from sign language videos;

[0034] f Cls =G se (f HPT ) = Fc(Bigru(f HPT ));

[0035] —The category probability matrix of class |C|, where |C| represents the total number of words in the sign language corpus;

[0036] Bigru—BI-GRU layer;

[0037] Fc — Fully connected layer;

[0038] One possible approach is that, in the step of training the continuous sign language recognition and segmentation system using the textual features of the sign language words and the multi-scale visual features of sign language,

[0039] A target function is constructed and incorporated into the text extraction model and the multi-scale visual feature extraction model to train the continuous sign language recognition and word segmentation system.

[0040] One possible approach is that the objective function includes:

[0041] The steps for constructing cross-modal alignment constraints and CTC objective functions include:

[0042] The objective function is:

[0043]

[0044] λ — a hyperparameter controlling the contribution of the cross-modal alignment constraint component.

[0045] θ sp —The ResNet18 parameters of the multi-scale visual feature extraction model;

[0046] θ hpt —Parameters of the multi-scale visual feature extraction model;

[0047] θ se —Parameters of the BI-GRU unit and fully connected layer in the video frame sequence network;

[0048] L ctc —CTC objective function.

[0049] L sDTW —Cross-modal alignment constraint cost function.

[0050] One possible approach is that the CTC objective function includes:

[0051] L CTC = -log p(Y|X);

[0052] kog p(Y|X) — The sum of probabilities of all feasible alignment paths given X;

[0053] The conditional probability of the alignment path π is calculated as follows:

[0054]

[0055] π—the set of alignment paths between all frames X in the video and their corresponding words;

[0056] C—Number of all word categories in the continuous sign language recognition dataset;

[0057] blank—blank category;

[0058] Given X, the sum of probabilities of all feasible alignment paths is obtained using the following formula.

[0059]

[0060] Β — A mapping that removes duplicate labels and blank classes from π.

[0061] One possible approach is that the cross-modal alignment constraint cost function includes:

[0062] —Sign language word text features;

[0063] —Visual features of sign language;

[0064] L sDTW —Cost function;

[0065] D i,j =d i,j +min(D i-1,j D i,j-1 D i-1,j-1 ), i∈L, j∈T

[0066]

[0067] In order to make D i,j The computation is differentiable, which introduces the minimum operator min. γ (a1,...,a n );

[0068]

[0069] γ—smoothing coefficient;

[0070] Secondly, this application provides a continuous sign language recognition and word segmentation device, applied to a continuous sign language recognition and word segmentation system, wherein the continuous sign language recognition and word segmentation system includes a text extraction model and a line multi-scale visual feature extraction model, including:

[0071] The first extraction module is used to input the continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset.

[0072] The second extraction module is used to determine the sign language recognition data video using the continuous sign language recognition dataset, and input the sign language recognition data video into the parallel multi-scale visual feature extraction model to segment the sign language recognition data video according to different time spans and extract multi-scale sign language visual features.

[0073] Training module: used to train the continuous sign language recognition and segmentation system using the text features of the sign language words and the visual features of the multi-scale sign language.

[0074] The embodiments of the present invention bring about the following beneficial effects:

[0075] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0076] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0077] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0078] Figure 1 A flowchart of a continuous sign language recognition and word segmentation method provided in this application embodiment;

[0079] Figure 2 This is a structural diagram of a continuous sign language recognition and word segmentation device provided in an embodiment of this application. Detailed Implementation

[0080] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0081] Currently, methods combining long and short temporal receptive fields are used to extract visual information from sign language movements. However, these constructed temporal receptive fields are fixed, limiting the extraction results to these two types of receptive fields and failing to finely characterize sign language movements of every temporal length. Therefore, they face the challenge of effectively capturing sign language movements of arbitrary temporal lengths. Furthermore, since most sign language movements have similar appearances and trajectories, training the model solely using CTC also faces the problem of inaccurate probability labels for similar sign language movements, making it difficult to effectively capture sign language movements of multiple temporal lengths.

[0082] Based on this, the present invention provides a continuous sign language recognition and word segmentation method and apparatus that can capture sign language actions of different time lengths and perform accurate word segmentation of sign language.

[0083] To facilitate understanding of this embodiment, a detailed description of the continuous sign language recognition and word segmentation method disclosed in this embodiment of the invention will be provided first.

[0084] Please refer to Figure 1 , Figure 1 A flowchart of a continuous sign language recognition and word segmentation method is provided for embodiments of this application. This method is applied to a continuous sign language recognition and word segmentation system, which includes a text extraction model and a line multi-scale visual feature extraction model. Specifically, it includes the following steps:

[0085] S101: Input the continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset;

[0086] Specifically, in this step, the text extraction model includes: a text feature extraction sub-model and a mapping sub-model. The step of inputting the continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset includes:

[0087] The continuous sign language recognition dataset is input into the text feature extraction sub-model to extract continuous sign language text features.

[0088] The continuous sign language text features are input into the mapping sub-model, the continuous sign language text features are dimensionally transformed, and the sign language word text features are output.

[0089] Furthermore, a method that can be implemented by those skilled in the art is to construct a text feature extraction sub-model using a BILSTM model, while simultaneously constructing a mapping sub-model using three fully connected layers;

[0090] In other words, the mapping sub-model includes:

[0091] o mlp1 =f mlp1 (W mlp1 ·o t +b mlp1 );

[0092] o mlp2 =f mlp2 (W mlp2 ·o mlp1 +b mlp2 );

[0093] o mlp3 =f mlp3 (W mlp3 ·o mlp2 +b mlp3 ).

[0094] W mlp —Parameters of the mapping sub-model;

[0095] b mlp —The bias of the mapping sub-model;

[0096] o mlp — Output of the mapping sub-model;

[0097] S102: Use the continuous sign language recognition dataset to determine the sign language recognition data video, and input the sign language recognition data video into the parallel multi-scale visual feature extraction model to segment the sign language recognition data video according to different time spans and extract multi-scale sign language visual features.

[0098] It should be noted that different time spans represent different time intervals. For example, assuming a video is 100 seconds long, in this application, the video can be segmented into time spans of 2 seconds, 5 seconds, 10 seconds, and 20 seconds.

[0099] In other words, the 100-second video is segmented every two seconds, then every five seconds, and then every ten seconds. Of course, this is merely an example; in this application, the time span is set to k+2. τ -1, {τ0,...,τ r-1},τ r-1 =2 r-1, r represents the r-th 1D dilated convolutional layer with kernel k in the H-th parallel temporal network. Based on experience, the range of r is set to r = 1, 2, 3, 4, 5;

[0100] This method allows for the segmentation of continuous sign language videos according to different time spans, enabling precise characterization of sign language actions for each time length during subsequent training, and fine-grained word segmentation of sign language actions.

[0101] Optionally, in this step, the parallel multi-scale visual feature extraction model includes ResNet18G. sp (.;θ sp Parallel multi-scale temporal network G hpt (.;θ hpt ) and video frame sequence network G se (.;θ se ).

[0102] One possible approach is that the parallel multi-scale temporal network comprises an H-layer parallel temporal network;

[0103] The operation of the H-layer parallel temporal network is as follows:

[0104]

[0105] —The r-th one-dimensional dilated convolutional layer in the H-th layer of the parallel temporal network;

[0106] —Input features of the Hth layer parallel temporal network, f in =f Sp This represents the input features of the first PT network structure;

[0107] —Output characteristics of the Hth layer parallel temporal network;

[0108] —Output features of a single one-dimensional dilated convolutional layer in the Hth layer of the parallel temporal network;

[0109] * — Convolution operation;

[0110] W r ∈ d×3 ,b r ∈ d —Refers to the weights and biases of the dilated convolutional layer.

[0111] d —Feature dimension;

[0112] R 1×1 —A 1D convolutional layer with a kernel of 1;

[0113] BN—Batch Normalization Layer;

[0114] ReLU — ReLU activation function;

[0115] —Multi-scale sign language temporal visual features.

[0116] Meanwhile, the video frame sequence network G se (.;θ se It consists of BI-GRU units and a fully connected layer, with the following specific structure:

[0117] f GRU =G se (f HPT );

[0118] —The output features of BI-GRU represent the visual features of sign language that integrate multi-scale temporal and sequence information from sign language videos;

[0119] f Cls =G se (f HPT ) = Fc(Bigru(f HPT ));

[0120] —The category probability matrix of class |C|, where |C| represents the total number of words in the sign language corpus;

[0121] Bigru—BI-GRU layer;

[0122] Fc — Fully connected layer;

[0123] S103: The continuous sign language recognition and segmentation system is trained using the text features of the sign language words and the visual features of the multi-scale sign language.

[0124] Optionally, in this step, an objective function is constructed and added to the text extraction model and the multi-scale visual feature extraction model to train the continuous sign language recognition and word segmentation system.

[0125] Specifically, the objective function includes:

[0126] The steps for constructing cross-modal alignment constraints and CTC objective functions include: the objective function is:

[0127]

[0128] λ — a hyperparameter controlling the contribution of the cross-modal alignment constraint component.

[0129] θ sp—The ResNet18 parameters of the parallel multi-scale visual feature extraction model;

[0130] θ hpt —Parameters of the parallel multi-scale visual feature extraction model;

[0131] θ se —Parameters of the BI-GRU unit and fully connected layer in the video frame sequence network; L ctc —CTC objective function.

[0132] L sDTW —Cross-modal alignment constraint cost function.

[0133] One possible approach is that the CTC objective function includes:

[0134] L CTC = -log p(Y|X);

[0135] log p(Y|X) — the sum of probabilities of all feasible alignment paths given X;

[0136] The conditional probability of the alignment path π is calculated as follows:

[0137]

[0138] π—the set of alignment paths between all frames X in the video and their corresponding words;

[0139] C—Number of all word categories in the continuous sign language recognition dataset;

[0140] blank—blank category;

[0141] Given X, the sum of probabilities of all feasible alignment paths is obtained using the following formula.

[0142] Β — A mapping that removes duplicate labels and blank classes from π.

[0143] One possible approach is that the cross-modal alignment constraint cost function includes:

[0144] —Sign language word text features;

[0145] —Visual features of sign language;

[0146] L sDTW —Cost function;

[0147] D i,j =d i,j +min(Di-1,j D i,j-1 D i-1,j-1 ), i∈L, j∈T

[0148]

[0149] In order to make D i,j The computation is differentiable, which introduces the minimum operator min. γ (a1,...,a n );

[0150]

[0151] γ—smoothing coefficient;

[0152] Through steps S101 to S103, the continuous sign language video can be segmented according to different time spans, expanding the dataset. Using this as the dataset improves accuracy, finely characterizes sign language actions of each time length, and performs fine word segmentation of sign language actions.

[0153] Figure 2 This application provides a structural diagram of a continuous sign language recognition and word segmentation device, applied to a continuous sign language recognition and word segmentation system. The continuous sign language recognition and word segmentation system includes a text extraction model and a line multi-scale visual feature extraction model, comprising:

[0154] The first extraction module is used to input the continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset.

[0155] The second extraction module is used to determine the sign language recognition data video using the continuous sign language recognition dataset, and input the sign language recognition data video into the parallel multi-scale visual feature extraction model to segment the sign language recognition data video according to different time spans and extract multi-scale sign language visual features.

[0156] Training module: used to train the continuous sign language recognition and segmentation system using the text features of the sign language words and the visual features of the multi-scale sign language.

[0157] The following is a specific example of this application. In the embodiments provided by this invention, a multi-scale visual feature extraction and cross-modal alignment model is trained. First, data augmentation is performed on the samples, and all video frames are resized, randomly cropped at multiple scales, and randomly horizontally flipped. In the resizing, the video frame size is redefined to 256×256 (width * height). Multi-scale cropping is performed based on the resized frame, randomly selecting a cropping ratio from (1.0, 0.875, 0.75, 0.66), multiplying it by the target cropping size (224×224), and then generating a cropping region for cropping. Finally, the video frame size is redefined to 224×224. Random horizontal flipping involves flipping all video frames horizontally by 180 degrees with a 50% probability. During testing, this invention only redefines the size (width * height) of the video frame to 256×256 and performs center cropping of the cropping area to (224×224). Center cropping is based on the redefinition size (width * height) of 256×256, and calculates and crops the center area of ​​the video frame according to the final video frame size (224×224).

[0158] To verify the effectiveness of this invention, it was evaluated on the continuous sign language recognition datasets RWTH-2014, RWTH-2014T, and CSL500. This invention uses CTC Beamsearch to predict the sign language sentences corresponding to sign language videos and employs the word error rate (WER) as an evaluation metric, which is widely used in continuous sign language recognition tasks [5,6,7,12,13]. WER measures the minimum number of "insertion," "deletion," and "replacement" operations required to convert a predicted sentence into a standard reference sentence; a lower WER indicates better evaluation performance.

[0159]

[0160] Where, n I ,n D ,n S These represent the number of "insert", "delete", and "replace" operations, respectively, and L is the number of words in the standard reference sentence.

[0161] To effectively mitigate video frame redundancy during training, similar to the frame selection strategy of SFL, this invention uniformly and randomly extracts half of all video frames for each video during training on the RWTH-2014 and RWTH-2014T datasets. For the CSL500 dataset, 20 frames are uniformly and randomly extracted from all video segments during training. For these three datasets, the experiment consisted of 80 training epochs. The model used the Adam optimizer with an initial learning rate of 1e–4, multiplied by 0.1 at the 30th and 60th epochs, and the hyperparameter λ was set to 5e-5. The training batch size was set to 4. During model testing, the beam width of CTC Beamsearch was set to 10.

[0162] The performance comparison of other advanced continuous sign language recognition algorithms with the present invention is shown in Tables 1, 2 and 3. The ablation experiment of the multi-scale visual feature extraction and cross-modal alignment model on the RWTH-2014 dataset is shown in Table 4.

[0163] Table 1 compares the results with other continuous sign language recognition algorithms on the RWTH-2014 dataset.

[0164]

[0165] Table 2 compares the results with other continuous sign language recognition algorithms on the RWTH-2014T dataset.

[0166]

[0167] Table 3 compares the results with other continuous sign language recognition algorithms on the CSL500 Split-II dataset.

[0168]

[0169] Table 4 Ablation experiments of the multi-scale visual feature extraction and cross-modal alignment model on the RWTH-2014 dataset.

[0170]

[0171] As shown in Tables 1, 2, and 3, the multi-scale visual feature extraction and cross-modal alignment model proposed in this invention exhibits advanced recognition performance on multiple publicly available continuous sign language recognition datasets. It has been demonstrated that the parallel multi-scale visual feature extraction model HPT proposed in this invention can effectively capture sign language actions of different temporal lengths in a sign language video, resulting in a significant improvement in word segmentation performance. Furthermore, the cross-modal alignment constraint proposed in this invention can effectively align the visual features of video frames with their corresponding sign language word features in a high-dimensional feature space, effectively improving the generalization ability of video visual features.

[0172] The device provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.

[0173] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0174] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0175] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A continuous sign language recognition and word segmentation method, characterized in that, An application is made in a continuous sign language recognition and word segmentation system, which includes a text extraction model and a parallel multi-scale visual feature extraction model, specifically including the following steps: The continuous sign language recognition dataset is input into a text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset; The sign language recognition data video is determined using a continuous sign language recognition dataset. The sign language recognition data video is then input into the parallel multi-scale visual feature extraction model to segment the sign language recognition data video according to different time spans and extract multi-scale sign language visual features. The continuous sign language recognition and segmentation system is trained using the text features of the sign language words and the visual features of the multi-scale sign language. The text extraction model includes: a text feature extraction sub-model and a mapping sub-model; The steps of inputting the continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset include: The continuous sign language recognition dataset is input into the text feature extraction sub-model to extract continuous sign language text features; The continuous sign language text features are input into the mapping sub-model, the continuous sign language text features are dimensionally transformed, and the sign language word text features are output. In the step of determining sign language recognition data video using a continuous sign language recognition dataset, and inputting the sign language recognition data video into the parallel multi-scale visual feature extraction model to segment the sign language recognition data video according to different time spans and extract multi-scale sign language visual features, the parallel multi-scale visual feature extraction model includes ResNet18. Parallel multi-scale temporal networks and video frame sequence network ; The parallel multi-scale temporal network comprises an H-layer parallel temporal network; The operation of the H-layer parallel temporal network is as follows: ; —The r-th one-dimensional dilated convolutional layer in the H-th layer of the parallel temporal network; —Input features of the Hth layer parallel temporal network, This represents the input features of the first PT network structure; —Output characteristics of the Hth layer parallel temporal network; —Output features of a single one-dimensional dilated convolutional layer in the Hth layer of the parallel temporal network; — Convolution operation; —Refers to the weights and biases of the dilated convolutional layer. —Feature dimension; —A 1D convolutional layer with a kernel of 1; —Batch normalization layer; —ReLU activation function; —Multi-scale visual features of sign language; The video frame sequence network It consists of BI-GRU units and a fully connected layer, with the following specific structure: ; —The output features of BI-GRU represent the visual features of sign language that integrate multi-scale temporal and sequential information from sign language videos; ; — The class probability matrix, This represents the total number of words in the sign language corpus. —BI-GRU layer; —Fully connected layer.

2. The method according to claim 1, characterized in that, In the step of training the continuous sign language recognition and segmentation system using the textual features of the sign language words and the visual features of the multi-scale sign language, An objective function is constructed and incorporated into the text extraction model and the parallel multi-scale visual feature extraction model to train the continuous sign language recognition and word segmentation system.

3. The method according to claim 2, characterized in that, The objective function includes: The steps for constructing cross-modal alignment constraints and CTC objective functions include: The objective function is: ; —A hyperparameter that controls the contribution of cross-modal alignment constraint components. —The ResNet18 parameters in the parallel multi-scale visual feature extraction model; —Parameters of the parallel multi-scale temporal network in the parallel multi-scale visual feature extraction model; —Parameters of the BI-GRU unit and fully connected layer in the video frame sequence network; —CTC objective function; —Cross-modal alignment constraint cost function.

4. The method according to claim 3, characterized in that, The CTC objective function includes: ; —The sum of probabilities of all feasible alignment paths given X; Among them, alignment path The conditional probability is calculated as follows: ; —The set of alignment paths of all frames X in the video with their corresponding words; —Number of all word categories in the continuous sign language recognition dataset; —Blank category; Given X, the sum of probabilities of all feasible alignment paths is obtained using the following formula. ; -Will Duplicate tags and The mappings removed from the class.

5. The method according to claim 4, characterized in that... The cross-modal alignment constraint cost function includes: ; —Sign language word text features; —Visual features of sign language; —Cost function; ; ; In order to make The computation is differentiable, which introduces the minimum operator. ; ; —Smoothing coefficient.

6. A continuous sign language recognition and word segmentation device, characterized in that, Applied to continuous sign language recognition and word segmentation system, The continuous sign language recognition and word segmentation system includes a text extraction model and a parallel multi-scale visual feature extraction model, including: The first extraction module is used to input the continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset. The second extraction module is used to determine the sign language recognition data video using the continuous sign language recognition dataset, and input the sign language recognition data video into the parallel multi-scale visual feature extraction model to segment the sign language recognition data video according to different time spans and extract multi-scale sign language visual features. Training module: used to train the continuous sign language recognition and segmentation system using the text features of the sign language words and the visual features of the multi-scale sign language; The text extraction model includes: a text feature extraction sub-model and a mapping sub-model; The steps of inputting the continuous sign language recognition dataset into the text extraction model to extract the text features of sign language words in the continuous sign language recognition dataset include: The continuous sign language recognition dataset is input into the text feature extraction sub-model to extract continuous sign language text features; The continuous sign language text features are input into the mapping sub-model, the continuous sign language text features are dimensionally transformed, and the sign language word text features are output. In the step of determining sign language recognition data video using a continuous sign language recognition dataset, and inputting the sign language recognition data video into the parallel multi-scale visual feature extraction model to segment the sign language recognition data video according to different time spans and extract multi-scale sign language visual features, the parallel multi-scale visual feature extraction model includes ResNet18. Parallel multi-scale temporal networks and video frame sequence network ; The parallel multi-scale temporal network comprises an H-layer parallel temporal network; The operation of the H-layer parallel temporal network is as follows: ; —The r-th one-dimensional dilated convolutional layer in the H-th layer of the parallel temporal network; —Input features of the Hth layer parallel temporal network, This represents the input features of the first PT network structure; —Output characteristics of the Hth layer parallel temporal network; —Output features of a single one-dimensional dilated convolutional layer in the Hth layer of the parallel temporal network; — Convolution operation; —Refers to the weights and biases of the dilated convolutional layer. —Feature dimension; —A 1D convolutional layer with a kernel of 1; —Batch normalization layer; —ReLU activation function; —Multi-scale visual features of sign language; The video frame sequence network It consists of BI-GRU units and a fully connected layer, with the following specific structure: ; —The output features of BI-GRU represent the visual features of sign language that integrate multi-scale temporal and sequential information from sign language videos; ; — The class probability matrix, This represents the total number of words in the sign language corpus. —BI-GRU layer; —Fully connected layer.

Citation Information

Patent Citations

  • Sign language recognition method and system

    CN111340006A

  • Scene character recognition system and method based on parallel iterative imitation decoding

    CN113963340A