Sign language data augmentation method and device, electronic equipment and storage medium

By augmenting multiple isolated word sign language data and combining action continuity and spatial difference loss functions, continuous sentence sign language data is generated, which solves the problem of ineffective augmentation in existing technologies and improves the robustness and training effect of sign language recognition models.

CN115909502BActive Publication Date: 2025-11-18IFLYTEK SOUTH CHINA ARTIFICIAL INTELLIGENCE RES INST GUANGZHOU CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211635638.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-11-18
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

Existing technologies cannot effectively augment new continuous sign language data, resulting in limited improvement in the training effect of sign language recognition models. Furthermore, traditional data augmentation methods may lead to overfitting problems.

Method used

By identifying multiple isolated word sign language data, augmentation processing is performed using a data augmentation model. Combining the action continuity scoring loss function and the spatial difference loss function, continuous sentence sign language data is generated to ensure coherence in both the temporal and spatial dimensions.

Benefits of technology

It improves the efficiency of sign language data acquisition, and the generated continuous sentence sign language data is highly correlated with actual sign language actions, thereby enhancing the robustness and training effect of the sign language recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115909502B_ABST
    Figure CN115909502B_ABST
Patent Text Reader

Abstract

The application provides a sign language data augmentation method and device, electronic equipment and a storage medium. The method comprises: determining a plurality of isolated word sign language data, any isolated word sign language data comprising a plurality of image frames; performing augmentation processing on the plurality of isolated word sign language data based on a data augmentation model to obtain continuous sentence sign language data, wherein the loss function of the data augmentation model is determined based on an action continuity scoring loss function and / or a spatial difference loss function. The method, device, electronic equipment and storage medium provided by the application can generate a large amount of continuous sentence sign language data by only collecting isolated word sign language data, efficiently expand the continuous sentence sign language data, and train the data augmentation model through the action continuity scoring loss function and / or the spatial difference loss function, thereby improving the effectiveness of the continuous sentence sign language data, so that the model training effect can be improved when training a sign language recognition model based on the continuous sentence sign language data, and the robustness of the sign language recognition model is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, electronic device, and storage medium for sign language data augmentation. Background Technology

[0002] With the rapid development of artificial intelligence technology, the application scenarios of sign language recognition are becoming increasingly widespread. Sign language recognition is based on a sign language recognition model, which performs sign language recognition on collected sign language videos to obtain the sign language recognition results; in order to improve the recognition accuracy of the sign language recognition model, a massive amount of sign language data is needed as sample data to train, test and validate the sign language recognition model.

[0003] Currently, most sign language data is augmented in either the temporal or spatial dimensions. Specifically, it scales the data temporally to simulate the speed of sign language movements, or it scales spatially to simulate the differences in limb movements among different individuals. However, augmentation in either the temporal or spatial dimensions can only augment an existing continuous sentence of sign language data; it cannot augment new continuous sentence sign language data. Summary of the Invention

[0004] This invention provides a sign language data augmentation method, apparatus, electronic device, and storage medium to overcome the shortcomings of existing technologies that cannot augment new continuous sentence sign language data, thereby achieving efficient sign language data augmentation.

[0005] This invention provides a method for augmenting sign language data, comprising:

[0006] Determine multiple isolated word sign language data, each of which includes multiple frames of images;

[0007] Based on the data augmentation model, the sign language data of the multiple isolated words are augmented to obtain sign language data of continuous sentences.

[0008] The loss function of the data augmentation model is determined based on the action continuity scoring loss function and / or the spatial difference loss function.

[0009] The action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data. The spatial difference loss function is determined based on the difference between each two adjacent frames of the sample continuous sentence sign language data. The sample continuous sentence sign language data is obtained by augmenting multiple sample isolated word sign language data using the data augmentation model.

[0010] According to a sign language data augmentation method provided by the present invention, the step of augmenting multiple isolated word sign language data based on a data augmentation model to obtain continuous sentence sign language data includes:

[0011] The sign language data of the multiple isolated words are concatenated to obtain the sign language data of the concatenated sentence;

[0012] Based on the data augmentation model, the spliced ​​sentence sign language data is augmented to obtain continuous sentence sign language data;

[0013] The loss function is determined based on the data difference loss function, the action continuity scoring loss function, and / or the spatial difference loss function.

[0014] The data difference loss function is determined based on the difference between the sample continuous sentence sign language data and the sample spliced ​​sentence sign language data, which is obtained by splicing together the multiple sample isolated word sign language data.

[0015] According to a sign language data augmentation method provided by the present invention, the data difference loss function is determined based on multiple word difference loss values;

[0016] Any of the aforementioned word difference loss values ​​is determined based on the word difference between the target isolated word sign language data in the sample spliced ​​sentence sign language data and the target sign language data corresponding to the target isolated word sign language data in the sample continuous sentence sign language data, wherein the target isolated word sign language data and the target sign language data correspond in the same frame number.

[0017] The word difference is determined based on multiple image differences. Each image difference is determined based on the difference between a first target image frame in the target isolated word sign language data and a second target image frame corresponding to the first target image frame in the target sign language data. The first target image frame and the second target image frame correspond to each other in that they have the same number of frames.

[0018] According to a sign language data augmentation method provided by the present invention, any of the image differences is determined based on the difference between a first target image frame in the target isolated word sign language data and a second target image frame corresponding to the first target image frame in the target sign language data, and the target weighting weight corresponding to the first target image frame;

[0019] The target isolated word sign language data includes first sign language data, second sign language data, and third sign language data, wherein the number of frames in the first sign language data is less than the number of frames in the second sign language data, and the number of frames in the second sign language data is less than the number of frames in the third sign language data;

[0020] The second weighted weight corresponding to the image frame in the second sign language data is greater than the first weighted weight corresponding to the image frame in the first sign language data, and the second weighted weight corresponding to the image frame in the second sign language data is greater than the third weighted weight corresponding to the image frame in the third sign language data.

[0021] According to a sign language data augmentation method provided by the present invention, the target weighting weight is determined based on the number of image frames of the target isolated word sign language data, a preset edge frame control parameter, and the target number of the first target image frame;

[0022] The preset edge frame control parameter is used to control the number of frames in the first sign language data and the third sign language data, and the target frame number is used to characterize the frame time of the first target image frame in the target isolated word sign language data;

[0023] If the target number of frames is less than the first threshold, then the target weight is the first weight.

[0024] If the target number of frames is greater than or equal to the first threshold and the target number of frames is less than or equal to the second threshold, then the target weighting weight is the second weighting weight.

[0025] If the target number of frames is greater than the second threshold, then the target weight is the third weight.

[0026] The first threshold is determined based on a first ratio of the number of image frames to the preset edge frame control parameter. The second threshold is determined based on the product of the number of image frames and a first difference. The first difference is determined based on the difference between the preset parameter and a second ratio. The second ratio is determined based on the ratio of the preset parameter to the preset edge frame control parameter.

[0027] According to a sign language data augmentation method provided by the present invention, the motion continuity scoring result is determined based on the following steps:

[0028] Identify multiple sign language data points to be scored within the sample continuous sentence sign language data;

[0029] Based on the action continuity scoring model, the multiple sign language data to be scored are scored respectively, resulting in multiple scoring results;

[0030] Based on the multiple scoring results, the action continuity scoring result is determined;

[0031] The action continuity scoring model is trained based on sample sign language data to be scored, and the sample sign language data to be scored is determined based on positive and negative sample data.

[0032] According to a sign language data augmentation method provided by the present invention, the positive sample data is obtained by sampling the sign language data of a first isolated word;

[0033] The negative sample data is obtained by replacing at least one image frame in the positive sample data with an image frame to be replaced, wherein the image frame to be replaced is an image frame in the second isolated word sign language data.

[0034] According to a sign language data augmentation method provided by the present invention, the difference between any two adjacent frames in the sample continuous sentence sign language data is determined based on the difference between multiple key points, and the two adjacent frames include a first image frame and a second image frame.

[0035] The difference degree of any of the key points is determined based on the difference degree between the first target key point in the first image frame and the second target key point corresponding to the first target key point in the second image frame. The correspondence between the first target key point and the second target key point is that the limb positions are the same. The first target key point is a posture key point related to sign language actions.

[0036] The present invention also provides a sign language data augmentation device, comprising:

[0037] A determination module is used to determine multiple isolated word sign language data, wherein any isolated word sign language data includes multiple frames of images;

[0038] The processing module is used to perform augmentation processing on the multiple isolated word sign language data based on the data augmentation model to obtain continuous sentence sign language data;

[0039] The loss function of the data augmentation model is determined based on the action continuity scoring loss function and / or the spatial difference loss function.

[0040] The action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data. The spatial difference loss function is determined based on the difference between each two adjacent frames of the sample continuous sentence sign language data. The sample continuous sentence sign language data is obtained by augmenting multiple sample isolated word sign language data using the data augmentation model.

[0041] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the sign language data augmentation method as described above.

[0042] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the sign language data augmentation method as described above.

[0043] The sign language data augmentation method, apparatus, electronic device, and storage medium provided by this invention, based on a data augmentation model, can augment multiple isolated word sign language data to obtain continuous sentence sign language data. This allows for the generation of a large amount of continuous sentence sign language data by only collecting isolated word sign language data, efficiently expanding continuous sentence sign language data, significantly reducing the time and cost of sign language data collection, improving the efficiency of sign language data acquisition, and achieving efficient sign language data augmentation. This improves the training effect of sign language recognition models based on continuous sentence sign language data, thereby enhancing the robustness of the sign language recognition model. Simultaneously, by training the data augmentation model using an action continuity scoring loss function, the series of sign language actions represented by the continuous sentence sign language data generated by the data augmentation model are made more continuous; in other words, the continuous sentence sign language data is made more continuous. The data augmentation model is more coherent in both time and space dimensions, thus exhibiting a stronger correlation with actual sign language movements. This enhances the effectiveness of continuous sentence sign language data, improving the training performance and robustness of sign language recognition models when trained based on it. Furthermore, training the data augmentation model using a spatial difference loss function further strengthens the coherence of the sign language movements represented by the generated continuous sentence sign language data. In other words, it makes the continuous sentence sign language data more coherent in both time and space dimensions, resulting in a stronger correlation with actual sign language movements and thus enhancing its effectiveness. This, in turn, improves the training performance and robustness of sign language recognition models when trained based on continuous sentence sign language data. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0045] Figure 1 One of the flowcharts for the sign language data augmentation method provided by the present invention;

[0046] Figure 2 The second flowchart illustrating the sign language data augmentation method provided by the present invention;

[0047] Figure 3 The third flowchart illustrating the sign language data augmentation method provided by the present invention;

[0048] Figure 4 A schematic diagram of the layout of attitude key points provided by the present invention;

[0049] Figure 5This is a schematic diagram of the structure of the sign language data augmentation device provided by the present invention;

[0050] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0052] Sign language is an important language for communication among deaf people. It mainly relies on hand shapes, hand positions, movements, facial expressions, and body movements to express specific meanings and convey information. Traditionally, the training of sign language interpreters has been used to ensure communication among hearing-impaired individuals and between hearing-impaired and hearing people. However, there is a significant shortage of sign language interpreters, and effective sign language interpretation services are still lacking in most places. Therefore, with the rapid development of computer vision technology, the application scenarios of sign language recognition are becoming increasingly widespread. Using technologies such as deep learning, it enables the translation between sign language and written text, as well as between sign language and speech. Sign language recognition, based on a sign language recognition model, performs sign language recognition on captured sign language videos to obtain the results, thereby achieving barrier-free communication and exchange between hearing-impaired and hearing people.

[0053] Based on the above, in order to improve the recognition accuracy of sign language recognition models, a massive amount of sign language data is needed as sample data for training, testing, and validation of the model. Traditionally, this involves investing significant manpower and time in collecting sign language data, which is extremely time-consuming and labor-intensive. Therefore, sign language data augmentation is necessary.

[0054] Currently, most sign language data augmentation involves scaling the data in either the temporal or spatial dimensions. This means scaling the data temporally to simulate the speed of sign language movements, or scaling it spatially to simulate differences in limbs or torsos among different individuals. However, temporal or spatial augmentation can only augment existing continuous sentence sign language data or existing isolated word sign language data; it cannot augment new continuous sentence sign language data. Furthermore, data augmentation by concatenating multiple isolated word sign language data may result in disjointed continuous sentence sign language data in either the temporal or spatial dimensions, lacking a clear correlation with actual sign language movements. Consequently, training models with sign language data augmented in this way often suffers from severe overfitting, offering limited improvement to the training performance of sign language recognition models.

[0055] To address the above problems, the present invention proposes the following embodiments. Figure 1 This is one of the flowcharts illustrating the sign language data augmentation method provided by the present invention, such as... Figure 1 As shown, this sign language data augmentation method includes:

[0056] Step 110: Determine multiple isolated word sign language data, each of which includes multiple frames of images.

[0057] Here, isolated word sign language data refers to the sign language data corresponding to a single sign language word, for example, the sign language data corresponding to the sign language word "tomorrow". This isolated word sign language data can be obtained from existing corpora.

[0058] In one specific embodiment, the target continuous sentence to be generated is split into sign language words to obtain multiple sign language words; based on these multiple sign language words, the corresponding multiple isolated word sign language data are obtained. For example, if the target continuous sentence is "Friend, are you coming home for dinner tonight?", then splitting the target continuous sentence can yield the sign language words "friend", "today", "evening", "home", and "dinner", thereby obtaining the isolated word sign language data for "friend", "today", "evening", "home", and "dinner". In another embodiment, the multiple isolated word sign language data can also be randomly determined or determined through other methods.

[0059] In some embodiments, isolated word sign language data is sign language video captured by a camera device, that is, isolated word sign language data includes multiple frames of sign language images captured by a camera device.

[0060] The sign language video includes multiple frames of sign language actions. Each frame of the sign language video represents the posture information of the sign language action. This posture information may include, but is not limited to, at least one of the following: hand shape, hand position, limb shape, limb position, facial expression, facial position, etc. Based on this, the information represented by each frame of the video may include, but is not limited to: hand movements, facial movements, limb movements, etc., thereby expressing a sign language word.

[0061] In other embodiments, the isolated word sign language data includes multi-frame pose keypoint maps. These pose keypoint maps can be used to characterize the spatial location of each pose keypoint, i.e., they encompass the spatial coordinates of each pose keypoint. These multi-frame pose keypoint maps are obtained by detecting pose keypoints in each frame of the sign language video corresponding to the isolated word.

[0062] Among them, postural key points are the key points related to sign language movements, that is, the key points in the limbs and trunk that are related to sign language movements. These postural key points play a crucial role in conveying the meaning of sign language. The location or position of each postural key point in the human body can be set according to actual needs.

[0063] In one embodiment, the attitude key point map is a heat map, and the peak position of the heat map is the spatial position of the corresponding attitude key point.

[0064] In other embodiments, the isolated word sign language data includes multi-frame pose maps. These multi-frame pose maps are obtained by classifying and labeling multiple pose keypoints in each pose keypoint map based on the relationships between these keypoints.

[0065] In this context, the pose map is used to represent the pose information and state of the subject in the current frame. This pose information may include, but is not limited to, at least one of the following: hand shape, hand position, limb shape, limb position, facial expression, facial position, etc. Based on this, the information represented by the pose maps in each frame may include, but is not limited to: hand movements, facial movements, limb movements, etc. That is, by processing each pose keypoint map frame by frame, pose stream data formed by each pose map can be obtained. Based on the pose stream data, the sign language expressor's sign language movement state information can be represented. In other words, the pose stream data records the expressor's movement direction, hand shape, and other key information.

[0066] Step 120: Based on the data augmentation model, augmentation processing is performed on the multiple isolated word sign language data to obtain continuous sentence sign language data.

[0067] Considering that a single isolated word sign language dataset can often represent a large number of consecutive sentences, collecting sign language videos for each consecutive sentence would be extremely resource-intensive. Therefore, a data augmentation model can be used to augment consecutive sentence sign language data from multiple isolated word sign language datasets, thereby improving the efficiency of sign language data augmentation.

[0068] Here, the data augmentation model is used to transform multiple isolated word sign language data into coherent continuous sentence sign language data in the time or space dimensions, that is, to generate continuous sentence sign language data with coherent sign language actions, thereby greatly expanding the continuous sentence sign language data. The specific structure of this data augmentation model can be set according to actual needs, and the embodiments of this invention do not impose specific limitations on it. For example, the data augmentation model includes an encoding layer and a decoding layer, which can be composed of LSTM (long short-term memory) network layers. Of course, other network layers can also be used.

[0069] Here, continuous sentence sign language data refers to the sign language data corresponding to a continuous sentence, such as the sign language data corresponding to the continuous sentence "Friend, are you coming home for dinner tonight?". This continuous sentence sign language data can be a continuous sentence sign language video, which includes multiple frames of images of continuous sign language actions. Each frame of the continuous sentence sign language video represents the posture information of the sign language action, which may include, but is not limited to, at least one of the following: hand shape, hand position, limb shape, limb position, facial expression, facial position, etc. Based on this, the information represented by each frame may include, but is not limited to: hand movements, facial movements, limb movements, etc.

[0070] The loss function of the data augmentation model is determined based on the action continuity scoring loss function and / or the spatial difference loss function.

[0071] The action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data. The spatial difference loss function is determined based on the difference between each two adjacent frames of the sample continuous sentence sign language data. The sample continuous sentence sign language data is obtained by augmenting multiple sample isolated word sign language data using the data augmentation model.

[0072] Here, the multiple isolated word sign language data samples are sign language data used to train the data augmentation model. These isolated word sign language data samples are essentially the same as the isolated word sign language data described above, and will not be repeated here.

[0073] Here, the action continuity score is used to characterize the continuity of a series of sign language actions represented by the sample continuous sentence sign language data. It is understandable that the series of sign language actions represented by the sample continuous sentence sign language data generated from multiple isolated word sign language data samples may not be very continuous. Therefore, the data augmentation model is trained using an action continuity score loss function to make the series of sign language actions represented by the continuous sentence sign language data generated by the data augmentation model more continuous, thus making the continuous sentence sign language data more coherent in both time and spatial dimensions. The higher the action continuity score, the smaller the loss value of the action continuity score loss function.

[0074] In one embodiment, the action continuity score is obtained by scoring sample continuous sentence sign language data based on an action continuity scoring model. Furthermore, the continuity of sign language actions considers sign language data within a shorter timeframe. Therefore, multiple sign language data to be scored are first determined from the sample continuous sentence sign language data. Then, based on the action continuity scoring model, each of the multiple sign language data to be scored is scored, resulting in multiple score results. The action continuity score is then determined based on these multiple score results.

[0075] Here, the difference degree is used to characterize the degree of difference between any two adjacent frames in the continuous sign language data of the sample. It can be determined based on the similarity between any two adjacent frames, that is, the higher the similarity, the smaller the difference degree. It can also be determined based on the mean square error between any two adjacent frames. Of course, it can also be determined by other methods, which will not be elaborated here.

[0076] It is understandable that the series of sign language actions represented by the sample continuous sentence sign language data generated based on multiple isolated word sign language data may not be very coherent. Based on this, the data augmentation model is trained by the spatial difference loss function to make the series of sign language actions represented by the continuous sentence sign language data generated by the data augmentation model more coherent, even if the continuous sentence sign language data is more coherent in both time and space dimensions.

[0077] In one embodiment, the loss function of the data augmentation model is determined based on the action continuity rating loss function and the spatial difference loss function. Based on this, the loss value of the loss function is obtained by weighted summation of the loss value of the action continuity rating loss function, the weighted weight corresponding to the action continuity rating loss function, the loss value of the spatial difference loss function, and the weighted weight corresponding to the spatial difference loss function.

[0078] In one embodiment, the continuous sentence sign language data can be used to train a sign language recognition model to improve the model's training performance and thus enhance its robustness. In another embodiment, the continuous sentence sign language data can be used in the field of sign language synthesis, for example, to render the continuous sentence sign language data to generate sign language animation.

[0079] The sign language data augmentation method provided in this invention, based on a data augmentation model, can augment multiple isolated word sign language data to obtain continuous sentence sign language data. This allows for the generation of a large amount of continuous sentence sign language data by only collecting isolated word sign language data, efficiently expanding the continuous sentence sign language data, significantly reducing the time and cost of sign language data collection, improving the efficiency of sign language data acquisition, and achieving efficient sign language data augmentation. This improves the training effect of the sign language recognition model when training based on continuous sentence sign language data, thereby enhancing the robustness of the sign language recognition model. Simultaneously, by training the data augmentation model using an action continuity scoring loss function, the series of sign language actions represented by the continuous sentence sign language data generated by the data augmentation model are made more continuous. In other words, the continuous sentence sign language data is made more continuous in both time and... The data augmentation model is more coherent in the spatial dimension, thus exhibiting a stronger correlation with actual sign language movements. This enhances the effectiveness of continuous sentence sign language data, improving the training performance and robustness of sign language recognition models when trained based on it. Furthermore, training the data augmentation model using a spatial difference loss function further strengthens the coherence of the sign language movements represented by the generated continuous sentence sign language data. In other words, it makes the continuous sentence sign language data more coherent in both the temporal and spatial dimensions, resulting in a stronger correlation with actual sign language movements and further enhancing its effectiveness. This, in turn, improves the training performance and robustness of sign language recognition models when trained based on continuous sentence sign language data.

[0080] Based on the above embodiments, Figure 2 This is a second flowchart illustrating the sign language data augmentation method provided by the present invention, as shown below. Figure 2 As shown, step 120 above includes:

[0081] Step 121: Concatenate the sign language data of the multiple isolated words to obtain concatenated sentence sign language data.

[0082] Here, the concatenated sign language data is sign language data composed of multiple isolated word sign language data. For example, concatenating the isolated word sign language data of "friend," "today," "evening," "go home," and "eat" yields the concatenated sign language data corresponding to "Friend, go home for dinner tonight?" However, the series of sign language actions represented by this concatenated sign language data may not be continuous, and may be disjointed in the time or space dimensions, showing no obvious correlation with actual sign language actions. Based on this, if the concatenated sign language data is used to train the sign language recognition model, there is often a serious overfitting problem, resulting in limited improvement in the training effect of the sign language recognition model.

[0083] In one embodiment, based on the target continuous sentence to be generated, multiple isolated word sign language data are concatenated to obtain concatenated sentence sign language data. For example, if the target continuous sentence is "Friend, are you going home for dinner tonight?", then the isolated word sign language data of "friend", "today", "evening", "home", and "dinner" are concatenated to obtain the concatenated sentence sign language data corresponding to "Friend, are you going home for dinner tonight?".

[0084] In another embodiment, multiple isolated word sign language data are randomly concatenated to obtain concatenated sentence sign language data. For example, randomly concatenating the isolated word sign language data of "friend," "today," "evening," "go home," and "eat" can yield concatenated sentence sign language data corresponding to "Friend, go home for dinner tonight?" and "Tonight, friend, go home for dinner?" and so on.

[0085] Step 122: Based on the data augmentation model, the spliced ​​sentence sign language data is augmented to obtain continuous sentence sign language data.

[0086] Considering that the series of sign language actions represented by spliced ​​sentence sign language data may not be continuous, and may be disjointed in the time or space dimensions, and have no obvious correlation with actual sign language actions, a data augmentation model is used to augment the spliced ​​sentence sign language data, transforming it into continuous sentence sign language data that is coherent in the time or space dimensions. In other words, continuous sentence sign language data with coherent sign language actions is generated to greatly expand the continuous sentence sign language data. In other words, the data augmentation model is used to reconstruct the spliced ​​sentence sign language data to fill in or repair the data and obtain continuous sentence sign language data.

[0087] The loss function is determined based on the data difference loss function, the action continuity scoring loss function, and / or the spatial difference loss function.

[0088] The data difference loss function is determined based on the difference between the sample continuous sentence sign language data and the sample spliced ​​sentence sign language data, which is obtained by splicing together the multiple sample isolated word sign language data.

[0089] Here, the splicing method of the sample spliced ​​sentence sign language data is basically the same as that of the spliced ​​sentence sign language data mentioned above, and will not be repeated here.

[0090] Here, the difference degree is used to characterize the degree of difference between the sample continuous sentence sign language data and the sample spliced ​​sentence sign language data; it can be determined based on the similarity between the sample continuous sentence sign language data and the sample spliced ​​sentence sign language data, that is, the higher the similarity, the smaller the difference degree. Of course, it can also be determined by other methods, which will not be elaborated here.

[0091] Understandably, in order to ensure that the main information of continuous sentence sign language data and spliced ​​sentence sign language data is as similar as possible, and to prevent the data augmentation model from modifying the main information of the spliced ​​sentence sign language data, the data augmentation model is trained using the data dissimilarity loss function to make the main information of continuous sentence sign language data and spliced ​​sentence sign language data as similar as possible. That is, to ensure that the continuous sentence sign language data and spliced ​​sentence sign language data match as closely as possible, without deviating from the target continuous sentence to be generated, i.e., to generate semantically consistent continuous sentence sign language data.

[0092] In one embodiment, the loss function of the data augmentation model is determined based on the action continuity rating loss function and the data dissimilarity loss function. Therefore, the loss value of this loss function is obtained by weighted summation of the loss value of the action continuity rating loss function, the corresponding weighted weights of the action continuity rating loss function, the loss value of the data dissimilarity loss function, and the corresponding weighted weights of the data dissimilarity loss function. It can be understood that the data augmentation model is trained using the action continuity rating loss function and the data dissimilarity loss function to generate coherent and semantically consistent continuous sign language data.

[0093] In another embodiment, the loss function of the data augmentation model is determined based on the spatial difference loss function and the data dissimilarity loss function. Therefore, the loss value of this loss function is obtained by weighted summation of the loss value of the spatial difference loss function, the corresponding weights of the spatial difference loss function, the loss value of the data dissimilarity loss function, and the corresponding weights of the data dissimilarity loss function. It can be understood that the data augmentation model is trained using the spatial difference loss function and the data dissimilarity loss function to generate coherent and semantically consistent continuous sign language data.

[0094] In another embodiment, the loss function of the data augmentation model is determined based on the action continuity rating loss function, the spatial difference loss function, and the data dissimilarity loss function. Therefore, the loss value of this loss function is obtained by weighted summation of the loss value of the action continuity rating loss function, the weighted weights corresponding to the action continuity rating loss function, the loss value of the spatial difference loss function, the weighted weights corresponding to the spatial difference loss function, and the loss value of the data dissimilarity loss function, and the weighted weights corresponding to the data dissimilarity loss function. For example, the loss function of this data augmentation model is as follows:

[0095] Loss = L1 + γ*L2 + K*L3;

[0096] In the formula, L1 represents the loss value of the action continuity rating loss function, L2 represents the loss value of the spatial difference loss function, L3 represents the loss value of the data difference loss function, γ represents the weighted weight of the spatial difference loss function, and k represents the weighted weight of the data difference loss function. The weighted weight of the action continuity rating loss function is 1.

[0097] The sign language data augmentation method provided in this invention, based on a data augmentation model, can augment concatenated sentence sign language data obtained by splicing multiple isolated word sign language data to obtain continuous sentence sign language data. This allows for the generation of a large amount of continuous sentence sign language data by only collecting isolated word sign language data, efficiently expanding the continuous sentence sign language data, greatly reducing the time and cost of sign language data collection, improving the efficiency of sign language data acquisition, and achieving efficient sign language data augmentation. This improves the training effect of the sign language recognition model when training based on continuous sentence sign language data, thereby further enhancing the robustness of the sign language recognition model. Furthermore, by training the data augmentation model using a data dissimilarity loss function, the main information of the continuous sentence sign language data and the concatenated sentence sign language data is made as similar as possible, ensuring that the continuous sentence sign language data and the concatenated sentence sign language data match as closely as possible, generating semantically consistent continuous sentence sign language data. This further enhances the effectiveness of the continuous sentence sign language data, improving the training effect of the sign language recognition model when training based on continuous sentence sign language data, and further enhancing the robustness of the sign language recognition model.

[0098] Based on any of the above embodiments, in this method, the data difference loss function is determined based on multiple word difference loss values;

[0099] Any of the aforementioned word difference loss values ​​is determined based on the word difference between the target isolated word sign language data in the sample spliced ​​sentence sign language data and the target sign language data corresponding to the target isolated word sign language data in the sample continuous sentence sign language data, wherein the target isolated word sign language data and the target sign language data correspond in the same frame number.

[0100] The word difference is determined based on multiple image differences. Each image difference is determined based on the difference between a first target image frame in the target isolated word sign language data and a second target image frame corresponding to the first target image frame in the target sign language data. The first target image frame and the second target image frame correspond to each other in that they have the same number of frames.

[0101] Here, the number of multiple-word dissimilarity loss values ​​is the same as the number of isolated-word sign language data included in the sample concatenated sentence sign language data; that is, the number of multiple-word dissimilarity loss values ​​is the same as the number of isolated-word sign language data. For example, the data dissimilarity loss function is shown below:

[0102] L3=∑ G L i ;

[0103] In the formula, L3 represents the data dissimilarity loss function, G represents the number of isolated word sign language data, and L... i This represents the difference loss value for the i-th word.

[0104] Here, the target isolated word sign language data is a portion of the sign language data in the sample spliced ​​sentence sign language data. This target isolated word sign language data is basically the same as the isolated word sign language data before splicing, and will not be described in detail here.

[0105] Here, the target sign language data is a portion of the sign language data in the sample continuous sentence sign language data. This target sign language data may be different from the isolated word sign language data before splicing.

[0106] Here, word dissimilarity is used to characterize the degree of difference between the target isolated word sign language data and the target sign language data; it can be determined based on the similarity between the target isolated word sign language data and the target sign language data, that is, the higher the similarity, the smaller the word dissimilarity. Of course, it can also be determined by other methods, which will not be elaborated here.

[0107] Here, the number of image dissimilarity values ​​is the same as the number of image frames included in the target isolated word sign language data. For example, the formula for calculating the word dissimilarity loss value is shown below:

[0108] L M =∑ g L m ;

[0109] In the formula, L M L represents the word difference loss value, g represents the number of differences among multiple images, and L represents the word difference loss value. m This represents the difference of the m-th image.

[0110] Based on the above, an exemplary data dissimilarity loss function is shown below:

[0111] L3=∑ G ∑ g L m ;

[0112] In the formula, L3 represents the data dissimilarity loss function, G represents the number of isolated word sign language data, g represents the number of image frames included in each isolated word sign language data, and L... m This represents the difference of the m-th image.

[0113] Here, the difference between the first target image frame and the second target image frame is used to characterize the degree of difference between the first target image frame and the second target image frame; it can be determined based on the similarity between the first target image frame and the second target image frame, that is, the higher the similarity, the smaller the difference. Of course, it can also be determined by other methods, which will not be elaborated here.

[0114] Based on the above, an exemplary data dissimilarity loss function is shown below:

[0115]

[0116] In the formula, L3 represents the data dissimilarity loss function, G represents the number of isolated word sign language data points, and g represents the number of image frames included in each isolated word sign language data point. This represents the second target image frame. This represents the first target image frame.

[0117] The sign language data augmentation method provided in this invention uses a data difference loss function determined based on multiple word difference loss values. This function determines the difference between continuous sentence sign language data and concatenated sentence sign language data at the word level, thereby ensuring that the continuous sentence sign language data and concatenated sentence sign language data match as closely as possible. This generates semantically consistent continuous sentence sign language data, further improving the effectiveness of the continuous sentence sign language data. When training a sign language recognition model based on this data, the model training effect can be improved, thus further enhancing the robustness of the sign language recognition model.

[0118] Based on any of the above embodiments, any of the image differences is determined based on the difference between the first target image frame in the target isolated word sign language data and the second target image frame corresponding to the first target image frame in the target sign language data, and the target weighting weight corresponding to the first target image frame;

[0119] The target isolated word sign language data includes first sign language data, second sign language data, and third sign language data, wherein the number of frames in the first sign language data is less than the number of frames in the second sign language data, and the number of frames in the second sign language data is less than the number of frames in the third sign language data;

[0120] The second weighted weight corresponding to the image frame in the second sign language data is greater than the first weighted weight corresponding to the image frame in the first sign language data, and the second weighted weight corresponding to the image frame in the second sign language data is greater than the third weighted weight corresponding to the image frame in the third sign language data.

[0121] It should be noted that, considering the isolated sign language information represented by the target isolated word sign language data is mainly represented by the intermediate segment data of the target isolated word sign language data, this intermediate segment data is designated as the second sign language data, and the two edge data of the target isolated word sign language data are designated as the first sign language data and the third sign language data, respectively. To weaken the constraint of the data difference loss function on the edge transition image frames of the edge data, the weighting weights corresponding to the image frames of the intermediate segment data are made larger, while the weighting weights corresponding to the edge transition image frames are made smaller. In other words, the second weighting weight is greater than the first weighting weight, and the second weighting weight is greater than the third weighting weight.

[0122] Understandably, by using a data augmentation model, the spliced ​​sign language data is reconstructed to fill in or repair the first and third sign language data, thus obtaining continuous sign language data.

[0123] For example, the formula for calculating image difference is shown below:

[0124]

[0125] In the formula, L N Let w(t,g) represent the image difference degree, w(t,g) represent the target weighting weight, t represent the number of frames in the first target image, and g represent the number of image frames included in the target isolated word sign language data. This represents the second target image frame. This represents the first target image frame.

[0126] Based on the above, an exemplary data dissimilarity loss function is shown below:

[0127]

[0128] In the formula, L3 represents the data dissimilarity loss function, G represents the number of isolated word sign language data, g represents the number of image frames included in each isolated word sign language data, w(t,g) represents the weighting weight, and t represents the number of target image frames. This represents the second target image frame. This represents the first target image frame.

[0129] The sign language data augmentation method provided in this invention further determines the image dissimilarity based on the target weighting weight corresponding to the first target image frame. This allows the data dissimilarity loss function to have different constraints on different image frames. Considering that the isolated word sign language information represented by the target isolated word sign language data is mainly represented by the second sign language data of the target isolated word sign language data, the second weighting weight corresponding to the image frame in the second sign language data is greater than the first weighting weight corresponding to the image frame in the first sign language data, and the second weighting weight corresponding to the image frame in the second sign language data is greater than the third weighting weight corresponding to the image frame in the third sign language data. This weakens the constraint of the data dissimilarity loss function on the first and third sign language data, thereby focusing the training objective of the data augmentation model on the second sign language data. This improves the model training efficiency and the training effect of the data dissimilarity loss function on the data augmentation model, further enhancing the effectiveness of continuous sentence sign language data. When training a sign language recognition model based on continuous sentence sign language data, this method can improve the model training effect and further enhance the robustness of the sign language recognition model.

[0130] Based on any of the above embodiments, the target weighting is determined based on the number of image frames of the target isolated word sign language data, the preset edge frame control parameters, and the target number of the first target image frame;

[0131] The preset edge frame control parameter is used to control the number of frames in the first sign language data and the third sign language data, and the target frame number is used to characterize the frame time of the first target image frame in the target isolated word sign language data;

[0132] If the target number of frames is less than the first threshold, then the target weight is the first weight.

[0133] If the target number of frames is greater than or equal to the first threshold and the target number of frames is less than or equal to the second threshold, then the target weighting weight is the second weighting weight.

[0134] If the target number of frames is greater than the second threshold, then the target weight is the third weight.

[0135] The first threshold is determined based on a first ratio of the number of image frames to the preset edge frame control parameter. The second threshold is determined based on the product of the number of image frames and a first difference. The first difference is determined based on the difference between the preset parameter and a second ratio. The second ratio is determined based on the ratio of the preset parameter to the preset edge frame control parameter.

[0136] Here, the number of image frames is used to characterize the number of image frames included in the target isolated word sign language data.

[0137] Here, the preset edge frame control parameter is used to control the number of frames for the first sign language data and the third sign language data, that is, to control the range of the edge transition image frames, so as to determine the range of the edge transition image frames for which the constraints need to be weakened. The larger the range of the edge transition image frames, the smaller the range of the intermediate segment data, that is, the fewer the number of frames for the second sign language data.

[0138] The preset edge frame control parameter can be set according to actual needs. In some embodiments, the preset edge frame control parameter is determined based on the number of image frames of the target isolated word sign language data. In one embodiment, the smaller the value of the preset edge frame control parameter, the more frames the first and third sign language data have, and the fewer frames the second sign language data has.

[0139] Here, the target frame number is used to characterize which frame of the target image is the first target image frame in the target isolated word sign language data.

[0140] Here, the first ratio can be directly used as the first threshold, or the first threshold can be obtained by further data processing of the first ratio. The product of the number of image frames and the first difference can be directly used as the second threshold, or the second threshold can be obtained by further data processing of the product. The difference between the preset parameter and the second ratio can be directly used as the first difference, or the first difference can be obtained by further data processing of the difference between the preset parameter and the second ratio. The ratio between the preset parameter and the preset edge frame control parameter can be directly used as the second difference, or the second difference can be obtained by further data processing of the ratio. The preset parameter can be set according to actual needs, for example, 1.

[0141] For example, the formula for calculating the first threshold is as follows:

[0142] T1 = g / β;

[0143] In the formula, T1 represents the first threshold, g represents the number of image frames, and β represents the preset edge frame control parameter.

[0144] For example, the formula for calculating the second threshold is as follows:

[0145]

[0146] In the formula, T2 represents the second threshold, g represents the number of image frames, and β represents the preset edge frame control parameter, which is 1.

[0147] In one embodiment, considering that the Gaussian distribution function can well characterize the distribution of isolated word sign language information represented by the target isolated word sign language data, the first weighting weight is as follows:

[0148]

[0149] In the formula, w(t,g) represents the first weighting weight, t represents the number of target frames in the first target image frame, g represents the number of image frames in the target isolated word sign language data, β represents the preset edge frame control parameter, and σ represents the standard deviation of the Gaussian distribution function.

[0150] For example, the third weighting is shown below:

[0151]

[0152] In the formula, w(t,g) represents the third weighting weight, t represents the number of target frames in the first target image frame, g represents the number of image frames in the target isolated word sign language data, β represents the preset edge frame control parameter, and σ represents the standard deviation of the Gaussian distribution function.

[0153] Based on the above, the formula for determining the target weighted weight is as follows:

[0154]

[0155] In the formula, w(t,g) represents the target weighting weight, t represents the target frame number of the first target image frame, g represents the number of image frames of the target isolated word sign language data, β represents the preset edge frame control parameter, and σ represents the standard deviation of the Gaussian distribution function.

[0156] The sign language data augmentation method provided in this invention determines the target weighting weight based on the number of image frames of the target isolated word sign language data, preset edge frame control parameters, and the target frame number of the first target image frame. This controls the number of frames in the first and third sign language data based on the preset edge frame control parameters, thus controlling the range of edge transition image frames. This allows for a more precise determination of the target weighting weight and a more accurate weakening of the constraints of the data difference loss function on the first and third sign language data. Consequently, the training objective of the data augmentation model is more precisely focused on the second sign language data, further improving model training efficiency and the training effect of the data difference loss function on the data augmentation model. This further enhances the effectiveness of continuous sentence sign language data, improving the model training effect and robustness when training a sign language recognition model based on continuous sentence sign language data.

[0157] Based on any of the above embodiments Figure 3 This is the third flowchart illustrating the sign language data augmentation method provided by the present invention, as shown below. Figure 3 As shown, the action continuity score is determined based on the following steps:

[0158] Step 310: Determine multiple sign language data to be scored in the sample continuous sentence sign language data.

[0159] Here, the sign language data to be scored is a subset of the continuous sentence sign language data from the sample. Since the continuity of sign language actions is determined based on sign language data within a short timeframe, multiple sign language data sets to be scored are identified from the continuous sentence sign language data. The number of frames in this set of sign language data should be relatively small to improve the accuracy of the action continuity scoring.

[0160] In one embodiment, multiple sign language data sets to be scored have overlapping data, thereby improving the accuracy of the movement continuity scoring. Of course, multiple sign language data sets to be scored may also have no overlapping data.

[0161] Step 320: Based on the action continuity scoring model, score the multiple sign language data to be scored to obtain multiple scoring results.

[0162] Specifically, based on the action continuity scoring model, any sign language data to be scored is scored to obtain the scoring result of that sign language data.

[0163] The action continuity scoring model is trained based on sample sign language data to be scored, and the sample sign language data to be scored is determined based on positive and negative sample data.

[0164] The method for obtaining the sign language data to be scored in this sample is basically the same as the method for obtaining the sign language data to be scored mentioned above, and will not be repeated here.

[0165] It should be noted that the sample sign language data to be scored is determined based on both positive and negative sample data. That is, the sample sign language data to be scored is derived from both positive and negative sample data to enable unsupervised training of the action continuity scoring model. The loss function of this action continuity scoring model can be set according to actual needs, such as the cross-entropy loss function.

[0166] In one embodiment, the label of the sample sign language data to be scored is 1, determined based on positive sample data, and the label of the sample sign language data to be scored is 0, determined based on negative sample data.

[0167] In another embodiment, the label of the sign language data to be scored determined based on positive sample data is the corresponding sample scoring result, and the label of the sign language data to be scored determined based on negative sample data is the corresponding sample scoring result.

[0168] In another embodiment, the first isolated word sign language data and the second isolated word sign language data are obtained as positive sample data, and the first isolated word sign language data and the second isolated word sign language data are shuffled as negative sample data, thereby performing comparative learning training.

[0169] Step 330: Based on the multiple scoring results, determine the action continuity scoring result.

[0170] Specifically, multiple scoring results can be summed to obtain the action continuity score; alternatively, multiple scoring results can be summed to obtain a summed result, which can then be further processed to obtain the action continuity score.

[0171] The sign language data augmentation method provided in this embodiment of the invention supports the determination of action continuity scoring results. It scores multiple sign language data to be scored in the sample continuous sentence sign language data separately, thereby improving the accuracy of action continuity scoring. At the same time, the action continuity scoring model is trained based on the sample sign language data to be scored, and the sample sign language data to be scored is determined based on positive sample data and negative sample data, thereby enabling unsupervised training and improving model training efficiency.

[0172] Based on any of the above embodiments, the positive sample data is obtained by sampling the first isolated word sign language data; the negative sample data is obtained by replacing at least one frame of the positive sample data with a frame to be replaced, wherein the frame to be replaced is an image frame in the second isolated word sign language data.

[0173] Here, the first and second isolated word sign language data are basically the same as the isolated word sign language data mentioned above, and will not be repeated here. The first and second isolated word sign language data are different isolated word sign language data.

[0174] The first isolated word sign language data can be used directly as positive sample data, or a portion of the sign language data from the first isolated word sign language data can be selected as positive sample data.

[0175] The sign language data augmentation method provided in this invention improves the efficiency of sample data acquisition by replacing at least one frame of the positive sample data with the image frame to be replaced in the sign language data of the second isolated word.

[0176] Based on any of the above embodiments, the difference between any two adjacent frames in the sample continuous sentence sign language data is determined based on the difference of multiple key points, and the two adjacent frames include a first image frame and a second image frame.

[0177] The difference degree of any of the key points is determined based on the difference degree between the first target key point in the first image frame and the second target key point corresponding to the first target key point in the second image frame. The correspondence between the first target key point and the second target key point is that the limb positions are the same. The first target key point is a posture key point related to sign language actions.

[0178] Here, the key point difference degree is used to characterize the degree of difference between the first target key point and the second target key point; it can be determined based on the similarity between the first target key point and the second target key point, that is, the higher the similarity, the smaller the key point difference degree; it can also be determined based on the mean square error of each key point in the first image frame and the second image frame, that is, the mean value of the offset of each key point; of course, it can also be determined by other methods, which will not be elaborated here.

[0179] Here, the postural key points are the key points related to sign language movements, that is, the key points in the limbs related to sign language movements. These postural key points play a crucial role in conveying the meaning of sign language. The location or position of multiple postural key points in the human body can be set according to actual needs.

[0180] In one embodiment, such as Figure 4As shown, multiple postural key points of the human body may include, but are not limited to, at least one of the following: key point sets for both thumbs, key point sets for both index fingers, key point sets for both middle fingers, key point sets for both ring fingers, key point sets for both little fingers, key point sets for the head and shoulders, key point sets for the main body, key point sets for both main bodies, etc. The key point set for both thumbs is illustrated using the key point set for one thumb as an example. This key point set may include, but is not limited to, at least one of the following: first palmar base key point 14, second palmar base key point 15, first thumb joint key point 16, second thumb joint key point 17, third thumb joint key point 18, etc. The key point set for both index fingers is illustrated using the key point set for one index finger as an example. This key point set may include, but is not limited to, at least one of the following: first palmar base key point 14, first index finger joint key point 19, second index finger joint key point 20, third index finger joint key point 21, fourth index finger joint key point 22, etc. The key point set for the middle fingers of both hands is illustrated using the key point set for the middle fingers of one hand as an example. This key point set may include, but is not limited to, at least one of the following: first palmar base key point 14, first middle finger joint key point 23, second middle finger joint key point 24, third middle finger joint key point 25, fourth middle finger joint key point 26, etc. The key point set for the ring fingers of both hands is illustrated using the key point set for the ring fingers of one hand as an example. This key point set may include, but is not limited to, at least one of the following: first palmar base key point 14, first ring finger joint key point 27, second ring finger joint key point 28, third ring finger joint key point 29, fourth ring finger joint key point 30, etc. The key point set for the little fingers of both hands is illustrated using the key point set for the little fingers of one hand as an example. This key point set may include, but is not limited to, at least one of the following: first palmar base key point 14, first little finger joint key point 31, second little finger joint key point 32, third little finger joint key point 33, fourth little finger joint key point 34, etc. The set of key points for the head and shoulders may include, but is not limited to, at least one of the following: key point 1 for the right eye, key point 2 for the left eye, key point 3 for the nose, key point 4 for the right ear, key point 5 for the left ear, key point 6 for the right shoulder, key point 7 for the left shoulder, etc. The set of key points for the main body may include, but is not limited to, at least one of the following: key point 6 for the right shoulder, key point 7 for the left shoulder, key point 12 for the right hip joint, key point 13 for the left hip joint, etc. The set of key points for the main body of both hands may include, but is not limited to, at least one of the following: key point 6 for the right shoulder, key point 8 for the right elbow joint, key point 10 for the right wrist joint, key point 7 for the left shoulder, key point 9 for the left elbow joint, key point 11 for the left wrist joint, etc.

[0181] The sign language data augmentation method provided in this invention determines the difference between any two adjacent frames of continuous sentence sign language data based on the difference of multiple key points. These key points are posture key points related to sign language actions. This removes redundant information from the continuous sentence sign language data and compares the posture information related to sign language actions, thereby improving the accuracy of spatial difference loss determination. The data augmentation model is then trained using the spatial difference loss function to make the series of sign language actions represented by the continuous sentence sign language data generated by the data augmentation model more coherent. In other words, the continuous sentence sign language data is more coherent in both the temporal and spatial dimensions, thus having a stronger correlation with actual sign language actions. This improves the effectiveness of the continuous sentence sign language data and enhances the training effect of the sign language recognition model when training it based on the continuous sentence sign language data, thereby further improving the robustness of the sign language recognition model.

[0182] In practical applications, the above embodiments enable sign language data augmentation, i.e., sign language data expansion, even with limited and small amounts of isolated word sign language data. This efficiently expands continuous sentence sign language data, effectively reducing the difficulty of collecting continuous sentence sign language data, thereby improving the training effect of the sign language recognition model, enhancing its robustness, and ultimately promoting the application of sign language recognition to achieve barrier-free communication.

[0183] The sign language data augmentation device provided by the present invention is described below. The sign language data augmentation device described below can be referred to in correspondence with the sign language data augmentation method described above.

[0184] Figure 5 This is a schematic diagram of the structure of the sign language data augmentation device provided by the present invention, as shown below. Figure 5 As shown, the sign language data augmentation device includes:

[0185] The determining module 510 is used to determine multiple isolated word sign language data, wherein any of the isolated word sign language data includes multiple frames of images;

[0186] Processing module 520 is used to perform augmentation processing on the multiple isolated word sign language data based on the data augmentation model to obtain continuous sentence sign language data;

[0187] The loss function of the data augmentation model is determined based on the action continuity scoring loss function and / or the spatial difference loss function.

[0188] The action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data. The spatial difference loss function is determined based on the difference between each two adjacent frames of the sample continuous sentence sign language data. The sample continuous sentence sign language data is obtained by augmenting multiple sample isolated word sign language data using the data augmentation model.

[0189] The sign language data augmentation device provided in this invention, based on a data augmentation model, can augment multiple isolated word sign language data to obtain continuous sentence sign language data. This allows for the generation of a large amount of continuous sentence sign language data simply by collecting isolated word sign language data, efficiently expanding the continuous sentence sign language data, significantly reducing the time and cost of sign language data collection, improving the efficiency of sign language data acquisition, and achieving efficient sign language data augmentation. This improves the training effect of the sign language recognition model when training based on continuous sentence sign language data, thereby enhancing the robustness of the sign language recognition model. Simultaneously, by training the data augmentation model using an action continuity scoring loss function, the series of sign language actions represented by the continuous sentence sign language data generated by the data augmentation model are made more continuous. In other words, the continuous sentence sign language data is made more continuous in both time and... The data augmentation model is more coherent in the spatial dimension, thus exhibiting a stronger correlation with actual sign language movements. This enhances the effectiveness of continuous sentence sign language data, improving the training performance and robustness of sign language recognition models when trained based on it. Furthermore, training the data augmentation model using a spatial difference loss function further strengthens the coherence of the sign language movements represented by the generated continuous sentence sign language data. In other words, it makes the continuous sentence sign language data more coherent in both the temporal and spatial dimensions, resulting in a stronger correlation with actual sign language movements and further enhancing its effectiveness. This, in turn, improves the training performance and robustness of sign language recognition models when trained based on continuous sentence sign language data.

[0190] Based on any of the above embodiments, the processing module 520 includes:

[0191] The data splicing unit is used to splice the sign language data of the multiple isolated words to obtain spliced ​​sentence sign language data;

[0192] The data augmentation unit is used to augment the spliced ​​sentence sign language data based on the data augmentation model to obtain continuous sentence sign language data;

[0193] The loss function is determined based on the data difference loss function, the action continuity scoring loss function, and / or the spatial difference loss function.

[0194] The data difference loss function is determined based on the difference between the sample continuous sentence sign language data and the sample spliced ​​sentence sign language data, which is obtained by splicing together the multiple sample isolated word sign language data.

[0195] Based on any of the above embodiments, the data difference loss function is determined based on multiple word difference loss values;

[0196] Any of the aforementioned word difference loss values ​​is determined based on the word difference between the target isolated word sign language data in the sample spliced ​​sentence sign language data and the target sign language data corresponding to the target isolated word sign language data in the sample continuous sentence sign language data, wherein the target isolated word sign language data and the target sign language data correspond in the same frame number.

[0197] The word difference is determined based on multiple image differences. Each image difference is determined based on the difference between a first target image frame in the target isolated word sign language data and a second target image frame corresponding to the first target image frame in the target sign language data. The first target image frame and the second target image frame correspond to each other in that they have the same number of frames.

[0198] Based on any of the above embodiments, any of the image differences is determined based on the difference between the first target image frame in the target isolated word sign language data and the second target image frame corresponding to the first target image frame in the target sign language data, and the target weighting weight corresponding to the first target image frame;

[0199] The target isolated word sign language data includes first sign language data, second sign language data, and third sign language data, wherein the number of frames in the first sign language data is less than the number of frames in the second sign language data, and the number of frames in the second sign language data is less than the number of frames in the third sign language data;

[0200] The second weighted weight corresponding to the image frame in the second sign language data is greater than the first weighted weight corresponding to the image frame in the first sign language data, and the second weighted weight corresponding to the image frame in the second sign language data is greater than the third weighted weight corresponding to the image frame in the third sign language data.

[0201] Based on any of the above embodiments, the target weighting is determined based on the number of image frames of the target isolated word sign language data, the preset edge frame control parameters, and the target number of the first target image frame;

[0202] The preset edge frame control parameter is used to control the number of frames in the first sign language data and the third sign language data, and the target frame number is used to characterize the frame time of the first target image frame in the target isolated word sign language data;

[0203] If the target number of frames is less than the first threshold, then the target weight is the first weight.

[0204] If the target number of frames is greater than or equal to the first threshold and the target number of frames is less than or equal to the second threshold, then the target weighting weight is the second weighting weight.

[0205] If the target number of frames is greater than the second threshold, then the target weight is the third weight.

[0206] The first threshold is determined based on a first ratio of the number of image frames to the preset edge frame control parameter. The second threshold is determined based on the product of the number of image frames and a first difference. The first difference is determined based on the difference between the preset parameter and a second ratio. The second ratio is determined based on the ratio of the preset parameter to the preset edge frame control parameter.

[0207] Based on any of the above embodiments, the device further includes a scoring determination module, which is used for:

[0208] Identify multiple sign language data points to be scored within the sample continuous sentence sign language data;

[0209] Based on the action continuity scoring model, the multiple sign language data to be scored are scored respectively, resulting in multiple scoring results;

[0210] Based on the multiple scoring results, the action continuity scoring result is determined;

[0211] The action continuity scoring model is trained based on sample sign language data to be scored, and the sample sign language data to be scored is determined based on positive and negative sample data.

[0212] Based on any of the above embodiments, the positive sample data is obtained by sampling the sign language data of the first isolated word;

[0213] The negative sample data is obtained by replacing at least one image frame in the positive sample data with an image frame to be replaced, wherein the image frame to be replaced is an image frame in the second isolated word sign language data.

[0214] Based on any of the above embodiments, the difference between any two adjacent frames in the sample continuous sentence sign language data is determined based on the difference of multiple key points, and the two adjacent frames include a first image frame and a second image frame.

[0215] The difference degree of any of the key points is determined based on the difference degree between the first target key point in the first image frame and the second target key point corresponding to the first target key point in the second image frame. The correspondence between the first target key point and the second target key point is that the limb positions are the same. The first target key point is a posture key point related to sign language actions.

[0216] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute a sign language data augmentation method, the method comprising: determining multiple isolated word sign language data, each of the isolated word sign language data comprising multiple frames of images; augmenting the multiple isolated word sign language data based on a data augmentation model to obtain continuous sentence sign language data; wherein, the loss function of the data augmentation model is determined based on an action continuity scoring loss function and / or a spatial difference loss function; the action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data, the spatial difference loss function is determined based on the difference between each two adjacent frames of the sample continuous sentence sign language data, and the sample continuous sentence sign language data is obtained by augmenting the multiple sample isolated word sign language data based on the data augmentation model.

[0217] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0218] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the sign language data augmentation method provided by the above methods. The method includes: determining a plurality of isolated word sign language data, each of the isolated word sign language data including multiple frames of images; performing augmentation processing on the plurality of isolated word sign language data based on a data augmentation model to obtain continuous sentence sign language data; wherein, the loss function of the data augmentation model is determined based on an action continuity scoring loss function and / or a spatial difference loss function; the action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data, the spatial difference loss function is determined based on the difference between each two adjacent frames of images in the sample continuous sentence sign language data, and the sample continuous sentence sign language data is obtained by augmenting the plurality of sample isolated word sign language data based on the data augmentation model.

[0219] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the sign language data augmentation method provided by the above methods. The method includes: determining a plurality of isolated word sign language data, each of the isolated word sign language data comprising multiple frames of images; augmenting the plurality of isolated word sign language data based on a data augmentation model to obtain continuous sentence sign language data; wherein the loss function of the data augmentation model is determined based on an action continuity scoring loss function and / or a spatial difference loss function; the action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data, the spatial difference loss function is determined based on the difference between adjacent frames of images in the sample continuous sentence sign language data, and the sample continuous sentence sign language data is obtained by augmenting the plurality of sample isolated word sign language data based on the data augmentation model.

[0220] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0221] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for augmenting sign language data, characterized in that, include: Determine multiple isolated word sign language data, each of which includes multiple frames of images; Based on the data augmentation model, the sign language data of the multiple isolated words are augmented to obtain sign language data of continuous sentences. The loss function of the data augmentation model is determined based on the action continuity scoring loss function and / or the spatial difference loss function. The action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data. The spatial difference loss function is determined based on the difference between each two adjacent frames of the sample continuous sentence sign language data. The sample continuous sentence sign language data is obtained by augmenting multiple sample isolated word sign language data using the data augmentation model.

2. The sign language data augmentation method according to claim 1, characterized in that, The data augmentation model is used to augment the multiple isolated word sign language data to obtain continuous sentence sign language data, including: The sign language data of the multiple isolated words are concatenated to obtain the sign language data of the concatenated sentence; Based on the data augmentation model, the spliced ​​sentence sign language data is augmented to obtain continuous sentence sign language data; The loss function is determined based on the data difference loss function, the action continuity scoring loss function, and / or the spatial difference loss function. The data difference loss function is determined based on the difference between the sample continuous sentence sign language data and the sample spliced ​​sentence sign language data, which is obtained by splicing together the multiple sample isolated word sign language data.

3. The sign language data augmentation method according to claim 2, characterized in that, The data difference loss function is determined based on multiple word difference loss values; Any of the aforementioned word difference loss values ​​is determined based on the word difference between the target isolated word sign language data in the sample spliced ​​sentence sign language data and the target sign language data corresponding to the target isolated word sign language data in the sample continuous sentence sign language data, wherein the target isolated word sign language data and the target sign language data correspond in the same frame number. The word difference is determined based on multiple image differences. Each image difference is determined based on the difference between a first target image frame in the target isolated word sign language data and a second target image frame corresponding to the first target image frame in the target sign language data. The first target image frame and the second target image frame correspond to each other in that they have the same number of frames.

4. The sign language data augmentation method according to claim 3, characterized in that, The image difference degree is determined based on the difference degree between the first target image frame in the target isolated word sign language data and the second target image frame corresponding to the first target image frame in the target sign language data, and the target weighting weight corresponding to the first target image frame; The target isolated word sign language data includes first sign language data, second sign language data, and third sign language data, wherein the number of frames in the first sign language data is less than the number of frames in the second sign language data, and the number of frames in the second sign language data is less than the number of frames in the third sign language data; The second weighted weight corresponding to the image frame in the second sign language data is greater than the first weighted weight corresponding to the image frame in the first sign language data, and the second weighted weight corresponding to the image frame in the second sign language data is greater than the third weighted weight corresponding to the image frame in the third sign language data.

5. The sign language data augmentation method according to claim 4, characterized in that, The target weighting is determined based on the number of image frames of the target isolated word sign language data, preset edge frame control parameters, and the target frame number of the first target image frame; The preset edge frame control parameter is used to control the number of frames in the first sign language data and the third sign language data, and the target frame number is used to characterize the frame time of the first target image frame in the target isolated word sign language data; If the target number of frames is less than the first threshold, then the target weight is the first weight. If the target number of frames is greater than or equal to the first threshold and the target number of frames is less than or equal to the second threshold, then the target weighting weight is the second weighting weight. If the target number of frames is greater than the second threshold, then the target weight is the third weight. The first threshold is determined based on a first ratio of the number of image frames to the preset edge frame control parameter. The second threshold is determined based on the product of the number of image frames and a first difference. The first difference is determined based on the difference between the preset parameter and a second ratio. The second ratio is determined based on the ratio of the preset parameter to the preset edge frame control parameter.

6. The sign language data augmentation method according to claim 1, characterized in that, The action continuity score is determined based on the following steps: Identify multiple sign language data points to be scored within the sample continuous sentence sign language data; Based on the action continuity scoring model, the multiple sign language data to be scored are scored respectively, resulting in multiple scoring results; Based on the multiple scoring results, the action continuity scoring result is determined; The action continuity scoring model is trained based on sample sign language data to be scored, and the sample sign language data to be scored is determined based on positive and negative sample data.

7. The sign language data augmentation method according to claim 6, characterized in that, The positive sample data is obtained by sampling the sign language data of the first isolated word; The negative sample data is obtained by replacing at least one image frame in the positive sample data with an image frame to be replaced, wherein the image frame to be replaced is an image frame in the second isolated word sign language data.

8. The sign language data augmentation method according to claim 1, characterized in that, The difference between any two adjacent frames in the sample continuous sign language data is determined based on the difference of multiple key points. The two adjacent frames include a first image frame and a second image frame. The difference degree of any of the key points is determined based on the difference degree between the first target key point in the first image frame and the second target key point corresponding to the first target key point in the second image frame. The correspondence between the first target key point and the second target key point is that the limb positions are the same. The first target key point is a posture key point related to sign language actions.

9. A sign language data augmentation device, characterized in that, include: A determination module is used to determine multiple isolated word sign language data, wherein any isolated word sign language data includes multiple frames of images; The processing module is used to perform augmentation processing on the multiple isolated word sign language data based on the data augmentation model to obtain continuous sentence sign language data; The loss function of the data augmentation model is determined based on the action continuity scoring loss function and / or the spatial difference loss function. The action continuity scoring loss function is determined based on the action continuity scoring results of the sample continuous sentence sign language data. The spatial difference loss function is determined based on the difference between each two adjacent frames of the sample continuous sentence sign language data. The sample continuous sentence sign language data is obtained by augmenting multiple sample isolated word sign language data using the data augmentation model.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the sign language data augmentation method as described in any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the sign language data augmentation method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Continuous sign language recognition method based on cross-modal data augmentation

    CN112149603A

  • Gesture language teaching method, device and system based on gesture action generation and recognition

    CN114842547A