Method, apparatus, storage medium and device for training clue language automatic recognition model
Through federated learning and encryption sharing training mechanisms, combined with the collaborative training of servers and clients, the problem of data distribution imbalance and privacy protection in automatic clue language recognition model training is solved, and efficient and privacy-protected model training effect is achieved.
Patent Information
- Application Number
- CN202111601585.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-24
AI Technical Summary
The prior art is difficult to effectively identify and decode Cued Speech, especially in terms of unbalanced data distribution, noisy annotation and privacy protection.
A federated learning-based method is adopted to automatically recognize the clue language by jointly training the server and the client. The method includes obtaining an open source clue video dataset, issuing server model parameters to the client, and the client performs model training and parameter aggregation until the server model converges. At the same time, encrypted sharing training mechanism and secure hybrid algorithm are used to protect user privacy.
It realizes efficient automatic clue language recognition model training, solves the problems of data distribution imbalance and noise labeling, and ensures the protection of user privacy and promotes communication convenience for deaf and dumb people.
Smart Images

Figure CN114445909B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of model training, and particularly relates to a method, device, storage medium and equipment for training an automatic recognition model of cue speech. Background Art
[0002] With the development of current society and the improvement of living quality, the communication problem among disabled people has attracted more and more attention from society. According to the report of the World Health Organization (WHO), there are now about 466 million people with hearing impairments globally, of which 34 million are children.
[0003] Lip reading is an earlier way to help deaf-mutes perceive speech. However, according to research, since different speeches in lip reading may have similar lip shapes, such as [u] and [y], this confusion results in poor speech recognition relying solely on lip shapes. To address this deficiency, Professor R. Orin Cornett of Gallaudet University in the United States invented a communication method in 1967 that uses gestures to assist lip reading, called Cued Speech (CS), which is translated into Chinese as cue speech. In this system, the position of the hand is used to encode vowels, and the shape of the hand is used to encode consonants. Specifically, in English CS, 4 hand positions are used to encode monophthongs, 2 hand movements encode diphthongs, and 8 hand shapes encode consonants.
[0004] Given the wide distribution of the hearing-impaired population globally and the increasingly widespread application of CS, in order to make the communication of deaf-mutes more convenient and efficient, the urgent development of an automatic recognition model for CS videos to text based on artificial intelligence has received more and more attention from countries, society, industry and academia. Using artificial intelligence technology for intelligent analysis and modeling of CS image and video content has become a major key technology in the development and application of the cause of disabled people in China, and will play an important role in promoting the auxiliary wisdom system for disabled people and ensuring the innovative development of intelligent health products. Summary of the Invention
[0005] Embodiments of the present invention provide a method, device, storage medium and equipment for training an automatic recognition model of cue speech, aiming to solve at least one of the problems in the background art.
[0006] Embodiments of the present invention are implemented as follows. A method for training an automatic recognition model of cue speech is applied to a server, and the server is communicatively connected to multiple clients. The method includes:
[0007] S01, obtaining an open-source cue speech video dataset, and training the server model using the open-source cue speech video dataset;
[0008] S02. Randomly select multiple target clients, and send the current global model parameters of the server model to each of the target clients, so that the target clients train their respective client models based on the current global model parameters and a preset training set. The architectures of the client model and the server model are the same;
[0009] S03. Receive the current client model parameters uploaded by each of the target clients, aggregate the current client model parameters, and use the aggregated model parameters to update the server model to obtain a new server model;
[0010] Repeat steps S02 - S03 until the server model converges to train a clue language automatic recognition model.
[0011] Preferably, before the step of randomly selecting multiple target clients and sending the current global model parameters of the server model to each of the target clients, it further includes:
[0012] Obtain clue language videos encrypted by a preset encryption algorithm uploaded by multiple clients to obtain a shared data set composed of the obtained encrypted clue language videos.
[0013] Preferably, the step of randomly selecting multiple target clients and sending the current global model parameters of the server model to each of the target clients includes:
[0014] Randomly select multiple target clients, send the current global model parameters of the server model to each of the target clients, and select a part of the shared data from the shared data set according to a preset ratio and send it to each of the target clients;
[0015] Among them, the preset training set includes the local data set of the target client and the part of the shared data currently sent by the server.
[0016] Preferably, before uploading the current client model parameters, the target client trims the current client model and introduces Laplace noise.
[0017] Preferably, the architectures of the client model and the server model both include:
[0018] A self-supervised incremental learning joint model for extracting hand shape features and lip features from the data set;
[0019] A hand position recognition model for extracting hand position features from the data set;
[0020] An asynchronous heterogeneous multi-modal temporal alignment model is used to statistically model and predict the lead of the movements of hand shapes and hand positions relative to lip movements based on hand shape features, lip features, and hand position features, and use the predicted lead to adjust the hand shape features and hand position features in time series to obtain preliminarily temporally aligned hand shape features, lip features, and hand position features;
[0021] A fusion model is used to predict and identify the phoneme categories of the cue language based on the preliminarily temporally aligned hand shape features, lip features, and hand position features.
[0022] Preferably, the fusion model is a multi-modal fusion model based on generative adversarial training composed of an identifier, a generator, and a discriminator.
[0023] An embodiment of the present invention also provides a training device for a cue language automatic recognition model, which is applied to a server. The server is communicatively connected to multiple clients. The device includes:
[0024] A pre-training module is used to obtain an open-source cue language video data set and train the server model using the open-source cue language video data set;
[0025] A data distribution module is used to randomly select multiple target clients and send the current global model parameters of the server model to each of the target clients, so that the target clients train their respective client models based on the current global model parameters and a preset training set. The architectures of the client models and the server model are the same;
[0026] A model update module is used to receive the current client model parameters uploaded by each of the target clients, aggregate the current client model parameters, and update the server model using the aggregated model parameters to obtain a new server model;
[0027] The data distribution module and the model update module are repeatedly executed until the server model converges to train a cue language automatic recognition model.
[0028] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the cue language automatic recognition model training method as described above.
[0029] An embodiment of the present invention also provides a server, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the cue language automatic recognition model training method as described above.
[0030] The beneficial effects achieved by the present invention are as follows: By proposing a method for training a clue language automatic recognition model, it is expected to promote the research results of CS automatic recognition to truly move towards application and benefit deaf-mute people and more people in need. At the same time, this method jointly completes the training of the clue language automatic recognition model by the server and the client, so that the private clue language videos of the client (such as self-recorded CS videos) do not need to be uploaded to the server, which well protects the privacy of users. In this way, more clients will be willing to take out their private clue language videos for model training, providing a sufficient data basis for model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a flowchart of the clue language automatic recognition model training method in the first embodiment of the present invention;
[0032] Figure 2 is a technical circuit diagram of the clue language automatic recognition model training method provided by the embodiment of the present invention;
[0033] Figure 3 is a structural diagram of the self-supervised incremental learning joint model provided by the embodiment of the present invention;
[0034] Figure 4 are several important statistics in the case where the variable is the target time t provided by the embodiment of the present invention;
[0035] Figure 5 is a structural diagram of the multi-modal fusion model based on generative adversarial training provided by the embodiment of the present invention;
[0036] Figure 6 is a structural diagram of the multi-modal feature fusion module provided by the embodiment of the present invention;
[0037] Figure 7 is the overall model framework of CS audio-visual recognition based on federated learning provided by the embodiment of the present invention;
[0038] Figure 8 is a benchmark model architecture diagram of the server model and the client model provided by the embodiment of the present invention;
[0039] Figure 9 is a legend of the secure hybrid algorithm when m = 3 provided by the embodiment of the present invention;
[0040] Figure 10 is a structural block diagram of the clue language automatic recognition model training device in the third embodiment of the present invention;
[0041] Figure 11 is a structural block diagram of the server in the fourth embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0043] Although the research on CS automatic recognition has achieved relatively good development in recent years, there is still a long way to go to an efficient CS automatic recognition model with privacy protection. To sum up, there are still the following three deficiencies and challenges in the theoretical research in this field. The present invention also aims to solve these three key scientific problems.
[0044] (1) Limited CS multimodal feature representation ability under noisy annotation and data distribution imbalance: Previous automatic CS handshape feature extraction methods either have high requirements for external environmental conditions or rely on a large amount of accurately annotated data. However, the annotation of a large amount of CS video data requires a huge cost, and at the same time requires the annotator to master the relevant professional background knowledge of CS. Using automatic annotation methods will generate noisy labels. In addition, due to the unbalanced nature of data distribution in the real world, the CS data collected often shows the phenomenon of unbalanced handshape category distribution. It should be noted that as the time series of the video data stream increases, the imbalance phenomenon of the video data will gradually worsen. It is difficult for deep learning models to learn fair and high-quality features when training on unbalanced video data sets, and it is easy to ignore the categories with fewer samples. Therefore, it is necessary to conduct in-depth research on the collaborative mechanism of unsupervised learning independent of annotation and unbiased learning of data distribution. Using unsupervised deep neural networks for stronger feature representation learning is the first important challenge faced by CS automatic recognition research.
[0045] (2) Poor CS asynchronous multimodal alignment and fusion effect: The asynchronous nature of the time series between CS video multimodals will cause interference between the time series information of different modalities in information fusion, and it is difficult for the information of multiple heterogeneous modalities to complement and enhance each other. Affected by the small amount of CS data (i.e., small samples), noisy annotation and unbalanced category distribution, previous CS fusion research and existing explicit and implicit alignment fusion methods cannot effectively reduce the interference caused by information misalignment between asynchronous multimodals. Therefore, it is necessary to conduct in-depth research on the optimization mechanism of asynchronous multimodal information alignment and information fusion methods. Making full use of the complementary and enhancing properties of multimodal data to fully fuse asynchronous heterogeneous multimodal CS video data is the second important challenge faced by CS automatic recognition research.
[0046] (3) Lack of a CS automatic recognition model with data privacy protection: Previous CS automatic recognition methods mainly focused on how to improve model performance and lacked in-depth research on the privacy protection of CS training data. The privacy protection of deaf people's data is an urgent problem to be solved for the further development, practical implementation, and application promotion of the CS automatic recognition model. Directly uploading the data in the client to the server for shared training easily exposes users' privacy information. At the same time, the non-independent and identically distributed video data from different CS coders also poses certain difficulties for video privacy information protection. We point out that if the privacy of data providers during model training can be protected, it will greatly encourage data recorders or institutions to provide us with more data. This will largely help us solve the above-mentioned difficulties in data acquisition. Therefore, it is necessary to conduct in-depth research on the training mechanism of the CS automatic recognition model with privacy protection. Exploring the encrypted shared training mechanism of federated learning to protect data privacy in CS automatic recognition while ensuring high accuracy of CS automatic recognition is the third important challenge of this invention.
[0047] In response to the above three deficiencies and challenges in CS automatic recognition research, the embodiments of this invention propose a method for training a sign language automatic recognition model, aiming to establish a CS automatic recognition model with privacy protection based on the theoretical framework of federated learning, so as to promote the research results of CS automatic recognition to truly move towards application and benefit deaf people and more people in need. This has important social and economic significance. At the same time, the three scientific problems that this invention intends to solve also receive much attention in other fields. For example, early education of hearing-impaired people, speech correction and treatment, privacy protection in audio-visual recognition, robots, audio-visual conversion, and human-computer interaction, etc.
[0048] Example 1
[0049] Please refer to Figure 1 , which shows the method for training a sign language automatic recognition model in Embodiment 1 of this invention, applied to a server. The server is communicatively connected to multiple clients. The method specifically includes steps S01 - S04.
[0050] Step S01, obtain an open-source sign language video dataset, and use the open-source sign language video dataset to train the server model.
[0051] Among them, the open-source sign language video dataset can be publicly available sign language videos on the network, that is, use the collected public sign language videos to pre-train the server model first, so that the model has initial global model parameters.
[0052] Step S02: Randomly select multiple target clients and send the current global model parameters of the server model to each of the target clients, so that the target clients train their respective client models based on the current global model parameters and a preset training set. The architectures of the client model and the server model are the same.
[0053] Among them, before uploading the current client model parameters, the target client trims the current client model and introduces Laplace noise, that is, trims and segments the client model parameters for uploading and adds Laplace noise, so as to prevent malicious servers or external attackers from stealing data information through the model parameters uploaded by the participants.
[0054] Step S03: Receive the current client model parameters uploaded by each target client, aggregate the current client model parameters, and update the server model with the aggregated model parameters to obtain a new server model.
[0055] Step S04: Repeat steps S02 - S03 until the server model converges to train a clue language automatic recognition model.
[0056] Specifically, in some preferred embodiments of this embodiment, as Figure 5 and Figure 8 shown, the benchmark architectures of the client model and the server model both include:
[0057] A self-supervised incremental learning joint model for extracting hand shape features and lip features from a data set;
[0058] A hand position recognition model for extracting hand position features from a data set. The hand position recognition model is specifically implemented by the YOLO algorithm;
[0059] An asynchronous heterogeneous multi-modal temporal alignment model for statistically modeling and predicting the lead of the movement of the hand shape and hand position relative to the lip movement based on the hand shape features, lip features, and hand position features, and using the predicted lead to adjust the hand shape features and hand position features in time series to obtain preliminarily temporally aligned hand shape features, lip features, and hand position features;
[0060] A fusion model for predicting and recognizing the phoneme categories of the clue language based on the preliminarily temporally aligned hand shape features, lip features, and hand position features. The fusion model is specifically a multi-modal fusion model based on generative adversarial training composed of an identifier, a generator, and a discriminator, as Figure 5 shown.
[0061] Example 2
[0062] Embodiment 2 of the present invention also proposes a method for training a clue language automatic recognition model. The difference between the clue language automatic recognition model training method in this embodiment and the clue language automatic recognition model training method in Embodiment 1 lies in:
[0063] Before the step of randomly selecting multiple target clients and sending the current global model parameters of the server model to each of the target clients, it further includes:
[0064] Obtain clue language videos encrypted by a preset encryption algorithm uploaded by multiple clients to obtain a shared dataset composed of the obtained encrypted clue language videos.
[0065] Based on this, the step of randomly selecting multiple target clients and sending the current global model parameters of the server model to each of the target clients includes:
[0066] Randomly select multiple target clients, send the current global model parameters of the server model to each of the target clients, and select a part of the shared data from the shared dataset according to a preset ratio and send it to each of the target clients;
[0067] Wherein, the preset training set includes the local dataset of the target client and the part of the shared data currently sent by the server.
[0068] Please refer to Figure 2 , the technical route of the clue language automatic recognition model training method in the embodiment of the present invention is as follows: First, based on the collaborative modeling of self-supervised learning mechanism and unbiased learning, efficiently extract hand shape features and mouth shape features in CS videos; then, explore the optimal fusion method for these asynchronous and heterogeneous multi-modal feature streams; finally, under the theoretical framework of federated learning, construct an efficient CS automatic recognition model with privacy protection. Specifically, it is elaborated in the following three parts.
[0069] I. Aiming at the first deficiency and challenge in the above CS automatic recognition research, this embodiment proposes the collaborative modeling of self-supervised contrast learning and unbiased learning of CS multi-modal features. This research is divided into three stages. The first stage is the theoretical analysis and modeling of self-supervised contrast learning, the second stage is to construct a self-supervised incremental learning joint model (as Figure 3 shown), and the third stage is test verification and evaluation.
[0070] (1) Method for efficiently generating positive and negative sample pairs in self-supervised contrast learning: The invention intends to adopt the self-supervised contrast learning method to reduce the influence of noisy data annotation in the process of CS handshape feature extraction. Among them, the most crucial step in contrast learning is to generate appropriate positive and negative sample pairs in the data transformation module for contrast learning. Most previous studies were based on heuristic selection methods, without quantitative indicators specifying the specific criteria for selecting positive and negative sample pairs. The present invention will explore the spatio-temporal constraint relationship of video data and propose a spatio-temporal consistency transformation method controlled by mutual information, in order to find the optimal data transformation method for positive and negative sample pairs. The specific technical route and key technologies are described as follows:
[0071] Spatio-temporal consistency transformation controlled by mutual information: Most existing contrast learning algorithms adopt manually designed (i.e., selecting adjacent frames as positive samples and frames far apart in time as negative samples), as well as data transformation methods of automatic search to generate positive and negative sample pairs. In order to obtain a good set of data transformation methods, a large number of experiments are often required, and these data transformation methods usually have randomness and cannot controllably generate the required positive and negative sample pairs. Therefore, the present invention proposes a method of spatio-temporal similarity consistency transformation based on mutual information control for video data streams. This method can be applied to various traditional data transformation methods and can generate the best sample pairs by controlling mutual information. First, in the scenario of self-supervised contrast learning, starting from the time dimension of the video, the video is semantically clustered, and two frames with the same semantics in time series are extracted for feature transformation as positive sample pairs, so as to overcome the problem that traditional algorithms perform data transformation on each frame in the video to generate positive sample pairs without considering whether the semantics (i.e., labels) of the video frames are the same. Secondly, from the perspective of the spatial features of images, samples with appropriate mutual information are obtained by controlling the intensity of different data transformations.
[0072] In the InfoMin theory, there is the following Proposition 1, where I≥0 is the mutual information.
[0073] Proposition 1: Given the picture sample x with the label y, the best positive sample pair u * and v * for x are:
[0074] (u * , v * ) = argmin u,u I(u, v),
[0075] s.t. I(u, y) = I(v, y) = I(x, y).
[0076] According to Proposition 1, the method given by InfoMin is for static images and only considers the relationship of the spatial features of the images, which is not applicable to the characteristics of our own task in the present invention. Now, considering the spatio-temporal consistency constraint simultaneously, we obtain the following proposition:
[0077] Proposition 2: Assume that f i ∈A is an arbitrary data transformation, and A is a set containing various data transformations. τ i , τ j (τ i ≠τ j ) are two different moments for sampling images from the video, V represents the video data, respectively represent the images from two semantic clusters sampled at τ i , τ j moments, and these two images have the same latent label y i and y j . Then the generation of good positive and negative sample pairs by V is equivalent to solving the following optimization problem:
[0078]
[0079]
[0080]
[0081] andy 1 = y 2 , y 3 = y 4 , y 1 ≠y 3 .
[0082] Therefore, the good positive sample pairs generated by the video V are and The negative sample pairs are combinations of any two samples with different semantic labels (from clusters with different semantic labels), such as Different from the traditional method, here we no longer perform two different data transformations on the same image, but perform two data transformations on two different images with the same semantics. This can avoid the problem that the feature representations of positive samples in the feature space are too similar and increase the diversity of positive samples. It should be noted that the above optimization process can consider multiple moments at once, and multiple good positive and negative sample pairs can be obtained through one optimization.
[0083] The above optimization problem is supported by the InfoMin theory. Considering the characteristics of video data, temporal semantic information and information collaborative constraints of spatial features are added. Video frames are semantically partitioned (clustered) in the time domain, and samples with the same latent label are used as positive samples to achieve semantic-balanced contrastive learning, avoiding ineffective learning of multiple samples for a label. At the same time, this method reduces the redundant information between positive sample pairs by minimizing the mutual information of samples, aiming to further improve the effectiveness of model training. Thus, we transform the problem of generating good positive and negative sample pairs in self-supervised contrastive learning into the optimization problem in Proposition 2 above, and generate positive and negative sample pairs by using mutual information to control the spatio-temporal information volume between sample pairs.
[0084] (2) Construct a self-supervised incremental learning joint model: In the initial stage, first, based on the CS video data stream, data processing is performed through resampling technology to construct an unbiased validation sample set, and then, combined with the above self-supervised contrastive learning based on the spatio-temporal similarity consistency transformation controlled by mutual information, high-quality sample features are extracted. In the middle stage, for the video data stream at different times, two gradient constraints are introduced for incremental training of the model: 1) Constraint is performed through the inner product of gradients between parameters at different times to prevent catastrophic forgetting of the model; 2) Constraint is performed through the inner product of gradients between parameters on the training set and the validation set at the same time to ensure unbiased learning of the model. In the later stage, as the video data stream is continuously updated, the feature extraction model continuously learns in the above incremental manner and can be extended to different data or different tasks until the video data stream stops or no new tasks appear. Based on this incremental learning framework, we expect to obtain a feature representation learning model that is more adaptable to the real CS multimodal data stream. The specific technical route and key technologies can be elaborated as follows:
[0085] For a biased data distribution affected by unbalanced data distribution or noisy labels, first, the incremental update of the model by the video data stream should be constrained. For the data streams at different times, at time t, the sample is represented as (x t , y t ), and the loss function is represented as l(f(x t , θ t ), y t ). Then, the gradient with respect to the model parameter θ t can be expressed as:
[0086]
[0087] Similarly, we can obtain the gradient representation of the model parameter at time k (k < t). To prevent the problem of catastrophic forgetting of the model caused by newly added data (that is, once an existing model is trained with a new data set, the model will lose the ability to recognize the original data set), we let the model parameter at time t be the solution of the following optimization problem:
[0088] minl(f(x t ,θ t ),y t ),
[0089] s.t.l(f(x t ,θ t ),y t )≤l(f(x k ,θ k ),y k ), for all k < t.
[0090] To avoid saving the model parameter states at all times and simplify the above constraints, it can be rewritten in the following gradient form:
[0091]
[0092] To prevent a biased data distribution from disrupting the normal learning of the model, we first need to obtain a clean subset (x clean , y clean ), abbreviated as (x c , y c ), that is, a set of samples with a balanced data distribution and accurate labels. According to the working mechanism of the validation set, the learning and update of the model on the training set should not degrade the performance of the model on the validation set, that is, the learning directions of the model on the training set and the validation set do not conflict, so as to ensure the unbiasedness of the model's learning on the training set. Therefore, the gradients of the validation subset with respect to the model should satisfy the following constraints:
[0093]
[0094] where, is the parameter gradient of the model on the validation subset:
[0095]
[0096] Finally, the model parameter gradient at time t can be obtained by solving the following quadratic optimization problem
[0097] Through the constraints of the above two conditions, during the incremental update process of the model and the video data stream, the learning of the model can constrain the update of the model to proceed in an unbiased direction, and at the same time prevent the problem of catastrophic forgetting caused by new data. When implementing this method, possible uncertainty problems may include complex sample distribution overlap or difficult samples (such as noise samples), etc. According to the actual situation, we will also explore algorithms such as sample distribution decoupling and difficult sample mining (such as gradient harmony mechanism) to overcome these possible challenges.
[0098] (3) Model Validation and Evaluation - Feature Classification Task: In the model validation and evaluation phase, we plan to divide all CS data into a training set (80%) and a test set (20%). After completing self-supervised contrastive incremental learning using unlabeled data, we plan to fine-tune the downstream hand shape feature classification model with 10% manually labeled data. On the one hand, we can obtain feature representation model parameters that are more suitable for CS data. On the other hand, we can fairly validate and evaluate the model through the prediction results given by the fine-tuned model. We plan to use a two-layer non-linear fully connected layer as the neural network model for this module and use the cross-entropy loss function to fine-tune the model parameters of both the self-supervised contrastive incremental learning feature representation learning module and the feature classification task module simultaneously.
[0099] II. To address the deficiencies and challenges in point (2) of the above CS automatic recognition research, this embodiment proposes the design and analysis of an asynchronous heterogeneous multimodal temporal alignment and generative adversarial fusion model.
[0100] Two key points of CS multimodal fusion are: 1) Analyze the alignment relationship between asynchronous multimodals; 2) Fully fuse asynchronous heterogeneous multimodals. Therefore, this research is divided into the following two stages. The specific elaboration is as follows:
[0101] (1) Multimodal Asynchronous Quantification Analysis Based on Statistical Analysis: Previous research has shown that during CS encoding, hand movements usually precede lip movements. This asynchronous problem is a difficulty in CS multimodal fusion. Previous methods could not well solve the CS asynchronous heterogeneous multimodal fusion problem due to issues such as insufficient CS data volume, noisy annotations, and unbalanced data distribution. To avoid the above problems and explore the law of the asynchronous amount between CS multimodals in a statistically interpretable manner, the present invention plans to combine the applicant's prior knowledge of CS encoding to conduct a fine-grained quantitative statistical analysis of the asynchronous lead amount between CS multimodals.
[0102] Please refer to Figure 4 , which shows several important statistics when the variable is the target time t. Among them, t v represents the target time when the sound signal reaches the vowel v, t tar_v represents the target time when the hand position reaches the vowel, and Δ v = t v - t tar_v represents the asynchronous amount by which the hand position precedes the lips when realizing the vowel v. Similarly, for the consonant vowel c, Δ c = t c - t tar_c is the asynchronous amount by which the hand shape precedes the lips.
[0103] We first manually annotate the hand position, hand shape, and lip position at the target moment of realizing the phoneme for 10% of the sentences in the dataset, and calculate their corresponding lead amounts (see Figure 4 ), and then, based on the applicant's prior knowledge of CS coding before, starting from three perspectives, using statistical variance analysis, hypothesis testing, and stepwise regression modeling, we perform backward stepwise regression modeling on the obtained asynchrony amounts corresponding to the phonemes and the normalized moments (t), phoneme categories (y), and context-related triphone categories (y tri ) that the phonemes correspond to in the sentences in the following steps:
[0104] 1) Establish a regression equation for t, y, y tri and the asynchrony amount Δ. Conduct an F-test on the 3 independent variables in the equation and take the minimum value If where n is the number of samples, then eliminate Without loss of generality, we can let and proceed to the following step 2); otherwise, there are no independent variables to eliminate, and at this time the regression equation is the optimal one.
[0105] 2) Establish a regression equation for t, y and the asynchrony amount Δ. Conduct an F-test on the regression coefficients in the equation and take the minimum value If then eliminate Without loss of generality, we can let Otherwise, there are no independent variables to eliminate, and at this time the regression equation is the optimal one.
[0106] 3) Keep iterating until the F-values of the regression coefficients of all variables are greater than the critical value, that is, until there are no variables in the equation that can be eliminated. At this time, the regression equation is the optimal one.
[0107] Based on the optimal regression equation obtained above, we can perform statistical modeling and prediction on the lead amounts of the movements of the hand shape and hand position relative to the lip movement, and then use the predicted lead amounts to adjust the hand features in time series to obtain the preliminarily aligned features. The present invention will further refine and explore the form of the regression equation to be established (such as linear regression, logistic regression, or polynomial regression, etc.) in combination with the distribution characteristics of the actual statistical data in the subsequent work.
[0108] (2) Design of Multi-modal Fusion Model Based on Generative Adversarial Training: Further aiming at the characteristics of CS multi-modal asynchronous heterogeneity, in order to make full use of existing knowledge and explore the implicit spatial and semantic knowledge within the data to achieve the purpose of mutual enhancement of different modal information, the present invention intends to propose a multi-modal fusion model based on generative adversarial training. Different from the traditional generative adversarial network that only generates better outputs through the two-party game between the generator G and the discriminator D, we will innovatively introduce a third-party recognizer to cooperate and compete with the generator and the discriminator respectively, forming a three-party game problem, and its framework is as shown in Figure 5 shown. This framework consists of three parties (i.e., the generator, the discriminator, and the recognizer). During the multi-modal fusion process, the training of the recognizer model contains phoneme category information, which can be regarded as a semantic-level constraint. The adversarial training between the generator and the discriminator pays more attention to image detail information, which can be regarded as an image structure-level constraint. This framework introduces semantic information and structural information into multi-modal fusion at the same time, making the fused features more robust.
[0109] In the proposed multi-modal fusion model, first, the hand shape and lip features are extracted using the previously proposed self-supervised incremental joint model, and the hand position features are extracted based on the YOLO algorithm. Through the feature fusion model designed by us (see Figure 6 ), the lip-hand position features and lip-hand shape features are obtained respectively, and these two features are concatenated and fused as the final classification features for predicting and recognizing phoneme categories. The generator synthesizes lip and hand shape images, the discriminator discriminates between real images and synthesized images to form a "competition" relationship, while the recognizer "cooperates" with the generator to obtain more data to improve the recognition ability of the model, and the multi-modal heterogeneous small sample data fusion problem is solved uniformly under the "cooperation competition" mechanism. In the optimized objective function, the category constraint of the recognizer is used as a regularization term, and the different modal samples of online fusion are used for iterative training to improve the quality of phoneme category prediction both in terms of data appearance and semantic features. The total loss function is defined as follows:
[0110] L = L G + L D + βL R ,
[0111] where,
[0112]
[0113] Here, L G , L D , L R are the optimization losses of the generator, the discriminator, and the recognizer respectively, x 1 , x 2 represent lip and hand shape data respectively, G(x 1 , x 2) is the image synthesized from the real image I. w is the parameter of the recognizer f, which is used to predict the fused feature vector o i for prediction, y is the corresponding phoneme class label, and the constants α and β are used to balance the influence of each item.
[0114] The following elaborates in detail on the three important components of this model (recognizer, generator, and discriminator):
[0115] 1) Recognizer: In the CS multi-modal learning process, fusing complementary information between different modalities plays a very important role in improving the performance of the model. The present invention intends to propose a recognizer model based on efficient multi-modal feature fusion, mining the relationships between modalities from the multi-modal data after the initial alignment in the previous step to achieve efficient and highly robust phoneme class prediction. Specifically, after obtaining the initially aligned multi-modal data through the multi-modal asynchronous quantization analysis based on data analysis in the previous step, using the feature extraction model pre-trained on natural images as initialization, the self-supervised incremental learning joint model proposed in the first research scheme is used to obtain the lip feature, hand position feature, and hand shape feature respectively, and two of these features are combined (such as lip - hand position, lip - hand shape), and then the combined features are concatenated and input into the fully connected layer for phoneme prediction.
[0116] In addition, to make full use of the fused multi-modal data, we will design an efficient multi-modal feature fusion model (see Figure 6 ). For data of different modalities, the fusion model uses element-wise addition +, element-wise multiplication ×, and the maximization operation Max to achieve adaptive and efficient joint multi-modal feature fusion. The fused features use concatenation to retain the original information and then are restored to the original dimension through the convolutional layer Conv. Since the above multi-modal feature fusion model only involves element-wise operations and simple convolutional operations, it can be embedded into the recognizer to achieve end-to-end training to enhance the effect of feature fusion. Further, introducing a cross-modal self-attention mechanism into this fusion model to further implicitly align the above initially aligned features will be explored in subsequent research.
[0117] 2) Generator: The purpose of the generator is to synthesize available data. The synthesized data can not only make up for the lack of CS data, but also provide rich knowledge for the update and iteration of the recognizer. To this end, different from the traditional generative adversarial network that only uses a single sample as input, our generator will input data of both lip and hand shape modalities simultaneously, forcing the network to synthesize images of both modalities. The network of the generator adopts the encoder-decoder deep network structure, and uses dilated convolution as the convolutional layer to capture more local information of the lips and hand shapes. In the training stage, the generator and the recognizer form a "cooperative" relationship. Since the generator needs to generate more realistic images, it pays more attention to the local information of the images, which complements the semantic knowledge of the recognizer model. The two jointly promote the effectiveness of feature fusion in the unified multi-modal fusion framework. The trained generator model can be used to generate unlabeled samples. Considering that there may be a risk of mode collapse (i.e., the generator generates a single or limited mode) in the adversarial training process of the proposed multi-modal fusion model, we will specifically refine and explore this in subsequent research, and study a more stable tripartite "cooperation-competition" adversarial training framework from two perspectives: mini-batch discrimination and feature mapping in data diversity methods.
[0118] 3) Discriminator: On the one hand, the discriminator here is used to judge whether the synthesized images output by the generator are real, and forms a game adversarial relationship with the generator. In order to compete with the generator, it continuously improves its ability to distinguish the authenticity of images through iterative training. On the other hand, the generator can continuously improve its forgery ability to bypass the discrimination of the discriminator. The discriminator is a binary classification network and will initialize the model using a pre-trained deep convolutional network. Its role is to form a "competitive" relationship with the generator to learn a better generator, thereby providing more accurate semantic constraints and high-quality training samples.
[0119] III. Aiming at the deficiencies and challenges in the above-mentioned CS automatic recognition research in point (3), this embodiment proposes the design and implementation of an efficient CS automatic recognition overall model with privacy protection.
[0120] This stage belongs to the overall modeling stage of the present invention. It will combine the achievements of the previous feature learning and multi-modal fusion stages to construct an efficient CS automatic recognition neural network model with privacy protection. In view of the urgent problem of protecting the data privacy of data providers, the present invention mainly studies the federated learning framework with privacy protection (see Figure 7) Among them, a global server model with strong generalization ability plays a crucial role in the quality of the final model. It can automatically identify the parameters of the model by processing and aggregating the CS uploaded by each local client device, ensuring its communication with all devices participating in the training. The present invention intends to innovatively propose a method for constructing a shared dataset for training by using sample interpolation based on a secure hybrid algorithm to obtain a robust and strongly generalized global model. It can be specifically elaborated as the following two points.
[0121] The client, server, and the benchmark model of the final test model here are all structured as shown at the lower end of the figure. Among them, Bi-LSTM is a phoneme decoder. CTC is used as the loss function, directly outputting the probability of phoneme sequence prediction.
[0122] (1) Establishment of the basic federated learning CS automatic recognition framework
[0123] According to the existing federated learning framework, we first initialize the global server model using the pre-training method and introduce the client differential privacy algorithm to perturb the client parameters to increase the concealment of the data and avoid external malicious attacks. The process of the initially designed basic CS automatic recognition federated learning model is as follows:
[0124] 1) Considering the strong correlation between CS automatic recognition, sign language recognition, and lip reading recognition, we will collect multiple open-source sign language video datasets and lip reading video datasets and integrate them into a pre-training dataset with sufficient samples. Then, this dataset is used to pre-train and initialize the global server model;
[0125] 2) The server randomly selects a part of the devices and downloads the model initialization parameters to the selected client device models through the network;
[0126] 3) After downloading the model parameters, the client devices use the local CS video data for the training of their models (to ensure the efficiency of the CS automatic recognition model, the benchmark models of the client and the server here will utilize the feature learning model and the multi-modal fusion model proposed in the first two research schemes), and calculate the loss value and the gradient;
[0127] 4) Update the model parameters and upload them to the server. Here, in order to strengthen the protection of the data privacy of the participants, we propose to apply the client differential privacy algorithm during the model training stage, that is, to clip the model parameters and introduce Laplace noise before the client model uploads, so as to prevent malicious servers or external attackers from stealing data information through the model parameters uploaded by the participants;
[0128] 5) After the server collects the model parameters uploaded by all the selected devices or after a certain time limit, it aggregates the obtained model parameters to obtain a new global server model;
[0129] 6) Repeat steps 2 - 5 until convergence.
[0130] (2) Design of encrypted shared training based on sample interpolation of secure hybrid algorithm
[0131] In the important step 5) of the above basic framework, since the devices, methods, and backgrounds of video shooting by different coders CS may be different, there are significant differences in the video samples on each device participating in the training in the above algorithm, that is, P i (x|y) ≠ P j (x|y), where i, j represent devices, x represents data, and y represents the label corresponding to the data. In addition, the sample distributions of different devices may also be different, that is, P i (y) ≠ P j (y). These two points make the performance of the global server model aggregated by the models trained by these clients poor, and even the training cannot converge. This is the problem of non - independent and identically - distributed data faced by federated learning. To address this problem, most current research works adopt customized strategies, that is, on the basis of training the global model, each client has a variant of the global model to better adapt to its own data distribution. Although this method can have good performance on the clients participating in the training, its generalization ability for datasets with unknown distributions is weak. There are also methods that use the shared dataset on the server to participate in the training, but this way of directly taking data from each client will expose some privacy of users, and it is difficult to ensure that the distribution of other data is close to the distribution of the overall dataset (the sum of each client's dataset). Therefore, the present invention intends to propose a method for constructing a shared dataset to participate in the training based on sample interpolation of a secure hybrid algorithm. The specific description is as follows:
[0132] 1) According to the idea of federated learning, since the data comes from different clients, the distribution of the shared dataset constructed on the server can be approximately regarded as a linear combination of the dataset distributions on the clients. Combining prior knowledge, the linear interpolation of feature vectors can be used to expand the data distribution of the global model training set, thereby alleviating the problem of non - independent and identically - distributed data and making the obtained global model easier to converge. However, if the samples and labels of different clients are directly taken for random linear interpolation and mixing, it will lead to the leakage of client data privacy. Therefore, we intend to use the Diffie - Hellma (DH) key exchange protocol to implement a secure hybrid algorithm to encrypt the samples, so that the server can only obtain the final mixed data generated by the clients and cannot know the real data participated in the mixing by each client. The purpose of the DH protocol is to enable both parties to complete the key exchange without directly transmitting the key. The process is roughly as follows:
[0133] A and B negotiate the parameters of the DH protocol: a large prime number p and a base g;
[0134] A generates a random number a and calculates K A = g a mod p, and then sends K A to B. B obtains K B in the same way and sends it to A. K A , K B are the public keys, and a, b are the private keys;
[0135] After A receives K B , it calculates S AB = K B a mod p, and B calculates S BA = K A b mod p, and S AB = S BA .
[0136] After executing the first step in the above basic framework, based on this security protocol, the server selects m clients each time and allows these m clients to communicate with each other and generate their own keys {S ij} respectively, where i, j represent different devices selected from them. Then each client encrypts the data Q i = {x i , y i} it will upload as follows using DH encryption (as Figure 9 ) to obtain
[0137]
[0138] 2) The server linearly interpolates and mixes the b encrypted data from m clients as follows:
[0139]
[0140] where λ 1 + λ 2 + … + λ m = 1, 0 ≤ λ i ≤ 1. λ i is generated by the server and downloaded to the client, and its value will gradually approach 1 / m as the number of iterations increases.
[0141] After the above operations, b new samples and labels can be obtained. Then, the server repeats the above steps until N new data are generated to obtain the shared dataset Q'. This step will be carried out after the first step of the above basic framework. Based on the generated shared dataset Q', the initialized global server model is first preheated on Q' and then formally enters the model training stage. After the server selects clients in each communication round, a part of the data is randomly selected from Q' according to the ratio α and sent to the selected clients to participate in the training together with the local dataset. The optimal value of α will vary in different datasets and scenarios. Note that in this algorithm, since the whole process only needs to be carried out for several rounds and the clients participating in data mixing only need to communicate with the server for one round, this circumvents the communication cost problem and the problem of client disconnection to a certain extent. However, in the actual scenario, this algorithm may be at risk of malicious attacks. Therefore, we will subsequently explore methods such as gradient-based anomaly detection to deal with attacks on the model such as data poisoning.
[0142] In summary, the method for training the clue language automatic recognition model in this embodiment has at least the following beneficial effects:
[0143] 1) Efficient and robust multi-modal feature representation learning: a) It is intended to extract features from the perspective of unsupervised learning to get rid of the strong dependence on data labels and the influence brought by noisy annotations in previous methods; b) Further, a collaborative modeling method of self-supervised contrast learning and data distribution unbiased learning based on the spatio-temporal information constraint of video pictures is proposed, and a CS video feature unbiased learning model based on multi-modal information and self-supervised features is established, in order to solve the problem of unbalanced data category distribution from the perspective of model optimization. This method can be extended to other problems with noisy annotations and data biased distributions in the future, and has good continuity and potential influence.
[0144] 2) Fine and optimized asynchronous heterogeneous multi-modal feature fusion: a) First, combined with the prior knowledge of CS coding theory and the relationship between asynchronous delays mined from multi-modal data by statistical analysis methods, the pairwise asynchronous modalities can be initially aligned in an interpretable manner; b) Then, innovatively, an identifier, a discriminator and a generator are introduced to form a tripartite "competition-cooperation" relationship, and a set of practical multi-modal fusion solutions for generative adversarial training are proposed. The fusion method based on generative adversarial training introduced by this model will open a new perspective for multi-modal fusion, and the sequential optimization mechanism of initial alignment first and then fusion also has good inspiration for revealing the internal mechanism of the difficult problem of asynchronous fusion.
[0145] 3) Efficient and Privacy-Preserving Overall Modeling for CS Automatic Recognition: The overall audio-visual recognition model based on federated learning proposed in this project will, for the first time, consider the issue of protecting the privacy of CS coders, construct a method for sample interpolation based on a secure hybrid algorithm to generate a shared dataset for participation in training, and apply it to the training of the CS automatic recognition model under the federated learning framework. This research will provide valuable ideas for the exploration of current artificial intelligence models in terms of privacy protection.
[0146] Example 3
[0147] On the other hand, the present invention also proposes a training device for a clue language automatic recognition model. Please refer to Figure 10 , which shows the training device for the clue language automatic recognition model provided in Embodiment 3 of the present invention, applied in a server. The server is communicatively connected to multiple clients. The device includes:
[0148] A pre-training module 11, configured to obtain an open-source clue language video dataset and use the open-source clue language video dataset to train the server model;
[0149] A data distribution module 12, configured to randomly select multiple target clients and send the current global model parameters of the server model to each of the target clients, so that the target clients train their respective client models based on the current global model parameters and a preset training set. The architectures of the client model and the server model are the same;
[0150] A model update module 13, configured to receive the current client model parameters uploaded by each of the target clients, aggregate the current client model parameters, and use the aggregated model parameters to update the server model to obtain a new server model;
[0151] The data distribution module 12 and the model update module 13 are repeatedly executed until the server model converges to train a clue language automatic recognition model.
[0152] Preferably, in some optional embodiments of the present invention, the training device for the clue language automatic recognition model further includes:
[0153] A shared dataset construction module, configured to obtain clue language videos encrypted by a preset encryption algorithm uploaded by multiple clients to obtain a shared dataset composed of the obtained encrypted clue language videos;
[0154] Wherein, the data distribution module 12 is further configured to randomly select multiple target clients, send the current global model parameters of the server model to each of the target clients, and select a part of the shared data from the shared dataset and send it to each of the target clients according to a preset ratio;
[0155] Among them, the preset training set includes the local data set of the target client and a part of the shared data currently sent by the server.
[0156] Preferably, in some alternative embodiments of the present invention, before uploading the current client model parameters, the target client trims the current client model and introduces Laplace noise.
[0157] Preferably, in some alternative embodiments of the present invention, the architectures of the client model and the server model both include:
[0158] A self-supervised incremental learning joint model for extracting hand shape features and lip features from the data set;
[0159] A hand position recognition model for extracting hand position features from the data set;
[0160] An asynchronous heterogeneous multi-modal temporal alignment model for statistically modeling and predicting the leading amount of the movement of the hand shape and hand position relative to the lip movement based on the hand shape features, lip features, and hand position features, and using the predicted leading amount to adjust the hand shape features and hand position features in time sequence to obtain the hand shape features, lip features, and hand position features with preliminary temporal alignment;
[0161] A fusion model for predicting and recognizing the phoneme categories of the cue language based on the hand shape features, lip features, and hand position features with preliminary temporal alignment.
[0162] Preferably, in some alternative embodiments of the present invention, the fusion model is a multi-modal fusion model based on generative adversarial training composed of an identifier, a generator, and a discriminator.
[0163] The functions or operation steps implemented when the above-mentioned modules and units are executed are substantially the same as those in the above method embodiments, and will not be elaborated here.
[0164] Example 4
[0165] Please refer to Figure 11 , Embodiment 4 of the present invention proposes a server, including a processor 10, a memory 20, and a computer program 30 stored on the memory and executable on the processor. When the processor 10 executes the program 30, it implements the above-mentioned cue language automatic recognition model training method.
[0166] Among them, in some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips, which are used to run the program code stored in the memory 20 or process data, such as executing an access restriction program, etc.
[0167] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 20 may be an internal storage unit of the server, such as the hard disk of the server. In some other embodiments, the memory 20 may also be an external storage device of the server, such as a plug-in hard disk equipped on the server, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Preferably, the memory 20 may also include both an internal storage unit of the server and an external storage device. The memory 20 can be used not only to store application software installed on the server and various types of data, but also to temporarily store data that has been output or will be output.
[0168] It should be noted that Figure 11 the shown structure does not constitute a limitation on the server. In other embodiments, the server may include fewer or more components than shown in the figure, or combine some components, or have different component arrangements.
[0169] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the clue word automatic recognition model training method as described above.
[0170] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0171] More specific examples (a non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0172] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0173] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0174] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for training a clue language automatic recognition model, characterized in that, applied to a server, the server is communicatively connected to multiple clients, and the method includes: S01, obtaining an open-source clue language video dataset, and training a server model using the open-source clue language video dataset; S02, randomly selecting multiple target clients, and sending the current global model parameters of the server model to each of the target clients, so that the target clients train their respective client models based on the current global model parameters and a preset training set, and the architectures of the client model and the server model are the same; S03, receiving the current client model parameters uploaded by each of the target clients, aggregating the current client model parameters, and updating the server model using the aggregated model parameters to obtain a new server model; Repeat steps S02 - S03 until the server model converges to train a clue language automatic recognition model; wherein, the architectures of the client model and the server model both include: A self-supervised incremental learning joint model for extracting hand shape features and lip features from a dataset; A hand position recognition model for extracting hand position features from a dataset; An asynchronous heterogeneous multi-modal temporal alignment model for statistically modeling and predicting the lead amount of the movement of the hand shape and hand position relative to the lip movement based on the hand shape features, lip features, and hand position features, and using the predicted lead amount to temporally adjust the hand shape features and hand position features to obtain preliminarily temporally aligned hand shape features, lip features, and hand position features; A fusion model for predicting and recognizing the phoneme categories of the clue language based on the preliminarily temporally aligned hand shape features, lip features, and hand position features.
2. The method for training a clue language automatic recognition model according to claim 1, characterized in that, before the step of randomly selecting multiple target clients and sending the current global model parameters of the server model to each of the target clients, it further includes: Obtaining clue language videos encrypted by a preset encryption algorithm uploaded by multiple clients to obtain a shared dataset composed of the obtained encrypted clue language videos.
3. The method for training a clue language automatic recognition model according to claim 2, characterized in that, the step of randomly selecting multiple target clients and sending the current global model parameters of the server model to each of the target clients includes: Randomly selecting multiple target clients, sending the current global model parameters of the server model to each of the target clients, and selecting a part of the shared data from the shared dataset according to a preset ratio and sending it to each of the target clients; wherein, the preset training set includes the local dataset of the target client and the part of the shared data currently sent by the server.
4. The method for training a clue language automatic recognition model according to claim 1, characterized in that, before uploading the current client model parameters, the target client trims the current client model and introduces Laplace noise.
5. The training method of the clue language automatic recognition model according to claim 1, characterized in that, the fusion model is a multi-modal fusion model based on generative adversarial training composed of an identifier, a generator and a discriminator.
6. A training device for the clue language automatic recognition model, characterized in that, it is used to implement the method according to claim 1, and is applied to a server. The server is communicatively connected to multiple clients. The device includes: A pre-training module, configured to obtain an open-source clue language video data set, and use the open-source clue language video data set to train the server model; A data distribution module, configured to randomly select a plurality of target clients, and send the current global model parameters of the server model to each of the target clients, so that the target clients train their respective client models based on the current global model parameters and a preset training set. The architectures of the client model and the server model are the same; A model update module, configured to receive the current client model parameters uploaded by each of the target clients, aggregate the current client model parameters, and use the aggregated model parameters to update the server model to obtain a new server model; The data distribution module and the model update module are repeatedly executed until the server model converges, so as to train a clue language automatic recognition model.
7. The training device for the clue language automatic recognition model according to claim 6, characterized in that, it further includes: A shared data set construction module, configured to obtain the clue language videos encrypted by a preset encryption algorithm uploaded by multiple clients, and obtain a shared data set composed of the obtained encrypted clue language videos; wherein, the data distribution module is further configured to randomly select a plurality of target clients, send the current global model parameters of the server model to each of the target clients, and select a part of the shared data from the shared data set according to a preset ratio and send it to each of the target clients; wherein, the preset training set includes the local data set of the target client and a part of the shared data currently sent by the server.
8. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the training method of the clue language automatic recognition model according to any one of claims 1-5.
9. A server, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the program, it implements the training method of the clue language automatic recognition model according to any one of claims 1-5.
Citation Information
Patent Citations
Precision feedback federated learning method for privacy protection
CN113762530A