A method and device suitable for underwater posture recognition
By combining ResNet-50, transformer encoder and Wasserstein GAN to generate unseen gesture features, the problem of insufficient data for underwater gesture recognition is solved, efficient underwater gesture recognition is achieved, and the communication accuracy between automatic underwater robots and human divers is improved.
Patent Information
- Application Number
- CN202411399265.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-10-09
AI Technical Summary
Underwater gesture recognition has problems such as low contrast, blur, and color distortion in computer vision. The existing models lack annotated data sets and cannot recognize no gestures, which makes it difficult for human divers to communicate with automatic underwater robots.
A feature extraction network based on ResNet-50 is adopted, combined with transformer encoder and pre-trained CLIP model, and a conditional Wasserstein GAN is used to generate unseen gesture features, and zero-sample gesture recognition is achieved through a classifier.
It improves the accuracy of underwater gesture recognition, enhances the non-verbal communication ability between automatic underwater robots and human divers, and improves the efficiency of underwater tasks.
Smart Images

Figure CN119206872B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater detection, and specifically provides a method and device for underwater posture recognition. Background Art
[0002] Due to problems such as increasing hydrostatic pressure and oxygen, human divers face difficulties in collecting data from deep water. In the past few decades, many tasks have been handed over to autonomous underwater vehicles (AUVs), which can capture underwater images / videos and can also withstand the problems faced by human divers well. Therefore, they have been widely used in the fields of oceanography, naval warfare, information navigation, ocean scene understanding, etc.
[0003] In many underwater tasks, AUVs are accompanied by human divers. Since they cannot speak, they communicate non-verbally through different gestures. However, due to the lack of annotated datasets, underwater gesture recognition is a relatively underdeveloped field in computer vision, mainly facing the following two challenges:
[0004] First, underwater images have problems such as low contrast, blurriness, color distortion, and fuzziness. Therefore, traditional gesture recognition methods face difficulties in analyzing them. Second, existing gesture recognition models are mainly supervised and can only recognize gestures from a predefined set used for training the model. Obviously, it is impossible to list all gestures, and it is impossible to collect thousands of labeled images for every possible gesture that a human diver may use in the wild.
[0005] Therefore, it is urgent to improve this shortcoming. The present invention studies and improves the existing technologies and deficiencies, and provides a method and device for underwater posture recognition. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and device for underwater posture recognition to solve the problems raised in the above background art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] In the first aspect, a method for underwater posture recognition is provided, including the following steps:
[0009] S1. Input an image , where C is the number of channels, H is the height, W is the width, and R is the real number space. Extract visual features through a feature extraction network ;
[0010] S2. Construct a position embedding , with the vector dimension the same as , and the dimension is denoted as , and and Add them according to their positions, and send the result to the Transformer encoder E (the encoder E needs the spatial position of the visual feature as a token when calculating self-attention);
[0011] S3. Use the pre-trained model CLIP to extract image features;
[0012] S4. After decoding the output result of E, combine it with the features extracted by CLIP as the real features and input them into the GAN structure, and the GAN uses conditional Wasserstein GAN, that is, WGAN;
[0013] S5. The trained WGAN is used to generate visual features of unseen classes, and the classifier is trained using visible class data and synthetic unseen class data to achieve zero-shot gesture recognition.
[0014] Furthermore, in the step S1, the extraction network of the picture visual feature adopts a network based on ResNet-50 as the Backbone, and ResNet-50 belongs to a configurable option. Selecting other structures as the Backbone does not affect the results of this case.
[0015] Furthermore, in the step S2, the input of the Transformer encoder E is denoted as: The Transformer calculates self-attention for the input token, and the output is , and the output carries the picture context information and is sent to the decoder.
[0016] Furthermore, in the step S3, the visual features extracted by CLIP , and the visual feature extraction includes mapping projection to achieve projecting onto the feature domain space of CLIP, and the parameters and are both the vector dimensions output by the model and are preset constants. In specific implementation, the parameters and usually have default preset values (the configurable values are multiples of 64 and are usually less than 1024), and and The constant values of are determined by the specific implementation method of the selected pre-trained CLIP.
[0017] Furthermore, the Transformer network part adopts a dual-branch cross-attention mechanism, and the operations of the two branches are the same, except for the values of different Q (query), K (key), and V (value) (QKV are terms related to the calculation of the attention mechanism in the Transformer). The two branches are denoted as L (left) and R (right) respectively:
[0018] And the output of the attention adopts layer normalization method CorssAtt is the cross-attention operation. For specific reference to the relevant content of the Transformer, it belongs to a fixed operation;
[0019] Subsequently, the outputs of the left and right branches pass through a convolutional layer, whose role is to add weight adjustment. Finally, after applying the GELU activation function, it is used as the final output:
[0020] Among them, W and b are the network parameters of the corresponding convolutional layer, and the output represents how much mutual attention information is retained; the left and right branch attentions are fused through a feedforward network (FFN) to form the final output .
[0021] Furthermore, in step S4, regarding the WGAN adversarial generation network part:
[0022] The visual features are used as the input , and the semantic vector a is used as a conditional variable (a part in the WGAN, which is a trainable parameter). The generator G attempts to simulate the real distribution, and then generates a , which is sent into the adversarial network as the discriminator D and is compared;
[0023] The output of D is a classifier, outputting true or false; as data is continuously input for training, D can improve its prediction performance through training, continuously making the generated features closer to the real feature distribution. Among them, the definition of the loss function is as follows: Among them, λ is the penalty coefficient, is the feature generated by the network, z is the noise; the generator is G, the discriminator is D, and D(G(z,a)) represents the score of the discriminator for the generated feature; E represents the expectation, represents the distance, generally using L2, the Euclidean distance;
[0024] And to avoid mode collapse (that is, the result generated by the generator is true, but the diversity is insufficient), another mode-seeking loss function is introduced: Among them, different noises introduced when generating different features represents the L1 distance, which is the absolute value;
[0025] The optimization objective of the final loss function:
[0026] Among them, is an adjustable parameter.
[0027] Furthermore, in the step S5, regarding the classifier part:
[0028] Construct a classifier network. By default, a multi-layer fully connected network is adopted, and the cross-entropy is used as the loss function. For the prediction of known gestures, input the network of this case. The input of the classification network is the output of the gesture decoder, and the judgment is directly made according to the classification; then the features generated by the GAN, which belong to the synthesized unseen features and the visual features of the visible classes extracted from the trained structure, are used as the input of the classifier to realize the prediction of unseen type samples; and the results of the above two types of inputs are the same. The function is to use the generation network to generate many unseen samples according to the visible samples to train the network and make up for the problem of insufficient samples.
[0029] In the second aspect, a device applicable to underwater gesture recognition is provided, which is applied to the method applicable to underwater gesture recognition as described above. The device includes: a memory, a processor, and computer program instructions stored on the memory and executable on the processor. When the processor executes the computer program instructions, the method applicable to underwater gesture recognition as described above is realized.
[0030] In the third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the method applicable to underwater gesture recognition as described above is realized.
[0031] The present invention provides a method and a device applicable to underwater gesture recognition, having the following beneficial effects:
[0032] The present invention proposes a method, which combines the strong representation learning ability of the Transform structure as a visual feature extractor, and then uses a generative adversarial network to synthesize the visual features of unseen gestures, enabling us to train a gesture classifier using visible and unseen data, overcoming the weaknesses of supervised learning and improving the performance of gesture recognition; thereby improving the accuracy of gesture communication between the AUV and human divers in underwater tasks, facilitating the AUV to better capture underwater images / videos, and bringing more convenience to underwater detection work. Description of the Drawings
[0033] Figure 1 This is the overall flowchart of a method for underwater posture recognition according to the present invention;
[0034] Figure 2 This is the flowchart of the transformer network part of a method for underwater posture recognition according to the present invention. Specific embodiments
[0035] The following further describes in detail the embodiments of the present invention in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0036] As Figure 1 - Figure 2 shown, a method for underwater posture recognition includes the following steps:
[0037] S1. Input an image , where C is the number of channels, H is the height, W is the width, and R is the real number space. Visual features are extracted through a feature extraction network ; The extraction network of the visual features of the picture uses a network based on ResNet-50 as the Backbone, and ResNet-50 belongs to a configurable option. Selecting other structures as the Backbone does not affect the results of this case;
[0038] S2. Construct a position embedding , the vector dimension is the same as , and the dimension is denoted as . Add and according to the position, and the result is sent to the transformer encoder E (the encoder E needs the spatial position of the visual features as tokens when calculating self-attention);
[0039] The input of the transformer encoder E is denoted as: The transformer calculates self-attention for the input tokens, and the output is , and the output carries the picture context information and is sent to the decoder;
[0040] S3. Use a pre-trained model CLIP to extract image features (refer to the paper Radford, A., Kim, et al.: Learning transferable visual models from natural language supervision. In: International conference on machine learning, this case only uses it);
[0041] Visual features extracted by CLIP , and the visual feature extraction includes mapping projection to project onto the feature domain space of CLIP;
[0042] The transformer network part adopts a dual-branch cross-attention mechanism, and the operations of the two branches are the same, except for different values of Q (query), K (key), and V (value) (QKV are terms related to the calculation of the attention mechanism in the transformer). The two branches are denoted as L (left) and R (right):
[0043] And the output of the attention adopts layer normalization method
[0044] CorssAtt is the cross-attention operation. For specific reference to the relevant content of the transformer, it belongs to a fixed operation;
[0045] The subsequent outputs of the left and right branches pass through a convolutional layer, whose function is to add weight adjustment. Finally, after applying the GELU activation function, it is used as the final output:
[0046] Among them, W and b are the network parameters of the corresponding convolutional layer, and the output represents how much mutual attention information is retained; the left and right branch attentions are fused through a feedforward network (Feedforward network FFN) to form the final output ;
[0047] After the outputs of S4 and E are decoded, they are combined with the features extracted by CLIP as the real features and input into the structure of the GAN. And the GAN adopts conditional Wasserstein GAN, that is, WGAN; for the main details, refer to (Arjovsky, M., Chintala, S., Bottou, L.: Wasserstein generative adversarial networks). This case only uses it to mainly generate semantic features of classes;
[0048] Regarding the adversarial generation network part of WGAN:
[0049] The visual features are used as the input , and the semantic vector a is used as a conditional variable (a part in WGAN, which is a trainable parameter). The generator G tries to simulate the real distribution and then generates a , which is sent to the adversarial network as the discriminator D for comparison with ;
[0050] The output of D is a classifier that outputs true or false. As data is continuously input for training, D can improve its prediction performance through training, making the generated features increasingly close to the true feature distribution. The loss function is defined as follows:
[0051] where λ is the penalty coefficient, is the feature generated by the network, and z is the noise; the generator is G, the discriminator is D, and D(G(z,a)) represents the score of the discriminator for the generated feature; E represents the expectation, represents the distance, usually the L2, Euclidean distance;
[0052] And to avoid mode collapse (i.e., the result generated by the generator is true but lacks diversity), another mode-seeking loss function is introduced:
[0053] where, is the different noise introduced when generating different features, represents the L1 distance, which is the absolute value;
[0054] The optimization objective of the final loss function: where, is an adjustable parameter;
[0055] S5. The trained WGAN is used to generate visual features of unseen classes, and a classifier is trained using visible class data and synthetic unseen class data to achieve zero-shot gesture recognition;
[0056] Regarding the classifier part: A classifier network is constructed, defaulting to a multi-layer fully connected network, and the loss function uses cross-entropy. For the prediction of known gestures, the input to the network in this case is the output of the gesture decoder, and the classification is directly determined based on the classification. Subsequently, the features generated by the GAN, which belong to the synthetic unseen features and the visual features of visible classes extracted from the trained structure, are used as the input to the classifier to achieve the prediction of unseen type samples; and the results of the above two types of inputs are the same. The role is to use the generation network to generate many unseen samples based on visible samples to train the network and supplement the problem of insufficient samples.
[0057] A device applicable to underwater gesture recognition is applied to the method for underwater gesture recognition as described above. The device includes: a memory, a processor, and computer program instructions stored on the memory and executable on the processor. When the processor executes the computer program instructions, the method for underwater gesture recognition as described above is implemented.
[0058] A computer-readable storage medium stores computer-executable instructions that, when executed by a processor, are used to implement the method for underwater posture recognition as described above.
[0059] Embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to best explain the principles of the invention and its practical application, and to enable those of ordinary skill in the art to understand the invention and design various embodiments with various modifications suitable for a particular purpose.
Claims
1. A method for underwater gesture recognition, characterized in that: The following steps are involved: S1. Input image X∈R C×H×W , C is the channel, H is the height, W is the width, and R is the real space. The visual feature V is extracted through the feature extraction network b ; S2. Build a position embedding X pos , vector dimensions and V b Similarly, the dimensions are recorded as C'×H'×W', and X pos and V b Add by position and send the result to transformer encoder E; S3, use the pre-trained model CLIP to extract image features; After decoding, the output results of S4 and E are combined with the features extracted by CLIP and input into the structure of GAN as real features. The GAN adopts conditional Wasserstein GAN, i.e. WGAN. S5. The trained WGAN is used to generate visual features of the unseen class, and the classifier is trained using the visible class data and the synthesized unseen class data to achieve zero-shot gesture recognition. In step S3, the visual feature V extracted by CLIP c ∈R C”×k , and visual feature extraction includes mapping projection to achieve O e Projected to the feature domain space of CLIP, the parameters C" and k are the vector dimensions of the model output and are preset constants; In step S4, regarding the WGAN adversarial generative network part: Visual features as input O T , the semantic vector a is used as a conditional variable, the generator G tries to simulate the real distribution, and then generates a Send it to the adversarial network as the judge D, and O T Make a comparison; The output of D is a classifier, which outputs true or false. As data is continuously input for training, D can improve the prediction performance through training, and continuously make the generated features close to the real feature distribution. The loss function is defined as follows: Among them, λ is the penalty coefficient, is the feature generated by the network, z is the noise; the generator is G, the discriminator is D, D(G(z,a)) represents the score of the discriminator on the generated feature; E represents expectation, ||x,y|| represents the distance, using L2, Euclidean distance; And in order to avoid mode collapse, another mode-seeking loss function is introduced: LMS=E[||G(z1,a)-G(z2,a)||1 / ||z1-z2||1] Among them, z1 and z2 introduce different noises when generating different features, and ||·||1 represents the distance of L1, which is the absolute value; The final loss function optimization goal is: Among them, σ is an adjustable parameter.
2. A method for underwater gesture recognition according to claim 1, characterized in that: In step S1, the network for extracting the visual features of the image uses a network based on ResNet-50 as the backbone, and ResNet-50 is a configurable option.
3. A method for underwater gesture recognition according to claim 1, characterized in that: In step S2, the input of the transformer encoder E is recorded as: X e =X pos +V b The transformer performs self-attention calculations on the input token, and the output is O e ∈RC'×H'×W', the output carries the image context information and is fed into the decoder.
4. A method for underwater gesture recognition according to claim 3, characterized in that: The transformer network uses a two-branch cross attention mechanism, and the operations of the two branches are the same, but with different values of Q, K, and V. The two branches are denoted by L and R respectively: Q L =Q e' K L =V L =V C' Q R =V C' K R =V R =O e And the output of attention adopts layer normalization method A L =LN(Q L +CorssAtt(Q L ,K L ,V L )) A R =LN(Q R +CorssAtt(Q R ,K R ,V R )) CorssAtt is the cross attention operation; The outputs of the left and right branches are then passed through a 1×1 convolutional layer to add weight adjustment, and finally the GELU activation function is applied as the final output: g(A L )=GELU(W L A L +b L ) g(A R )=GELU(W R A R +b R ) Among them, W and b are the network parameters of the corresponding convolutional layer, and the output indicates how much mutual attention information is retained; the left and right branch attentions are fused through a forward feedback network to form the final output O T .
5. The method for underwater gesture recognition according to claim 1, characterized in that: In step S5, regarding the classifier part: A classifier network is constructed. By default, a multi-layer fully connected network is used, and the loss function uses cross entropy. For the prediction of known gestures, the network of this case is input. The input of the classification network is the output of the gesture decoder, and the judgment is made directly according to the classification. Then the features generated by GAN, which belong to the synthesized unseen features and the visual features of the visible class extracted from the trained structure, are used as the input of the classifier to realize the prediction of unseen type samples. The results of the above two types of inputs are the same.
6. A device for underwater gesture recognition, applied to the method for underwater gesture recognition as claimed in any one of claims 1 to 5, characterized in that: The device comprises: a memory, a processor, and computer program instructions stored in the memory and executable on the processor, wherein when the processor executes the computer program instructions, the method for underwater gesture recognition as described in any one of claims 1 to 5 is implemented.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method for underwater gesture recognition as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Image augmentation model training method and image classification method based on variational auto-encoder and generative adversarial network
CN114386534A
Underwater target identification method based on multi-modal fusion
CN118485908A