An Underwater Human Pose Recognition Method and System Based on MediaPipe
The MediaPipe algorithm and image restoration technology improve the quality of underwater images, combined with long-term memory networks and support vector regression networks, the problem of reduced posture recognition accuracy caused by poor underwater image quality is solved, and high-precision underwater human posture recognition is achieved.
Patent Information
- Application Number
- CN202510599816.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-12
AI Technical Summary
Poor underwater image quality leads to a decrease in the accuracy of human posture recognition, and it is difficult for the prior art to effectively extract the posture feature points of underwater operators.
The MediaPipe algorithm is used to invert the image degradation process frame by frame through image restoration technology, extract key point information frame by frame, and perform data distortion correction and enhancement processing, and establish a human posture recognition prediction model based on long-term and short-term memory networks and support vector regression networks.
It significantly improves the accuracy of underwater image quality and human posture recognition, improves the accuracy of recognition and prediction, enhances the generalization ability and robustness of the model, and solves the problem of poor image quality caused by underwater light absorption and scattering.
Smart Images

Figure CN120126060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of feature recognition, and particularly to an underwater human body pose recognition method and system based on MediaPipe. Background Art
[0002] Underwater operations are of great significance in the fields of marine resource exploration, oil and gas development, underwater infrastructure maintenance, and environmental monitoring, providing key data for environmental protection, climate change detection, and seabed resource development. In these scenarios, accurately identifying the pose of underwater operators is crucial. For example, in underwater engineering installation tasks, it is necessary to judge the operation standardization and safety based on the pose of underwater operators; in marine ecological monitoring, the pose of underwater operators may affect the accuracy of work such as sample collection. At the same time, there are also urgent needs in the fields of underwater rescue and underwater archaeology for quickly and accurately identifying the pose of underwater operators to improve the operation efficiency and safety.
[0003] Currently, human body pose recognition is generally based on underwater images. However, water has the effects of absorbing and scattering light, resulting in low contrast, color distortion, and blurriness of underwater images, seriously affecting the overall image quality. Moreover, there are complex water flows underwater, which will cause irregular shaking of personnel, increasing the difficulty of pose recognition. This makes it difficult for traditional human body pose recognition algorithms based on vision to effectively extract human feature points. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide an underwater human body pose recognition method and system based on MediaPipe, aiming to solve the technical problem in the prior art that due to the absorption and scattering effects of water on light, the quality of underwater images is poor, resulting in a decrease in the accuracy of human body pose recognition.
[0005] On the one hand, the present invention provides an underwater human body pose recognition method based on MediaPipe, and the method includes:
[0006] Obtain underwater video stream data, and inversely infer the image degradation process frame by frame through image restoration technology to obtain ideal state underwater images frame by frame;
[0007] Based on the MediaPipe algorithm, extract the frame-by-frame key point information of the human body pose of the frame-by-frame underwater images;
[0008] Perform data distortion correction on the frame-by-frame key point information and perform data enhancement processing through Gaussian noise to obtain frame-by-frame human body pose information;
[0009] Add a time series to the frame-by-frame human body pose information, and establish a human body pose recognition prediction model based on the long short-term memory network and the support vector regression network;
[0010] Identify and predict the pose information of underwater operators through the human pose recognition and prediction model.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the underwater human pose recognition method based on MediaPipe provided by the present invention, the image degradation process is inversely calculated frame by frame through the image restoration technology to obtain the ideal frame-by-frame underwater image, improving the quality of the image and the accuracy of human pose recognition. Optimize the key point positions based on the frame-by-frame underwater images to further improve the recognition ability and accuracy of the subsequent model. Then correct the distortion of the key point information, significantly improving the recognition and prediction accuracy. Finally, combine the long short-term memory network and the support vector regression network to establish a human pose recognition and prediction model to process time series, regression, and classification, enhancing the generalization ability of this model, improving the prediction accuracy and robustness, and reducing overfitting, thus solving the technical problem in the prior art that due to the absorption and scattering of light by water, the quality of underwater images is poor, resulting in a decrease in the accuracy of human pose recognition.
[0012] According to one aspect of the above technical solution, the steps of obtaining underwater video stream data and inversely calculating the image degradation process frame by frame through the image restoration technology to obtain the ideal frame-by-frame underwater image specifically include:
[0013] Collect underwater video stream data through a camera, perform frame-by-frame processing on the underwater video stream data to obtain frame-by-frame distorted underwater images;
[0014] Calculate the non-degraded underwater image from the frame-by-frame distorted underwater images through the image restoration technology. The calculation formula is:
[0015] ,
[0016] where, is the ideal frame-by-frame underwater image, is the frame-by-frame distorted underwater image, is the pixel point of the frame-by-frame distorted underwater image, is various forms of light entering the camera, is the distance from the target to the camera, is the background light, is the transmittance, which decays exponentially with , is the attenuation coefficient of the reflected light, is the energy after normalization loss.
[0017] According to one aspect of the above technical solution, the method further includes:
[0018] Based on the frame-by-frame distorted underwater image, a color space is constructed, the initial values in the color channels of the color space are obtained, the average value of the initial values is calculated, and the initial values of each pixel point in the color channels are corrected to obtain the corrected values of the color channels. The calculation formula is as follows:
[0019] ,
[0020] where, 、 are the corrected values of the pixel point in the color channels 、 respectively, 、 are the initial values of the pixel point in the color channels 、 respectively, 、 are the average values of the corresponding color channels 、 in the entire frame-by-frame distorted underwater image respectively, 、 are parameters, which are the correction levels for evaluating the two initial color channels 、 , is the influence value of the pixel point on the background light position;
[0021] Perform scattering light processing on the frame-by-frame distorted underwater image, and construct a scattered underwater image with the scattered light diffusing outward from the center of the background light. Therefore, the scattered underwater image is represented in the form of spherical coordinates centered on the background light, and is expressed as follows:
[0022] ,
[0023] where, is the scattered underwater image, is the state in which the scattered light is uniformly scattered at a pre-given different angle, is the distance from each pixel point in the scattered light to the background light, is the pixel point in the scattered light;
[0024] Based on the scattered underwater image, calculate the pixel point mapping value of the scattered light. The calculation formula is as follows:
[0025] ,
[0026] where, is the pixel point mapping value of the scattered light, is the spherical coordinate range centered on respectively, is half of the distance from the center point of the background light to the dark pixel point, is the maximum distance from the pixel point in the scattered light to the background light, is the scattering coefficient;
[0027] According to the pixel point mapping value, calculate the transmittance, and the calculation formula is as follows:
[0028] ,
[0029] wherein, is the pixel point in the fog line at each pixel point to the background light distance, is the pixel point pixel point mapping value of the scattered light at the location;
[0030] Combine the scattered underwater image with the color channel to calculate the background light, and the calculation formula is as follows:
[0031] ,
[0032] wherein, is the position of the pixel point with the maximum brightness, is centered on the pixel point is a local positive direction channel, is centered on the pixel point is the channel point in a local positive direction channel, , , are respectively the channel point red, green, and blue light intensity components in the color channel, , , are respectively , , different weights, is each local area corresponding to each pixel point in the frame-by-frame distorted underwater image.
[0033] According to one aspect of the above technical solution, the steps of extracting the frame-by-frame key point information of the human pose of the frame-by-frame underwater image based on the MediaPipe algorithm specifically include:
[0034] According to the underwater video stream data, determine the key point information required to be marked by the underwater operator, and optimize the key point positions of the human pose in the Blazepose model;
[0035] Extract the frame-by-frame key point information of the human pose in the frame-by-frame underwater image according to the optimized key point positions. The frame-by-frame key point information includes frame-by-frame key point coordinates and the corresponding confidence information.
[0036] According to one aspect of the above technical solution, the step of performing data distortion correction on the frame-by-frame key point information and performing data enhancement processing through Gaussian noise to obtain frame-by-frame human pose information specifically includes:
[0037] For the frame-by-frame key point coordinates , with pixel coordinates , perform alignment processing on the frame-by-frame key point coordinates and the pixel coordinates respectively to obtain homogeneous key point coordinates , homogeneous pixel coordinates , and perform 3D projection calculation. The calculation formula is as follows:
[0038] ,
[0039] where is the scale factor, is the intrinsic matrix 's rotation and translation parameters, is the intersection coordinate of the camera and the imaging plane, , is the focal length of the camera relative to the pixel with pixel coordinates , is the parameter describing the tilt of the image axis;
[0040] Assume that all feature points of the homogeneous key point coordinates are located in the plane with the vertical coordinate being zero. Then the calculation formula for 3D projection is expressed as:
[0041] ,
[0042] Use the homography matrix to replace to obtain , where , are the first two columns of the homography matrix, , are 's first two columns. Since the rotation matrix has orthogonality, that is: , , the following expression can be obtained:
[0043] ,
[0044] Solve for the intrinsic matrix to obtain the value of the parameter describing the tilt of the image axis;
[0045] According to the parameters describing the tilt of the image axis, the calibrated radial distortion parameters, and the tangential distortion parameters, the coordinates of the key points in each frame are corrected for distortion to obtain the corrected key point coordinates. The calculation formula is as follows:
[0046] ,
[0047] where 、 are respectively 、 the corrected coordinate values, is the radial distortion parameter within the central region of the underwater image for each frame, is the radial distortion parameter within the edge region of the underwater image for each frame, is the horizontal tangential distortion parameter, is the vertical distortion parameter;
[0048] Replace the coordinates of the key points in each frame with the corrected key point coordinates to obtain the corrected key point information, and perform geometric transformation on the corrected key point information to simulate the changes in vision and distance;
[0049] And perform data augmentation processing through Gaussian noise to obtain the human body pose information for each frame.
[0050] According to one aspect of the above technical solution, the steps of adding a time series to the human body pose information for each frame and establishing a human body pose recognition and prediction model based on a long short-term memory network and a support vector regression network specifically include:
[0051] Add a time series to the human body pose information for each frame, obtain the changes in the coordinates of a certain corrected key point in the time series of the human body pose information for each frame, and construct an action data set, expressed as:
[0052] ,
[0053] where is the action data set, is the coordinate of the corrected key point at the th time series;
[0054] Use the action data set as the input of the human body pose recognition and prediction model, and output the hidden state through the gate mechanism of the long short-term memory network, expressed as:
[0055] ,
[0056] where is the set of hidden states, is the hidden state at the th time series;
[0057] Select is the final feature vector, where the loss function of the long short-term memory network is:
[0058] ,
[0059] where, is the observation value of the long short-term memory network at the time series, is the average observation value of all time series of the long short-term memory network, = 1, 2,... ;
[0060] Input the final feature vector into the support vector regression network for regression calculation to train the support vector regression network for prediction. The calculation formula is as follows:
[0061] ,
[0062] where, is the output result of the support vector regression network, and the loss function of the support vector regression network is:
[0063] ,
[0064] where, is the observation value of the support vector regression network at the time series, the average observation value of all time series of the support vector regression network, is a preset insensitive threshold.
[0065] According to one aspect of the above technical solution, the steps of identifying and predicting the pose information of underwater operators through the human pose recognition and prediction model specifically include:
[0066] Obtain the pose information of the underwater operator, and calculate the key limb angle, key limb elevation angle, key elbow rotation angle, and key elbow swing angle.
[0067] The calculation formula of the key limb angle is as follows:
[0068] ,
[0069] where, is the distance between the angle key point and the angle key point , is the key limb angle between the angle key point and the connection line of the angle key point , is from the angle key point pointing to the angle key point The vector is the value of the coordinate dimension in the vector, is the value of the coordinate dimension in the vector, is the vector pointing from the included angle key point to the included angle key point ; is the coordinate dimension, is the total number of dimensions, is the coordinate of the included angle key point ; is the coordinate of the included angle key point ;
[0070] The calculation formula for the key limb lifting angle is as follows:
[0071] ,
[0072] where is the key limb lifting angle, and are respectively the abscissa and ordinate of the lifting key point , and are respectively the abscissa and ordinate of the lifting key point , and are respectively the abscissa and ordinate of the lifting key point ;
[0073] The calculation formula for the key elbow joint rotation angle is as follows:
[0074] ,
[0075] where is the key elbow joint rotation angle, and are respectively the abscissa and ordinate of the rotation key point , and are respectively the abscissa and ordinate of the rotation key point , and are respectively the abscissa and ordinate of the rotation key point ;
[0076] The calculation formula for the key elbow joint swing angle is as follows:
[0077] ,
[0078] Among them, is the key elbow joint swing angle, and are respectively the vertical coordinate and the vertical coordinate of the swing key point ; and are respectively the vertical coordinate and the vertical coordinate of the swing key point ; and are respectively the vertical coordinate and the vertical coordinate of the swing key point ;
[0079] According to the key limb angle, the key limb lifting angle, the key elbow joint rotation angle, and the key elbow joint swing angle, through the human posture recognition and prediction model, the posture actions of the underwater operator are recognized and predicted.
[0080] According to one aspect of the above technical solution, after the step of recognizing and predicting the posture actions of the underwater operator according to the key limb angle, the key limb lifting angle, the key elbow joint rotation angle, and the key elbow joint swing angle through the human posture recognition and prediction model, it further includes:
[0081] Comparing the posture action with the standard posture action, calculating the deviation, and obtaining the action deviation value. The calculation formula is as follows:
[0082] ,
[0083] Among them, is the action deviation value, is the coordinate of the key point of the posture action, is the coordinate of the key point of the standard posture action, is the total number of key points;
[0084] According to the action deviation value, combined with the dynamic actions and speed of the underwater operator, the correctness of the posture of the underwater operator is detected.
[0085] According to one aspect of the above technical solution, the step of detecting the correctness of the posture of the underwater operator according to the action deviation value, combined with the dynamic actions and speed of the underwater operator, specifically includes:
[0086] According to the action deviation value, combined with the dynamic actions and speed of the underwater operator, it is judged whether the posture of the underwater operator is correct;
[0087] If so, continue to perform the next posture information recognition and prediction;
[0088] Otherwise, remind the underwater operator to correct the attitude information.
[0089] Another aspect of the present invention is to provide an underwater human pose recognition system based on MediaPipe, which is used to implement the above-mentioned underwater human pose recognition method based on MediaPipe. The system includes:
[0090] An image restoration module, which is used to obtain underwater video stream data, inversely infer the image degradation process frame by frame through image restoration technology, and obtain ideal state underwater images frame by frame;
[0091] An information extraction module, which is used to extract the key point information of the human pose of each frame of the underwater image frame by frame based on the MediaPipe algorithm;
[0092] An information correction module, which is used to correct the data distortion of the key point information of each frame, and perform data enhancement processing through Gaussian noise to obtain the human pose information of each frame;
[0093] A model construction module, which is used to add a time series to the human pose information of each frame, and establish a human pose recognition prediction model based on the long short-term memory network and the support vector regression network;
[0094] A recognition and prediction module, which is used to recognize and predict the pose information of the underwater operator through the human pose recognition prediction model. Description of the Drawings
[0095] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, where:
[0096] Figure 1 It is a schematic flowchart of the underwater human pose recognition method based on MediaPipe in Embodiment 1 of the present invention;
[0097] Figure 2 It is the key point positions of the human pose in the optimized Blazepose model in Embodiment 1 of the present invention;
[0098] Figure 3 It is a structural block diagram of the underwater human pose recognition system based on MediaPipe in Embodiment 2 of the present invention;
[0099] Explanation of the symbols of the components in the drawings:
[0100] Image restoration module 100, information extraction module 200, information correction module 300, model construction module 400, recognition and prediction module 500. Detailed Embodiments
[0101] To make the objectives, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0102] Embodiment 1
[0103] Please refer to Figure 1 - Figure 2 , a method for underwater human pose recognition based on MediaPipe provided by the first embodiment of the present invention is shown, and the method includes steps S10 - S14:
[0104] Step S10, obtain underwater video stream data, and inversely calculate the image degradation process frame by frame through image restoration technology to obtain underwater images in an ideal state frame by frame;
[0105] Among them, due to underwater microparticles or plankton, the absorption and refraction of light by water bodies, the forward and backward scattering of light, moving objects, etc., all of which will cause the details of underwater images to be blurred, the color to fade, and the contrast to decrease. These phenomena will cause distortion, blurring, distortion, and additional noise in the process of image acquisition, transmission, storage, and processing. Therefore, it is necessary to process and analyze the underwater video stream data through image restoration technology and enhancement technology to improve the quality of the images to support the effective progress of underwater tasks.
[0106] Specifically, collect underwater video stream data through a camera, process the underwater video stream data frame by frame to obtain distorted underwater images frame by frame;
[0107] Calculate the non-degraded underwater image from the distorted underwater images frame by frame through image restoration technology, and the calculation formula is:
[0108] ,
[0109] Among them, is the underwater image in an ideal state frame by frame, is the distorted underwater image frame by frame, is the pixel point of the distorted underwater image frame by frame, is various forms of light entering the camera, is the distance from the target to the camera, is the background light, is the transmittance, and it decays exponentially with , is the attenuation coefficient of the reflected light, is the energy after normalization loss.
[0110] It can be understood that is the background light, that is, the per-frame distorted underwater image with respect to depth to infinity, that is, the ambient light (such as water body scattered light) when the light propagates to infinity.
[0111] In the formula, and The obtaining steps are as follows:
[0112] Based on the per-frame distorted underwater image, construct a color space, obtain the initial values of the color space in the color channels, calculate the average value of the initial values, correct the initial values of each pixel point in the color channels, and obtain the corrected values of the color channels. The calculation formula is as follows:
[0113] ,
[0114] Wherein, , are respectively the corrected values of the pixel point in the color channels , , , are respectively the initial values of the pixel point in the color channels , , , are respectively the average values of the corresponding color channels , in the entire per-frame distorted underwater image, , are parameters, which are the correction levels for evaluating the two initial color channels , , usually set to 1, is the influence value of the pixel point on the background light position;
[0115] It should be noted that underwater color correction is established in the color space. Regarding the average value of each color channel as the neutral color, by subtracting the average value of the color channel in the entire per-frame distorted underwater image from the color channel itself, the color deviation of the per-frame distorted underwater image can be significantly reduced.
[0116] Perform scattered light processing on the per-frame distorted underwater image, and construct a scattered underwater image with the scattered light diffusing outward from the center of the background light. Therefore, the scattered underwater image is represented in the form of spherical coordinates centered on the background light, and is expressed as follows:
[0117] ,
[0118] Wherein, is the scattered underwater image, is a state where scattered light is uniformly scattered at different preset angles. is the distance from each pixel point in the scattered light to the background light. is a pixel point in the scattered light.
[0119] Based on the scattered underwater image, calculate the pixel point mapping value of the scattered light. The calculation formula is as follows:
[0120] ,
[0121] where, is the pixel point mapping value of the scattered light. is the spherical coordinate range with as the center point. is half of the distance from the center point of the background light to the dark pixel point. is the maximum distance from the pixel point in the scattered light to the background light.
[0122] It should be noted that the larger the value, the clearer the pixel point. However, due to the influence of water body absorbing light, the brightness value of the pixel point farthest from the background light will approach 0, resulting in this part of the area being darker and the image being blurred. Therefore, it is necessary to optimize the pixel point mapping value of the scattered light. When is greater than optimize the pixel point mapping value so that is mapped to a pixel point position with higher brightness. Generally, = 0.833.
[0123] According to the pixel point mapping value, calculate the transmittance. The calculation formula is as follows:
[0124] ,
[0125] where, is the distance from each pixel point in the fog line at pixel point to the background light . is the pixel point mapping value of the scattered light at pixel point .
[0126] Combine the scattered underwater image with the color channel to calculate the background light. The calculation formula is as follows:
[0127] ,
[0128] where, is the position of the pixel point with the maximum brightness. is with pixel point A local positive direction channel centered on is a pixel point is a channel point in a local positive direction channel centered on , , are respectively the light intensity components of red, green, and blue in the color channel , , are respectively , , different weights of are the respective local regions corresponding to each pixel point in the frame-by-frame distorted underwater image .
[0129] Step S11, based on the MediaPipe algorithm, extract the frame-by-frame key point information of the human pose in the frame-by-frame underwater image;
[0130] Specifically, according to the underwater video stream data, determine the key point information required to be marked by the underwater operator, and optimize the key point positions of the human pose in the Blazepose model;
[0131] According to the optimized key point positions, extract the frame-by-frame key point information of the human pose in the frame-by-frame underwater image, and the frame-by-frame key point information includes frame-by-frame key point coordinates and the corresponding confidence information.
[0132] Furthermore, based on the underwater video stream data, extract the outline and bone key points of the underwater operator and integrate a custom SSD MobileNet object detection model, and at the same time extract the information of the person and surrounding objects in the picture. To reduce the data bandwidth and reduce the signal transmission delay, the present invention will redefine the 33 key points on the human joints marked by default in the Blazepose model to optimize the inference process and inference rate. Considering that the underwater operator wears diving goggles and a diving cap on the head when moving underwater, covering the two ears of the underwater operator, wears a diving mask and a breathing bottle on the face, so that the inner and outer corners of the eyes and the inner and outer corners of the mouth cannot be precisely refined, and wears a diving suit on the body, covering the chest and abdominal areas, and these areas that cannot be seen clearly cannot provide effective learning features for the follow-up; as Figure 2 shown, when optimizing the key point positions of the human pose in the Blazepose model, subtract the two ears, the nose, and the inner and outer sides of the left and right eyes, merge the left and right side mouth corners, add a neck joint point and redefine the left and right foot toes as the left and right foot flippers. Since the chest and abdominal areas of the underwater operator are covered by the diving suit, add a new neck joint and connect it to the left and right iliacs to optimize the detection of the chest and abdominal areas by the original Blazepose model, and redefine the detection toe area as the detection of the toes in order to optimize and obtain more accurate joint positions.
[0133] Therefore, by selecting new joint definitions, expanding or modifying the positions of existing key points, the ability of the Balzepose model to identify the positions of key points of underwater operators in the underwater environment is optimized, enhancing the accuracy and robustness of the data, simplifying steps such as coordinate transformation and post-processing, thereby accelerating the computing power of the Blazepose model, reducing the real-time transmission of data size, and enabling more effective underwater-to-land data communication. Compared with other complex vision algorithms, the per-frame key point information extracted by the Blazepose model has less computational complexity and can run efficiently on mobile devices in an offline environment, achieving faster, more accurate, and more robust detection and extraction of per-frame key point information in the underwater environment.
[0134] In addition, due to the problems of bandwidth limitation, high latency, and signal attenuation in underwater communication, it is difficult to support the transmission of large amounts of data, and it is vulnerable to the influence of factors such as water flow, temperature, salinity, and depth. The amount of data calculated by the Mediapipe algorithm is greatly reduced, which can significantly reduce the bandwidth and energy consumption required for transmission. In addition, the key point information for each frame is easier to compress and encrypt, which is beneficial to improving the transmission efficiency and transmission security. The present invention uses a technology that combines Li-Fi (Light Fidelity) technology and blockchain to achieve underwater communication, that is, uses visible light for data transmission. Compared with traditional radio communication, it has the advantages of high bandwidth, easy deployment, high transmission efficiency, low interference, and high security. The blockchain provides the ability to record decentralized and tamper-proof data. The key point information for each frame includes the key point coordinates for each frame and the corresponding confidence information. Compared with the original video, the amount of data information is significantly reduced, which is suitable for low-bandwidth transmission. Specifically, the data can be serially sent by the Arduino module. After being converted into a serial bit stream, the data is sent from the output pin of the Arduino Mega to the input of the transistor, and the encoded data is emitted in the form of an optical signal through the light-emitting diode light source. When the optical signal reaches the node, the photodetector will convert it into a voltage or current signal and transmit it to the computer efficiently and quickly through the terrestrial transmission system. However, in the actual transmission process, there is a risk of data leakage in the data sent from the underwater system. To avoid this problem, blockchain technology is used at the same time to provide additional security for the data. The data transmits this information by setting up a smart contract. The smart contract is used to automate the hashing process, aiming to provide guidance for data verification and blockchain data storage. The data is first sent to the serial monitor of the Arduino IDE, and then Python code is used to retrieve the data using blockchain technology and encrypt it using the Advanced Encryption Standard (AES) algorithm. Encryption technology (such as RSA) can be used to create a digital signature for each data packet. During the transmission process, the digital signature will be added to the encrypted data. In the present invention, the blockchain is used to securely retrieve the data and view the data in the mobile interface. Each access is stored in the form of a block to ensure that the data has not been tampered with or lost. In this way, compared with traditional wireless communication methods, this embodiment provides a faster and more secure way to send and receive data information.
[0135] Step S12: Perform data distortion correction on the key point information for each frame, and perform data enhancement processing through Gaussian noise to obtain the human body posture information for each frame;
[0136] It should be noted that during the acquisition process, there are problems such as noise, inaccurate frame-by-frame key-point information positioning, or environmental light and occlusion. To obtain good imaging effects, a lens and a protective cover are generally added in front of the underwater camera, and the lens and the protective cover will cause light to be distorted during propagation. In addition, the frame-by-frame key-point information will also be blocked or visually interfered, or the frequency of occurrence of some actions or postures may be low, etc., which will overfit the more common actions, resulting in unbalanced data categories. Therefore, this data needs to be corrected and enhanced for data distortion to improve its quality, reliability, and the robustness and generalization ability of subsequent neural networks for action classification and prediction.
[0137] Specifically, for the frame-by-frame key-point coordinates , with pixel coordinates , alignment processing is respectively performed on the frame-by-frame key-point coordinates and the pixel coordinates to obtain homogeneous key-point coordinates , homogeneous pixel coordinates , and 3D projection calculation is carried out. The calculation formula is as follows:
[0138] ,
[0139] where is the scale factor, is the intrinsic matrix 's rotation and translation parameters, is the intersection coordinate of the camera and the imaging plane, , is the focal length of the camera relative to the pixel with pixel coordinates , is the parameter describing the tilt of the image axis;
[0140] Assuming that all feature points of the homogeneous key-point coordinates are located in the plane with the vertical coordinate being zero, the calculation formula of 3D projection is expressed as:
[0141] ,
[0142] Using the homography matrix to replace to obtain , where , are the first two columns of the homography matrix, , are 's first two columns. Since the rotation matrix has orthogonality, that is: , , the following expression can be obtained:
[0143] ,
[0144] Solving the intrinsic matrix to obtain the parameters describing the tilt of the image axis of the numerical value;
[0145] According to the parameters describing the tilt of the image axis, the calibrated radial distortion parameters, and the tangential distortion parameters, the distortion correction is performed on the coordinates of the key points frame by frame to obtain the corrected key point coordinates. The calculation formula is as follows:
[0146] ,
[0147] wherein, , are respectively , the corrected coordinate values, is the radial distortion parameter within the central region of the underwater image frame by frame, is the radial distortion parameter within the edge region of the underwater image frame by frame, is the horizontal tangential distortion parameter, is the vertical distortion parameter;
[0148] Replace the coordinates of the key points frame by frame with the corrected key point coordinates to obtain the corrected key point information, and perform geometric transformation on the corrected key point information to simulate the changes in vision and distance;
[0149] And perform data augmentation processing through Gaussian noise to obtain the human pose information frame by frame.
[0150] Therefore, the key point information frame by frame can significantly improve the accuracy of recognition and prediction after correction and enhancement, reduce the dependence on high-quality labeled data, and make the subsequent analysis and prediction of the heading of underwater operators more accurate and efficient.
[0151] Step S13, add a time series to the human pose information frame by frame, and establish a human pose recognition and prediction model based on the long short-term memory network and the support vector regression network;
[0152] It should be noted that adding a time series to the human pose information frame by frame will present a complex non-linear dynamic pattern with the changes of actions and environments, and the relationships between data such as the angles and positions of joints are not linear. Especially when underwater operators make various movements underwater (such as turning around, kicking, rolling, etc.), these movements are often non-linear. The long short-term memory network processes the input at different time steps through its activation functions (sigmoid, tanh), and these activation functions themselves are also non-linear; even if the input data is linear, the internal state update of the long short-term memory network and the modeling process of the time series will introduce non-linear features. At the same time, the present invention introduces the idea of the support vector regression network to find the optimal hyperplane to maximize the interval from the data points to this hyperplane.
[0153] Specifically, a time series is added to the frame-by-frame human pose information, the change of the coordinates of a certain corrected key point in the frame-by-frame human pose information on the time series is obtained, and an action data set is constructed, expressed as:
[0154] ,
[0155] Among them, is the action data set, is the coordinate of the corrected key point at the time series;
[0156] The action data set is used as the input of the human pose recognition prediction model, and the hidden state is output through the gate mechanism of the long short-term memory network, expressed as:
[0157] ,
[0158] Among them, is the hidden state set, is the hidden state at the time series;
[0159] Select as the final feature vector, where the loss function of the long short-term memory network is:
[0160] ,
[0161] Among them, is the observation value of the long short-term memory network at the time series, is the average observation value of all time series of the long short-term memory network, = 1, 2,... ;
[0162] The final feature vector is input into the support vector regression network for regression calculation to train the support vector regression network for prediction. The calculation formula is as follows:
[0163] ,
[0164] Among them, is the output result of the support vector regression network, and the loss function of the support vector regression network is:
[0165] ,
[0166] Among them, is the observation value of the support vector regression network at the time series, the average observation value of all time series of the support vector regression network, is a preset insensitive threshold.
[0167] Therefore, by combining the long short-term memory network and the support vector regression network, a human pose recognition and prediction model is established to handle time series, regression, and classification, enhancing the generalization ability of this model, improving the prediction accuracy and robustness, and reducing overfitting.
[0168] Step S14, through the human pose recognition and prediction model, identify and predict the pose information of the underwater operator.
[0169] Specifically, obtain the pose information of the underwater operator, and calculate the key limb angle, the key limb lifting angle, the key elbow rotation angle, and the key elbow swing angle.
[0170] The calculation formula for the key limb angle is as follows:
[0171] ,
[0172] Among them, is the angle key point to the angle key point the distance between, is the angle key point and the angle key point the key limb angle between the connecting lines, is from the angle key point pointing to the angle key point vector, is the value of the coordinate dimension in the vector, is the value of the coordinate dimension in the vector, is from the angle key point pointing to the angle key point vector, is the coordinate dimension, is the total number of dimensions, is the angle key point coordinates, is the angle key point coordinates;
[0173] The calculation formula for the key limb lifting angle is as follows:
[0174] ,
[0175] Among them, is the key limb lifting angle, and are respectively the lifting key points The abscissa and ordinate of and are respectively the abscissa and ordinate of the lifting key point ; and are respectively the abscissa and ordinate of the lifting key point ;
[0176] The calculation formula for the key elbow joint rotation angle is as follows:
[0177] ,
[0178] wherein, is the key elbow joint rotation angle, and are respectively the abscissa and ordinate of the rotation key point ; and are respectively the abscissa and ordinate of the rotation key point ; and are respectively the abscissa and ordinate of the rotation key point ;
[0179] The calculation formula for the key elbow joint swing angle is as follows:
[0180] ,
[0181] wherein, is the key elbow joint swing angle, and are respectively the ordinate and vertical coordinate of the swing key point ; and are respectively the ordinate and vertical coordinate of the swing key point ; and are respectively the ordinate and vertical coordinate of the swing key point ;
[0182] According to the key limb angle, the key limb lifting angle, the key elbow joint rotation angle, and the key elbow joint swing angle, the posture and movement of the underwater operator are recognized and predicted through the human posture recognition and prediction model.
[0183] The method further includes:
[0184] Comparing the posture and movement with the standard posture and movement, calculating the deviation to obtain the action deviation value, and the calculation formula is as follows:
[0185] ,
[0186] Among them, is the action deviation value, is the key point of the posture action coordinates, is the key point of the standard posture action coordinates, is the total number of key points;
[0187] According to the action deviation value, combined with the dynamic actions and speeds of underwater operators, the correctness of the postures of underwater operators is detected.
[0188] Specifically, according to the action deviation value, combined with the dynamic actions and speeds of underwater operators, it is judged whether the postures of underwater operators are correct;
[0189] If so, continue to identify and predict the next posture information;
[0190] If not, remind to correct the posture information of the underwater operator.
[0191] Furthermore, in the face of complex scenarios, the instructor or expert gives operation suggestions and technical guidance according to the real-time posture information and heading information of the underwater operator, and combines the decision-making data of the system to generate a voice signal and transmit it to the underwater operator through the underwater communication device.
[0192] Compared with the prior art, the underwater human posture recognition method based on MediaPipe shown in this embodiment is adopted. By using the image restoration technology to inversely deduce the image degradation process frame by frame, the ideal state of the underwater image frame by frame is obtained, the quality of the image is improved, and the accuracy of human posture recognition is improved. Optimize the key point positions based on the underwater image frame by frame, further improve the recognition ability and accuracy of the subsequent model, and then correct the distortion of the key point information, significantly improving the recognition and prediction accuracy. Finally, combine the long short-term memory network and the support vector regression network to establish a human posture recognition and prediction model to process time series, regression and classification, enhance the generalization ability of this model, improve the prediction accuracy and robustness and reduce overfitting, thus solving the technical problem in the prior art that due to the absorption and scattering of light by water, the quality of underwater images is poor, resulting in the decline of the accuracy of human posture recognition.
[0193] Embodiment 2
[0194] Please refer to Figure 3 , which shows a MediaPipe-based underwater human posture recognition system provided by the second embodiment of the present invention. The system includes:
[0195] An image restoration module 100, configured to obtain underwater video stream data, and inversely deduce the image degradation process frame by frame through image restoration technology to obtain an ideal state of the underwater image frame by frame;
[0196] An information extraction module 200, configured to extract frame-by-frame key point information of the human body posture of the frame-by-frame underwater image based on the MediaPipe algorithm;
[0197] An information correction module 300, configured to correct data distortion of the frame-by-frame key point information and perform data enhancement processing through Gaussian noise to obtain frame-by-frame human body posture information;
[0198] A model construction module 400, configured to add a time series to the frame-by-frame human body posture information and establish a human body posture recognition and prediction model based on a long short-term memory network and a support vector regression network;
[0199] An identification and prediction module 500, configured to identify and predict the posture information of an underwater operator through the human body posture recognition and prediction model.
[0200] Compared with the prior art, the underwater human body posture recognition system based on MediaPipe shown in this embodiment restores the image degradation process frame by frame through an image restoration module to obtain an ideal frame-by-frame underwater image, improves the quality of the image, and improves the accuracy of human body posture recognition. The information extraction module optimizes the key point positions according to the frame-by-frame underwater image, further improving the subsequent model recognition ability and accuracy. Then, the information correction module corrects the distortion of the key point information, significantly improving the recognition and prediction accuracy. Finally, the model construction module combines a long short-term memory network and a support vector regression network to establish a human body posture recognition and prediction model to process time series, regression, and classification, enhancing the generalization ability of this model, improving the prediction accuracy and robustness, and reducing overfitting, thereby solving the technical problem in the prior art that due to the absorption and scattering of light by water, the quality of underwater images is poor, resulting in a decrease in the accuracy of human body posture recognition.
[0201] The technical features of each of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0202] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable storage medium for an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can obtain instructions
[0203] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0204] The above-described embodiments merely represent several implementation manners of the present invention. The descriptions are relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. An underwater human body pose recognition method based on MediaPipe, characterized in that, The method includes: Obtain underwater video stream data, inversely infer the image degradation process frame by frame through image restoration technology, and obtain the underwater images frame by frame in an ideal state; Based on the MediaPipe algorithm, extract the key point information frame by frame of the human pose in the underwater images frame by frame; Perform data distortion correction on the key point information frame by frame, and perform data augmentation processing with Gaussian noise to obtain the human pose information frame by frame; Add a time series to the human pose information frame by frame, and establish a human pose recognition and prediction model based on the long short-term memory network and the support vector regression network; Identify and predict the pose information of underwater operators through the human pose recognition and prediction model.
2. The underwater human body posture recognition method based on MediaPipe according to claim 1, characterized in that, The step of obtaining underwater video stream data and inversely inferring the image degradation process frame by frame through image restoration technology to obtain the underwater images frame by frame in an ideal state specifically includes: Collect underwater video stream data through a camera, perform frame-by-frame processing on the underwater video stream data, and obtain the distorted underwater images frame by frame; Calculate the undegraded underwater images from the distorted underwater images frame by frame through image restoration technology. The calculation formula is: , Among them, is the frame-by-frame underwater image in the ideal state, is the frame-by-frame distorted underwater image, is the pixel point of the frame-by-frame distorted underwater image, is various forms of light entering the camera, is the distance from the target to the camera, is the background light, is the transmittance, and it decays exponentially with in an exponential law, is the attenuation coefficient of the reflected light, is the energy after normalization loss.
3. The underwater human body posture recognition method based on MediaPipe according to claim 2, characterized in that, The method further includes: Based on the distorted underwater images frame by frame, construct a color space, obtain the initial values in the color channels of the color space, calculate the average value of the initial values, and correct the initial values of each pixel point in the color channels to obtain the corrected values of the color channels. The calculation formula is as follows: , Among them, , are the correction values of the pixel in the color channels , respectively. , are the initial values of the pixel in the color channels , respectively. , are the average values of the corresponding color channels , in the entire frame-by-frame distorted underwater image. , are parameters, which are the correction levels for evaluating the two initial color channels , respectively. is the influence value of the pixel on the background light position; Perform scattered light processing on the distorted underwater images frame by frame, and construct scattered underwater images with scattered light diffusing outward from the center of the background light. Therefore, the scattered underwater images are represented in the form of spherical coordinates centered on the background light, and are expressed as follows: , Among them, is a scattered underwater image, is the state where scattered light is uniformly scattered at different pre-given angles, is the distance from each pixel point in the scattered light to the background light, is the pixel point in the scattered light; Based on the scattered underwater images, calculate the pixel point mapping values of the scattered light. The calculation formula is as follows: , Among them, is the pixel point mapping value of the scattered light, is centered on is the sphere coordinate range, is half of the distance from the center point of the background light to the dark pixel point, is the maximum distance from the pixel point in the scattered light to the background light, is the scattering coefficient; Calculate the transmittance according to the pixel point mapping values. The calculation formula is as follows: , Among them, is each pixel in the fog line at the distance from each pixel in the fog line at to the background light is the pixel mapping value of the scattered light at the pixel Combine the scattered underwater images with the color channels to calculate the background light. The calculation formula is as follows: , Among them, is the position of the pixel point with the maximum brightness, is a local positive direction channel centered on the pixel point , is a channel point in a local positive direction channel centered on the pixel point , , , are respectively the light intensity components of red, green, and blue of the channel point in the color channel, , , are respectively , , 's different weights, are each local area corresponding to each pixel point in the frame-by-frame distorted underwater image.
4. The underwater human body pose recognition method based on MediaPipe according to claim 1, wherein The step of extracting the key point information frame by frame of the human pose in the underwater images frame by frame based on the MediaPipe algorithm specifically includes: According to the underwater video stream data, determine the key point information to be marked by the underwater operator, and optimize the key point positions of the human pose in the Blazepose model; According to the optimized key point positions, extract the key point information frame by frame of the human pose in the underwater images frame by frame. The key point information frame by frame includes the key point coordinates frame by frame and the corresponding confidence information.
5. The underwater human body posture recognition method based on MediaPipe according to claim 4, characterized in that The step of performing data distortion correction on the key point information frame by frame and performing data augmentation processing with Gaussian noise to obtain the human pose information frame by frame specifically includes: Align the frame-by-frame key point coordinates , with pixel coordinates being . Respectively perform alignment processing on the frame-by-frame key point coordinates and the pixel coordinates to obtain homogeneous key point coordinates , homogeneous pixel coordinates , and perform 3D projection calculation. The calculation formula is as follows: , Among them, is the scale factor, is the rotation and translation parameter of the intrinsic matrix , is the intersection coordinate of the camera and the imaging plane, , is the focal length of the camera relative to the pixel with pixel coordinates , is the parameter describing the tilt of the image axis; Assume that all feature points of the homogeneous key point coordinates are located in a plane with a vertical coordinate of zero, then the calculation formula for 3D projection is expressed as: , Using the homography matrix Replace Obtain , where 、 are the first two columns of the homography matrix, 、 are 's first two columns. Since the rotation matrix is orthogonal, that is: , , the following expression can be obtained: , Solve for the intrinsic matrix , and obtain the numerical values of the parameters that describe the tilt of the image axis; According to the parameters describing the tilt of the image axis, the calibrated radial distortion parameters, and the tangential distortion parameters, perform distortion correction on the key point coordinates frame by frame to obtain the corrected key point coordinates. The calculation formula is as follows: , Among them, and are respectively and the corrected coordinate values, is the radial distortion parameter within the central region of the underwater image frame by frame, is the radial distortion parameter within the edge region of the underwater image frame by frame, is the horizontal tangential distortion parameter, is the vertical distortion parameter; Replace the key point coordinates frame by frame with the corrected key point coordinates to obtain the corrected key point information, and perform geometric transformation on the corrected key point information to simulate the changes in vision and distance; And perform data augmentation processing through Gaussian noise to obtain frame-by-frame human body pose information.
6. The underwater human body posture recognition method based on MediaPipe according to claim 5, wherein The steps of adding a time series to the frame-by-frame human body pose information and establishing a human body pose recognition and prediction model based on a long short-term memory network and a support vector regression network specifically include: Add a time series to the frame-by-frame human body pose information, obtain the change of the coordinates of a certain corrected key point in the frame-by-frame human body pose information on the time series, and construct an action data set, expressed as: , Among them, is the action data set, is the coordinate of the corrected key point at the coordinate in the time series; Use the action data set as the input of the human body pose recognition and prediction model, and output the hidden state through the gate mechanism of the long short-term memory network, expressed as: , Among them, is the hidden state set, is the hidden state at the th time sequence. Select as the final feature vector, where the loss function of the long short-term memory network is: , Among them, is the observation value of the long short-term memory network at the time series, is the average observation value of all time series of the long short-term memory network, = 1, 2,..., ; Input the final feature vector into the support vector regression network for regression calculation to train the support vector regression network for prediction. The calculation formula is as follows: , wherein, is the output result of the support vector regression network, and the loss function of the support vector regression network is: , Among them, is the observed value of the support vector regression network at the time series, is the average observed value of all time series of the support vector regression network, is a preset insensitive threshold.
7. The underwater human body pose recognition method based on MediaPipe according to claim 6, characterized in that, The steps of identifying and predicting the pose information of underwater operators through the human body pose recognition and prediction model specifically include: Obtain the pose information of underwater operators, and calculate the key limb angle, the key limb elevation angle, the key elbow rotation angle, and the key elbow swing angle. The calculation formula of the key limb angle is as follows: , Among them, is the included angle key point to the included angle key point the distance between, is the included angle key point and the included angle key point the key limb included angle between the connections, is from the included angle key point pointing to the included angle key point vector, is in the vector the value of the coordinate dimension, is in the vector the value of the coordinate dimension, is from the included angle key point pointing to the included angle key point vector, is the coordinate dimension, is the total number of dimensions, is the included angle key point coordinates, is the included angle key point coordinates; The calculation formula of the key limb elevation angle is as follows: , Among them, is the key limb lifting angle, and are the abscissa and ordinate of the lifting key point respectively, and are the abscissa and ordinate of the lifting key point respectively, and are the abscissa and ordinate of the lifting key point respectively; The calculation formula of the key elbow rotation angle is as follows: , Among them, is the key elbow joint rotation angle, and are the abscissa and ordinate of the rotation key point respectively, and are the abscissa and ordinate of the rotation key point respectively, and are the abscissa and ordinate of the rotation key point respectively; The calculation formula of the key elbow swing angle is as follows: , Among them, is the key elbow joint swing angle, and are the ordinate and abscissa of the swing key point respectively, and are the ordinate and abscissa of the swing key point respectively, and are the ordinate and abscissa of the swing key point respectively; According to the key limb angle, the key limb elevation angle, the key elbow rotation angle, and the key elbow swing angle, identify and predict the pose actions of underwater operators through the human body pose recognition and prediction model.
8. The underwater human body pose recognition method based on MediaPipe according to claim 7, wherein, After the step of identifying and predicting the pose actions of underwater operators through the human body pose recognition and prediction model according to the key limb angle, the key limb elevation angle, the key elbow rotation angle, and the key elbow swing angle, it further includes: Compare the pose action with the standard pose action, perform deviation calculation to obtain an action deviation value. The calculation formula is as follows: , Among them, is the action deviation value, is the key point of the attitude action coordinates, is the key point of the standard attitude action coordinates, is the total number of key points; According to the action deviation value, combine the dynamic actions and speeds of underwater operators to detect the pose correctness of underwater operators.
9. The underwater human body pose recognition method based on MediaPipe according to claim 8, wherein The steps of detecting the pose correctness of underwater operators according to the action deviation value, combining the dynamic actions and speeds of underwater operators specifically include: According to the action deviation value, combine the dynamic actions and speeds of underwater operators to judge whether the pose of underwater operators is correct; If so, continue to perform the next pose information recognition and prediction; If not, remind the underwater operator to correct the pose information.
10. An underwater human body pose recognition system based on MediaPipe, characterized in that, The system is used to implement the MediaPipe-based underwater human body pose recognition method described in any one of claims 1 to 9. The system includes: An image restoration module for obtaining underwater video stream data and inversely inferring the image degradation process frame by frame through image restoration technology to obtain ideal frame-by-frame underwater images; An information extraction module for extracting frame-by-frame key point information of the human body pose of the frame-by-frame underwater images based on the MediaPipe algorithm; An information correction module for correcting data distortion of the frame-by-frame key point information and performing data augmentation processing through Gaussian noise to obtain frame-by-frame human body pose information; A model construction module, which is used to add a time series to the frame-by-frame human body pose information and establish a human body pose recognition and prediction model based on a long short-term memory network and a support vector regression network; An identification and prediction module, which is used to identify and predict the pose information of underwater operators through the human body pose recognition and prediction model.
Citation Information
Patent Citations
Capture method based on video stream attitude simulation
CN118629085A
Digital human video generation method and device, equipment and medium
CN118842975A