Posture recognition method and equipment based on neural network and medium
By capturing the three-dimensional posture of the human body in real time, adjusting the lighting, using quadrilateral data detection methods and combining multimodal information in the posture recognition system, the problems of lighting changes, poor image quality and single modal information are solved, and the accuracy and robustness of posture recognition are improved.
Patent Information
- Application Number
- CN202510161152.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Neural network-based pose recognition has problems such as lighting changes, poor image quality and single mode information, which affects the recognition accuracy and robustness.
By configuring the posture recognition remote control area server IP address information, enter the posture capture end to capture the three-dimensional posture of the human body in real time, judge the lighting abnormality and adjust the lighting angle; enter the masking end to use quadrilateral data detection means to determine the true recognition posture; enter the multimodal combination end to combine the sound and shape information of the human body to achieve real-time interaction and data association.
The accuracy of pose key position recognition is improved, the problem of poor image quality is reduced, and the accuracy and robustness of pose recognition is enhanced.
Smart Images

Figure CN120107997A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gesture recognition, and in particular to a gesture recognition method, device and medium based on a neural network. Background Art
[0002] Posture recognition based on neural networks is the process of estimating and identifying human posture through deep learning technology. Posture recognition usually involves locating the positions of various joints of the human body to infer the posture of the human body. Deep neural networks play a key role in this process. They can automatically extract human features in images and predict the positions of human joints based on these features. The realization of posture recognition based on neural networks mainly includes steps such as data preprocessing, model training, prediction and post-processing.
[0003] Application No. CN202210303698.9 discloses posture recognition, posture recognition model training method, device, equipment and medium, the method includes: obtaining a first image of the target object; obtaining image position information of a first key point based on the first image, the image position information is used to indicate the position of the first key point in the first image, the first key point is a key point used to describe the two-dimensional posture of the target object; extracting features from the image position information to obtain the target three-dimensional posture parameters. In this way, the three-dimensional posture of the target object is identified based on the posture representation features, the process is relatively simple, which is conducive to saving calculations, thereby improving the efficiency of posture recognition.
[0004] After searching the above patents, it was found that posture recognition based on neural networks still has some shortcomings: 1. Due to the diversity and complexity of human postures, there may be factors such as lighting changes in the image during posture recognition, which makes it impossible to accurately estimate the positions of key points of the human body, affecting the accuracy of posture key position recognition; 2. Poor image quality may be encountered during posture recognition, especially when the key position is blocked and cannot be effectively processed in time; 3. Posture recognition technology usually only focuses on image or video data, and ignores information in other modes, such as sound, text, etc., and cannot guarantee the accuracy and robustness of posture recognition.
[0005] Therefore, a posture recognition method, device and medium based on neural network are proposed to solve the above problems. Summary of the invention
[0006] The main purpose of the present invention is to provide a gesture recognition method, device and medium based on a neural network to solve the problems raised in the above background.
[0007] To achieve the above object, the technical solution adopted by the present invention is: a posture recognition method, device and medium based on neural network, the method comprising the following implementation steps:
[0008] Step 1: Configure the IP address information of the posture recognition remote control area server;
[0009] Step 2: Enter the posture capture terminal, use the camera to capture multiple parameters of the human body's three-dimensional posture in real time, generate an image and determine whether the image lighting is abnormal. If abnormal, track the occlusion area of the key points of the human body, adjust the lighting angle in real time, and determine the optimal lighting parameters;
[0010] Step 3: Enter the mask processing end, obtain the real required recognition posture in the image based on the quadrilateral data detection method, and determine the real shape of the posture according to the real required recognition posture;
[0011] Step 4: Enter the multimodal combination terminal, combine posture recognition with the sound and shape information of the human body, realize real-time interaction and data association of posture recognition, which is suitable for multi-parameter posture recognition and solves the singleness of posture recognition.
[0012] The posture capture terminal includes a posture acquisition module, an illumination parameter module and a key capture module;
[0013] The posture acquisition module includes a standard posture unit and a posture capture unit;
[0014] The standard posture unit is used to set standard parameters of human posture through a processor, the standard parameters include three-dimensional human body, human organ and joint posture parameters, and set the three-dimensional position, organ position, joint position and human posture standard position corresponding to the human posture, and save them in real time through a memory;
[0015] The posture capture unit is used to capture the dynamic changes of human posture in real time through a camera;
[0016] The illumination parameter module includes an image generation unit and an illumination acquisition unit, wherein the image generation unit is used to automatically generate a captured image of a human body posture through an image converter;
[0017] The illumination collection unit includes a multi-region comparison unit and a region compensation unit;
[0018] The multi-region comparison unit is used to compare the illumination intensity of multiple regions of the generated image to determine whether the detection of the standard position of the human body posture is abnormal. The comparison method is as follows:
[0019] According to the different areas of image generation, standard values of different light intensities are set, and the standard values of the corresponding parts of different light intensities are compared with the basic values of the light intensity of the corresponding parts on the image. The comparison formula is as follows:
[0020]
[0021] Among them, Q nIndicates the abnormal value of the light intensity on the image, N 1 Indicates the standard value of the light intensity of the corresponding parts of different areas at the corresponding time, N 2 represents the basic value of the illumination intensity of the corresponding part of the image at the corresponding moment, p represents the illumination energy consumption of the tracked human posture detection parameters, P represents the standard illumination energy consumption of human posture detection, if Q n If Q is equal to 0, it means that there is no abnormal value in the light intensity on the image, and the standard position detection of human posture is normal. n If it is not equal to 0, it means that there is an abnormal value in the light intensity on the image, and the standard position detection of the human body posture is abnormal;
[0022] The area compensation unit is used to report to the system when it is determined that the standard position of the human body posture is abnormal, and according to the calculated Q n Automatically compensate for light intensity parameters during human gesture recognition detection.
[0023] The key capture module includes a position tracking unit, which is used to capture human posture changes in real time through a data tracker combined with light intensity parameters, the changes including three-dimensional position, organ position, joint position and standard position of human posture.
[0024] The mask processing end includes a posture recognition module, a shape determination module and a shape self-tracking module;
[0025] The gesture recognition module includes an image analysis unit and a position positioning unit;
[0026] The image analysis unit is used to analyze the best recognition point of the image according to the generated image through the human body posture shape and the human body part model;
[0027] The position positioning unit is used to determine the specific position of the human body posture through the best recognition point;
[0028] The shape determination module includes a four-side detection unit and a shape determination unit;
[0029] The four-side detection unit is used to detect the posture and shape of the human body in real time, as follows:
[0030] The human body posture quadrilateral annotation box in the generated image is represented by [a, b, c, d], (X1, Y2) and (x1, y1) represent the coordinates of the upper left corner and lower right corner of the quadrilateral box respectively, and the coordinates of the four vertices of the non-quadrilateral annotation box obtained by perspective transformation become:
[0031] a1=(x1,y1);
[0032] b1=(x2,y2);
[0033] c1=(x3,y3);
[0034] d1=(x4,y4);
[0035] Let X min=min{x1,x2,x3,x4};
[0036] Xmax=max{x1,x2,x3,x4};
[0037] Y min = min{y1,y2,y3,y4};
[0038] Y max = max{y1,y2,y3,y4};
[0039] Then {X min, X max, Y min, Y max} represents the four-side annotation box of the transformed image;
[0040] If any point in the image exceeds the image range, calculate the corresponding four-point coordinates after converting the image into a regular quadrilateral frame;
[0041] Calculate the perspective transformation matrix M from the image to the image, obtain the four vertices of the image quadrilateral through the width and height of the image quadrilateral, transform the quadrilateral using the perspective transformation matrix M to obtain the four-point coordinate quadrilateral of the image after normalization, calculate the four-point cross mark of the minimum outer quadrilateral of the intersection of the quadrilateral surrounded by the quadrilateral and the quadrilateral surrounded by the quadrilateral, and the cross mark is the actual quadrilateral area of the human body posture to be identified;
[0042] The shape determination unit is used to perform an inverse transformation M of the perspective transformation matrix M on the four-point cross mark to obtain the last point, and the last point is the true human posture shape in the generated image, thereby obtaining the true shape of human posture recognition.
[0043] The shape self-tracking module includes a shape tracking unit and a shape early warning unit;
[0044] The shape tracking unit is used to track the real-time status of the true shape of human body posture recognition in real time through a camera and a data tracker;
[0045] The shape warning unit is used to calculate the difference between the true shape of human posture recognition and the standard shape. If the difference is equal to 0, it means that the human posture recognition is normal. If the difference is not equal to 0, it means that the human posture recognition is abnormal, and the reporting system issues a voice alarm reminder.
[0046] The multi-modal combining end includes a sound information module, a shape receiving module and a multi-parameter determining module;
[0047] The sound information module includes a sound receiving unit and a voice interaction unit, wherein the sound receiving unit is used to receive the sound generated by the current human posture change in real time through a sound sensor and a microphone;
[0048] The voice interaction unit is used to perform real-time voice interaction on a computer page through a virtual AI character, and the voice interaction uses a microphone and a sound sensor for real-time detection and interactive recognition.
[0049] The shape receiving module includes a shape receiving unit and a data association unit;
[0050] The shape receiving unit is used to receive the abnormal result and sound of the real shape of the human body posture recognition in real time through the data receiver;
[0051] The data association unit is used to perform data association between the sound and the real shape of human posture recognition. The specific association is: the keywords with the highest repetition rate obtained through the virtual AI character voice interaction text are similarly identified with the keywords obtained from the real shape of human posture recognition. The identification formula is as follows:
[0052]
[0053] Among them, A represents the keyword with the highest repetition rate obtained from the virtual AI character voice interaction text, and B represents the keyword obtained from the true shape of human posture recognition. If J(A,B) is equal to 0, it means that there is no correlation between the sound and the true shape of human posture recognition. If J(A,B) is not equal to 0, it means that there is a correlation between the sound and the true shape of human posture recognition.
[0054] The multi-parameter determination module includes a parameter comparison unit and a model matching unit;
[0055] The parameter comparison unit is used to receive the real shape parameters and voice interaction information of human posture recognition in real time through the data receiver, and perform difference calculation between the keyword data parameter with the highest repetition rate obtained from the virtual AI character voice interaction text and the keyword data parameter with the highest repetition rate obtained from the virtual AI character voice interaction text. If the difference is equal to 0, it indicates that the human posture recognition is correct, and if the difference is not equal to 0, it indicates that the human posture recognition is wrong;
[0056] The model matching unit is used to automatically generate a human posture recognition model based on a neural network according to the human posture recognition result, and to add or delete corresponding posture recognition data parameters according to different postures.
[0057] An electronic device comprises: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory.
[0058] A computer storage medium stores a computer program.
[0059] The present invention has the following beneficial effects:
[0060] 1. In the present invention, by setting a posture capture end, during posture recognition based on a neural network, by setting standard posture parameters, and by using a camera to capture multiple parameters of a human body's three-dimensional posture in real time, an image is generated and image information is compared in real time in multiple regions, and it is determined whether the illumination of the image generated after posture capture is abnormal. When the illumination is abnormal, the occluded area of the key points of the human body is tracked, and the illumination angle is adjusted in real time to avoid the image affecting the accuracy of the key point detection of the human body according to the illumination change during posture recognition, so that the optimal position of the key points of the human body can be accurately estimated, and the accuracy of posture key position recognition is improved;
[0061] 2. In the present invention, by setting a mask processing end, during the posture recognition based on the neural network, the posture that is really needed in the image can be determined in real time by means of a detection method based on quadrilateral data, and the true shape of the posture can be determined according to the really needed posture, thereby reducing the problem of poor image quality during posture recognition. In particular, when the key position is blocked or covered, the true shape of the posture to be recognized can be obtained timely and accurately by means of a four-side detection method;
[0062] 3. In the present invention, by setting a multimodal combination terminal, in the posture recognition based on the neural network, by combining the posture recognition with the sound and shape information of the human body, the human body sound is monitored in real time during the posture recognition process to realize voice interaction, and data association is realized after receiving the true shape of the posture. By means of multi-parameter comparison, the multi-parameter stability of the posture recognition is fused and detected, and the fused data is automatically matched to the posture recognition model in real time after the multi-parameter fusion, so that the posture recognition technology can combine the sound and shape to further increase the accuracy of the posture recognition while paying attention to the image or video data, thereby ensuring the accuracy and robustness of the posture recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 The present invention is a flowchart of a method, device and medium for gesture recognition based on a neural network;
[0064] Figure 2 It is a schematic diagram of the structure of a posture recognition method, device and medium posture capture terminal based on a neural network of the present invention;
[0065] Figure 3 It is a schematic diagram of the structure of a posture recognition method, device and medium mask processing end based on a neural network of the present invention;
[0066] Figure 4 The present invention is a schematic diagram of the architecture of a multi-modal combination terminal of a posture recognition method, device and medium based on a neural network. DETAILED DESCRIPTION
[0067] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.
[0068] Embodiment 1
[0069] Please refer to Figure 1 to Figure 2 As shown: A gesture recognition method, device and medium based on a neural network, the method includes the following implementation steps:
[0070] Step 1: Configure the IP address information of the posture recognition remote control area server;
[0071] Step 2: Enter the posture capture terminal, use the camera to capture multiple parameters of the human body's three-dimensional posture in real time, generate an image and determine whether the image lighting is abnormal. If abnormal, track the occlusion area of the key points of the human body, adjust the lighting angle in real time, and determine the optimal lighting parameters;
[0072] Step 3: Enter the mask processing end, obtain the real required recognition posture in the picture based on the quadrilateral data detection method, determine the real shape of the posture according to the real required recognition posture, and the mask processing end is the shape recognition processing area of the key point, that is, the real posture to be recognized;
[0073] Step 4: Enter the multimodal combination terminal, combine posture recognition with the sound and shape information of the human body, realize real-time interaction and data association of posture recognition, which is suitable for multi-parameter posture recognition and solves the singleness of posture recognition.
[0074] The gesture capture end includes a gesture acquisition module, a lighting parameter module and a key capture module;
[0075] The posture acquisition module includes a standard posture unit and a posture capture unit;
[0076] The standard posture unit is used to set the standard parameters of human posture through the processor. The standard parameters include the three-dimensional, organ and joint posture parameters of the human body, and set the three-dimensional position, organ position, joint position and standard position of the human body posture corresponding to the human body posture, and save them in real time through the memory;
[0077] The posture capture unit is used to capture the dynamic changes of human posture in real time through a camera;
[0078] The illumination parameter module includes an image generation unit and an illumination acquisition unit, wherein the image generation unit is used to automatically generate a captured image of a human body posture through an image converter;
[0079] The illumination acquisition unit includes a multi-region comparison unit and a region compensation unit;
[0080] The multi-region comparison unit is used to compare the illumination intensity of multiple regions of the generated image to determine whether the detection of the standard position of the human body posture is abnormal. The comparison method is as follows:
[0081] According to the different areas of image generation, standard values of different light intensities are set, and the standard values of the corresponding parts of different light intensities are compared with the basic values of the light intensity of the corresponding parts on the image. The comparison formula is as follows:
[0082]
[0083] Among them, Q n Indicates the abnormal value of the light intensity on the image, N 1 Indicates the standard value of the light intensity of the corresponding parts of different areas at the corresponding time, N 2 represents the basic value of the illumination intensity of the corresponding part of the image at the corresponding moment, p represents the illumination energy consumption of the tracked human posture detection parameters, P represents the standard illumination energy consumption of human posture detection, if Q n If Q is equal to 0, it means that there is no abnormal value in the light intensity on the image, and the standard position detection of human posture is normal. n If it is not equal to 0, it means that there is an abnormal value in the light intensity on the image, and the standard position detection of the human body posture is abnormal;
[0084] The area compensation unit is used to report to the system when the standard position detection of human posture is abnormal. n Automatically compensate for the light intensity parameters in the process of human posture recognition and detection, capture multiple parameters of the human body's three-dimensional posture in real time through the camera, generate images and compare image information in real time in multiple areas, and determine whether the lighting of the image generated after posture capture is abnormal. If the light intensity is normal, the optimal position for human key point posture recognition can be accurately estimated.
[0085] The key capture module includes a position tracking unit, which is used to capture changes in human posture in real time through a data tracker combined with light intensity parameters. The changes include three-dimensional position, organ position, joint position and standard position of human posture. When the lighting is abnormal, it tracks the occluded area of the human body's key points and adjusts the lighting angle in real time to avoid the image affecting the accuracy of human body key point detection due to changes in lighting during posture recognition, and can accurately estimate the optimal position of the human body's key points.
[0086] Embodiment 2
[0087] Please refer to Figure 3 As shown: Based on the first embodiment, the mask processing end includes a posture recognition module, a shape determination module and a shape self-tracking module;
[0088] The gesture recognition module includes an image analysis unit and a position positioning unit;
[0089] The image analysis unit is used to analyze the best recognition point of the image according to the generated image through the human body posture shape and the human body part model;
[0090] The position positioning unit is used to determine the specific position of the human body posture through the best recognition point;
[0091] The shape determination module includes a four-side detection unit and a shape determination unit;
[0092] The four-side detection unit is used to detect the human body posture shape in real time, as follows:
[0093] The human body posture quadrilateral annotation box in the generated image is represented by [a, b, c, d], (X1, Y2) and (x1, y1) represent the coordinates of the upper left corner and lower right corner of the quadrilateral box respectively, and the coordinates of the four vertices of the non-quadrilateral annotation box obtained by perspective transformation become:
[0094] a1=(x1,y1);
[0095] b1=(x2,y2);
[0096] c1=(x3,y3);
[0097] d1=(x4,y4);
[0098] Let X min=min{x1,x2,x3,x4};
[0099] Xmax=max{x1,x2,x3,x4};
[0100] Y min = min{y1,y2,y3,y4};
[0101] Y max = max{y1,y2,y3,y4};
[0102] Then {X min, X max, Y min, Y max} represents the four-side annotation box of the transformed image;
[0103] If any point in the image exceeds the image range, calculate the corresponding four-point coordinates after converting the image into a regular quadrilateral frame;
[0104] Calculate the perspective transformation matrix M from the image to the image, obtain the four vertices of the image quadrilateral through the width and height of the image quadrilateral, transform the quadrilateral using the perspective transformation matrix M to obtain the four-point coordinate quadrilateral of the image after normalization, calculate the four-point cross mark of the minimum outer quadrilateral of the intersection of the quadrilateral surrounded by the quadrilateral and the quadrilateral surrounded by the quadrilateral, and the cross mark is the actual quadrilateral area of the human body posture to be identified;
[0105] The shape determination unit is used to perform an inverse transformation M of the perspective transformation matrix M on the four-point cross mark to obtain the last point. The last point is the true human posture shape in the generated image, and the true shape of the human posture recognition is obtained. Through the quadrilateral data detection method, the truly required recognition posture in the image is determined in real time, and the true shape of the posture is determined based on the truly required recognition posture.
[0106] The shape self-tracking module includes a shape tracking unit and a shape early warning unit;
[0107] The shape tracking unit is used to track the real-time status of the true shape of human posture recognition in real time through a camera and a data tracker;
[0108] The shape warning unit is used to calculate the difference between the true shape of human posture recognition and the standard shape. If the difference is equal to 0, it means that the human posture recognition is normal. If the difference is not equal to 0, it means that the human posture recognition is abnormal. The reporting system will issue a voice alarm reminder. When the key position is blocked or covered, the true shape of the posture to be recognized can be obtained in a timely and accurate manner through four-side detection.
[0109] Embodiment 3
[0110] Please refer to Figure 4 As shown: Based on the first embodiment, the multi-modal combination end includes a sound information module, a shape receiving module and a multi-parameter determination module;
[0111] The sound information module includes a sound receiving unit and a voice interaction unit. The sound receiving unit is used to receive the sound generated by the current human posture change in real time through a sound sensor and a microphone.
[0112] The voice interaction unit is used to conduct real-time voice interaction on a computer page through a virtual AI character. The voice interaction uses a microphone and sound sensor for real-time detection and interaction recognition.
[0113] The shape receiving module includes a shape receiving unit and a data association unit;
[0114] The shape receiving unit is used to receive the abnormal result and sound of the real shape of the human body posture recognition in real time through the data receiver;
[0115] The data association unit is used to associate the sound with the real shape of human posture recognition. The specific association is: the keywords with the highest repetition rate obtained through the virtual AI character voice interaction text are similarly identified with the keywords obtained from the real shape of human posture recognition. The identification formula is as follows:
[0116]
[0117] Among them, A represents the keyword with the highest repetition rate obtained from the virtual AI character voice interaction text, and B represents the keyword obtained from the true shape of human posture recognition. If J(A,B) is equal to 0, it means that there is no correlation between the sound and the true shape of human posture recognition. If J(A,B) is not equal to 0, it means that there is a correlation between the sound and the true shape of human posture recognition. During the posture recognition process, the human body sound is monitored in real time to realize voice interaction, and data association is realized after receiving the true shape of the posture. Through the method of multi-parameter comparison, the multi-parameter stability of posture recognition is integrated and detected.
[0118] The multi-parameter determination module includes a parameter comparison unit and a model matching unit;
[0119] The parameter comparison unit is used to receive the real shape parameters and voice interaction information of human posture recognition in real time through the data receiver, and perform difference calculation between the keyword data parameter with the highest repetition rate obtained from the virtual AI character voice interaction text and the keyword data parameter with the highest repetition rate obtained from the virtual AI character voice interaction text. If the difference is equal to 0, it indicates that the human posture recognition is correct, and if the difference is not equal to 0, it indicates that the human posture recognition is wrong.
[0120] The model matching unit is used to automatically generate a human posture recognition model based on a neural network according to the human posture recognition results, and to add or delete corresponding posture recognition data parameters according to different postures. After multi-parameter fusion, the fused data is automatically matched to the posture recognition model in real time, so that the posture recognition technology can combine sound and shape to further increase the accuracy of posture recognition while paying attention to image or video data.
[0121] In the present invention, a posture recognition method, device and medium based on neural network are provided. First, the IP address information of the posture recognition remote control area server is configured; the posture capture end is entered, multiple parameters of the three-dimensional posture of the human body are captured in real time by a camera, an image is generated and it is determined whether the image illumination is abnormal, and when it is abnormal, the occlusion area of the key points of the human body is tracked, the illumination angle is adjusted in real time, and the optimal illumination parameters are determined. In the posture recognition based on neural network, standard posture parameters are set, and multiple parameters of the three-dimensional posture of the human body are captured in real time by a camera, an image is generated, and image information is compared in real time in multiple areas, it is determined whether the illumination of the image generated after the posture capture is abnormal, and when the illumination is abnormal, the occlusion area of the key points of the human body is tracked, and the illumination angle is adjusted in real time, so as to avoid the image affecting the accuracy of the key point detection of the human body according to the illumination change during the posture recognition, and the optimal position of the key point of the human body can be accurately estimated, so as to improve the accuracy of the posture key position recognition; the mask processing end is entered, the truly required recognition posture in the picture is obtained based on the quadrilateral data detection method, and the true shape of the posture is determined according to the truly required recognition posture, and the true shape of the posture is determined based on the neural network. When the network recognizes the posture, the real required recognition posture in the image is determined in real time based on the quadrilateral data detection method, and the real shape of the posture is determined according to the real required recognition posture, so as to reduce the problem of poor image quality in posture recognition. In particular, when the key position is blocked or covered, the real shape of the posture to be recognized can be obtained timely and accurately through the quadrilateral detection method; entering the multimodal combination end, the posture recognition is combined with the sound and shape information of the human body to realize the real-time interaction and data association of the posture recognition, which is suitable for multi-parameter posture recognition and solves the singleness of posture recognition. By combining the posture recognition with the sound and shape information of the human body, the human body sound is monitored in real time during the posture recognition process to realize voice interaction, and data association is realized after receiving the real shape of the posture. By means of multi-parameter comparison, the multi-parameter stability of posture recognition is fused and detected, and the fused data is automatically matched to the posture recognition model in real time after multi-parameter fusion, so that the posture recognition technology can combine the sound and shape to further increase the accuracy of posture recognition while paying attention to the image or video data, and ensure the accuracy and robustness of posture recognition.
[0122] The above shows and describes the basic principles and main features of the present invention and the advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A gesture recognition method based on a neural network, characterized in that: The method comprises the following implementation steps: Step 1: Configure the IP address information of the posture recognition remote control area server; Step 2: Enter the posture capture terminal, use the camera to capture multiple parameters of the human body's three-dimensional posture in real time, generate an image and determine whether the image lighting is abnormal. If abnormal, track the occlusion area of the key points of the human body, adjust the lighting angle in real time, and determine the optimal lighting parameters; Step 3: Enter the mask processing end, obtain the real required recognition posture in the image based on the quadrilateral data detection method, and determine the real shape of the posture according to the real required recognition posture; Step 4: Enter the multimodal combination terminal, combine posture recognition with the sound and shape information of the human body, realize real-time interaction and data association of posture recognition, which is suitable for multi-parameter posture recognition and solves the singleness of posture recognition.
2. The method according to claim 1, characterized in that: The posture capture terminal includes a posture acquisition module, an illumination parameter module and a key capture module; The posture acquisition module includes a standard posture unit and a posture capture unit; The standard posture unit is used to set standard parameters of human posture through a processor, the standard parameters include three-dimensional human body, human organ and joint posture parameters, and set the three-dimensional position, organ position, joint position and human posture standard position corresponding to the human posture, and save them in real time through a memory; The posture capture unit is used to capture the dynamic changes of human posture in real time through a camera; The illumination parameter module includes an image generation unit and an illumination acquisition unit, wherein the image generation unit is used to automatically generate a captured image of a human body posture through an image converter; The illumination collection unit includes a multi-region comparison unit and a region compensation unit; The multi-region comparison unit is used to compare the illumination intensity of multiple regions of the generated image to determine whether the detection of the standard position of the human body posture is abnormal. The comparison method is as follows: According to the different areas of image generation, standard values of different light intensities are set, and the standard values of the corresponding parts of different light intensities are compared with the basic values of the light intensity of the corresponding parts on the image. The comparison formula is as follows: Among them, Q n represents the abnormal value of the light intensity on the image, N1 represents the standard value of the light intensity of the corresponding parts of different regions at the corresponding time, N2 represents the basic value of the light intensity of the corresponding parts on the image at the corresponding time, p represents the light energy consumption of the tracked human posture parameters, P represents the standard light energy consumption of human posture detection, if Q n If Q is equal to 0, it means that there is no abnormal value in the light intensity on the image, and the standard position detection of human posture is normal. n If it is not equal to 0, it means that there is an abnormal value in the light intensity on the image, and the standard position detection of the human body posture is abnormal; The area compensation unit is used to report to the system when it is determined that the standard position of the human body posture is abnormal, and according to the calculated Q n Automatically compensate for light intensity parameters during human gesture recognition detection.
3. The method according to claim 2, characterized in that: The key capture module includes a position tracking unit, which is used to capture human posture changes in real time through a data tracker combined with light intensity parameters, the changes including three-dimensional position, organ position, joint position and standard position of human posture.
4. The method according to claim 3, characterized in that: The mask processing end includes a posture recognition module, a shape determination module and a shape self-tracking module; The gesture recognition module includes an image analysis unit and a position positioning unit; The image analysis unit is used to analyze the best recognition point of the image according to the generated image through the human body posture shape and the human body part model; The position positioning unit is used to determine the specific position of the human body posture through the best recognition point; The shape determination module includes a four-side detection unit and a shape determination unit; The four-side detection unit is used to detect the posture and shape of the human body in real time, as follows: The human body posture quadrilateral annotation box in the generated image is represented by [a, b, c, d], (X1, Y2) and (x1, y1) represent the coordinates of the upper left corner and lower right corner of the quadrilateral box respectively, and the coordinates of the four vertices of the non-quadrilateral annotation box obtained by perspective transformation become: a1=(x1,y1); b1=(x2,y2); c1=(x3,y3); d1=(x4,y4); Let Xmin=min{x1,x2,x3,x4}; Xmax=max{x1,x2,x3,x4}; Ymin=min{y1,y2,y3,y4}; Ymax=max{y1,y2,y3,y4}; Then {Xmin, Xmax, Ymin, Ymax} represents the four-side annotation box of the transformed image; If any point in the image exceeds the image range, calculate the corresponding four-point coordinates after converting the image into a regular quadrilateral frame; Calculate the perspective transformation matrix M from the image to the image, obtain the four vertices of the image quadrilateral through the width and height of the image quadrilateral, transform the quadrilateral using the perspective transformation matrix M to obtain the four-point coordinate quadrilateral of the image after normalization, calculate the four-point cross mark of the minimum outer quadrilateral of the intersection of the quadrilateral surrounded by the quadrilateral and the quadrilateral surrounded by the quadrilateral, and the cross mark is the actual quadrilateral area of the human body posture to be identified; The shape determination unit is used to perform an inverse transformation M of the perspective transformation matrix M on the four-point cross mark to obtain the last point, and the last point is the true human posture shape in the generated image, thereby obtaining the true shape of human posture recognition.
5. The method according to claim 4, characterized in that: The shape self-tracking module includes a shape tracking unit and a shape early warning unit; The shape tracking unit is used to track the real-time status of the true shape of human body posture recognition in real time through a camera and a data tracker; The shape warning unit is used to calculate the difference between the true shape of human posture recognition and the standard shape. If the difference is equal to 0, it means that the human posture recognition is normal. If the difference is not equal to 0, it means that the human posture recognition is abnormal, and the reporting system issues a voice alarm reminder.
6. The method according to claim 1, characterized in that: The multi-modal combining end includes a sound information module, a shape receiving module and a multi-parameter determining module; The sound information module includes a sound receiving unit and a voice interaction unit, wherein the sound receiving unit is used to receive the sound generated by the current human posture change in real time through a sound sensor and a microphone; The voice interaction unit is used to perform real-time voice interaction on a computer page through a virtual AI character, and the voice interaction uses a microphone and a sound sensor for real-time detection and interactive recognition.
7. The method according to claim 6, characterized in that: The shape receiving module includes a shape receiving unit and a data association unit; The shape receiving unit is used to receive the abnormal result and sound of the real shape of the human body posture recognition in real time through the data receiver; The data association unit is used to perform data association between the sound and the real shape of human posture recognition. The specific association is: the keywords with the highest repetition rate obtained through the virtual AI character voice interaction text are similarly identified with the keywords obtained from the real shape of human posture recognition. The identification formula is as follows: Among them, A represents the keyword with the highest repetition rate obtained from the virtual AI character voice interaction text, and B represents the keyword obtained from the true shape of human posture recognition. If J(A,B) is equal to 0, it means that there is no correlation between the sound and the true shape of human posture recognition. If J(A,B) is not equal to 0, it means that there is a correlation between the sound and the true shape of human posture recognition.
8. The method according to claim 7, characterized in that: The multi-parameter determination module includes a parameter comparison unit and a model matching unit; The parameter comparison unit is used to receive the real shape parameters and voice interaction information of human posture recognition in real time through the data receiver, and perform difference calculation between the keyword data parameter with the highest repetition rate obtained from the virtual AI character voice interaction text and the keyword data parameter with the highest repetition rate obtained from the virtual AI character voice interaction text. If the difference is equal to 0, it indicates that the human posture recognition is correct, and if the difference is not equal to 0, it indicates that the human posture recognition is wrong; The model matching unit is used to automatically generate a human posture recognition model based on a neural network according to the human posture recognition result, and to add or delete corresponding posture recognition data parameters according to different postures.
9. An electronic device, characterized in that: include: Processor and memory; The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes a gesture recognition method based on a neural network as described in any one of claims 1 to 8.
10. A computer storage medium storing a computer program, characterized in that: The computer program is executed by a processor to implement a gesture recognition method based on a neural network as claimed in any one of claims 1 to 8.
Citation Information
Patent Citations
Posture recognition method and device, posture recognition model training method and device, equipment and medium
CN116863460A
Cited By
Human body posture recognition method and system based on double-attention structured position coding
CN121600558A
Human pose recognition method and system based on double-attention structured position encoding
CN121600558B