An image processing and recognition system and method based on a personnel training system
Patent Information
- Application Number
- CN202611080087.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明提供了一种基于人员训练系统的图像处理识别系统及方法,用于解决训练场景中灯影镜像干扰导致的人像分割不准、复杂光照与背景条件下人体骨架与关节点识别不稳定、现有训练评估系统缺乏成像条件自适应调节和闭环动作评价的问题
本发明,通过引入可旋转线偏振片及灰度-偏振角曲线分析,自动区分实体角与镜像角,并在两种偏振角单帧图像之间逐像素计算灰度差生成二值掩膜,实现实体区域与灯影镜像在成像物理层面的显式分离,相比依赖颜色差异、边缘特征的传统图像分割方式,能够在固定补光下针对强反射光斑、光滑地面反光、训练器材高亮表面等场景稳定抑制伪影,对受训者真实轮廓保留度高,显著提高后续人像分割和骨架提取的信噪比与可靠性。
Smart Images

Figure CN122821074A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to an image processing and recognition system and method based on a personnel training system. Background Technology
[0002] In training scenarios, camera-based human motion capture and recognition has become a crucial component of personnel training systems. Typical systems use fixed or semi-fixed cameras to capture image sequences of trainees performing specified actions. They then utilize image segmentation, pose estimation, and skeleton tracking algorithms to extract key joints and combine this with standard action templates to quantitatively evaluate the quality of action execution. These systems can objectively score repetitive training movements without increasing the burden on wearable devices, helping to reduce subjective bias in human evaluation and providing data support for subsequent intelligent training plan recommendations. The demand for these systems continues to grow.
[0003] In existing technologies, most personnel training image recognition schemes directly perform foreground segmentation and pose estimation on ordinary visible light images. They lack specific handling for strong reflections caused by fixed supplementary lighting in the training area, bright ground reflections, and specular reflections from equipment surfaces. This easily leads to the formation of light shadows and mirror images in the images, causing human contours to be truncated or adhered to by artifacts. Skeleton extraction results may also exhibit joint drift and limb loss. Furthermore, common systems typically employ a one-time imaging condition calibration strategy. Even if the action recognition effect significantly decreases during subsequent recognition processes, they do not actively adjust imaging conditions such as polarization state and exposure parameters. Action similarity evaluation metrics are mostly based on simple distance threshold judgments, lacking a closed-loop optimization mechanism linked to the imaging physical process. This makes it difficult to maintain stable and reliable recognition accuracy under complex lighting environments and significant differences in trainees. Summary of the Invention
[0004] This invention provides an image processing and recognition system and method based on a personnel training system, which solves the problems of inaccurate human image segmentation caused by light and shadow mirror interference in training scenarios, unstable recognition of human skeleton and joint points under complex lighting and background conditions, and the lack of adaptive adjustment of imaging conditions and closed-loop motion evaluation in existing training evaluation systems.
[0005] An image processing and recognition method based on a personnel training system includes the following steps: S1: With the fixed supplementary lighting on in the training area, a rotatable linear polarizer is simultaneously installed on the front of the camera lens to capture a single frame image of the trainee completing a specified action. The single frame image is subjected to directional grayscale statistics. The grayscale-polarization angle curve is obtained by using the polarization angle as the scanning variable. The polarization angles corresponding to the two local peaks in the curve are recorded as the entity angle and the mirror angle, respectively. A binary mask is generated based on the grayscale difference between the two angles. The white area of the mask represents the entity and the black area represents the light shadow mirror. S2: Multiply the binary mask with the original image pixel by pixel to obtain the entity image with the light shadows and reflections filtered out. Perform human figure segmentation on the entity image to obtain the entity outline, and extract the skeleton of the entity outline to obtain a two-dimensional coordinate sequence containing only the real human joints. S3: Match the two-dimensional coordinate sequence with the standard action template to calculate the action similarity. If the action similarity is lower than a set threshold, feed back the current polarization angle error to S1, drive the micro motor to rotate the polarizer to the new solid angle, re-acquire and update the binary mask, forming a closed loop of "polarization orientation - mirror rejection - action recognition" until continuous The process ends when the frame similarity meets the standard.
[0006] Optionally, S1 includes: S11: With the fixed supplementary lighting on in the training area, a rotatable linear polarizer is synchronously installed on the front of the camera lens. The polarizer is controlled to rotate at equal angular intervals. At each polarization angle, a single-frame image of the trainee completing the specified action is acquired. The directional grayscale statistics of each single-frame image are then performed to obtain the single-frame image and its directional grayscale statistics value corresponding to each polarization angle. S12: Fit the polarization angles obtained in S11 and their corresponding directional gray-scale statistical values as data points to construct a gray-scale-polarization angle curve with polarization angle as the abscissa and directional gray-scale statistical values as the ordinate. S13: Perform peak detection on the generated grayscale-polarization angle curve, identify two local peak points in the curve, and record the polarization angles corresponding to these two local peak points as the solid angle and the mirror angle, respectively. S14: Based on the determined entity angle and mirror angle, calculate the difference in grayscale value of each pixel in the original single-frame image at these two polarization angles. Based on the comparison result of the grayscale value difference with the preset threshold, generate a binary mask, where the white area of the binary mask represents the entity and the black area represents the light shadow mirror.
[0007] Optionally, the directional grayscale statistics for each single-frame image specifically involves: for each single-frame image acquired at a specific polarization angle, calculating the average grayscale value of all pixels in the single-frame image, and using this average grayscale value as the directional grayscale statistics value corresponding to this polarization angle.
[0008] Optionally, the peak detection of the gray-scale-polarization angle curve to identify two local peak points in the curve specifically involves: traversing every data point on the gray-scale-polarization angle curve, determining that a data point is a local peak point if its directional gray-scale statistical value is simultaneously greater than the directional gray-scale statistical values of the preceding and following data points, and selecting the two points with the largest directional gray-scale statistical values from all local peak points as target local peak points.
[0009] Optionally, S2 includes: S21: Perform a pixel-by-pixel multiplication operation between the binary mask generated in S1 and the corresponding original image, so that the original image pixels corresponding to the white areas in the binary mask are retained and the original image pixels corresponding to the black areas are suppressed, thereby obtaining a physical image with the light shadow image filtered out. S22: Perform semantic segmentation-based human image segmentation processing on the entity image obtained in S21, classify the pixels in the image into human body and background, extract the pixel set of the human body region, and the pixel set of the human body region constitutes the entity outline of the entity image. S23: Apply a skeletonization algorithm to the extracted entity contour, iteratively remove contour boundary pixels until its topology converges into a line structure with a single pixel width. This line structure is the human skeleton. Then locate the key joints on the human skeleton and output the two-dimensional coordinate sequence of these key joints.
[0010] Optionally, the step of performing pixel-by-pixel multiplication between the binary mask and the corresponding original image specifically involves multiplying the pixel value of each pixel in the binary mask with the gray value of the corresponding pixel in the original image, wherein the pixel value of the white area of the binary mask is 1 and the pixel value of the black area is 0.
[0011] Optionally, the semantic segmentation-based human image segmentation process specifically involves: inputting the entity image into a pre-trained semantic segmentation neural network model; the semantic segmentation neural network model outputs a segmentation map with the same size as the input image; the category label of each pixel in the segmentation map identifies whether it belongs to the human body or the background; and all pixels labeled as the human body constitute the entity outline.
[0012] Optionally, S3 includes: S31: Compare the spatial position of the two-dimensional coordinate sequence with the pre-stored standard action template, calculate the overall matching degree between the two sets of coordinates, and obtain a quantified action similarity. S32: Compare the action similarity with a preset threshold. If the action similarity is lower than the threshold, calculate a polarization angle error to correct the imaging conditions based on the difference between the current action similarity and the threshold. S33: The calculated polarization angle error is used as a control signal and fed back to the micro motor that drives the rotatable linear polarizer. The micro motor rotates the rotatable linear polarizer to a new polarization angle according to the polarization angle error. Based on this new polarization angle, S1 is triggered again to acquire a new single-frame image and update the binary mask. Then, the processing flow of S2 and S31-S32 is repeated. S34: Continuously execute the closed-loop feedback and reprocessing process of S33, and monitor the similarity of actions obtained each time. If the calculated similarity of actions is not lower than the set threshold, the current action is deemed to have met the standard and the process ends.
[0013] Optionally, the step of comparing the spatial position of the two-dimensional coordinate sequence with the pre-stored standard action template and calculating the overall matching degree between the two sets of coordinates to obtain a quantified action similarity specifically involves: calculating the Euclidean distance between each joint point in the two-dimensional coordinate sequence and the corresponding joint point in the standard action template, summing or averaging the Euclidean distances of all joint points, and converting the sum or average of the distances into an action similarity value between 0 and 1 through a preset mapping function.
[0014] An image processing and recognition system based on a personnel training system, used to implement the aforementioned image processing and recognition method based on a personnel training system, includes the following modules: The polarization imaging control module is used to control the camera and its front-end rotatable linear polarizer and micro motor when the fixed supplementary lighting in the training area is on, so as to acquire single-frame images of the trainee performing a specified action at different polarization angles. The mirror separation module is used to receive a single-frame image acquired by the polarization imaging control module, perform directional grayscale statistics and construct a grayscale-polarization angle curve, identify the entity angle and the mirror angle, and generate a binary mask for separating the lamp shadow mirror. The entity processing module is used to multiply the binary mask generated by the mirror separation module with the corresponding original image to obtain an entity image, and to perform human portrait segmentation and skeleton extraction on the entity image, outputting a two-dimensional coordinate sequence of real human joints; The evaluation and closed-loop control module is used to compare the two-dimensional coordinate sequence output by the entity processing module with the standard action template to calculate the action similarity. When the similarity is lower than a set threshold, the polarization angle error is calculated, and a control command is sent to the polarization imaging control module to drive the polarizer to rotate to a new angle, triggering a new round of processing until the action meets the standard.
[0015] The beneficial effects of this invention are: This invention, by introducing a rotatable linear polarizer and grayscale-polarization angle curve analysis, automatically distinguishes between entity angles and mirror angles. It then calculates the grayscale difference pixel by pixel between single-frame images of the two polarization angles to generate a binary mask, achieving explicit separation of entity regions and light shadow mirrors at the imaging physical level. Compared with traditional image segmentation methods that rely on color differences and edge features, this invention can stably suppress artifacts in scenarios such as strong reflective spots, smooth ground reflections, and bright surfaces of training equipment under fixed lighting conditions. It also preserves the true contours of trainees with high accuracy, significantly improving the signal-to-noise ratio and reliability of subsequent portrait segmentation and skeleton extraction.
[0016] 2. This invention inputs the entity image screened by a binary mask into a semantic segmentation neural network based on an encoder-decoder structure. It obtains well-connected entity contours through pixel-level human / background classification combined with morphological closing operations. Then, it generates a single-pixel-width human skeleton using an iterative morphological thinning algorithm. Finally, a pose estimation joint detection model outputs a two-dimensional coordinate sequence of key joints. Compared to pose extraction processes that rely solely on edge detection or simple threshold segmentation followed by geometric fitting, this solution can stably output a structurally complete skeleton and joint coordinates even in complex training environments, with diverse clothing colors and rich background textures. It significantly reduces the impact of background interference on pose recognition accuracy and provides a unified and standardized coordinate representation for action quantification evaluation.
[0017] 3. This invention constructs a quantifiable index for motion similarity based on the normalized Euclidean distance between the two-dimensional coordinate sequence of key joints and the standard motion template. The difference between the similarity and a set threshold is mapped to a polarization angle error via proportional gain. This error drives a stepper motor to dynamically adjust the polarization angle of a rotatable linear polarizer, forming a closed-loop control mechanism of "polarization orientation - image rejection - motion recognition." Compared to training systems that rely on manual lighting adjustments, fixed exposure, or single-shot imaging calibration, this solution can adaptively optimize imaging conditions for different trainees' body types, clothing materials, and variations in training site lighting. The similarity criterion ensures stable and reliable action recognition results, and provides a recognition failure report when convergence is not achieved, which helps coaches and system maintenance personnel to quickly locate problems and greatly improves the automation and robustness of the training evaluation process. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the system flow according to an embodiment of the present invention. Detailed Implementation
[0020] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. It should also be noted that, to make the embodiments more comprehensive, the following embodiments are the best and preferred embodiments, and those skilled in the art can use other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.
[0021] It should be noted that the use of terms such as "an embodiment," "an embodiment," "an exemplary embodiment," and "some embodiments" in the specification indicates that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, when a specific feature, structure, or characteristic is described in connection with an embodiment, implementing such a feature, structure, or characteristic in conjunction with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the art.
[0022] Generally, terms can be understood at least partly from their use in context. For example, depending at least partly on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or a combination of features, structures, or characteristics in a plural sense. Additionally, the term "based on" can be understood not necessarily to convey an exclusive set of factors, but rather, alternatively, depending at least partly on the context, to allow for the presence of other factors that are not necessarily explicitly described.
[0023] like Figure 1 As shown, an image processing and recognition method based on a personnel training system includes the following steps: S1: With the fixed supplementary lighting on in the training area, a rotatable linear polarizer is simultaneously installed on the front of the camera lens to capture single-frame images of the trainee performing a specified action. Oriented grayscale statistics are performed on the single-frame images, and a grayscale-polarization angle curve is obtained using the polarization angle as the scanning variable. The polarization angles corresponding to the two local peaks in the curve are recorded as the entity angle and the mirror angle, respectively. A binary mask is generated based on the grayscale difference between the two angles. The white area of the mask represents the entity, and the black area represents the light shadow mirror image. Specifically: S11: With the fixed supplementary lighting on in the training area, a rotatable linear polarizer is synchronously installed on the front of the camera lens. The polarizer is controlled to rotate at equal angular intervals, and single-frame images of the trainee performing a specified action are acquired at each polarization angle. Oriented grayscale statistics are then performed on each single-frame image to obtain the single-frame image and its directional grayscale statistical value corresponding to each polarization angle. Specifically: Fixed supplementary lights are pre-installed in the training area, with the central optical axis of the lights pointing towards the area where the trainee is located. The supplementary lights are connected to a constant voltage power supply to keep the illumination level stable during training. An industrial camera with manual exposure lock function is selected and fixed on a support stable relative to the trainee's position. A fixed focal length and fixed exposure parameters are set while ensuring that the field of view covers the entire action area.
[0024] A rotatable linear polarizer is fixed to the front of the camera lens using a retaining ring structure, ensuring that the central optical axis of the rotatable linear polarizer is coaxial with the camera's optical axis. The outer ring of the rotatable linear polarizer is connected to the output shaft of a stepper motor, which is driven by a controller.
[0025] Controller setting polarization angle scanning range and equal angular intervals By sending a fixed number of pulse signals to the stepper motor, a rotatable linear polarizer is made to rotate between two adjacent shots. , corresponding to the The polarization angles are: ; in, For the first The polarization angle of the linear polarizer can be rotated during each acquisition. The initial polarization angle, The polarization angle step interval, This is the polarization angle index.
[0026] While the trainee maintains the designated action, the rotatable linear polarizer rotates to a certain polarization angle. After stabilization, the controller sends a trigger signal to the camera, which then acquires a single-frame image at that polarization angle. This single-frame image is recorded as... .
[0027] For a single frame image When performing directional grayscale statistics, first read the image resolution and record the number of rows in the image as... The number of columns in the image is denoted as And record the grayscale value of each pixel in the image as ,in This is the pixel row index, with values ranging from 1 to... , The pixel column index has a value range of 1 to 1. The controller calculates the first... Oriented grayscale statistics of a single frame image ; in, For at the polarization angle The directional grayscale statistics of a single frame image acquired below. The number of rows in the image. Number of columns in the image For at the polarization angle The single-frame image captured below is located at the first Line number The pixel grayscale value of the column.
[0028] The controller will collect the polarization angle from each acquisition. With the corresponding directional grayscale statistics The data is stored in pairs in memory to form a set of polarization angle-orientation grayscale statistical values. , .
[0029] S12: Using the polarization angles obtained in S11 and their corresponding directional grayscale statistical values as data points, fit the data to construct a grayscale-polarization angle curve with the polarization angle as the abscissa and the directional grayscale statistical values as the ordinate. Specifically: The controller first processes the polarization angle-orientation grayscale statistical data set obtained in S11. According to polarization angle Sort the values by size to ensure that: ; in, This represents the number of polarization angles collected. Each of the sorted values... These are considered as discrete sampling points on the grayscale-polarization angle curve.
[0030] When constructing a continuous grayscale-polarization angle curve, the controller operates at any two adjacent data points. and The interpolation polarization angle is calculated using a linear interpolation method. Corresponding directional grayscale statistics The interpolation formula is: ; in, For located in the interval Any polarization angle within, This is the polarization angle. The corresponding directional grayscale statistical interpolation results, and At polarization angles and The directional grayscale statistics of a single frame image acquired below.
[0031] The controller operates in fixed angular step sizes within each adjacent polarization angle interval. Sampling is performed, and the corresponding interpolation formula is used to calculate the interpolation result. and all Connect them in order of polarization angle to obtain a line in A continuous grayscale-polarization angle curve within the range.
[0032] The grayscale-polarization angle curve is stored in array form, where the array index corresponds to the polarization angle sampling point and the array element corresponds to the directional grayscale statistical value.
[0033] S13: Perform peak detection on the generated grayscale-polarization angle curve to identify two local peak points in the curve. Record the polarization angles corresponding to these two local peak points as the solid angle and the mirror angle, respectively. Specifically: The controller performs a local comparison on each sampling point except for the beginning and end endpoints, based on the grayscale-polarization angle curve array stored in S12. For index 1... sampling points Obtain its previous sampling point and the next sampling point Calculate the directional grayscale statistical value and determine whether it meets the following requirements: ; When the above conditions are met, the index will be... Corresponding data points Mark it as a local peak point and add it to the list of local peak points.
[0034] After completing the traversal of all sampling points, the controller counts the directional grayscale statistical value of each local peak point in the local peak point list, sorts these directional grayscale statistical values according to their numerical values, and selects the two local peak points with the largest values as the target local peak points.
[0035] Set the polarization angle corresponding to the local peak point of the target with a large gray-scale statistical value as The corresponding directional grayscale statistical value is The polarization angle corresponding to the local peak point of the target with a smaller directional gray-scale statistical value is The corresponding directional grayscale statistical value is And satisfy: ; The controller will adjust the polarization angle. Recorded as the solid angle, the polarization angle The solid angle and the mirror angle are recorded as the mirror angle, and the solid angle and the mirror angle are written into the parameter storage area to be used as input parameters in the subsequent binary mask generation process.
[0036] S14: Based on the determined entity angle and mirror angle, calculate the difference in grayscale value of each pixel in the original single-frame image at these two polarization angles. Based on the comparison between this grayscale value difference and a preset threshold, generate a binary mask, specifically: In all the single-frame images detected by S11, the controller determines the entity angle based on the... A single frame image acquired at this polarization angle is selected and denoted as the solid angle single frame image. According to the mirror angle A single frame image acquired at this polarization angle is selected and denoted as the mirror angle single frame image. .
[0037] Single-frame image of solid corner and mirror angle single frame image Consistent with the original acquisition resolution, the controller targets each pixel coordinate. Read single-frame image of entity corner grayscale value at this coordinate and mirrored single-frame images grayscale value at this coordinate The difference in grayscale values of the pixel is calculated using the following formula. ; in, This is the pixel row index, with values ranging from 1 to... , The pixel column index has a value range of 1 to 1. , For a single frame image of a solid corner, in pixel coordinates grayscale value at that location For a single frame image with a mirrored angle in pixel coordinates grayscale value at that location This is the absolute difference in grayscale value of the pixel at the solid angle and the mirror angle.
[0038] The system pre-determines a global threshold through statistical analysis of multiple sets of training samples. The global threshold is stored in the threshold register. The controller bases the threshold on the difference in grayscale values. With global threshold Generate a binary mask based on the size relationship. The specific rules can be written as follows: ; in, For the binary mask at pixel coordinates Pixel value at that location, This indicates that the pixel belongs to a solid region. This indicates that the pixel belongs to either the mirrored area of the light source or the background area.
[0039] After completing the calculation of all pixel coordinates After calculation, the controller output is... The resulting binary mask image contains white areas with a pixel value of 1, which correspond to the main distribution areas of the trainee's physical outline, and black areas with a pixel value of 0, which correspond to the reflected light and shadow areas and the background area of the training venue. This provides accurate pixel-level physical region markings for subsequent human figure segmentation and skeleton extraction of the physical image.
[0040] S2: Multiply the binary mask pixel-by-pixel with the original image to obtain a solid image with filtered light shadows and reflections. Perform human figure segmentation on this solid image to obtain the solid contour, and extract the skeleton from the solid contour to obtain a two-dimensional coordinate sequence containing only the real human body joints, specifically: S21: Perform pixel-by-pixel multiplication between the binary mask generated in S1 and the corresponding original image, so that the original image pixels corresponding to the white areas in the binary mask are preserved, while the original image pixels corresponding to the black areas are suppressed, thus obtaining a solid image with the shadow reflections filtered out, specifically: In S1, a binary mask with the same resolution as the original image has been obtained. This binary mask is denoted as... ,in This is the pixel row index, with values ranging from 1 to... , The pixel column index has a value range of 1 to 1. .
[0041] The original image corresponding to this binary mask is denoted as... The original image is a grayscale image, and its pixel grayscale values have been preprocessed to be uniformly limited to 0, 255 or uniformly normalized to the 0, 1 range.
[0042] According to the generation rules in S1, pixels belonging to the entity region in the binary mask are assigned a value of 1, and pixels belonging to the light shadow mirror region and the background region are assigned a value of 0. Therefore, we have: ; The controller iterates through the binary mask and the original image one by one according to the pixel coordinates, and performs a check on each pixel coordinate. Performing a multiplication operation converts the pixel values of the entity image into... Defined as: ; in, For entity images in pixel coordinates The pixel grayscale value at that location. When that pixel location is white in the binary mask. The corresponding original image pixels are completely preserved; when the pixel position is black, ... The corresponding original image pixels are suppressed to 0.
[0043] By performing the above pixel-by-pixel multiplication operation on all pixels, a physical image containing only the trainee's entity region and with the interference of light shadows and reflections filtered out is obtained. .
[0044] S22: Perform semantic segmentation-based human image segmentation on the entity image obtained in S21. Classify the pixels in the image into human body and background, extract the pixel set of the human body region. The pixel set of the human body region constitutes the entity contour of the entity image. Perform morphological closing operation on the entity contour to fill the small holes inside the contour. Specifically: The entity image obtained in S21 As input, a pre-trained semantic segmentation neural network model is fed in. This semantic segmentation neural network model adopts a fully convolutional neural network with an encoder-decoder structure. The encoder part extracts multi-scale features from the entity image through multiple convolutions and downsampling operations, while the decoder part restores the semantic information to the same spatial resolution as the input image through upsampling and skip connections, and generates an entity image identical to the input image at the output layer. Segmented images of the same size.
[0045] Let the semantic segmentation neural network model be a function. Its input is an entity image. The output is a probability map of each pixel belonging to each category. ,in For category indexing, the value set in this invention is: Category 0 represents the background, and category 1 represents the human body. Therefore: ; For each pixel coordinate The category label of a pixel is obtained by selecting the category with the highest probability along the category dimension. : ; Among them, when When, the pixel is classified as a human body, when At that time, the pixel was classified as background. In order to extract entity contours from the segmentation results, all pixels marked as human body were combined into a pixel set of human body region, and an entity contour mask was constructed. , is defined as: ; in, The set of pixels represents the entity contour region of the entity image. Considering that there may be small holes inside the contour due to noise in the semantic segmentation output, in order to obtain entity contours with good connectivity, a entity contour mask is applied. Performing a morphological closing operation, specifically: Let the structural element be Morphological dilation operation is denoted as The morphological erosion operation is denoted as e, and the result of the closing operation of the solid contour mask is... Defined as: ; Among them, expansion operation In structural elements Under the influence of this process, the solid contour boundary is expanded, causing the boundaries of internal holes to fill inwards, followed by an erosion operation. The outward portion of the outline is shrunk to a position close to the original outline, thus preserving the overall shape of the outline and filling the small holes inside the outline.
[0046] After the closing operation is completed, The set of pixels forms a solid outline with a complete topological structure and filled internal holes.
[0047] S23: Apply a skeletonization algorithm to the entity contour extracted in S22, iteratively removing contour boundary pixels until its topological structure converges to a line structure with a single pixel width. This line structure is the human skeleton. Then, locate the key joints on the human skeleton and output the two-dimensional coordinate sequence of these key joints, specifically: The solid contour mask obtained in S22 As input to the skeletonization algorithm, an iterative morphological thinning algorithm is executed, and the intermediate results during the skeletonization process are denoted as... ,in The index is the number of iterations, at the initial time. At that time, there were: ; In each iteration, for All boundary pixels that meet the preset thinning conditions are deleted. For any candidate pixel... A pixel is considered a deletable boundary pixel if and only if it meets the following condition: (1) This means that the pixel belongs to the current entity outline; (2) Centered on this pixel The number of foreground pixels in the neighborhood is denoted as ,satisfy: ; (3) Count the number of transitions from the background pixel to the foreground pixel in the neighborhood in a clockwise direction centered on the pixel, and record it as . ,satisfy: ; (4) Deleting the pixel will not cause any connection in the entity contour to break, that is, the connected region where the pixel is located will still maintain a single connected component after deletion.
[0048] For all pixels that satisfy the above conditions In the current iteration, set its pixel value to 0 to obtain the skeleton image for the next iteration: ; Repeat the above iterative process until, in a certain iteration, no pixels satisfying the thinning condition are deleted, i.e., for all... All of the following are available: ; At this point, it is assumed that the skeletonization process has converged, and the final result is... A human skeleton mask with a width of one pixel is denoted as . .
[0049] To locate key joints in the human skeleton, a human skeleton mask is used. Input a pre-trained joint detection model. This joint detection model is a deep learning-based pose estimation model, denoted as the function. Based on the topological structure and local features of the human skeleton, this model outputs the two-dimensional coordinates of each preset key joint in the entity image coordinate system. Let the number of preset key joints in the human body be... The output of the joint detection model is then expressed as: ; in, For the first The two-dimensional coordinates of key joints in the entity image. The horizontal pixel coordinates are... These are the pixel coordinates in the vertical direction.
[0050] After obtaining the two-dimensional coordinates of all key joints, the system sorts these joints according to a preset human joint topology order, arranging them into an ordered list: ; List This is the two-dimensional coordinate sequence of key joints used in subsequent action matching processing in this invention, which fully represents the human posture structure corresponding to the current trainee's entity image.
[0051] S3: Match the two-dimensional coordinate sequence with the standard action template to calculate the action similarity. If the action similarity is lower than a set threshold, feed back the current polarization angle error to S1, drive the micro motor to rotate the polarizer to the new solid angle, re-acquire and update the binary mask, forming a closed loop of "polarization orientation - mirror rejection - action recognition" until continuous The process ends when the frame similarity meets the standard, specifically: S31: Compare the spatial position of the two-dimensional coordinate sequence obtained in step S2 with the pre-stored standard action template, calculate the overall matching degree between the two sets of coordinates, and obtain a quantified action similarity, specifically: The two-dimensional coordinate sequence of key joints output by S2 is denoted as...
[0052] in, The number of critical nodes. For the first The two-dimensional coordinates of key joints in the entity image.
[0053] The two-dimensional coordinate sequence of the corresponding key joints in the pre-stored standard motion template is denoted as... ; To eliminate the effects of differences in overall human body size and global translation, we first separately... and Normalization is performed.
[0054] Calculate the geometric center of the current action coordinate sequence: ; And calculate the geometric center of the standard motion template coordinate sequence: ; For the current motion, calculate the average distance from each joint to the geometric center as a scale reference. : ; For a standard action template, calculate a scale reference of the same form. : ; After obtaining the above parameters, the current action and the standard action template are respectively decentered and scale-normalized to obtain normalized coordinates: ; ; After normalization, for each key node Calculate the Euclidean distance between the current action and the standard action template in the normalized space, as follows: ; The overall position matching error is obtained by aggregating the Euclidean distances of all key joints. The overall position matching error can be defined as the average distance: ; in, The smaller the value, the closer the current action is to the standard action template.
[0055] The overall position matching error is input into a preset mapping function to obtain an action similarity score between 0 and 1. : ; Mapping function During the system calibration phase, based on the training samples, it is pre-determined to be a monotonically decreasing function from the overall position matching error to the action similarity value. When the overall position matching error is zero, the action similarity is close to 1, and as the overall position matching error increases, the action similarity gradually approaches 0.
[0056] After completing the above calculations, the controller will The similarity of the current action used in step S3 is stored in the cache.
[0057] S32: Compare the motion similarity calculated in S31 with a preset threshold. If the motion similarity is lower than the threshold, calculate a polarization angle error to correct the imaging conditions based on the difference between the current motion similarity and the threshold. Specifically: Pre-set the movement similarity threshold based on the training requirements and scoring criteria. This threshold is used to distinguish between actions that meet the standard and those that do not. The controller reads the current action similarity from step S31. Calculate the similarity difference : ; when When the similarity of the current action is no less than the set threshold, the polarization angle error can be set to zero. ; when When the similarity of the current action is lower than the set threshold, it means that compensation needs to be made by adjusting the imaging conditions.
[0058] In this case, the similarity difference is multiplied by a preset proportional gain coefficient. The polarization angle error used to correct imaging conditions is obtained: ,in, The proportional gain coefficient is a constant obtained during system calibration based on the rotational resolution of the rotatable linear polarizer and the sensitivity of the grayscale-polarization angle curve under current brightness conditions. Its unit is radians per similarity unit or degrees per similarity unit. A larger proportional gain coefficient results in a larger polarization angle error for a given similarity difference, and a faster adjustment speed. To prevent excessive polarization angle adjustment, upper and lower limits are set for the polarization angle error, with the maximum allowable polarization angle adjustment range set as follows: ,right Perform saturation treatment: ; Will This is ultimately used to control the polarization angle error of the micro motor and is then passed on to the next step.
[0059] S33: The polarization angle error calculated in S32 is used as a control signal and fed back to the micromotor driving the rotatable linear polarizer. The micromotor rotates the rotatable linear polarizer to a new polarization angle according to the polarization angle error. Based on this new polarization angle, step S1 is triggered again to acquire a new single-frame image and update the binary mask. Then, the processing flow of S2 and S31-S32 is repeated. Specifically: The micromotor driving the rotatable linear polarizer is a stepper motor, with a minimum step angle of . The controller receives the saturated polarization angle error from S32. And calculate the number of steps to be performed based on the motor step angle: ; in, The number of step pulses to be executed. This is for rounding up to the nearest integer.
[0060] The direction of rotation is determined by the sign of the polarization angle error. At that time, the controller outputs in a preset clockwise direction. A step pulse, when When the controller outputs counterclockwise... A step pulse, when Sometimes step pulses are not output.
[0061] Let the current polarization angle be... After executing step control, the new polarization angle is: ; After detecting the stepper motor's position signal, the controller updates the current polarization angle to... It immediately sends a trigger signal to the camera, causing the camera to switch to a new polarization angle. Next, acquire a new single-frame image.
[0062] Subsequently, the controller re-executes S11-S14 under the new polarization angle, recalculates the grayscale-polarization angle curve, re-identifies the entity angle and mirror angle, and generates an updated binary mask to obtain an entity image that matches the new imaging conditions. After updating the binary mask and entity image, the controller again executes the entity image processing flow described in S2, generates a new two-dimensional coordinate sequence of key joints, and repeats the motion similarity calculation and polarization angle error calculation flows of S31 and S32.
[0063] The above process constitutes a closed-loop control link of "action similarity evaluation - polarization angle adjustment - re-imaging - re-recognition".
[0064] S34: Continuously execute the closed-loop feedback and reprocessing process of step S33, and monitor the similarity of actions obtained each time. If the calculated action similarity is not lower than the set threshold, then the current action is deemed to have met the standard and the closed loop ends. Specifically: At the start of the closed loop, the system sets the iteration counter. Initialize to 0, and set the continuous target counter. Initialize to 0 and set the maximum allowed number of iterations. After each completion of the action similarity calculation between step S31 and step S32, the iteration counter is incremented by 1. .
[0065] Read the action similarity calculated in the current iteration , and the set threshold Compare. If Then increment the continuous achievement counter by 1, which is represented as: .
[0066] like Then the continuous compliance counter will be reset to 0, which is represented as:
[0067] The system pre-sets the number of consecutive successes based on the stability requirements of different training subjects. , Let be a positive integer, ranging from 3 to 10. At any given moment, the following condition must be met: At that time, the controller determines that the current action is the most recent continuous action. If the set threshold is reached in each iteration, the trainee's current action is considered stable and up to standard. The subsequent polarization angle adjustment and re-imaging process is immediately stopped, the closed-loop control ends, and the current action similarity and the final two-dimensional coordinate sequence of key joints are output to the upper-level training evaluation system.
[0068] During closed-loop execution, when the iteration counter reaches the maximum allowed number of iterations, that is: And still not satisfied with continuous When the similarity of the actions is not lower than the set threshold, the system determines that the current recognition process has failed to converge within the limited number of attempts and is regarded as a recognition failure.
[0069] In this state, the controller stops further polarization angle adjustment and image acquisition, and outputs an error report containing the reason for failure, the similarity value of the last action, and the two-dimensional coordinate sequence of key joints, providing a basis for subsequent manual review and training process adjustments.
[0070] like Figure 2As shown, an image processing and recognition system based on a personnel training system is used to implement the aforementioned image processing and recognition method based on a personnel training system, and includes the following modules: The polarization imaging control module is used to control the camera and its front-end rotatable linear polarizer and micro motor when the fixed supplementary lighting in the training area is on, so as to acquire single-frame images of the trainee performing a specified action at different polarization angles. The mirror separation module is used to receive single-frame images acquired by the polarization imaging control module, perform directional grayscale statistics and construct grayscale-polarization angle curves, identify entity angles and mirror angles, and generate a binary mask for separating lamp shadow mirrors. The entity processing module is used to multiply the binary mask generated by the mirror separation module with the corresponding original image to obtain the entity image, and to perform human portrait segmentation and skeleton extraction on the entity image, outputting a two-dimensional coordinate sequence of real human joints; The evaluation and closed-loop control module compares the two-dimensional coordinate sequence output by the entity processing module with the standard action template to calculate the action similarity. When the similarity is lower than a set threshold, it calculates the polarization angle error and sends a control command to the polarization imaging control module to drive the polarizer to rotate to a new angle, triggering a new round of processing until the action meets the standard.
[0071] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0072] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An image processing and recognition method based on a personnel training system, characterized in that, Includes the following steps: S1: With the fixed supplementary lighting on in the training area, a rotatable linear polarizer is simultaneously installed on the front of the camera lens to capture a single frame image of the trainee completing a specified action. The single frame image is subjected to directional grayscale statistics. The grayscale-polarization angle curve is obtained by using the polarization angle as the scanning variable. The polarization angles corresponding to the two local peaks in the curve are recorded as the entity angle and the mirror angle, respectively. A binary mask is generated based on the grayscale difference between the two angles. The white area of the mask represents the entity and the black area represents the light shadow mirror. S2: Multiply the binary mask with the original image pixel by pixel to obtain the entity image with the light shadows and reflections filtered out. Perform human figure segmentation on the entity image to obtain the entity outline, and extract the skeleton of the entity outline to obtain a two-dimensional coordinate sequence containing only the real human joints. S3: Match the two-dimensional coordinate sequence with the standard action template to calculate the action similarity. If the action similarity is lower than a set threshold, feed back the current polarization angle error to S1, drive the micro motor to rotate the polarizer to the new solid angle, re-acquire and update the binary mask, forming a closed loop of "polarization orientation - mirror rejection - action recognition" until continuous The process ends when the frame similarity meets the standard.
2. The image processing and recognition method based on a personnel training system according to claim 1, characterized in that, S1 includes: S11: With the fixed supplementary lighting on in the training area, a rotatable linear polarizer is synchronously installed on the front of the camera lens. The polarizer is controlled to rotate at equal angular intervals. At each polarization angle, a single-frame image of the trainee completing the specified action is acquired. The directional grayscale statistics of each single-frame image are then performed to obtain the single-frame image and its directional grayscale statistics value corresponding to each polarization angle. S12: Fit the polarization angles obtained in S11 and their corresponding directional gray-scale statistical values as data points to construct a gray-scale-polarization angle curve with polarization angle as the abscissa and directional gray-scale statistical values as the ordinate. S13: Perform peak detection on the generated grayscale-polarization angle curve, identify two local peak points in the curve, and record the polarization angles corresponding to these two local peak points as the solid angle and the mirror angle, respectively. S14: Based on the determined entity angle and mirror angle, calculate the difference in grayscale value of each pixel in the original single-frame image at these two polarization angles. Based on the comparison result of the grayscale value difference with the preset threshold, generate a binary mask, where the white area of the binary mask represents the entity and the black area represents the light shadow mirror.
3. The image processing and recognition method based on a personnel training system according to claim 2, characterized in that, The specific method for performing directional grayscale statistics on each single-frame image is as follows: for each single-frame image acquired at a specific polarization angle, calculate the average grayscale value of all pixels in the single-frame image, and use the average grayscale value as the directional grayscale statistical value corresponding to this polarization angle.
4. The image processing and recognition method based on a personnel training system according to claim 3, characterized in that, The peak detection of the gray-polarization angle curve and the identification of two local peak points in the curve are specifically as follows: traverse every data point on the gray-polarization angle curve, and determine that if the directional gray-level statistical value of a certain data point is greater than the directional gray-level statistical values of the preceding and following data points, the data point is a local peak point. Select the two points with the largest directional gray-level statistical values from all local peak points as target local peak points.
5. The image processing and recognition method based on a personnel training system according to claim 4, characterized in that, S2 includes: S21: Perform a pixel-by-pixel multiplication operation between the binary mask generated in S1 and the corresponding original image, so that the original image pixels corresponding to the white areas in the binary mask are retained and the original image pixels corresponding to the black areas are suppressed, thereby obtaining a physical image with the light shadow image filtered out. S22: Perform semantic segmentation-based human image segmentation processing on the entity image obtained in S21, classify the pixels in the image into human body and background, extract the pixel set of the human body region, and the pixel set of the human body region constitutes the entity outline of the entity image. S23: Apply a skeletonization algorithm to the extracted entity contour, iteratively remove contour boundary pixels until its topology converges into a line structure with a single pixel width. This line structure is the human skeleton. Then locate the key joints on the human skeleton and output the two-dimensional coordinate sequence of these key joints.
6. The image processing and recognition method based on a personnel training system according to claim 5, characterized in that, The step of performing pixel-by-pixel multiplication between the binary mask and the corresponding original image specifically involves multiplying the pixel value of each pixel in the binary mask with the gray value of the corresponding pixel in the original image, wherein the pixel value of the white area of the binary mask is 1 and the pixel value of the black area is 0.
7. The image processing and recognition method based on a personnel training system according to claim 5, characterized in that, The semantic segmentation-based human image segmentation process specifically involves: inputting the entity image into a pre-trained semantic segmentation neural network model; the semantic segmentation neural network model outputs a segmentation map with the same size as the input image; the category label of each pixel in the segmentation map indicates whether it belongs to the human body or the background; and all pixels labeled as the human body constitute the entity outline.
8. The image processing and recognition method based on a personnel training system according to claim 7, characterized in that, S3 includes: S31: Compare the spatial position of the two-dimensional coordinate sequence with the pre-stored standard action template, calculate the overall matching degree between the two sets of coordinates, and obtain a quantified action similarity. S32: Compare the action similarity with a preset threshold. If the action similarity is lower than the threshold, calculate a polarization angle error to correct the imaging conditions based on the difference between the current action similarity and the threshold. S33: The calculated polarization angle error is used as a control signal and fed back to the micro motor that drives the rotatable linear polarizer. The micro motor rotates the rotatable linear polarizer to a new polarization angle according to the polarization angle error. Based on this new polarization angle, S1 is triggered again to acquire a new single-frame image and update the binary mask. Then, the processing flow of S2 and S31-S32 is repeated. S34: Continuously execute the closed-loop feedback and reprocessing process of S33, and monitor the similarity of actions obtained each time. If the calculated similarity of actions is not lower than the set threshold, the current action is deemed to have met the standard and the process ends.
9. The image processing and recognition method based on a personnel training system according to claim 8, characterized in that, The step of comparing the spatial position of the two-dimensional coordinate sequence with the pre-stored standard action template and calculating the overall matching degree between the two sets of coordinates to obtain a quantified action similarity is as follows: calculate the Euclidean distance between each joint point in the two-dimensional coordinate sequence and the corresponding joint point in the standard action template, sum or average the Euclidean distances of all joint points, and convert the sum or average of the distances into an action similarity value between 0 and 1 through a preset mapping function.
10. An image processing and recognition system based on a personnel training system, used to implement the image processing and recognition method based on a personnel training system as described in any one of claims 1-9, characterized in that, Includes the following modules: The polarization imaging control module is used to control the camera and its front-end rotatable linear polarizer and micro motor when the fixed supplementary lighting in the training area is on, so as to acquire single-frame images of the trainee performing a specified action at different polarization angles. The mirror separation module is used to receive a single-frame image acquired by the polarization imaging control module, perform directional grayscale statistics and construct a grayscale-polarization angle curve, identify the entity angle and the mirror angle, and generate a binary mask for separating the lamp shadow mirror. The entity processing module is used to multiply the binary mask generated by the mirror separation module with the corresponding original image to obtain an entity image, and to perform human portrait segmentation and skeleton extraction on the entity image, outputting a two-dimensional coordinate sequence of real human joints; The evaluation and closed-loop control module is used to compare the two-dimensional coordinate sequence output by the entity processing module with the standard action template to calculate the action similarity. When the similarity is lower than a set threshold, the polarization angle error is calculated, and a control command is sent to the polarization imaging control module to drive the polarizer to rotate to a new angle, triggering a new round of processing until the action meets the standard.