An image processing method

By combining multi-scale feature extraction and prior hand shape features, the problems of low accuracy and poor generalization ability in hand shape target image segmentation of the morning inspection machine are solved, and accurate segmentation under different conditions is achieved, which can meet the detection needs of different groups of people and hand postures.

CN122392126APending Publication Date: 2026-07-14SUZHOU DEWO INTELLIGENT SYST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUZHOU DEWO INTELLIGENT SYST
Filing Date
2026-04-13
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing morning inspection machines suffer from low accuracy and poor generalization ability in the process of segmenting hand-shaped target images. In particular, when there is uneven lighting, cluttered background, and changes in hand posture, general image segmentation models are prone to misjudging background noise or losing fine areas such as hand edges and fingertips.

Method used

By combining multi-scale feature extraction and hand shape prior features, the first image features and the first hand shape prior features of the target image are obtained. The hand shape prior feature set is used to match and adapt the target image for image segmentation. Hand shape structure constraints are added to ensure that the segmentation results conform to the hand shape structure rules.

Benefits of technology

It improves the accuracy and robustness of hand image segmentation, and can accurately restore fine areas such as hand edges and fingertips under different lighting and background conditions. It enhances the robustness and generalization ability of the segmentation results, and adapts to the detection needs of different groups of people and hand poses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122392126A_ABST
    Figure CN122392126A_ABST
Patent Text Reader

Abstract

The application discloses an image processing method. The method comprises the following steps: acquiring a target image containing a hand image; performing multi-scale feature extraction on the target image to obtain first image features of the target image at each scale, and determining first fusion features corresponding to each scale according to each first image feature and a first hand shape prior feature; wherein the first hand shape prior feature is determined according to the target image and a pre-constructed hand shape prior feature set; and determining a hand image segmentation result of the target image according to each first fusion feature and the first hand shape prior feature. By using the method, the first hand shape prior feature of the target image is matched and adapted based on the hand shape prior feature set, and the first hand shape prior feature is integrated into the whole image segmentation process, thereby realizing deep guidance of prior knowledge to the segmentation process, improving the image segmentation accuracy and robustness in a complex shooting environment, and providing strong support for subsequent intelligent health screening of hand signs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more particularly to an image processing method. Background Technology

[0002] As an intelligent health screening device for various scenarios, the morning health check machine needs to collect hand images to detect signs such as nail abnormalities. Accurate segmentation of hand-shaped targets is the prerequisite and key to subsequent sign analysis.

[0003] Currently, when morning inspection machines segment hand-shaped targets in images, they mostly use general image segmentation models such as UNet and FCN. These models rely solely on the pixel and texture features of the image for segmentation, which has the following technical drawbacks: 1. Hand images captured by the morning inspection machine are easily affected by the shooting environment, such as uneven lighting, cluttered background, slight changes in hand posture, and partial adhesion between the hand and the morning inspection machine table. General image segmentation models are prone to misidentifying background noise as hand-shaped targets or losing fine areas such as hand edges and fingertips, resulting in low segmentation accuracy and problems such as target breakage and blurred outlines. 2. The general model lacks shape constraints in the segmentation process and has poor generalization ability for hand images of different poses and age groups, which cannot meet the high-precision segmentation requirements of the morning inspection machine. Summary of the Invention

[0004] This invention provides an image processing method to solve the problems of low accuracy and poor generalization ability of existing morning inspection machines when performing image segmentation on hand-shaped targets.

[0005] In a first aspect, embodiments of the present invention provide an image processing method, the method comprising: Obtain the target image containing the hand image; Multi-scale feature extraction is performed on the target image to obtain the first image features of the target image at each scale, and the first fusion features corresponding to each scale are determined based on each first image feature and the first hand shape prior feature; wherein, the first hand shape prior feature is determined based on the target image and a pre-constructed set of hand shape prior features; Based on each of the first fusion features and the first hand shape prior features, the hand image segmentation result of the target image is determined.

[0006] The image processing method provided in this invention involves acquiring a target image containing a hand image; extracting multi-scale features from the target image to obtain first image features at each scale; and determining first fusion features corresponding to each scale based on each first image feature and a first hand shape prior feature. The first hand shape prior feature is determined based on the target image and a pre-constructed set of hand shape prior features. The hand image segmentation result of the target image is determined based on each first fusion feature and the first hand shape prior feature. This method, by matching and adapting the first hand shape prior feature of the target image to the set of hand shape prior features and integrating the first hand shape prior feature into the entire image segmentation process, achieves deep guidance of the segmentation process by prior knowledge. This effectively solves the problems of low segmentation accuracy and poor generalization ability caused by the significant influence of lighting and background in existing methods, providing accurate hand region localization for subsequent intelligent health screening of hand signs such as nail abnormalities and rashes. Specifically, by obtaining the first image features of the target image at various scales, the different granular features of the hand shape can be comprehensively captured, taking into account both detailed and global features, laying the foundation for subsequent accurate segmentation. Based on the hand shape prior feature set, the first hand shape prior features of the target image are matched and adapted, and hand shape structure constraints are added to the first image features, effectively avoiding problems such as contour distortion and missing fingertips caused by illumination and background when segmenting hand images. Then, based on each first fusion feature and the first hand shape prior feature, the hand image segmentation result of the target image is determined, further ensuring that the hand image segmentation result conforms to the hand shape structure rules, avoiding segmentation deviations caused by background noise, uneven illumination, and other interferences, achieving accurate restoration of fine areas such as hand shape edges, fingertips, and finger gaps, improving the robustness, accuracy, and scene generalization ability of the segmentation result, and being able to meet the hand detection needs of different groups such as children and adults under various hand postures.

[0007] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 A flowchart of an image processing method provided in an embodiment of the present invention; Figure 2This is a schematic diagram of the structure of an image processing device provided in an embodiment of the present invention; Figure 3 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation

[0010] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0011] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0012] It is understood that before using the technical methods disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0013] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as the electronic device, application, server, or storage medium performing the operations of this disclosed technology.

[0014] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0015] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0016] Based on this, embodiments of the present invention provide an image processing method. Figure 1 This is a flowchart of an image processing method provided by an embodiment of the present invention. The embodiment of the present invention is applicable to scenarios where hand-shaped targets in images are segmented, especially to scenarios where hand-shaped targets in hand images collected by a morning inspection machine are accurately segmented. The method can be executed by an image processing device, which can be implemented in the form of software and / or hardware, and optionally, by an electronic device, preferably a morning inspection machine.

[0017] like Figure 1 As shown, the image processing method provided in this embodiment of the invention may specifically include: S101. Obtain the target image containing the hand image.

[0018] The target image can be understood as an image containing a hand and requiring segmentation. The target image can include the hand image itself, as well as background regions such as the detection platform area. The hand image can be considered as an image containing only the hand region and is the target object for image segmentation. The hand image can include parts such as the palm, fingers, fingertips, finger gaps, and palm edges.

[0019] In this embodiment, the method for obtaining the target image containing the hand image can be as follows: obtaining the original image containing the hand image, preprocessing the original image to eliminate shooting noise and unify image specifications to obtain the target image. For example, obtaining the original image containing the hand image can be done by calling the camera configured on the device to capture the hand placed on the detection table area in real time, thus obtaining the original image containing the hand image; obtaining the original image containing the hand image can also be done by retrieving the original image of the hand placed on the detection table area from the device's photo album at a historical moment in response to a selection operation; obtaining the original image containing the hand image can also be done by retrieving the original image containing the hand image input by the interactive object through the provided interactive interface in response to an input operation.

[0020] For example, preprocessing the original image can include: 1. Gaussian Denoising. A Gaussian filter is used to denoise the original image, suppressing salt-and-pepper noise and Gaussian noise. Specifically, this can be represented as: ;in, The original image after denoising; Original image; For Gaussian kernel function, The standard deviation of the Gaussian kernel can be optionally taken as... ; This is a convolution operation.

[0021] 2. Normalization. The pixel values ​​of the denoised original image are normalized, mapping the pixel values ​​to... An interval, specifically, can be represented as: ;in, The normalized image at the pixel level Pixel value at; For the original image after denoising, at the pixel level Pixel value at; The minimum pixel value of the original image after denoising; This represents the maximum pixel value of the original image after denoising.

[0022] 3. Size Adjustment: The normalized image is adjusted to a fixed size to obtain the preprocessed target image. Optionally, the fixed size is the same as the size of the original image acquired by the morning inspection machine. ;in, The height of the original image; The width of the original image.

[0023] S102. Perform multi-scale feature extraction on the target image to obtain the first image features of the target image at each scale, and determine the first fusion features corresponding to each scale based on each first image feature and the first hand shape prior feature; wherein, the first hand shape prior feature is determined based on the target image and the pre-constructed hand shape prior feature set.

[0024] The first image feature can be considered as image features of different scales (i.e., different resolutions) obtained by multi-scale feature extraction of the target image. The first hand shape prior feature can be understood as the hand shape prior feature selected from the hand shape prior feature set that best matches the target image. It is used as a reference standard to describe what the hand in the target image should look like, thereby constraining the first image feature and avoiding hand shape distortion during segmentation. The hand shape prior feature set can be considered as a set of hand shape prior features corresponding to different postures (fingers open, half-closed), different age groups (children, adults), and different angles. It is used to provide all possible hand shape reference standards to guide the accurate segmentation of the hand image. The first fusion feature can be considered as the first image feature that fuses the first hand shape prior feature. It retains the real details of the hand in the target image and conforms to the shape rules of the hand.

[0025] In this embodiment, the method for multi-scale feature extraction of the target image to obtain the first image features of the target image at each scale can be as follows: the target image is input into an encoder, and multiple cascaded convolutional layers in the encoder extract features from the target image according to the scale set relative to each convolutional layer, thereby obtaining the first image features of the corresponding scale output by each convolutional layer. It can be understood that shallow convolutional layers extract fine-grained detail features, i.e., they can be used to extract the shallow texture features of the target image, while deep convolutional layers extract coarse-grained global features, i.e., they can be used to extract the deep semantic features of the target image.

[0026] In this embodiment, the method for determining the first fusion features corresponding to each scale based on each first image feature and the first hand shape prior feature can be as follows: the first hand shape prior feature is downsampled according to the scale of each first image feature to obtain the downsampled feature, and the first image feature at each scale is fused with the downsampled feature at the corresponding scale to obtain the first image feature at each scale.

[0027] In this embodiment, the method for determining the first hand shape prior feature based on the target image and the pre-constructed hand shape prior feature set can be as follows: extract the features of the target image and the features of each hand shape prior feature in the hand shape prior feature set, compare the features of the target image with the features of each hand shape prior feature respectively, and determine the hand shape prior feature that best matches the target image as the first hand shape prior feature, so as to adapt to the hand image of the corresponding posture and age group in the target image.

[0028] As one implementation method, the steps for constructing the hand shape prior feature set include: Obtain a set of hand sample images; For each first hand sample image in the hand sample image set, global shape features and local contour features of the first hand sample image are extracted, and the global shape features and local contour features are encoded respectively to obtain global encoded features and local encoded features. The global encoded features and local encoded features are weighted and fused based on the fusion weight to generate the hand shape prior features of the first hand sample image. The hand shape prior features of each of the first hand sample images are integrated to generate the hand shape prior feature set.

[0029] The hand sample image set can be understood as a collection of pre-collected first hand sample images containing hand images. These first hand sample images cover different situations, such as different postures (open / half-closed), different age groups, and different lighting conditions, ensuring sample diversity. The first hand sample images can be considered as images used as valid samples to obtain prior features of hand shape. For example, the first hand sample images can be images collected in the actual shooting scenario of a morning inspection machine. Global shape features can be considered as overall shape features, while local contour features can be considered as detailed contour features.

[0030] In this embodiment, firstly, hand shape features can be annotated for each first hand sample image to extract global shape features and local contour features of the hand in the first hand sample image; then, the global shape features are normalized and encoded to obtain global encoded features, and the local contour features are interpolated and filled to obtain local encoded features; then, the global encoded features and local encoded features are weighted and fused based on a pre-determined fusion weight for the adaptation scene to obtain the hand shape prior features of the first hand sample image, which have both global shape constraints and local detail constraints.

[0031] Optionally, the global shape features include the circumscribed convex contour of the hand, the area ratio of the palm to the fingers, and the topological connection relationship of the five fingers; wherein, the topological connection relationship of the five fingers may include the connection position between the fingers and the palm, and the relative distribution between the fingers; the local contour features include a set of fine contour points of the hand and a slender strip-shaped convex constraint of a single finger, wherein the set of fine contour points of the hand includes the contour points of the fingertips, finger gaps, and palm edges.

[0032] For example, the specific method for obtaining the hand shape prior features of the first hand sample image by weighted fusion of the global and local coding features based on fusion weights can be expressed as follows: ; in, These are prior features of hand shape. ; The global encoding features are generated through normalized encoding based on the circumscribed convex contour of the hand, the area ratio of the palm to the fingers, and the topological connectivity of the five fingers. ; The local encoding features are generated by interpolation and filling from the fine contour point set of the hand and the slender strip-shaped convex constraint of a single finger. The fusion weights are used to balance the contributions of global shape features and local contour features.

[0033] In an optional embodiment, the process of extracting multi-scale features from the target image to determine the hand image segmentation result can be implemented by a trained image segmentation model. Based on this, the value of the fusion weights can be adaptively adjusted according to actual sample training. The method for determining the fusion weights can be: the training and updating of the fusion weights are performed synchronously with the overall training of the image segmentation model. Specifically, when the image segmentation model performs batch iterative training using the sample training set, the fusion weights participate in the loss calculation as one of the learnable parameters; based on the target loss function... The calculation results are used to calculate the gradient of the target loss function value (total loss value) with respect to the fusion weights using the backpropagation algorithm. The fusion weights are updated based on the gradient direction using optimizers such as Adam. The update formula can be expressed as: ; in, The learning rate used to train the image segmentation model; The fusion weights before the update; These are the updated fusion weights. When the image segmentation model training converges, the values ​​of the fusion weights at this point are the optimal fusion weights for the current application scenario.

[0034] In this embodiment, the method for integrating the hand shape prior features of each of the first hand sample images to generate the hand shape prior feature set can be as follows: clustering and optimizing the hand shape prior features of all the first hand sample images to remove redundant features, thereby constructing a standardized hand shape prior feature set. For example, the hand shape prior feature set may include... There are hand shape prior features, and the set of hand shape prior features can be specifically represented as: ;in, In this embodiment, the number of hand shape prior features in the hand shape prior feature set is [number]. Take the number of clusters after clustering.

[0035] The above-described technical solution in this embodiment effectively suppresses the influence of background noise and uneven illumination by constructing hand shape prior features that combine global shape constraints and local detail constraints. This enables accurate reconstruction of fine areas such as hand shape edges, fingertips, and finger gaps, thereby improving the accuracy of image segmentation. At the same time, the generated hand shape prior feature set covers hand shape prior features of different age groups and different postures, making the image processing process less affected by changes in hand posture and shooting angle deviations, improving the generalization ability of image processing, and making it more suitable for actual shooting scenarios.

[0036] S103. Based on each first fusion feature and the first hand shape prior feature, determine the hand image segmentation result of the target image.

[0037] The hand image segmentation result can be understood as a result image that clearly distinguishes the hand image from the background, and can be a binary image with the same size as the target image.

[0038] In this embodiment, the method for determining the hand image segmentation result of the target image based on each first fusion feature and the first hand shape prior feature can be as follows: the first fusion feature corresponding to the deepest convolutional layer of the encoder is input into the decoder, and the decoder performs upsampling operation layer by layer to enlarge the feature size. After each layer of upsampling, the first decoded feature obtained by upsampling is matched and corrected in combination with the first hand shape prior feature to eliminate upsampling distortion and constrain the hand shape structure. The correction effect is passed down layer by layer. The first decoded feature after the last layer of correction is input into the convolutional layer to map and obtain the target segmentation probability map. After binarization processing, the final result that distinguishes the hand shape target area from the background area is obtained, which is the hand image segmentation result.

[0039] Optionally, after obtaining the hand image segmentation results, they can be sent to the vital sign detection module of the morning check machine for health screening such as nail abnormalities.

[0040] The image processing method provided in this invention involves acquiring a target image containing a hand image; extracting multi-scale features from the target image to obtain first image features at each scale; and determining first fusion features corresponding to each scale based on each first image feature and a first hand shape prior feature. The first hand shape prior feature is determined based on the target image and a pre-constructed set of hand shape prior features. The hand image segmentation result of the target image is determined based on each first fusion feature and the first hand shape prior feature. This method, by matching and adapting the first hand shape prior feature of the target image to the set of hand shape prior features and integrating the first hand shape prior feature into the entire image segmentation process, achieves deep guidance of the segmentation process by prior knowledge. This effectively solves the problems of low segmentation accuracy and poor generalization ability caused by the significant influence of lighting and background in existing methods, providing accurate hand region localization for subsequent intelligent health screening of hand signs such as nail abnormalities and rashes. Specifically, by obtaining the first image features of the target image at various scales, the different granular features of the hand shape can be comprehensively captured, taking into account both detailed and global features, laying the foundation for subsequent accurate segmentation. Based on the hand shape prior feature set, the first hand shape prior features of the target image are matched and adapted, and hand shape structure constraints are added to the first image features, effectively avoiding problems such as contour distortion and missing fingertips caused by illumination and background when segmenting hand images. Then, based on each first fusion feature and the first hand shape prior feature, the hand image segmentation result of the target image is determined, further ensuring that the hand image segmentation result conforms to the hand shape structure rules, avoiding segmentation deviations caused by background noise, uneven illumination, and other interferences, achieving accurate restoration of fine areas such as hand shape edges, fingertips, and finger gaps, improving the robustness, accuracy, and scene generalization ability of the segmentation result, and being able to meet the hand detection needs of different groups such as children and adults under various hand postures.

[0041] As a first optional embodiment of the present invention, based on the above embodiments, the multi-scale feature extraction of the target image can be performed to obtain the first image features of the target image at each scale, and the first fusion features corresponding to each scale can be determined according to each first image feature and the first hand shape prior feature, which is specifically implemented as follows: a1) Using the encoder in the image segmentation model, feature extraction is performed on the target image according to a set scale based on the multiple cascaded convolutional layers included in the encoder to obtain the first image features at each scale. The convolutional layers include at least one convolutional layer for texture feature extraction and at least one convolutional layer for semantic feature extraction.

[0042] In this embodiment, image processing of the target image can be considered to be mainly achieved through an image segmentation model. The image segmentation model may include a constructed encoder, which is used for feature extraction and is lightweight by reducing the number of channels compared to an existing encoder (such as the UNet encoder). The encoder performs downsampling processing on the target image, and the downsampling method is used to extract and retain more critical image features.

[0043] For example, the encoder contains four convolutional layers, each implemented based on a convolutional block. Each convolutional block consists of convolution, batch normalization, and an activation function. By inputting the target image into the encoder, shallow texture features can be extracted sequentially through each convolutional layer. , and deep semantic features , This allows us to obtain the first image features at the corresponding scales output by the first to fourth convolutional layers, respectively. , , and ,in, , , and This represents the number of channels for the first image feature in each layer.

[0044] b1) The first hand shape prior features are input into the downsampling network structure contained in the prior feature embedding sub-model in the image segmentation model, and the first hand shape prior features are subjected to multi-scale downsampling processing to obtain the second hand shape prior features that match the size of each of the first image features.

[0045] In this embodiment, the image segmentation model may further include sub-models for performing different image processing logic. The prior feature embedding sub-model can be considered as a sub-model built within the image segmentation model, specifically a network model that participates in adding hand-shaped prior feature constraints. Specifically, a downsampling network structure and a first feature fusion module can be created to construct the prior feature embedding sub-model.

[0046] The prior feature embedding sub-model in this embodiment can be considered as having been pre-trained and is a usable sub-model that can be directly downsampled and fused. In this embodiment, the first hand shape prior feature can be input into the downsampling network structure, and multi-scale downsampling processing can be performed on the first hand shape prior feature to obtain a second hand shape prior feature that matches the size of the first image feature output by each convolutional layer. For example, the first hand shape prior feature can be represented as... ; relative to the first image features , , and The prior features of the second hand shape can be represented as , , and .

[0047] c1) Input each of the first image features and each of the second hand shape prior features into the first feature fusion module included in the prior feature embedding sub-model, and fuse each of the second hand shape prior features with the first image features of the corresponding scale to obtain the first fused features corresponding to each scale.

[0048] In this embodiment, this step can be considered as one of the steps executed by the prior feature embedding sub-model in the logical processing. Specifically, it can be executed by the first feature fusion module in the prior feature embedding sub-model. The first feature fusion module adaptively fuses the second hand shape prior features of the same scale with the first image features to obtain the first fused feature corresponding to that scale.

[0049] For example, the method of fusing each of the second hand shape prior features with the first image features at the corresponding scale to obtain the first fused features corresponding to each scale can be expressed as follows: ; in, For the first The first image feature output by each convolutional layer The first fused feature is obtained by fusing the second chiral prior features of the corresponding scale; The attention weight matrix consists of learnable parameters that can be updated as the image segmentation model is trained. For the corresponding number The first image feature output by each convolutional layer The second hand shape prior features, compared with the first image features Size matching; This represents element-wise multiplication. This represents matrix multiplication. The Sigmoid activation function is used to generate prior attention weights, enabling adaptive feature fusion.

[0050] The above technical solution in this embodiment provides an implementation method for fusing the second hand shape prior feature into the first image feature. The first fused feature obtained after fusion includes the pixel, texture and semantic features of the target image, and also incorporates the second hand shape prior feature, which reduces the impact of background noise and ensures that the feature extraction direction always conforms to the shape law of the hand.

[0051] As a second optional embodiment of the present invention, based on the above embodiments, the first hand shape prior feature can be determined according to the target image and the pre-constructed hand shape prior feature set, specifically optimized as follows: a2) The target image is input into the pooling module of the prior feature selection sub-model in the image segmentation model. Global average pooling is performed on the target image to obtain the first feature vector. Global average pooling is then performed on each hand shape prior feature in the hand shape prior feature set to obtain the second feature vector corresponding to each hand shape prior feature.

[0052] The prior feature selection sub-model can be considered as a sub-model used to select hand shape prior features from the hand shape prior feature set that are suitable for the target image. This prior feature selection sub-model can be constructed by creating a pooling processing module and a feature selection module. The first feature vector can be understood as a vector representing hand-related information such as hand pose contained in the target image; the second feature vector can be considered as a vector representing hand-related information such as hand pose corresponding to the hand shape prior features.

[0053] In this embodiment, this step can be considered as one of the steps executed in the logical processing of the prior feature selection sub-model. Specifically, it can be executed by the pooling processing module in the prior feature selection sub-model. The pooling processing module performs global average pooling processing on each hand shape prior feature in the target image and the hand shape prior feature set, thereby obtaining the first feature vector and each second feature vector that can be used to further calculate the matching degree.

[0054] b2) Input the first feature vector and each of the second feature vectors into the feature filtering module in the prior feature filtering sub-model, determine the cosine similarity value between the first feature vector and each of the second feature vectors, and determine the hand shape prior feature corresponding to the maximum cosine similarity value as the first hand shape prior feature adapted to the target image.

[0055] In this embodiment, the feature selection module selects the optimal hand shape prior feature as the first hand shape prior feature based on the input first feature vector and each of the second feature vectors through cosine similarity matching from the hand shape prior feature set. For example, the method of determining the cosine similarity value between the first feature vector and each of the second feature vectors, and determining the hand shape prior feature corresponding to the maximum cosine similarity value as the first hand shape prior feature adapted to the target image, can be expressed as follows: ; in, This is the first eigenvector; This is the second feature vector; This refers to the dot product operation of vectors. Let L2 be the norm of the vector.

[0056] The above-described technical solution in this embodiment provides a specific implementation for determining the first hand shape prior features of the target image, enabling the subsequent addition of effective and suitable hand shape structure constraints to the first image features. This lays the foundation for effectively avoiding problems such as contour distortion and fingertip loss caused by the influence of lighting, background, etc. when segmenting hand images.

[0057] As a third optional embodiment of the present invention, based on the above embodiments, the determination of the hand image segmentation result of the target image according to each of the first fusion features and the first hand shape prior features can be specified as follows: a3) Using the decoder and cross-layer feature fusion sub-model in the image segmentation model, determine the first decoding feature at each scale based on each of the first fusion features.

[0058] In this embodiment, the image segmentation model includes a lightweight decoder corresponding to the encoder. The encoder extracts and retains more critical image features through downsampling, while the decoder upsampling progressively restores the feature size from its smallest value to a size matching the target image. The image segmentation model also includes a cross-layer feature fusion sub-model, used to fuse the first fused features of the encoder's corresponding layer and the upsampled features of the decoder's corresponding layer.

[0059] Optionally, the method of determining the first decoding features at each scale based on each of the first fused features through the decoder and cross-layer feature fusion sub-model in the image segmentation model can be as follows: The first fusion feature with the smallest size among the first fusion features is input into the decoder. According to the upsampling network structure included in the decoder, the first fusion feature is upsampled layer by layer according to a set scale to enlarge the feature size and obtain upsampled features at each scale. Each of the first fusion features and each of the upsampled features are input into the cross-layer feature fusion sub-model. The cross-layer feature fusion sub-model concatenates each of the upsampled features with the first fusion features of the corresponding scale in the channel dimension to obtain each of the first decoding features.

[0060] In this embodiment, the decoder uses the included upsampling network structure to perform a first fusion feature with the smallest size output from the encoder, which is also the deepest first fusion feature (such as...). Upsampling is performed layer by layer. The upsampling network structure can be constructed using transposed convolutions to sequentially generate upsampling features (such as...). , , and Next, the cross-layer feature fusion sub-model directly concatenates each upsampled feature with the first fusion feature of the corresponding convolutional layer of the encoder in the channel dimension to complete the cross-layer fusion and obtain the first decoding feature corresponding to each layer (e.g., The first decoding feature can be considered as the feature to be corrected output by the decoder after upsampling and cross-layer fusion in a network layer.

[0061] b3) Using the feature correction sub-model in the image segmentation model, each of the first decoding features is corrected according to the first hand shape prior features to obtain each of the second decoding features.

[0062] The feature correction sub-model can be understood as a sub-model used to match and correct the first decoded features based on the first hand shape prior features. The second decoded features can be considered as decoded features constrained and corrected by the first hand shape prior features. It not only preserves the original texture and semantics of the target image, but more importantly, it eliminates topological distortions and ensures the physical rationality of the hand shape contour.

[0063] In this embodiment, it can be determined whether the first decoding feature of each layer needs to be corrected based on a preset evaluation rule. If the evaluation indicates that correction is needed, the first decoding feature is corrected based on the first hand shape prior feature to ensure that the second decoding feature obtained after correcting the first decoding feature conforms to the shape pattern of the hand.

[0064] Optionally, the method of modifying each of the first decoding features based on the first hand shape prior features to obtain each of the second decoding features can be as follows: For each first decoded feature, a similarity value is determined between the first decoded feature and the corresponding scale of the second hand shape prior feature, which is obtained based on the first hand shape prior feature and the prior feature embedding sub-model in the image segmentation model. If the similarity value is less than a preset similarity threshold, the first decoding feature is corrected based on the similarity value and the second hand shape prior feature, and the corrected first decoding feature is used as the second decoding feature; if the similarity value is greater than or equal to the preset similarity threshold, the first decoding feature is used as the second decoding feature.

[0065] Optionally, the second hand-shaped prior features can be obtained by acquiring the output of the downsampled network structure in the prior feature embedding submodel of the image segmentation model.

[0066] In this embodiment, the similarity between the first decoding feature corresponding to each layer of the decoder and the second hand shape prior feature at the corresponding scale can be calculated. Based on a preset similarity threshold, the first decoding feature with low matching degree is determined, and the first decoding feature is corrected according to the similarity value and the second hand shape prior feature. The correction method can be specifically expressed as follows: ; in, For the corresponding encoder in the decoder The first decoding feature of the layer The second decoding feature obtained after correction; The correction coefficients are learnable parameters that can be updated as the image segmentation model is trained. The feature similarity function can be optionally used, and the range of similarity values ​​is [range missing]. The lower the similarity value, the greater the correction force.

[0067] Understandably, the correction operation is performed on each layer of the decoder to eliminate upsampling distortion layer by layer and solidify the global and local prior constraints of the hand shape layer by layer. The correction results are passed progressively layer by layer, ensuring that the decoded features passed at each step conform to the shape rules of the hand and avoiding the accumulation of biases in the lower-level decoded features. The second decoded feature corresponding to the last layer of the encoder is the highest resolution feature, which integrates all semantics and details. It is consistent with the size of the input target image and can directly generate pixel-by-pixel segmentation results.

[0068] c3) Using the output sub-model in the image segmentation model, based on the convolutional layer and decision module included in the output sub-model, the second decoding feature with the same size as the target image is mapped to obtain a target segmentation probability map, and the hand image segmentation result of the target image is determined according to the preset probability threshold and the target segmentation probability map.

[0069] The target segmentation probability map can be understood as a map that can represent the probability that each pixel is a hand image or a hand-shaped target.

[0070] In this embodiment, the second decoding feature corresponding to the last layer of the encoder is the highest resolution feature, which integrates all semantics and details, and has the same size as the input target image, allowing for direct generation of pixel-by-pixel segmentation results. Based on this, the second decoding feature corresponding to the last layer of the encoder (such as...) can be used to... Input a 1×1 convolutional layer, and map the output value of the convolutional layer to a sigmoid activation function. The target segmentation probability map is obtained by dividing the interval. For example, the target segmentation probability map can be represented as... ,in, Represents pixels The probability of a hand-shaped target or hand image is determined. Next, the target segmentation probability map is input to the determination module, which includes a preset probability threshold. ,like The method for determining the hand image segmentation result of the target image based on the preset probability threshold and the target segmentation probability map can be as follows: when When a pixel is identified as a hand-shaped target or a hand image, it is determined to be background, thus obtaining the hand image segmentation result. The hand image segmentation result can be presented as a binary image, where 1 represents a hand-shaped target or a hand image, and 0 represents background.

[0071] The above-described technical solution in this embodiment uses an encoder to gradually restore the multi-scale first fusion features to the target image size, while supplementing detailed features through cross-layer fusion, thus avoiding the problems of feature distortion and loss of details. Through the feature correction sub-model, the first decoded features at each scale are corrected layer by layer by combining the first hand shape prior features, realizing the layer-by-layer transmission of hand shape prior constraints, effectively avoiding the problem of feature distortion, strengthening the rationality of hand shape structure, thereby ensuring the integrity of hand contour and the accuracy of details in the hand image segmentation results, and thus improving the accuracy of subsequent health screening.

[0072] As one implementation method, based on the above embodiments, the training steps of the image segmentation model can be further optimized as follows: a4) Obtain a sample training set and an initial segmentation model, wherein the sample training set includes at least one second hand sample image and the corresponding label segmentation result.

[0073] In this embodiment, a sample training set for model training can be obtained first, and an initial segmentation model can be constructed. The sample training set includes second hand sample images specifically determined for model training, along with corresponding labeled segmentation results. The second hand sample images can be understood as real-world sample images containing hand images that are used as input during training; the labeled segmentation results can be considered as segmentation results labeled relative to the second hand sample images, presented as binary images. The initial segmentation model can be understood as a pre-constructed network model containing an encoder, decoder, and various sub-model network structures. The encoder, decoder, and each sub-model in the initial segmentation model can be considered to possess initial learnable network parameters. The included sub-models may include prior feature embedding sub-models and feature correction sub-models, etc.

[0074] b4) Input the second hand sample image into the initial segmentation model to obtain the current segmentation result output by the initial segmentation model and the third hand shape prior feature that matches the second hand sample image.

[0075] The current segmentation result can be understood as the segmentation result obtained by the initial segmentation model through image segmentation processing of the input second hand sample image; the third hand shape prior feature can be considered as the hand shape prior feature in the initial hand shape prior feature set that matches the second hand sample image. The initial hand shape prior feature set is determined based on the hand sample image set and the initial fusion weights.

[0076] In this embodiment, the process of training the image segmentation model can be regarded as an iterative training process. In each iteration, one or more second hand sample images and corresponding label segmentation results can be selected from the sample training set to participate in the model training. Specifically, the selected second hand sample images can be used as input data to the initial segmentation model.

[0077] It is known that the initial segmentation model can process the second hand sample image in sequence according to its encoder, decoder and sub-models to obtain the current segmentation result output by the initial segmentation model.

[0078] c4) Based on the current segmentation result, the label segmentation result, and the third hand shape prior feature, determine the first loss function value of the pixel loss function, the second loss function value of the contour loss function, and the third loss function value of the shape loss function.

[0079] In this embodiment, the iterative training process of the image segmentation model can be considered as a process of inverse adaptive adjustment of the learnable network parameters, and the loss function value can be used as the information value required for inverse adaptive adjustment. This embodiment can achieve the determination of the loss function value required for inverse adaptive adjustment of the learnable network parameters through this step and step d4) below.

[0080] In this embodiment, binary cross-entropy loss can be used to measure the pixel-level difference between the current segmentation result and the label segmentation result. Based on this, the first loss function value of the pixel loss function is determined. The method can be expressed as: ; in, The label segmentation results at the pixel level The value at that location, For the current segmentation result at the pixel point The value at this location is 1, which represents a hand-shaped target or hand image, and 0, which represents the background.

[0081] In this embodiment, a distance transform loss can be used to measure the contour difference between the current segmentation result and the label segmentation result, thereby enhancing the segmentation accuracy of fine regions such as hand edges and fingertips. Based on this, the second loss function value of the contour loss function is determined. The method can be expressed as: ; in, This is a distance transformation function used to convert a binary image into a distance-transformed image, where the value of each pixel is the distance from that point to the nearest target contour.

[0082] In this embodiment, segmentation regions that do not conform to the third chiral prior feature can be penalized by measuring the shape difference between the current segmentation result and the third chiral prior feature. Based on this, the value of the third loss function of the shape loss function is determined. The methods can be: ; in, The third chiral prior feature; the third loss function value The range of values ​​is The higher the shape similarity between the current segmentation result and the third hand shape prior feature, the higher the value of the third loss function. The smaller it is, the larger it is.

[0083] d4) Determine the target loss function value based on the first loss function value, the second loss function value, and the third loss function value.

[0084] In this embodiment, the first loss function value, the second loss function value, and the third loss function value are weighted and fused to obtain the target loss function value. Specifically, it can be expressed as follows: ; in, , and These are loss weights, used to balance the loss contribution of each loss function value, satisfying... ;Optionally, take , , .

[0085] e4) Based on the target loss function value, perform reverse learning to adjust the learnable network parameters in the initial segmentation model to obtain the adjusted initial segmentation model, and return to re-execute the relevant steps of inputting the second hand sample image into the initial segmentation model until the training termination condition is met; determine the initial segmentation model obtained after training as the image segmentation model.

[0086] In this embodiment, the target loss function value determined above can be used as a learning metric and fed back to the encoder, decoder and sub-models of the initial segmentation model through gradient descent and other methods to achieve adaptive adjustment of the available learnable network parameters.

[0087] It is known that a maximum number of training epochs and a learning rate are preset. The training termination condition can be that the number of training epochs reaches the maximum or the target loss function value converges to a preset threshold. Based on this, the adjusted initial segmentation model is equivalent to completing one iteration. Since the training termination condition has not yet been met, it is necessary to return to step b4 to perform a new round of iteration. This process is repeated until the training termination condition is met. The initial segmentation model obtained after training can be considered as an image segmentation model capable of participating in image segmentation processing of target images containing hand images involved in practical applications.

[0088] The above-described technical solution in this embodiment provides a training implementation of the image segmentation model. The image segmentation model formed through training can better achieve effective and high-precision segmentation of target images containing hand images under complex shooting environments, guided by prior hand features.

[0089] Optionally, after obtaining the image segmentation model at the end of training, optimization may be performed by: determining the importance score of each convolutional layer channel according to the convolutional kernel parameters of the image segmentation model, and pruning redundant channels with importance scores lower than the preset score threshold according to the preset score threshold to obtain the pruned image segmentation model; and performing weight quantization processing on the pruned image segmentation model to obtain the lightweight image segmentation model.

[0090] In this embodiment, channel pruning and weight quantization can also be performed on the trained image segmentation model to further reduce the number of model parameters and computational load. As one implementation, the L1 norm of the convolutional kernel can be used to measure channel contribution; the smaller the L1 norm, the lower the channel importance. For example, based on the convolutional kernel parameters of the image segmentation model, the importance score of each convolutional layer channel can be determined as follows: For each convolutional layer, the L1 norm of the included convolutional kernels is determined, and the L1 norm is used as the importance score of the corresponding convolutional layer channel. Then, the importance score is compared with a preset score threshold. If the importance score is less than the preset score threshold, the channel can be considered a redundant channel and pruned to retain the core channels.

[0091] In this embodiment, the weight quantization process for the pruned image segmentation model can be performed by quantizing the 32-bit floating-point weights of all learnable network parameters in the pruned image segmentation model into 8-bit integer weights to reduce the storage overhead and computational load of the weights. The quantization process can use a linear quantization method to ensure that the loss of model accuracy is within an acceptable range.

[0092] This embodiment presents a specific implementation of a lightweight image segmentation model obtained by performing channel pruning and weight quantization on a trained image segmentation model. Compared to existing improved segmentation methods that often increase network layers and feature dimensions to improve accuracy, resulting in large parameter counts and time-consuming inference, this lightweight model significantly reduces parameter count and computational cost, while significantly improving inference speed without a noticeable decrease in segmentation accuracy. It is also compatible with embedded hardware environments, meeting the embedded deployment requirements of devices such as morning inspection machines and satisfying the high demands for real-time performance and accuracy. Furthermore, the lightweight image segmentation model is easy to deploy, requiring no major modifications to the hardware in devices like morning inspection machines, resulting in low modification costs and strong applicability.

[0093] Figure 2 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of the present invention. Figure 2 As shown, the device includes: an acquisition module 21, a first determination module 22, and a second determination module 23, wherein, Acquisition module 21 is used to acquire a target image containing a hand image; The first determining module 22 is used to perform multi-scale feature extraction on the target image to obtain the first image features of the target image at each scale, and to determine the first fusion features corresponding to each scale based on each first image feature and the first hand shape prior feature; wherein, the first hand shape prior feature is determined based on the target image and a pre-constructed set of hand shape prior features. The second determining module 23 is used to determine the hand image segmentation result of the target image based on each of the first fusion features and the first hand shape prior features.

[0094] The image processing apparatus provided in this invention acquires a target image containing a hand image; performs multi-scale feature extraction on the target image to obtain first image features of the target image at each scale; and determines first fusion features corresponding to each scale based on each first image feature and a first hand shape prior feature; wherein the first hand shape prior feature is determined based on the target image and a pre-constructed set of hand shape prior features; and determines the hand image segmentation result of the target image based on each first fusion feature and the first hand shape prior feature. Using this apparatus, by matching and adapting the first hand shape prior feature of the target image based on the set of hand shape prior features, and integrating the first hand shape prior feature into the entire image segmentation process, it achieves deep guidance of the segmentation process by prior knowledge, effectively solving the problems of low segmentation accuracy and poor generalization ability caused by the large influence of lighting and background in existing methods. This provides accurate hand region localization for subsequent intelligent health screening of hand signs such as nail abnormalities and rashes. Specifically, by obtaining the first image features of the target image at various scales, the different granular features of the hand shape can be comprehensively captured, taking into account both detailed and global features, laying the foundation for subsequent accurate segmentation. Based on the hand shape prior feature set, the first hand shape prior features of the target image are matched and adapted, and hand shape structure constraints are added to the first image features, effectively avoiding problems such as contour distortion and missing fingertips caused by illumination and background when segmenting hand images. Then, based on each first fusion feature and the first hand shape prior feature, the hand image segmentation result of the target image is determined, further ensuring that the hand image segmentation result conforms to the hand shape structure rules, avoiding segmentation deviations caused by background noise, uneven illumination, and other interferences, achieving accurate restoration of fine areas such as hand shape edges, fingertips, and finger gaps, improving the robustness, accuracy, and scene generalization ability of the segmentation result, and being able to meet the hand detection needs of different groups such as children and adults under various hand postures.

[0095] The image processing apparatus provided in the embodiments of the present invention can execute the image processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0096] Figure 3 A schematic diagram of an electronic device 30 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0097] like Figure 3 As shown, the electronic device 30 includes at least one processor 31 and a memory, such as a read-only memory (ROM) 32 or a random access memory (RAM) 33, communicatively connected to the at least one processor 31. The memory stores computer programs executable by the at least one processor. The processor 31 can perform various appropriate actions and processes based on the computer program stored in the ROM 32 or loaded from storage unit 38 into the RAM 33. The RAM 33 can also store various programs and data required for the operation of the electronic device 30. The processor 31, ROM 32, and RAM 33 are interconnected via a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.

[0098] Multiple components in electronic device 30 are connected to I / O interface 35, including: input unit 36, such as keyboard, mouse, etc.; output unit 37, such as various types of monitors, speakers, etc.; storage unit 38, such as disk, optical disk, etc.; and communication unit 39, such as network card, modem, wireless transceiver, etc. Communication unit 39 allows electronic device 30 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0099] Processor 31 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 31 performs the various methods and processes described above, such as image processing methods.

[0100] In some embodiments, the image processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 38. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 30 via ROM 32 and / or communication unit 39. When the computer program is loaded into RAM 33 and executed by processor 31, one or more steps of the image processing method described above may be performed. Alternatively, in other embodiments, processor 31 may be configured to perform the image processing method by any other suitable means (e.g., by means of firmware).

[0101] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0102] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0103] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0104] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0105] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0106] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0107] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0108] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An image processing method, characterized in that, include: Obtain the target image containing the hand image; Multi-scale feature extraction is performed on the target image to obtain the first image features of the target image at each scale, and the first fusion features corresponding to each scale are determined based on each first image feature and the first hand shape prior feature; wherein, the first hand shape prior feature is determined based on the target image and a pre-constructed set of hand shape prior features; Based on each of the first fusion features and the first hand shape prior features, the hand image segmentation result of the target image is determined.

2. The method according to claim 1, characterized in that, The step of performing multi-scale feature extraction on the target image to obtain the first image features of the target image at each scale, and determining the first fusion features corresponding to each scale based on each of the first image features and the first hand shape prior features, includes: The image segmentation model uses an encoder to extract features from the target image at a set scale based on multiple cascaded convolutional layers included in the encoder, thereby obtaining first image features at each scale. The convolutional layers include at least one convolutional layer for texture feature extraction and at least one convolutional layer for semantic feature extraction. The first hand shape prior features are input into the downsampling network structure contained in the prior feature embedding sub-model in the image segmentation model, and the first hand shape prior features are subjected to multi-scale downsampling processing to obtain second hand shape prior features that match the size of each of the first image features. Each of the first image features and each of the second hand shape prior features are input into the first feature fusion module included in the prior feature embedding sub-model. Each of the second hand shape prior features is fused with the first image features of the corresponding scale to obtain the first fused features corresponding to each scale.

3. The method according to claim 1, characterized in that, The first hand shape prior feature is determined based on the target image and a pre-constructed set of hand shape prior features, including: The target image is input into the pooling module of the prior feature selection sub-model in the image segmentation model. Global average pooling is performed on the target image to obtain the first feature vector. Global average pooling is then performed on each hand shape prior feature in the hand shape prior feature set to obtain the second feature vector corresponding to each hand shape prior feature. The first feature vector and each of the second feature vectors are input into the feature filtering module in the prior feature filtering sub-model to determine the cosine similarity value between the first feature vector and each of the second feature vectors, and the hand shape prior feature corresponding to the maximum cosine similarity value is determined as the first hand shape prior feature to adapt to the target image.

4. The method according to claim 1 or 3, characterized in that, The steps for constructing the hand shape prior feature set include: Obtain a set of hand sample images; For each first hand sample image in the hand sample image set, the global shape features and local contour features of the first hand sample image are extracted, and the global shape features and local contour features are encoded respectively to obtain global encoded features and local encoded features. The global encoded features and local encoded features are weighted and fused based on the fusion weight to generate the hand shape prior features of the first hand sample image. The hand shape prior features of each of the first hand sample images are integrated to generate the hand shape prior feature set.

5. The method according to claim 4, characterized in that, The global shape features include the circumscribed convex contour of the hand, the area ratio of the palm to the fingers, and the topological connection relationship of the five fingers. The local contour features include a set of fine contour points of the hand and a slender strip-shaped convex constraint of a single finger. The set of fine contour points of the hand includes fine contour points of the fingertip, finger gaps and palm edge.

6. The method according to claim 1, characterized in that, The step of determining the hand image segmentation result of the target image based on each of the first fusion features and the first hand shape prior features includes: By using the decoder and cross-layer feature fusion sub-model in the image segmentation model, the first decoding feature at each scale is determined based on each of the first fusion features; The first decoding features are modified according to the first hand shape prior features through the feature correction sub-model in the image segmentation model to obtain the second decoding features. The output sub-model of the image segmentation model is used to map the second decoding feature with the same size as the target image to obtain a target segmentation probability map based on the convolutional layer and the decision module included in the output sub-model. The hand image segmentation result of the target image is determined based on the preset probability threshold and the target segmentation probability map.

7. The method according to claim 6, characterized in that, The step of determining the first decoding features at each scale based on each of the first fused features through the decoder and cross-layer feature fusion sub-model in the image segmentation model includes: The first fusion feature with the smallest size among the first fusion features is input into the decoder. According to the upsampling network structure included in the decoder, the first fusion feature is upsampled layer by layer according to a set scale to enlarge the feature size and obtain upsampled features at each scale. Each of the first fusion features and each of the upsampled features are input into the cross-layer feature fusion sub-model. The cross-layer feature fusion sub-model concatenates each of the upsampled features with the first fusion features of the corresponding scale in the channel dimension to obtain each of the first decoding features.

8. The method according to claim 6, characterized in that, The step of correcting each of the first decoding features based on the first hand shape prior features to obtain each of the second decoding features includes: For each first decoded feature, a similarity value is determined between the first decoded feature and the corresponding scale of the second hand shape prior feature, which is obtained based on the first hand shape prior feature and the prior feature embedding sub-model in the image segmentation model. If the similarity value is less than a preset similarity threshold, the first decoding feature is corrected based on the similarity value and the second hand shape prior feature, and the corrected first decoding feature is used as the second decoding feature. If the similarity value is greater than or equal to the preset similarity threshold, then the first decoding feature is used as the second decoding feature.

9. The method according to any one of claims 2-3 and 6-8, characterized in that, The training steps of the image segmentation model include: Obtain a sample training set and an initial segmentation model, wherein the sample training set includes at least one second hand sample image and the corresponding label segmentation result; The second hand sample image is input into the initial segmentation model to obtain the current segmentation result output by the initial segmentation model and the third hand shape prior feature that matches the second hand sample image; Based on the current segmentation result, the label segmentation result, and the third hand shape prior feature, determine the first loss function value of the pixel loss function, the second loss function value of the contour loss function, and the third loss function value of the shape loss function; The target loss function value is determined based on the first loss function value, the second loss function value, and the third loss function value. Based on the target loss function value, the learnable network parameters in the initial segmentation model are back-learned and adjusted to obtain the adjusted initial segmentation model. Then, the relevant steps of inputting the second hand sample image into the initial segmentation model are re-executed until the training termination condition is met. The initial segmentation model obtained after training is determined as the image segmentation model.

10. The method according to claim 9, characterized in that, Also includes: Based on the convolution kernel parameters of the image segmentation model, the importance score of each convolutional layer channel is determined, and redundant channels with importance scores lower than the preset score threshold are pruned according to the preset score threshold to obtain the pruned image segmentation model. The pruned image segmentation model is subjected to weight quantization to obtain a lightweight image segmentation model.