Skin texture depth determination method, system, electronic device and storage medium based on image analysis

By acquiring and fusing multimodal images, and utilizing generative adversarial networks and cascaded classifiers, the shortcomings in accuracy and comprehensiveness of existing skin texture depth measurement methods have been addressed, achieving accurate measurement of skin texture depth.

CN121353364BActive Publication Date: 2026-04-10SHENG EN (BEIJING) PHARM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In the existing technology, the skin texture depth measurement method based on a single visible light image has insufficient accuracy and lacks deep features, resulting in insufficient comprehensiveness and reliability of the measurement results.

Method used

Visible light, near-infrared, and cross-polarized light images of the target object's skin surface are acquired, multimodal image fusion is performed, super-resolution reconstruction is carried out using a generative adversarial network, and feature selection and 3D analysis are performed using a cascaded classifier to obtain multi-dimensional skin features such as epidermal texture, subcutaneous blood vessels, and oil distribution.

Benefits of technology

By using multimodal image fusion and feature reconstruction, the accuracy and comprehensiveness of skin texture depth measurement are improved, the reliability of the measurement results is enhanced, and the impact of image quality issues is reduced by comprehensively considering the surface and deep features of the skin.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353364B_ABST
    Figure CN121353364B_ABST
Patent Text Reader

Abstract

The application provides a skin texture depth determination method and system based on image analysis, an electronic device and a storage medium, and relates to the technical field of image analysis. The application forms a multi-modal image by collecting a visible light image, a near-infrared image and a cross-polarized light image of the skin surface of a target object. The multi-modal image is spatially aligned, and the spatially aligned multi-modal image is fused to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels and oil distribution. Based on a generative adversarial network, the features of pixels with a quality value lower than a preset quality threshold in the initial composite feature map are super-resolution reconstructed to obtain a target composite feature map. The features in the target composite feature map are screened by a front-stage screening unit of a cascaded classifier, and based on the screened features, three-dimensional analysis is performed by a rear-stage analysis unit of the cascaded classifier to realize the determination of skin texture depth and improve the accuracy and comprehensiveness of the determination of skin texture depth.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image analysis, and in particular to a skin texture depth determination method and system based on image analysis, an electronic device and a storage medium. BACKGROUND

[0002] In the development of skin care products, clinical diagnosis of dermatology, and the formulation of personalized makeup programs, skin texture depth is a key indicator for evaluating the degree of skin aging, barrier function, and health status. These scenarios require precise determination of skin texture depth through non-invasive means combined with image analysis technology, and need to reflect multi-dimensional characteristics such as epidermal texture morphology, subcutaneous blood vessel distribution, and oil secretion to support more comprehensive skin condition assessment.

[0003] Currently, there is a skin texture depth determination scheme based on a single visible light image to meet the above needs. This scheme acquires a visible light image of the skin surface, extracts texture contour features from the image, combines a pre-set gray value and depth mapping relationship, calculates the relative depth of the raised and recessed areas of the texture, and then obtains the determination result of the skin texture depth.

[0004] This scheme relies only on visible light images for analysis, which has obvious limitations: on the one hand, visible light images are easily affected by factors such as skin surface reflection, keratin layer state, etc., leading to inaccurate texture feature extraction in low-quality areas (such as blurred, reflective areas), which in turn affects the depth determination accuracy; on the other hand, it does not consider deep features related to skin texture depth such as subcutaneous blood vessel distribution and oil secretion, and only calculates the depth by the gray level change of the epidermal texture, making it difficult to reflect the correlation between the texture and the subcutaneous tissue, resulting in insufficient comprehensiveness and reliability of the determination result. SUMMARY

[0005] The present application aims to provide a skin texture depth determination method and system based on image analysis, an electronic device and a storage medium to solve the problem of low accuracy and comprehensiveness of skin texture depth determination caused by insufficient accuracy, lack of deep features, and insufficient comprehensiveness and reliability in the prior art.

[0006] To solve the above technical problems, in a first aspect, the present application provides a skin texture depth determination method based on image analysis, comprising:

[0007] Acquiring a visible light image, a near-infrared image, and a cross-polarized light image of the skin surface of a target object to form a multi-modal image;

[0008] Spatially aligning the multi-modal image and performing fusion processing on the spatially aligned multi-modal image to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels, and oil distribution;

[0009] reconstructing, based on the generative adversarial network, features of pixels in the initial composite feature map with a quality value lower than a preset quality threshold, to obtain a target composite feature map;

[0010] screening features in the target composite feature map through a front-stage screening unit of the cascade classifier, and performing three-dimensional analysis on the screened features through a rear-stage analysis unit of the cascade classifier to determine the skin texture depth.

[0011] Optionally, the reconstructing, based on the generative adversarial network, features of pixels in the initial composite feature map with a quality value lower than a preset quality threshold, to obtain a target composite feature map, comprises:

[0012] regarding a continuous region composed of pixels in the initial composite feature map with a quality value lower than a preset quality threshold as a low-quality region;

[0013] extending features of the low-quality region through a generative module of the generative adversarial network to generate a candidate reconstruction region;

[0014] comparing, through a discriminative module of the generative adversarial network, features of the candidate reconstruction region with features of a surrounding region formed by pixels in the initial composite feature map with a quality value not lower than a preset quality threshold, and outputting a comparison result;

[0015] adjusting features of the candidate reconstruction region according to the comparison result to obtain an adaptive reconstruction region;

[0016] replacing the adaptive reconstruction region with the low-quality region to obtain a target composite feature map.

[0017] Optionally, the extending features of the low-quality region through a generative module of the generative adversarial network to generate a candidate reconstruction region, comprises:

[0018] associating features of the low-quality region with features of the surrounding region to determine a feature correspondence relationship;

[0019] complementing features of the low-quality region according to the feature correspondence relationship to increase the number and details of features of the low-quality region;

[0020] integrating the complemented features of the low-quality region to form a candidate reconstruction region with a size consistent with that of the low-quality region.

[0021] Optionally, the spatially aligning the multi-modal images and fusing the spatially aligned multi-modal images to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels and oil distribution, comprises:

[0022] extracting salient features of each image in the multi-modal images, wherein the salient feature of the visible light image is epidermal texture, the salient feature of the near-infrared image is subcutaneous blood vessel, and the salient feature of the cross-polarized light image is oil distribution;

[0023] establishing a positional relationship of features among the multi-modal images based on the epidermal texture, the subcutaneous blood vessel and the oil distribution;

[0024] adjusting positions of each image in the multi-modal images according to the positional relationship, so as to spatially align the images;

[0025] performing fusion processing on the spatially aligned images to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessel and oil distribution.

[0026] Optionally, the establishing of the positional relationship of features among the multi-modal images based on the epidermal texture, the subcutaneous blood vessel and the oil distribution comprises:

[0027] selecting a texture intersection point from the epidermal texture as a first reference point, selecting a blood vessel branch point from the subcutaneous blood vessel as a second reference point, and selecting an oil distribution boundary point from the oil distribution as a third reference point;

[0028] comparing positions of the first reference point, the second reference point and the third reference point in each image, calculating spatial distances between each reference point in different images, and screening reference points with spatial distances less than a preset distance threshold, wherein the reference points from the visible light image, the near-infrared image and the cross-polarized light image respectively are taken as a group of corresponding points;

[0029] establishing the positional relationship of features among the multi-modal images based on each group of corresponding points.

[0030] Optionally, the screening of features in the target composite feature map by the front-stage screening unit of the cascade classifier, and the three-dimensional analysis of the screened features by the back-stage analysis unit of the cascade classifier to realize the determination of skin texture depth, comprises:

[0031] determining, by the front-stage screening unit of the cascade classifier, a feature type related to skin texture depth in the target composite feature map, wherein the feature type comprises raised and recessed features of epidermal texture, distribution features of subcutaneous blood vessels under the texture, and connection features of oil distribution and texture edges;

[0032] integrating the raised and recessed features, the distribution features and the connection features to generate a set of features to be analyzed, and taking the set of features to be analyzed as the screened features;

[0033] The features in the feature set to be analyzed are layered according to spatial positions by a back-stage analysis unit of the cascade classifier to distinguish skin surface texture features and associated features located below the skin surface;

[0034] The distance between the skin surface texture features and the corresponding associated features is taken as a measurement result of skin texture depth.

[0035] Optionally, the integration of the convex and concave features, the distribution features and the connection features generates a feature set to be analyzed, including:

[0036] The convex and concave features, the distribution features and the connection features at the same spatial position are identified and associated to form a feature group;

[0037] All the feature groups corresponding to the spatial positions are summarized to generate the feature set to be analyzed.

[0038] In a second aspect, the present application provides a skin texture depth measurement system based on image analysis, including:

[0039] A collection module is configured to collect a visible light image, a near-infrared image and a cross-polarized light image of a skin surface of a target object to form a multi-modal image;

[0040] A fusion module is configured to perform spatial alignment on the multi-modal image and perform fusion processing on the spatially aligned multi-modal image to obtain an initial composite feature map containing epidermis texture, subcutaneous blood vessels and oil distribution;

[0041] A reconstruction module is configured to perform super-resolution reconstruction on features of pixels with a quality value lower than a preset quality threshold in the initial composite feature map based on a generative adversarial network to obtain a target composite feature map;

[0042] A screening module is configured to screen features in the target composite feature map through a front-stage screening unit of a cascade classifier, and perform three-dimensional analysis on the screened features through a back-stage analysis unit of the cascade classifier to realize measurement of skin texture depth.

[0043] In a third aspect, the present application provides an electronic device, including:

[0044] A memory is configured to store a computer program;

[0045] A processor is configured to execute the computer program to realize steps of the skin texture depth measurement method based on image analysis according to the first aspect.

[0046] In a fourth aspect, the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program, when executed by a processor, enables the steps of the skin texture depth determination method based on image analysis according to the first aspect.

[0047] In the present application, a skin texture depth determination method based on image analysis is provided, which comprises: acquiring a visible light image, a near-infrared image and a cross-polarized light image of a skin surface of a target object to form multi-modal images; performing spatial alignment on the multi-modal images, and performing fusion processing on the spatially aligned multi-modal images to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels and oil distribution; performing super-resolution reconstruction on the features of pixels with a quality value lower than a preset quality threshold in the initial composite feature map based on a generative adversarial network to obtain a target composite feature map; performing screening on the features in the target composite feature map through a front-stage screening unit of a cascaded classifier, and performing three-dimensional analysis based on the screened features through a rear-stage analysis unit of the cascaded classifier to realize determination of the skin texture depth.

[0048] The skin texture depth determination method based on image analysis provided in the present application can obtain multi-dimensional skin feature information covering epidermal texture, subcutaneous blood vessels and oil distribution by acquiring a visible light image, a near-infrared image and a cross-polarized light image of a skin surface of a target object to form multi-modal images, thereby providing a data basis for subsequent comprehensive analysis; the spatial matching and integration of different modal features can be realized by performing spatial alignment and fusion processing on the multi-modal images to obtain an initial composite feature map, thereby forming a unified feature map containing various skin features; the features of low-quality pixels in the initial composite feature map can be repaired by performing super-resolution reconstruction based on a generative adversarial network to obtain a target composite feature map, thereby improving the overall quality and detail integrity of the composite feature map; the key features related to the texture depth can be accurately extracted by screening the features through a cascaded classifier and performing three-dimensional analysis to realize determination of the skin texture depth, thereby improving the accuracy of the skin texture depth determination.

[0049] Further, after identifying the low-quality region in the initial composite feature map, the generative module of the generative adversarial network associates the features of the low-quality region and the surrounding high-quality region to determine the corresponding relationship, supplements and integrates the features to generate a candidate reconstruction region, and the discriminative module adjusts the candidate region and the surrounding region features to obtain an adaptive reconstruction region, and replaces the low-quality region to obtain the target composite feature map; through the synergistic effect of the generative module and the discriminative module, the low-quality region is accurately repaired in combination with the surrounding high-quality features, the coordination between the reconstruction region and the surrounding features is improved, and the quality and consistency of the target composite feature map are enhanced. BRIEF DESCRIPTION OF DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the accompanying drawings required to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings described below are only some embodiments of the present application, and all other embodiments obtained by a person of ordinary skill in the art without creative labor on the basis of the embodiments in the present application shall fall within the scope of protection of the present application.

[0051] Figure 1 A flowchart of a skin texture depth determination method based on image analysis provided by an embodiment of the present application is shown in the figure.

[0052] Figure 2 A specific implementation diagram of a skin texture depth determination method based on image analysis provided by an embodiment of the present application is shown in the figure.

[0053] Figure 3 A structure diagram of a skin texture depth determination system based on image analysis provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0054] In order to solve the problems of insufficient precision, lack of deep features, and insufficient comprehensiveness and reliability in the prior art, an embodiment of the present application provides a skin texture depth determination method based on image analysis, which adopts the following design concept: first, visible light images, near-infrared images, and cross-polarized light images of the skin surface are collected, and various information such as skin surface texture, subcutaneous blood vessels, and oil distribution is obtained through these different images; then, the images are adjusted to appropriate positions and fused together to form a comprehensive image containing multiple information; next, for the part of the image that is not clear enough, special techniques are used to repair it, so that the overall image is clearer and the details are more complete; finally, special tools are used to select useful information related to texture depth from the comprehensive image, and then the information is analyzed in depth to determine the depth of the skin texture. In this way, the relevant information of the deep skin is included, the defects in the image are repaired, and the pertinence and accuracy of the results are ensured through screening and analysis, effectively solving the problems existing in the prior art.

[0055] In order to make the person in the technical field better understand the present application, the present application will be further described in detail below in combination with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor shall fall within the scope of protection of the present application.

[0056] The core of the present application is to provide a skin texture depth determination method based on image analysis, and a flowchart of a specific implementation thereof is shown in Figure 1 The method comprises:

[0057] S11, collect a visible light image, a near-infrared image and a cross-polarized light image of a skin surface of a target object to form a multi-modal image.

[0058] The target object is an individual whose skin texture depth needs to be determined; the visible light image is an image of the skin surface taken by ordinary light, which can present surface features such as epidermal texture; the near-infrared image is an image taken by near-infrared light, which can show deeper features such as subcutaneous blood vessels; the cross-polarized light image is an image taken by cross-polarized light technology, which can reflect the distribution of skin oil; and the multi-modal image is a set of images formed by combining the above three images.

[0059] In the embodiments of the present application, first, a target object is determined, then visible light images, near-infrared images and cross-polarized light images of the same skin region of the object are taken by corresponding devices, and finally the three images are integrated into a multi-modal image. For example, a volunteer D is selected as the target object, and visible light images, near-infrared images and cross-polarized light images of the same position on the cheeks of the volunteer D are taken to record epidermal texture, present subcutaneous blood vessels and reflect oil distribution, and then the three images are integrated into a multi-modal image.

[0060] S12, spatially align the multi-modal image, and fuse the spatially aligned multi-modal image to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels and oil distribution.

[0061] The spatial alignment is to adjust the positions of the images in the multi-modal image so that they correspond to the same skin region; the fusion processing is to combine the spatially aligned images into one image; and the initial composite feature map is an image containing epidermal texture, subcutaneous blood vessels and oil distribution after fusion.

[0062] In the embodiments of the present application, the multi-modal image obtained in S11 is first spatially aligned, the positions of the images are adjusted by recognizing the same skin feature points (such as obvious pores) so that the same skin position corresponds in the three images; then the aligned images are fused to combine the feature information at the corresponding positions into one image. For example, in the multi-modal image of the cheeks of the volunteer D, the positions are adjusted with three obvious pores as feature points, and then the epidermal texture, subcutaneous blood vessels and oil distribution are superimposed at the corresponding positions to obtain an initial composite feature map.

[0063] S13, based on a generative adversarial network, performing super-resolution reconstruction on the features of pixels with a quality value lower than a preset quality threshold in the initial composite feature map to obtain a target composite feature map.

[0064] The generative adversarial network is an image processing technology including a generative module and a discriminative module; the preset quality threshold is a standard value for judging pixel quality; the pixel is a basic unit of an image; the super-resolution reconstruction is to optimize a low-quality part of the image to make it clearer; and the target composite feature map is a composite feature map with higher quality after reconstruction.

[0065] In the embodiment of the present application, first, a preset quality threshold is set, and the pixel quality value in the initial composite feature map is checked, and the continuous pixel region below the threshold is defined as a low-quality region; then the generative module supplements the low-quality region features to generate a candidate reconstruction region by referring to the features of the surrounding high-quality region; the discriminative module compares the candidate region with the surrounding features and outputs the results, and accordingly adjusts the candidate region to obtain an adaptive reconstruction region; finally, the adaptive reconstruction region is used to replace the low-quality region to obtain the target composite feature map. For example, if the threshold is set to 80 (the quality value is 0 to 100), the pixel quality value of a certain region in the cheek initial map of the volunteer D is 60 to 70, the generative module supplements the details by referring to the regions with 85 to 90, and the discriminative module adjusts to obtain an adaptive region, and after replacement, the target composite feature map is formed.

[0066] S14, screening features in the target composite feature map through the front-stage screening unit of the cascade classifier, and based on the screened features, performing three-dimensional analysis through the back-stage analysis unit of the cascade classifier to realize the determination of the skin texture depth.

[0067] The cascade classifier is a feature processing tool including a front-stage screening unit and a back-stage analysis unit; the front-stage screening unit is used to select features related to the skin texture depth; and the back-stage analysis unit is used to in-depth analyze the selected features; and the three-dimensional analysis is to determine the texture depth by analyzing the spatial distribution of the features.

[0068] In the embodiment of the present application, first, the front-stage screening unit selects the features related to the texture depth from the target composite feature map, and excludes irrelevant features; and then the back-stage analysis unit analyzes the spatial hierarchical relationship of these features to determine the distance between the epidermal texture and the subcutaneous related features, and realizes the depth determination. For example, in the target map of the cheek of the volunteer D, the front-stage screening unit selects the texture protrusion / depression and the corresponding subcutaneous blood vessels and oil interface features, and the back-stage analysis unit analyzes the spatial positions of these features to determine the texture depth.

[0069] As shown in Figure 2 , the embodiment of the present application provides a specific implementation schematic diagram of a skin texture depth determination method based on image analysis, based on Figure 1 and Figure 2In the content, the present application provides the following specific examples: researchers select volunteer D as the target object, first take visible light images, near-infrared images and cross-polarized light images of the same position of the cheeks of the volunteer D, integrate into multi-modal images; then take three obvious pores as feature points, adjust the positions of the three images to realize spatial alignment, and then superimpose the epidermal texture, subcutaneous blood vessels and oil distribution at the corresponding positions to obtain an initial composite feature map; then set the preset quality threshold to 80, find that the pixel quality value of a region in the initial map is 60 to 70, determine that it is a low-quality region, generate a generation module of the generative adversarial network to supplement the details of the region to generate a candidate reconstruction region, the discriminant module adjusts to obtain an adaptive region after comparison, and the target composite feature map is formed after replacement; finally, the front stage of the cascade classifier screens out the texture protrusions / depressions, corresponding subcutaneous blood vessel distribution and oil joint features of the region, and the rear stage analyzes the spatial hierarchical relationship of these features to determine the texture depth of the cheek skin of volunteer D.

[0070] By performing S11-S14, the embodiment of the present application acquires skin multi-dimensional features by collecting multi-modal images, forms an integrated initial composite feature map through spatial alignment and fusion, then repairs low-quality regions through super-resolution reconstruction to improve image quality and details, and finally screens key features through a cascade classifier and analyzes them in three dimensions to realize accurate determination of skin texture depth. The entire process comprehensively considers the surface and deep features of the skin, reduces the influence of image quality problems, improves the comprehensiveness and reliability of the determination results, and provides an effective basis for skin state evaluation.

[0071] In a possible embodiment, S13, based on the generative adversarial network, performing super-resolution reconstruction on the features of pixels with a quality value lower than the preset quality threshold in the initial composite feature map to obtain a target composite feature map, comprising:

[0072] Step 131, regarding a continuous region composed of pixels with a quality value lower than the preset quality threshold in the initial composite feature map as a low-quality region.

[0073] Wherein, the initial composite feature map is an integrated image containing epidermal texture, subcutaneous blood vessels and oil distribution, the quality value is an index for measuring the clarity of pixels and the integrity of features in the image, the preset quality threshold is a standard value set by a person to judge whether the quality of pixels meets the standard, the pixel is a basic unit of the image, the continuous region is a region formed by a plurality of adjacent pixels, and the low-quality region is a continuous region composed of pixels with a quality value lower than the preset quality threshold.

[0074] In the embodiments of the present application, after the initial composite feature map is obtained, the quality value of each pixel in the map is determined, for example, the numerical value corresponding to the clarity and feature completeness of each pixel is read by an image analysis tool, the quality value of each pixel is compared with a preset quality threshold, and the pixels with a quality value lower than the threshold are screened out, and it is observed whether these low-quality pixels are adjacent to each other to form a continuous region and are determined as a low-quality region, for example, in the initial composite feature map of the face of volunteer E, the preset quality threshold is set to 70 (assuming that the quality value range is 0 to 100), and it is found by analysis that there is a region of pixels in the cheek with a quality value mostly between 50 and 60 and connected to each other to form a piece, and the region is determined as a low-quality region.

[0075] In step 132, the feature of the low-quality region is expanded by the generation module of the generative adversarial network to generate a candidate reconstruction region.

[0076] In the embodiments of the present application, the generation module of the generative adversarial network is a processing part for supplementing and perfecting the features of the image, the feature of the low-quality region is information such as the epidermal texture, subcutaneous blood vessels or oil distribution contained in the region, the expansion processing is an operation of supplementing feature details and enriching feature content, and the candidate reconstruction region is a preliminary repair region generated after the expansion processing and used to replace the low-quality region.

[0077] In the embodiments of the present application, the generation module first obtains the existing features of the low-quality region, for example, the blurred epidermal texture direction or the incomplete subcutaneous blood vessel fragment in the region, then refers to the features of the region (i.e., the high-quality region) around the low-quality region in the initial composite feature map, for example, the clear texture form and complete blood vessel distribution around the low-quality region, and finally supplements and perfects the features of the low-quality region according to the features of the high-quality region around the low-quality region, increases the detailed information, and generates a candidate reconstruction region consistent in size with the low-quality region, for example, the epidermal texture of the high-quality region around the low-quality region of the face of volunteer E is radial, and the subcutaneous blood vessels have obvious branches, and the generation module supplements the radial texture details and complete blood vessel branches for the low-quality region according to these features to form a candidate reconstruction region.

[0078] In step 133, the feature of the candidate reconstruction region is compared with the feature of the surrounding region formed by the pixels with a quality value not lower than the preset quality threshold in the initial composite feature map by the discriminant module of the generative adversarial network, and an output is output.

[0079] The discrimination module of the generative adversarial network is a processing part for judging whether the image features are coordinated and consistent, the features of the candidate reconstruction region are information such as epidermal texture, subcutaneous blood vessels and oil distribution contained in the region, the surrounding region is a region in the initial composite feature map near the low-quality region composed of pixels with a quality value not lower than a preset quality threshold, the comparison is a comparison of the consistency of different region features, and the comparison result is the judgment information output by the discrimination module about whether the features of the candidate reconstruction region and the surrounding region are coordinated.

[0080] In the embodiments of the present application, the discrimination module first extracts the features of the candidate reconstruction region, such as the direction of the epidermal texture, the distribution density of the subcutaneous blood vessels, and the edge shape of the oil distribution, then extracts the features of the surrounding region to obtain the corresponding texture direction, blood vessel density and oil edge features in the region, and finally compares the two types of features to judge whether they are coordinated in terms of texture continuity, blood vessel distribution matching degree, oil edge connection, etc. and outputs specific comparison results, for example, in the case of volunteer E, after comparing the candidate reconstruction region with the surrounding region, the discrimination module found that the epidermal texture direction of the candidate region deviated from the reticular texture of the surrounding region and the oil edge connection was unnatural, and outputted specific information about these inconsistencies.

[0081] Step 134, adjusting the features of the candidate reconstruction region according to the comparison result to obtain an adaptive reconstruction region.

[0082] The comparison result is the information output by the discrimination module about whether the features of the candidate reconstruction region and the surrounding region are coordinated, the adjustment is to modify the features of the candidate reconstruction region according to the comparison result to make them more coordinated, and the adaptive reconstruction region is a reconstruction region that is coordinated and consistent with the features of the surrounding region after adjustment.

[0083] In the embodiments of the present application, the comparison result output by the discrimination module is first obtained to determine in which features the candidate reconstruction region and the surrounding region are inconsistent, such as texture direction deviation, blood vessel distribution mismatch, etc. Then, the inconsistent features are modified, the corresponding part of the candidate reconstruction region is adjusted with reference to the features of the surrounding region, such as adjusting the texture direction to make it consistent with the surrounding, supplementing the blood vessel branches to make their density match the surrounding, and finally obtaining an adaptive reconstruction region that is coordinated with the features of the surrounding region after multiple fine-tuning, for example, in the case of volunteer E, the epidermal texture of the candidate reconstruction region is adjusted to be a close reticular texture, the number of subcutaneous blood vessel branches is increased, and the oil edge is smoothed to make it consistent with the features of the surrounding region, forming an adaptive reconstruction region.

[0084] Step 135, replacing the low-quality region with the adaptive reconstruction region to obtain a target composite feature map.

[0085] Wherein, the adaptive reconstruction region is a repaired region coordinated with the characteristics of the surrounding region, the low-quality region is a continuous region of substandard quality in the initial composite feature map, and the target composite feature map is a composite feature map with higher overall quality obtained by replacing the low-quality region with the adaptive reconstruction region.

[0086] In the embodiments of the present application, the position and range of the low-quality region in the initial composite feature map are first determined, and then the adaptive reconstruction region is overlaid in the initial composite feature map according to the position and range to replace the original low-quality region. Finally, the integrated image after replacement ensures that the adaptive reconstruction region is naturally connected with the surrounding region, forming a target composite feature map. For example, in the initial composite feature map of the face of volunteer E, the adaptive reconstruction region at the nasal ala is accurately overlaid at the position of the original low-quality region, and after integration, the texture, blood vessels and oil distribution in this region are coordinated with the surrounding region, obtaining a clear and complete target composite feature map.

[0087] The present application provides the following specific examples: when the technical personnel process the initial composite feature map of the face of volunteer E, the quality value of each pixel is obtained through the image tool, which is compared with the preset 70, and the pixels with a quality value lower than 70 are screened out. It is found that these pixels form a continuous region near the nasal ala, which is determined as a low-quality region. The generation module extracts the blurred oil edge and epidermal texture of the region, and refers to the fine mesh texture and smooth oil edge of the surrounding high-quality region to supplement the details to generate a candidate reconstruction region. After comparison, the discrimination module finds that the candidate region has loose texture, few blood vessel branches, and oil edge burrs, which do not match the surrounding tight mesh texture, dense blood vessel branches and smooth edge, and outputs the comparison result. The technical personnel adjust accordingly: change the texture to tight mesh, supplement 3 blood vessel branches, and eliminate the oil edge burrs to form an adaptive reconstruction region. Finally, the adaptive region is overlaid in the low-quality region according to the original position and range, and the target composite feature map of the face of volunteer E is obtained after integration, and the characteristics of the region are naturally connected with the surrounding region.

[0088] By performing steps 131-135, the embodiments of the present application accurately locate the low-quality region, generate and optimize the reconstruction region based on the surrounding high-quality features, and finally complete the image repair. This process clearly defines the repair target, supplements the feature details, ensures the coordination of the repair region with the surrounding environment, and improves the overall quality and feature integrity of the image, providing a clear and reliable image basis for subsequent skin texture depth measurement.

[0089] In one possible embodiment, step 132, the generation module of the generative adversarial network performs expansion processing on the features of the low-quality region to generate a candidate reconstruction region, including:

[0090] a1, associate the features of the low-quality region with the features of the surrounding region to determine the feature correspondence relationship.

[0091] Wherein, the feature of the low-quality area is that the information such as epidermis texture, subcutaneous blood vessel, oil distribution contained in the area is possibly blurred or incomplete, the feature of the surrounding area is that the information such as epidermis texture, subcutaneous blood vessel, oil distribution in the high-quality area near the low-quality area is clear and complete, the correlation is to compare and connect the features of the low-quality area and the surrounding area, and the feature correspondence is to determine the features in the low-quality area and the features in the surrounding area which have matching relationship in aspects such as property and position.

[0092] In the embodiment of the present application, the features of the low-quality area are extracted, such as the blurred epidermis texture, the incomplete subcutaneous blood vessel, and the unclear oil distribution edge, and then the features of the surrounding area are extracted, such as the clear epidermis texture, the complete subcutaneous blood vessel, and the clear oil distribution edge, and then the two types of features are compared to find out their matching relationship in aspects such as texture, blood vessel distribution, and oil edge shape, so as to determine the feature correspondence, for example, in the low-quality area of the volunteer E's ala nasi, there is a blurred horizontal epidermis texture, and the surrounding area has clear horizontal epidermis texture, and through comparison, it is found that the two have consistent direction and position connection, so it is determined as the corresponding feature.

[0093] a2, according to the feature correspondence, supplement the features of the low-quality area to increase the number and details of the features of the low-quality area.

[0094] Wherein, the feature correspondence is the matching relationship between the features of the low-quality area and the features of the surrounding area determined in the step a1, the supplement is to add the missing feature information to the low-quality area according to the correspondence, the number of features is how many information such as epidermis texture, subcutaneous blood vessel, and oil distribution is contained in the low-quality area, and the details of the features are detailed information such as specific shape, direction, and distribution of the features.

[0095] In the embodiment of the present application, according to the feature correspondence determined in the step a1, it is clear which features in the low-quality area are missing or not detailed enough, such as the epidermis texture lacking branches and the incomplete subcutaneous blood vessel, and then the features of the low-quality area are supplemented to increase the number and details by referring to the corresponding features in the surrounding area, such as adding corresponding number of branches to the epidermis texture of the low-quality area according to the branch density of the epidermis texture in the surrounding area, and supplementing the missing blood vessel fragments in the low-quality area according to the direction of the complete blood vessel in the surrounding area, for example, in the low-quality area of the volunteer E's ala nasi, 2 branches per millimeter are supplemented to the area according to the condition that there are 3 epidermis texture branches per millimeter in the surrounding area, and 2 missing blood vessels are supplemented according to the direction of the surrounding blood vessel, so that the features are more complete.

[0096] a3, integrating the supplemented features of the low-quality area to form a candidate reconstruction area consistent with the size of the low-quality area.

[0097] Wherein, the supplemented features are richer and more complete epidermal texture, subcutaneous blood vessels, oil distribution and other information in the low-quality area after the a2 step, the integration is to arrange and combine these features to make them coordinate with each other, and the candidate reconstruction area is the preliminary repair area formed after the integration, which is the same size as the low-quality area and is used to replace the low-quality area.

[0098] In the embodiments of the present application, all the features supplemented in the a2 step are collected, including the supplemented epidermal texture, subcutaneous blood vessels, oil distribution and the like, and then these features are arranged to adjust their positions and shapes to ensure that they coordinate with each other, for example, the direction of the epidermal texture is not in conflict with the distribution of the subcutaneous blood vessels, and the edge of the oil distribution is naturally connected with the edge of the epidermal texture. Finally, the integrated features are combined together to form a candidate reconstruction area which is the same size as the low-quality area. For example, in the processing of the low-quality area of the volunteer E's ala nasi, the supplemented oblique epidermal texture, complete subcutaneous blood vessels and smooth oil edge are adjusted in position so that the texture branches do not cover important blood vessels, and the oil edge is aligned with the texture edge, and then combined into a candidate reconstruction area which is the same size as the low-quality area.

[0099] The present application provides the following specific examples: when the technician processes the low-quality area of the volunteer E's ala nasi, the blurred oblique epidermal texture and the intermittent subcutaneous blood vessels in the area are extracted, and the clear oblique epidermal texture and the continuous subcutaneous blood vessels in the surrounding high-quality area are extracted. By comparison, it is found that the directions of both are 45 degrees and the blood vessel end points can be naturally connected, and the feature correspondence relationship is determined. Then, according to the correspondence relationship, the density of 3 epidermal texture branches per millimeter in the surrounding area is referred to, 2 branches per millimeter are supplemented in the low-quality area, 2 missing blood vessels are supplemented according to the direction of the surrounding blood vessels, and 3 oil edges are supplemented according to the curvature and 2 millimeter length of the surrounding oil edge. Finally, the supplemented features are collected, the positions of the epidermal texture branches are adjusted to avoid the key nodes of the blood vessels, the distance between the oil edge and the texture edge is adjusted to 0.2 millimeters to achieve natural connection, and finally a candidate reconstruction area which is the same size as the low-quality area is formed by combination.

[0100] By performing a1-a3, the embodiments of the present application clarify the matching relationship by associating the features of the low-quality area with the surrounding area, and provide a basis for feature supplementation. After supplementing the features based on the correspondence relationship, the number and details of the features in the low-quality area are increased, making them richer and more complete. Finally, the candidate reconstruction area formed by integration is coordinated and the same size as the low-quality area, which lays a reliable foundation for further optimization of subsequent image repair and ensures the adaptability of the repair area to the surrounding environment.

[0101] In a possible embodiment, S12, the multi-modal images are spatially aligned, and the spatially aligned multi-modal images are fused to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels and oil distribution, including:

[0102] Step 121, extracting the salient features of each image in the multi-modal images, wherein the salient features of the visible light image are epidermal texture, the salient features of the near-infrared image are subcutaneous blood vessels, and the salient features of the cross-polarized light image are oil distribution.

[0103] Wherein, the multi-modal images are a set of images composed of visible light images, near-infrared images and cross-polarized light images of the skin surface of the target object, the salient features are the information in the image that best represents its own characteristics and is related to the skin state, the epidermal texture is the morphological characteristics such as the skin surface texture and groove, and is the salient feature of the visible light image, the subcutaneous blood vessels are the blood vessel distribution under the skin surface, and are the salient feature of the near-infrared image, the oil distribution is the distribution state of the skin surface oil, and is the salient feature of the cross-polarized light image, and the extraction is the operation of finding and separating these salient features from the image.

[0104] In the embodiments of the present application, multi-modal images are obtained, including visible light images, near-infrared images and cross-polarized light images, for example, the multi-modal images of the face of volunteer F, for the visible light image, the epidermal texture is extracted by identifying the skin surface texture, groove and other morphologies, such as identifying the fine texture of the cheeks and the obvious groove of the forehead from the visible light image of the face of volunteer F as the epidermal texture, for the near-infrared image, the subcutaneous blood vessels are extracted by identifying the blood vessel morphology and distribution, such as finding the blood vessel direction and distribution on both sides of the nose bridge from the near-infrared image of the face of volunteer F as the subcutaneous blood vessel feature, and for the cross-polarized light image, the oil distribution is extracted by identifying the oil aggregation area and range, such as determining the T-zone oil aggregation area from the cross-polarized light image of the face of volunteer F as the oil distribution feature.

[0105] Step 122, establishing the positional relationship of the features between the multi-modal images based on the epidermal texture, subcutaneous blood vessels and oil distribution.

[0106] Wherein, the epidermal texture, subcutaneous blood vessels and oil distribution are the three salient features extracted in step 121, the positional relationship is the corresponding relationship of the three features in space, that is, the epidermal texture at a certain position corresponds to the subcutaneous blood vessels below and the oil distribution on the surface, and the establishment is the process of determining this spatial corresponding relationship.

[0107] In the embodiments of the present application, a specific point on the epidermal texture is selected as a reference point, such as the intersection of a certain line, for example, the intersection of two lines on the cheek of the epidermal texture of the face of the volunteer F is selected as the reference point A, the position of the subcutaneous blood vessel corresponding to the position of the reference point A in the near-infrared image is found, such as the branch point of the blood vessel directly below the intersection point is recorded as the reference point B, the position of the oil distribution corresponding to the position of the reference point A in the cross-polarized light image is found, such as the oil edge point around the intersection point is recorded as the reference point C, and the spatial position relationship of the three features is determined through a plurality of such corresponding reference points, that is, the specific positions of the subcutaneous blood vessel and the oil distribution corresponding to the epidermal texture at a certain position are determined, for example, through the corresponding relationship of the reference points A, B and C, it is determined that the epidermal texture intersection point below the region is a specific blood vessel branch and the surrounding is a specific oil edge.

[0108] Step 123, adjusting the positions of the images in the multi-modal images according to the position relationship to make the images spatially aligned.

[0109] Wherein, the position relationship is the spatial correspondence of the three features established in step 122, the images in the multi-modal images refer to the visible light image, the near-infrared image and the cross-polarized light image, the adjustment is to move or rotate the image to change its spatial position, and the spatial alignment is that the significant features in the adjusted images correspond to each other in spatial position, that is, the three features at the same position are at the same coordinate position in each image.

[0110] In the embodiments of the present application, the difference between the current position of the significant feature in each image and the target corresponding position is determined according to the position relationship established in step 122, for example, the position of a certain blood vessel branch point in the near-infrared image of the volunteer F is 5 mm to the left of the position of the corresponding epidermal texture intersection point, and the near-infrared image and the cross-polarized light image are moved or rotated according to this position difference, for example, the near-infrared image is moved 5 mm to the right to make the blood vessel branch point consistent with the position of the corresponding epidermal texture intersection point, and the angle of the cross-polarized light image is adjusted to make the oil distribution boundary point aligned with the position of the corresponding epidermal texture intersection point, and the adjusted images are checked to ensure that the three significant features accurately correspond in space to achieve spatial alignment, for example, the same position of the epidermal texture, the subcutaneous blood vessel and the oil distribution in the adjusted images of the face of the volunteer F are at the same position.

[0111] Step 124, performing fusion processing on the spatially aligned images to obtain an initial composite feature map containing the epidermal texture, the subcutaneous blood vessel and the oil distribution.

[0112] The space-aligned images are visible light images, near-infrared images and cross-polarized light images whose significant feature space positions correspond to each other after step 123 adjustment, and the fusion processing is an operation of merging the three images into one image containing all significant features. The initial composite feature image is an image containing epidermal texture, subcutaneous blood vessels and oil distribution after the fusion processing.

[0113] In the embodiments of the present application, the space-aligned visible light images, near-infrared images and cross-polarized light images are obtained, for example, three images of the face of volunteer F after space alignment. The fusion processing is performed on the three images to superimpose the significant features at the corresponding positions of each image, for example, the epidermal texture in the visible light image, the subcutaneous blood vessels in the near-infrared image and the oil distribution in the cross-polarized light image are superimposed at the same coordinate position to make the position simultaneously present the three features. The initial composite feature image containing all significant features is formed after superposition and integration, for example, the same position in the initial composite feature image of the face of volunteer F can see the epidermal texture, the subcutaneous blood vessels below and the oil distribution around.

[0114] The present application provides the following specific examples: when the technician processes the multi-modal images of the face of volunteer F, the eye corner fine lines and the lower jaw lines are extracted from the visible light image as epidermal texture, the blood vessel branches on both sides of the cheeks are extracted from the near-infrared image as subcutaneous blood vessels, and the oil accumulation areas of the forehead and the alae nasi are extracted from the cross-polarized light image as oil distribution. Then, taking one intersection point of the lines on the forehead as a reference point X, the bifurcation point of the blood vessels directly below the point in the near-infrared image is found as a corresponding point Y, and the oil distribution boundary point around the point in the cross-polarized light image is found as a corresponding point Z, thereby establishing the positional relationship of the three features through a plurality of similar X, Y, Z corresponding points. Then, according to the positional relationship, it is found that a certain blood vessel point in the near-infrared image is 3 mm above the corresponding epidermal texture point, and a certain oil point in the cross-polarized light image is 2 mm to the right of the corresponding epidermal texture point, so the near-infrared image is moved down by 3 mm and the cross-polarized light image is moved left by 2 mm to align the feature positions of the images. Finally, the three space-aligned images are fused to superimpose the eye corner fine lines, the blood vessel distribution below the eye corner and the oil distribution around the eye corner at the corresponding positions, thereby obtaining the initial composite feature image containing these features.

[0115] By performing steps 121-124, the embodiments of the present application lay a foundation for focusing on key information for subsequent processing by extracting significant features of multi-modal images, make clear the spatial correspondence of different features by establishing the positional relationship between features, ensure accurate matching of features in position by adjusting images to achieve space alignment, and provide comprehensive and coordinated feature information for subsequent skin texture depth measurement by forming an initial composite feature image through fusion processing to integrate various skin features in the same image.

[0116] In a possible embodiment, step 122, based on the epidermal texture, subcutaneous blood vessels and oil distribution, establishes the positional relationship of features between multi-modal images, including:

[0117] b1, selecting a texture intersection point from the epidermal texture as a first reference point, selecting a blood vessel branch point from the subcutaneous blood vessels as a second reference point, and selecting an oil distribution boundary point from the oil distribution as a third reference point.

[0118] Wherein, the epidermal texture is the morphology of the skin surface such as texture and groove, the texture intersection point is the point formed by the intersection of different textures in the epidermal texture, the first reference point is the texture intersection point selected from the epidermal texture for positioning; the subcutaneous blood vessels are the distribution of blood vessels under the skin surface, the blood vessel branch point is the point formed by the bifurcation of blood vessels in the subcutaneous blood vessels, and the second reference point is the blood vessel branch point selected from the subcutaneous blood vessels for positioning; the oil distribution is the distribution range of oil on the skin surface, the oil distribution boundary point is the point at the edge of the oil distribution area, and the third reference point is the boundary point selected from the oil distribution for positioning.

[0119] In the embodiments of the present application, the epidermal texture is observed to find the texture intersection points formed by the intersection of textures, and these points are taken as the first reference points. For example, in the epidermal texture of the face of volunteer G, two points at which three textures intersect on the cheeks are selected as the first reference points a1 and a2. The subcutaneous blood vessels are observed to find the blood vessel branch points formed by the bifurcation of blood vessels, and these points are taken as the second reference points. For example, in the subcutaneous blood vessels of the face of volunteer G, two points at which blood vessels bifurcate on the bridge of the nose are selected as the second reference points b1 and b2. The oil distribution is observed to find the oil distribution boundary points at the edges thereof, and these points are taken as the third reference points. For example, in the oil distribution of the face of volunteer G, two points at the edges of the oil region on the forehead are selected as the third reference points c1 and c2.

[0120] b2, comparing the positions of the first reference points, the second reference points and the third reference points in the respective images, calculating the spatial distances between the reference points in different images, and screening out the reference points with a spatial distance less than a preset distance threshold. The reference points from the visible light image, the near-infrared image and the cross-polarized light image respectively screened out are taken as a group of corresponding points.

[0121] Wherein, the first reference points, the second reference points and the third reference points are the positioning points selected from the epidermal texture, the subcutaneous blood vessels and the oil distribution in step b1, the respective images are the visible light image, the near-infrared image and the cross-polarized light image in which the respective reference points are located; the spatial distance is the positional difference between the reference points in different images; the preset distance threshold is the maximum positional difference standard artificially set for judging whether the reference points correspond to each other; and the corresponding points are a group of first, second and third reference points respectively from the three images and having a spatial distance less than the preset distance threshold.

[0122] In the embodiments of the present application, the positions of the first reference point, the second reference point and the third reference point in the respective images are determined and the coordinates are recorded, for example, the coordinates of the first reference point a1 of the face of the volunteer G in the visible light image are (2, 3), the coordinates of the second reference point b1 in the near-infrared image are (2.2, 3.1), and the coordinates of the third reference point c1 in the cross-polarized light image are (2.3, 3.2); the spatial distances between the reference points in different images are calculated, and the calculation method is the square root of the sum of the squares of the coordinate differences, for example, the distance between a1 and b1 is calculated: first, the x coordinate difference 2-2.2=-0.2 is calculated, and the square is 0.04; the y coordinate difference 3-3.1=-0.1 is calculated, and the square is 0.01; the sum of the two is 0.05, and the square root is about 0.22; the calculated spatial distance is compared with the preset distance threshold (such as 0.5), and the reference points with a distance less than the threshold are screened out; the reference points from the three images are grouped as a group of corresponding points, for example, the distance between a1 and b1 is 0.22, and the distance between a1 and c1 is 0.36, both of which are less than 0.5, so a1, b1 and c1 become a group of corresponding points.

[0123] b3, based on each group of corresponding points, establishing the position relationship of features between multi-modal images.

[0124] Among them, each group of corresponding points is a group of reference points from the three images and with close spatial distances screened out in the b2 step; the multi-modal image is a set including visible light images, near-infrared images and cross-polarized light images; the position relationship is the corresponding relationship of the epidermal texture, subcutaneous blood vessels and oil distribution in space; the establishment is the process of clarifying this spatial relationship through corresponding points.

[0125] In the embodiments of the present application, all groups of corresponding points obtained in the b2 step are collected, for example, the corresponding point groups (a1, b1, c1) and (a2, b2, c2) of the face of the volunteer G; the position relationship of the three reference points in each group of corresponding points is analyzed, for example, in (a1, b1, c1), the position of a1 (the intersection point of the epidermal texture) corresponds to b1 (the branch point of the blood vessels below) and c1 (the oil boundary point around); the position relationship of the epidermal texture, subcutaneous blood vessels and oil distribution in space is summarized by comprehensively analyzing the position relationship of multiple groups of corresponding points, for example, the intersection point of the epidermal texture usually corresponds to the branch point of the blood vessels within a certain range below and the oil boundary point within a certain range around; based on these rules, the position relationship of features between multi-modal images is established, and the approximate position range of the subcutaneous blood vessels and oil distribution corresponding to the epidermal texture at any position is clarified.

[0126] The present application provides the following specific examples: when the skilled person processes the multi-modal images of the face of the volunteer G, two intersection points of the lines on the cheeks are selected as the first reference points a1 and a2 from the epidermal texture, two bifurcation points of the blood vessels on the bridge of the nose are selected as the second reference points b1 and b2 from the subcutaneous blood vessels, and two points on the edge of the oil distribution area on the forehead are selected as the third reference points c1 and c2 from the oil distribution; then the coordinates of the reference points are recorded, and the distance between a1 (2, 3) and b1 (2.2, 3.1) is calculated: the square of the x difference (2-2.2)²=0.04 is calculated first, the square of the y difference (3-3.1)²=0.01, the total is 0.05, the square root is about 0.22, the distance between a1 and c1 (2.3, 3.2) is: the square of the x difference (2-2.3)²=0.09, the square of the y difference (3-3.2)²=0.04, the total is 0.13, the square root is about 0.36, both of which are less than the preset threshold 0.5, a1, b1 and c1 are grouped into a corresponding point, and the same is true for the group (a2, b2, c2); finally, through the two groups of corresponding points, it is found that the intersection point of the epidermal texture is about 0.2-0.3 below the branch point of the blood vessels, and the boundary point of the oil is about 0.3-0.4 around, thereby establishing the positional relationship of the multi-modal image features in the region.

[0127] By performing b1~b3, the embodiments of the present application provide clear markers for feature positioning by selecting representative reference points; the spatial correlation of the reference points is ensured by calculating the spatial distance and selecting corresponding points, and irrelevant points are excluded; the spatial correspondence rules of the epidermal texture, the subcutaneous blood vessels and the oil distribution are clearly reflected based on the positional relationship established by the multiple groups of corresponding points, which provides a reliable basis for the spatial alignment of the multi-modal images and helps to improve the accuracy of subsequent processing.

[0128] In a possible embodiment, S14, the features in the target composite feature map are screened by the front-stage screening unit of the cascaded classifier, based on the screened features, the three-dimensional analysis is performed by the rear-stage analysis unit of the cascaded classifier, so as to realize the determination of the skin texture depth, including:

[0129] Step 141, determine the feature types related to the skin texture depth in the target composite feature map by the front-stage screening unit of the cascaded classifier, the feature types include the convex and concave features of the epidermal texture, the distribution features of the subcutaneous blood vessels below the texture, and the connection features of the oil distribution and the edge of the texture.

[0130] The pre-stage screening unit of the cascade classifier is a tool part for selecting features related to a specific target from an image. The target composite feature map is an image that has been optimized and contains epidermal texture, subcutaneous blood vessels, and oil distribution. The skin texture depth is the depth of the skin surface texture. The feature type is a different feature category related to the skin texture depth, including the convex and concave features of the epidermal texture (the morphology of the uneven skin surface), the distribution features of the subcutaneous blood vessels under the texture (the positional relationship between the blood vessels under the skin surface and the texture), and the connection features of the oil distribution and the texture edge (the connection state of the skin surface oil and the texture edge). The determination is a process of identifying and selecting these related feature types.

[0131] In the embodiments of the present application, the target composite feature map, for example, the facial target composite feature map of volunteer H, is obtained. The pre-stage screening unit of the cascade classifier analyzes all features in the image to determine which features are related to the skin texture depth, for example, excluding uniform skin color area features that are not related to the texture depth. According to the analysis result, the feature types related to the texture depth are determined, including the convex and concave features of the epidermal texture, such as the convex and concave texture of the cheeks, the distribution features of the subcutaneous blood vessels under the texture, such as the blood vessel density under the concave, and the connection features of the oil distribution and the texture edge, such as the oil aggregation state of the convex edge. For example, in the facial image of volunteer H, the pre-stage screening unit identifies the convex and concave texture of the forehead, the blood vessel distribution under these textures, and the oil connection of the texture edge, and determines them as related feature types.

[0132] Step 142, integrating the convex and concave features, the distribution features, and the connection features to generate a set of features to be analyzed, and taking the set of features to be analyzed as the screened features.

[0133] The convex and concave features are the morphology features of the uneven epidermal texture, the distribution features are the position and distribution state of the subcutaneous blood vessels under the texture, the connection features are the connection state of the oil distribution and the texture edge, the integration is to combine different types of features according to certain rules, the set of features to be analyzed is the set formed after integrating all features related to the texture depth, and the screened features are the features selected by the pre-stage screening unit.

[0134] In the embodiments of the present application, the convex and concave features, distribution features and junction features determined in the collecting step 141, such as the convex texture of a certain region of the face of the volunteer H, the blood vessel distribution below the convex texture, and the oil junction of the edge of the convex texture, are associated to form a feature group, such as a certain convex texture and the blood vessel distribution directly below the convex texture and the oil junction of the edge are grouped together, and all such feature groups are summarized together to form a feature set to be analyzed as screening features for subsequent analysis, for example, in the facial image of the volunteer H, the convexity / concavity of each texture on the forehead is combined with the corresponding subcutaneous blood vessel and oil junction features to form a feature set to be analyzed.

[0135] In step 143, the features in the feature set to be analyzed are layered according to the spatial position by the post-stage analysis unit of the cascade classifier to distinguish the skin surface texture features and the associated features below the skin surface.

[0136] In the embodiments of the present application, the features in the feature set to be analyzed, such as the feature set to be analyzed of the face of the volunteer H, are analyzed by the post-stage analysis unit according to the spatial position of each feature in the feature set to be analyzed to determine whether they belong to the surface layer of the skin or below the surface layer, such as the convexity and concavity of the epidermis are located on the surface layer, and the subcutaneous blood vessels below the surface layer are located below the surface layer, and the features are layered according to the determination results to distinguish the skin surface texture features and the associated features below the surface layer, such as the convexity and concavity of the texture are grouped as the surface texture features in the facial feature set of the volunteer H, and the blood vessel distribution and deep oil junction features below the texture are grouped as the associated features.

[0137] In the embodiments of the present application, the features in the feature set to be analyzed, such as the feature set to be analyzed of the face of the volunteer H, are analyzed by the post-stage analysis unit according to the spatial position of each feature in the feature set to be analyzed to determine whether they belong to the surface layer of the skin or below the surface layer, such as the convexity and concavity of the epidermis are located on the surface layer, and the subcutaneous blood vessels below the surface layer are located below the surface layer, and the features are layered according to the determination results to distinguish the skin surface texture features and the associated features below the surface layer, such as the convexity and concavity of the texture are grouped as the surface texture features in the facial feature set of the volunteer H, and the blood vessel distribution and deep oil junction features below the texture are grouped as the associated features.

[0138] In step 144, the distance between the skin surface texture features and the corresponding associated features is taken as the measurement result of the skin texture depth.

[0139] In the embodiments of the present application, the features in the feature set to be analyzed, such as the feature set to be analyzed of the face of the volunteer H, are analyzed by the post-stage analysis unit according to the spatial position of each feature in the feature set to be analyzed to determine whether they belong to the surface layer of the skin or below the surface layer, such as the convexity and concavity of the epidermis are located on the surface layer, and the subcutaneous blood vessels below the surface layer are located below the surface layer, and the features are layered according to the determination results to distinguish the skin surface texture features and the associated features below the surface layer, such as the convexity and concavity of the texture are grouped as the surface texture features in the facial feature set of the volunteer H, and the blood vessel distribution and deep oil junction features below the texture are grouped as the associated features.

[0140] In the embodiments of the present application, the skin surface texture features and the corresponding associated features are determined, for example, the convex features of a certain texture on the face of the volunteer H correspond to the blood vessel distribution features directly below the texture, and the spatial distance between the two features is calculated as the determination result of the skin texture depth at the position.

[0141] The present application provides the following specific examples: when the technical personnel determine the skin texture depth of the face of the volunteer H, the feature types related to the texture depth are determined from the target composite feature map by the front-stage screening unit of the cascade classifier, including the epidermal convexity and concavity of the left cheek, the subcutaneous blood vessel distribution directly below the texture, and the grease junction features of the texture edge; then the features are collected, the convexity / concavity, the corresponding blood vessel distribution, and the grease junction at the same position are associated into a feature group, and the generated feature set to be analyzed is summarized; then the feature set is analyzed by the rear-stage analysis unit, the convexity and concavity are classified as skin surface texture features, and the blood vessel distribution and deep grease junction features directly below the texture are classified as associated features; finally, a certain concave feature (z=0) and the blood vessel feature (z=0.2) directly below the concave feature are selected, the distance is calculated as the determination result of the texture depth at the concave position, and the rest of the positions are calculated in the same way to complete the overall determination.

[0142] By performing steps 141-144, the embodiments of the present application accurately select the features related to the texture depth through the front-stage screening unit, exclude irrelevant interference, realize systematic management of the features by integrating the generated feature set to be analyzed, clearly define the spatial hierarchical relationship of the features through the rear-stage analysis unit, and obtain the determination result by calculating the distance, which accurately reflects the depth of the skin texture and provides reliable and intuitive data support for skin state evaluation. The whole process is logically coherent, ensuring the pertinence and accuracy of the determination.

[0143] In a possible embodiment, step 142 integrates the convexity and concavity features, the distribution features, and the junction features to generate the feature set to be analyzed, including:

[0144] c1, identifying the convexity and concavity features, the distribution features, and the junction features at the same spatial position, and performing association processing to form a feature group.

[0145] Among them, the convexity and concavity features are the height and depth forms of the skin surface texture, the distribution features are the position and distribution state of the subcutaneous blood vessels directly below the texture, the junction features are the connection of the grease distribution and the texture edge, the same spatial position refers to the area at the same position on the skin surface, the association processing refers to associating different features at the same position to form a whole, and the feature group is a set containing the convexity and concavity features, the distribution features, and the junction features at the same spatial position after association processing.

[0146] In the embodiments of the present application, the spatial positions of each feature in the target composite feature map are determined, for example, by marking the coordinates of the features in the image to determine the positions, for example, the convex feature of a certain region of the face of volunteer I is located at coordinates (5, 6), the spatial positions of different features are compared, and the convex and concave features, distribution features and connection features at the same spatial position are identified, for example, it is found that there is a convex feature at the position (5, 6), and there are corresponding blood vessel distribution features below it, and there are corresponding oil connection features at the edge, the features at the same position are associated and processed, and a feature group is formed, for example, the convex feature, the corresponding blood vessel distribution and the oil connection feature at the position (5, 6) are combined into a feature group.

[0147] c2, all feature groups corresponding to the spatial positions are summarized to generate a feature set to be analyzed.

[0148] Among them, the feature group corresponding to the spatial position is the set of all related features of a certain spatial position formed in step c1, the summary means collecting and integrating all such feature groups together, and the feature set to be analyzed is the overall set of all related features after summarizing, which is used for subsequent depth analysis.

[0149] In the embodiments of the present application, all feature groups formed in step c1 are collected, for example, the (5, 6) position feature group of the face of volunteer I, the (3, 4) position feature group, etc., and these feature groups are integrated together in the order of the distribution of the spatial positions to form a complete set, for example, all feature groups of the face of volunteer I are arranged in the order of coordinates from left to right and from top to bottom, and are summarized into a feature set to be analyzed.

[0150] The present application provides the following specific examples: when the technician processes the features of the face of volunteer I, the spatial positions of each feature are first determined through step c1, a convex feature, a blood vessel distribution feature below it and an oil connection feature at the edge are identified at coordinates (2, 3), and the three are associated into a feature group; at coordinates (5, 6), a concave feature, a corresponding blood vessel distribution feature and an oil connection feature are identified, and are also associated into a feature group. Subsequently, through step c2, these feature groups are collected and integrated in the order of coordinates from left to right and from top to bottom, and finally a feature set to be analyzed containing all related features of the face of volunteer I is generated.

[0151] By performing c1-c2, the embodiments of the present application associate different features at the same spatial position into a group through step c1, and the spatial correlation between the features is clear, which avoids the analysis confusion caused by scattered features; through step c2, the feature groups are summarized to form a feature set to be analyzed, which realizes the systematic integration of the features, ensures that all related features can be comprehensively and orderly used in the subsequent analysis process, and provides clear structure and clear correlation of the basic data for accurate determination of the depth of skin texture.

[0152] Figure 3 A structural schematic diagram of a skin texture depth determination system based on image analysis provided by an embodiment of the present application is shown in FIG. 1. As shown in the figure, the system comprises: Figure 3

[0153] An acquisition module 31 is configured to acquire a visible light image, a near-infrared image and a cross-polarized light image of a skin surface of a target object to form a multi-modal image.

[0154] A fusion module 32 is configured to perform spatial alignment on the multi-modal image and perform fusion processing on the spatially aligned multi-modal image to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels and oil distribution.

[0155] A reconstruction module 33 is configured to perform super-resolution reconstruction on features of pixels with a quality value lower than a preset quality threshold in the initial composite feature map based on a generative adversarial network to obtain a target composite feature map.

[0156] A screening module 34 is configured to screen features in the target composite feature map through a front-stage screening unit of a cascaded classifier, and perform three-dimensional analysis on the screened features through a rear-stage analysis unit of the cascaded classifier to realize determination of skin texture depth.

[0157] The skin texture depth determination system based on image analysis of the embodiment of the present application is used to realize the aforementioned skin texture depth determination method based on image analysis, and thus the specific implementation of the skin texture depth determination system based on image analysis can be seen from the foregoing embodiment part of the skin texture depth determination method based on image analysis. The specific implementation can be referred to the description of the corresponding embodiment part, and will not be described here again.

[0158] The present application also provides an electronic device comprising a memory for storing a computer program and a processor for executing the computer program to realize the steps of the skin texture depth determination method based on image analysis.

[0159] The present application also provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to realize the steps of the skin texture depth determination method based on image analysis.

[0160] In an exemplary embodiment, the aforementioned computer readable storage medium can include but is not limited to a U disk, a read-only memory, a random access memory, a mobile hard disk, a magnetic disk or an optical disk and various media that can store computer programs.

[0161] ​The embodiment of the present application further provides a computer program product, the computer program product comprising a computer program, the computer program being executed by a processor to implement the steps in any of the above image analysis based skin texture depth determination method embodiments.

[0162] Those skilled in the art will further appreciate that the functions of the examples described herein, including any related steps of a method, can be implemented using electronic hardware, computer software, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various examples have been described herein generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0163] The above provides a kind of based on image analysis skin texture depth determination method, system, electronic equipment and storage medium provided in the present application in detail.The principle and implementation of the present application are described in this paper by applying specific examples, the above example is only used to help understand the method of the present application and its core idea.It should be pointed out that, for the ordinary skilled person in the art, without departing from the principle of the present application, the present application can be improved and modified, and these improvements and modifications also fall within the scope of the present application.

Claims

1. A method for determining skin texture depth based on image analysis, characterized in that, include: Acquire visible light, near-infrared, and cross-polarized light images of the target object's skin surface to form a multimodal image; The multimodal images are spatially aligned, and the spatially aligned multimodal images are fused to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels, and oil distribution. Based on a generative adversarial network, super-resolution reconstruction is performed on the features of pixels with quality values ​​lower than a preset quality threshold in the initial composite feature map to obtain the target composite feature map. The features in the target composite feature map are filtered by the front-stage filtering unit of the cascade classifier. Based on the filtered features, the three-dimensional analysis is performed by the back-stage parsing unit of the cascade classifier to determine the depth of skin texture. The target composite feature map is filtered by the pre-stage filtering unit of the cascaded classifier. Based on the filtered features, the post-stage parsing unit of the cascaded classifier performs three-dimensional analysis to determine the skin texture depth, including: The cascade classifier uses a pre-filtering unit to determine the feature types related to skin texture depth in the target composite feature map. These feature types include the convex and concave features of epidermal texture, the distribution features of subcutaneous blood vessels below the texture, and the connection features between oil distribution and texture edges. The protrusion and depression features, the distribution features, and the connection features are integrated to generate a feature set to be analyzed, and the feature set to be analyzed is used as the filter features. The subsequent parsing unit of the cascaded classifier stratifies the features in the feature set to be parsed according to their spatial location, so as to distinguish the skin surface texture features and the associated features located below the skin surface. The distance between the skin surface texture features and the corresponding associated features is used as the measurement result of the skin texture depth.

2. The method according to claim 1, characterized in that, The method based on a generative adversarial network (GAN) involves super-resolution reconstruction of the features of pixels with quality values ​​below a preset quality threshold in the initial composite feature map to obtain a target composite feature map, including: The continuous region composed of pixels whose quality values ​​are lower than a preset quality threshold in the initial composite feature map is defined as a low-quality region. The features of the low-quality region are expanded by the generative module of the generative adversarial network to generate candidate reconstruction regions. The discriminant module of the generative adversarial network compares the features of the candidate reconstruction region with the features of the surrounding region formed by pixels with quality values ​​not lower than a preset quality threshold in the initial composite feature map, and outputs the comparison result. Based on the comparison results, the features of the candidate reconstruction regions are adjusted to obtain the adapted reconstruction regions; The low-quality region is replaced by the adapted reconstruction region to obtain the target composite feature map.

3. The method according to claim 2, characterized in that, The generation module of the generative adversarial network expands the features of the low-quality region to generate candidate reconstruction regions, including: The features of the low-quality region are associated with the features of the surrounding region to determine the feature correspondence; Based on the feature correspondence, the features of the low-quality region are supplemented to increase the number and detail of the features in the low-quality region; The supplemented features of the low-quality region are integrated to form a candidate reconstruction region of the same size as the low-quality region.

4. The method according to claim 1, characterized in that, The step of spatially aligning the multimodal images and fusing the spatially aligned multimodal images to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels, and oil distribution includes: The salient features of each image in the multimodal image are extracted, wherein the salient feature of the visible light image is the epidermal texture, the salient feature of the near-infrared image is the subcutaneous blood vessels, and the salient feature of the cross-polarized light image is the distribution of oil. Based on the epidermal texture, the subcutaneous blood vessels, and the oil distribution, the positional relationship of features among the multimodal images is established; According to the positional relationship, adjust the position of each image in the multimodal image to make the images spatially aligned; The spatially aligned images are fused to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels, and oil distribution.

5. The method according to claim 4, characterized in that, The step of establishing the positional relationship of features among the multimodal images based on the epidermal texture, subcutaneous blood vessels, and oil distribution includes: The first reference point is selected from the intersection of the epidermal texture, the second reference point is selected from the branching point of the subcutaneous blood vessels, and the third reference point is selected from the boundary point of the oil distribution. Compare the positions of the first reference point, the second reference point, and the third reference point in their respective images, calculate the spatial distance between each reference point in different images, filter out reference points whose spatial distance is less than a preset distance threshold, and take the filtered reference points from the visible light image, the near-infrared image, and the cross-polarized light image as a set of corresponding points. Based on the corresponding points in each group, the positional relationship of features between the multimodal images is established.

6. The method according to claim 1, characterized in that, The process of integrating the protrusion and depression features, the distribution features, and the connection features to generate a feature set to be parsed includes: Identify the protrusions and depressions, the distribution features, and the connection features that are in the same spatial position, and perform association processing to form a feature group; The feature sets corresponding to all spatial locations are aggregated to generate a feature set to be parsed.

7. A skin texture depth measurement system based on image analysis, used in the skin texture depth measurement method based on image analysis as described in any one of claims 1 to 6, characterized in that, include: The acquisition module is used to acquire visible light images, near-infrared images, and cross-polarized light images of the target object's skin surface to form a multimodal image; The fusion module is used to spatially align the multimodal images and fuse the spatially aligned multimodal images to obtain an initial composite feature map containing epidermal texture, subcutaneous blood vessels and oil distribution. The reconstruction module is used to perform super-resolution reconstruction of the features of pixels with quality values ​​lower than a preset quality threshold in the initial composite feature map based on a generative adversarial network, so as to obtain the target composite feature map. The filtering module is used to filter features in the target composite feature map through the pre-filtering unit of the cascaded classifier, and perform three-dimensional analysis through the post-parse unit of the cascaded classifier based on the filtered features to achieve the determination of skin texture depth.

8. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the image analysis-based skin texture depth measurement method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the skin texture depth measurement method based on image analysis as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • 3D finger vein extraction method and system based on multispectral image

    CN110298273A

  • Face skin feature analysis method, device, equipment, medium and product

    CN119867646A