A Real-time Detection Method for Human Skin Color Region Based on ZYNQ

By developing a real-time detection method for human skin tone areas on the ZYNQ hardware platform, the problems of inconvenience, high cost and poor real-time performance of human skin tone detection on traditional PC platforms are solved, and efficient and real-time detection of human skin tone areas on the FPGA hardware platform is achieved.

CN114708340BActive Publication Date: 2025-07-01NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111668529.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-01
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

Traditional human skin tone detection has problems such as inconvenience, high cost and poor real-time performance on PC platforms, especially in application scenarios where real-time performance of image processing is high.

Method used

Using the FPGA hardware platform based on ZYNQ, real-time detection method for human skin color areas is developed. By building the skin color detection algorithm IP core in HLS, and C simulation and optimization instruction processing are carried out to realize real-time acquisition, processing and display of image data.

Benefits of technology

It realizes real-time detection of human skin color areas on the ZYNQ hardware platform, has low resource utilization, can meet daily real-time image data acquisition and processing needs, has broad application prospects, and has strong system redevelopmentability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708340B_ABST
    Figure CN114708340B_ABST
Patent Text Reader

Abstract

The present invention relates to a real-time detection method for human skin color regions based on ZYNQ. The method includes the following steps: Step (1): Collect images; Step (2): Process the image data collected in Step (1) through a human skin color region detection IP core, convert the color space, perform dynamic threshold processing to obtain the skin color region and perform detection and marking; Step (3): Perform the minimum connected region algorithm processing on the skin color part of the image detected in Step (2); Step (4): Perform image data transfer and extraction, write the image data into DDR3 using VDMA, and use FIFO for pipelining processing; Step (5): Display the processed image on the LCD screen in real time. Since the present invention detects human skin color regions through ZYNQ, it only needs to be powered on to start working, collect images in real time, identify the corresponding skin color regions for display, and at the same time, secondary image processing can be performed during the operation of the system, and the system has strong re-developability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a real-time detection method for human skin color regions based on ZYNQ. Background Technique

[0002] FPGA (Field Programmable Gate Array) is a further development product based on programmable devices such as PAL and GAL. It appears as a semi-custom circuit in the field of application-specific integrated circuits (ASICs). It not only solves the deficiencies of custom circuits but also overcomes the drawback of limited gate circuits in the original programmable devices. Due to the parallel processing mechanism of the FPGA itself, the FPGA is often applied in hardware algorithm acceleration and is also increasingly widely used in the field of image processing. Especially in the context of high requirements for real-time image processing today, more stringent requirements are imposed on the hardware.

[0003] Traditional human skin color detection is directly carried out on a PC for acquisition and processing, which will mainly result in two major limitations: one is that it is not convenient to carry a PC and it is not miniaturized enough. At the same time, in actual engineering applications, directly mounting the entire PC on the device will also greatly increase the cost; the other is that image processing on the CPU of the PC will occasionally experience frame drops and the real-time performance is very poor. If image processing is carried out on the GPU of the PC, although the performance can be completely comparable or even exceed that of the FPGA, its high power consumption is also a very influential factor. Summary of the Invention

[0004] The purpose of the present invention is to provide a real-time detection method for human skin color regions based on ZYNQ.

[0005] The technical solution for achieving the purpose of the present invention is: a real-time detection method for human skin color regions based on ZYNQ, including the following steps:

[0006] Step (1): Collect images;

[0007] Step (2): Process the image data collected in step (1) through a human skin color region detection IP core, convert the color space, perform dynamic threshold processing to obtain the skin color region and perform detection and marking;

[0008] Step (3): Perform the minimum connected region elimination algorithm processing on the skin color partial image detected in step (2);

[0009] Step (4): Perform image data transfer and extraction, write the image data into DDR3 using VDMA, and perform pipelining processing using FIFO;

[0010] Step (5): Real-time display the processed image on the LCD screen.

[0011] Further, step (2) "detect the IP of the human skin area, process the image data collected in step (1), convert the color space, perform dynamic threshold processing to obtain the skin area and perform detection and marking" is specifically as follows:

[0012] Step (21): It is necessary to preprocess the image data collected by the OV5640 camera, that is, RGB565 -> RGB888:

[0013] pixel_data[7:0] = {source_data[4:0], 0} (1)

[0014] pixel_data[7:0] = {source_data[6:0], 0} (2)

[0015] Equation (1) represents the conversion of the image data of the R and B channels, and equation (2) represents the conversion of the G image data channel, that is, extracting the high bits and filling the low bits with 0;

[0016] Step (22): Convert the image data converted in step (21) into YCbCr image data and perform skin color extraction, specifically as follows:

[0017] Y = 0.299R + 0.587G + 0.114B (3)

[0018] Cb = 0.564(B - Y) (4)

[0019] Cr = 0.713(R - Y) (5) where Y represents the luminance information, and Cb and Cr represent the color difference information of the blue and red channels;

[0020] Step (23): Perform threshold determination on the three converted variables. When equations (6), (7), and (8) are simultaneously satisfied, the image discrimination flag bit skin_flag is obtained;

[0021] y_upper > Y > y_lower (6)

[0022] Cb_upper > Cb > Cb_lower (7)

[0023] Cr_upper > Cr > Cr_lower (8)

[0024] Step (24): Perform pure white marking on the area divided by the flag bit in step (23), as shown in the following formula:

[0025] temp_pixel = (skin_flag)? 255 : pixel_data (9)

[0026] Among them, temp_pixel represents the processed image pixel value, generally referring to the three-channel image data R, G, B, and pixel_data is the pixel value of the original pixel point.

[0027] Further, step (3) "performing the minimum connected region algorithm processing on the skin-color part image detected in step (2)" is specifically as follows:

[0028] Set the search area of the central point pixel (x, y) as the array R, and set the search direction of the central point as 8 directions, that is, the pixel block of 3×3 for the first-layer outward expansion search area, and so on; give the search boundary conditions to finally obtain the search area R of this central pixel point;

[0029] In the inclined direction:

[0030]

[0031]

[0032] In the horizontal and vertical directions:

[0033] x′ = x + cosθ θ ∈ θ2

[0034] y′ = y - sinθ θ ∈ θ2

[0035] Among them x′ and y′ are the coordinates of the next search pixel point, where θ is the eight angular values around the central point, and four angular values are taken from θ1 in the inclined direction, and four angular values are taken from θ2 in the horizontal and vertical directions;

[0036] At the same time, limit the threshold judgment condition:

[0037] G(x′, y′) - G(x, y) < T

[0038] In the above formula, the gray difference between the front and back pixels is used as the judgment basis, and the threshold T is set as the termination condition for the regional search iteration;

[0039] Finally, obtain the connected region R and perform region size elimination:

[0040]

[0041] Judge the region R to see if its value is less than the threshold A (the upper limit value of the connected region area). If it is less, eliminate it; otherwise, keep it.

[0042] Further, in step (1), the OV5640 camera is used for image acquisition.

[0043] Further, the interface definition of the human skin color region detection IP core in step (2) is as follows: The interface types of the input and output are defined as axis register both, and the interface types of the eight parameters rows, cols, y_lower, y_upper, cb_lower, cb_upper, cr_lower, and cr_upper are defined as s_axilite. The optimization instruction dataflow is used to perform pipelining on the global function processing flow. The return value interface of the IP and core top-level function is defined as ap_ctrl_none, and this interface is used to implement a design without any block-level I / O protocols.

[0044] Compared with the prior art, the present invention has the following remarkable advantages:

[0045] (1) The ZYNQ hardware platform is used to detect the human skin color region in real time. By pre-building the corresponding skin color detection algorithm IP core in HLS and performing C simulation to ensure the correct function, then considering the corresponding parameters and video data stream interface definition, and taking into account the timing of data processing, reasonable optimization instructions are used to perform pipelining operations on the input and output image data, and the corresponding IP core is integrated and exported. Then, the IP core block diagram of the entire system is built and subsequent effectiveness analysis and comprehensive analysis are carried out to obtain the underlying hardware platform. The resource occupancy of the entire IP core system is very low, and it can fully meet the daily real-time image data acquisition. The identified image data is stored in DDR3 in real time after post-processing, which is convenient for secondary processing of image data, enabling it to be connected to many other fields and having broad application prospects.

[0046] (2) Since this system detects the human skin color region through ZYNQ, it only needs to be powered on to start working, and can collect images in real time, identify the corresponding skin color region for display, and at the same time, secondary image processing can be performed during the operation of the system. The system has strong reusability. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is the system flow chart of the present invention.

[0048] Figure 2 is the detailed diagram of the intermediate data flow direction.

[0049] Figure 3 is the implementation flow chart of the HLS human skin color detection IP core.

[0050] Figure 4 is the simulation test diagram in the HLS stage.

[0051] Figure 5 is the simulation test result diagram in the HLS stage.

[0052] Figure 6 For the system hardware implementation platform, the ZYNQ chip model XC7Z020clg400-2's Navigator development board is selected.

[0053] Figure 7 It is the effect diagram of the human hand skin color area detected in actual use.

[0054] Figure 8 It is the HLS timing synthesis result and the ZYNQ resource utilization rate.

[0055] Figure 9 It is the interface after the final synthesis of HLS. Specific implementation manners

[0056] The following refers to Figure 1 , and further describes the specific technical implementation solutions of the present invention, so as to facilitate those skilled in the art to further understand the present invention without constituting a limitation to its rights.

[0057] The latest heterogeneous chip technology ZYNQ, adopting the architecture of FPGA (PL side) + ARM (PS side), can flexibly adapt to different application requirements, thus greatly improving its engineering applicability. Therefore, the present invention mainly builds a high-precision real-time human skin color area detection system on ZYNQ. The overall system design and the skin color detection IP core design are the main invention points. By optimizing the overall data flow mode and the IP core algorithm instructions in the system application, the accuracy and real-time performance of image processing are greatly improved, and it has broad application prospects.

[0058] Please refer to Figure 1 As shown, a high-precision real-time human skin color area detection system based on ZYNQ, the method includes the following steps:

[0059] Step 1: Image acquisition, using the OV5640 camera to acquire images and transfer them to the subsequent image processing.

[0060] Step 2: Based on Step 1, perform human skin color area detection. In this part, the color space is transformed, and dynamic threshold processing is performed to obtain the skin color area and detect and mark it.

[0061] Step 3: Based on Step 2, perform image post-processing, and perform post-processing algorithms such as removing the smallest connected area on the detected skin color part of the image.

[0062] Step 4: Based on Step 3, perform image data transfer and extraction, use VDMA to write the image data into DDR3 (memory) and use FIFO for pipelining processing.

[0063] Step 5: Finally, use the LCD screen to display the processed image on the LCD screen in real time.

[0064] Detailed implementation of each step:

[0065] Step 1: Use an OV5640 camera with 5 million pixels to collect images and transfer them to subsequent image processing.

[0066] Step 2: Process the image data collected in Step 1 through a human skin color region detection IP core. Development process of the human skin color region detection IP core: In Figure 3 The leftmost is the main algorithm implementation flowchart of skin color detection in the HLS implementation process. First, the AXI4-Stream video stream image data is converted into RGB888 image data. For the convenience of image data processing and calculation, the RGB888 image data is then converted into YCbCr image data to separate the luminance information and chrominance information. The YCbCr image data can then be judged by six independent parameters to obtain the judgment flag bit skin_flag, and the corresponding skin color region is marked as pure white. Finally, it is converted back into AXI4-Stream video stream image data for image data output; Figure 3 In the middle is the HLS simulation process. The main idea is to use the opencv library imported by HLS to verify the skin color detection algorithm for the image to be detected (see Figure 4 ). The verification result is shown in Figure 5 . It can be analyzed that this skin color detection algorithm is feasible; in Figure 3 On the far right, the interface definition of this IP core is listed, and the function data processing is optimized using corresponding optimization instructions. Among them, the input and output of the entrance are defined as the axis register both interface type, and the eight parameters rows, cols, y_lower, y_upper, cb_lower, cb_upper, cr_lower, and cr_upper are defined as s_axilite to facilitate the PS side to configure the corresponding parameters. The optimization instruction dataflow is used to pipeline the global function processing flow. The return value interface of the IP and core top-level function is defined as ap_ctrl_none, and this interface is used to implement a design without any block-level I / O protocol; Figure 8 shows the HLS timing synthesis result and the ZYNQ resource utilization rate. It can be seen that the resource occupancy rate of this IP core is extremely low, leaving more resources for secondary development of other image processing. Figure 9 is the interface list after the final synthesis of HLS.

[0067] The human skin color detection algorithm in Step 2 is as follows:

[0068] First, the image data collected by the OV5640 camera needs to be preprocessed, that is, RGB565 -> RGB888:

[0069] pixel_data[7:0] = {source_data[4:0], 0} (1)

[0070] pixel_data[7:0] = {source_data[6:0], 0} (2)

[0071] Equation (1) represents the conversion of image data for the R and B channels, and equation (2) represents the conversion of the G image data channel, that is, extracting the high bits and padding 0 for the low bits.

[0072] Then, the above image data is converted into YCbCr image data, and skin color extraction is performed, as shown in the following formula:

[0073] Y = 0.299R + 0.587G + 0.114B (3)

[0074] Cb = 0.564(B - Y) (4)

[0075] Cr = 0.713(R - Y) (5)

[0076] y_upper > Y > y_lower (6)

[0077] Cb_upper > Cb > Cb_lower (7)

[0078] Cr_upper > Cr > Cr_lower (8)

[0079] Among them, Y represents the luminance information, and Cb and Cr represent the chrominance difference information of the blue and red channels. Since the original RGB image does not distinguish between luminance information and chrominance difference information, such conversion is for facilitating image data processing; at the same time, threshold determination is performed on the three converted variables. When equations (6), (7), and (8) are satisfied simultaneously, the image discrimination flag bit skin_flag is obtained;

[0080] In order to demarcate the skin color area, the area demarcated by the above flag bit is marked with pure white, as shown in the following formula:

[0081] temp_pixel = (skin_flag)? 255 : pixel_data (9)

[0082] Among them, temp_pixel represents the processed image pixel value, generally referring to the three-channel image data R, G, and B.

[0083] Step 3: On the basis of Step 2, directly develop the image post-processing part of the skin color area in the IP core. Using the parallel computing method of the FPGA, the core algorithm in the image post-processing stage, the algorithm for removing the smallest connected region, is also developed and synthesized through HLS.

[0084] The algorithm for removing the smallest connected region in Step 3 is as follows:

[0085] Let the search area of the central pixel (x, y) be the array R. Due to the parallel computing characteristics of the FPGA, the search direction of the central point is set to 8 directions, that is, the pixel block of 3×3 in the first-layer outer expansion search area, and so on. And the boundary conditions for the search are given, and finally the search area R of this central pixel point is obtained.

[0086] In the inclined direction:

[0087]

[0088]

[0089] In the horizontal and vertical directions:

[0090] x′ = x + cosθ θ ∈ θ2

[0091] y′ = y - sinθ θ ∈ θ2

[0092] Where x′ and y′ are the coordinates of the next search pixel. Among them, θ is the eight angular values around the central point. Four angular values are taken from θ1 in the inclined direction, and four angular values are taken from θ2 in the horizontal and vertical directions;

[0093] At the same time, the threshold judgment condition is defined:

[0094] G(x′, y′) - G(x, y) < T

[0095] In the above formula, the gray difference between the front and back pixels is used as the judgment basis, and the threshold T is set as the termination condition for the regional search iteration;

[0096] Finally, the connected region R is obtained, and the region size is removed:

[0097]

[0098] Judge the region R. Whether its value is less than the threshold A (the upper limit value of the area of the connected region). If it is less, it is removed; otherwise, it is retained.

[0099] Step 4: After the above image processing is completed, each processed frame of the image is temporarily stored in DDR3 in real time, which is convenient for secondary development during the system operation. In the system, the Video In to AXI4-Stream core and the AXI4-Stream to Video Out core are mainly used to convert the video stream data into AXI4-Stream data, and the VDMA core is used to allocate different address spaces to temporarily store the acquired real-time images in DDR3. At the same time of temporary storage, the images read from DDR3 are cached by using FIFO to realize pipeline data processing and improve the single-cycle image processing ability of the FPGA.

[0100] Step 5: The image data read in Step 4 is displayed on the LCD screen in real time. The developed rgb to lcd core is used to adapt to different models of external LCD screens and drive the LCD screen to work and display images.

[0101] Overall system IP core framework: Camera (OV5640) image data acquisition IP core, RGB565 to AXI4-Stream video stream IP core, skin color detection IP core, VDMA data storage IP core (directly accessing the DDR3 on the PS side), AXI4-Stream video stream to RGB888 image output IP core, rgb to lcd core for adapting to external LCD screens with different IDs. The above is the main image data processing flow line on the PL side; in terms of PS side control, the ZYNQ7 Processing System core mainly provides the corresponding clock (100MHZ) for the PL side. At the same time, in order to consider the dynamic timing matching problems of the VTC core and the AXI4-Stream to Video Out core, etc., the dynamic clock generation IP core is used to provide clocks for the above several types of cores, thereby improving the stability and reliability of the system. After building the IP system diagram, the external ports of the two ends of the IP core are led out and constrained with the corresponding on-board I / O, thus forming the entire block diagram. For the detailed data flow transfer diagram description, see Figure 2 ;

[0102] Finally, the above system block diagram is subjected to IP core validity analysis and comprehensive analysis to generate the underlying hardware platform. After generating the bit stream, it is imported into the SDK for on-board verification. The verification results are shown in Figure 7 。

[0103] After the above steps, the exported bit stream is opened in the SDK, and the corresponding IP is configured using the axilite interface, initialized and driven to work with the corresponding peripherals. Finally, the program is downloaded into the ZYNQ development board for on-board verification.

[0104] This system detects the human skin color area through ZYNQ. Therefore, it only needs to be powered on to start working, collect images in real time, identify the corresponding skin color areas for display, and at the same time, secondary image processing can be carried out during the operation of the system. The system has strong redevelopability.

Claims

1. A real-time detection method for human skin color regions based on ZYNQ, characterized in that, It includes the following steps: Step (1): Collect images; Step (2): Use the human skin color area detection IP core to process the data of the images collected in step (1), convert the color space, perform dynamic threshold processing to obtain the skin color area and perform detection and marking; Step (3): Perform the minimum connected region elimination algorithm processing on the skin color part of the images detected in step (2); Step (4): Perform image data transfer and extraction, use VDMA to write the image data into DDR3, and use FIFO for pipelining processing; Step (5): Display the processed images on the LCD screen in real time; Step (2) "Use the human skin color area detection IP core to process the data of the images collected in step (1), convert the color space, perform dynamic threshold processing to obtain the skin color area and perform detection and marking" specifically is: Step (21): Preprocess the image data collected by the OV5640 camera, specifically RGB565 -> RGB888: pixel_data[7:0] = {source_data[4:0], 0} (1) pixel_data[7:0] = {source_data[5:0], 0} (2) Equation (1) represents the conversion of the image data of the R and B channels, and equation (2) represents the conversion of the G image data channel. Specifically, the high bits are extracted and the low bits are filled with 0; Step (22): Convert the image data converted in step (21) into YCbCr image data and perform skin color extraction, specifically: Y = 0.299R + 0.587G + 0.114B (3) Cb = 0.564(B - Y) (4) Cr = 0.713(R - Y) (5) where Y represents the luminance information, and Cb and Cr represent the chromatic aberration information of the blue and red channels; Step (23): Perform threshold determination on the three converted variables. When equations (6), (7), and (8) are simultaneously satisfied, the image discrimination flag bit skin_flag is obtained; y_upper > Y > y_lower (6) Cb_upper > Cb > Cb_lower (7) Cr_upper > Cr > Cr_lower (8) Step (24): Perform pure white marking on the area divided by the flag bit in step (23), as shown in the following formula: temp_pixel = (skin_flag)? 255 : pixel_data (9) where temp_pixel represents the processed image pixel value, which is the three-channel image data R, G, B, and pixel_data is the original pixel point pixel value; Step (3) "Perform the minimum connected region elimination algorithm processing on the skin color part of the images detected in step (2)" specifically is: Set the search area of the central point pixel (x, y) as the array R, set the search direction of the central point as 8 directions, which is the pixel block of 3x3 for the first-layer outer expansion search area, and so on; give the search boundary conditions to finally obtain the search area of this central pixel point; In the inclined direction: In the horizontal and vertical directions: x′ = x + cosθ, θ ∈ θ2 y′ = y - sinθ, θ ∈ θ2 wherein x' and y' are the coordinates of the next search pixel point, where θ is the eight angular values around the center point, four angular values in the four directions of θ1 are taken in the tilt direction, and four angular values in the four directions of θ2 are taken in the horizontal and vertical directions; Meanwhile, the threshold judgment condition is defined as follows: G(x′, y′) - G(x, y) < T In the above formula, the gray difference between the previous and subsequent pixels is used as the judgment basis, and the threshold T is set as the termination condition for the regional search iteration; Finally, the connected search region array R is obtained, and the region size is eliminated: The search region array R is judged to determine whether its value is less than the threshold A. A is the upper limit of the area of the connected region. If it is less than A, it is eliminated; Otherwise, it is retained.

2. The method according to claim 1, wherein In step (1), the OV5640 camera is used for image acquisition.