Monocular human body distance measurement method and system based on AI depth and segmentation model
Through parallel processing and morphological operation of AI depth and segmentation model, monocular visual ranging problems are solved in complex scenarios and real-time human target ranging accuracy and real-time nature, thereby achieving efficient monocular human ranging.
Patent Information
- Application Number
- CN202510821811.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
AI Technical Summary
The existing monocular visual ranging technology is difficult to adapt to the dynamic characteristics of non-rigid targets such as the human body. It has problems such as low ranging accuracy, low efficiency and poor scene adaptability, especially in complex scenarios, background interference and occlusion lead to large ranging errors.
The parallel processing method based on AI depth and segmentation model is adopted, and the binary mask and depth estimation model are generated through the human body segmentation model to generate the initial depth feature map. The target distance of the human body is calculated by combining the camera imaging geometric model, and the depth features are optimized using morphological operations and dynamic Gaussian fuzzy processing to position the geometric center point of the bottom boundary of the human body mask and the ground contact area.
It improves the ranging accuracy and real-time performance of non-rigid human targets in complex scenarios, reduces the calculation load, adapts to interference and occlusion in different scenarios, and achieves high-precision monocular ranging.
Smart Images

Figure CN120339296A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and image processing, and particularly relates to a monocular human body ranging method and system based on an AI depth and segmentation model. Background Art
[0002] Existing monocular vision ranging technologies mostly rely on rigid target assumptions, multi-view geometric constraints or complex deep learning models, and it is difficult to adapt to the dynamic characteristics of non-rigid targets such as the human body. Traditional methods often face the following limitations: The geometric prior-based solutions are sensitive to pose changes and are prone to ranging jumps due to human joint movements; The solutions based on stereo vision or ToF sensors require additional hardware, with high costs and limited deployment; While the end-to-end deep learning models rely on a large amount of labeled data, with insufficient generalization and high computational loads. Especially in complex scenarios, background interference, occlusion and illumination changes will couple depth noise in the human body area, further amplifying the ranging error. The existing technologies have not effectively balanced the contradictions among accuracy, efficiency and scene adaptability, restricting the practical application process of monocular ranging in fields such as security, robotics, AR / VR, etc. Summary of the Invention
[0003] The main object of the present invention is to provide a monocular human body ranging method and system based on an AI depth and segmentation model, and through parallel processing of dual models and mask-guided depth optimization, to achieve the purpose of improving the ranging accuracy and real-time performance of non-rigid human body targets in complex scenarios.
[0004] To achieve the above object, the present invention provides a monocular human body ranging method based on an AI depth and segmentation model, including the following steps: Simultaneously input the acquired monocular RGB image into a pre-trained human body segmentation model and a depth estimation model, and respectively generate a binary mask of the human body area and an initial depth feature map; Optimize the initial depth feature map according to the binary mask of the human body area to obtain a depth feature map; Based on the depth feature map and the binary mask of the human body area, locate the three-dimensional coordinates of the geometric center point of the lowest connected area on the bottom boundary of the human body mask in contact with the ground, and calculate the distance of the human body target in combination with the camera imaging geometric model.
[0005] Further, the step of simultaneously inputting the acquired monocular RGB image into a pre-trained human body segmentation model and a depth estimation model, and respectively generating a binary mask of the human body area and an initial depth map, includes: Input the RGB image into the human body segmentation model and the depth estimation model respectively through parallel computing threads; Output a pixel-level binary classification result through the human body segmentation model to generate a binary mask of the human body area; Meanwhile, an initial depth feature map with the same resolution as the input image is output through the depth estimation model.
[0006] Furthermore, the training steps of the human body segmentation model include: Extract multi-scale human body features at each downsampling stage of the encoder; Set a binary supervision branch at the corresponding upsampling stage of the decoder; Optimize the model parameters by weighted fusion of the supervision losses at each scale.
[0007] Furthermore, the training steps of the depth estimation model include: Generate depth pseudo-labels using the optical flow information of consecutive frames in the monocular video sequence; Construct a joint loss function that includes an image edge alignment constraint; Improve the depth estimation accuracy through a multi-stage training strategy.
[0008] Furthermore, the steps of optimizing the initial depth feature map according to the binary mask of the human body region to obtain the depth feature map include: For the human body region covered by the binary mask of the human body region, perform median filtering on the depth values according to a preset window value; For the non-human body region outside the binary mask of the human body region, perform dynamic Gaussian blur processing where the blur intensity weakens as the distance of the pixel point from the human body edge increases.
[0009] Furthermore, the steps of the dynamic Gaussian blur processing include: Determine the Euclidean distance between each pixel point in the non-human body region and the human body edge; Dynamically calculate the standard deviation of the Gaussian blur kernel corresponding to the pixel according to the Euclidean distance, where the larger the distance, the smaller the standard deviation; Perform Gaussian blur processing on the corresponding pixel using the dynamically calculated standard deviation.
[0010] Furthermore, the steps of locating the three-dimensional coordinates of the geometric center point of the lowest connected region on the contact area between the bottom boundary of the human body mask and the ground include: Perform a morphological dilation operation in the vertically downward direction on the binary mask of the human body region; Extract the lowest connected region of the binary mask of the human body region after dilation; Calculate the two-dimensional projection of the coordinates of the geometric center point of the lowest connected region on the three-dimensional coordinates of the contact area.
[0011] Furthermore, the steps of the morphological dilation operation include: Use a rectangular structuring element with a preset length-width ratio to perform a single dilation process on the binary mask of the human body region in the vertical direction of the image; Restrict the direction of the inflation effect to extend only vertically downward until the bottom boundary of the mask intersects with the preset ground area of the image.
[0012] Furthermore, the steps of calculating the distance of the human target in combination with the camera imaging geometric model include: Obtain camera parameters, including monocular focal length parameters and installation height parameters; Calculate the vertical offset from the image center point according to the vertical pixel position of the three-dimensional coordinates of the geometric center point in the image; Output the target distance value based on the geometric proportional relationship between the vertical offset and the camera parameters.
[0013] The present invention also provides a monocular human body ranging system based on an AI depth and segmentation model, including: A parallel computing unit for simultaneously inputting the acquired monocular RGB image into a pre-trained human body segmentation model and a depth estimation model to respectively generate a binary mask of the human body area and an initial depth feature map; A depth optimization unit for optimizing the initial depth feature map according to the binary mask of the human body area to obtain a depth feature map; A ranging calculation unit for positioning the three-dimensional coordinates of the geometric center point of the lowest connected area on the contact area between the bottom boundary of the human body mask and the ground based on the depth feature map and the binary mask of the human body area, and calculating the distance of the human target in combination with the camera imaging geometric model.
[0014] The monocular human body ranging method and system based on the AI depth and segmentation model provided by the present invention have the following beneficial effects: By adopting a dual-model parallel inference architecture, the present invention synchronously generates a human body mask and a depth map, avoiding error accumulation in traditional serial processing; designs a mask-guided depth optimization mechanism to adaptively denoise the human body area and simultaneously implement distance-aware blurring of the background, retaining key geometric features while suppressing interference; and based on a contact point positioning method combining morphological operations and physical constraints, it can stably extract projection geometric parameters without relying on human body pose estimation; the present invention also integrates lightweight model design and hardware acceleration strategies, taking into account the computing power limitations of embedded devices. The overall solution demonstrates strong robustness in complex dynamic scenarios, providing a reliable technical path for low-cost and high-precision monocular ranging. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a schematic flowchart of a monocular human body ranging method based on an AI depth and segmentation model in an embodiment of the present invention; Figure 2 is a structural block diagram of a monocular human body ranging system based on an AI depth and segmentation model in an embodiment of the present invention.
[0016] The realization, functional features, and advantages of the present invention will be further described in conjunction with embodiments with reference to the accompanying drawings. Detailed Embodiment
[0017] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0018] Referring to Figure 1 , which is a schematic flowchart of a monocular human body ranging method based on an AI depth and segmentation model proposed by the present invention, includes the following steps: S1. Simultaneously input the obtained monocular RGB image into a pre-trained human body segmentation model and a depth estimation model to respectively generate a human body region binary mask and an initial depth feature map; S2. Optimize the initial depth feature map according to the human body region binary mask to obtain a depth feature map; S3. Based on the depth feature map and the human body region binary mask, locate the three-dimensional coordinates of the geometric center point of the lowest connected region on the bottom boundary of the human body mask in contact with the ground area, and calculate the distance of the human body target in combination with the camera imaging geometric model.
[0019] In one embodiment, for step S1, The step of simultaneously inputting the obtained monocular RGB image into a pre-trained human body segmentation model and a depth estimation model to respectively generate a human body region binary mask and an initial depth map includes: Input the RGB image into the human body segmentation model and the depth estimation model respectively through parallel computing threads; Output a pixel-level binary classification result through the human body segmentation model to generate a human body region binary mask; Simultaneously output an initial depth feature map with the same resolution as the input image through the depth estimation model.
[0020] In the specific implementation process, after uniformly preprocessing the obtained monocular RGB image, it is synchronously input into the human body segmentation model and the depth estimation model through parallel computing threads. The human body segmentation model is based on a lightweight encoding-decoding structure, integrating a multi-scale feature supervision mechanism, and outputs a pixel-level binary mask to accurately locate the human body area; the depth estimation model adopts a global attention mechanism and a resolution restoration design to generate a depth feature map with the same resolution as the input image. The two models achieve zero-copy transmission of input data through shared video memory, and use the intermediate layer feature selective fusion technology to enhance the depth representation consistency of the human body boundary area. To balance computational efficiency and accuracy, the human body segmentation model compresses the computational amount through lightweight modules, while the depth estimation model relies on the Transformer architecture to capture long-range depth dependencies. Under the parallel computing architecture, the inference time of the dual models is significantly reduced compared to the serial scheme. At the same time, through the feature complementary mechanism, the depth estimation error of the human body area is greatly reduced. In monocular human body ranging, the collaborative processing of human body segmentation and depth estimation is the core basis for ensuring ranging accuracy. The traditional serial processing mode is difficult to meet the dual requirements of real-time performance and accuracy due to problems such as cumulative timing errors and lack of feature interaction. Step S1 breaks through the coupling error bottleneck of segmentation and depth estimation in monocular ranging through the dual-model parallel collaboration and feature interaction design, forms an efficient closed-loop from data input to feature generation, and realizes the efficient collaboration of human body area segmentation and scene depth estimation.
[0021] In one embodiment, the training steps of the human body segmentation model include: Extract multi-scale human body features at each downsampling stage of the encoder; Set a binarization supervision branch at the corresponding upsampling stage of the decoder; Optimize the model parameters by weighted fusion of the multi-scale supervision losses.
[0022] Specifically, the encoder extracts multi-level human features through a cascaded downsampling module, captures detailed features such as human contours and clothing textures through a shallow network, and learns semantic information such as human postures and proportions through a deep network. In each upsampling stage of the decoder, a binary supervision branch is correspondingly set up to form a multi-granularity supervision flow from coarse to fine. Among them, the deep supervision branch focuses on the localization of the human body's main area, the middle branch optimizes the integrity of limb connections, and the shallow branch strengthens the discrimination ability of fine edges such as hair tips and hem lines. During the training process, the contribution ratio of the supervision losses at each scale is dynamically adjusted through an adaptive weight allocation algorithm, so that the model focuses on global structure learning in the initial stage of training and gradually strengthens local detail optimization in the later stage. To improve the generalization performance of the model, enhancement strategies such as light perturbation, occlusion simulation, and multi-human interaction are introduced into the training data, and a two-stage training process is adopted: first, the basic feature extraction ability is pre-trained based on a large-scale general human dataset, and then it is fine-tuned through small-sample target scene data to adapt to a specific environment. The traditional single-stage supervision training mode is prone to problems such as blurred human body edges and missed detection of small targets due to a single feature scale; the multi-scale progressive supervision mechanism proposed in this solution improves the robustness of human body segmentation in complex scenarios through an encoding-decoding feature alignment and dynamic loss fusion strategy.
[0023] In one embodiment, the training steps of the depth estimation model include: Generate depth pseudo-labels using the optical flow information of consecutive frames in a monocular video sequence; Construct a joint loss function that includes an image edge alignment constraint; Improve the depth estimation accuracy through a multi-stage training strategy.
[0024] Specifically, based on the spatio-temporal consistency constraint between consecutive frames of a monocular video, an optical flow-guided depth inference mechanism is adopted: calculate the dense pixel displacement field between adjacent frames through an optical flow estimation model (such as RAFT); Combine the relative motion parameters between frames output by a camera pose estimation network (such as ORB-SLAM); construct a polar geometry constraint equation, inversely infer the pixel depth value to generate pseudo-labels, and filter out the unreliable optical flow in the dynamic object area by introducing a motion confidence mask to ensure the reliability of the static background area of the pseudo-labels. To simultaneously constrain the depth geometric rationality and edge alignment, construct a multi-task loss function to enforce the spatial consistency between the current frame depth map and the depth map of the next frame after optical flow transformation; calculate the cosine similarity between the depth map gradient and the Sobel edge detection result of the input image to penalize the depth edge misalignment; apply an anisotropic smoothing constraint in the low-texture area to retain the sharpness of the depth step boundary. Dynamically balance the contribution ratio of each loss term through a learnable weight coefficient to adapt to the optimization objectives at different training stages. Adopt a phased optimization strategy to gradually improve the model performance. Train the basic depth perception ability on a large-scale synthetic dataset (such as Virtual KITTI) in the pre-training stage; perform domain adaptation training on the monocular video data of the target scene with the optical flow pseudo-labels as the main supervision signal; freeze the optical flow and pose estimation network parameters, and jointly optimize the depth network and the edge alignment loss term; continuously fine-tune the model parameters online through the temporal consistency constraint in the real-time video stream at the deployment stage. Traditional supervised learning relies on expensive sensors to obtain depth ground truth. The self-supervised dynamic pseudo-label generation and edge-aware joint optimization strategy of this solution realizes unsupervised high-precision depth estimation through the internal correlation between the temporal information of monocular video and the image geometric features, and has the ability to adapt to dynamic scenes, providing a highly robust depth perception basis for the ranging system.
[0025] In one embodiment, for step S2, The step of optimizing the initial depth feature map according to the human body region binary mask to obtain a depth feature map includes: For the human body region covered by the human body region binary mask, perform median filtering on the depth values according to a preset window value; For the non-human body region outside the human body region binary mask, perform dynamic Gaussian blur processing with the blur intensity decreasing as the distance from the pixel point to the human body edge increases.
[0026] In the specific implementation process, a pre-trained human segmentation model is used to generate an accurate binary mask of the human body region. For the human body region covered by the mask, an adaptive median filtering strategy is adopted, and the filtering window is dynamically adjusted according to the human body size, while smoothing the noise and retaining important depth step features. A gradient filtering intensity attenuation mechanism is designed for the human body edge region to ensure a natural transition between the human body and the background. For the non-human body region, a distance-aware dynamic blur algorithm is adopted, and the blur intensity is automatically adjusted according to the distance between the pixel point and the human body edge. A stronger blur is applied to the background region close to the human body to suppress interference, while more details are retained in the region far from the human body. To balance the processing effect and real-time performance, the system adopts a parallel computing architecture to accelerate the optimization process and introduces an edge protection mechanism to avoid the loss of scene structure caused by excessive blur. Step S2 effectively solves the problems of depth jump and background interference existing in the traditional method through the refined processing of the depth map guided by the human segmentation mask by means of a dual-region differential optimization method based on semantic priors.
[0027] In one embodiment, the steps of the dynamic Gaussian blur processing include: Determine the Euclidean distance between each pixel point in the non-human body region and the human body edge; Dynamically calculate the standard deviation of the Gaussian blur kernel corresponding to the pixel according to the Euclidean distance, and the larger the distance, the smaller the standard deviation; Perform Gaussian blur processing on the corresponding pixel using the dynamically calculated standard deviation.
[0028] Specifically, through the binary mask distance transformation algorithm, calculate each pixel point in the non-human body region The Euclidean distance to the nearest human body edge : ; In the formula, Represents the set of edge pixels of the human body mask. During actual deployment, a parallelized chamfer matching algorithm is used to accelerate the calculation to ensure millisecond-level processing on the GPU.
[0029] Design a non-linear attenuation function to dynamically adjust the standard deviation of the Gaussian kernel : ; In the formula, Is the preset maximum blur intensity (such as the typical value 5.0), which controls the strongest blur effect at the human body edge; Is the attenuation rate coefficient, which determines the rate of decrease of the blur intensity with the increase of the distance; Is the normalized distance value (mapped to the [0,1] interval), that is, the Euclidean distance.
[0030] For each non-human pixel , according to Generate the corresponding Gaussian kernel , perform a convolution operation: ; In the formula, is the kernel radius, usually taking to cover 99.7% of the energy; is the standard Gaussian kernel function. In monocular ranging, the depth noise in the background area is likely to interfere with the ranging result. In this solution, by fusing distance changes and adaptive Gaussian kernel regulation, intelligent suppression of background interference is achieved in the monocular ranging system. By deeply coupling semantic information (human mask) with underlying image processing, the mathematical model quantifies the dynamic relationship between fuzzy intensity and spatial position, providing key guarantee for ranging of the monocular vision system in complex environments.
[0031] In one embodiment, for step S3, The step of positioning the three-dimensional coordinates of the geometric center point of the lowest connected area on the contact area between the bottom boundary of the human mask and the ground includes: Perform a morphological dilation operation in the vertically downward direction on the binary mask of the human body area; Extract the lowest connected area of the binary mask of the human body area after dilation; Calculate the two-dimensional projection of the three-dimensional coordinates of the contact area of the geometric center point coordinates of the lowest connected area.
[0032] In the specific implementation process, perform a vertically oriented morphological dilation operation on the binary mask, use an asymmetric structuring element to strengthen the downward extension characteristic of the mask, and simulate the natural diffusion process of the human body contacting the ground under the action of gravity. The number of dilation times is dynamically adjusted according to the mask height to ensure that the contact areas of human bodies at different distances can effectively cover the ground projection range. The lowest connected area of the dilated mask is quickly extracted through parallel connected component analysis, combined with area preference and geometric constraint rules to exclude noise interference, and finally the two-dimensional projection position of the contact area is determined by calculating the mean value of the pixel coordinates of the connected component: ; In the formula Let the pixel coordinates be within the connected region, and N be the total number of pixels. After being corrected by the camera pitch angle, this coordinate is further mapped to the three-dimensional space for ranging calculation. To improve the anti-interference ability, a dilation-erosion balance strategy is introduced. Before vertical dilation, the mask edge is smoothed and preprocessed to avoid false contact areas caused by burr noise. At the same time, the ground equation estimated in real time by visual SLAM is fused to verify the candidate points in the three-dimensional space, ensuring that the positioning result conforms to the physical space constraints. In monocular human body ranging, the accurate positioning of the contact point between the human body and the ground is the key link to realize geometric ranging. This embodiment proposes a robust positioning method based on mask morphological operations, which solves the problem of contact point ambiguity in complex pose and occlusion scenarios through image processing technology driven by physical constraints.
[0033] In one embodiment, the steps of the morphological dilation operation include: Using a rectangular structuring element with a preset length-width ratio, perform a single dilation process on the binary mask of the human body region along the vertical direction of the image; Limit the dilation action direction to extend only vertically downward until the bottom boundary of the mask intersects with the preset ground region of the image.
[0034] Specifically, use rectangular structuring elements with significantly different length-width ratios, and its mathematical representation is: ; In the formula, B is an asymmetric rectangular structuring element, and are the width and height of the structuring element respectively, and satisfy . Perform a unidirectional dilation operation on the binary mask M: ; When performing the unidirectional dilation operation, only expand the mask area in the vertically downward direction to avoid false touch phenomena caused by horizontal diffusion. The dilation termination condition is dynamically determined by the preset ground region of the image: when the bottom boundary of the dilated mask enters the preset ground region G of the image (defined as the pixel band at the bottom of the image), it terminates immediately to ensure that the contact point positioning conforms to the physical constraints of the scene. The vertical coordinate threshold of the ground region can be calculated in real time according to the camera installation height H and the pitch angle : ; where is the vertical center coordinate of the image, is the focal length, is the minimum ranging range.
[0035] Traditional isotropic expansion is prone to introducing interference due to multi-directional diffusion. In contrast, the directional expansion in this embodiment is deeply integrated with physical constraints. Through the collaborative optimization of the structural element design and the expansion termination condition, the underlying image operation is elevated to the level of spatial geometric reasoning, ensuring that the mask extends only vertically downward and precisely simulating the physical contact process between the human body and the ground.
[0036] In one embodiment, the steps of calculating the distance of a human target in combination with the camera imaging geometric model include: Obtain camera parameters, including the monocular focal length parameter and the installation height parameter; Calculate the vertical offset from the image center point based on the vertical pixel position of the three-dimensional coordinates of the geometric center point in the image; Based on the geometric proportional relationship between the vertical offset and the camera parameters, output the target distance value.
[0037] In a specific implementation, obtain calibration parameters, including the camera focal length (in pixel units) and the installation height H (in physical units), and determine the coordinates of the image center point . Calculate the vertical offset from the image center based on the two-dimensional coordinates output by the contact area positioning module ; Based on the pinhole imaging principle and the relationship of similar triangles, the geometric calculation formula for the target distance is: ; Map the pixel-level offset to the actual physical distance through the focal length , which directly reflects the projection position difference of the target in the optical axis direction. To improve the robustness of the model, the system introduces a dynamic correction mechanism for the camera pitch angle θ. When the camera is tilted, the vertical offset needs to be corrected to: ; to ensure the ranging accuracy in non-horizontal installation scenarios. At the same time, fuse the preliminary distance estimation value of the depth optimization module and adopt an adaptive weight fusion strategy, where is the adaptive fusion weight coefficient (range 0 - 1): Balance the reliability of the geometric model and the data-driven results through Kalman filtering, significantly reducing the ranging error in extreme scenarios. In this embodiment, through the ingenious modeling of vertical projection geometry, the monocular ranging is transformed into a lightweight linear calculation, with the advantages of physical interpretability and real-time performance, providing a theoretically complete and highly efficient implementation solution for the deployment of low-cost embedded devices.
[0038] Refer to Figure 2, which is a structural block diagram of a monocular human body ranging system based on an AI depth and segmentation model in an embodiment of the present invention, includes: A parallel computing unit for simultaneously inputting the acquired monocular RGB image into a pre-trained human body segmentation model and a depth estimation model to respectively generate a binary mask of the human body region and an initial depth feature map; A depth optimization unit for optimizing the initial depth feature map according to the binary mask of the human body region to obtain a depth feature map; A ranging calculation unit for positioning the three-dimensional coordinates of the geometric center point of the lowest connected region on the contact area between the bottom boundary of the human body mask and the ground based on the depth feature map and the binary mask of the human body region, and calculating the distance of the human body target in combination with the camera imaging geometric model.
[0039] For the specific implementation of each unit in the above device example, please refer to that described in the above method embodiment, and details are not repeated here.
[0040] In summary, the present invention simultaneously inputs the acquired monocular RGB image into a pre-trained human body segmentation model and a depth estimation model to respectively generate a binary mask of the human body region and an initial depth feature map; optimizes the initial depth feature map according to the binary mask of the human body region to obtain a depth feature map; positions the three-dimensional coordinates of the geometric center point of the lowest connected region on the contact area between the bottom boundary of the human body mask and the ground based on the depth feature map and the binary mask of the human body region, and calculates the distance of the human body target in combination with the camera imaging geometric model, so as to achieve the purpose of improving the ranging accuracy and real-time performance of non-rigid human body targets in complex scenarios.
[0041] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided by the present invention and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0042] It should be noted that in this document, the terms "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, apparatus, article, or method including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, apparatus, article, or method including that element.
[0043] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A monocular human body ranging method based on an AI depth and segmentation model, characterized in that Including the following steps: Input the acquired monocular RGB image into a pre-trained human body segmentation model and a depth estimation model simultaneously, and respectively generate a binary mask of the human body region and an initial depth feature map; Optimize the initial depth feature map according to the binary mask of the human body region to obtain a depth feature map; Based on the depth feature map and the binary mask of the human body region, locate the three-dimensional coordinates of the geometric center point of the lowest connected region on the contact area between the bottom boundary of the human body mask and the ground, and calculate the distance of the human body target in combination with the camera imaging geometric model.
2. The monocular human body ranging method based on the AI depth and segmentation model according to claim 1, wherein The step of inputting the acquired monocular RGB image into a pre-trained human body segmentation model and a depth estimation model simultaneously to respectively generate a binary mask of the human body region and an initial depth map includes: Input the RGB image into the human body segmentation model and the depth estimation model respectively through parallel computing threads; Output a pixel-level binary classification result through the human body segmentation model to generate a binary mask of the human body region; Simultaneously output an initial depth feature map with the same resolution as the input image through the depth estimation model.
3. The monocular human body ranging method based on the AI depth and segmentation model according to claim 2, wherein The training steps of the human body segmentation model include: Extract multi-scale human body features at each downsampling stage of the encoder; Set a binarization supervision branch at the corresponding upsampling stage of the decoder; Optimize the model parameters by weighted fusion of the supervision losses at each scale.
4. The monocular human body ranging method based on the AI depth and segmentation model according to claim 2, wherein The training steps of the depth estimation model include: Generate depth pseudo-labels using the optical flow information of consecutive frames in a monocular video sequence; Construct a joint loss function including an image edge alignment constraint; Improve the depth estimation accuracy through a multi-stage training strategy.
5. The monocular human body ranging method based on the AI depth and segmentation model according to claim 1, wherein, The step of optimizing the initial depth feature map according to the binary mask of the human body region to obtain a depth feature map includes: Perform median filtering on the depth values of the human body region covered by the binary mask of the human body region according to a preset window value; Perform dynamic Gaussian blur processing on the non-human body region outside the binary mask of the human body region, where the blur intensity weakens as the distance from the pixel point to the human body edge increases.
6. The monocular human body ranging method based on the AI depth and segmentation model according to claim 5, characterized in that, The steps of the dynamic Gaussian blur processing include: Determine the Euclidean distance between each pixel point in the non-human body region and the human body edge; Dynamically calculate the standard deviation of the Gaussian blur kernel corresponding to the pixel according to the Euclidean distance, and the larger the distance, the smaller the standard deviation; Perform Gaussian blur processing on the corresponding pixel using the dynamically calculated standard deviation.
7. The monocular human body ranging method based on the AI depth and segmentation model according to claim 1, wherein The step of locating the three-dimensional coordinates of the geometric center point of the lowest connected region on the contact area between the bottom boundary of the human body mask and the ground includes: Perform a morphological dilation operation on the binary mask of the human body region in the vertical downward direction; Extract the lowest connected region of the dilated binary mask of the human body region; Calculate the two-dimensional projection of the geometric center point coordinates of the lowest connected region on the three-dimensional coordinates of the contact area.
8. The monocular human body ranging method based on the AI depth and segmentation model according to claim 7, characterized in that, The steps of the morphological dilation operation include: Use a rectangular structuring element with a preset length-width ratio to perform a single dilation process on the binary mask of the human body region along the vertical direction of the image; Limit the dilation action direction to extend only vertically downward until the bottom boundary of the mask intersects with the preset ground region of the image.
9. The monocular human body ranging method based on the AI depth and segmentation model according to claim 1, wherein The step of calculating the distance of the human body target in combination with the camera imaging geometric model includes: Obtain camera parameters, including the monocular focal length parameter and the installation height parameter; Calculate the vertical offset from the center point of the image based on the vertical pixel position of the three-dimensional coordinates of the geometric center point in the image; Output the target distance value based on the geometric proportional relationship between the vertical offset and the camera parameters.
10. A monocular human body ranging system based on an AI depth and segmentation model, characterized in that, It includes: A parallel computing unit for simultaneously inputting the obtained monocular RGB image into a pre-trained human segmentation model and a depth estimation model to generate a binary mask of the human body region and an initial depth feature map respectively; A depth optimization unit for optimizing the initial depth feature map according to the binary mask of the human body region to obtain a depth feature map; A ranging calculation unit for positioning the three-dimensional coordinates of the geometric center point of the lowest connected region on the contact area between the bottom boundary of the human body mask and the ground based on the depth feature map and the binary mask of the human body region, and calculating the human target distance in combination with the camera imaging geometric model.
Citation Information
Patent Citations
Portrait background automatic replacement method combining convolutional network and neighborhood similarity
CN110956681A
Human body image detection method and device, electronic equipment and readable storage medium
CN111126300A
Monocular image depth estimation method in complex environment based on domain adaptation
CN113436240A
Monocular distance measurement method based on camera imaging geometrical relationship and system thereof
CN113465572A
Monocular three-dimensional instance segmentation method based on depth information guidance
CN116258734A