Machine vision target positioning method and system based on bounding box adaptive adjustment

Through the improved lightweight object detection model and adaptive width measurement and distance measurement algorithm, the problem of background interference in object detection and positioning is solved, and more accurate acquisition of target position, distance, width and height is achieved.

CN120279252AActive Publication Date: 2025-07-08QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510393405.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-08
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

Existing object detection and positioning algorithms can easily confuse the background and foreground in complex backgrounds, resulting in inaccurate target positioning and poor bounding box identification, which affects the calculation accuracy of target width and height.

Method used

Using an improved lightweight object detection model and adaptive width measurement and ranging algorithm, the bounding box adaptive adjustment will improve detection accuracy and efficiency, accurately distinguish the target area and background area, and calculate the target position, distance, width and height.

Benefits of technology

It achieves more accurate and reliable target positioning, reduces background interference, improves detection accuracy and efficiency, and ensures the rigor of the target bounding box and the accuracy of calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279252A_ABST
    Figure CN120279252A_ABST
Patent Text Reader

Abstract

The invention discloses a machine vision target positioning method and system based on bounding box adaptive adjustment, and relates to the technical field of target detection and positioning, and the method comprises the steps: obtaining a real-time image and a depth image of a current monitoring environment; inputting the real-time image into an improved YOLO-based lightweight target detection model, identifying a target in the image and extracting a target bounding box; aligning the real-time image with the depth image, and determining the depth value of each pixel point according to the coordinate of each pixel point of the target bounding box area; and based on the depth value of each pixel point in the target bounding box region, distinguishing a target region and a background region in the target bounding box region by adopting an adaptive width measurement and distance measurement algorithm, extracting the overall depth value of the target region, determining the target distance and the target position, and then identifying the peripheral boundaries of the target in the target bounding box region. And determining the width and height of the target so as to complete target positioning. According to the invention, more accurate and reliable acquisition of target positioning information can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection and positioning, and particularly to a machine vision object positioning method and system based on adaptive adjustment of bounding boxes. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] Object detection and positioning is one of the research hotspots in the current field of computer vision. With the systematic use of artificial intelligence machine deep learning in image and object detection, currently, object detection algorithms based on deep learning such as YOLO series algorithms like YOLOv4 and YOLOv5 are widely used in the field of object detection and positioning, and object positioning information including object distance, object width and height, object position, etc. can be obtained.

[0004] However, affected by factors such as the complexity of the image background, the existing object detection and positioning algorithms inevitably have confusion between the characteristics of the background and the foreground. The presence of the background will lead to inaccurate object positioning, and the obtained distance information may also involve the background area, thus resulting in inaccurate positioning information such as the final object position and object distance. In addition, limited by the accuracy of the object detection algorithm, when generating the object boundary, it may not be possible to make it airtight, and there are non-object gaps in the boundary, and the existence of these gaps may affect the subsequent acquisition of the object width and height, making the calculation of the object width and height inaccurate. Summary of the Invention

[0005] To solve the above deficiencies of the prior art, the present invention provides a machine vision object positioning method and system based on adaptive adjustment of bounding boxes, which uses an improved lightweight object detection model to track and monitor the object and extract the object bounding box, effectively improving the detection accuracy and efficiency; then adopts an adaptive width measurement and distance measurement algorithm to optimize the extracted object bounding box, and calculates and obtains more accurate and reliable object positioning information such as object position, object distance, object width and height, etc., to solve the problems of inaccurate identification of the object bounding box area and low accuracy caused by background interference in object positioning in the prior art.

[0006] In the first aspect, the present invention provides a machine vision object positioning method based on adaptive adjustment of bounding boxes.

[0007] A machine vision object positioning method based on adaptive adjustment of bounding boxes includes: Obtaining a real-time image and a depth image of the current monitoring environment; Inputting the real-time image into a lightweight object detection model based on improved YOLO to identify the object in the image and extract the object bounding box; Align the real-time image with the depth image, map the target bounding box area in the real-time image to the depth image, and determine the depth value of each pixel point in the target bounding box area according to the coordinate of each pixel point in the target bounding box area; wherein, the pixel points without depth values are labeled as the NAN data type. Based on the depth value of each pixel point in the target bounding box area, adopt an adaptive width measurement and distance measurement algorithm to distinguish the target area and the background area in the target bounding box area, extract the overall depth value of the target area, determine the target distance and the target position, and then identify the four perimeters of the target in the target bounding box area to determine the target width and height, so as to complete the target positioning.

[0008] In a second aspect, the present invention provides a machine vision target positioning system based on adaptive adjustment of the bounding box.

[0009] A machine vision target positioning system based on adaptive adjustment of the bounding box includes: An image acquisition module, configured to acquire a real-time image and a depth image of the current monitoring environment; A target bounding box extraction module, configured to input the real-time image into a lightweight target detection model based on improved YOLO to identify the target in the image and extract the target bounding box; A depth value mapping module, configured to align the real-time image with the depth image, map the target bounding box area in the real-time image to the depth image, and determine the depth value of each pixel point in the target bounding box area according to the coordinate of each pixel point in the target bounding box area; wherein, the pixel points without depth values are labeled as the NAN data type. A target positioning module, configured to based on the depth value of each pixel point in the target bounding box area, adopt an adaptive width measurement and distance measurement algorithm to distinguish the target area and the background area in the target bounding box area, extract the overall depth value of the target area, determine the target distance and the target position, and then identify the four perimeters of the target in the target bounding box area to determine the target width and height, so as to complete the target positioning.

[0010] In a third aspect, the present invention further provides an electronic device, including: a memory for storing executable instructions; a processor, when executing the executable instructions stored in the memory, implementing the above-mentioned machine vision target positioning method based on adaptive adjustment of the bounding box.

[0011] In a fourth aspect, the present invention further provides a computer-readable storage medium storing executable instructions, which when causing a processor to execute the executable instructions, implement the above-mentioned machine vision target positioning method based on adaptive adjustment of the bounding box.

[0012] In a fifth aspect, the present invention also provides a computer program product, which includes executable instructions stored in a computer-readable storage medium; wherein, when a processor of an electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, a machine vision target localization method based on adaptive adjustment of bounding boxes as described above is implemented.

[0013] The above one or more technical solutions have the following beneficial effects: 1. The present invention provides a machine vision target localization method and system based on adaptive adjustment of bounding boxes, which uses an improved lightweight object detection model to track and monitor the target and extract the target bounding box, effectively improving the detection accuracy and efficiency; then, an adaptive width measurement and distance measurement algorithm is used to optimize the extracted target bounding box, and more accurate and reliable target localization information is calculated and obtained, solving the problems of inaccurate identification of the target bounding box area and low accuracy caused by background interference in target recognition and localization in the prior art.

[0014] 2. Considering that the accurate and rapid extraction of the target detection box is an important prerequisite for realizing accurate target localization, therefore, in the present invention, to reduce the amount of computation and the number of parameters of the deep learning model and improve the detection accuracy, based on the existing YOLOv11 model, a lightweight object detection model based on improved YOLO is proposed. By introducing the SaE attention module into the neck network, the feature extraction ability of the model is strengthened, the ability of the network to capture channel patterns and master global knowledge is enhanced, deeper feature representations are extracted, and the recognition accuracy of the final target detection box is improved; by improving the SCDown module, a TDDown module is proposed and introduced into the backbone network and neck network of the improved model to further lightweight the network model, which can reduce the number of parameters and the amount of computation of the model while maintaining channel information; through the above improvements, the efficiency and accuracy of the final detection are effectively improved.

[0015] 3. The present invention uses an adaptive width measurement and distance measurement algorithm to optimize the extracted target bounding box. By analyzing the depth value of each pixel point in the target bounding box, it is judged whether there is background interference in the target detection box area, and the background and foreground are further separated based on this to ensure the accuracy of the final target depth distance calculation; through column-by-column or row-by-row NAN judgment, the target area is further refined to calculate more accurate target height and width, and more accurate and reliable target localization information is obtained.

[0016] The advantages of the additional aspects of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. Description of the Drawings

[0017] The accompanying drawings of the specification, which form a part of the present invention, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0018] Figure 1 It is a flowchart of the machine vision target localization method based on boundary box adaptive adjustment according to the embodiment of the present invention; Figure 2 It is a schematic structural diagram of the vision sensor in the embodiment of the present invention; Figure 3 It is a schematic structural diagram of the lightweight target detection model based on improved YOLO in the embodiment of the present invention; Figure 4 It is a schematic diagram of the adaptive width and distance measurement algorithm for the detection target in the embodiment of the present invention; Figure 5 It is a schematic diagram of the machine vision target localization system based on boundary box adaptive adjustment according to the embodiment of the present invention.

[0019] Wherein, 1. Vision sensor; 2. Photosensitive sensor; 3. Color camera; 4. Infrared camera; 4-1. Left infrared camera; 4-2. Right infrared camera; 5. Structured light projector. Specific embodiments

[0020] It should be noted that the following detailed description is exemplary only for describing specific embodiments, aiming to provide a further explanation of the present invention and not intended to limit the exemplary embodiments of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0021] Embodiment 1 To solve the problem of inaccurate target localization in the prior art, this embodiment provides a machine vision target localization method based on boundary box adaptive adjustment. By combining technologies such as image processing, deep learning, and target detection, an improved target detection model and an adaptive width and distance measurement algorithm are proposed to achieve more accurate and reliable target localization and obtain more accurate target localization information. The machine vision target localization method based on boundary box adaptive adjustment proposed in this embodiment, as Figure 1 shown, includes the following steps: Step S1, obtaining real-time images and depth images of the current monitoring environment.

[0022] In this embodiment, a vision sensor is set, which is arranged at an appropriate position facing the monitoring direction and communicatively connected to the industrial control computer, and the vision sensor is used to monitor the current monitoring environment in real time. As Figure 2 shown, the set vision sensor 1 includes an infrared camera 4 (including a left infrared camera 4-1 and a right infrared camera 4-2), a color camera 3, and a structured light projector 5. The structured light projector is used to provide illumination to make the images captured by the camera clearer; the color camera is used to collect RGB real-time images; the infrared camera is used to collect infrared real-time images. At the same time, binocular ranging can be performed through the left and right dual infrared cameras to obtain the depth image of the current monitoring environment, and more accurate depth image information can be obtained in combination with the structured light projector.

[0023] In the design of the above vision sensor, considering the influence of the light intensity on image acquisition, two types of cameras are set in this embodiment for image acquisition. Considering that the RGB color image obtained by the color camera can provide more information for target detection under strong light conditions, and the infrared image obtained by the infrared camera can provide more information for target detection under weak light conditions. Therefore, a photosensitive sensor 2 is also set in this embodiment, which is arranged on one side of the vision sensor and is used to obtain the light intensity of the current monitoring environment in real time. The set photosensitive sensor is communicatively connected to the vision sensor and the industrial control computer, and the industrial control computer makes a judgment according to the received light intensity and selects the corresponding image capture method to obtain real-time images, that is: when the light intensity is greater than the set value, the color camera is started to capture real-time images; when the light intensity is less than the set value, the infrared camera is started to capture real-time images.

[0024] Step S2: Input the real-time image into the lightweight target detection model based on the improved YOLO to identify the targets in the image and extract the target bounding boxes.

[0025] In this embodiment, the improved YOLO target detection model is used to monitor the targets in the image. It learns the image features through a multi-layer neural network, extracts the target bounding boxes according to the image features, and then combines the depth information obtained by the depth camera based on the target bounding boxes for subsequent precise target positioning.

[0026] Among them, in order to reduce the amount of computation and the number of parameters of the deep learning model, improve the detection accuracy, and be more suitable for the deployment of edge devices, this embodiment is based on the existing YOLOv11 model, and by improving the existing model, a lightweight target detection model based on the improved YOLO is proposed, as Figure 3 shown, specifically: First, clarify the structure of the existing YOLOv11 model, which mainly includes a backbone network, a neck network, and a head network. The image is input into the backbone network, and through the downsampling feature extraction and feature fusion operations in the backbone network, deep features are extracted. The deep features are then subjected to feature fusion and upsampling through the neck network, and finally, the detected object detection boxes are output through the head network.

[0027] Among them, the YOLOv11 network model includes Input (input layer), Conv (Convolution, convolution module), C3k2 (Faster Implementation of CSP Bottleneck with 2 convolutions, feature extraction module), SPPF (Spatial Pyramid Pooling - Fast, multi-scale feature extraction module), C2PSA (feature extraction and processing module based on attention mechanism), Upsample (upsampling module), Concat (merging module), Detect (object detection box output module), etc. Further, Conv is used for downsampling; C3k2 is used for feature extraction and feature fusion; SPPF extracts multi-scale features by performing pooling operations at different scales to help the model better detect objects of different sizes; C2PSA introduces a self-attention mechanism to make up for the lack of long-range dependence in the CNN-Based model, that is, the basic convolution model, enhances the global modeling ability of the network, and extracts the features weighted by self-attention; Concat is used to merge features; Detect is used to output the object detection box. In addition, in the figure k represents the convolutional kernel ( kernel ), s represents the stride ( stride ), C3k = False means not enabling the dynamic convolutional kernel size control function in the C3k2 structure, and conversely, C3k = True means enabling; FC (Full Connected Layer) represents the fully connected layer, and C is Concat, representing feature splicing and merging.

[0028] In this embodiment, to further enhance the YOLO detection ability and reduce the computational amount and the number of parameters, the existing YOLOv11 model is improved: (1) To improve the detection accuracy, the SaE (Squeeze Aggregated Excitation) attention module based on the SENetV2 attention mechanism, namely the SaEBlock, is introduced into the neck network. The SaE attention module is combined with the C3k2 module to propose the C3k2SaE module. In this module, the SENetV2 attention mechanism is placed after two convolutions in the Bottleneck (i.e., the bottleneck layer) of the C3k2 module to enhance the feature extraction ability of the module, making the module pay more attention to features rich in expressive ability, reducing redundancy, and thus enhancing the model detection performance.

[0029] Specifically, in the SaE attention module, the input features are compressed by two convolutional layers and then input into the SaE attention module. After the Adaptive AvgPool2d operation, the importance weights of each channel of the compressed features are extracted. The extracted importance weights are aggregated through the fully connected layer FC, and then channel weighting is performed according to the importance weights of each channel and the input features to obtain the deep feature representation. Its calculation formula can be expressed as: ; Among them, represents the excitation operation, represents the squeeze operation, represents the input features, represents the activation function.

[0030] By introducing the above SaE attention module, the ability of the network to capture channel patterns and master global knowledge can be effectively enhanced, thus achieving better feature representation.

[0031] (2) To achieve lightweight, the TDDown (TrioDepth-Wise Convolution Down) module is also introduced into the backbone network and the neck network in this embodiment. This module is an improvement based on the SCDown module (a type of downsampling module). In the TDDown module, the input features are processed in parallel through two parallel Depth-Wise Convolution (DW Conv) operations, and the two processed features are fused and then downsampled through one more depth convolution. In this embodiment, the above method replaces the traditional Point-WiseConvolution (PW Conv) operation used by the original SCDown module to further lightweight the SCDown, and thus further lightweight downsampling of features is achieved through the TDDown module.

[0032] By using the above-mentioned depth convolution for downsampling, the spatial dimension of the features can be reduced, and low-dimensional features can be extracted, which can help effectively reduce the spatial dimension of the input tensor, thereby reducing the number of model parameters and the amount of computation, while maintaining the channel information.

[0033] Step S3: Align the real-time image and the depth image, map the target bounding box area in the real-time image to the depth image, and determine the depth value of each pixel point in the target bounding box area according to the coordinate of each pixel point in the target bounding box area.

[0034] In the above steps, considering that the real-time image can be obtained through a color camera or an infrared camera respectively, when obtaining the RGB real-time image using the color camera, the RGB image and the depth image obtained by the depth camera are subjected to synchronization and image alignment processing, and the alignment processing includes: First, calibrate the RGB camera and the depth camera to obtain the internal parameters (including focal length such as 、 and the principal point, etc.) and external parameters (including the rotation matrix and the translation vector ) of the two cameras to ensure that their spatial relationship is known.

[0035] Then, convert the pixels in the depth image into 3D space points, that is, according to the internal parameters of the depth camera, combined with the coordinate of each depth pixel point and its corresponding depth value, the 3D position coordinates in the camera coordinate system can be calculated.

[0036] After that, use the external parameter matrix and the translation vector to convert the above-converted 3D points from the depth camera coordinate system to the RGB camera coordinate system so that they are in the same spatial reference system. Considering that the converted 3D points need to be projected back to the RGB image coordinate system, therefore, using the internal parameters of the RGB camera again, map the 3D position point coordinates to the 2D pixel coordinates of the RGB image. Further, due to the different resolutions of the RGB image and the depth image, there may be holes or pixel misalignments in the projected depth data, so the bilinear interpolation method can be used for correction to improve the alignment quality.

[0037] Through the above method, the mapping from the depth image to the RGB view is completed, and the aligned RGB-D data is obtained. In fact, most current depth cameras usually have a built-in synchronous alignment mode, which can be directly called through the SDK (a kind of software development kit) to complete.

[0038] When obtaining an infrared real-time image using an infrared camera, since the depth camera is implemented based on a binocular infrared camera, when using the infrared image, only determining whether the depth image is from the left camera or the right camera is required to complete the alignment, and there is no need to perform the above-mentioned image alignment process.

[0039] After aligning the real-time image with the depth image, according to the coordinates of each pixel point of the bounding box of the target obtained in the image, the corresponding area of the depth image is sliced and analyzed, that is, the depth value of each pixel point in the target bounding box is obtained. Among them, considering the limitation of the maximum and minimum detection distances of the depth camera, a typical range is, for example, the minimum working distance is 200 mm and the maximum working distance is 20,000 mm. When the object is within the minimum working distance or outside the maximum working distance, unavailable data is returned. For this reason, in this embodiment, the NAN (Not a Number) data type in the Numpy library of Python is used to identify the pixel points without depth values to indicate the default.

[0040] Step S4: Based on the depth value of each pixel point in the target bounding box area, use an adaptive width measurement and distance measurement algorithm to distinguish the target area and the background area in the target bounding box area, extract the overall depth value of the target area, determine the target distance and the target position, and then identify the four surrounding boundaries of the target in the target bounding box area to determine the target width and height, so as to complete the target positioning.

[0041] Step S4.1: Based on the depth value of each pixel point in the target bounding box area, use an adaptive width measurement algorithm to distinguish the target area and the background area in the target bounding box area, extract the overall depth value of the target area, and determine the target distance and the target position.

[0042] Considering the background interference existing in the distance measurement by the target detection algorithm, it may wrongly regard the background distance as the target distance, which may lead to inaccurate final target distance measurement. For this reason, this embodiment proposes a more accurate vision-based distance measurement method, including: Step S4.1.1: After obtaining the coordinates and depth values of each pixel point of the target bounding box, first determine whether there is background interference in the target bounding box. Specifically, ignore all NAN values in the obtained target bounding box area, and then make a comparison and judgment in combination with a preset difference threshold (usually select the maximum width of the measured object category at each angle as the difference threshold). When the difference between the minimum depth value and the maximum depth value in the target bounding box area is greater than the difference threshold, it indicates that there is a background in the image that is deeper than the detected object, and at this time, it is determined that there is background interference in the current area, and the subsequent step S4.1.2 is executed; otherwise, it is considered that there is no background interference in the image, and the subsequent step S4.1.3 is executed.

[0043] Step S4.1.2: After determining the existence of background interference, since the depth value of the background is usually much larger than that of the target, the minimum depth value and the maximum depth value of all pixel points in this area can be extracted. Take the average value of the minimum depth value and the maximum depth value, and use this average value as the depth boundary between the background and the target. Use the depth value less than or equal to the average value as the depth value of the target pixel, and use the depth value greater than the average value as the background pixel depth value to distinguish the target area and the background area in the target bounding box area. Then, take the average value of all target pixel depth values less than or equal to the average value to obtain the overall average depth distance (i.e., the average depth value) of the target area. This distance is the target distance, which has stronger robustness and accuracy.

[0044] Preferably, select the corresponding comparison threshold according to the distance between the foreground and the background in the actual environment. For example, when the distance between the background and the foreground is relatively small, as small values as possible such as 1 / 3, 1 / 4 of their sum, etc. should be used. Otherwise, larger values such as the average value are used. In this way, more accurate distinction between the background and the foreground can be realized.

[0045] Step S4.1.3: When it is determined that there is no background interference in the image, take the average value of all target pixel depth values within the target detection box area to obtain the overall average depth distance (i.e., the average depth value) of the target area.

[0046] Step S4.1.4: According to the determined target pixel coordinates, after converting the camera pixel coordinates to the world coordinates, obtain the actual position of the target in the world coordinates, that is, the target position. Among them, the coordinate conversion formula is: ; Among them, represents the internal parameter matrix of the camera, is the principal point position, that is, the origin of the depth image coordinate system (also known as the center of the pixel coordinate system), 、 respectively represent the rotation matrix and the translation vector in the external parameter matrix, 、 respectively represent the pixel point coordinates to be converted, 、 、 represent the converted world coordinate values.

[0047] Step S4.2: Based on the depth value of each pixel point in the target bounding box area, use the adaptive width measurement algorithm to identify the four perimeters of the target in the target bounding box area and determine the target width and height.

[0048] Considering that when the target detection model draws the target bounding box, the target bounding box may not fit perfectly, and there are non-target gaps at the boundaries. The existence of these gaps may affect the subsequent acquisition and calculation of the target width. Moreover, as a default value and invalid value, NAN cannot be used for operations. Therefore, the existence of NAN may cause a fatal error in directly measuring the width using the left and right breakpoints. Therefore, to achieve more accurate target width measurement, this embodiment proposes a method for determining the target width and height, as follows: Figure 4 As shown in: First, based on the depth values of each pixel point in the target bounding box area, perform column-by-column NAN judgment from left to right and row-by-row NAN judgment from top to bottom respectively to determine the four-sided boundaries of the target and optimize the target bounding box. Among them, the column-by-column or row-by-row NAN judgment is as follows: if all pixel points in the current column or row are NAN, then ignore this column or row and perform the NAN judgment on the next column or row until the current column or row is not all NAN. Through this method, the bounding box that cannot fit perfectly can be further optimized to more accurately represent the left, right, upper, and lower boundaries of the target.

[0049] Second, based on the optimized target bounding box, find the minimum value for the columns of the left and right boundaries and the minimum value for the rows of the upper and lower boundaries respectively to determine the left and right endpoint coordinates and the upper and lower endpoint coordinates of the target, so as to exclude the interference of the background. After the conversion between the camera pixel coordinates and the world coordinates, the world coordinate of the left endpoint of the target is obtained and the world coordinate of the right endpoint . Calculate to obtain the actual width of the target; similarly, the world coordinates of the upper endpoint and the lower endpoint of the target can also be obtained, and then the actual height of the target can be calculated.

[0050] Embodiment 2 This embodiment provides a machine vision target positioning system based on adaptive adjustment of the bounding box, including: An image acquisition module for acquiring real-time images and depth images of the current monitoring environment; A target bounding box extraction module for inputting the real-time image into a lightweight target detection model based on improved YOLO to identify the target in the image and extract the target bounding box; A depth value mapping module for aligning the real-time image and the depth image, corresponding the target bounding box area in the real-time image to the depth image, and determining the depth value of each pixel point in the target bounding box area according to the coordinates of each pixel point in the target bounding box area; among them, the pixel points without depth values are marked as the NAN data type; A target positioning module, which is used to distinguish the target area and the background area in the target bounding box area based on the depth value of each pixel point in the target bounding box area, adopt an adaptive width measurement and ranging algorithm, extract the overall depth value of the target area, determine the target distance and target position, and then identify the four perimeters of the target in the target bounding box area to determine the target width and height, so as to complete target positioning.

[0051] As Figure 5 shown, in the system proposed in this embodiment, it specifically includes an industrial control computer, a vision sensor for real-time monitoring of the presence of a target (including a color camera, an infrared camera, and a depth camera), a photosensitive sensor for detecting the environmental light intensity, other peripheral expansion modules, and a human-computer interaction module for facilitating user monitoring and custom adjustment. Among them, the image acquisition module includes a vision sensor and a photosensitive sensor. The vision sensor uses a color camera and an infrared camera to capture images of the surrounding environment, obtains depth images through the left and right infrared cameras and a structured light projector, and is connected to the industrial control computer through a communication signal to transmit image data.

[0052] In addition, the industrial control computer is equipped with a target bounding box extraction module, a depth value mapping module, and a target positioning module. It stores a target detection model, trains the model through a training data set to improve the detection accuracy, and then uses the target detection model to perform inference and decision on real-time images, extracts the target bounding box, combines the depth map obtained by the depth camera, and applies an adaptive width measurement and ranging algorithm to obtain accurate world coordinates (i.e., target position), the closest distance (i.e., target distance), and target positioning information such as the target width and height of the target. Further, the control module communicates with other peripheral modules through the communication module according to the obtained target positioning information, can realize the remote acquisition of target positioning information, and can also realize other expansion functions.

[0053] Embodiment III This embodiment provides an electronic device, including: a memory for storing executable instructions; a processor for implementing the above method provided in this embodiment when executing the executable instructions stored in the memory.

[0054] Embodiment IV This embodiment also provides a computer-readable storage medium storing executable instructions, which when executed by a processor will cause the processor to execute the above method provided in this embodiment.

[0055] Embodiment V This embodiment provides a computer program product, which includes executable instructions. The executable instructions are a type of computer instructions, and the executable instructions are stored in a computer-readable storage medium. When the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the electronic device is caused to execute the above method provided by this embodiment.

[0056] The steps involved in the above Embodiments 2 to 5 correspond to those in Method Embodiment 1. For the specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.

[0057] Those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. The present invention is not limited to any specific combination of hardware and software.

[0058] The above are only the preferred embodiments of the present invention. Although the specific implementation manners of the present invention are described in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications or deformations that can be made without creative efforts on the basis of the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. A machine vision target positioning method based on adaptive adjustment of bounding boxes, characterized in that Including: Obtain the real-time image and depth image of the current monitoring environment; Input the real-time image into the lightweight object detection model based on improved YOLO to identify the objects in the image and extract the object bounding boxes; Align the real-time image and the depth image, map the object bounding box region in the real-time image to the depth image, and determine the depth value of each pixel point in the object bounding box region according to the coordinate of each pixel point in the object bounding box region; wherein, the pixel points without depth values are labeled as NAN data type; Based on the depth value of each pixel point in the object bounding box region, adopt the adaptive width measurement and distance measurement algorithm to distinguish the object region and the background region in the object bounding box region, extract the overall depth value of the object region, determine the object distance and object position, and then identify the four perimeters of the object in the object bounding box region to determine the object width and height, so as to complete object positioning.

2. The machine vision target positioning method based on boundary box adaptive adjustment according to claim 1, characterized in that Obtain the real-time image of the current monitoring environment, including: Obtain the light intensity of the current monitoring environment in real time; According to the light intensity, select the image capture method to obtain the real-time image, including: When the light intensity is greater than the set value, start the color camera to capture the real-time image; When the light intensity is less than the set value, start the infrared camera to capture the real-time image.

3. The machine vision target positioning method based on adaptive adjustment of bounding boxes according to claim 1, characterized in that, The lightweight object detection model based on improved YOLO is based on the YOLO11 model, including a backbone network, a neck network and a head network. Among them, the SaE attention module is introduced into the neck network, and the SaE attention module is combined with the S3k2 module to form the C3k2SaE module; the TDDown module is introduced into the backbone network and the neck network; In the SaE attention module, the input features are compressed by two convolutional layers, and after compression, the importance weights of each channel of the compressed features are extracted through the adaptive two-dimensional average pooling operation in the SaE module. After the importance weights extracted are aggregated by the fully connected layer, channel weighting is performed on the input features according to the importance weights of each channel to obtain the depth feature representation; In the TDDown module, the input features are processed in parallel through two parallel depth convolutional operations. After the two processed features are fused, depth convolutional operation is performed for feature downsampling to extract low-dimensional features.

4. A machine vision target positioning method based on adaptive adjustment of bounding boxes according to claim 1, characterized in that, The determination of the object distance and object position includes: According to the coordinate and depth value of each pixel point of the object bounding box, judge whether there is background interference in the object bounding box; wherein, all NAN values in the obtained object bounding box region are ignored, and then the minimum depth value and the maximum depth value of all pixel points in this region are extracted. The difference between the minimum depth value and the maximum depth value in the object bounding box region is compared with the preset difference threshold, and whether there is background interference is judged according to the comparison result; If it is determined that there is background interference, based on the depth value of each pixel point in the target bounding box area, the average value of the minimum depth value and the maximum depth value is taken, and the depth value of each pixel point is compared with this average value. The depth value less than or equal to the average value is used as the depth value of the target pixel, and the depth value greater than the average value is used as the background pixel depth value, so as to distinguish the target area and the background area in the target bounding box area; then the average value of all target pixel depth values less than or equal to the average value is taken to obtain the overall average depth distance of the target area, that is, the target distance. If it is determined that there is no background interference, the average value of all target pixel depth values within the target detection box area is taken to obtain the overall average depth distance of the target area, that is, the target distance. According to the determined target pixel coordinates, after the transformation between the camera pixel coordinates and the world coordinates, the actual position of the target is obtained, that is, the target position.

5. The machine vision target positioning method based on adaptive adjustment of bounding boxes according to claim 1, characterized in that, The determination of the target width and height includes: Based on the depth value of each pixel point in the target bounding box area, perform column-by-column NAN judgment from left to right and row-by-row NAN judgment from top to bottom respectively to determine the four-side boundaries of the target and optimize the target bounding box. Based on the optimized target bounding box, find the minimum value of the columns on the left and right boundaries and the minimum value of the rows on the upper and lower boundaries respectively to determine the coordinates of the left and right endpoints and the upper and lower endpoints of the target. After the transformation between the camera pixel coordinates and the world coordinates, the world coordinate values of the left and right endpoints and the upper and lower endpoints of the target are obtained, so as to determine the actual width and height of the target.

6. The machine vision target positioning method based on boundary box adaptive adjustment according to claim 5, wherein The column-by-column or row-by-row NAN judgment is as follows: if all pixel points in the current column or row are NAN, then ignore this column or row and perform the NAN judgment on the next column or row until the current column or row is not all NAN.

7. A machine vision target positioning system based on adaptive adjustment of bounding boxes, characterized in that, It includes: An image acquisition module, which is used to acquire the real-time image and depth image of the current monitoring environment. A target bounding box extraction module, which is used to input the real-time image into a lightweight target detection model based on improved YOLO to identify the target in the image and extract the target bounding box. A depth value mapping module, which is used to align the real-time image and the depth image, map the target bounding box area in the real-time image to the depth image, and determine the depth value of each pixel point in the target bounding box area according to the coordinate of each pixel point in the target bounding box area; among them, the pixel points without depth values are marked as NAN data type. A target positioning module, which is used to distinguish the target area and the background area in the target bounding box area based on the depth value of each pixel point in the target bounding box area, adopt an adaptive width measurement and ranging algorithm, extract the overall depth value of the target area, determine the target distance and target position, and then identify the four-side boundaries of the target in the target bounding box area to determine the target width and height, so as to complete the target positioning.

8. An electronic device, characterized in that, It includes: A memory, which is used to store executable instructions. A processor, which is used to implement the machine vision target positioning method according to any one of claims 1-6 when executing the executable instructions stored in the memory.

9. A computer-readable storage medium, characterized in that, Stored with executable instructions, when the processor is caused to execute the executable instructions, it implements a machine vision target positioning method based on adaptive adjustment of a bounding box according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes executable instructions, and the executable instructions are stored in a computer-readable storage medium; When the processor of the electronic device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, it implements a machine vision target positioning method based on adaptive adjustment of a bounding box according to any one of claims 1-6.

Citation Information

Patent Citations

  • Target detection method and device

    CN111950543A

  • Robot target detection method and system based on DW-SEnet model

    CN114842320A

  • Method for detecting and positioning blast beads of high-speed moving filter stick

    CN115526908A

  • Robot positioning method and system based on binocular vision and laser scanning fusion

    CN118050734A

  • Dynamic scene RGB-D SLAM method based on target detection and deep clustering

    CN118918184A