Indoor target positioning method and system based on machine vision and deep learning

By applying machine vision and deep learning technology in the workshop, combined with Bluetooth beacon positioning, the three-dimensional representation of the product is constructed and tracked its motion trajectory, the problem of target positioning and tracking in complex lighting and occlusion environments in the workshop is solved, and high-precision target positioning and intelligent production management are achieved.

CN119941854AActive Publication Date: 2025-05-06GUANGDONG MECHANICAL & ELECTRICAL COLLEGE

Patent Information

Application Number
CN202510016300.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

In the complex lighting and occlusion environment of the workshop, traditional object detection algorithms are difficult to achieve accurate positioning and real-time tracking of product targets, especially when production equipment and workers are dynamically occluding.

Method used

Using an indoor target positioning method based on machine vision and deep learning, a suitable image preprocessing algorithm is selected by analyzing the characteristics of occluded objects in real-time video; combining Bluetooth beacon positioning data, the occluded products are detected and spatially positioned; a three-dimensional representation of the product is constructed and its motion trajectory is tracked; when occlusion causes tracking failure, a target recapture mechanism based on Bluetooth positioning is triggered.

Benefits of technology

It realizes high-precision positioning and continuous tracking of product targets in complex workshop environments, solves the problem of interference between occlusions and positioning tracking, and improves the automation and intelligence level of the production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941854A_ABST
    Figure CN119941854A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target positioning, in particular to an indoor target positioning method and system based on machine vision and deep learning, and the method comprises the steps: selecting a proper image preprocessing algorithm through analyzing the features of a shielding object in a real-time video; in combination with Bluetooth beacon positioning data, detection and spatial positioning are carried out on the shielded product; three-dimensional representation of the product is constructed, and the motion trail of the product is When the shielding causes tracking failure, triggering a target recapture mechanism based on Bluetooth positioning; according to the invention, the problem of interference of shielding objects on product positioning and tracking in a workshop environment is solved, high-precision positioning and continuous tracking of products under complex conditions are realized through multi-source data fusion and an intelligent algorithm, reliable real-time position information is provided for workshop production management, and the automation and intelligent levels of the production process are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target positioning technology, and in particular to an indoor target positioning method and system based on machine vision and deep learning. Background Art

[0002] In the production workshops in the factory, especially in the workshops of electronics factories and chemical plants, high-position capture cameras are usually installed. These cameras are mainly used to monitor the production process, ensure the quality and safety of the products, and locate and track the target products in real time. At the same time, based on the video images collected by the camera, the products in the production workshop in the video images can be captured to realize the intelligent, digital, and information-based tracking and positioning of the production workshop, so as to timely grasp the production progress and quality status of the products. However, the lighting conditions inside the workshop are complex and changeable, and the frequent changes in light cause serious interference to the target positioning system based on machine vision. In addition, there are many production equipment in the workshop, and the movement and operation of workers often block the target products. These blockages not only hinder the visual acquisition of the target products, but also introduce a lot of background noise, which brings great difficulties to target positioning. Since the equipment and workers on the production line are not fixed, their blockage of the target products is dynamically changing, and the size, shape, and position of the blockages are constantly changing. Traditional target detection algorithms are difficult to adapt to such complex application scenarios. At the same time, the target products move very fast on the production line, which places very high demands on real-time performance. Frequent occlusions will lead to target tracking failures. How to quickly recapture the target after occlusion occurs is also a major technical challenge. In summary, in the complex lighting and occlusion environment of the workshop, achieving accurate positioning and real-time tracking of product targets is a technical problem that needs to be solved urgently. Summary of the invention

[0003] The purpose of the present invention is to provide an indoor target positioning method and system based on machine vision and deep learning, so as to achieve accurate positioning and real-time tracking of product targets.

[0004] In order to achieve the above object, the present invention provides the following technical solutions:

[0005] In a first aspect, an embodiment of the present invention provides an indoor target positioning method based on machine vision and deep learning, the method comprising the following steps:

[0006] S100, obtaining a real-time video image in the workshop, analyzing the movement trajectory and size of the obstruction in the image, determining the frequency and duration of the obstruction, and determining the degree of influence of the obstruction on the target product;

[0007] S200, based on the degree of influence of the occlusion on the target product, preprocess the video image, analyze the transparency and refractive index properties of the pixels in the occluded area, infer the material type of the occluded object, and select the occluded area processing method of the material type to obtain the preprocessed video image;

[0008] S300, obtaining an estimated distance and azimuth between the target product and the Bluetooth beacon, detecting the position of the blocked target product according to the preprocessed video image, spatially registering the position detection result of the target product with the Bluetooth beacon data, obtaining an estimated distance and azimuth of the Bluetooth beacon relative to the target product according to Bluetooth positioning, detecting and locating the target product by constructing a spatial topological relationship between the product and the beacon, and obtaining the position coordinates and bounding box information of the target product;

[0009] S400, tracking and predicting the motion state of the target product according to the position coordinates and bounding box information of the target product in combination with the Bluetooth positioning data, and obtaining the motion trajectory and speed information of the target product;

[0010] S500, during the target product tracking process, if an obstruction blocks the target product, resulting in tracking failure, a local search is performed in an area where the target product may appear based on the historical motion trajectory of the target product and the current Bluetooth positioning estimated position, until the motion trajectory and speed information of the target product are recaptured;

[0011] S600, based on the target product's motion trajectory and speed information, obtains the target product's depth information through a depth camera, combines the distance and azimuth estimated by Bluetooth positioning, constructs a three-dimensional representation of the target product, generates the position and tracking motion information of the indoor target product, and transmits it to the indoor workshop production management system to achieve target positioning of indoor products that are blocked during the production process.

[0012] Optionally, in S100, the real-time video image in the workshop is acquired, and the frequency and duration of the occlusion are determined by analyzing the movement trajectory and size of the occlusion in the image, so as to determine the impact of the occlusion on the target product, including:

[0013] S110, acquiring real-time video streams collected by multiple cameras, and performing image preprocessing on the video streams, wherein the image preprocessing includes median filtering, histogram equalization, and Sobel operator edge sharpening;

[0014] S120, constructing a scene background using a Gaussian mixture model, performing a difference operation between the current frame and the scene background, and obtaining an initial outline of the occluder;

[0015] S130, matching and associating the initial outline of the occluder between consecutive frames to obtain a motion trajectory of the occluder on the image plane;

[0016] S140, performing a time series analysis on the movement trajectory of the obstruction, calculating the number of times the obstruction appears in a specific area and the length of time it stays there, and obtaining statistical data on the frequency and duration of the obstruction;

[0017] S150, according to the spatial position and time distribution of the obstruction, find the affected target products and analyze the impact of the obstruction on each target product; set differentiated impact weights for different production stages to determine the impact of the obstruction on product quality and production efficiency.

[0018] Optionally, in S200, based on the degree of influence of the occlusion on the target product, the video image is preprocessed, the transparency and refractive index properties of the pixels in the occluded area are analyzed, the material type of the occluded object is inferred, and the occluded area processing method of the material type is selected to obtain the preprocessed video image, including:

[0019] S210, acquiring a real-time video image captured by a camera, locating the outline of the obstruction, and obtaining a pixel set of the obstruction area;

[0020] S220, calculating brightness distribution characteristics and gradient information according to the pixel set of the occluded area, analyzing the spatial relationship between the occluded area and the target product in combination with the location information of the target product, and determining a preliminary impact range;

[0021] S230, extracting a material feature vector according to the texture features and color histogram of the occluded area, and mapping the feature vector to a predefined material category through a support vector machine classifier; if the material category is glass, using an Alpha blending algorithm to process the occluded area; if the material category is translucent plastic, using a transparency-based image enhancement algorithm to process the occluded area; if the material category is opaque metal, using an edge-preserving filtering algorithm to process the occluded area;

[0022] S240, optimizing the parameters of the occluded area processing method, traversing the parameter space using a grid search method, evaluating the quality of the processing result using a structural similarity index, and applying the optimized method to the processing of the occluded area of ​​the original image to obtain a preprocessed image.

[0023] Optionally, in S300, the step of obtaining an estimated distance and azimuth between the target product and the Bluetooth beacon, detecting the position of the blocked target product according to the preprocessed video image, spatially registering the position detection result of the target product with the Bluetooth beacon data, obtaining an estimated distance and azimuth of the Bluetooth beacon relative to the target product according to Bluetooth positioning, detecting and locating the target product by constructing a spatial topological relationship between the product and the beacon, and obtaining the position coordinates and bounding box information of the target product includes:

[0024] S310, acquiring signal strength data transmitted by Bluetooth beacons, selecting a signal strength attenuation model according to the signal strength data; and calculating an estimated distance between a target product and each Bluetooth beacon according to the signal strength attenuation model;

[0025] S320, analyzing the preprocessed image to obtain the position and size information of the target product in the image, and determining the initial position of the target product by combining the estimated distance and the position and size information of the target product in the image;

[0026] S330, calculating the distance and azimuth of the target product relative to each beacon, and using an adaptive weighting method based on signal strength to perform weighted averaging on the distance and azimuth of the target product relative to each beacon to obtain a positioning result of the target product; wherein the weight in the adaptive weighting method is positively correlated with the signal strength and is mapped to a range of 0 to 1 through a sigmoid function;

[0027] S340, establishing a spatial topological relationship between the target product and the Bluetooth beacon according to the initial position and positioning result of the target product, and using an extended Kalman filter algorithm to estimate and track the position of the target product in real time based on the spatial topological relationship and in combination with the initial position and positioning result of the target product; wherein the state vector includes position, velocity and acceleration, and the observation vector includes image detection position and Bluetooth positioning result;

[0028] S350, using the extended Kalman filter algorithm to estimate the initial position of the target product in real time to optimize the position coordinates of the target product; and introducing the particle filter algorithm for occlusion, by maintaining multiple hypothetical particles to handle multimodal distribution, to improve tracking stability under occlusion; wherein the state vector of the extended Kalman filter algorithm includes position, velocity and acceleration, and the observation vector includes image detection position and Bluetooth positioning results;

[0029] S360, maps the optimized position coordinates back to the image plane, and combines them with the bounding box information of the target detection to obtain the precise position and size of the target product in the image.

[0030] Optionally, in S400, the tracking and predicting the motion state of the target product according to the position coordinates and bounding box information of the target product in combination with the Bluetooth positioning data to obtain the motion trajectory and speed information of the target product includes:

[0031] S410, obtaining the location coordinates, bounding box information and Bluetooth positioning data of the target product;

[0032] S420, using a Kalman filter algorithm to fuse the position coordinates, the bounding box information, and the Bluetooth positioning data to obtain fused position data; wherein the state vector of the Kalman filter algorithm includes a three-dimensional position and a speed, and the observation vector includes a two-dimensional coordinate and a three-dimensional coordinate;

[0033] S430, fitting the motion trajectory of the target product by least square method according to the fused position data; constructing a motion state model of the target product based on the motion trajectory and the fused position data; wherein the motion state model includes a linear regression model and a long short-term memory network model;

[0034] S440, predicting the future trajectory of the target product according to the motion state model; wherein the linear regression model is used to predict the short-term future position, and the long short-term memory network model is used to capture the long-term motion pattern; and optimizing the future trajectory using a particle filter algorithm to obtain an optimized future trajectory.

[0035] Optionally, in S500, during the target product tracking process, if an obstruction blocks the target product, resulting in tracking failure, a local search is performed in an area where the target product may appear according to the historical motion trajectory of the target product and the position estimated by the current Bluetooth positioning, until the motion trajectory and speed information of the target product are recaptured, including:

[0036] S510, receiving continuous frame images and target detection confidence information, and determining whether occlusion has occurred and caused tracking failure according to the continuous frame images and the target detection confidence information, and triggering a target recapture mechanism if the target detection confidence is lower than a preset threshold or the number of consecutive undetected frames exceeds a preset number of frames;

[0037] S520, obtaining the location and time information of the last successful tracking of the target product, and using the Kalman filter algorithm to predict the current location of the target product;

[0038] S530, receiving target product position estimation information, fusing the Kalman filter prediction result and the Bluetooth positioning result through a covariance-based adaptive weight allocation method to obtain a fused estimate of the target product position;

[0039] S540, determining a local search area according to the fusion estimation of the target product position, and recapturing the target using an adaptive grid search algorithm in the local search area; if the target is detected in a certain grid, refining the search of the grid; receiving continuous multi-frame tracking results, and confirming that the target is captured successfully if the number of continuous tracking frames exceeds a preset frame number threshold;

[0040] S550, obtaining historical trajectory information, Bluetooth positioning data and new visual tracking results of the target product, back-estimating the motion state of the target product during the occlusion period, and obtaining the motion trajectory and speed information of the target product.

[0041] Optionally, in S600, based on the target product motion trajectory and speed information, the depth information of the target product is obtained through a depth camera, and the distance and azimuth estimated by Bluetooth positioning are combined to construct a three-dimensional representation of the target product, generate the position and tracking motion information of the indoor target product, and transmit it to the indoor workshop production management system to achieve the target positioning of the indoor product blocked during the production process, including:

[0042] S610, obtaining the motion trajectory and speed information of the target product, and smoothing the information using a Kalman filter algorithm to obtain filtered motion data;

[0043] S620, collecting a depth image of the target product through a depth camera, and extracting a depth profile of the target product;

[0044] S630, performing illumination unevenness correction on the depth image using a three-scale Retinex algorithm, and improving detail performance of the depth image using a CLAHE algorithm;

[0045] S640, constructing a three-dimensional point cloud representation of the target product based on the processed depth information and Bluetooth positioning data;

[0046] S650, registering the three-dimensional point cloud representation with a pre-established CAD model of the target product to determine the precise position and posture of the target product; if the mean absolute error of the precise position and posture exceeds a preset threshold, triggering anomaly detection;

[0047] S660: Pack the precise position and filtered motion data and transmit them to an indoor workshop production management system via industrial Ethernet.

[0048] In a second aspect, an embodiment of the present invention provides an indoor target positioning system based on machine vision and deep learning, the system comprising:

[0049] at least one processor;

[0050] at least one memory for storing at least one program;

[0051] When the at least one program is executed by the at least one processor, the at least one processor implements the method as described in any one of the first aspects.

[0052] The beneficial effects of the present invention are as follows: the present invention discloses an indoor target positioning method and system based on machine vision and deep learning, which selects a suitable image preprocessing algorithm by analyzing the characteristics of obstructions in real-time video; detects and spatially locates obstructed products in combination with Bluetooth beacon positioning data; constructs a three-dimensional representation of the product and tracks its motion trajectory; and triggers a target recapture mechanism based on Bluetooth positioning when tracking fails due to occlusion. The present invention solves the problem of interference of obstructions on product positioning and tracking in a workshop environment, and realizes high-precision positioning and continuous tracking of products under complex conditions through multi-source data fusion and intelligent algorithms, providing reliable real-time location information for workshop production management, and effectively improving the automation and intelligence level of the production process. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0054] Figure 1 It is a flow chart of an indoor target positioning method based on machine vision and deep learning in an embodiment of the present invention;

[0055] Figure 2 It is a structural diagram of an indoor target positioning system based on machine vision and deep learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects of the present invention, so as to fully understand the purpose, scheme and effect of the present invention. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict.

[0057] In the process of Bluetooth target positioning based on machine vision and deep learning technology for the production workshop in the factory, the problems encountered are: the lighting environment in the workshop is unstable, and the light is sometimes strong and sometimes weak; at the same time, the moving equipment and workers on the production line will frequently block the target object, thereby interfering with the detection of the target product. In order to solve this technical problem, the present invention provides an indoor target positioning method and system based on machine vision and deep learning.

[0058] See also Figure 1 , Figure 1 The present invention provides an indoor target positioning method based on machine vision and deep learning, and the method comprises the following steps:

[0059] S110, obtaining a real-time video image in the workshop, analyzing the movement trajectory and size of the obstruction in the image, determining the frequency and duration of the obstruction, and determining the degree of influence of the obstruction on the target product;

[0060] S120, preprocessing the video image based on the degree of influence of the occlusion on the target product, analyzing the transparency and refractive index properties of the pixels in the occluded area, inferring the material type of the occluded object, and selecting an occluded area processing method of the material type to obtain a preprocessed video image;

[0061] S130, obtaining an estimated distance and azimuth between the target product and the Bluetooth beacon, performing position detection on the blocked target product according to the preprocessed video image, spatially registering the position detection result of the target product with the Bluetooth beacon data, obtaining an estimated distance and azimuth of the Bluetooth beacon relative to the target product according to Bluetooth positioning, detecting and locating the target product by constructing a spatial topological relationship between the product and the beacon, and obtaining the position coordinates and bounding box information of the target product;

[0062] S140, tracking and predicting the motion state of the target product according to the position coordinates and bounding box information of the target product in combination with the Bluetooth positioning data, to obtain the motion trajectory and speed information of the target product;

[0063] S150, during the target product tracking process, if an obstruction blocks the target product, resulting in tracking failure, a local search is performed in an area where the target product may appear based on the historical motion trajectory of the target product and the current Bluetooth positioning estimated position, until the motion trajectory and speed information of the target product are recaptured;

[0064] S160, based on the target product's motion trajectory and speed information, obtains the target product's depth information through a depth camera, combines the distance and azimuth estimated by Bluetooth positioning, constructs a three-dimensional representation of the target product, generates the position and tracking motion information of the indoor target product, and transmits it to the indoor workshop production management system to achieve target positioning of indoor products that are blocked during the production process.

[0065] In the embodiments provided by the present invention, by analyzing the features of obstructions in real-time video, a suitable image preprocessing algorithm is selected; in combination with Bluetooth beacon positioning data, the obstructed product is detected and spatially positioned; a three-dimensional representation of the product is constructed and its motion trajectory is tracked; when the occlusion causes tracking failure, a target recapture mechanism based on Bluetooth positioning is triggered; the present invention solves the problem of interference of obstructions on product positioning and tracking in a workshop environment, and through multi-source data fusion and intelligent algorithms, high-precision positioning and continuous tracking of products under complex conditions are achieved, providing reliable real-time location information for workshop production management, and effectively improving the automation and intelligence level of the production process.

[0066] In some embodiments, in S100, the real-time video image in the workshop is acquired, and the frequency and duration of the occlusion are determined by analyzing the movement trajectory and size of the occlusion in the image, and the degree of influence of the occlusion on the target product is determined, including:

[0067] S110, acquiring real-time video streams collected by multiple cameras, and performing image preprocessing on the video streams, wherein the image preprocessing includes median filtering, histogram equalization, and Sobel operator edge sharpening;

[0068] S120, constructing a scene background using a Gaussian mixture model, performing a difference operation between the current frame and the scene background, and obtaining an initial outline of the occluder;

[0069] S130, matching and associating the initial outline of the occluder between consecutive frames to obtain a motion trajectory of the occluder on the image plane;

[0070] S140, performing a time series analysis on the movement trajectory of the obstruction, calculating the number of times the obstruction appears in a specific area and the length of time it stays there, and obtaining statistical data on the frequency and duration of the obstruction;

[0071] S150, according to the spatial position and time distribution of the obstruction, find the affected target products and analyze the impact of the obstruction on each target product; set differentiated impact weights for different production stages to determine the impact of the obstruction on product quality and production efficiency.

[0072] Exemplarily, real-time video streams captured by multiple high-definition cameras in the workshop are obtained, and image preprocessing is performed on the video frames, including using median filtering to denoise, histogram equalization to enhance contrast, and Sobel operator to sharpen edges to improve image quality. The Gaussian mixture model is used to construct the scene background, and the current frame is differentially calculated from the background to extract the foreground target and obtain the initial outline of the occluder. For the extracted outline of the occluder, the KCF algorithm is used to match and associate it between consecutive frames to obtain the motion trajectory of the occluder on the image plane. According to the calibration parameters of multiple cameras, the three-dimensional position and actual size of the occluder are calculated by the triangulation principle. The trajectory is smoothed using the Kalman filter to eliminate the influence of noise.

[0073] Perform time series analysis on the movement trajectory of the occlusion, calculate the number of times it appears and the length of time it stays in a specific area, and obtain the frequency and duration statistics of the occlusion. Set the frequency threshold and duration threshold. If the occlusion frequency exceeds the threshold or the duration of a single occlusion exceeds the threshold, it is determined to be abnormal occlusion. According to the movement characteristics of the occlusion, distinguish the types of occlusion caused by personnel movement and equipment movement, and preliminarily determine the impact range of the abnormal occlusion. Combined with the product production line layout information, build a mapping relationship between the product position and the camera field of view. Find the affected target products according to the spatial location and time distribution of the abnormal occlusion. Establish an impact assessment model, consider factors such as the size of the occlusion, the duration of the occlusion, the production line process and key control points, and quantify the impact score of the occlusion on each target product. Set differentiated impact weights for different production stages to comprehensively evaluate the impact of occlusion on product quality and production efficiency. Install multiple high-definition network cameras in the workshop with a resolution of 1920x1080 pixels and a frame rate of 30fps. The collected video stream is denoised by median filtering, and the filter window size is 3x3 pixels. Then, the image contrast is enhanced by histogram equalization, and the pixel intensity values ​​are redistributed to the range of 0-255. The Sobel operator is used for edge sharpening to highlight the target contour. The Gaussian mixture model uses 5 Gaussian components and the learning rate is set to 0.01 to construct the scene background. The foreground target is obtained by background subtraction, and the morphological opening operation is applied to remove small noise points. The structure element size is 5x5 pixels. The KCF algorithm uses the HOG feature descriptor, the feature unit size is 4x4 pixels, and the block size is 2x2 units. The bandwidth parameter σ of the Gaussian kernel function is set to 0.5. The filter weights are updated by minimizing the objective function, and the learning rate is set to 0.02. The Zhang Zhengyou calibration method is used for multi-camera calibration. The chessboard size is 9x6, and the actual size of each grid is 25mmx25mm. During triangulation, the baseline length is 1 meter, the focal length is 8mm, and the parallax accuracy is 0.1 pixel. The initial values ​​of the process noise covariance matrix Q and the measurement noise covariance matrix R of the Kalman filter are set to the unit matrix and adjusted through iterative optimization. The trajectory analysis of the occluder uses a sliding time window with a window size of 10 minutes and a step size of 1 minute. The number of occluder occurrences and the cumulative duration are counted in each window. The frequency threshold is set to 5 times every 10 minutes, and the duration threshold is set to 2 minutes cumulatively. Exceeding any threshold is considered an abnormal occlusion. Based on the average speed and acceleration characteristics of the occluder, a threshold is set to distinguish between people walking, with a speed of <1.5m / s and an acceleration of <0.5m / s 2and equipment movement. The initial impact range is determined by the projected area and movement range of the obstruction. The production line layout is represented in a grid with a grid size of 0.5mx0.5m. The position of each product in the production process is tracked in real time via Bluetooth. The impact assessment model takes into account the size of the obstruction, which accounts for 20% of the area, the duration of the obstruction, which accounts for 30%, the process stage of the product, which accounts for 30%, and the location of the obstruction, which accounts for 20%. The weights of different production stages are set according to the importance of the process, such as the weight of key quality inspection points is 1.5, and the weight of general processing points is 1.0. Finally, the impact score of 0-100 is obtained. A score greater than 80 is judged as a serious impact and needs to be handled immediately.

[0074] In some embodiments, in S200, the preprocessing of the video image based on the degree of influence of the occlusion on the target product, analyzing the transparency and refractive index properties of the pixels in the occluded area, inferring the material type of the occluded object, and selecting the occluded area processing method of the material type to obtain the preprocessed video image includes:

[0075] S210, acquiring a real-time video image captured by a camera, locating the outline of the obstruction, and obtaining a pixel set of the obstruction area;

[0076] S220, calculating brightness distribution characteristics and gradient information according to the pixel set of the occluded area, analyzing the spatial relationship between the occluded area and the target product in combination with the location information of the target product, and determining a preliminary impact range;

[0077] S230, extracting a material feature vector according to the texture features and color histogram of the occluded area, and mapping the feature vector to a predefined material category through a support vector machine classifier; if the material category is glass, using an Alpha blending algorithm to process the occluded area; if the material category is translucent plastic, using a transparency-based image enhancement algorithm to process the occluded area; if the material category is opaque metal, using an edge-preserving filtering algorithm to process the occluded area;

[0078] S240, optimizing the parameters of the occluded area processing method, traversing the parameter space using a grid search method, evaluating the quality of the processing result using a structural similarity index, and applying the optimized method to the processing of the occluded area of ​​the original image to obtain a preprocessed image.

[0079] For example, real-time video images captured by a high-definition camera in the workshop are obtained, and the GrabCut algorithm is used to accurately locate the outline of the obstruction to obtain a pixel set of the obstruction area. The brightness distribution characteristics and gradient information of the pixels in the obstruction area are calculated, and the spatial relationship between the obstruction area and the target product is analyzed in combination with the known location information of the target product to determine the initial impact range. The material feature vector is extracted based on the texture features and color histogram of the obstruction area. The feature vector is mapped to a predefined material category, such as glass, plastic, metal, etc., using a pre-trained support vector machine classifier. The approximate transparency and reflectivity of the obstruction are estimated based on the ambient lighting conditions and the reflective characteristics of the obstruction surface. For the inferred material type, a suitable processing method is selected from the preset obstruction area processing algorithm library. If the material is glass, the Alpha blending algorithm is used; if the material is translucent plastic, the transparency-based image enhancement algorithm is used; if the material is opaque metal, the edge-preserving filtering algorithm is applied. According to the material characteristics and the degree of occlusion, the initial parameters of the algorithm are set, such as the Alpha value, the enhancement strength, the filter kernel size, etc. The parameters of the selected occlusion area processing algorithm are optimized, and the grid search method is used to traverse the parameter space. The structural similarity index is used to evaluate the quality of the processing result after each iteration. When the evaluation index reaches the preset threshold or the number of iterations exceeds the upper limit, the optimization process is terminated. The optimized algorithm is applied to the occluded area of ​​the original image to obtain the preprocessed image. By comparing the visibility changes of the target product before and after the treatment, the occlusion impact score is calculated for subsequent product quality control and production scheduling decisions. A high-definition camera with a resolution of 1920x1080 pixels is installed in the workshop to collect real-time video streams with a frame rate of 30fps. The GrabCut algorithm is used to segment the occluder for each frame of the image, and the number of iterations is set to 5. The Gaussian mixture model of the foreground and background uses 5 components each. The brightness histogram of the pixels in the occluded area is calculated, and the 256-level grayscale value is divided into 16 bins to obtain the brightness distribution feature vector. The Sobel operator is used to calculate the gradient information, and the gradient amplitude threshold is set to 50. Through the pre-calibrated camera parameters, the image coordinate system is converted to the world coordinate system, and the minimum Euclidean distance between the occluded area and the product is calculated in combination with the known three-dimensional position (x, y, z) of the target product. If the distance is less than 0.5 meters, it is judged as a direct impact. The grayscale co-occurrence matrix of the occluded area is extracted as the texture feature, and four statistics of energy, contrast, homogeneity and entropy are calculated. The color features are represented by the RGB color histogram of the 32bin interval. The texture and color features are spliced ​​into a 64-dimensional vector and input into the pre-trained SVM classifier, which uses the RBF kernel function, C=10, gamma=0.1, to classify the occluder as glass, plastic or metal.For glass materials, the Alpha blending algorithm is used, and the initial Alpha value is set to 0.7; for translucent plastics, the contrast limited adaptive histogram equalization (CLAHE) algorithm is used, the block size is 8x8 pixels, and the contrast limit threshold is 3; for opaque metals, bilateral filtering is applied, the spatial domain standard deviation is set to 3, and the value domain standard deviation is 50. The algorithm parameters are optimized by grid search, and the Alpha value is searched in the range of 0.5-0.9 with a step size of 0.1. The block size of CLAHE takes 4 values ​​from 4x4 to 16x16, and the two standard deviations of bilateral filtering are searched in the range of 1-5 and 30-70 respectively. The structural similarity index (SSIM) is calculated after each iteration, and it stops when the SSIM is greater than 0.9 or the number of iterations reaches 50. Finally, the edge strength and texture complexity of the target product area before and after the treatment are compared, and the visibility improvement percentage is calculated. A value greater than 30% is considered to be significantly improved, which is used as the basis for evaluating the degree of occlusion impact.

[0080] For opaque occluders, the true value of the occluded pixel is estimated by searching for similar image blocks in the neighborhood of the occluded area, and the original image content is replaced with the repair result. For semi-transparent and transparent occluders, multiple frames of images of the occluded area are obtained, and the transmittance of the occluded object is estimated by using the image changes caused by the movement of the occluder between different frames. The transmittance is combined with the background image information to restore the target product image without the influence of the occlusion.

[0081] A real-time video sequence captured by a camera in a workshop is obtained, wherein the video sequence includes a plurality of continuous frame images; the continuous frame images are processed by a U-Net semantic segmentation algorithm to obtain an occluded area in each frame image; a pre-trained ResNet model is used to classify the occluded area to determine the transparency attribute of the occluded object; if the occluded object is an opaque occluded object, the PatchMatch algorithm is used to perform image restoration on the occluded area; if the occluded object is a semi-transparent or transparent occluded object, a multi-scale Lucas-Kanade method is used to estimate the optical flow field of the continuous frame images; background modeling is performed by a Gaussian mixture model, and a weighted linear mixture model is established according to the optical flow field estimation result and the background model; the inverse problem of the weighted linear mixture model is solved to obtain a target product image with occlusion removed; adaptive histogram equalization is applied to the target product image to obtain a restored image with enhanced contrast; a structural similarity index between the restored image and the original image is calculated to determine whether the structural similarity index is lower than a preset threshold; if the structural similarity index is lower than the preset threshold, reprocessing is performed.

[0082] For example, a real-time video sequence captured by a high-definition camera in a workshop is obtained, and the occluded area in each frame is extracted by the U-Net semantic segmentation algorithm. The occluded area is classified using a pre-trained ResNet model to determine the transparency attribute of the occluded object. According to the classification results, the occluded objects are divided into three categories: opaque, semi-transparent, and transparent, providing a basis for subsequent processing. For opaque occluded objects, a search radius is set around the occluded area, and the PatchMatch algorithm is used for image restoration. The nearest neighbor field is constructed, and the matching results are iteratively optimized by random sampling and local propagation. The image block with the highest similarity score is selected as the restoration source, and the true value of the occluded pixel is estimated by bilinear interpolation. The Poisson image editing technology is used to seamlessly merge the restoration result into the original image and replace the content of the occluded area. For semi-transparent and transparent occluded objects, continuous multi-frame images containing the target product are extracted from the video sequence. An image pyramid is constructed, and the multi-scale Lucas-Kanade method is used to estimate the optical flow field and calculate the motion vector of the occluded object between consecutive frames. The problem is solved layer by layer from coarse to fine to improve the accuracy and robustness of motion estimation. Based on the motion vector, the pixel changes of the occluded area between different frames are analyzed, and the transmittance estimation equation group is established. The transmittance estimation value of each pixel of the occluded object is obtained by solving the least squares method. The Gaussian mixture model is used for background modeling, and multiple Gaussian components are initialized to represent the background distribution of each pixel. The model parameters are iteratively updated by the expectation maximization algorithm to adapt to scene changes. A weighted linear mixture model is established by combining the estimated transmittance and background model. The inverse problem of the mixture model is solved to separate the influence of the occluded object and restore the target product image without the occlusion effect. Adaptive histogram equalization is applied to the restored result to improve the image contrast. Finally, the structural similarity index between the restored image and the original image is calculated to evaluate the restoration quality. If the index is lower than the preset threshold, the reprocessing mechanism is triggered. A 4K high-definition camera with a resolution of 1920x1080 pixels is installed in the workshop to collect real-time video streams with a frame rate of 30fps. Each frame of the image is processed using the U-Net semantic segmentation algorithm. The network input size is 512x512 pixels, and the occluded area is extracted through a 5-layer encoder and decoder structure. Subsequently, the pre-trained ResNet-50 model was used to classify the occluded areas, with an input size of 224x224 pixels, and the last layer was replaced with 3 output nodes, corresponding to the three categories of opaque, semi-transparent, and transparent. The classification threshold was set to 0.7, and the category with a confidence level higher than the threshold was taken as the final judgment result. For opaque occluders, the PatchMatch algorithm set the search radius to 1.5 times the side length of the occluded area, the patch size to 8x8 pixels, and the number of iterations to 5. In each iteration, 100 candidate matching points were randomly sampled, and the best match was selected as the repair source. The occluded pixel values ​​were estimated using bilinear interpolation, and then Poisson image editing was applied, and the number of iterations to solve the Poisson equation was set to 50.For semi-transparent and transparent occluders, a 4-layer image pyramid is constructed with a size ratio of 0.5 for each layer. The Lucas-Kanade method uses a local window of 3x3 pixels, sets the maximum number of iterations to 20, and the convergence threshold to 0.01 pixels. The transmittance is estimated using the least squares method, and the regularization parameter λ is set to 0.1. Five Gaussian components are used for background modeling, the learning rate α is 0.01, and the background threshold T is 0.7. The expectation maximization algorithm is iterated 50 times or stopped when the log-likelihood function changes less than 1e-6. When restoring the image, the weights of the mixed model are solved by the ADMM algorithm, with a maximum number of iterations of 100 and a convergence threshold of 1e-4. The block size of the adaptive histogram equalization is 8x8 pixels, and the contrast limit parameter is set to 3. Finally, the SSIM index is calculated. If it is lower than 0.85, the process is retriggered and retried up to 3 times. The restoration process can be run on a workstation equipped with an NVIDIA RTX3090 GPU, with an average processing time of 50ms per frame, meeting the real-time requirements.

[0083] In some embodiments, in S300, the step of obtaining the estimated distance and azimuth between the target product and the Bluetooth beacon, detecting the position of the blocked target product according to the preprocessed video image, spatially registering the position detection result of the target product with the Bluetooth beacon data, obtaining the estimated distance and azimuth of the Bluetooth beacon relative to the target product according to Bluetooth positioning, detecting and locating the target product by constructing a spatial topological relationship between the product and the beacon, and obtaining the position coordinates and bounding box information of the target product includes:

[0084] S310, acquiring signal strength data transmitted by Bluetooth beacons, selecting a signal strength attenuation model according to the signal strength data; and calculating an estimated distance between a target product and each Bluetooth beacon according to the signal strength attenuation model;

[0085] Exemplarily, the signal strength data transmitted by the Bluetooth beacons arranged in the workshop is obtained, and the signal strength attenuation model is adaptively selected according to environmental factors. The free space path loss model is used in open space, and the logarithmic distance path loss model is used in a multi-obstacle environment to calculate the estimated distance between the target product and each Bluetooth beacon. The formula of the free space path loss model is:

[0086]

[0087] Where PL fs represents the free space path loss, d represents the distance between the transmitting point and the receiving point, f represents the frequency of the signal, and c represents the speed of light.

[0088] The formula for the logarithmic distance path loss model is:

[0089]

[0090] Where PL log represents the logarithmic distance path loss, PL 0 Indicates the reference distance d 0 where n is the path loss exponent and d is the distance between the transmitting point and the receiving point.

[0091] S320, analyzing the preprocessed image to obtain the position and size information of the target product in the image, and determining the initial position of the target product by combining the estimated distance and the position and size information of the target product in the image;

[0092] S330, calculating the distance and azimuth of the target product relative to each beacon, and using an adaptive weighting method based on signal strength to perform weighted averaging on the distance and azimuth of the target product relative to each beacon to obtain a positioning result of the target product; wherein the weight in the adaptive weighting method is positively correlated with the signal strength and is mapped to a range of 0 to 1 through a sigmoid function;

[0093] S340, establishing a spatial topological relationship between the target product and the Bluetooth beacon according to the initial position and positioning result of the target product, and using an extended Kalman filter algorithm to estimate and track the position of the target product in real time based on the spatial topological relationship and in combination with the initial position and positioning result of the target product; wherein the state vector includes position, velocity and acceleration, and the observation vector includes image detection position and Bluetooth positioning result;

[0094] S350, using the extended Kalman filter algorithm to estimate the initial position of the target product in real time to optimize the position coordinates of the target product; and introducing the particle filter algorithm for occlusion, by maintaining multiple hypothetical particles to handle multimodal distribution, to improve tracking stability under occlusion; wherein the state vector of the extended Kalman filter algorithm includes position, velocity and acceleration, and the observation vector includes image detection position and Bluetooth positioning results;

[0095] S360, maps the optimized position coordinates back to the image plane, and combines them with the bounding box information of the target detection to obtain the precise position and size of the target product in the image.

[0096] Exemplarily, the preprocessed image is used as input, and the YOLOv5 target detection algorithm is used to detect the obscured target product to obtain the position and size information of the target product in the image. Combined with the signal strength estimated distance and image detection results, the initial position of the target product is roughly estimated by triangulation. According to the internal and external parameter matrix of the camera, the target product detection result in the image coordinate system is converted to the world coordinate system. Using the known coordinate information of the Bluetooth beacon, combined with the estimated distance calculated by the signal strength, the initial position estimate of the target product is optimized by the least squares method. For the deviation between the optimization result and the detection result, the iterative nearest point algorithm is used for spatial registration. Feature points in the image and three-dimensional space are extracted, and the corresponding relationship is established. The rigid body transformation matrix is ​​calculated by SVD decomposition, and iterative optimization is performed until convergence or the maximum number of iterations is reached. Using the signal strength data of at least three Bluetooth beacons, the precise distance of the target product relative to each beacon is calculated in combination with the triangulation method, and the azimuth of the target product relative to each beacon is calculated by combining the data of the Bluetooth beacon and the triangulation method. The formula for solving the azimuth of the target product relative to each beacon is:

[0097]

[0098] In the formula, θ i represents the azimuth of the target product relative to the i-th beacon, x i and i Represent the x and y coordinates of the ith beacon, x 0 and 0 Represents the x and y coordinates of the target product.

[0099] An adaptive weighting method based on signal strength is used to perform weighted averaging on multiple positioning results; the weight is positively correlated with the signal strength and is mapped to a range of 0 to 1 through a sigmoid function. Based on the fused positioning results and the registered detection results, a spatial topological relationship diagram between the target product and the Bluetooth beacon is established; based on the established spatial topological relationship, combined with the image detection results and Bluetooth positioning results, the extended Kalman filter algorithm is used to estimate and track the position of the target product in real time; the state vector contains position, velocity and acceleration, and the observation vector contains image detection position and Bluetooth positioning results; the position coordinates of the target product are continuously optimized through the prediction and update iterative process; for occlusion situations, a particle filter algorithm is introduced to handle multimodal distribution by maintaining multiple hypothetical particles to improve tracking stability under occlusion; the optimized position coordinates are mapped back to the image plane, and combined with the bounding box information of the target detection, the precise position and size of the target product in the image are obtained.

[0100] Ten Bluetooth 5.1 beacons were deployed in the workshop, with a beacon spacing of 5 meters and a transmission power of 0dBm. An adaptive signal strength attenuation model was used. A free space path loss model was used in open areas with a path loss exponent of n=2, and a logarithmic distance path loss model was used in areas with multiple obstacles with a path loss exponent of n=3.5. The estimated distance between the target product and each beacon was calculated by the received signal strength indication (RSSI) value, and the measurement error was controlled within ±0.5 meters. A high-definition camera with a resolution of 1920x1080 pixels was used to capture images at a frame rate of 30fps. The YOLOv5s model was used to detect the target product, with an input image size of 640x640 pixels, and the bounding box coordinates of each detected target product were output. The confidence threshold was set to 0.5 and the NMS threshold was 0.4. Combining the RSSI estimated distance and the image detection results, the triangulation method was used to preliminarily estimate the target product position with an accuracy of approximately ±1 meter. The camera calibration adopts Zhang Zhengyou calibration method, the chessboard size is 9x6, and the actual size of each square is 25mmx25mm. The internal and external parameter matrix of the camera is calculated by DLT algorithm, and the reprojection error is controlled within 0.1 pixel. The ICP algorithm of SVD decomposition is used for spatial registration, and the maximum number of iterations is set to 50, and the convergence threshold is 0.001 meters. The trilateral measurement method and the triangle centroid method are integrated for precise positioning, and the weight is calculated by the sigmoid function based on RSSI, with function parameters k=0.1 and x0=-70dBm. In the extended Kalman filter, the state vector contains three-dimensional position, velocity and acceleration, a total of 9 variables. The observation vector contains two-dimensional coordinates of image detection and three-dimensional coordinates of Bluetooth positioning, a total of 5 variables. The process noise covariance matrix Q and the measurement noise covariance matrix R are obtained by offline data analysis. For occlusion, the SIR particle filter algorithm with 500 particles is used, and the resampling threshold is set to 50% of the number of effective particles. The final target product positioning accuracy reaches ±0.2 meters, and the bounding box detection accuracy reaches above IoU0.85.

[0101] In some embodiments, in S400, the tracking and predicting of the motion state of the target product based on the position coordinates and bounding box information of the target product in combination with the Bluetooth positioning data to obtain the motion trajectory and speed information of the target product includes:

[0102] S410, obtaining the location coordinates, bounding box information and Bluetooth positioning data of the target product;

[0103] S420, using a Kalman filter algorithm to fuse the position coordinates, the bounding box information, and the Bluetooth positioning data to obtain fused position data; wherein the state vector of the Kalman filter algorithm includes a three-dimensional position and a speed, and the observation vector includes a two-dimensional coordinate and a three-dimensional coordinate;

[0104] S430, fitting the motion trajectory of the target product by least square method according to the fused position data. Constructing a motion state model of the target product based on the motion trajectory and the fused position data; wherein the motion state model includes a linear regression model and a long short-term memory network model;

[0105] S440, predicting the future trajectory of the target product according to the motion state model; wherein the linear regression model is used to predict the short-term future position, and the long short-term memory network model is used to capture the long-term motion pattern. The future trajectory is optimized using a particle filter algorithm to obtain an optimized future trajectory.

[0106] Exemplarily, the position coordinates, bounding box information and Bluetooth positioning data of the target product are obtained; the image detection results and Bluetooth positioning data are fused by the Kalman filter algorithm, the state vector contains the three-dimensional position and speed, and the observation vector contains the two-dimensional coordinates of the image detection and the three-dimensional coordinates of the Bluetooth positioning. The sliding window method is used to process continuous multi-frame data, the window size is set to 10 frames, the step size is 1 frame, and a time series model is established to capture the movement trend of the target product. According to the fused position data, the least squares method is used to fit the movement trajectory of the target product. By comparing the position differences of adjacent time points, the speed components of the target product in the three directions of x, y, and z are calculated. The calculated speed is smoothed by the exponentially weighted moving average algorithm, and the weight factor is set to 0.2 to eliminate the influence of noise and obtain a more stable speed estimate. Based on the historical trajectory and speed information, a motion state model of the target product is constructed. The linear regression algorithm is used to predict the short-term future position, and the long short-term memory network is used to capture the long-term movement pattern to achieve the prediction of the future trajectory of the target product. The LSTM network contains two hidden layers, each with 128 neurons, and the input features are the position and speed data of the past 20 time steps. The comprehensive prediction trajectory is obtained by weighted fusion of the two prediction results. The prediction accuracy is evaluated using the root mean square error and mean absolute percentage error. If the error exceeds the preset threshold, the anomaly detection mechanism is triggered. The prediction trajectory is optimized using the particle filter algorithm, taking into account environmental constraints such as factory layout and equipment location. 1000 particles are set to represent possible future states, and the particle weights are continuously updated through importance sampling and resampling processes. Combined with the predicted trajectory and actual observation data, the prediction model parameters are dynamically adjusted to improve the prediction accuracy. The optimized target product motion trajectory and speed information are output for subsequent production scheduling and quality control. For sudden changes in motion patterns, the sliding variance detection method is used to identify anomalies. When the variance exceeds 3 times the normal range, the model retraining mechanism is triggered. A high-precision camera system and Bluetooth positioning network are deployed in the workshop. The camera resolution is 4K, which is 3840x2160 pixels, the acquisition frame rate is 60fps, and the Bluetooth beacon uses BLE5.1 technology with a signal transmission interval of 50ms. The target product location coordinates are obtained through image processing with centimeter-level accuracy; the bounding box information is provided by the YOLOv5 target detection algorithm, the model input size is 416x416 pixels, and the detection confidence threshold is 0.6. The Kalman filter state vector contains six variables [x, y, z, vx, vy, vz], and the observation vector contains five variables [x_img, y_img, x_ble, y_ble, z_ble]. The sliding window size is set to 10 frames, a total of about 167ms, and the quadratic polynomial curve is fitted as the local trajectory by the least squares method. The speed calculation adopts the central difference method with a time step of 16.7ms. The smoothing factor of the exponentially weighted moving average is α=0.2 to achieve noise suppression of the speed.The input of the LSTM network is a sequence of [x, y, z, vx, vy, vz] of 20 time steps. The two hidden layers each contain 128 LSTM units, and the output layer is a fully connected layer that predicts the position of the next 5 time steps. The linear regression and LSTM prediction results are weighted and fused in a ratio of 6:4. The prediction performance is evaluated using RMSE and MAPE indicators, with thresholds set to 0.1 meters and 5%, respectively. The particle filter uses 1000 particles, each of which contains position and velocity information. The resampling threshold is 50% of the number of valid particles. The environmental constraints consider the 3D model of the factory, and the space occupied by the equipment is used as a no-go zone. Anomaly detection uses a 20-frame sliding window to calculate the variance, and it is considered an anomaly when the variance exceeds 3 times the mean.

[0107] In some embodiments, in S500, during the target product tracking process, if an obstruction blocks the target product, resulting in tracking failure, a local search is performed in an area where the target product may appear according to the historical motion trajectory of the target product and the current Bluetooth positioning estimated position until the motion trajectory and speed information of the target product are recaptured, including:

[0108] S510, receiving continuous frame images and target detection confidence information, and determining whether occlusion has occurred and caused tracking failure according to the continuous frame images and the target detection confidence information, and triggering a target recapture mechanism if the target detection confidence is lower than a preset threshold or the number of consecutive undetected frames exceeds a preset number of frames;

[0109] S520, obtaining the location and time information of the last successful tracking of the target product, and using the Kalman filter algorithm to predict the current location of the target product;

[0110] S530, receiving target product position estimation information, fusing the Kalman filter prediction result and the Bluetooth positioning result through a covariance-based adaptive weight allocation method to obtain a fused estimate of the target product position;

[0111] S540, determining a local search area according to the fusion estimation of the target product position, and recapturing the target using an adaptive grid search algorithm in the local search area; if the target is detected in a certain grid, refining the search of the grid; receiving continuous multi-frame tracking results, and confirming that the target is captured successfully if the number of continuous tracking frames exceeds a preset frame number threshold;

[0112] S550, obtaining historical trajectory information, Bluetooth positioning data and new visual tracking results of the target product, back-estimating the motion state of the target product during the occlusion period, and obtaining the motion trajectory and speed information of the target product.

[0113] Exemplarily, during the target tracking process, continuous frame image analysis and target detection confidence evaluation are used to determine whether tracking failure caused by occlusion occurs. If the confidence output by the target detection algorithm is lower than 0.5 or the target is not detected for 5 consecutive frames, the target recapture mechanism based on Bluetooth positioning is triggered. The location and time information of the last successful tracking of the target product is obtained as the starting point for recapture. Based on the historical motion trajectory data of the target product, the Kalman filter algorithm is used to predict the current possible location of the target product. At the same time, the current location of the target product estimated by the Bluetooth positioning system is obtained. Through the covariance-based adaptive weight allocation method, the Kalman filter prediction results and the Bluetooth positioning results are fused to obtain the most likely location estimate of the target product. Based on the fused position estimate and its uncertainty, a local search area centered on the position is determined. In the determined local search area, an adaptive grid search algorithm is used to recapture the target.

[0114] The initial grid size is set to 1m×1m and is dynamically adjusted according to the uncertainty of the predicted position. Target detection is performed on each grid center point. The improved YOLOv5 algorithm is used, and the spatial attention mechanism is introduced to improve the detection ability of small targets, and the detection threshold is reduced to 0.3 to improve the recall rate. If the target is detected in a certain grid, the area is used as the key search area, and the grid is refined to 0.2m×0.2m for precise positioning. The search results are preliminarily evaluated, and the deviation between the detected target and the expected position is calculated. Once the target product is re-detected in the local search, the multi-frame confirmation mechanism is immediately started, and the target is tracked continuously for at least 5 frames to verify the reliability of the capture. If the capture is confirmed to be successful, the normal tracking process is resumed and the motion trajectory and speed information of the target product are updated. At the same time, the particle filter algorithm is used, 1000 particles are used, and the historical trajectory, Bluetooth positioning data and new visual tracking results are combined to back-estimate the motion state of the target product during the occlusion period. The particle weight is updated based on the consistency with the observed data, and a resampling strategy is adopted, and the threshold is set to 50% of the number of valid particles. If the target is not recaptured for 20 consecutive frames, the search range is expanded to twice the original size, and other available sensor data, such as infrared or ultrasonic auxiliary positioning, are called. A high-definition camera array with a resolution of 4K, 3840x2160 pixels, and a frame rate of 60fps is deployed in the factory environment to cover the entire production area. At the same time, a Bluetooth 5.1 beacon network is installed, arranged at intervals of 5 meters to achieve centimeter-level positioning accuracy. The target tracking adopts the YOLOv5 algorithm, introduces the channel attention mechanism, and the model input size is 640x640 pixels. During normal tracking, the detection confidence threshold is set to 0.5. When the detection confidence is lower than the threshold or the target is not detected for 5 consecutive frames, a total of about 83ms, the recapture mechanism is triggered. The Kalman filter state vector contains 6 variables, including position and velocity, and the observation vector is 3 variables, including position. Bluetooth positioning data is updated every 50ms. Adaptive weight fusion uses a weight calculation method based on Mahalanobis distance, and the weight is inversely proportional to the covariance matrix. The initial search area is set to a range of 5 meters x 5 meters around the last known position, with a grid size of 1 meter x 1 meter. The YOLOv5 model reduces the detection threshold to 0.3 during the recapture phase, and adds additional attention weights to targets that are less than 50% of their original size. The grid refinement strategy is to divide the grid into four equal parts after the target is detected, with the minimum refinement being 0.2m x 0.2m. The multi-frame confirmation mechanism requires 5 consecutive frames and about 83ms of stable detection to be considered a successful capture. The particle filter uses 1000 particles, each of which contains position and velocity information, and the weight update frequency is consistent with the image frame rate of 60Hz. The resampling threshold is 500 valid particles. If recapture is not achieved within 20 frames, about 333ms, the search range will be expanded to 10m x 10m, and an infrared thermal imaging camera with a resolution of 640x480 and a frame rate of 30fps will be started to assist in the search.The recapture process runs on an edge server equipped with NVIDIA RTX3090 GPU, with an average processing delay of less than 20ms, meeting real-time requirements.

[0115] If the target product is partially blocked, the position and motion information of the target product are predicted based on the particle filter algorithm, and the visible part features of the target product are matched to correct the prediction results; if the target product is completely blocked, the position and motion information of the target product are predicted only based on the particle filter algorithm. When the target product is blocked, the motion area of ​​the occluder is predicted based on the prediction results and the position estimated by Bluetooth positioning, and a local search is performed within this area.

[0116] The real-time image data of the target product is obtained, and the target product is divided into regions through the U-Net semantic segmentation network to obtain the area ratio of the visible part of the target product. If the area ratio of the visible part is less than the preset threshold and greater than zero, it is determined to be partially occluded; if the area of ​​the visible part is zero, it is determined to be completely occluded. According to the judgment result, for partial occlusion, the particle filter algorithm is used to predict the position and motion information of the target product. The number of particles is set to a preset value, and the particles contain position, velocity and acceleration information. The particle swarm is initialized by historical motion data, and the state of the particles is updated. For complete occlusion, the weight fusion method is used to predict the position and motion information of the target product in combination with Bluetooth positioning data. The Bluetooth positioning results are introduced into the particle filter update process as additional observations. The particles are updated by the sequential importance resampling method to maintain particle diversity. According to the prediction results of the particle filter algorithm, the position estimated by Bluetooth positioning and the motion prediction of the occluder, a probability distribution map of the motion area of ​​the target product is constructed. The probability distribution map is discretized into a grid, and the grid size is set to a preset value. The grid with a preset percentage before the probability value is selected as the key search area. A spiral search mode with adaptive step length is used, expanding outward from the highest probability point, wherein the probability distribution map is updated during the search process until the target product is recaptured or the preset maximum number of searches is reached.

[0117] Exemplarily, real-time image data of the target product is obtained, and the target product is divided into regions through the U-Net semantic segmentation network, and the area ratio of the visible part of the target product is calculated. If the visible area ratio is less than 30% but greater than zero, it is determined to be partially occluded; if the visible area is zero, it is determined to be completely occluded. According to the judgment result, the corresponding tracking strategy is selected. For partial occlusion, the particle filter algorithm is used to predict the position and motion information of the target product. The number of particles is set to 1000, and each particle contains position, velocity and acceleration information. The particle swarm is initialized by historical motion data, and the state of each particle is updated using the motion model. The features of the visible part of the target product are extracted, including color histogram and HOG features. The cell size in the HOG feature is 8x8 pixels and the block size is 2x2cells. They are matched with the predicted position of the particle and the weight of each particle is calculated. The root mean square error is used to evaluate the prediction accuracy. If the error exceeds the preset threshold, the number of particles is increased or the noise parameter is adjusted. For complete occlusion, prediction is made only by the particle filter algorithm. The number of particles is increased to 2000, and the particle distribution range is expanded to cope with greater uncertainty. Combined with Bluetooth positioning data, the adaptive weight fusion method of covariance intersection criterion is adopted to introduce Bluetooth positioning results as additional observations into the particle filter update process. Particles are updated through sequential importance resampling to maintain particle diversity. At the same time, The optical flow algorithm estimates the motion of the occluder and predicts its possible motion area. Based on the prediction results of the particle filter algorithm, the position estimated by Bluetooth positioning and the motion prediction of the occluder, a probability distribution map of the possible motion area of ​​the target product is constructed. The probability distribution map is discretized into grids, and the size of each grid is set to 0.5 meters × 0.5 meters. The grids with the top 10% of the probability values ​​are selected as the key search areas, and local searches are performed in these areas. A spiral search mode with an adaptive step size is adopted, with an initial step size of 0.5 meters, and the step size increases by 0.1 meters after each turn, expanding outward from the highest probability point. The probability distribution map is continuously updated during the search process until the target product is recaptured or the maximum number of searches is reached. A 4K resolution is deployed in the factory environment, which is a high-speed camera with 3840x2160 pixels and a capture frame rate of 60fps. A U-Net semantic segmentation network is used, with an input size of 512x512 pixels, a 4-layer encoder and decoder structure, and one frame of image is processed every 30ms. Partial occlusion processing is triggered when the visible area of ​​the target product is less than 30%, and full occlusion processing is triggered when it is 0. The particle filter algorithm initializes 1000 particles, and the state vector contains three-dimensional position, velocity and acceleration, a total of 9 variables. When extracting HOG features, the cell size is set to 8x8 pixels, the block size is 2x2cells, and the step size is 4 pixels, resulting in a 3780-dimensional feature vector. The color histogram uses three RGB channels, 16 bins per channel, and a total of 48-dimensional features. The particle weight is calculated by the Euclidean distance of feature matching, normalized and then resampled. The prediction RMSE threshold is set to 0.5 meters, and the number of particles is increased to 1500 if it exceeds. When fully occluded, the number of particles increases to 2000, and the particle distribution range is expanded to a 3-meter area around the last known position. The Bluetooth 5.1 positioning system provides centimeter-level accuracy data, and uses the covariance intersection criterion to calculate the adaptive weight, with a weight range of 0.1-0.9. The parameters of the optical flow algorithm are set to 3 pyramid levels, 15 window size, and 3 iterations. The resolution of the probability distribution map is centimeter-level, and a 100x100 grid covers an area of ​​5mx5m. The initial step length of the spiral search is 0.5 meters, and it increases by 0.1 meters per circle. The maximum search is 100 times or covers an area of ​​25 square meters. It can be run on NVIDIA Jetson AGX Xavier, and the average processing delay is less than 16.7ms (1 / 60 seconds), meeting the real-time requirements.

[0118] In some embodiments, in S600, based on the motion trajectory and speed information of the target product, the depth information of the target product is obtained through a depth camera, and the distance and azimuth estimated by Bluetooth positioning are combined to construct a three-dimensional representation of the target product, and the position and tracking motion information of the indoor target product is generated and transmitted to the indoor workshop production management system to achieve target positioning of indoor products obscured during the production process.

[0119] S610, obtaining the motion trajectory and speed information of the target product, and smoothing the information using a Kalman filter algorithm to obtain filtered motion data;

[0120] S620, collecting a depth image of the target product through a depth camera, and extracting a depth profile of the target product;

[0121] S630, performing illumination unevenness correction on the depth image using a three-scale Retinex algorithm, and improving detail performance of the depth image using a CLAHE algorithm;

[0122] S640, constructing a three-dimensional point cloud representation of the target product based on the processed depth information and Bluetooth positioning data;

[0123] S650, aligning the three-dimensional point cloud representation with the pre-established CAD model of the target product to determine the precise position and posture of the target product. If the mean absolute error of the precise position and posture exceeds a preset threshold, anomaly detection is triggered;

[0124] S660: Pack the precise position and filtered motion data and transmit them to an indoor workshop production management system via industrial Ethernet.

[0125] Exemplarily, the motion trajectory and speed information of the target product are obtained, and the Kalman filter algorithm is used to smooth the data to eliminate the influence of noise. The depth image of the target product is collected by the depth camera, and the depth contour of the target product is extracted by the adaptive threshold segmentation algorithm. The window size is set to 51x51 pixels and the C value is 2. Combined with the distance information estimated by the Bluetooth positioning system, the initial three-dimensional coordinates of the target product are calculated by the triangulation principle. In view of the complex lighting conditions in the workshop, the three-scale Retinex algorithm is used to correct the uneven illumination of the depth image. The scale sizes are 15, 80 and 250 pixels respectively, and the weight ratio is 1:2:1. The CLAHE algorithm is used to improve the detail performance of the depth image. The block size is set to 8x8 pixels and the contrast limit threshold is 3. The morphological opening operation is used to remove noise and small areas in the depth image, and the structure element size is 3x3 pixels. The signal-to-noise ratio of the processed depth image is evaluated. If it is lower than the preset threshold of 15dB, the re-collection mechanism is triggered. Based on the processed depth information and Bluetooth positioning data, a three-dimensional point cloud representation of the target product is constructed. The iterative closest point algorithm is used to align the multi-frame point cloud data, with the maximum number of iterations set to 50 and the convergence threshold set to 0.001 meters. The voxel grid filtering algorithm is used to downsample the point cloud, with the voxel size set to 0.01 meters, to reduce the data volume and retain the geometric features. The generated 3D model is aligned with the pre-established target product CAD model, and the robust model matching method based on RANSAC is used. The internal point threshold is set to 0.005 meters and the number of iterations is 1000. The principal axis direction of the model is determined by the principal component analysis method, and the rigid body transformation matrix is ​​calculated to obtain the precise position and posture of the target product. Combined with the motion trajectory and velocity information processed by the Kalman filter, high-precision position and tracking motion information are generated. The mean absolute error and root mean square error are used to evaluate the positioning accuracy, and the thresholds are set to 0.05 meters and 0.08 meters respectively. If the error exceeds the threshold, the anomaly detection mechanism is triggered, and the sliding variance method is used to identify sudden positioning errors. The data is packaged into a specified format and transmitted to the indoor workshop production management system through industrial Ethernet to achieve the target positioning of indoor products that are blocked during the production process. Intel Real Sense D455 depth cameras are deployed in the workshop with a resolution of 1920x1080 pixels, a frame rate of 30fps, and a depth range of 0.6-6 meters. A Bluetooth 5.1 positioning system is also installed, with a beacon spacing of 3 meters and a positioning accuracy of ±10 cm. The target product motion trajectory data acquisition frequency is 50Hz, and the Kalman filter state vector contains three-dimensional position and velocity, a total of 6 variables. Adaptive threshold segmentation uses a 51x51 pixel window, a C value of 2, and a processing time of less than 5 milliseconds. The scales of the three-scale Retinex algorithm are set to 15, 80, and 250 pixels, with a weight of 1:2:1, and a running time of about 20 milliseconds. The CLAHE algorithm uses 8x8 pixel blocks, a contrast limit threshold of 3, and a processing time of 7 milliseconds. The morphological opening operation structure element is a 3x3 pixel square.The signal-to-noise ratio threshold of the depth image is 15dB. If it is lower than this value, resampling is triggered and the maximum number of retries is 3. The maximum iteration of the ICP algorithm is 50 times, the convergence threshold is 0.001 meters, and the average registration time is 100 milliseconds. The voxel size of the voxel grid filter is 0.01 meters, and the point cloud downsampling rate is about 80%. The threshold of the internal point matching of the RANSAC model is 0.005 meters, the iteration is 1000 times, and the matching time is about 200 milliseconds. The PCA main axis direction calculation uses singular value decomposition to take the eigenvector corresponding to the maximum singular value. The positioning accuracy evaluation uses a 20-frame sliding window, the MAE threshold is 0.05 meters, and the RMSE threshold is 0.08 meters. Anomaly detection uses a 5-second sliding window, and the variance exceeds 3 times the standard deviation. It is judged as an anomaly. Data packaging uses JSON format and is transmitted through 100Mbps industrial Ethernet with a delay of less than 10 milliseconds. The indoor workshop production management system can run on edge computing devices equipped with NVIDIA Jetson XavierNX, with an average processing delay of 33 milliseconds, meeting the real-time requirements.

[0126] and Figure 1 Corresponding to the method, refer to Figure 2 , an embodiment of the present invention provides an indoor target positioning system based on machine vision and deep learning, comprising:

[0127] at least one processor;

[0128] at least one memory for storing at least one program;

[0129] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0130] It can be seen that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0131] In addition, an embodiment of the present invention further discloses a computer program product or a computer program, which is stored in a computer-readable storage medium. A processor of a computer device can read the computer program from a computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above method. Similarly, the contents of the above method embodiment are all applicable to the storage medium embodiment, and the functions specifically implemented by the storage medium embodiment are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those achieved by the above method embodiment.

[0132] It will be appreciated by those skilled in the art that all or some of the methods disclosed above, the system can be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0133] The above is a specific description of the preferred implementation of the present disclosure, but the present disclosure is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions without violating the spirit of the present disclosure. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present disclosure.

Claims

1. A method for indoor target positioning based on machine vision and deep learning, characterized in that: The method comprises the following steps: S100, obtaining a real-time video image in the workshop, analyzing the movement trajectory and size of the obstruction in the image, determining the frequency and duration of the obstruction, and determining the degree of influence of the obstruction on the target product; S200, based on the degree of influence of the occlusion on the target product, preprocess the video image, analyze the transparency and refractive index properties of the pixels in the occluded area, infer the material type of the occluded object, and select the occluded area processing method of the material type to obtain the preprocessed video image; S300, obtaining an estimated distance and azimuth between the target product and the Bluetooth beacon, detecting the position of the blocked target product according to the preprocessed video image, spatially registering the position detection result of the target product with the Bluetooth beacon data, obtaining an estimated distance and azimuth of the Bluetooth beacon relative to the target product according to Bluetooth positioning, detecting and locating the target product by constructing a spatial topological relationship between the product and the beacon, and obtaining the position coordinates and bounding box information of the target product; S400, tracking and predicting the motion state of the target product according to the position coordinates and bounding box information of the target product in combination with the Bluetooth positioning data, and obtaining the motion trajectory and speed information of the target product; S500, during the target product tracking process, if an obstruction blocks the target product, resulting in tracking failure, a local search is performed in an area where the target product may appear based on the historical motion trajectory of the target product and the current Bluetooth positioning estimated position, until the motion trajectory and speed information of the target product are recaptured; S600, based on the target product's motion trajectory and speed information, obtains the target product's depth information through a depth camera, combines the distance and azimuth estimated by Bluetooth positioning, constructs a three-dimensional representation of the target product, generates the position and tracking motion information of the indoor target product, and transmits it to the indoor workshop production management system to achieve target positioning of indoor products that are blocked during the production process.

2. The method according to claim 1, characterized in that In S100, the real-time video image in the workshop is acquired, and the frequency and duration of the occlusion are determined by analyzing the movement trajectory and size of the occlusion in the image, and the degree of influence of the occlusion on the target product is determined, including: S110, acquiring real-time video streams collected by multiple cameras, and performing image preprocessing on the video streams, wherein the image preprocessing includes median filtering, histogram equalization, and Sobel operator edge sharpening; S120, constructing a scene background using a Gaussian mixture model, performing a difference operation between the current frame and the scene background, and obtaining an initial outline of the occluder; S130, matching and associating the initial outline of the occluder between consecutive frames to obtain a motion trajectory of the occluder on the image plane; S140, performing a time series analysis on the movement trajectory of the obstruction, calculating the number of times the obstruction appears in a specific area and the length of time it stays there, and obtaining statistical data on the frequency and duration of the obstruction; S150, according to the spatial position and time distribution of the obstruction, find the affected target products and analyze the impact of the obstruction on each target product; set differentiated impact weights for different production stages to determine the impact of the obstruction on product quality and production efficiency.

3. The method according to claim 1, characterized in that In S200, based on the degree of influence of the occlusion on the target product, the video image is preprocessed, the transparency and refractive index properties of the pixels in the occluded area are analyzed, the material type of the occluded object is inferred, and the occluded area processing method of the material type is selected to obtain the preprocessed video image, including: S210, acquiring a real-time video image captured by a camera, locating the outline of the obstruction, and obtaining a pixel set of the obstruction area; S220, calculating brightness distribution characteristics and gradient information according to the pixel set of the occluded area, analyzing the spatial relationship between the occluded area and the target product in combination with the location information of the target product, and determining a preliminary impact range; S230, extracting a material feature vector according to the texture features and color histogram of the occluded area, and mapping the feature vector to a predefined material category through a support vector machine classifier; if the material category is glass, using an Alpha blending algorithm to process the occluded area; if the material category is translucent plastic, using a transparency-based image enhancement algorithm to process the occluded area; if the material category is opaque metal, using an edge-preserving filtering algorithm to process the occluded area; S240, optimizing the parameters of the occluded area processing method, traversing the parameter space using a grid search method, evaluating the quality of the processing result using a structural similarity index, and applying the optimized method to the processing of the occluded area of ​​the original image to obtain a preprocessed image.

4. The method according to claim 1, characterized in that: In S300, the estimated distance and azimuth between the target product and the Bluetooth beacon are obtained, the position of the blocked target product is detected according to the preprocessed video image, the position detection result of the target product is spatially aligned with the Bluetooth beacon data, the estimated distance and azimuth of the Bluetooth beacon relative to the target product are obtained according to the Bluetooth positioning, and the target product is detected and located by constructing a spatial topological relationship between the product and the beacon to obtain the position coordinates and bounding box information of the target product, including: S310, acquiring signal strength data transmitted by Bluetooth beacons, selecting a signal strength attenuation model according to the signal strength data; and calculating an estimated distance between a target product and each Bluetooth beacon according to the signal strength attenuation model; S320, analyzing the preprocessed image to obtain the position and size information of the target product in the image, and determining the initial position of the target product by combining the estimated distance and the position and size information of the target product in the image; S330, calculating the distance and azimuth of the target product relative to each beacon, and using an adaptive weighting method based on signal strength to perform weighted averaging on the distance and azimuth of the target product relative to each beacon to obtain a positioning result of the target product; wherein the weight in the adaptive weighting method is positively correlated with the signal strength and is mapped to a range of 0 to 1 through a sigmoid function; S340, establishing a spatial topological relationship between the target product and the Bluetooth beacon according to the initial position and positioning result of the target product, and using an extended Kalman filter algorithm to estimate and track the position of the target product in real time based on the spatial topological relationship and in combination with the initial position and positioning result of the target product; wherein the state vector includes position, velocity and acceleration, and the observation vector includes image detection position and Bluetooth positioning result; S350, using the extended Kalman filter algorithm to estimate the initial position of the target product in real time to optimize the position coordinates of the target product; and introducing the particle filter algorithm for occlusion, by maintaining multiple hypothetical particles to handle multimodal distribution, to improve tracking stability under occlusion; wherein the state vector of the extended Kalman filter algorithm includes position, velocity and acceleration, and the observation vector includes image detection position and Bluetooth positioning results; S360, maps the optimized position coordinates back to the image plane, and combines them with the bounding box information of the target detection to obtain the precise position and size of the target product in the image.

5. The method according to claim 1, characterized in that In S400, the motion state of the target product is tracked and predicted based on the position coordinates and bounding box information of the target product in combination with the Bluetooth positioning data to obtain the motion trajectory and speed information of the target product, including: S410, obtaining the location coordinates, bounding box information and Bluetooth positioning data of the target product; S420, using a Kalman filter algorithm to fuse the position coordinates, the bounding box information, and the Bluetooth positioning data to obtain fused position data; wherein the state vector of the Kalman filter algorithm includes a three-dimensional position and a speed, and the observation vector includes a two-dimensional coordinate and a three-dimensional coordinate; S430, fitting the motion trajectory of the target product by least square method according to the fused position data; constructing a motion state model of the target product based on the motion trajectory and the fused position data; wherein the motion state model includes a linear regression model and a long short-term memory network model; S440, predicting the future trajectory of the target product according to the motion state model; wherein the linear regression model is used to predict the short-term future position, and the long short-term memory network model is used to capture the long-term motion pattern; and optimizing the future trajectory using a particle filter algorithm to obtain an optimized future trajectory.

6. The method according to claim 1, characterized in that In S500, during the target product tracking process, if an obstruction blocks the target product, resulting in tracking failure, a local search is performed in the area where the target product may appear according to the historical motion trajectory of the target product and the current Bluetooth positioning estimated position until the motion trajectory and speed information of the target product are recaptured, including: S510, receiving continuous frame images and target detection confidence information, and determining whether occlusion has occurred and caused tracking failure according to the continuous frame images and the target detection confidence information, and triggering a target recapture mechanism if the target detection confidence is lower than a preset threshold or the number of consecutive undetected frames exceeds a preset number of frames; S520, obtaining the location and time information of the last successful tracking of the target product, and using the Kalman filter algorithm to predict the current location of the target product; S530, receiving target product position estimation information, fusing the Kalman filter prediction result and the Bluetooth positioning result through a covariance-based adaptive weight allocation method to obtain a fused estimate of the target product position; S540, determining a local search area according to the fusion estimation of the target product position, and recapturing the target using an adaptive grid search algorithm in the local search area; if the target is detected in a certain grid, refining the search of the grid; receiving continuous multi-frame tracking results, and confirming that the target is captured successfully if the number of continuous tracking frames exceeds a preset frame number threshold; S550, obtaining historical trajectory information, Bluetooth positioning data and new visual tracking results of the target product, back-estimating the motion state of the target product during the occlusion period, and obtaining the motion trajectory and speed information of the target product.

7. The method according to claim 1, characterized in that In S600, based on the target product's motion trajectory and speed information, the depth information of the target product is obtained through a depth camera, and the distance and azimuth estimated by Bluetooth positioning are combined to construct a three-dimensional representation of the target product, generate the position and tracking motion information of the indoor target product, and transmit it to the indoor workshop production management system to achieve the target positioning of the indoor product blocked during the production process, including: S610, obtaining the motion trajectory and speed information of the target product, and smoothing the information using a Kalman filter algorithm to obtain filtered motion data; S620, collecting a depth image of the target product through a depth camera, and extracting a depth profile of the target product; S630, performing illumination unevenness correction on the depth image using a three-scale Retinex algorithm, and improving detail performance of the depth image using a CLAHE algorithm; S640, constructing a three-dimensional point cloud representation of the target product based on the processed depth information and Bluetooth positioning data; S650, registering the three-dimensional point cloud representation with a pre-established CAD model of the target product to determine the precise position and posture of the target product; if the mean absolute error of the precise position and posture exceeds a preset threshold, triggering anomaly detection; S660: Pack the precise position and filtered motion data and transmit them to an indoor workshop production management system via industrial Ethernet.

8. An indoor target positioning system based on machine vision and deep learning, characterized in that: The system comprises: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle tracking method combining target information and motion estimation

    CN103927764A

  • Positioning method, terminal and server

    CN110166501A

  • Method for dynamically tracking and positioning indoor personnel in large-scale place

    CN111080679A

  • Tracking method of video moving target

    CN116993774A

  • Human body posture classification method based on computer vision

    CN117593788A

Cited By

  • Indoor positioning method based on ADMM optimization

    CN120593753A

  • Cross-environment multi-mode fusion positioning method for emergency rescue

    CN120628101A

  • Target tracking method and system based on AI vision

    CN120976875A

  • Three-dimensional vision high-precision real-time positioning method and system for single agent

    CN121280688A

  • Digital twin motion track rendering method and system based on space-time linkage

    CN121346817A