Machine vision fusion-based vehicle and personnel cooperative positioning system for closed stock ground
By employing multimodal perception, environmental feature analysis, visual enhancement, and dynamic weight fusion technologies, the adaptability and accuracy issues of the closed material yard positioning system in high dust and strong light environments have been resolved, enabling collaborative positioning and safety management of vehicles and personnel.
Patent Information
- Application Number
- CN202511444237.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-01-02
AI Technical Summary
Existing closed material yard positioning technology has poor adaptability in high dust and strong light environments. Fixed weights of multiple sensors lead to reduced fusion positioning accuracy, and the positioning output is not well adapted to the scene, making it impossible to achieve human-vehicle collaborative management.
The system employs a multimodal sensing module to collect data in real time, filters dust concentration and light intensity through an environmental feature analysis module, enhances images using a visual enhancement module, adjusts sensor weights through a dynamic weight fusion module, and utilizes a digital twin mapping unit for visualization and early warning.
The system improved the adaptability and accuracy of the positioning system in high dust and strong light environments, enabling accurate positioning of vehicles and personnel and enhancing the safety management level of enclosed material yards.
Smart Images

Figure CN121254293A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of closed stockyard vehicle personnel cooperative positioning, in particular to a closed stockyard vehicle personnel cooperative positioning system based on machine vision fusion. BACKGROUND
[0002] In modern industrial production, closed stockyards, as important places for storing and processing raw materials, are widely used in many industries such as mines, ports, power, steel and so on. The closed stockyard has the characteristics of relatively closed space and complex and changeable environment, and there are usually a large number of material piles, operating vehicles (such as loaders, transport trucks, etc.) and workers inside. With the development of industrial automation and intelligentization, there are higher and higher requirements for the accurate positioning and cooperative management of vehicles and personnel in the closed stockyard. On the one hand, accurate vehicle positioning helps to realize automatic scheduling and operation process optimization, improve material handling efficiency and reduce operating costs. For example, in a coal storage stockyard, accurate positioning of transport trucks can reasonably arrange loading and unloading tasks, reduce vehicle waiting time and improve overall operation efficiency. On the other hand, accurate positioning of personnel is the key to ensuring safety in production. There are many safety hazards in the closed stockyard, such as vehicle collision and material collapse, and real-time understanding of personnel location can timely issue warnings to avoid safety accidents and quickly organize rescue in emergency situations.
[0003] The existing positioning technology of closed stockyard mainly has the following defects: Poor environmental adaptability: The closed stockyard has high concentration of dust for a long time, which causes the images collected by traditional visual sensors to have serious fogging and loss of details; at the same time, the light intensity in the stockyard varies greatly between day and night, and ordinary cameras are prone to overexposure or underexposure, which cannot provide effective image data for visual positioning. Although some technologies use de-fogging algorithms or HDR technology, they are not optimized for the characteristics of high dust and strong light intensity fluctuations in the stockyard.
[0004] Fixed weight of multiple sensors: Existing multi-sensor fusion positioning mostly uses fixed weight fusion strategy, without considering the dynamic influence of the closed stockyard environment on the performance of the sensors. For example, the positioning error of the UWB module in the steel structure shielding area of the stockyard increases from 10 cm to more than 50 cm, but the fixed weight is still set the same as that in the non-shielding area, resulting in a sharp decline in the accuracy of the fusion positioning; when the dust concentration is >60 mg / m³, the point cloud penetration of the laser radar decreases, and the effective detection distance shortens by 30%, but the existing technology does not adjust the weight proportion of the laser radar according to the dust concentration, causing resource waste or error accumulation.
[0005] Insufficient adaptation of positioning output to the scene: The existing positioning system cannot match the dynamic changes of vehicles and personnel in the closed stockyard, and the positioning results are not linked to the digital twin model of the closed stockyard, which cannot provide a visual spatial monitoring interface for the management personnel, and is not conducive to safety decision-making.
[0006] In summary, there is an urgent need for a system that can adapt to the characteristics of a closed feedlot environment, dynamically optimize perception strategies, and achieve human-vehicle collaborative positioning, to solve the problems of poor environmental adaptability, insufficient robustness, and low real-time performance of existing technologies. SUMMARY
[0007] The purpose of the present application is to provide a closed feedlot vehicle personnel collaborative positioning system based on machine vision fusion to solve the problems raised in the above background.
[0008] The purpose of the present application can be achieved by the following technical solution: a closed feedlot vehicle personnel collaborative positioning system based on machine vision fusion, comprising: A multi-modal perception module is used to collect multi-source visual positioning raw data, UWB positioning data, and laser radar positioning data of the closed feedlot in real time according to multi-modal perception sensors; An environmental feature analysis module is used to obtain environmental data of the closed feedlot, including dust concentration and illumination intensity, and filter the environmental data through an average filtering algorithm to obtain the dust concentration and illumination intensity of the closed feedlot after filtering; A vision enhancement module is used to enhance the images in the multi-source visual positioning data based on the dust concentration and illumination intensity of the closed feedlot after filtering, to obtain enhanced images; the vision enhancement module includes a deep learning defogging unit, an HDR image processing unit, and a motion artifact elimination unit; wherein the deep learning defogging unit is used to perform defogging processing analysis based on the images in the multi-source visual positioning raw data and the dust concentration of the closed feedlot to obtain defogging images; The HDR image processing unit is used to perform HDR processing on the defogging images that meet the image quality requirements based on the illumination intensity of the closed feedlot to obtain HDR processed images; The motion artifact elimination unit is used to determine whether the HDR images have artifact regions, and if so, perform motion artifact elimination processing to obtain enhanced images; A multi-modal feature extraction module is used to extract features from the positioning data of the multi-modal perception sensors to generate multi-modal perception sensor positioning feature information, and output respective reliability score values of the multi-modal perception sensors based on a reinforcement deep learning model; A dynamic weight fusion module is used to match the respective reliability score values of the multi-modal perception sensors with a preset basic weight rule base corresponding to each reliability score value to obtain the basic weights of each perception sensor, and dynamically adjust the basic weights through a Kalman filtering algorithm to obtain real-time dynamic weights of each perception sensor; The collaborative positioning and settlement module is used to calculate the final positioning results of vehicles and personnel in the enclosed material yard based on the real-time dynamic weights and positioning data of the multimodal sensing sensors. The digital twin mapping unit is used to map the final location results of vehicles and personnel in the enclosed material yard to a pre-built digital twin model of the enclosed material yard for visualization, while analyzing the corresponding distance between vehicles and personnel and providing corresponding early warnings.
[0009] The beneficial effects of this invention are: This invention employs an average filtering algorithm to filter dust concentration and light intensity in enclosed material yards, removing high-frequency noise caused by vehicle exhaust and headlight flicker, ensuring stable and reliable environmental parameters. This provides an accurate environmental parameter basis for subsequent visual enhancement processing, improving image processing quality and effectiveness, and ultimately enhancing positioning accuracy. Visual enhancement involves steps such as dehazing, HDR processing, and motion artifact removal to enhance the original images acquired by the visual sensor. Different global atmospheric light value estimation methods are used for different dust concentration ranges, combined with a pre-trained AOD-Net model for dehazing, effectively eliminating the influence of dust on the image. An appropriate HDR processing mode is selected based on light intensity to address image problems caused by insufficient or excessive lighting. By identifying and eliminating motion artifacts, image quality is further improved, enabling the system to adapt to image quality requirements under different dust concentrations and lighting conditions, providing reliable assurance for subsequent positioning feature extraction and enhancing the system's adaptability and stability in different environments.
[0010] This invention extracts multi-dimensional features from the collected multimodal positioning data and uses a reinforced deep learning model (DQN) to score the reliability of multimodal sensing sensors. Based on the reliability scores, the basic weights are dynamically adjusted using a Kalman filter algorithm to obtain the real-time dynamic weights of the multimodal sensing sensors. The Kalman filter algorithm can continuously adjust the weights of the multimodal sensing sensors based on new observation information, adapting to the dynamic changes of vehicles and personnel in a closed material yard environment and the potential performance fluctuations of the modal sensing sensors, thus ensuring the accuracy and robustness of subsequent positioning.
[0011] The application obtains the final positioning result of the vehicle and personnel in the closed material field by performing coordinate unification and format conversion on the positioning data of the multi-modal perception sensor, fusing multi-source positioning results by using the weighted least square method, fully utilizing the advantages of multi-source data, improving the positioning accuracy and reliability through weighted processing, and more accurately reflecting the actual position of the vehicle and personnel in the closed material field. The final positioning result of the vehicle and personnel in the closed material field is mapped to the pre-constructed digital twin model of the closed material field for visual display, the distance between the vehicle and personnel is analyzed, and corresponding warning display is performed, the distance is calculated by the Euclidean distance algorithm and compared with the preset value, a collision risk signal is generated and different color warning marks are triggered, and the closed material field management personnel terminal is synchronously sent. Through the visualization and warning mechanism, the management personnel can intuitively understand the situation in the closed material field, discover potential safety hazards in time, take corresponding measures, and effectively improve the safety management level of the closed material field. BRIEF DESCRIPTION OF DRAWINGS
[0012] The application will be further described below in conjunction with the drawings.
[0013] Figure 1 is a system block diagram of the application.
[0014] Figure 2 is an AOD-Net model structure diagram of the application.
[0015] Figure 3 is a logic diagram of the visual enhancement module of the application.
[0016] Figure 4 The logic diagram of the Kalman filter algorithm of the application for dynamically adjusting the basic weight. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the application will be described clearly and completely below in conjunction with the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the application.
[0018] Please refer to Figure 1 The application is a closed material field vehicle and personnel cooperative positioning system based on machine vision fusion, which comprises: The multi-modal perception module is used to collect multi-source visual positioning data, UWB positioning data and laser radar positioning data of the enclosed stockyard in real time according to the multi-modal perception sensor; specifically, industrial-grade dustproof and waterproof cameras are deployed at key points of the enclosed stockyard, such as entrances and exits, work channels, stockpile perimeters and turning blind spots, etc., vehicle-mounted fisheye cameras or infrared sensors are deployed at the front, rear, left and right of the vehicle, and intelligent safety helmet cameras are provided for personnel to collect and obtain multi-source visual positioning raw data of the enclosed stockyard, including stockyard scene images, vehicle images and personnel images, etc.; UWB positioning data is collected by UWB base stations deployed on stockyard pillars and stockpile edges; laser radar positioning data is obtained by solid-state laser radars deployed in dense stockpile areas and turning blind spots, etc.
[0019] The environmental feature analysis module is used to obtain environmental data of the enclosed stockyard, the environmental data including dust concentration and illumination intensity, and the environmental data is filtered by an average filtering algorithm to obtain the dust concentration and illumination intensity of the filtered enclosed stockyard; specifically, the specific steps of filtering are as follows: The enclosed stockyard is divided according to a preset spatial volume division method to obtain each sub-space region of the enclosed stockyard, the dust concentration of each sub-space region of the enclosed stockyard is detected by a laser scattering dust sensor to obtain dust concentration data of each sub-space region of the enclosed stockyard, and the illumination intensity of each sub-space region of the enclosed stockyard is detected by a digital illumination sensor to obtain illumination intensity data of each sub-space region of the enclosed stockyard. The dust concentration data and illumination concentration data of each sub-space region of the enclosed stockyard are filtered according to a preset sliding window size, and the average filtering algorithm formula is used for filtering The filtered dust concentration and illumination intensity of each sub-space region of the enclosed stockyard are calculated, where N represents the sliding window size, i.e. the sampling number, and represent the dust concentration sampling value and the illumination intensity sampling value of the i-th sub-space region of the enclosed stockyard in the j-th sliding window, respectively. The filtered dust concentration and illumination intensity of each sub-space region of the enclosed stockyard are subjected to mean value processing to obtain the mean value of the dust concentration and the mean value of the illumination intensity of the filtered enclosed stockyard, which are used as the dust concentration and the illumination intensity of the filtered enclosed stockyard.
[0020] It should be noted that the dust concentration data and the illumination intensity data of the enclosed stockyard collected are filtered by the average filtering algorithm, in order to remove high-frequency noise caused by vehicle exhaust, light flickering, etc. in the enclosed stockyard, and to ensure the stability of the environmental parameters, thereby providing reliable environmental parameter basis for subsequent visual enhancement.
[0021] a visual enhancement module, configured to perform enhancement processing on images in the multi-source visual positioning data according to the dust concentration and the illumination intensity of the filtered closed field, eliminate environmental interference, and obtain multi-source visual positioning data after image enhancement processing; specifically, the visual enhancement module comprises a deep learning defogging unit, an HDR image processing unit, and a motion artifact elimination unit; the deep learning defogging unit is configured to perform defogging processing analysis on the images in the multi-source visual positioning data and the dust concentration of the closed field, and obtain defogged images; specifically: 301-1: Obtain the dust concentration of the filtered closed field, and compare it with a preset low dust concentration interval, a medium dust concentration interval, and a high dust concentration interval; if the dust concentration of the filtered closed field is in the low dust concentration interval, obtain the average value of the first 0.1% pixels in the brightness of the dark channel of the image as the global atmospheric light value A; it should be noted that this step uses global dark channel priority estimation A, and the dark channel is calculated pixel by pixel for the input image, that is, the minimum value of RGB of each pixel, to obtain a dark channel matrix; the dark channel matrix is flattened into a one-dimensional array, sorted from high to low in brightness, and the first 0.1% bright pixels are taken to extract the RGB values of these pixels in the input image, and the average value is taken as the global atmospheric light value A; If the dust concentration of the filtered closed field is in the medium dust concentration interval, the image is divided into a preset grid, and the average value of the first 0.3% pixels in the dark channel of each grid is calculated as the atmospheric light value of the region, and the average value of the atmospheric light values of all regions is calculated as the global atmospheric light value A; it should be noted that this step uses a region adaptive method to estimate A, which avoids the deviation of A value caused by local dust accumulation; in a specific example, the preset grid can be a grid; If the dust concentration of the filtered closed field is in the high dust concentration interval, first calculate the basic global atmospheric light value A0 according to the region adaptive method in the case of the medium dust concentration interval, introduce a dust concentration compensation coefficient , the formula is , where represents the lower limit value of the high dust concentration interval, represents the dust concentration of the filtered closed field, and the basic global atmospheric light value A0 is multiplied by the dust concentration compensation coefficient to obtain the adjusted global atmospheric light value A; it should be noted that the dust concentration compensation coefficient formula is obtained by experiment fitting; under high dust concentration, scattering makes the brightness of the fog "underestimated", and by introducing the dust concentration compensation coefficient on the basis of the basic global atmospheric light value, the defogging intensity is enhanced, and fog residues are avoided; It should be noted that the global atmospheric light value A is a parameter for describing the global light of the environment, the higher the dust concentration in the closed material field, the stronger the atmospheric light scattering, and the global atmospheric light value A needs to be increased accordingly to enhance the fog removal intensity; 301-2: Construct a pre-trained AOD-Net model, the pre-trained AOD-Net model includes an input layer, an encoder, a decoder, a skip connection and an output layer; It should be noted that the image to be processed is taken as the input of the AOD-Net model, which is received by the input layer; the encoder gradually extracts the feature information of the image through the convolution layer and the pooling layer operation, and reduces the spatial dimension of the feature map; the decoder gradually recovers the definition of the image by using the up-sampling and transposed convolution layer; the skip connection splices the feature map extracted in the encoder with the corresponding level feature map in the decoder to retain the detail information of the image; and the output layer outputs the processed defogging image; As shown in Figure 2 , the original image is normalized and taken as the input of the AOD-Net model, the input image is processed through the Conv1 layer with a convolution kernel of 1*1 and the Conv2 layer with a convolution kernel of 3*3, the feature matrices obtained by the two layers are fused into the Concat1 layer, then the feature matrices obtained by the Conv2 layer and the Conv3 layer with a convolution kernel of 5*5 are fused into the Concat2 layer, then the feature matrices obtained by the previous four times of convolution are all fused into the Concat3 layer through the Conv4 layer with a convolution kernel of 7*7, and the unified parameter is calculated through the Conv5 layer with a convolution kernel of 3*3 by the formula , is and a unified form, is a medium transmission matrix, and finally the defogging image is obtained by the formula , wherein represents the original foggy image, is a constant with a default value of 1; 301-3: Obtain the peak signal-to-noise ratio and the structural similarity of the defogging image, and compare them with the peak signal-to-noise ratio threshold and the structural similarity threshold of the preset visual positioning image defogging quality requirement, if , the defogging image is sent to the HDR image processing unit as the input of the subsequent HDR synthesis; if , the defogging image is executed again in steps 301-1 to 301-3 until the image quality requirement is met, that is , wherein are represented as logical symbols, respectively, and and or.
[0022] The HDR image processing unit is configured to perform HDR processing on the defogged image satisfying the image quality requirement based on the light intensity of the enclosed field to obtain an HDR processed image; specifically: The light intensity of the filtered enclosed field is obtained and compared with a preset light intensity threshold value. If the light intensity of the filtered enclosed field is less than or equal to the preset light intensity threshold value, the defogged image is processed by the HDR image processing software in a single-frame HDR mode to obtain an HDR processed image. If the light intensity of the filtered enclosed field is greater than the preset light intensity threshold value, the defogged image is processed by the HDR image processing software in a multi-frame synthesis HDR mode to obtain an HDR processed image. The HDR image processing software can be HDRPhotoPro. It should be noted that if the light intensity of the filtered enclosed field is greater than the preset light intensity threshold value, the single-frame image is overexposed and needs to be complemented by multiple frames. The multi-source visual sensor pre-acquires 3 frames of images with different exposure times, obtains 3 frames of defogged images with different exposure times through the deep learning defogging unit, aligns and merges the 3 frames of defogged images with different exposure times through the HDR image processing software, and generates multi-frame HDR images.
[0023] The motion artifact elimination unit is configured to determine whether the HDR image has an artifact area. If so, motion artifact elimination processing is performed to obtain an enhanced processed image; specifically: 303-1: Calculate the absolute gray difference between the current HDR image and the previous frame of HDR image pixel by pixel The formula is: Compare the absolute gray difference with a preset gray difference threshold value. If the absolute gray difference is greater than the preset gray difference threshold value, it is determined that the pixel belongs to the artifact area. 303-2: Extract ORB feature points from the current HDR image and the previous frame HDR image respectively by the ORB algorithm, which is a fast and effective feature point detection and description algorithm that can extract key points and corresponding feature descriptors in the image; estimate the pixel motion vector using the ucas-Kanade optical flow method; the optical flow method estimates the motion of objects by analyzing the intensity changes of pixels in the image sequence, and the Lucas-Kanade method is a local optical flow estimation method, which assumes that within a set optical flow method window size, the motion of all pixels is similar, for each ORB feature point in the image, by establishing an equation system and using the least squares method to solve, the motion vector of each ORB feature point is obtained, forming an optical flow field; according to the calculated optical flow field, the current frame HDR image is transformed to align it with the previous frame HDR image in space, generating a preliminary corrected image; use Gaussian filtering to smooth the edges of the preliminary corrected image to obtain an image with motion artifacts removed; 303-3: Obtain the area of the processed artifact region and the area of the artifact region before processing, according to the formula Calculate the artifact removal rate, and compare the artifact removal rate with the preset artifact removal rate threshold, if the artifact removal rate is greater than or equal to the preset artifact removal rate threshold, output the enhanced image; otherwise, increase the optical flow method window size (such as from to ), re-execute steps 303-2 to 303-3 until the artifact removal quality requirement is met.
[0024] It should be noted that, as Figure 3 indicated, the visual enhancement module enhances the original image collected by the visual sensor through steps such as defogging processing, HDR processing and motion artifact removal, to adapt to the image quality requirements under different dust concentrations and lighting conditions, and to provide reliable guarantee for subsequent positioning feature extraction.
[0025] The multi-modal feature extraction module is configured to extract features from the positioning data of the multi-modal perception sensor to generate multi-modal perception sensor positioning feature information, and output respective reliability scores of the multi-modal perception sensor based on the reinforced deep learning model; specifically, the multi-modal perception sensor positioning feature information is generated, including: Obtain multi-source visual positioning data, UWB positioning data and laser radar positioning data after visual enhancement processing; Generate multi-modal perception sensor positioning feature information based on the multi-modal perception sensor positioning information feature extraction model and the preset multi-source visual positioning data, UWB positioning data and laser radar positioning data after visual enhancement processing.
[0026] It should be noted that the preset multi-modal perception sensor positioning information feature extraction model includes a preset visual sensor positioning information feature extraction model, a UWB positioning information feature extraction model and a laser radar positioning information feature extraction model, and respectively corresponds to a plurality of preset positioning information feature extraction sub-models, a plurality of preset positioning information feature extraction sub-model selection functions, wherein the preset multi-modal perception sensor positioning information feature extraction sub-model corresponds to the preset positioning information feature extraction sub-model selection function one by one; the preset multi-modal perception sensor positioning information feature extraction model is a model based on the transforme architecture, and is a trained model.
[0027] It should be further noted that the multi-modal perception sensor positioning feature information is generated according to the multi-source visual positioning data after visual enhancement processing, the UWB positioning data and the laser radar positioning data, and the preset multi-modal perception sensor positioning information feature extraction model, including the following steps: According to the multi-source visual positioning data after visual enhancement processing, the UWB positioning data and the laser radar positioning data, and the corresponding plurality of preset positioning information feature extraction sub-model selection functions, the corresponding plurality of preset positioning information feature extraction sub-model selection weight information is obtained; It should be noted that the preset positioning information feature extraction sub-model selection function is artificially set, which is used to quantify the adaptation degree of the multi-source visual positioning data after visual enhancement processing, the UWB positioning data and the laser radar positioning data to each positioning information feature extraction sub-model; the multi-source visual positioning data after visual enhancement processing, the UWB positioning data and the laser radar positioning data are converted into vectors, and the respective vectors are spliced, the spliced vector is taken as the independent variable of the corresponding plurality of positioning information feature extraction sub-model selection functions, and the function value obtained by calculating the corresponding plurality of positioning information feature extraction sub-model selection functions is taken as the corresponding plurality of positioning information feature extraction sub-model selection weight information.
[0028] The corresponding plurality of preset positioning information feature extraction sub-model selection weight information is standardized to obtain the positioning feature information plurality of preset positioning information feature extraction sub-model standardized selection weight information; It should be noted that the plurality of positioning information feature extraction sub-model selection weight information is standardized by Z-score standardization, which is used to scale the value range of the plurality of preset positioning information feature extraction sub-model selection weight information to interval, so as to avoid the dimensional error caused by the subsequent calculation process.
[0029] According to the multi-source visual positioning data after visual enhancement processing, the UWB positioning data and the laser radar positioning data, and the corresponding multiple preset multi-positioning information feature extraction sub-models, the corresponding multiple positioning information feature variable information is calculated; It should be noted that the multi-source visual positioning data after visual enhancement processing, the UWB positioning data and the laser radar positioning data are converted into vectors, and each vector is spliced, and the spliced vector is used as the input vector of the multiple positioning information feature extraction sub-models. Through the calculation of the multiple positioning information feature extraction sub-models, the output results of the multiple positioning information feature extraction sub-models are used as the multiple positioning information feature variable information.
[0030] According to the corresponding multiple preset positioning information feature extraction sub-model standardization selection weight information, the corresponding multiple positioning information feature variable information is weighted and summed to generate the corresponding positioning feature information.
[0031] It should be noted that the multi-modal perception sensor positioning data corresponding to each positioning information feature extraction sub-model standardization selection weight information corresponds to the output result of each positioning information feature extraction sub-model, and then weighted sum calculation is performed, and the calculation result is used as the positioning feature information, so that the positioning feature information of the multi-modal perception sensor is obtained.
[0032] Specifically, the reliability score values of the multi-modal perception sensors are output based on the reinforcement deep learning model, including: The reinforcement deep learning model DQN network structure is constructed, including an input layer, a hidden layer and an output layer, wherein the input layer receives the multi-dimensional features of the multi-modal perception sensors, the hidden layer performs nonlinear transformation and abstraction on the multi-dimensional features, and the output layer outputs the respective reliability score values of the multi-modal perception sensors; The feature information of the multi-modal perception sensor is defined as a state vector , and is used as the input of the reinforcement deep learning model, and the action of the reinforcement deep learning model is defined ; the reliability score vector of the multi-modal perception sensor combination at the current time step is set as , the reliability score vector of the multi-modal perception sensor combination at the last time step is , and the reliability score value transformation rate is calculated by the formula , which is used to measure the change degree of the reliability score value at different time steps, wherein represents the Euclidean norm of the vector, which is used to calculate the length or size of the vector; The reliability score value transformation rate a preset reliability score value change rate threshold If the comparison result is , an additional reward r is given, and if the comparison result is , no reward is given. The reinforcement deep learning model DQN calculates the Q value of each action by forward propagation according to the state vector , where the Q value represents the expected cumulative reward of performing the action on the state vector, and the reinforcement deep learning model selects the action with the maximum Q value as the output, thereby obtaining the respective reliability score values of the multi-modal perception sensors; The dynamic weight fusion module is configured to match the respective reliability score values of the multi-modal perception sensors with a preset basic weight rule library corresponding to each reliability score value, to obtain the basic weights of each perception sensor, and to dynamically adjust the basic weights by a Kalman filtering algorithm to obtain real-time dynamic weights of each perception sensor. Specifically, the dynamic adjustment of the basic weights by the Kalman filtering algorithm includes the following steps: 501: State definition and observation equation establishment: define the weights of the multi-modal perception sensors as a state vector , where represents the weight of the i-th perception sensor, and n is the number of perception sensors; set the observation equation of the i-th perception sensor as , where is the observation vector of the i-th perception sensor, is the observation matrix, maps the state vector to the observation space, is the observation noise, and has a covariance matrix Integrate the observation equations of all perception sensors to obtain the overall observation equation , where , is a block matrix composed of , and its covariance matrix is a block diagonal matrix composed of ; 502: Initialization: preset an initial state vector estimate and an initial state estimate covariance matrix ; It should be noted that the initial state vector estimate is set according to the basic weights of each perception sensor, and the initial state estimate covariance matrix represents the uncertainty of the initial state estimate, and the value of the covariance matrix can be set according to the confidence level of the initial estimate. If the initial estimate is relatively certain, the value of the covariance matrix can be set smaller; otherwise, the value of the covariance matrix can be set larger.
[0033] 503: Prediction: predict the state vector according to the dynamic model of the system, set the dynamic model of the system as wherein is the state vector at the current time, is the state vector at the last time, that is, the current time is , the last time is , is the state transition matrix, is the process noise, and the covariance matrix is , through the model, the state vector at the time is predicted , and the covariance matrix of the state estimation is predicted simultaneously.
[0034] 504: Update: calculate the Kalman gain by the formula , the Kalman gain determines the weight of the observation information in updating the state estimation, which considers the uncertainty of the predicted state and the influence of the observation noise; the state estimation is updated using the new observation value , and the update formula is , the state update combines the observation information and the prediction information to obtain a more accurate state estimation, that is, the updated multi-modal perception sensor weight; the covariance matrix of the state estimation is updated , and the updated covariance matrix of the state estimation reflects the uncertainty of the updated state estimation, providing a basis for the next prediction and update; 505: Iterative operation: take the state vector estimation at the current time and the covariance matrix of the state estimation as the input of the step 503 at the next time, and repeatedly execute the steps 503-504, thereby obtaining the real-time dynamic weight of each perception sensor.
[0035] It should be noted that, as shown in Figure 4 , the Kalman filtering algorithm can continuously adjust the weight of the multi-modal perception sensor according to the new observation information, so as to adapt to the dynamic changes of the vehicle and the personnel in the closed field environment and the possible performance fluctuations of the perception sensor. For example, if the observation quality of a certain perception sensor is reduced due to shielding or other interference, the Kalman filtering algorithm will gradually reduce the weight of the sensor in the subsequent iteration, and increase the weight of other stable performance sensors, thereby ensuring the accuracy and robustness of subsequent positioning.
[0036] The cooperative positioning settlement module is configured to calculate final positioning results of the vehicles and the personnel in the closed material yard according to the acquired real-time dynamic weights of the multi-modal perception sensors and the positioning data of the multi-modal perception sensors; specifically: The multi-source visual positioning data after the visual enhancement processing, the UWB positioning data and the laser radar positioning data are subjected to coordinate unification and format conversion to obtain three-dimensional coordinates of the vehicles and the personnel in a unified coordinate system of the perception sensors, denoted as , wherein f represents the number of the vehicles and the personnel, f = 1, 2, …, g, and g represents the total number of the numbers of the vehicles and the personnel; It should be noted that the local coordinate system of the closed material yard is adopted, for example, taking the northwest corner of the closed material yard as the origin, the X-axis along the east-west direction, the Y-axis along the south-north direction, and the Z-axis perpendicular to the ground, and the positioning data of the visual pixel coordinates, the UWB geodetic coordinates and the laser radar polar coordinates are converted into three-dimensional coordinates in the unified coordinate system .
[0037] The weighted least squares method is adopted to fuse the multi-source positioning results, and the formula is to obtain the final positioning results of the vehicles and the personnel in the closed material yard, wherein represents the real-time dynamic weight of each perception sensor.
[0038] The digital twin mapping unit is configured to map the final positioning results of the vehicles and the personnel in the closed material yard to the pre-constructed digital twin model of the closed material yard for visual display, analyze the corresponding vehicle personnel distances, and perform corresponding early warning display. It should be noted that the pre-constructed digital twin model of the closed material yard includes the following steps: Data acquisition: including geometric data acquisition, physical property data acquisition and operation data acquisition; specifically, using laser radar, total station and other measuring equipment, the building structure, equipment layout (such as stacker reclaimer, conveyor belt, etc.), pile shape and position of the closed material yard are accurately measured to obtain the geometric size and spatial position information; the physical properties of various objects in the material yard are collected, such as the material quality, density, particle size distribution of the pile, the weight and material quality of the equipment; the running state data of the equipment is collected in real time through sensors (such as temperature sensors, pressure sensors, displacement sensors, etc.), such as the speed and temperature of the motor, the pressure of the hydraulic system, etc.; at the same time, the production data of the material yard is recorded, such as the in-out quantity and storage time of the material.
[0039] Model construction: including geometric modeling, physical modeling and behavior modeling; specifically, using three-dimensional modeling software (such as AutoCAD, SolidWorks, 3dsMax, etc.), according to the collected geometric data to build a three-dimensional geometric model of the closed stockyard, accurately restore the shape and position of the buildings, equipment and stockpile; based on physical property data, give the geometric model physical properties. For example, set the mass, center of gravity and other physical parameters for the stockpile model, set the kinematics and dynamics parameters for the equipment model, so that it can simulate the real physical behavior; analyze the production process and equipment operation logic of the stockyard, and establish the corresponding behavior model. For example, simulate the working process of the stacker-reclaimer, including walking, rotating, pitching and other actions, as well as the stacking and taking process of the material; consider the cooperative work and logical relationship between the equipment, ensure that the model can reflect the actual operation of the stockyard.
[0040] Model verification and optimization: compare the running data simulated by the model with the actual collected running data, check the accuracy and reliability of the model, for example, compare the differences between the running time of the equipment in the model and the actual data, the material processing capacity; according to the comparison result, adjust and optimize the parameters of the model, so that the output of the model is closer to the actual situation; test the functions of the model, such as visual display, data analysis, early warning function, etc., to ensure that the model can meet the needs of the closed stockyard digital twin system.
[0041] Model integration and update: integrate the constructed closed stockyard digital twin model into the digital twin system, connect with the data acquisition module, visualization module, etc., realize real-time interaction of data and dynamic update of model; with the changes of the actual situation of the closed stockyard, such as equipment updating and modification, stockpile re-planning, etc., update and maintain the digital twin model in time, ensure that the model is always consistent with the actual stockyard.
[0042] Specifically, analyze the corresponding vehicle-personnel distance and perform corresponding warning display, including: calculating the distance between vehicles and vehicles, the distance between vehicles and personnel by the Euclidean distance algorithm formula, comparing the distance between vehicles and vehicles with the preset vehicle distance warning distance value respectively, if the distance between vehicles and vehicles is less than the preset vehicle distance warning distance value, a vehicle collision risk signal is generated, and a red warning identifier in the closed stockyard digital twin model is triggered at the same time; compare the distance between vehicles and personnel with the preset vehicle-personnel warning distance value respectively, if the distance between vehicles and personnel is less than the preset vehicle-personnel warning distance value, a vehicle-personnel collision risk signal is generated, and a blue warning identifier in the closed stockyard digital twin model is triggered at the same time; send the sound and light warning signal to the closed stockyard management personnel terminal synchronously.
[0043] The above merely illustrates and describes the structure of the present application, and those skilled in the art can make various modifications or supplements to the described specific embodiments or adopt similar ways to replace, as long as the modifications or supplements do not deviate from the structure of the present application or exceed the scope defined by the present claims, and should belong to the protection scope of the present application.
Claims
1. A closed yard vehicle personnel collaborative positioning system based on machine vision fusion, characterized in that, The method comprises the following steps: A multi-modal perception module is used to collect multi-source visual positioning raw data, UWB positioning data and laser radar positioning data of the enclosed material yard in real time according to a multi-modal perception sensor; An environment feature analysis module is used to obtain environment data of the enclosed material yard, wherein the environment data comprises dust concentration and illumination intensity, and the environment data is filtered by an average filtering algorithm to obtain filtered dust concentration and illumination intensity of the enclosed material yard; A visual enhancement module is used to enhance the image in the multi-source visual positioning data according to the filtered dust concentration and illumination intensity of the enclosed material yard, and obtain an enhanced image; A multi-modal feature extraction module is used to extract features of the positioning data of the multi-modal perception sensor to generate multi-modal perception sensor positioning feature information, and output respective reliability score values of the multi-modal perception sensor based on a reinforced deep learning model; A dynamic weight fusion module is used to match the respective reliability score values of the multi-modal perception sensor with a preset basic weight rule base corresponding to each reliability score value to obtain a basic weight of each perception sensor, and dynamically adjust the basic weight by a Kalman filtering algorithm to obtain a real-time dynamic weight of each perception sensor; A cooperative positioning settlement module is used to calculate the final positioning result of the vehicle and the personnel in the enclosed material yard according to the obtained real-time dynamic weight of the multi-modal perception sensor and the positioning data of the multi-modal perception sensor; A digital twin mapping unit is used to map the final positioning result of the vehicle and the personnel in the enclosed material yard to a pre-constructed digital twin model of the enclosed material yard for visual display, analyze the corresponding vehicle personnel distance, and display a corresponding warning.
2. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 1, wherein, The specific steps of the filtering process are as follows: The enclosed material yard is divided according to a preset spatial volume division method to obtain each sub-space region of the enclosed material yard, the dust concentration of each sub-space region of the enclosed material yard is detected by a laser scattering dust sensor to obtain dust concentration data of each sub-space region of the enclosed material yard, and the illumination intensity of each sub-space region of the enclosed material yard is detected by a digital illumination sensor to obtain illumination intensity data of each sub-space region of the enclosed material yard; The dust concentration data and the illumination intensity data of each sub-space region of the enclosed material yard are filtered according to a preset sliding window size, and the filtered dust concentration and the illumination intensity of each sub-space region of the enclosed material yard are calculated according to an average filtering algorithm formula; The filtered dust concentration and the illumination intensity of each sub-space region of the enclosed material yard are processed by mean value to obtain the mean value of the filtered dust concentration and the mean value of the filtered illumination intensity of the enclosed material yard as the filtered dust concentration and the filtered illumination intensity of the enclosed material yard.
3. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 1, wherein, The visual enhancement module comprises a deep learning defogging unit, an HDR image processing unit and a motion artifact elimination unit; wherein the deep learning defogging unit is used to perform defogging processing analysis based on the image in the multi-source visual positioning raw data and the dust concentration of the enclosed material yard to obtain a defogged image. The HDR image processing unit is configured to perform HDR processing on the defogged image meeting the image quality requirement based on the light intensity of the closed field to obtain an HDR processed image. The motion artifact elimination unit is configured to determine whether the HDR image has an artifact region, and perform motion artifact elimination processing if the artifact region exists to obtain an enhanced processed image.
4. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 3, wherein, The deep learning defogging unit includes: 301-1: Obtain the dust concentration of the filtered closed field, and compare it with the preset low dust concentration interval, medium dust concentration interval and high dust concentration interval. If the dust concentration of the filtered closed field is in the low dust concentration interval, obtain the average value of the first 0.1% pixels in the image dark channel as the global atmospheric light value A. If the dust concentration of the filtered closed field is in the medium dust concentration interval, divide the image into a preset grid, calculate the average value of the first 0.3% pixels in the dark channel of each grid as the atmospheric light value of the region, and calculate the average value of the atmospheric light values of all regions as the global atmospheric light value A. If the dust concentration of the filtered closed field is in the high dust concentration interval, first calculate the basic global atmospheric light value according to the region adaptive method in the medium dust concentration interval, introduce a dust concentration compensation coefficient, multiply the basic global atmospheric light value by the dust concentration compensation coefficient to obtain the adjusted global atmospheric light value A. 301-2: Construct a pre-trained AOD-Net model, normalize the original image and input it into the AOD-Net model. After the input image passes through the Conv1 layer with a convolution kernel of 1*1 and the Conv2 layer with a convolution kernel of 3*3, the feature matrices obtained by the two layers are fused into the Concat1 layer. Then, through the Conv3 layer with a convolution kernel of 5*5, the feature matrices obtained by the Conv2 layer and the Conv3 layer are fused into the Concat2 layer. Then, through the Conv4 layer with a convolution kernel of 7*7, all the feature matrices obtained by the previous four convolutions are fused into the Concat3 layer. Through the Conv5 layer with a convolution kernel of 3*3, a unified parameter is calculated. Finally, the defogged image is obtained through the formula. 301-3: Obtain the peak signal-to-noise ratio of the dehazed image and structural similarity and compare it with a preset peak signal-to-noise ratio threshold of the visual positioning image dehazing quality requirement and structural similarity threshold If , the dehazed image is sent to the HDR image processing unit as the input of subsequent HDR synthesis. If The defogged image is again executed steps 301-1 to 301-3 until the image quality requirements are met, i.e. wherein are represented as logical symbols, respectively, and and or.
5. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 3, wherein, The HDR image processing unit includes: Obtain the light intensity of the filtered closed field, and compare it with the preset light intensity threshold. If the light intensity of the filtered closed field is less than or equal to the preset light intensity threshold, process the defogged image through the HDR image processing software in the single-frame HDR mode to obtain an HDR processed image. If the light intensity of the filtered closed field is greater than the preset light intensity threshold, process the defogged image through the HDR image processing software in the multi-frame synthesis HDR mode to obtain an HDR processed image.
6. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 3, wherein, The motion artifact elimination unit includes: 303-1: Calculate the absolute gray difference pixel by pixel for the current HDR image and the previous frame of the HDR image, compare the absolute gray difference with the preset gray difference threshold, and determine that the pixel belongs to the artifact region if the absolute gray difference is greater than the preset gray difference threshold. 303-2: Extract ORB feature points from the current HDR image and the previous frame HDR image respectively by the ORB algorithm; estimate the pixel motion vector by using the ucas-Kanade optical flow method; for each ORB feature point in the image, the motion vector of each ORB feature point is obtained by establishing an equation system and using the least square method, forming an optical flow field; according to the calculated optical flow field, the current frame HDR image is transformed to align it with the previous frame HDR image in space, generating a preliminary corrected image; using Gaussian filtering to smooth the edges of the preliminary corrected image, obtaining the image after motion artifact elimination; 303-3: Obtain the area of the processed artifact region and the area of the artifact region before processing, calculate the artifact elimination rate, and compare the artifact elimination rate with the preset artifact elimination rate threshold; if the artifact elimination rate is greater than or equal to the preset artifact elimination rate threshold, output the enhanced image; otherwise, increase the window size of the optical flow method and re-execute steps 303-2 to 303-3 until the artifact elimination quality requirement is met.
7. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 1, wherein, The respective reliability score values of the multi-modal perception sensors based on the reinforced deep learning model include: The reinforced deep learning model DQN network structure is constructed, including an input layer, a hidden layer and an output layer, wherein the input layer receives multi-dimensional features of the multi-modal perception sensors, the hidden layer performs nonlinear transformation and abstraction on the multi-dimensional features, and the output layer outputs respective reliability score values of the multi-modal perception sensors; Define the feature information of the multi-modal perception sensor as a state vector , and take it as the input of the reinforcement deep learning model, and define the action of the reinforcement deep learning model ; Assign respective reliability score values to the multi-modal perception sensors, and design a reward function of the reinforcement deep learning model ; The reliability score vector of the multi-modal perception sensor combination at the current time step is set, and the reliability score vector of the multi-modal perception sensor combination at the previous time step is calculated to obtain the reliability score value transformation rate; reliability score value conversion rate a preset reliability score value conversion rate threshold is compared, if an additional reward r is given, if no reward is given; Reinforcement deep learning model DQN based on state vector Each action is calculated through forward propagation. The Q-value represents the expected cumulative reward for performing an action in the state vector. The deep reinforcement learning model selects the action with the largest Q-value as the output, thereby obtaining the reliability score of each multimodal sensing sensor.
8. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 1, wherein, The dynamic adjustment of the basic weight by the Kalman filtering algorithm includes the following steps: 501: State definition and observation equation establishment: define the weight of multi-modal perception sensor as a state vector wherein represents the weight of the i-th perception sensor, and n is the number of perception sensors; set the observation equation of the i-th perception sensor as wherein is the observation vector of the i-th perception sensor, is the observation matrix, and the state vector is mapped to the observation space, is the observation noise, having a covariance matrix Integrate the observation equations of all perception sensors together to obtain the overall observation equation wherein , is a block matrix composed of , and the covariance matrix is a block diagonal matrix composed of ; 502: initialization: preset initialization state vector estimate and covariance matrix of the initialization state estimate ; 503: prediction: predict the state vector according to the dynamic model of the system, set the dynamic model of the system as wherein is the state vector at the current time, is the state vector at the previous time, that is, the current time is , the previous time is , is the state transition matrix, is the process noise, and the covariance matrix is , through the model, the state vector at the time is predicted , and the covariance matrix of the state estimation is predicted 504: update: by formula The Kalman gain is calculated ; with the new observation The state estimate is updated, the update formula is The covariance matrix of the state estimate is updated ; 505: Iteratively run: the state vector estimation of the current time instant and the covariance matrix of the state estimation As the input of the step 503 of the next time instant, the execution operation of the steps 503-504 is repeatedly repeated, thereby obtaining the real-time dynamic weight of each perception sensor.
9. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 1, wherein, The final positioning result of the vehicle and the personnel in the closed material field is calculated, including: The multi-source visual positioning data after visual enhancement processing, the UWB positioning data and the laser radar positioning data are subjected to coordinate unification and format conversion, to obtain three-dimensional coordinates of the vehicle and the personnel in a unified coordinate system of each perception sensor, denoted as wherein f represents the number of the vehicle and the personnel, f = 1, 2, …, g, and g represents the total number of the vehicle and the personnel numbers; The weighted least square method is used to fuse the multi-source positioning results, and the formula is The final positioning results of the vehicles and personnel in the closed stockyard are obtained, wherein is the real-time dynamic weight of each perception sensor.
10. The machine vision fusion based closed lot vehicle personnel collaborative positioning system according to claim 1, wherein, The corresponding vehicle personnel distance is analyzed, and the corresponding warning display is performed, including: The distance between vehicles and the distance between vehicles and personnel are calculated by the Euclidean distance algorithm formula, and the distance between vehicles is compared with the preset vehicle distance warning distance value, if the distance between vehicles is less than the preset vehicle distance warning distance value, a vehicle collision risk signal is generated, and a red warning identifier in the closed material field digital twin model is triggered; the distance between vehicles and personnel is compared with the preset vehicle-person warning distance value, if the distance between vehicles and personnel is less than the preset vehicle-person warning distance value, a vehicle-person collision risk signal is generated, and a blue warning identifier in the closed material field digital twin model is triggered; a sound and light warning signal is sent to the closed material field management personnel terminal.