A device dynamic obstacle avoidance method and system based on behavior trend prediction, device, medium
By performing image frame detection and back-projection processing during drilling operations, combined with morphological dilation and machine learning models, the boundary uncertainty problem of equipment and personnel collision detection in drilling operations was solved, achieving higher accuracy in early warning and obstacle avoidance.
Patent Information
- Application Number
- CN202511386068.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing collision detection technologies based on camera images in drilling operations suffer from problems such as uncertain boundaries and insufficient spatial reconstruction, resulting in large collision judgment errors and poor early warning effects.
By acquiring image frames for target detection and device semantic segmentation, a ground two-dimensional coordinate system is established for back projection. Combined with morphological dilation correction and machine learning models, the location of equipment and personnel is predicted, and an early warning model is established to output early warning signals.
It improves the accuracy of restoring the two-dimensional spatial relationship between equipment and personnel and the stability of collision prediction, significantly enhancing the accuracy and range of early warning, especially improving the stability and interpretability of projection under high-angle oblique shooting conditions.
Smart Images

Figure CN120894832B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine vision recognition technology, and more specifically, to a method, system, device, and medium for dynamic obstacle avoidance based on behavioral trend prediction. Background Technology
[0002] With the rapid development of intelligent sensing and industrial computing technologies, image segmentation and physical projection technologies have shown great potential in fields such as intelligent manufacturing, safety monitoring, and human-machine collaboration. As a multimodal fusion method combining machine vision and equipment state modeling, the equipment-personnel collision prediction system, by introducing image back-projection modeling and a multi-scale dilation weighted regression mechanism, not only significantly improves the fitting accuracy of the prediction area but also ensures the consistency of ground projection and the accuracy of collision judgment, thereby promoting the development of drilling platform safety automation technology. In drilling operation environments, ensuring the accuracy of equipment spatial distribution and personnel behavior prediction is crucial for executing high-precision safety avoidance. Operational safety often requires spatiotemporal accuracy, boundary judgment stability, and predictive foresight to meet the needs of real-time obstacle avoidance and behavioral alerts. However, the limitations of existing technologies result in uncertain boundaries and insufficient spatial reconstruction in camera image-based collision detection, leading to large judgment errors and poor early warning effects. Summary of the Invention
[0003] The purpose of this invention is to provide a device dynamic obstacle avoidance method, system, device, and medium based on behavioral trend prediction to solve the above-mentioned problems in the prior art.
[0004] This invention is achieved through the following technical solution:
[0005] A device dynamic obstacle avoidance method based on behavioral trend prediction includes:
[0006] Obtain the current image input frame, and based on the current image input frame, obtain the object detection bounding box for person object detection and the device semantic segmentation mask:
[0007] Establish a ground two-dimensional coordinate system, calculate the coordinates of the center pixel of the bottom edge of the target detection box, and back-project the center pixel of the bottom edge onto the ground two-dimensional coordinate system;
[0008] Obtain all pixels of the device semantic segmentation mask and camera parameters, back-project the pixels to the ground two-dimensional coordinate system based on the camera parameters to obtain the outline polygon of the projection area, obtain the effective projection area of the device based on the outline polygon of the projection area, and perform morphological dilation correction on the effective projection area of the device to obtain the final effective projection area of the device.
[0009] Acquire several consecutive frames of images after the current image input frame. Based on the consecutive frames of images, obtain the state vectors of personnel and equipment for the personnel and equipment in the several consecutive frames of images. Use the state vectors and state transition matrix to perform k-step recursion to obtain the predicted positions of personnel and equipment.
[0010] An early warning model is established. Based on the predicted location of personnel and equipment, the current early warning score is output through the evaluation model. An early warning threshold is set. When the early warning score is greater than the early warning threshold, an alarm control signal is issued.
[0011] Preferably, the calculation of the coordinates of the center pixel of the bottom edge of the target detection box includes:
[0012] ;
[0013] ;
[0014] ;
[0015] In the formula, For the target detection bounding box, These are the pixel coordinates of the target detection bounding box. , These are the coordinates of the center pixel of the bottom edge. , These represent the coordinates of the personnel on the ground after back projection. The inverse function of the camera. The height is the 3D world coordinate of the center pixel of the bottom edge.
[0016] Preferably, obtaining the projection region outline polygon and obtaining the effective projection region of the device based on the projection region outline polygon includes:
[0017] ;
[0018] ;
[0019] In the formula, and The first The coordinates of the pixels of the semantic segmentation mask of a device projected onto the ground in a two-dimensional coordinate system. and These are the first three semantic segmentation masks in the device. The coordinates of each pixel For the set of pixels of the semantic segmentation mask, The projection region is a polygonal outline, specifically a set of projection points;
[0020] Calculate the polygon outline of the projection area Geometric centroid coordinates Select Greater than or equal to Using the target point as the objective point, calculate the minimum circumscribed convex polygon based on the objective point to construct the effective projection area of the device.
[0021] Preferably, the morphological dilation correction of the effective projection area of the device includes:
[0022] Set the expansion kernel to a square structuring element, and finally define the final expansion mask:
[0023] ;
[0024] Using machine learning models for prediction:
[0025] ;
[0026] Machine learning model input PLC status Output an integer value. Finally, based on the number of times learned by the machine learning model, the result of the inflation is obtained. Constructing new boundaries ;
[0027] ;
[0028] In the formula, For the final expansion mask, For an expanding core, The effective projection area of the device. The output of the machine learning model, PLC status. For machine learning models, For the largest scale, For convex hull algorithm, This is an XOR operation.
[0029] Preferably, the state vector includes:
[0030] ;
[0031] ;
[0032] ;
[0033] ;
[0034] ;
[0035] In the formula, For state vectors, , for The speed of time , for acceleration at any moment , for The acceleration of time, Inter-frame time, This is a disturbance term.
[0036] Preferably, the step of using the state vector and state transition matrix to perform k-step recursion to obtain the predicted positions of personnel and equipment includes:
[0037] ;
[0038] In the formula, Predicted state Here is the state transition matrix. To predict the step size.
[0039] Preferably, the establishment of the early warning model includes:
[0040] ;
[0041] ;
[0042] ;
[0043] ;
[0044] In the formula, To predict the intersection-union ratio (IU / UK) of the bounding boxes for personnel and the expanded bounding boxes for equipment, Distance from the center , The geometric center coordinates of the personnel. , for Corresponding geometric center coordinates As relative kinetic energy, Let V be the planar velocity vector of the person in the ground coordinate system. Let V be the planar velocity vector of the device in the ground coordinate system. For human equivalent mass For warning scores, , , Risk weights.
[0045] Secondly, the present invention also provides a device dynamic obstacle avoidance system based on behavior trend prediction, for executing the above-described device dynamic obstacle avoidance method based on behavior trend prediction, comprising:
[0046] The projection processing module is configured to acquire the current image input frame, and based on the current image input frame, acquire the target detection box for personnel target detection and the device semantic segmentation mask: establish a ground two-dimensional coordinate system, calculate the coordinates of the center pixel of the bottom edge of the target detection box, and back-project the center pixel of the bottom edge onto the ground two-dimensional coordinate system; acquire all pixels of the device semantic segmentation mask and camera parameters, back-project the pixels onto the ground two-dimensional coordinate system based on the camera parameters to obtain the projection area outline polygon, obtain the device effective projection area based on the projection area outline polygon, and perform morphological dilation correction on the device effective projection area to obtain the final device effective projection area;
[0047] The early warning module is configured to acquire several consecutive frames of images after the current image input frame, acquire the state vectors of personnel and equipment in the several consecutive frames of images, and perform k-step recursion using the state vectors and state transition matrix to obtain the predicted positions of personnel and equipment; establish an early warning model, output the current early warning score based on the predicted positions of personnel and equipment through the evaluation model, set an early warning threshold, and issue an alarm control signal when the early warning score is greater than the early warning threshold.
[0048] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described device dynamic obstacle avoidance method based on behavioral trend prediction.
[0049] Fourthly, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for dynamic obstacle avoidance of a device based on behavioral trend prediction.
[0050] The technical solution of the present invention has at least the following advantages and beneficial effects:
[0051] This invention introduces an image segmentation dilation scale regression mechanism driven by device PLC features. The image segmentation dilation prediction model can effectively adjust the projection range of the device on the ground. This not only improves the fitting accuracy of the true boundary but also ensures the consistency of the dilated boundary coverage of landmarks, providing stronger data support for image-physical fusion collision modeling. The introduction of a human positioning method based on the back projection of the image's bottom edge center point and the construction of the device's lower edge convex edge is a key step in collision modeling, enabling the reconstruction of the two-dimensional spatial relationship between the device and the person. Through these methods, not only is the diversity of different device forms considered, but the orientation and directionality characteristics of the device can also be accurately modeled, significantly improving the stability and interpretability of collision prediction. Especially under high-angle oblique camera conditions, this technology significantly improves projection stability and interpretability.
[0052] By combining an upgraded behavior state prediction method with a state variable update mechanism that includes velocity, acceleration, and momentum, continuous prediction of equipment or personnel behavior trends is achieved. Compared to traditional Kalman filtering, this method exhibits stronger dynamic response and high robustness under multi-source disturbances and non-stationary behavior, significantly improving prediction performance and early warning range. Attached Figure Description
[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a schematic diagram of the process of the present invention;
[0055] Figure 2 This is a schematic diagram of the system structure of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0057] The module divisions described in this application are logical divisions. In practical applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the connections, couplings, or communications in this application can be direct connections, couplings, or communications between related objects, or indirect connections, couplings, or communications through other devices. Moreover, the connections, couplings, or communications between objects can be electrical or other similar forms, and are not limited in this application.
[0058] The independently described modules or sub-modules may or may not be physically separated; they may be implemented in software or hardware, and some modules or sub-modules may be implemented in software, with the processor calling the software to implement the function of these modules or sub-modules, while other modules or sub-modules may be implemented in hardware, such as through hardware circuits. Furthermore, some or all of the modules can be selected to achieve the purpose of this application's solution according to actual needs.
[0059] Please refer to Figures 1-2 The present invention provides a device dynamic obstacle avoidance method based on behavioral trend prediction, comprising:
[0060] S101: Obtain the current image input frame, and based on the current image input frame, obtain the target detection bounding box for person detection and the device semantic segmentation mask:
[0061] S102: Establish a ground two-dimensional coordinate system, calculate the coordinates of the center pixel of the bottom edge of the target detection box, and back-project the center pixel of the bottom edge to the ground two-dimensional coordinate system;
[0062] Because drilling platform cameras typically capture images from a top-down perspective, the actual positions of personnel and equipment are distorted in the images, leading to misjudgments when directly performing collision detection. Therefore, it is necessary to back-project the image pixels back to the ground coordinate system to construct a unified two-dimensional space. The position of objects on the ground is obtained by back-projecting the camera imaging model. Homogeneous coordinate transformation is then performed using the camera's intrinsic and extrinsic parameters.
[0063] ;
[0064] The projected two-dimensional pixel coordinates are usually in homogeneous coordinate form; This represents the camera intrinsic parameter matrix, including focal length and principal point, and its size. ; Represents camera extrinsic parameters; represents spatial transformation. Points in the world coordinate system are represented in homogeneous form; devices and people are projected onto a two-dimensional projection plane in this way.
[0065] When shooting from a top-down view of an drilling platform, the bounding box of a person in the image does not represent their actual position; the center point of the bottom edge best represents their true location on the ground. Therefore, this point needs to be back-projected from the image pixels onto the ground plane as the coordinate basis for collision detection.
[0066] S103: Obtain all pixels of the device semantic segmentation mask and camera parameters, back-project the pixels to the ground two-dimensional coordinate system based on the camera parameters to obtain the projection area outline polygon, obtain the device effective projection area based on the projection area outline polygon, perform morphological dilation correction on the device effective projection area to obtain the final device effective projection area.
[0067] S104: Obtain several consecutive frame images after the current image input frame, obtain the state vectors of personnel and equipment in several consecutive frame images based on the consecutive frame images, and use the state vectors and state transition matrix to perform k-step recursion to obtain the predicted positions of personnel and equipment.
[0068] S105: Establish an early warning model, output the current early warning score based on the predicted location of personnel and equipment through the evaluation model, set an early warning threshold, and issue an alarm control signal when the early warning score is greater than the early warning threshold.
[0069] This invention introduces an image segmentation dilation scale regression mechanism driven by device PLC features. The image segmentation dilation prediction model can effectively adjust the projection range of the device on the ground. This not only improves the fitting accuracy of the true boundary but also ensures the consistency of the dilated boundary coverage of landmarks, providing stronger data support for image-physical fusion collision modeling. The introduction of a human positioning method based on the back projection of the image's bottom edge center point and the construction of the device's lower edge convex edge is a key step in collision modeling, enabling the reconstruction of the two-dimensional spatial relationship between the device and the person. Through these methods, not only is the diversity of different device forms considered, but the orientation and directionality characteristics of the device can also be accurately modeled, significantly improving the stability and interpretability of collision prediction. Especially under high-angle oblique camera conditions, this technology significantly improves projection stability and interpretability.
[0070] By combining an upgraded behavior state prediction method with a state variable update mechanism that includes velocity, acceleration, and momentum, continuous prediction of equipment or personnel behavior trends is achieved. Compared to traditional Kalman filtering, this method exhibits stronger dynamic response and high robustness under multi-source disturbances and non-stationary behavior, significantly improving prediction performance and early warning range.
[0071] In one exemplary embodiment of the present invention, calculating the coordinates of the center pixel of the bottom edge of the target detection box includes:
[0072] ;
[0073] ;
[0074] ;
[0075] In the formula, For the target detection bounding box, These are the pixel coordinates of the target detection bounding box. , These are the coordinates of the center pixel of the bottom edge. , These represent the coordinates of the personnel on the ground after back projection. The inverse function of the camera. The height of the 3D world coordinates of the center pixel of the bottom edge. This means that the center pixel of the bottom edge is on the ground.
[0076] In one exemplary embodiment of the present invention, under top-view camera conditions, the segmentation mask of the device in the image exhibits a perspective compression effect. To determine a collision on the ground, it is necessary to identify the "most likely contact area" of the device with the ground. The polygon formed by back-projecting the mask region in the image has its lower boundary, with its geometric centroid closest to the projection area generated by the camera's viewpoint. Therefore, a ground projection area is constructed.
[0077] Specifically, obtain all pixels in the device segmentation mask, back-project them onto the ground coordinate plane using camera parameters, back-project the mask points onto the ground, and construct the outline polygon of the device projection area.
[0078] The obtained projection region contour polygon, and the effective projection region of the device based on the projection region contour polygon, include:
[0079] ;
[0080] ;
[0081] In the formula, and The first The coordinates of the pixels of the semantic segmentation mask of a device projected onto the ground in a two-dimensional coordinate system. and These are the first three semantic segmentation masks in the device. The coordinates of each pixel For the set of pixels of the semantic segmentation mask, The projection region is a polygonal outline, specifically a set of projection points;
[0082] Calculate the polygon outline of the projection area Geometric centroid coordinates Select Greater than or equal to The point is taken as the target point, and the minimum circumscribed convex polygon is calculated based on the target point to construct the effective projection area of the device. The calculation method of the geometric center and the convex hull algorithm are existing technologies in this field, and will not be described in detail here.
[0083] ;
[0084] ;
[0085] Calculate the minimum bounded convex polygon for these points, define it as the effective projection region of the device, and construct the effective projection region of the device:
[0086] ;
[0087] This is the set of ground points located below (including) the center of gravity. This is the effective projection area of the device (minimum circumscribed convex polygon), used for subsequent collision detection and expansion modeling.
[0088] In one exemplary embodiment of the present invention, due to variations in camera angle and device height, the device outline segmented in the image is typically smaller than its actual footprint projection, failing to cover the actual area. Morphological dilation correction is employed. Based on the segmented image region, multiple dilation operations are performed using a fixed structuring element to expand its ground projection boundary. When using a multi-scale dilation scheme, the structuring element acts on the effective projection area of the segmented device, and weights are used. This represents the contribution of dilation at different scales. However, the device cannot be infinitely large in an image. Increasing the dilation kernel size is meaningless when the dilated region already covers the entire image. Therefore, a weight vector needs to be set based on the image size (H, W) and the dilation kernel size. The maximum dimension limit.
[0089] The morphological dilation correction of the effective projection area of the device includes:
[0090] The expansion kernel is set as a square structuring element with dimensions (2r+1)*(2r+1), where r is the radius of the circular kernel. Finally, the final expansion mask is defined.
[0091] ;
[0092] Using machine learning models for prediction:
[0093] ;
[0094] Machine learning model input PLC status Output an integer value. Finally, based on the number of times learned by the machine learning model, the result of the inflation is obtained. Constructing new boundaries PLC status The state vector represents the PLC, which is a set of continuous signals acquired and output by the controller. Examples include features such as multiple digital inputs, multiple digital outputs, multiple analog inputs, multiple program segment numbers, and multiple alarm flags.
[0095] ;
[0096] In the formula, For the final expansion mask, For an expanding core, The effective projection area of the device. The output of the machine learning model, PLC status. For machine learning models (Naive Bayes is selected). For the largest scale, For convex hull algorithm, This is an XOR operation.
[0097] In one exemplary embodiment of the present invention, to predict the future position of equipment or personnel, it is necessary to consider not only velocity but also acceleration, and even jerk, to improve the accuracy of short-term dynamic prediction. An eight-dimensional state vector is established, and the future position state is recursively derived using third-order Newtonian kinematics formulas.
[0098] The state vector includes:
[0099] ;
[0100] ;
[0101] ;
[0102] ;
[0103] ;
[0104] In the formula, For state vectors, , for The speed of time , for acceleration at any moment , for The acceleration of time, Inter-frame time, This is a disturbance term.
[0105] By using image segmentation models or key point detection, the pixel position of the center of a person's foot or the bottom edge of equipment protrusion in the image can be identified. After camera calibration, an imaging-backprojection matrix is obtained, converting pixel positions into ground projection positions. This point serves as the starting point for all dynamic states and is considered an observation; subsequent velocities and accelerations are estimated through difference / fitting. When the frame interval is... Estimate based on the position difference between two consecutive frames:
[0106] ;
[0107] and Let be the coordinates at time t. and The coordinates are at time t-1.
[0108] Use the past Frame position information is fitted with a linear trend using least squares:
[0109] ;
[0110] In the formula, For frame number, The speed to be solved is denoted as .
[0111] The location sequence obtained from the images requires at least 3 frames per target to stably estimate the trend, direction, and speed of movement of equipment or personnel on the ground, which is the primary factor in predicting future locations.
[0112] ;
[0113] Velocity changes are extracted from consecutive frames of images. Sudden acceleration of a person (such as running), sudden start-up of equipment, deceleration, and turning all correspond to obvious changes in acceleration; abrupt changes in acceleration are often important precursory information for behavioral intentions.
[0114] jerk. Definition of the rate of change of acceleration:
[0115] ;
[0116] Obtained by differencing the acceleration estimation sequence; high-frequency jitter can be removed by low-pass filtering while preserving trend changes; represents the "abruptness of motion behavior"; sudden turns, jumps, violent starts or stops of equipment all have large jerk values; the larger the jerk value, the more likely it is a "high-risk impact behavior". The derivation of these three variables follows the physical time derivative chain:
[0117] Location speed acceleration jerk.
[0118] In discrete-time systems, it is approximated as:
[0119] ;
[0120] In the formula, , and These represent the velocity, acceleration, and jerk in a discrete-time system, respectively.
[0121] In one exemplary embodiment of the present invention, the step of performing k-step recursion using state vectors and state transition matrices to obtain the predicted positions of personnel and equipment includes:
[0122] ;
[0123] In the formula, Predicted state Here is the state transition matrix. To predict the step size.
[0124] In this embodiment, the state transition matrix is: ;
[0125] It should be noted that the first two lines update the coordinates. Position is determined by velocity and acceleration; lines 3-4 update velocity: velocity is affected by acceleration; lines 5-6 keep acceleration constant; lines 7-8 are constant, indicating that angular velocity and height remain unchanged.
[0126] In one exemplary embodiment of the present invention, the establishment of the early warning model includes:
[0127] ;
[0128] ;
[0129] ;
[0130] ;
[0131] In the formula, To predict the intersection-union ratio (IU / UK) of the bounding boxes for personnel and the expanded bounding boxes for equipment, Distance from the center , The geometric center coordinates of the personnel (that is, the ones in front) , (human coordinates) , for Corresponding geometric center coordinates As relative kinetic energy, Let V be the planar velocity vector of the person in the ground coordinate system. Let V be the planar velocity vector of the device in the ground coordinate system. The equivalent mass of a human body (here set to the common 70 kg). For warning scores, , , Risk weights.
[0132] Weighting coefficient yes The weighted average of center distance and kinetic energy, on the labeled "safe / dangerous" samples, are hyperparameters that need to be adjusted and can be set according to specific circumstances.
[0133] It should be noted that the bottom midpoint of the target detection box selected in this embodiment can be understood as a small circle with a certain area, and the intersection-union ratio is calculated with the effective projection area of the device, which is also a two-dimensional polygon.
[0134] in, , The threshold can be flexibly set according to the on-site safety strategy. Alarm methods include: real-time voice prompts saying "Personnel approaching a danger zone"; recording alarm logs in the control system; and controlling equipment to automatically stop.
[0135] A device dynamic obstacle avoidance system based on behavior trend prediction, used to execute the above-mentioned device dynamic obstacle avoidance method based on behavior trend prediction, includes:
[0136] The projection processing module is configured to acquire the current image input frame, and based on the current image input frame, acquire the target detection box for personnel target detection and the device semantic segmentation mask: establish a ground two-dimensional coordinate system, calculate the coordinates of the center pixel of the bottom edge of the target detection box, and back-project the center pixel of the bottom edge onto the ground two-dimensional coordinate system; acquire all pixels of the device semantic segmentation mask and camera parameters, back-project the pixels onto the ground two-dimensional coordinate system based on the camera parameters to obtain the projection area outline polygon, obtain the device effective projection area based on the projection area outline polygon, and perform morphological dilation correction on the device effective projection area to obtain the final device effective projection area;
[0137] The early warning module is configured to acquire several consecutive frames of images after the current image input frame, acquire the state vectors of personnel and equipment in the several consecutive frames of images, and perform k-step recursion using the state vectors and state transition matrix to obtain the predicted positions of personnel and equipment; establish an early warning model, output the current early warning score based on the predicted positions of personnel and equipment through the evaluation model, set an early warning threshold, and issue an alarm control signal when the early warning score is greater than the early warning threshold.
[0138] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0139] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. This computer software product, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0140] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A device dynamic obstacle avoidance method based on behavior trend prediction, characterized in that, The method comprises: obtaining a current image input frame, obtaining a target detection box of personnel target detection and a device semantic segmentation mask based on the current image input frame; establishing a ground two-dimensional coordinate system, calculating the coordinates of the bottom edge center pixel point of the target detection box, and back-projecting the bottom edge center pixel point to the ground two-dimensional coordinate system; obtaining all pixel points of the device semantic segmentation mask and camera parameters, back-projecting the pixel points to the ground two-dimensional coordinate system based on the camera parameters to obtain a projection area contour polygon, obtaining a device effective projection area based on the projection area contour polygon, and performing morphological dilation correction on the device effective projection area to obtain a final device effective projection area; obtaining a plurality of continuous frame images after the current image input frame, obtaining state vectors of personnel and devices based on the continuous frame images, respectively, performing k-step recursion using the state vectors and a state transition matrix to obtain predicted positions of the personnel and the devices; establishing an early warning model, outputting a current early warning score through the early warning model based on the predicted positions of the personnel and the devices, setting an early warning threshold, and issuing an alarm control signal when the early warning score is greater than the early warning threshold; establishing the early warning model comprises: ; ; ; ; wherein, is the intersection over union of the bounding box of the person and the inflated bounding box of the device, is the target bounding box, is the inflated bounding box of the person, is the new bounding box, is the final inflated mask, is the center distance, , is the corresponding geometric center coordinate of the person, , is the corresponding geometric center coordinate, is the relative kinetic energy, is the planar velocity vector of the person in the ground coordinate system, is the planar velocity vector of the device in the ground coordinate system, is the equivalent mass of the human body, is the early warning score, , , is the risk weight.
2. The device dynamic obstacle avoidance method based on behavior trend prediction according to claim 1, characterized in that, the calculation of the coordinates of the bottom edge center pixel point of the target detection box comprises: ; ; ; In the formula, is a target detection frame, are pixel coordinates of the target detection frame, respectively, , are coordinates of the bottom center pixel point, respectively, , are coordinate positions of the personnel on the ground after back projection, respectively, is a camera inverse function, is the height of the three-dimensional world coordinate where the bottom center pixel point is located.
3. The device dynamic obstacle avoidance method based on behavior trend prediction according to claim 2, characterized in that, the obtaining of the projection area contour polygon, and the obtaining of the device effective projection area based on the projection area contour polygon comprise: ; ; In the formula, and respectively are the coordinates of the pixel points of the first device semantic segmentation mask projected to the ground two-dimensional coordinate system, and respectively are the coordinates of the first pixel point in the device semantic segmentation mask, is a pixel point set of the semantic segmentation mask, is a projection area contour polygon, specifically a projection point set; Computing a projection area polygon of a geometric barycentric coordinate , selecting a point greater than or equal to as a target point, calculating a minimum circumscribed convex polygon based on the target point, and constructing a device effective projection area.
4. The device dynamic obstacle avoidance method based on behavior trend prediction according to claim 3, characterized in that, the morphological dilation correction on the device effective projection area comprises; setting a square structural element as a dilation kernel, and finally defining a final dilation mask: ; using a machine learning model for prediction comprises: ; Machine learning model input PLC state , output an integer value , finally according to the number of times learned by the machine learning model, get the new boundary of the inflation structure ; ; wherein, is the final dilation mask, is the dilation kernel, is the device effective projection area, is the output of the machine learning model, is the PLC status, is the machine learning model, is the maximum scale, is the convex hull algorithm, is the XOR operation.
5. The device dynamic obstacle avoidance method based on behavior trend prediction according to claim 4, characterized in that, the state vector comprises: ; ; ; ; ; wherein is the state vector, , is the velocity at time, , is the acceleration at time, , is the jerk at time, is the inter-frame time, is the perturbation term.
6. The device dynamic obstacle avoidance method based on behavior trend prediction according to claim 5, characterized in that, the k-step recursion using the state vectors and the state transition matrix to obtain the predicted positions of the personnel and the devices comprises: wherein predicted state, is a state transition matrix, is a prediction step.
7. A device dynamic obstacle avoidance system based on behavior trend prediction, characterized in that, a device dynamic obstacle avoidance method based on behavior trend prediction for executing any one of claims 1-6, comprising: a projection processing module configured to obtain a current image input frame, obtain a target detection box of personnel target detection and a device semantic segmentation mask based on the current image input frame, establish a ground two-dimensional coordinate system, calculate the coordinates of the bottom edge center pixel point of the target detection box, and back-project the bottom edge center pixel point to the ground two-dimensional coordinate system, obtain all pixel points of the device semantic segmentation mask and camera parameters, back-project the pixel points to the ground two-dimensional coordinate system based on the camera parameters to obtain a projection area contour polygon, obtain a device effective projection area based on the projection area contour polygon, and perform morphological dilation correction on the device effective projection area to obtain a final device effective projection area; an early warning module configured to obtain a plurality of continuous frame images after the current image input frame, obtain state vectors of personnel and devices based on the continuous frame images, respectively, perform k-step recursion using the state vectors and a state transition matrix to obtain predicted positions of the personnel and the devices, establish an early warning model, output a current early warning score through the early warning model based on the predicted positions of the personnel and the devices, set an early warning threshold, and issue an alarm control signal when the early warning score is greater than the early warning threshold; establishing the early warning model comprises: ; ; ; ; wherein, is the intersection over union of the predicted bounding box of the person and the inflated bounding box of the device, is the target bounding box, is the inflated bounding box, is the constructed new bounding box, is the final inflated mask, is the center distance, , is the corresponding geometric center coordinate of the person, , is the corresponding geometric center coordinate, is the relative kinetic energy, is the planar velocity vector of the person in the ground coordinate system, is the planar velocity vector of the device in the ground coordinate system, is the equivalent mass of the human body, is the early warning score, , , is the risk weight.
8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the device dynamic obstacle avoidance method based on behavior trend prediction in any one of claims 1-6 when executing the computer program.
9. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium and is executed by the processor to implement the device dynamic obstacle avoidance method based on behavior trend prediction in any one of claims 1-6.
Citation Information
Patent Citations
Image analysis method and apparatus, model training method and apparatus, and device, medium and program
WO2023178951A1