Room obstacle detection optimization method and system based on cross-modal attention mechanism

By adopting a cross-modal attention mechanism in room obstacle detection, combining multi-spectral sensors and equipment motion states, dynamically adjusting the modal weights to generate obstacle topology maps, the problem of low detection accuracy of dynamic obstacles in complex environments is solved, and efficient obstacle distinction and path planning are achieved.

CN120178893AActive Publication Date: 2025-06-20BEIJING XINGWANG SHIP POWER TECH CO LTD

Patent Information

Application Number
CN202510661437.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The prior art has low detection accuracy of dynamic obstacles in complex environments and weak ability to distinguish multiple obstacles, resulting in the risk of missed detection or collision.

Method used

The room obstacle detection optimization method based on the cross-modal attention mechanism is adopted, and the visible light and near-infrared dual-band data are obtained through multi-spectral sensors, and the modal weight is dynamically adjusted in combination with the equipment's motion state to generate obstacle topology maps, identify living and static obstacles, and generate path planning parameters.

Benefits of technology

It significantly improves the ability of mobile devices to distinguish biological and static obstacles and the response speed of dynamic obstacle avoidance in complex scenarios, improves environmental adaptability, and ensures real-time and safe path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178893A_ABST
    Figure CN120178893A_ABST
Patent Text Reader

Abstract

The invention provides a room obstacle detection optimization method and system based on a cross-modal attention mechanism. Wherein a dynamic compensation coefficient is calculated by acquiring motion track data of the mobile equipment, and a stable visual coordinate system is constructed. Visible light and near-infrared wave band acquisition of the multispectral sensor is synchronously triggered, the visible light captures geometric deformation characteristics of a floor projection surface, and the near-infrared detects a local heat radiation abnormal area. And extracting a thermal feature connected domain conforming to periodic biological motion as a living obstacle, and marking a collision risk boundary of a static obstacle. And based on a living body motion law and a collision boundary deviation trend, synchronously optimizing the dual-band exposure parameters and updating the trajectory prediction model, and forming a closed-loop feedback link of sensing and planning. According to the technical scheme provided by the invention, the living body and the static obstacle are accurately distinguished in a dynamic environment, and the sensor parameters are adaptively adjusted to improve the real-time obstacle avoidance response capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent robot environmental perception and navigation optimization, and particularly to an optimized method and system for room obstacle detection based on a cross-modal attention mechanism. Background Art

[0002] In mobile robots such as floor-sweeping robots or smart home scenarios, real-time trajectory prediction and obstacle avoidance of dynamic obstacles such as pets and moving furniture need to solve core problems such as multi-modal perception data fusion, dynamic intention reasoning, and low-latency decision-making. The system needs to accurately capture the motion state and sudden behaviors of obstacles, such as a pet suddenly turning, to ensure the real-time and safety of decision-making.

[0003] Currently, some solutions adopt a lightweight method based on single-modal threshold segmentation and linear extrapolation. The moving target area is extracted through a background subtraction algorithm, and a historical trajectory sequence is constructed using the centroid coordinates of the target to directly extrapolate the future position. For multi-obstacle scenarios, simple clustering based on distance thresholds is used to classify adjacent pixel regions as the same obstacle, and finally, an obstacle avoidance instruction is generated in combination with the movement speed of the robot.

[0004] However, this solution has obvious shortcomings: poor adaptability to dynamic environments. The background subtraction algorithm is sensitive to light changes and shadow interference, and is prone to misjudging stationary objects as obstacles. The ability to distinguish multiple obstacles is weak. Clustering based on pixel distance is difficult to distinguish dense moving targets and is often misjudged as a single obstacle, leading to risks of missed detection or collision. Summary of the Invention

[0005] This application provides an optimized method and system for room obstacle detection based on a cross-modal attention mechanism to solve the problems of low prediction accuracy and weak multi-obstacle discrimination ability in the prior art.

[0006] In a first aspect, this application provides an optimized method for room obstacle detection based on a cross-modal attention mechanism, including: Obtain the motion trajectory data of a mobile device, calculate the compensation coefficients of the device's moving direction and the viewing angle jitter amplitude based on the motion trajectory data, and generate a visual coordinate system that matches the physical space; Activate the cross-modal acquisition mode of a multi-spectral sensor in the visual coordinate system, and divide the visible light and near-infrared dual-band overlapping perception area according to the moving forward area corresponding to the mobile device, where the area covered by the visible light band is used to detect the geometric deformation characteristics of ground projections, and the area covered by the near-infrared band is used to detect temperature anomaly areas related to biological signs; Input the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, and generate an obstacle topology map; In the obstacle topology map, extract the connected regions that conform to the periodic biological motion law in the temperature anomaly region as living obstacles, and determine the collision risk boundary of static obstacles according to the projection offset of the geometric deformation feature between consecutive frames; Generate the path planning parameters of the mobile device based on the periodic motion law of the living obstacle and the offset trend of the collision risk boundary.

[0007] Optionally, input the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into a cross-modal attention network, and generate dynamic allocation parameters based on the moving speed and direction change amount of the mobile device through the cross-modal attention network; Calculate the weight distribution ratio of the visible light data stream and the near-infrared data stream through the dynamic allocation parameters, and perform weighted fusion on the geometric deformation feature and the temperature anomaly region according to the weight distribution ratio; Perform cross-modal interaction on the weighted geometric deformation feature and the temperature anomaly region. Generate a deformation-thermal correlation mark by overlapping and comparing the geometric deformation contour and the temperature anomaly region in the spatial dimension, and construct a dynamic displacement chain by combining the relative offset of the deformation contour and the temperature anomaly region in consecutive frames with the displacement correction factor of the compensation coefficient in the time dimension; Generate an obstacle topology map that fuses biological features and deformations based on the dynamic displacement chain and the deformation-thermal correlation mark.

[0008] Optionally, when the direction deviation between the displacement direction of the geometric deformation contour and the migration direction of the temperature anomaly region in the dynamic displacement chain is less than a preset threshold, extract synchronous offset trajectory data from the dynamic displacement chain to construct a dynamic trajectory layer of the living obstacle; Through the spatial superposition relationship between the geometric deformation contour and the temperature anomaly region, mark the projection boundary of the geometric deformation contour and the outer edge boundary of the temperature anomaly region as the spatial distribution layer of static obstacles; Based on the fracture feature of the geometric deformation contour and the periodic intensity fluctuation of the temperature anomaly region in the deformation-thermal correlation mark, when the fracture position matches the fluctuation period phase, generate an association response layer of potential interaction behaviors through the coverage relationship between the extended region of the fracture boundary and the thermal fluctuation region; Superimpose the spatial distribution layer, the dynamic trajectory layer, and the association response layer in the order of priority to form an obstacle topology map that fuses geometric deformation features and biological thermal features.

[0009] Optionally, in the visual coordinate system, determine a moving forward area according to the moving direction of the mobile device, form a fan-shaped coverage area with the current position of the device as the vertex and extending along the moving direction in the moving forward area, and adjust the radius length of the fan-shaped coverage area according to the moving speed; Activate the dual-band synchronous acquisition mode of the visible light sensor and the near-infrared sensor within the adjusted fan-shaped coverage area, so that the physical acquisition ranges of both completely cover the fan-shaped coverage area, where the visible light sensor focuses on the continuous frame contour changes of the floor projection plane, and the near-infrared sensor focuses on the thermal intensity gradient changes between adjacent pixels; Map the continuous frame contour changes collected by the visible light sensor into geometric deformation features, calculate the position coordinates of the deformation boundary through the displacement difference of the adjacent frame contour boundaries, and at the same time map the thermal intensity gradient changes collected by the near-infrared sensor into temperature anomaly regions, and extract the boundary coordinates of the temperature anomaly regions through the spatial distribution of the thermal intensity mutation points; Based on the position coordinates and the boundary coordinates, divide the visible light and near-infrared overlapping perception regions within the adjusted fan-shaped coverage area.

[0010] Optionally, extract periodic parameters based on the periodic motion law of the living obstacle and calculate its predicted range of motion trajectory, and generate a first avoidance direction adjustment amount according to the geometric relationship between the predicted range of motion trajectory and the current position of the mobile device; Calculate the angle between the projection offset rate of the geometric deformation feature and the moving direction based on the offset trend of the collision risk boundary. When the projection offset rate exceeds a preset threshold, determine the dynamic safety distance through the product of the angle and the moving speed and generate a second avoidance direction adjustment amount; Perform direction conflict detection on the first avoidance direction adjustment amount and the second avoidance direction adjustment amount. If the direction angle is less than the preset angle, superimpose and generate a composite avoidance parameter. If the direction angle is greater than or equal to the preset angle, select the direction with a smaller dynamic safety distance as the dominant avoidance direction; Generate the path planning parameters of the mobile device based on the composite avoidance parameter or the dominant avoidance direction, and in combination with the moving speed.

[0011] Optionally, extract the periodic intensity fluctuation parameters of the temperature anomaly region from the obstacle topology map, calculate the fluctuation period through the intensity change curve of the temperature anomaly region in the continuous frames, and when the fluctuation period matches the preset biological motion period threshold range, mark the corresponding temperature anomaly region as the candidate region of the living obstacle; Within the candidate region, analyze the continuity of the movement trajectory of the temperature anomaly region, and verify the biological movement law through the displacement direction and velocity fluctuation pattern of the centroid of the temperature anomaly region in consecutive frames. When the random change rate of the displacement direction exceeds a preset threshold and the fluctuation amplitude of the velocity conforms to the biological acceleration characteristics, confirm that the candidate region is a living obstacle; Synchronously based on the projection offset of the geometric deformation feature in consecutive frames, calculate the included angle relationship between the deformation contour movement trajectory of the static obstacle and the movement direction of the mobile device. When the included angle between the projection offset direction and the device movement direction is less than a preset acute angle threshold, mark the region corresponding to the deformation contour as the collision risk boundary of the static obstacle.

[0012] Optionally, generate the expected steering angle of the mobile device according to the included angle between the avoidance direction in the composite avoidance parameter or the dominant avoidance direction and the current movement direction, and at the same time adjust the movement speed based on the speed decay coefficient in the composite avoidance parameter to generate the target movement speed in the avoidance state; Associate the expected steering angle with the direction speed component of the target movement speed, and combine to generate the steering curvature parameter; Fuse the steering curvature parameter with the movement speed to generate the path planning parameter including the speed adjustment instruction and the steering control instruction; The calculating the compensation coefficient of the device movement direction and the viewing angle jitter amplitude based on the movement trajectory data includes: Decompose the device movement direction into a horizontal component and a vertical component based on the movement trajectory data, generate a horizontal jitter compensation factor according to the movement speed projection value and the direction angle deviation of the horizontal component, and generate a vertical jitter compensation factor according to the movement speed projection value and the direction angle deviation of the vertical component; Generate the compensation coefficient through the non-linear fusion of the horizontal jitter compensation factor and the vertical jitter compensation factor.

[0013] In a second aspect, the present application provides an optimized system for room obstacle detection based on a cross-modal attention mechanism, including: An acquisition module, which acquires the movement trajectory data of the mobile device, calculates the compensation coefficient of the device movement direction and the viewing angle jitter amplitude based on the movement trajectory data, and generates a visual coordinate system matching the physical space; A collection module, which activates the cross-modal collection mode of the multi-spectral sensor under the visual coordinate system, and divides the visible light and near-infrared dual-band overlapping sensing regions according to the mobile forward region corresponding to the mobile device. The region covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection, and the region covered by the near-infrared band is used to detect the temperature anomaly region related to the biological signs; An adjustment module inputs the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the movement speed and direction change amount of the mobile device through the cross-modal attention network, so as to generate an obstacle topology map; An extraction module extracts, from the obstacle topology map, a connected region that conforms to the periodic biological movement law in the temperature anomaly region as a living obstacle, and determines a collision risk boundary of the static obstacle according to the projection offset amount of the geometric deformation feature between consecutive frames; A generation module generates path planning parameters of the mobile device based on the periodic movement law of the living obstacle and the offset trend of the collision risk boundary.

[0014] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an optimization method for room obstacle detection based on a cross-modal attention mechanism as described in the first aspect above.

[0015] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it implements an optimization method for room obstacle detection based on a cross-modal attention mechanism as described in the first aspect.

[0016] In an embodiment of the present application, motion trajectory data of a mobile device is obtained, a compensation coefficient of the movement direction and the view jitter amplitude of the device is calculated based on the motion trajectory data, and a visual coordinate system matching the physical space is generated; in the visual coordinate system, a cross-modal acquisition mode of a multispectral sensor is activated, and a visible light and near-infrared dual-band overlapping perception region is divided according to the forward movement region corresponding to the mobile device, wherein the region covered by the visible light band is used to detect the geometric deformation feature of the ground projection, and the region covered by the near-infrared band is used to detect the temperature anomaly region related to biological signs; the compensation coefficient, the geometric deformation feature, and the temperature anomaly region are input into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the movement speed and direction change amount of the mobile device through the cross-modal attention network, so as to generate an obstacle topology map; in the obstacle topology map, a connected region that conforms to the periodic biological movement law in the temperature anomaly region is extracted as a living obstacle, and a collision risk boundary of the static obstacle is determined according to the projection offset amount of the geometric deformation feature between consecutive frames; path planning parameters of the mobile device are generated based on the periodic movement law of the living obstacle and the offset trend of the collision risk boundary.

[0017] The embodiments of the present application have the following beneficial effects: By fusing the motion trajectory compensation of the mobile device and multi-spectral cross-modal perception, a visual coordinate system matching the physical space is dynamically constructed. The geometric deformation features (such as shape and contour) of the ground projection objects are detected in the visible light band, and the abnormal body temperature of biological signs is captured in the near-infrared band to achieve multi-dimensional perception of environmental obstacles. Further combined with the cross-modal attention network, the weights of visible light and near-infrared are adaptively allocated according to the device movement state to generate a topological map containing static and living obstacles, in which living targets (such as humans and animals) are accurately identified through periodic biological movement laws, and the collision risk boundary of static obstacles is estimated by the continuous frame projection offset of geometric deformation features. Finally, by synthesizing the dynamic behaviors and risk trends of the two types of obstacles, real-time and safe path planning parameters are generated, significantly improving the ability of the mobile device to distinguish between biological and static obstacles, the dynamic obstacle avoidance response speed, and the environmental adaptability in complex scenarios, and being applicable to application scenarios such as service robots and autonomous driving that require both safety and efficiency.

[0018] Furthermore, by dynamically allocating parameters to adaptively adjust the weights of visible light and near-infrared, the real-time performance of static obstacle detection in high-speed moving scenarios (relying on geometric deformation) and the accuracy of living body recognition in low-speed scenarios (relying on temperature anomalies) are improved; the deformation-thermal correlation marking in the spatial dimension enhances the ability to distinguish the mixed area of biological and static obstacles (such as separating the overlapping features when a human body approaches an obstacle), and the dynamic displacement chain in the time dimension combines with the displacement correction factor to eliminate the interference of device jitter, accurately estimating the motion trajectory and risk boundary of the obstacle; finally, the generated topological map significantly reduces the misjudgment rate (such as high-temperature non-biological interference and missed detection of static obstacles) by fusing spatio-temporal two-dimensional information, providing a high-precision and robust environmental perception basis for obstacle avoidance decisions in complex dynamic scenarios.

[0019] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 The flowchart of an optimization method for room obstacle detection based on a cross-modal attention mechanism provided by the present application is shown; Figure 2Shows a schematic structural diagram of an optimized system for room obstacle detection based on a cross-modal attention mechanism provided by this application; Figure 3 Shows a schematic structural diagram of a computing device provided by this application. Detailed implementation manners

[0022] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application.

[0023] In some processes described in the specification, claims and the above-mentioned drawings of this application, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish each different operation, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0024] Researchers have found that the existing obstacle detection methods for mobile devices have significant defects in complex scenarios: traditional single-modal sensors (such as pure vision or pure infrared) are difficult to take into account both the geometric features of static obstacles and the biological signs of living targets, resulting in missed detections or false detections; the perspective jitter and movement trajectory deviation during device movement are likely to cause coordinate system misalignment, reducing the accuracy of multi-modal data fusion; the ability to distinguish the attributes of obstacles (such as living / static) in real time and predict risks in a dynamic environment is insufficient, affecting the reliability of path planning. Based on this, an optimized method for room obstacle detection based on a cross-modal attention mechanism is provided. This method can construct a high-precision visual coordinate system through motion trajectory compensation and multi-spectral collaborative perception, dynamically allocate the modal weights of visible light and near-infrared in combination with the device motion state, and generate an obstacle topology map that combines biological signs and collision risks using spatio-temporal two-dimensional features (geometric deformation contours and temperature anomaly regions), and finally realize the periodic motion tracking of living obstacles and the real-time obstacle avoidance decision-making of static obstacles.

[0025] The technical solution of this application is applicable to scenarios such as indoor navigation of service robots, complex road condition perception of autonomous driving vehicles, and human body dynamic detection in security monitoring. It is especially suitable for application environments where there are coexisting biological and static obstacles, frequent device movement jitters, and drastic changes in environmental light or temperature.

[0026] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0027] Figure 1 The following is a flowchart of an optimized method for room obstacle detection based on a cross-modal attention mechanism provided by an embodiment of the present application. As Figure 1 shown, the method includes: 101. Obtain the motion trajectory data of the mobile device, calculate the compensation coefficient of the device movement direction and the viewing angle jitter amplitude based on the motion trajectory data, and generate a visual coordinate system that matches the physical space; In this step, the motion trajectory data refers to a continuous spatio-temporal coordinate sequence collected by the accelerometer, gyroscope, and GPS sensor of the mobile device. For example, 10 groups per second include position , speed and acceleration data, which is used to describe the device movement direction and the viewing angle jitter amplitude. The device movement direction includes the azimuth deviation , and the viewing angle jitter amplitude includes the pitch angle fluctuation . The compensation coefficient is a dynamic correction parameter calculated based on the trajectory data. For example, the horizontal compensation , the vertical compensation , which is generated after fusing the sensor noise through Kalman filtering, where the noise variance , which is used to offset the influence of device jitter on the visual coordinate system. The visual coordinate system is a reference system that maps the physical space to the device local coordinates. Its origin is the device centroid, and the Z-axis points to the gravity direction . The matching accuracy with the physical space is controlled within the error range meters through coordinate transformation. The coordinate transformation is implemented using the Euler angle rotation matrix: .

[0028] In the embodiments of the present application, first, the motion trajectory data is collected through the inertial measurement unit and GPS module of the mobile device, including acceleration angular velocity and displacement increment and other parameters. The Kalman filtering algorithm is used to smooth the noise data, and the device movement direction and the viewing angle jitter amplitude are calculated based on the kinematic model. The device movement direction is obtained through the heading angle calculation formula, where the heading angle:

[0029] The perspective jitter amplitude is integrated by angular velocity to obtain the pitch angle fluctuation:

[0030] Secondly, the device attitude data is converted into a visual coordinate system in physical space through a coordinate transformation algorithm, where the compensation coefficient:

[0031] is used to offset the perspective shift caused by jitter. Finally, the compensated visual coordinate system is written into memory, and the origin of this coordinate system is the current position of the device The X-axis and Y-axis are respectively aligned with the north-south and east-west directions of the physical space for use in calibrating the multispectral sensor.

[0032] In the autonomous navigation system of a certain intelligent floor cleaning robot, the acceleration of the device is collected in real time through a built-in nine-axis inertial measurement unit sensor angular velocity and magnetometer data B to construct a three-dimensional motion trajectory model. When the robot moves on an uneven ground, an improved Kalman filtering algorithm is used to denoise the original data, and the filtering equation is:

[0033] And combined with the quaternion attitude solution method: , the instantaneous offsets of the pitch angle and roll angle of the device are calculated. And through the dynamic weighted fusion algorithm: , the offset is converted into a compensation coefficient for perspective jitter, and a visual coordinate system with the centroid of the device as the origin is established based on this. This coordinate system is rigidly registered with the physical space grid map through SLAM point cloud data, and the registration error satisfies: , achieving centimeter-level spatial alignment.

[0034] 102. Activate the cross-modal acquisition mode of the multispectral sensor under the visual coordinate system, and divide the visible light and near-infrared dual-band overlapping perception area according to the moving forward area corresponding to the mobile device. The area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection objects, and the area covered by the near-infrared band is used to detect the temperature anomaly area related to biological signs; In this step, the cross-modal acquisition mode of the multispectral sensor is a cooperative working mode that synchronously activates the visible light sensor and the near-infrared sensor. The working wavelength of the visible light sensor is 400 - 700 nm, the working wavelength of the near-infrared sensor is 700 - 1000 nm, and the sampling interval is 33 milliseconds. The forward moving area is a fan-shaped detection area centered on the device's moving direction, with a horizontal opening angle of 60 degrees and a depth of 10 meters. The geometric deformation features of the ground projection objects are extracted from the visible light band coverage area through the Canny edge detection algorithm, such as the trapezoidal distortion caused by the inclination of a rectangular object. The calculation formula for the geometric deformation feature is that the distortion rate is equal to the Δ aspect ratio divided by the original aspect ratio. The temperature anomaly area is identified from the near-infrared band coverage area through the thermal radiation intensity threshold, and the threshold is set to be greater than or equal to 35 degrees Celsius or the temperature difference from the environment is greater than or equal to 3 degrees Celsius. The temperature anomaly area is the area marked as potential biological signs, with a confidence level of not less than 85%.

[0035] In the embodiment of the present application, first, based on the visual coordinate system generated in step 101, the synchronous acquisition mode of the multispectral sensor is activated. The working wavelength of the visible light camera is 400 - 700 nm, the working wavelength of the near-infrared sensor is 800 - 1400 nm, and the two are spatially aligned in the forward fan-shaped area along the device's moving direction, with a field of view angle of 120 degrees. The forward area is divided into a visible light band coverage area and a near-infrared band coverage area through the region segmentation algorithm, where the visible light area covers the central 60-degree field of view angle, and the near-infrared area covers 30-degree field of view angles on both sides. The geometric deformation features of the ground projection objects are extracted from the visible light area through the background difference method, including parameters such as edge curvature and aspect ratio. The temperature anomaly area is used to identify biological heat sources through temperature threshold segmentation, and the temperature threshold is set to be greater than 35 degrees Celsius as abnormal. The dual-band data is aligned in the overlapping area through the spatio-temporal registration module, which is based on the feature point matching algorithm and is finally marked as the cross-modal fusion area.

[0036] For example, after completing the coordinate system calibration, the sweeping robot activates the multispectral fusion camera mounted on the top and switches to the cross-modal synchronous acquisition mode. According to the current moving speed v of the robot equal to 0.3 m / s and the direction angle θ, the forward 60-degree field of view angle is divided into a visible light and near-infrared dual-band overlapping perception area. The working wavelength of the visible light is 400 - 700 nm, and the ground texture is captured at a frame rate of 30 fps. The geometric deformation of projections such as carpet edges and wires is detected through SIFT feature matching, and the deformation threshold is set to 2 mm. The working wavelength of the near-infrared is 850 - 1700 nm, and an uncooled microbolometer is used to scan the temperature distribution within 1.5 meters in front with an accuracy of thermal sensitivity of 0.05 degrees Celsius, and the biological heat source area that is 0.5 degrees Celsius higher than the ambient temperature, such as the feet of a pet or a human, is identified.

[0037] 103. Input the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into the cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, and generate an obstacle topology map; In this step, the cross-modal attention network is a dual-modal neural network composed of a visible light branch ResNet-18 and a near-infrared branch MobileNet-V2. The moving speed and direction change amount are obtained by calculating the first derivative of the trajectory data. For example, the speed v = 1.2 m / s and the direction change rate ω = 0.1 rad / s. The attention weight distribution dynamically adjusts the contribution ratio of visible light and near-infrared according to the moving speed. When the speed v > 2 m / s, the visible light weight = 0.8. The obstacle topology map is a grid map annotating the position, type, and safety distance of obstacles, and the resolution δ = 0.05 m / grid.

[0038] In the embodiment of the present application, first, input the compensation coefficient calculated in step 101 , the geometric deformation feature G of the ground projection object in the visible light band extracted in step 102, and the temperature anomaly region T detected in the near-infrared band into the cross-modal attention network. Through the multi-head attention mechanism, this network maps the geometric deformation feature and the temperature anomaly region into high-dimensional feature vectors and . According to the real-time speed of the mobile device and the direction change amount , dynamically calculate the attention weight distribution ratios and of the two modalities. The specific implementation is as follows: When the speed increases or the direction changes violently, that is, when the condition is satisfied, the network increases the weight of the near-infrared modality, which is increased from the reference value of 0.4 to 0.7, and preferentially responds to living obstacles; when moving stably at a low speed, that is, the weight of the visible light modality

[0039] dominates, strengthening the detection of static obstacles.

[0040] where is the sigmoid function, and are the speed and direction change sensitivity coefficients, is the bias term.

[0041] Finally, the network outputs an obstacle topology map that fuses the dual-modal features , where the nodes in the figure represent the positions of obstacles, and the edge weights represent the collision risk probability between obstacles. For dynamic obstacles, the velocity vector and the predicted trajectory are additionally marked. The compensation coefficient , the geometric deformation feature vector G, and the coordinates T of the temperature anomaly region are input into the pre-trained cross-modal attention network . This network contains two layers of bidirectional LSTM modules. When the robot's acceleration is detected, the weight of the near-infrared modality is dynamically increased to 0.7; when the robot is in a uniform motion state, the weight of the visible light modality is restored to 0.6. The network outputs an obstacle probability map , which is processed by morphological closing operation to generate a topological grid map containing the height and heat source intensity attributes, where dynamic obstacles are marked as red grids and static obstacles are marked as blue grids.

[0042] 104. In the obstacle topology map, extract the connected regions that conform to the periodic biological motion law in the temperature anomaly region as living obstacles, and determine the collision risk boundary of the static obstacle according to the projection offset of the geometric deformation feature between consecutive frames; In this step, the periodic biological motion law is the living motion feature extracted by near-infrared data spectrum analysis, where the fast Fourier transform is used to detect the characteristic frequency, and the typical step frequency range is The connected region is defined as the smallest circumscribed rectangle composed of continuous pixels in the temperature anomaly region, and it needs to satisfy the area and the aspect ratio conditions. The projection offset is calculated by the optical flow method, which represents the displacement of the static obstacle in the continuous visual coordinate system, and the calculation formula is . Where represents the coordinate of the feature point at time . The collision risk boundary is dynamically adjusted according to the projection offset, and its inflation radius The calculation formula of The safety factor

[0043] In the embodiment of the present application, first, 8-neighborhood clustering analysis is performed on the temperature anomaly region, and the motion periodicity feature is extracted by discrete Fourier transform. For the detected spectral peak , when it satisfies , it is determined as a biological motion feature. At the same time, an area threshold Perform connected component filtering. For geometric deformation features, a feature point projection calculation method based on the homography matrix is adopted. The calculation formula for the displacement is . When exceeds the threshold , it is determined as a static obstacle, and its collision risk boundary is generated through the convex hull algorithm. The boundary dilation amount adopts an adaptive algorithm, and the calculation formula is .

[0044] In specific implementation, multi-frame analysis is performed on the temperature anomaly area: the periodic temperature fluctuation area that persists is extracted through background difference method, and its frequency characteristic needs to satisfy to correspond to the biological signs. Combining the optical flow method to calculate the included angle between the regional movement vector and the robot's motion trajectory. When , it is determined as a living obstacle approaching positively, and a three-level avoidance strategy is triggered. For static obstacles, the Lucas-Kanade dense optical flow algorithm is used to calculate the projection offset of geometric feature points within 5 consecutive frames. When satisfies the radial distribution condition in the polar coordinate system: , it is determined that its collision risk boundary is a sector area centered on the obstacle with a radius , where is the radial-tangential displacement ratio threshold.

[0045] 105. Generate the path planning parameters of the mobile device based on the periodic motion law of the living obstacle and the offset trend of the collision risk boundary.

[0046] In this step, the periodic motion law of the living obstacle is used to predict its future position, such as the Kalman filter prediction error of ±0.1 meter. The offset trend of the collision risk boundary is a parameter obtained by linearly regressing the boundary expansion rate, such as increasing by 0.05 meter per second. The path planning parameters include the obstacle avoidance priority (living weight 0.9), the path curvature limit (radius ≥ 0.8 meter), and the speed adjustment coefficient (the speed decays to 50% when approaching the risk boundary), generating a safe path, and the path length increases by no more than 20%.

[0047] In the embodiments of the present application, first, according to the periodic motion law of the living obstacle extracted in step 104 (such as the pedestrian walking frequency of 1.2 Hz and the moving direction due north) and the offset trend of the collision risk boundary of the static obstacle (such as the boundary expanding westward at a rate of 0.1 m / s), the dynamic window algorithm is used to calculate the feasible path of the mobile device. Specifically, the future trajectory of the living obstacle is predicted by Kalman filtering to generate a dynamic avoidance area centered on the living body (radius 1.5 m); at the same time, according to the offset trend of the static boundary, a buffer zone with real-time expansion and contraction is generated around the obstacle (such as width 0.5 m). After fusing the constraints of the dynamic avoidance area and the buffer zone, based on the kinematic parameters of the mobile device (maximum steering angle of 30 degrees, acceleration upper limit of 1 m / s²), path planning parameters are generated, and the parameters synchronously control the exposure duration ratio of the visible light and near-infrared bands in the forward area of the mobile device before movement, and trigger the cross-modal attention network to perform incremental iterative updates on the trajectory prediction module of the connected domain of the living obstacle. It includes the target heading angle (such as adjusted to 25 degrees east of northeast), the speed curve (such as decreasing from 2 m / s to 0.8 m / s), and the steering angle sequence, and is sent to the drive controller for execution through the serial communication protocol.

[0048] Based on the motion period function of the living obstacle , where Determined by the heat source fluctuation frequency and the risk boundary gradient data of the static obstacle, the path planning module adopts an improved dynamic window algorithm (DWA). A time-varying repulsive force field is applied to the living obstacle, and the repulsive force coefficient decays with the inverse square of the distance; for the static obstacle, a direction constraint term is introduced to limit the feasible speed of the robot in the tangent direction of the risk boundary inside. Finally, the output contains the maximum safe speed , the steering angle and the emergency braking probability of the planning parameters, and the hub motor is controlled through the ROS system to achieve an S-shaped obstacle avoidance trajectory, and the measured obstacle avoidance success rate is improved.

[0049] In summary, through steps 101 to 105, the mobile device realizes high-precision autonomous obstacle avoidance and navigation in a complex environment. By fusing visible light and near-infrared multi-modal perception data, the system can simultaneously identify the geometric deformation characteristics of static obstacles and the biological sign signals of dynamic living bodies, and construct an obstacle topology map containing spatio-temporal information. The innovative cross-modal attention mechanism can dynamically adjust the perception weight according to the motion state of the device to ensure stable environmental modeling ability even during fast movement or perspective jitter. The finally generated path planning parameters comprehensively consider the periodic motion law of the living obstacle and the collision risk boundary of the static obstacle, enabling the device to have predictive obstacle avoidance ability and significantly improving the navigation reliability in challenging scenarios such as low light and complex backgrounds.

[0050] To solve the problem of cross-modal data fusion for obstacle detection in complex environments by mobile devices, in some embodiments, inputting the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into a cross-modal attention network in step 103 to adjust the attention weight distribution between the visible light and near-infrared modalities according to the movement speed and direction change amount of the mobile device through the cross-modal attention network to generate an obstacle topology map includes: 201. Input the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into a cross-modal attention network, and generate a dynamic allocation parameter through the cross-modal attention network based on the movement speed and direction change amount of the mobile device; In step 201, the compensation coefficient refers to a dynamic correction parameter calculated based on the device movement trajectory (for example, horizontal compensation 0.7, vertical compensation 0.5), which is used to offset the influence of perspective jitter on the visual coordinate system. The geometric deformation feature is the ground object shape distortion parameter extracted from visible light data (such as trapezoidal distortion rate = 0.15). The temperature anomaly region is the region where the thermal radiation intensity detected by the near-infrared sensor exceeds the threshold (such as temperature difference ≥ 3°C). The cross-modal attention network is a dual-modal neural network composed of a visible light branch (ResNet-18) and a near-infrared branch (MobileNet-V2). The movement speed and direction change amount of the mobile device are parameters calculated by the first derivative of the trajectory data (for example, speed = 1.5 m / s, direction change rate = 0.2 rad / s). The dynamic allocation parameter is the weight adjustment parameter generated by the cross-modal attention network according to the speed and direction change (for example, when the speed > 1 m / s, the visible light weight increases by 5%).

[0051] In the embodiments of the present application, first, input the compensation coefficient (such as K = 0.85) calculated in step 101, the geometric deformation features (such as edge curvature, aspect ratio) of the visible light band extracted in step 102, and the temperature anomaly region (such as heat source contour) of the near-infrared band into the cross-modal attention network. This network adopts a multi-head attention mechanism, and respectively performs high-dimensional encoding on the geometric deformation feature and the temperature anomaly region through a convolutional neural network (CNN) to generate feature vectors of the visible light and near-infrared modalities; subsequently, based on the real-time movement speed (such as 2 m / s) and direction change amount (such as deflecting 15 degrees per second) of the mobile device, calculate the initial attention scores of the two modalities through a fully connected layer, and use the compensation coefficient as a weight factor to dynamically weight and adjust the scores, and finally output the dynamic allocation parameters of the visible light and near-infrared modalities (such as visible light weight 0.7, near-infrared weight 0.3). This process realizes the adaptive allocation of modal weights through three-stage processing of feature embedding, dynamic parameter calculation, and compensation weighting.

[0052] 202. Calculate the weight allocation ratio of the visible light data stream and the near-infrared data stream based on the dynamic allocation parameter, and perform weighted fusion on the geometric deformation feature and the temperature anomaly region according to the weight allocation ratio; In step 202, the dynamic allocation parameter is an intermediate variable that controls the fusion ratio of visible light and near-infrared data (such as visible light weight 0.7, near-infrared weight 0.3). The visible light data stream is a continuous image sequence collected by a visible light sensor (resolution 1920×1080, 30fps). The near-infrared data stream is a thermal radiation intensity matrix output by a near-infrared sensor (resolution 640×480, 30fps). The weight allocation ratio is the contribution ratio of the visible light and near-infrared modalities determined according to the dynamic allocation parameter (such as 7:3). Weighted fusion is the process of superimposing the geometric deformation feature (visible light modality) and the temperature anomaly region (near-infrared modality) according to the weight ratio (for example, the fusion formula: 0.7×geometric feature + 0.3×temperature feature).

[0053] In the embodiment of the present application, first, based on the dynamic allocation parameter (such as visible light weight 0.7, near-infrared weight 0.3), the weight allocation ratio (visible light ratio ≈ 0.67, near-infrared ratio ≈ 0.33) is calculated through Softmax function normalization. Subsequently, spatial alignment processing is performed on the visible light and near-infrared data: both are mapped to the same visual coordinate system by affine transformation to eliminate the perspective shift caused by device movement; then, the geometric deformation feature (such as the edge gradient matrix) and the temperature anomaly region (such as the heat source probability map) are weighted and superimposed according to the weight ratio to generate a fused feature map (such as visible light feature × 0.67 + near-infrared feature × 0.33), and Gaussian filtering is used to suppress noise in the fusion result to eliminate discrete anomaly points. This process realizes the effective integration of cross-modal data through weight allocation, spatial alignment, and feature fusion.

[0054] 203. Perform cross-modal interaction on the weighted geometric deformation feature and the temperature anomaly region, generate a deformation heat correlation mark by overlapping and comparing the geometric deformation contour and the temperature anomaly region in the spatial dimension, and construct a dynamic displacement chain by combining the relative offset of the deformation contour and the temperature anomaly region in consecutive frames with the displacement correction factor of the compensation coefficient in the time dimension; In step 203, cross-modal interaction refers to the joint analysis of visible light and near-infrared data in the spatial and temporal dimensions. The geometric deformation contour is the contour of the ground object extracted by edge detection (such as the sequence of polygon vertex coordinates). The temperature anomaly region is the high-temperature connected domain calibrated in the thermal image (area ≥ 0.5 m²). The deformation-thermal correlation marker is the marker of the overlap between the geometric contour and the temperature region in the spatial dimension (for example, the region with a contour overlap rate > 80% is marked as highly correlated). The continuous frames refer to the adjacent visual coordinate system data in the time series (interval 33 ms). The relative offset is the displacement difference between the deformation contour and the temperature region in the continuous frames (such as a lateral offset of 0.2 meters). The displacement correction factor is a dynamic correction parameter calculated according to the compensation coefficient (for example, the horizontal displacement compensation factor is 0.8). The dynamic displacement chain is the displacement trajectory constructed by the offsets and correction factors of multiple frames in the time dimension (for example, the trajectory sequence: [0.1 m, 0.15 m, 0.2 m]).

[0055] In the embodiment of the present application, first in the spatial dimension, the sliding window algorithm is used to scan the fused feature map to locate the overlapping region between the geometric deformation contour and the temperature anomaly region (such as the overlapping area ratio > 50%), and assign a deformation-thermal correlation marker to the overlapping region (the binary mask is marked as 1, and the non-overlapping region is 0). In the time dimension, the optical flow method (Lucas-Kanade algorithm) is used to track the displacement vectors of the deformation contour and the temperature anomaly region in the continuous frames (such as ), and the displacement is corrected by combining the compensation coefficient (such as K = 0.9) ( ), and finally the corrected displacement vectors are stored in time series to construct a dynamic displacement chain (such as [(0.18, 0.09), (0.20, 0.10)]). Through the joint analysis of the spatial overlap marker and the time displacement chain, the spatio-temporal consistency modeling of cross-modal interaction is realized.

[0056] 204. Generate an obstacle topology map that fuses biometric features and deformations based on the dynamic displacement chain and the deformation-thermal correlation marker.

[0057] In step 204, the dynamic displacement chain is a trajectory chain that describes the movement trend of the obstacle in the time dimension (such as a displacement increment of 0.1 meter per second). The deformation-thermal correlation marker is the joint feature identifier of geometric deformation and temperature anomaly in the spatial dimension (for example, the confidence level of the overlapping region ≥ 90%). The biometric feature refers to the living body attribute identified by temperature anomaly and periodic motion law (such as the human walking frequency of 1.2 Hz). The obstacle topology map is a grid map that fuses deformation, temperature, and motion features (resolution 0.1 meter / grid), marking the positions and risk levels of static obstacles (gray) and living body obstacles (red).

[0058] In the embodiments of the present application, first, the center point of the region marked with deformation heat correlation is used as a topological node (such as coordinates (x = 3.5, y = 2.1)), and the motion correlation between nodes (such as cosine similarity) is calculated based on the dynamic displacement chain, and this is used as the edge weight (such as similarity 0.8 corresponding to weight 0.2). Subsequently, a topological model of nodes and edges is built through a graph neural network (GNN), and a biometric feature (such as temperature value 42°C) and a deformation feature (such as curvature 0.6) are bound to each node, and finally an obstacle topological map is generated. The nodes in the map represent the positions and attributes of obstacles, and the edge weights represent the intensity of the collision risk correlation (in the range of 0 - 1). Through node generation, motion correlation analysis, and attribute binding, this process realizes the structured expression of multi-dimensional obstacle information.

[0059] In summary, in steps 201 to 204, multi-modal intelligent perception and dynamic obstacle modeling of the mobile device in a complex environment are realized. Through the adaptive weight allocation mechanism of the cross-modal attention network, the system can optimize the fusion strategy of visible light and near-infrared data in real time according to the device motion state, ensuring the best environmental perception effect in different motion scenarios. The innovative deformation heat correlation marking technology realizes the precise spatial alignment of geometric features and biometric features, and the construction of the dynamic displacement chain effectively integrates the motion trajectory information in the time dimension. The finally generated obstacle topological map not only contains the accurate geometric contours of static obstacles but also completely retains the motion characteristics of dynamic biological targets, providing the mobile device with environmental cognition ability with both spatial accuracy and temporal continuity.

[0060] To solve problems such as multi-modal data conflict, motion interference, and lack of biometric features existing in traditional obstacle detection methods in complex dynamic scenarios, in some embodiments, generating an obstacle dynamic topological map that fuses biometric features and deformations in step 204 includes: 301. When the direction deviation between the displacement direction of the geometric deformation contour in the dynamic displacement chain and the migration direction of the temperature anomaly region is less than a preset threshold, extract synchronous offset trajectory data from the dynamic displacement chain to construct a dynamic trajectory layer of the living obstacle; In step 301, the dynamic displacement chain is a sequence of obstacle motion trajectories constructed by multi-frame displacement correction factors, such as a displacement increment of 0.1 meters per second. The geometric deformation contour is the edge contour of ground objects extracted from visible light data, such as a rectangular distortion contour composed of a set of polygon vertex coordinates. The migration direction of the temperature anomaly region is the moving direction of the thermal radiation region in consecutive frames, such as a moving path with an azimuth angle of 30°. The direction deviation refers to the angular difference between the displacement direction of the geometric deformation contour and the migration direction of the temperature region. Specifically, when the deviation angle ≤ 10 degrees, a threshold judgment is triggered. The preset threshold is a critical parameter for judging direction consistency. For example, the direction deviation angle threshold is 15 degrees. The synchronous offset trajectory data is a set of displacement trajectories extracted when the direction deviation is below the threshold, such as a trajectory chain containing the coordinates [0.1m, 0.2m, …, 1.0m] for 10 consecutive frames. A living obstacle is a dynamic object with periodic biological signs, such as a moving target with a pedestrian step frequency feature of 1.2 Hz. The dynamic trajectory layer is a rasterized map layer that marks the movement path of living obstacles, with a raster granularity of 0.1 meters per pixel.

[0061] In the embodiment of the present application, first, the direction deviation between the displacement direction of the geometric deformation contour (such as vector ΔV = (0.2, 0.1)) and the migration direction of the temperature anomaly region (such as vector ΔU = (0.18, 0.09)) in the dynamic displacement chain is calculated by the vector angle cosine algorithm. The formula is: direction deviation = 1 - (ΔV·ΔU) / (||ΔV||×||ΔU||). When the direction deviation is less than the preset threshold (such as 0.1), it is determined that the two motions are synchronized. Subsequently, the time series clustering algorithm (such as DBSCAN) is used to extract the synchronous offset trajectory data from the dynamic displacement chain, specifically including the trajectory starting point, the moving direction angle, and the average speed. Finally, a smooth dynamic trajectory layer of living obstacles is generated by B-spline curve fitting, and the trajectory data is stored in the spatio-temporal database according to the time stamp.

[0062] 302. Through the spatial superposition relationship between the geometric deformation contour and the temperature anomaly region, when the projection boundary of the geometric deformation contour coincides with the outer edge boundary of the temperature anomaly region, it is marked as the spatial distribution layer of static obstacles; In step 302, the spatial superposition relationship refers to the overlapping degree between the geometric deformation contour and the temperature anomaly region in the visual coordinate system, such as 80% area coincidence in the horizontal direction. The projection boundary of the geometric deformation contour is the set of two-dimensional projection outer contour vertices of ground objects in the visible light band, such as the four corner point coordinates of a rectangular object. The outer edge boundary of the temperature anomaly region refers to the contour polygon of the high-temperature region in thermal imaging, such as the side length data of an irregular shape generated by the thermal radiation intensity threshold. A static obstacle is an object without biological characteristics and with a fixed position, such as a roadblock stone with a diameter of 0.5 meters. The spatial distribution layer is a raster layer that marks the position and contour of static obstacles, and the raster is filled with gray coding.

[0063] In the embodiment of the present application, first, the intersection over union (IoU) algorithm is used to calculate the overlap degree between the projected boundary of the geometric deformation contour (such as the rectangular box coordinates (x1, y1, x2, y2)) and the outer boundary of the temperature anomaly region (such as the polygon vertex set). When the IoU value is greater than the threshold (such as 0.7), it is determined that the two coincide in space, and the coincidence region is marked as the spatial distribution layer of the static obstacle through rasterization processing. Specifically, the morphological dilation algorithm is used to expand the boundary of the coincidence region by 0.2 m to generate a safety buffer zone, and the spatial distribution layer is stored in a binary matrix (the obstacle area is 1, and the idle area is 0).

[0064] 303. Based on the fracture characteristics of the geometric deformation contour and the periodic intensity fluctuation of the temperature anomaly region in the deformation heat correlation mark, when the fracture position matches the phase of the fluctuation period, the correlation response layer of the potential interaction behavior is generated through the coverage relationship between the extended region of the fracture boundary and the thermal fluctuation region; In step 303, the deformation heat correlation mark is the overlapping identification of the geometric contour and the temperature anomaly region in space. For example, a red semi-transparent overlay layer is set in the overlapping region. The fracture characteristic is the discontinuous edge fracture in the geometric deformation contour caused by occlusion or damage, such as a contour notch with a length of 0.3 m. The periodic intensity fluctuation refers to the regular change of the thermal radiation intensity in the temperature region. The fracture position is the set of physical coordinates of the contour fracture points. The fluctuation period phase is the time stage mark of the thermal intensity fluctuation. For example, when t = 0.3T is divided within the period T, it is the rising edge stage. The extended region of the fracture boundary is a circular buffer zone with a radius of 0.5 m centered on the fracture point, which is used to detect the potential interaction region. The thermal fluctuation region is a sub-region with periodic changes in the temperature anomaly region, such as the core region of thermal radiation change with a radius of 0.8 m. The coverage relationship is the determination of the position overlap ratio between the extended region and the fluctuation region. The potential interaction behavior refers to the approaching event of a living obstacle to a static obstacle during movement. The correlation response layer is a yellow grid layer marking the interaction risk region.

[0065] In the embodiment of the present application, first, the fracture characteristics (such as the fracture length > 0.5 m) of the geometric deformation contour in the deformation heat correlation mark are extracted through the Canny edge detection algorithm, and at the same time, the Fourier transform is used to analyze the intensity fluctuation period (such as the main frequency of 1.2 Hz) of the temperature anomaly region. When the time difference between the fracture position and the fluctuation period phase is less than the preset value (such as 10% of the period duration), it is determined that the two match. Subsequently, the region growing algorithm is used to generate an extended region by expanding 0.3 m from the fracture boundary, and the coverage area ratio between it and the thermal fluctuation region is calculated. If the coverage ratio exceeds the threshold (such as 60%), an interaction behavior mark (such as "suspected living body contact") is assigned to this region, and finally, the thermal distribution map of the correlation response layer is generated.

[0066] 304. Superimpose the spatial distribution layer, the dynamic trajectory layer, and the association response layer in the order of priority to form an obstacle topology map that fuses geometric deformation features and bio-thermal features.

[0067] In step 304, the spatial distribution layer is a gray grid background layer that calibrates the spatial positions of static obstacles. The dynamic trajectory layer is a superimposed layer that marks the moving paths of living bodies with red line types, and the line width is 3 pixels. The association response layer is a semi-transparent yellow layer that highlights potential interaction areas, and its coverage priority is higher than that of the static layer. The priority order defines the display weight rules for layer superimposition. Fusing geometric deformation features and bio-thermal features refers to combining the contour distortion rate of visible light (deformation rate ≥ 15%) and the body temperature fluctuation period of near-infrared (period matching error ≤ 0.1 s). The obstacle topology map is the final grid map that integrates the information of the three layers, with an output resolution of 1024 × 768 pixels, containing the coordinates, types, and risk level parameters of various types of obstacles.

[0068] In the embodiment of the present application, first, according to the preset priority rules (the spatial distribution layer is the bottom layer, the dynamic trajectory layer is the middle layer, and the association response layer is the top layer), the three-layer data is superimposed and processed through an image fusion algorithm. Specifically, the binary matrix of static obstacles in the spatial distribution layer (the obstacle area is 1, and the idle area is 0) is used as the base map, and a rasterized rendering is used to generate the bottom layer framework; subsequently, the living body movement trajectory data in the dynamic trajectory layer (such as a B-spline curve path) is superimposed on the middle layer through a semi-transparent color band (such as a 50% transparency of the blue channel) to distinguish dynamic and static obstacles; finally, the thermal distribution map in the association response layer (such as a pseudo-color mapping: red indicates high risk, and yellow indicates early warning) is covered on the top layer through an Alpha blending algorithm to highlight potential interaction behavior areas. During the fusion process, the coordinate alignment and pixel-level synthesis of multi-layer data are performed based on the OpenGL rendering engine, and finally, an obstacle topology map containing geometric deformation features (such as contour positions, broken boundaries) and bio-thermal features (such as temperature intensity, periodic fluctuations) is generated, and it supports real-time query of the attribute data of each layer through an interactive interface (such as clicking on a node to obtain the temperature value, deformation curvature, and movement trajectory history).

[0069] In summary, through steps 301 to 304, the intelligent upgrade of obstacle dynamic modeling and behavior prediction in complex environments is achieved. Through multi-modal data fusion and spatio-temporal feature analysis, the system establishes a three-dimensional obstacle representation system that includes three-dimensional spatial distribution, dynamic movement trajectories, and potential interaction behaviors. This technology significantly improves the accuracy of obstacle recognition and the depth of environmental understanding in complex scenarios, especially demonstrating excellent perception capabilities in complex scenarios where dynamic targets interact with static environments.

[0070] To solve the problems of mismatched multi-modal perception data acquisition range and insufficient dynamic target detection accuracy of mobile devices in complex environments, in some embodiments, in step 102, the cross-modal acquisition mode of the multi-spectral sensor is activated in the visual coordinate system, and the visible light and near-infrared dual-band overlapping perception areas are divided according to the forward movement area corresponding to the mobile device, where the area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection object, and the area covered by the near-infrared band is used to detect the temperature anomaly area related to biological signs, including: 401. In the visual coordinate system, determine the forward movement area according to the movement direction of the mobile device, and form a fan-shaped coverage area extending along the movement direction with the current position of the device as the vertex in the forward movement area, and adjust the radius length of the fan-shaped coverage area according to the movement speed; In step 401, the visual coordinate system refers to a local coordinate reference system with the current position of the device as the origin and the Z-axis pointing in the movement direction, such as a three-dimensional space coordinate system (error ±0.1m) updated in real time by the SLAM algorithm. The forward movement area is the detection area directly in front of the movement direction of the device, with a horizontal opening angle of 60°, and the depth range is dynamically adjusted according to the movement speed. The fan-shaped coverage area is a fan-shaped space with the device position as the vertex and extending along the movement direction, and its radius length is linearly related to the movement speed (for example, when the speed is 1m / s, the radius is 5m, and the radius increases by 2m for every 0.5m / s increase in speed). The movement speed is a real-time speed parameter calculated by fusing the accelerometer and GPS, for example, the current speed is 2.3m / s.

[0071] In the embodiment of the present application, first, based on the visual coordinate system (with the origin as the current position of the device and the X / Y axes aligned with the physical space direction) generated in step 101, the forward movement area is determined through the heading angle (such as θ = 30°) of the mobile device. A fan-shaped coverage area is expanded with the current position of the device as the vertex along the movement direction (θ ± 60°), and its initial radius is set by a preset value (such as 5m). Subsequently, according to the real-time speed of the mobile device (such as v = 2m / s), the fan radius is dynamically adjusted through a linear interpolation formula: radius R = R0×(1 + 0.2v) (R0 is the initial radius, and the coefficient 0.2 is the speed scaling factor). For example, when the speed increases to 3m / s, the radius expands to 5×(1 + 0.2×3) = 8m. This adjustment mechanism ensures that a farther area is covered during high-speed movement to sense obstacles in advance.

[0072] 402. Activate the dual-band synchronous acquisition mode of the visible light sensor and the near-infrared sensor within the adjusted fan-shaped coverage area, so that their physical acquisition ranges completely cover the fan-shaped coverage area, where the visible light sensor focuses on the continuous frame contour changes of the floor projection plane, and the near-infrared sensor focuses on the thermal intensity gradient changes between adjacent pixels; In step 402, the adjusted sector coverage area refers to a sector detection range with a radius that varies dynamically with speed (e.g., when the speed = 3 m / s, the radius = 11 m). The visible light sensor is a device for collecting images in the 400 - 700 nm band (resolution 1920×1080, frame rate 30 Hz). The near-infrared sensor is a device for collecting the thermal radiation intensity in the 700 - 1000 nm band (resolution 640×480, frame rate 30 Hz). The dual-band synchronous acquisition mode is a data synchronization mechanism for aligning the timestamps of the visible light and near-infrared sensors (synchronization error ≤ 1 ms). The physical acquisition range is the physical space area actually covered by the sensor (e.g., the visible light horizontal viewing angle is 70°, and the near-infrared viewing angle is 65°). The continuous frame contour change of the floor projection plane is a sequence of the projected shapes of ground objects extracted from the visible light image (e.g., the rectangular contour shifts by 0.2 m between two frames). The thermal intensity gradient change is the difference in thermal radiation between adjacent pixels in the near-infrared data (e.g., a gradient ≥ 3 °C / pixel is marked as a mutation point).

[0073] In the embodiment of the present application, within the sector coverage area generated in step 401, the acquisition timings of the visible light sensor (400 - 700 nm) and the near-infrared sensor (800 - 1400 nm) are synchronized through a hardware trigger signal to ensure that the physical fields of view of both completely cover the sector area. The visible light sensor collects continuous images of the floor projection plane at a frame rate of 30 fps, and extracts the contour change through the background difference algorithm (such as the edge displacement = 0.2 m); the near-infrared sensor collects thermal radiation data at the same frame rate, and calculates the thermal intensity gradient between adjacent pixels through the Sobel operator (e.g., a gradient value > 10 °C / pixel is determined as abnormal). The data of the two sensors are aligned through timestamps and mapped to the sector area grid map in the same visual coordinate system.

[0074] 403. Map the continuous frame contour change collected by the visible light sensor into geometric deformation features, calculate the position coordinates of the deformation boundary through the displacement difference of the contour boundaries of adjacent frames, and at the same time map the thermal intensity gradient change collected by the near-infrared sensor into a temperature anomaly area, and extract the boundary coordinates of the temperature anomaly area through the spatial distribution of the thermal intensity mutation points; In step 403, the continuous frame contour change refers to the shape displacement of the ground object projection in the visible light image sequence (such as the aspect ratio change rate of the trapezoidal contour between two frames is 0.1). The geometric deformation feature is the deformation parameter calculated through the contour vertex coordinates (for example, the average vertex displacement difference = 0.15 m). The position coordinates of the deformation boundary are the set of contour vertices of the geometric deformation region (such as the rectangle vertex coordinates [x1, y1], [x2, y2],...). The thermal intensity mutation point refers to the pixel point coordinates in the near-infrared data where the thermal gradient exceeds the threshold (such as ≥ 5 °C / pixel). The boundary coordinates of the temperature anomaly region are the set of polygon vertices generated by clustering the mutation points (such as the clustering radius is 0.3 m and the minimum number of points is 10).

[0075] In the embodiment of the present application, first, for the continuous frame contour data of the visible light sensor, the displacement difference of the contour boundary between adjacent frames is calculated by the optical flow method (Lucas-Kanade algorithm) (such as ), and the displacement difference is mapped to the visual coordinate system based on the homography matrix to generate the absolute position coordinates of the geometric deformation boundary (such as (x = 3.2, y = 4.5)). Secondly, for the thermal intensity gradient data of the near-infrared sensor, the region growing algorithm is used to start from the gradient mutation point (such as the gradient > 8 °C / pixel) and expand outward to the pixels with the gradient value lower than the threshold (such as 2 °C / pixel) to generate the closed polygon boundary of the temperature anomaly region, and its boundary coordinates are calculated by the centroid coordinate method (such as the vertex set [(3.0, 4.3), (3.5, 4.8)]).

[0076] 404. Based on the position coordinates and the boundary coordinates, divide the visible light and near-infrared overlapping perception regions within the adjusted sector coverage area.

[0077] In step 404, the position coordinates are the contour vertex data of the geometric deformation region in the visible light modality (such as the polygon vertex sequence). The boundary coordinates are the contour vertex data of the temperature anomaly region in the near-infrared modality (such as the outer rectangle vertices of the thermal image). The visible light and near-infrared overlapping perception region is the spatial intersection region of the coverage ranges of the two sensors (such as the overlapping area ratio ≥ 85%), and the coordinate alignment algorithm (such as ICP registration) is used to ensure the spatial consistency between modalities, with an error ≤ 0.05 m.

[0078] In the embodiments of the present application, first, the geometric deformation boundary position coordinates of the visible light sensor and the temperature anomaly region boundary coordinates of the near-infrared sensor are mapped to the grid map of the same fan-shaped coverage area, and the overlapping area between the two is calculated by the polygon intersection algorithm (such as the Weiler-Atherton algorithm). If the proportion of the intersection area in the area of any one region exceeds the threshold (such as 30%), the overlapping area is marked as the overlapping perception area of visible light and near-infrared, and a priority label is assigned (such as priority 1 is the dual-modal confirmation area). Finally, the boundary coordinate set (such as the polygon vertex sequence) of the overlapping area and its priority attribute are output for subsequent cross-modal fusion processing.

[0079] In summary, through steps 401 to 404, the intelligent dynamic optimization of the forward environmental perception area of the mobile device and the precise alignment of multi-modal data are realized. This technology breaks through the spatial alignment problem in traditional multi-sensor data fusion, enabling pixel-level position association between geometric contour changes and thermal anomaly regions, and providing an accurate spatial reference for subsequent cross-modal feature fusion. This method of dynamically adjusting the perception area and precisely aligning data significantly improves the capture ability and perception accuracy of the mobile device for environmental features in complex motion states.

[0080] To solve the problems of inaccurate prediction of the movement of living obstacles and conflict in obstacle avoidance decisions in a dynamic environment of a mobile device, in some embodiments, in step 105, generating the path planning parameters of the mobile device based on the periodic movement law of the living obstacle and the offset trend of the collision risk boundary includes: 501. Extract periodic parameters based on the periodic movement law of the living obstacle and calculate the prediction range of its movement trajectory, and generate a first avoidance direction adjustment amount according to the geometric relationship between the prediction range of the movement trajectory and the current position of the mobile device; In step 501, the periodic movement law of the living obstacle refers to the biometric parameters extracted by near-infrared data spectrum analysis, such as a step frequency of 1.2 Hz or a breathing frequency of 0.3 Hz. The periodic parameters include the movement period (such as period T = 0.83 s) and the amplitude (such as displacement amplitude ±0.5 m). The prediction range of the movement trajectory is the future position distribution area calculated based on the periodic parameters. For example, the fan-shaped area (radius 2 m, opening angle 90°) where the living body may move within the next 3 seconds is predicted by Kalman filtering. The geometric relationship is the relative position parameter between the current position of the mobile device and the prediction range. For example, the distance between the device and the boundary of the prediction area is 1.5 m and the azimuth deviation is 30°. The first avoidance direction adjustment amount is the path correction parameter generated according to the geometric relationship. For example, it deflects 15° to the left to avoid the prediction area.

[0081] In the embodiments of the present application, first, the periodic motion parameters of the living obstacle are extracted through the Fast Fourier Transform (FFT) (such as the pedestrian step frequency of 1.2 Hz and the limb swing amplitude of 0.5 m), and the motion trajectory range within the next 3 seconds is predicted based on the Kalman filtering algorithm to generate a set of trajectory points (such as ). Subsequently, the shortest distance between the current position of the mobile device and each point of the predicted trajectory is calculated through the Euclidean distance formula. If the shortest distance is less than the safety threshold (such as 1.5 m), the first avoidance direction adjustment amount is generated according to the direction of the line connecting the device position and the nearest point on the trajectory. For example, when the direction of the line connecting the device and the nearest point on the trajectory is the northeast direction, the heading angle correction amount is calculated through the arctangent function , and is mapped to the device control instruction to adjust the moving direction to avoid overlapping with the trajectory of the living obstacle.

[0082] 502. Calculate the angle between the projection offset rate of the geometric deformation feature and the moving direction based on the offset trend of the collision risk boundary. When the projection offset rate exceeds the preset threshold, determine the dynamic safety distance through the product of the angle and the moving speed and generate the second avoidance direction adjustment amount; In step 502, the offset trend of the collision risk boundary refers to the linear regression result of the static obstacle contour expansion rate. For example, the boundary expands outward by 0.1 m per second. The projection offset rate is the displacement speed of the geometric deformation feature in the visual coordinate system. For example, the contour vertex moving rate is 0.3 m / s. The angle is the angle between the projection offset rate direction and the device moving direction. For example, the offset direction forms an angle of 45° with the device forward direction. The preset threshold is the minimum rate for determining the offset risk. For example, avoidance is triggered when the rate ≥ 0.2 m / s. The dynamic safety distance is the buffer distance calculated through the product of the angle and the moving speed. For example, when the angle is 45° and the speed is 1 m / s, the safety distance = 1 m × sin(45°) ≈ 0.7 m. The second avoidance direction adjustment amount is the path correction parameter generated according to the dynamic safety distance. For example, deflect 20° to the right to maintain the safety distance.

[0083] In the embodiments of the present application, first, the projection offset rate of the geometric deformation feature (such as the collapsed wall contour) is calculated through the optical flow method (such as ), and the angle between its offset direction and the device moving direction is calculated using the vector dot product formula:

[0084] where V1 is the device moving direction vector and V2 is the projection offset direction vector. When the offset rate exceeds the preset threshold (such as 0.2 m / s), the dynamic safety distance S is calculated through the formula Calculate (where K is the empirical coefficient 0.5 and v is the device speed). If S is less than the preset safety value (such as 1 m), then generate a second avoidance direction adjustment amount. Specifically, by scaling the included angle α proportionally (such as =α / 2) to determine the course angle correction direction. For example, when α = 60°, generate a correction instruction to turn left by 30° to make the device move away from the offset path of the geometric deformation area.

[0085] 503. Perform a direction conflict detection on the first avoidance direction adjustment amount and the second avoidance direction adjustment amount. If the direction angle is less than the preset angle, then generate a composite avoidance parameter by superposition. If the direction angle is greater than or equal to the preset angle, then select the direction with a smaller dynamic safety distance as the dominant avoidance direction; In step 503, the direction conflict detection is a process of determining whether the included angle between the first and second avoidance direction adjustment amounts will cause a path conflict. For example, the included angle between the two directions is 30°. The preset angle is the critical value for determining direction conflict. For example, the included angle threshold is set to 45°. The composite avoidance parameter is the vector superposition result of the two when the direction angle is less than the threshold. For example, the first adjustment amount of turning left by 15° and the second adjustment amount of turning right by 20° are combined to be a 5° right turn. The dominant avoidance direction is the direction with lower risk selected when the direction angle is greater than or equal to the threshold. For example, if the included angle between the two directions is 60° and the safety distance of the second direction is smaller (0.5 m < 0.7 m), then the second avoidance direction is preferentially adopted.

[0086] In the embodiment of the present application, first calculate the first avoidance direction adjustment amount and the second adjustment amount of the direction vector included angle: If it satisfies: Then through weighted average: Generate a composite avoidance parameter; if it satisfies: Then compare the dynamic safety distances and , select the smaller value direction as the dominant avoidance direction. For example, when (avoiding birds) and (avoiding swinging wires), if , then preferentially select as the dominant direction.

[0087] 504. Based on the composite avoidance parameter or the dominant avoidance direction, and in combination with the moving speed, generate the path planning parameter of the mobile device.

[0088] In step 504, the composite avoidance parameter is the weighted vector sum of two adjustment amounts. For example, the synthesis direction deflects by 10° and the speed is reduced to 0.8 m / s. The dominant avoidance direction is the direct application of a single adjustment amount. For example, it is selected to turn right by 20° and maintain the speed of 1 m / s. The path planning parameters include the direction deflection angle (such as 10°), the speed decay coefficient (such as reduced to 70% of the original speed), and the path curvature limit (such as the curvature radius ≥ 1 m).

[0089] In the embodiment of the present application, first, the dynamic window algorithm (DWA) is used to fuse the avoidance direction and speed constraints: According to a candidate heading angle set is generated (such as Δθ ± 10°), and the feasible speed range is calculated in combination with the maximum acceleration of the device (such as 1 m / s²) ( ). Subsequently, a B-spline curve is used to fit the optimal combination of the heading angle and speed to generate a sequence of path points (such as ), and the optimal path is selected based on the cost function (the farther away from the obstacle and the more stable the speed, the lower the cost). Finally, the target heading angle, speed curve, and steering command are output, and are sent to the actuator through ROS (Robot Operating System) to drive the device to travel along the planned path while maintaining a dynamic safety distance from the obstacle.

[0090] The following is a specific example: In the dynamic obstacle avoidance scenario of a domestic service robot, the system realizes intelligent avoidance decision-making through multi-modal perception: When it detects that a pet dog (a living obstacle) swings left and right periodically at a speed of 0.8 m / s (the period is 1.2 s ± 0.15 s), its motion law parameters are extracted and the trajectory prediction range for the next 3 seconds is calculated (the fan-shaped area ± 45°), and the first avoidance direction adjustment amount is generated (deflecting 25° to the right); at the same time, it is monitored that the left sofa leg (a static obstacle) has a geometric deformation feature that offsets outward at a rate of 0.3 m / s due to the risk of collision by the robot's side brush (the angle with the current moving direction is 50°), triggering the second avoidance direction adjustment amount (deflecting 15° to the left). After direction conflict detection, it is found that the included angle between the two adjustment amounts is 40° < the preset threshold of 45°, so they are superimposed to generate a composite avoidance parameter (the synthesis direction deflects 10° to the right), and the safety distance (0.5 m) is dynamically calculated in combination with the current moving speed of 0.6 m / s. Finally, the path planning parameters including decelerating to 0.4 m / s and correcting the heading by 10° are output. This solution enables the robot to avoid living pets (predicted trajectory matching degree) and dynamic deformation obstacles (anti-collision success rate) synchronously within 1.2 seconds, reducing path detours compared to a single avoidance strategy.

[0091] In summary, through steps 501 to 504, intelligent obstacle avoidance and adaptive path planning of the mobile device in a complex dynamic environment are achieved. By integrating the prediction of the motion trajectory of living obstacles and the analysis of the collision risk of static obstacles, the system establishes a multi-dimensional collaborative decision-making mechanism, which can intelligently balance the requirements of dynamic and static obstacle avoidance. When a direction conflict is detected, the system automatically selects the optimal obstacle avoidance strategy based on the safety distance priority to ensure the best passing efficiency while guaranteeing safety. This technology significantly improves the autonomous navigation ability of the device in complex scenarios with both dynamic and static obstacles, and realizes an integrated intelligent decision-making of environmental perception, risk assessment, and path planning.

[0092] To solve the problems of inaccurate recognition of living obstacles and difficult assessment of the collision risk of static obstacles in a dynamic environment, in some embodiments, in the obstacle topology map in step 104, extracting a connected domain that conforms to the periodic biological motion law in the temperature anomaly region as a living obstacle, and determining the collision risk boundary of the static obstacle according to the projection offset of the geometric deformation feature between consecutive frames includes: 601. Extract the periodic intensity fluctuation parameter of the temperature anomaly region from the obstacle topology map, calculate the fluctuation period through the intensity change curve of the temperature anomaly region in consecutive frames, and when the fluctuation period matches the preset biological motion period threshold range, mark the corresponding temperature anomaly region as a candidate region for living obstacles; In step 601, the periodic intensity fluctuation parameter of the temperature anomaly region refers to the regular parameter of the thermal radiation intensity changing with time extracted from the near-infrared sensor data, such as a periodic fluctuation of 1.2 times per second ± 0.1 Hz. The intensity change curve of the temperature anomaly region in consecutive frames is an intensity-time series generated from multi-frame thermal imaging data (such as a sampling rate of 30 Hz), such as 30 thermal intensity sampling points per second. The fluctuation period is the period value corresponding to the main frequency of the Fourier transform of the intensity change curve (such as the period T = 0.83 seconds). The preset biological motion period threshold range is the typical motion frequency range of living obstacles (such as the human walking frequency of 0.8 - 2.0 Hz, the animal breathing frequency of 0.2 - 0.5 Hz). The candidate region for living obstacles is the temperature anomaly region whose fluctuation period conforms to the biological threshold, such as marked as a red highlighted block (confidence level ≥ 90%).

[0093] In the embodiment of the present application, first, time-series sampling is performed on the intensity data of the temperature anomaly region in the obstacle topology map to generate an intensity change curve of consecutive frames (for example, the interval between each frame is 100 ms), and periodic fluctuation parameters (such as the main frequency of 1.2 Hz and the amplitude of ±5 °C) are extracted through fast Fourier transform (FFT). Subsequently, the calculated fluctuation period is matched with the preset biological motion period threshold range (such as the human walking frequency of 0.8 - 2 Hz and the animal breathing frequency of 0.3 - 1 Hz). If the main frequency falls within the threshold range, a living obstacle candidate label is assigned to the temperature anomaly region. For example, when the main frequency of temperature fluctuation in a certain region is 1.5 Hz, it is determined that it conforms to the human walking characteristics and is marked as a candidate region. During this process, the continuous frame data is segmented through a sliding window algorithm to eliminate noise interference (such as environmental temperature drift) to ensure the robustness of periodic parameter extraction.

[0094] 602. Within the candidate region, analyze the continuity of the movement trajectory of the temperature anomaly region, and verify the biological motion law through the displacement direction and speed fluctuation pattern of the centroid of the temperature anomaly region in consecutive frames. When the random change rate of the displacement direction exceeds the preset threshold and the fluctuation amplitude of the speed conforms to the biological acceleration characteristics, confirm that the candidate region is a living obstacle; In step 602, the candidate region refers to the suspected living region preliminarily screened in step 601 (such as a circular region with a radius of 0.5 meters). The continuity of the movement trajectory of the temperature anomaly region is calculated by the motion smoothness index (such as the standard deviation of displacement direction change ≤ 10°) of the displacement sequence of the region centroid in consecutive frames (such as the centroid coordinates [x1, y1] → [x2, y2] → [x3, y3]). The displacement direction of the centroid is the azimuth angle of the line connecting the centroids between adjacent frames (such as the direction angle changes from 30° to 35° between two frames, and the change rate is 5° / frame). The speed fluctuation pattern is the standard deviation of the centroid movement speed (such as the fluctuation amplitude of the speed sequence [0.8 m / s, 1.0 m / s, 0.9 m / s] is ±0.1 m / s). The verification of the biological motion law is through the comprehensive determination of the randomness of the displacement direction (such as the direction change rate ≥ 5° / frame) and the speed fluctuation amplitude (such as conforming to the human walking acceleration of ±0.3 m / s²). The random change rate of the displacement direction is the variance of the direction angle change per unit time (such as triggering the living body determination when the variance ≥ 15°). The biological acceleration characteristic is the unique speed change pattern of living body movement (such as the acceleration peak conforms to a sine curve, R² ≥ 0.8).

[0095] In the embodiment of the present application, for the candidate region marked in step 601, first, the centroid position of the temperature anomaly region in consecutive frames is tracked through the optical flow method (Lucas-Kanade algorithm), and the displacement vector is calculated (such as )(e.g., v = 0.3 m / s). Subsequently, the random change rate of the displacement direction (e.g., the standard deviation of the direction angle σ = 25° within 10 frames) and the speed fluctuation amplitude (e.g., the root mean square acceleration of 0.1 m / s²) are statistically calculated and compared with the biological motion feature library: If the direction change rate exceeds the preset threshold (e.g., σ > 20°) and the acceleration conforms to the biological characteristics (e.g., the acceleration of human walking is 0.1 - 0.5 m / s²), it is determined as a living obstacle. During this process, the centroid trajectory is smoothed through Kalman filtering to eliminate the error caused by sensor jitter, and finally the topological map label is updated to confirm the candidate area as a living obstacle.

[0096] 603. Synchronize the projection offset of the geometric deformation feature in consecutive frames, calculate the included angle relationship between the deformation contour movement trajectory of the static obstacle and the movement direction of the mobile device. When the included angle between the projection offset direction and the device movement direction is less than the preset acute angle threshold, mark the area corresponding to the deformation contour as the collision risk boundary of the static obstacle.

[0097] In step 603, the projection offset of the geometric deformation feature refers to the position difference of the static obstacle in consecutive frames in the visual coordinate system (e.g., the offset between two frames is 0.2 meters). The deformation contour movement trajectory of the static obstacle is a displacement sequence constructed by the projection offsets of multiple frames (e.g., the trajectory points [0.1m, 0.3m, 0.5m]). The movement direction of the mobile device is the azimuth angle of the device's current movement (e.g., the due north direction is 0°). The included angle relationship is the included angle between the deformation contour movement trajectory direction and the device movement direction (e.g., the trajectory direction is 30°, the device direction is 45°, and the included angle is 15°). The preset acute angle threshold is the critical angle for determining the conflict between the trajectory and the device direction (e.g., 30°). The projection offset direction refers to the vector direction of the deformation contour movement trajectory (e.g., the azimuth angle is 30°). The collision risk boundary is the safety buffer area of the static obstacle marked according to the included angle threshold (e.g., when the included angle < 30°, the inflated contour is 0.5 meters to generate a red warning boundary).

[0098] In the embodiment of the present application, first, the projection offset of the geometric deformation feature in consecutive frames is calculated through the homography matrix (e.g., ), extract its movement trajectory direction vector (e.g., ), and obtain the device's own movement direction vector based on the IMU data of the mobile device (e.g., ). Calculate the included angle between the two through the vector included angle formula: , when θ is less than the preset acute angle threshold (e.g., 30°), it is determined that the deformation contour offset direction is close to coinciding with the device movement direction, and there is a collision risk. At this time, the morphological dilation algorithm is used to expand the deformation contour by 0.5 m to generate a collision risk boundary, which is marked as a high-risk area of the static obstacle in the topological map. This process improves the adaptability to different scenarios through dynamic threshold adjustment (e.g., increasing the dilation distance according to the device speed).

[0099] The following is a specific example: In the multi-obstacle recognition scenario of a domestic service robot, the system achieves accurate classification through thermal-deformation multimodal analysis: First, extract the temperature anomaly area (such as a hot spot of 39.5°C ± 0.8°C) from the topological map, and analyze its intensity fluctuation curve through Fourier transform. When a periodic fluctuation of 1.2Hz ± 0.3Hz (matching the movement characteristics of felines) is detected, it is marked as a living candidate area; then analyze the movement trajectory of the centroid of the hot spot in the candidate area. If there is a direction mutation > 45° / 0.1s and a speed fluctuation of 0.3 - 0.7m / s (meeting the start-stop characteristics of pets) within 8 consecutive frames (133ms), it is confirmed as a living obstacle (confidence level 94%). At the same time, calculate the projection offset of the furniture deformation contour (such as a coffee table leg tilted 5°) in the robot's moving direction (0.5m / s, due east) through RGB-D data. When the westward movement rate of the deformation area is detected to be 0.05m / s (angle < 15°), it is marked as a static collision risk boundary and a virtual anti-collision area expanded by 8cm is generated. This solution improves the accuracy of living body recognition and reduces the misjudgment rate of static obstacles. In the composite scenario of a pet suddenly rushing out (speed 0.8m / s) and furniture displacement (0.1m / s), the system can complete the classification response within 80ms.

[0100] In summary, through steps 601 to 603, the mobile device realizes intelligent recognition and risk assessment of dynamic and static obstacles in a complex environment. The system establishes a dual verification mechanism based on biological motion characteristics and geometric deformation characteristics through the fusion analysis of multimodal perception data, and can accurately distinguish living obstacles from static obstacles. This technology significantly improves the accuracy and reliability of obstacle recognition. Especially in a complex scenario with a mixture of dynamic and static obstacles, it can provide accurate environmental perception and safety warning capabilities for the mobile device, effectively ensuring the safety and stability of autonomous navigation.

[0101] To solve the problems of inaccurate steering control and insufficient motion jitter compensation during the dynamic obstacle avoidance process of the mobile device, in some embodiments, generating the path planning parameters of the mobile device based on the composite avoidance parameter or the dominant avoidance direction and combining with the moving speed in step 504 includes: 701. Generate the expected steering angle of the mobile device according to the angle between the avoidance direction in the composite avoidance parameter or the dominant avoidance direction and the current moving direction, and at the same time adjust the moving speed based on the speed decay coefficient in the composite avoidance parameter to generate the target moving speed in the avoidance state; In step 701, the composite avoidance parameter is the path correction parameter synthesized after the direction conflict detection in step 503, such as a vector including a steering angle of 10° and a speed attenuation coefficient of 0.7. The dominant avoidance direction is a single avoidance direction selected after the conflict detection, such as turning right by 20°. The included angle between the avoidance direction and the current moving direction is the deflection angle of the correction direction from the original moving direction of the device (e.g., included angle = 15°). The expected steering angle is the device steering instruction parameter generated according to the included angle (e.g., turning left by 15°). The speed attenuation coefficient is the deceleration ratio calculated according to the dynamic safety distance (e.g., the original speed of 1 m / s is reduced to 0.7 m / s). The target moving speed is the adjusted real-time speed in the avoidance state (e.g., 0.7 m / s ± 0.1 m / s).

[0102] In the embodiment of the present application, first, the included angle between the composite avoidance direction (or the dominant avoidance direction) and the current moving direction is calculated through the vector included angle formula:

[0103] where is the avoidance direction vector, is the current moving direction vector. According to the included angle the expected steering angle is generated: And the current speed is adjusted through the speed attenuation coefficient (e.g., ) to generate the target moving speed: For example, when the current speed is and the included angle , the expected steering angle is , and the target speed is reduced to . This process ensures that the avoidance action is smooth and conforms to the kinematic constraints of the device by dynamically adjusting the steering and speed.

[0104] 702. Correlate the direction and speed components of the expected steering angle and the target moving speed, and combine them to generate a steering curvature parameter; In step 702, the expected steering angle is the input parameter for device steering control (e.g., the steering angle of the steering gear is 15°). The target moving speed is the adjusted real-time speed value (e.g., 0.7 m / s). The direction and speed component correlation is to decompose the speed into components in the steering direction (e.g., the tangential speed component along the steering angle of 15° is 0.68 m / s, and the normal component is 0.18 m / s). The steering curvature parameter is the theoretical turning radius calculated according to the steering angle and the speed (e.g., the curvature radius R = 1 m / s² / (0.7 m / s × tan15°) ≈ 3.8 m).

[0105] In the embodiment of the present application, first, the expected steering angle and the target moving speed Associate through the curvature formula: Calculate the steering curvature parameter, where is the control period (such as seconds). For example, when (converted to radians: 0.42 rad), then: .

[0106] Subsequently, truncate the curvature exceeding the threshold through the curvature limit condition (such as ) to ensure that the steering action is within the allowable range of the device's mechanical structure. This process generates curvature parameters under physical constraints, providing geometric feasibility guarantees for path planning.

[0107] 703. Integrate the steering curvature parameter and the moving speed to generate path planning parameters including speed adjustment instructions and steering control instructions; In step 703, the steering curvature parameter is a geometric constraint parameter for path planning (such as the curvature radius ≥ 1.5 m). The moving speed is the actual current speed of the device (such as 0.7 m / s). The speed adjustment instruction is a pulse width modulation (PWM) signal for controlling the motor output (such as a 30% reduction in the duty cycle). The steering control instruction is a control command for driving the steering mechanism (such as the pulse width of the servo PWM signal being 1500 μs). The path planning parameters are a set including speed and steering instructions (such as the JSON format instruction: {"speed": 0.7, "steer_angle": 15}).

[0108] In the embodiment of the present application, first, based on the steering curvature parameter C, generate a sequence of path points through a B-spline curve (such as an arc path corresponding to the curvature radius R = 1 / C), and combine it with the target moving speed to generate a speed curve (such as the tangential speed along the path). Subsequently, fuse the steering and speed constraints through the dynamic window algorithm (DWA) to screen the optimal path point and speed combination, and finally generate path planning parameters, including a sequence of target heading angles (such as ), a sequence of speed instructions (such as ), and a sequence of steering angles (such as ). These parameters are sent to the drive controller through ROS (Robot Operating System) to drive the device to perform avoidance actions according to the planned path.

[0109] 704. Decompose the device's moving direction into horizontal and vertical components based on the motion trajectory data, generate a horizontal jitter compensation factor according to the projection value and direction angle deviation of the moving speed of the horizontal component, and generate a vertical jitter compensation factor according to the projection value and direction angle deviation of the moving speed of the vertical component; In step 704, the motion trajectory data is a sequence of spatio-temporal coordinates during the device's movement (such as 10 groups per second Data). The horizontal component is the X-axis projection of the device's moving direction on a two-dimensional plane (e.g., velocity x = 0.8 m / s, direction angle θ = 30°). The vertical component is the Y-axis projection of the device's moving direction on a two-dimensional plane (e.g., velocity y = 0.6 m / s). The projected value of the moving speed is the magnitude of the component of the speed in the horizontal or vertical direction (e.g., horizontal projection 0.8 m / s, vertical projection 0.6 m / s). The direction angle deviation is the deflection angle between the actual moving direction of the device and the target direction (e.g., deviation angle = 5°). The horizontal jitter compensation factor is a correction parameter calculated based on the horizontal projected speed and the angle deviation (e.g., horizontal compensation factor = 0.8 × cos5° ≈ 0.796). The vertical jitter compensation factor is a correction parameter for the vertical projected speed and the angle deviation (e.g., vertical compensation factor = 0.6 × sin5° ≈ 0.052).

[0110] In the embodiment of the present application, first, the device motion trajectory data is kinematically decomposed into a horizontal component (X-axis velocity , direction angle deviation ) and a vertical component ( axis velocity , direction angle deviation ).

[0111] The horizontal jitter compensation factor is calculated by the formula: , and the vertical jitter compensation factor is calculated by the formula: .

[0112] For example: when the horizontal speed and the direction deviation (converted to radians: 0.087 rad): ; When the vertical speed and (converted to radians: 0.052 rad): , this process provides a quantitative basis for jitter compensation by quantifying the coupled influence of the direction deviation and the speed.

[0113] 705. Generate a compensation coefficient through the non-linear fusion of the horizontal jitter compensation factor and the vertical jitter compensation factor.

[0114] In step 705, the horizontal jitter compensation factor is a dynamic parameter that cancels the horizontal jitter of the device (e.g., 0.796). The vertical jitter compensation factor is a dynamic parameter that cancels the vertical jitter of the device (e.g., 0.052). The non-linear fusion combines the horizontal and vertical compensation factors through a weighted sum of squares formula (e.g., compensation coefficient = ). The compensation coefficient is the comprehensive correction parameter finally used to stabilize the visual coordinate system (such as horizontal compensation 0.78 and vertical compensation 0.12).

[0115] In the embodiment of the present application, first, the non - linear weighted fusion algorithm is used to combine the horizontal jitter compensation factor with the vertical jitter compensation factor into the global compensation coefficient , where = 0.6, = 0.4 are the weight coefficients, and λ = 0.2 is the interaction term coefficient. For example, when = 0.92, = 0.94, = 0.6×0.92 + 0.4×0.94 + 0.2×0.92×0.94≈0.93. This coefficient is used to adjust the control gain of the device actuator (such as motor torque and steering servo response speed) to suppress the trajectory deviation caused by motion jitter.

[0116] In summary, through steps 701 to 705, an integrated solution for intelligent motion planning and stability control of a mobile device during dynamic obstacle avoidance is achieved. By establishing a collaborative optimization mechanism for steering angle and speed adjustment, the system can generate a smooth obstacle - avoidance trajectory that takes into account both safety and motion efficiency in real time. This technology not only realizes the precise coordination of direction control and speed control to ensure the smooth execution of obstacle - avoidance actions, but also can adaptively adjust the compensation parameters according to the real - time motion state, significantly improving the motion stability and control accuracy in complex obstacle - avoidance scenarios, providing a reliable motion control guarantee for the autonomous navigation of mobile devices.

[0117] Figure 2 FIG. shows a schematic structural diagram of an optimized system for room obstacle detection based on a cross - modal attention mechanism provided by an embodiment of the present application. As Figure 2 shown, the system includes: An acquisition module 21, which acquires the motion trajectory data of the mobile device, calculates the compensation coefficient of the device moving direction and the viewing angle jitter amplitude based on the motion trajectory data, and generates a visual coordinate system matching the physical space; A collection module 22, which activates the cross - modal acquisition mode of the multispectral sensor in the visual coordinate system, and divides the visible light and near - infrared dual - band overlapping perception area according to the moving forward area corresponding to the mobile device. The area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection, and the area covered by the near - infrared band is used to detect the temperature anomaly area related to biological signs; Adjustment module 23 inputs the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into the cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, and generate an obstacle topology map; Extraction module 24 extracts the connected regions that conform to the periodic biological movement law in the temperature anomaly region in the obstacle topology map as living obstacles, and determines the collision risk boundary of static obstacles according to the projection offset amount of the geometric deformation feature between consecutive frames; Generation module 25 generates the path planning parameters of the mobile device based on the periodic movement law of the living obstacle and the offset trend of the collision risk boundary.

[0118] Figure 2 The described room obstacle detection optimization system based on cross-modal attention mechanism can execute Figure 1 The described room obstacle detection optimization method based on cross-modal attention mechanism in the illustrated embodiment, its implementation principle and technical effects will not be elaborated. For the room obstacle detection optimization system based on cross-modal attention mechanism in the above embodiment, the specific ways for each module and unit to execute operations have been described in detail in the embodiment related to the method, and will not be elaborated here.

[0119] In a possible design, Figure 2 The room obstacle detection optimization system based on cross-modal attention mechanism in the illustrated embodiment can be implemented as a computing device, such as Figure 3 shown, this computing device can include a storage component 31 and a processing component 32; The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32.

[0120] The processing component 32 is used for the Figure 1 room obstacle detection optimization method based on cross-modal attention mechanism in the above

[0121] embodiment. Among them, the processing component 32 can include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.

[0122] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0123] Of course, the computing device may also necessarily include other components, such as input / output interfaces, display components, communication components, etc.

[0124] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above-mentioned peripheral interface module can be an output device, an input device, etc.

[0125] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.

[0126] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device can refer to a cloud server, and the above-mentioned processing component, storage component, etc. can be basic server resources leased or purchased from a cloud computing platform.

[0127] The embodiments of the present application also provide a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 An optimization method for room obstacle detection based on a cross-modal attention mechanism shown in the embodiments.

[0128] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0130] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An optimized method for room obstacle detection based on cross-modal attention mechanism, characterized in that, Including: Obtaining the motion trajectory data of a mobile device, calculating a compensation coefficient for the device's moving direction and the amplitude of view jitter based on the motion trajectory data, and generating a visual coordinate system that matches the physical space; Activating the cross-modal acquisition mode of a multispectral sensor under the visual coordinate system, and dividing a visible light and near-infrared dual-band overlapping perception area according to the forward moving area corresponding to the mobile device, where the area covered by the visible light band is used to detect the geometric deformation characteristics of ground projections, and the area covered by the near-infrared band is used to detect temperature anomaly areas related to biological signs; Inputting the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly area into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, in order to generate an obstacle topology map; In the obstacle topology map, extracting the connected regions that conform to the periodic biological motion law in the temperature anomaly area as living obstacles, and determining the collision risk boundary of static obstacles according to the projection offset of the geometric deformation characteristics between consecutive frames; Generating path planning parameters for the mobile device based on the periodic motion law of the living obstacles and the offset trend of the collision risk boundary; 2. The method according to claim 1, characterized in that, Inputting the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly area into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, in order to generate an obstacle topology map, including: Inputting the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly area into a cross-modal attention network, and generating a dynamic allocation parameter based on the moving speed and direction change amount of the mobile device through the cross-modal attention network; Calculating the weight distribution ratio of the visible light data stream and the near-infrared data stream through the dynamic allocation parameter, and performing weighted fusion on the geometric deformation characteristics and the temperature anomaly area according to the weight distribution ratio; Performing cross-modal interaction on the weighted geometric deformation characteristics and the temperature anomaly area, generating a deformation heat correlation mark by overlapping and comparing the geometric deformation contour and the temperature anomaly area in the spatial dimension, and constructing a dynamic displacement chain by combining the relative offset of the deformation contour and the temperature anomaly area in consecutive frames with the displacement correction factor of the compensation coefficient in the time dimension; Generating an obstacle topology map that fuses biological characteristics and deformation based on the dynamic displacement chain and the deformation heat correlation mark; 3. The method according to claim 2, characterized in that, Generating an obstacle dynamic topology map that fuses biological characteristics and deformation based on the dynamic displacement chain and the deformation heat correlation mark, including: When the direction deviation between the displacement direction of the geometric deformation contour and the migration direction of the temperature anomaly area in the dynamic displacement chain is less than a preset threshold, extracting synchronous offset trajectory data from the dynamic displacement chain to construct a dynamic trajectory layer of the living obstacle; By using the spatial superposition relationship between the geometric deformation contour and the temperature anomaly region, when the projection boundary of the geometric deformation contour coincides with the outer boundary of the temperature anomaly region, it is marked as the spatial distribution layer of static obstacles; Based on the fracture characteristics of the geometric deformation contour and the periodic intensity fluctuation of the temperature anomaly region in the deformation-thermal correlation mark, when the fracture position matches the phase of the fluctuation period, an association response layer of potential interaction behaviors is generated through the coverage relationship between the extended region of the fracture boundary and the thermal fluctuation region; The spatial distribution layer, the dynamic trajectory layer, and the association response layer are superimposed in the order of priority to form an obstacle topology map that fuses geometric deformation features and biothermal features.

4. The method according to claim 1, characterized in that, In the visual coordinate system, activate the cross-modal acquisition mode of the multispectral sensor. According to the mobile forward region corresponding to the mobile device, divide the visible light and near-infrared dual-band overlapping perception region, where the visible light band coverage region is used to detect the geometric deformation features of the ground projection objects, and the near-infrared band coverage region is used to detect the temperature anomaly regions related to biological signs, including: In the visual coordinate system, determine the mobile forward region according to the moving direction of the mobile device, and in the mobile forward region, form a fan-shaped coverage region with the current position of the device as the vertex along the moving direction, and adjust the radius length of the fan-shaped coverage region according to the moving speed; In the adjusted fan-shaped coverage region, activate the dual-band synchronous acquisition mode of the visible light sensor and the near-infrared sensor, so that the physical acquisition ranges of both completely cover the fan-shaped coverage region, where the visible light sensor focuses on the continuous frame contour changes of the floor projection surface, and the near-infrared sensor focuses on the thermal intensity gradient changes between adjacent pixels; Map the continuous frame contour changes collected by the visible light sensor into geometric deformation features, calculate the position coordinates of the deformation boundary through the displacement difference between the contour boundaries of adjacent frames. At the same time, map the thermal intensity gradient changes collected by the near-infrared sensor into temperature anomaly regions, and extract the boundary coordinates of the temperature anomaly regions through the spatial distribution of thermal intensity mutation points; Based on the position coordinates and the boundary coordinates, divide the visible light and near-infrared overlapping perception regions in the adjusted fan-shaped coverage region.

5. The method according to claim 1, characterized in that, Based on the periodic motion law of the living obstacle and the offset trend of the collision risk boundary, generate the path planning parameters of the mobile device, including: Extract the periodic parameters based on the periodic motion law of the living obstacle and calculate the prediction range of its motion trajectory. Generate the first avoidance direction adjustment amount according to the geometric relationship between the motion trajectory prediction range and the current position of the mobile device; Calculate the angle between the projection offset rate of the geometric deformation feature and the moving direction based on the offset trend of the collision risk boundary. When the projection offset rate exceeds the preset threshold, determine the dynamic safety distance through the product of the angle and the moving speed and generate the second avoidance direction adjustment amount; Perform direction conflict detection on the first avoidance direction adjustment amount and the second avoidance direction adjustment amount. If the direction angle is less than the preset angle, generate a composite avoidance parameter by superposition. If the direction angle is greater than or equal to the preset angle, select the direction with a smaller dynamic safety distance as the dominant avoidance direction; Based on the composite avoidance parameter or the dominant avoidance direction, and in combination with the moving speed, generate the path planning parameter of the mobile device.

6. The method according to claim 1, characterized in that, In the obstacle topology map, extract the connected domain that conforms to the periodic biological movement law in the temperature anomaly area as a living obstacle, and determine the collision risk boundary of the static obstacle according to the projection offset amount of the geometric deformation feature between consecutive frames, including: Extract the periodic intensity fluctuation parameter of the temperature anomaly area from the obstacle topology map, calculate the fluctuation period through the intensity change curve of the temperature anomaly area in consecutive frames. When the fluctuation period matches the preset biological movement period threshold range, mark the corresponding temperature anomaly area as the candidate area of the living obstacle; In the candidate area, analyze the continuity of the movement trajectory of the temperature anomaly area, and verify the biological movement law through the displacement direction and speed fluctuation mode of the centroid of the temperature anomaly area in consecutive frames. When the random change rate of the displacement direction exceeds the preset threshold and the fluctuation amplitude of the speed conforms to the biological acceleration feature, confirm that the candidate area is a living obstacle; Synchronously based on the projection offset amount of the geometric deformation feature in consecutive frames, calculate the included angle relationship between the deformation contour movement trajectory of the static obstacle and the movement direction of the mobile device. When the included angle between the projection offset direction and the device movement direction is less than the preset acute angle threshold, mark the area corresponding to the deformation contour as the collision risk boundary of the static obstacle.

7. The method according to claim 5, characterized in that, Based on the composite avoidance parameter or the dominant avoidance direction, and in combination with the moving speed, generate the path planning parameter of the mobile device, including: According to the included angle between the avoidance direction in the composite avoidance parameter or the dominant avoidance direction and the current movement direction, generate the expected steering angle of the mobile device. At the same time, adjust the moving speed based on the speed attenuation coefficient in the composite avoidance parameter to generate the target moving speed in the avoidance state; Perform direction-speed component association on the expected steering angle and the target moving speed, and combine them to generate a steering curvature parameter; Fuse the steering curvature parameter and the moving speed to generate a path planning parameter including a speed adjustment instruction and a steering control instruction; The calculating the compensation coefficient of the device movement direction and the viewing angle jitter amplitude based on the movement trajectory data includes: Decompose the device movement direction into a horizontal component and a vertical component based on the movement trajectory data, generate a horizontal jitter compensation factor according to the moving speed projection value and the direction angle deviation of the horizontal component, and generate a vertical jitter compensation factor according to the moving speed projection value and the direction angle deviation of the vertical component; Generate a compensation coefficient through the non-linear fusion of the horizontal jitter compensation factor and the vertical jitter compensation factor.

8. An optimized system for room obstacle detection based on a cross-modal attention mechanism, characterized in that, Including: An acquisition module, which acquires the movement trajectory data of the mobile device, calculates the compensation coefficient of the device movement direction and the viewing angle jitter amplitude based on the movement trajectory data, and generates a visual coordinate system that matches the physical space; The acquisition module activates the cross-modal acquisition mode of the multispectral sensor in the visual coordinate system, and divides the visible light and near-infrared dual-band overlapping perception regions according to the moving forward region corresponding to the mobile device, where the visible light band coverage region is used to detect the geometric deformation characteristics of the ground projection object, and the near-infrared band coverage region is used to detect the temperature anomaly region related to biological signs; The adjustment module inputs the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly region into the cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities through the cross-modal attention network according to the moving speed and direction change amount of the mobile device, so as to generate an obstacle topology map; The extraction module extracts the connected regions that conform to the periodic biological movement law in the temperature anomaly region in the obstacle topology map as living obstacles, and determines the collision risk boundary of the static obstacle according to the projection offset amount of the geometric deformation characteristics between consecutive frames; The generation module generates the path planning parameters of the mobile device based on the periodic movement law of the living obstacle and the offset trend of the collision risk boundary.

9. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an optimized method for room obstacle detection based on cross-modal attention mechanism according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a computer, it implements an optimized method for room obstacle detection based on cross-modal attention mechanism according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Autonomous decision-making unmanned driving method, device and system based on multi-modal perception

    CN117406719A

  • Intelligent agent enhanced path planning method based on multi-modal information fusion

    CN117784776A

  • Obstacle speed detection method and device, computer equipment and storage medium

    CN117930220A

  • Multi-sensor fusion obstacle detection method

    CN117968860A

  • Campus inspection robot navigation method based on large model fusion environment and biological multi-modal information

    CN118329044A

Cited By

  • Multi-model migration optimization method and system for aviation large model, computing device and storage medium

    CN120354639A

  • Humanoid robot multi-mode dynamic jump test system, method, equipment and medium

    CN120538863A

  • Robot obstacle recognition method and system based on artificial intelligence

    CN120715888A

  • Plastic packaging barrel online defect detection method and system based on machine vision and deep learning

    CN120726012A

  • Building equipment operation data intelligent analysis method and system

    CN120822840A