An Optimization Method and System for Room Obstacle Detection Based on Cross-Modal Attention Mechanism

Through the obstacle detection method of cross-modal attention mechanism, a visual coordinate system is constructed using multi-spectral sensors and motion trajectory data, modal weights are dynamically adjusted, and the obstacle topology map is generated, which solves the accuracy of dynamic obstacle detection and multiple obstacle distinction problems in mobile robots, and improves the stability and efficiency of obstacle avoidance decisions.

CN120178893BActive Publication Date: 2025-07-25BEIJING XINGWANG SHIP POWER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510661437.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-25
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The prior art has problems in the detection of dynamic obstacles in mobile robots with poor adaptability to dynamic environments and weak distinction ability of multiple obstacles. It is easy to misjudgment of stationary objects as obstacles, and it is difficult to accurately distinguish between living objects and static obstacles in complex scenarios, resulting in unstable obstacle avoidance decisions.

Method used

The obstacle detection method based on the cross-modal attention mechanism is adopted, and the visual coordinate system is calculated by obtaining the motion trajectory data of the mobile device, combining the cross-modal acquisition mode of multi-spectral sensors, the geometric deformation characteristics of the ground projection are detected using visible light and the temperature abnormal areas of the near-infrared detection of biological signs. The attention weight allocation between modes is adjusted through the cross-modal attention network to generate obstacle topology maps, and path planning parameters are generated based on the living body's motion laws and collision risk boundaries.

Benefits of technology

It realizes the accurate distinction between living bodies and static obstacles in complex environments, improves obstacle avoidance response speed and environmental adaptability, reduces the misjudgment rate, and provides high-precision path planning decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178893B_ABST
    Figure CN120178893B_ABST
Patent Text Reader

Abstract

The present application provides an optimization method and system for room obstacle detection based on a cross-modal attention mechanism. Among them, a dynamic compensation coefficient is calculated by obtaining the motion trajectory data of a mobile device to construct a stable visual coordinate system. The visible light and near-infrared bands of a multispectral sensor are synchronously triggered for acquisition. The visible light captures the geometric deformation characteristics of the floor projection plane, and the near-infrared detects local thermal radiation abnormal areas. The thermal feature connected regions that conform to periodic biological movements are extracted as living obstacles, and the collision risk boundaries of static obstacles are marked. Based on the living movement rules and the trend of collision boundary offset, the dual-band exposure parameters are synchronously optimized and the trajectory prediction model is updated to form a closed-loop feedback link for perception and planning. The technical solution provided by the present application realizes the accurate distinction between living and static obstacles in a dynamic environment and adaptively adjusts the sensor parameters to improve the real-time obstacle avoidance response ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent robot environmental perception and navigation optimization, and particularly relates to an optimization method and system for room obstacle detection based on a cross-modal attention mechanism. Background Art

[0002] In mobile robots such as floor-sweeping robots or smart home scenarios, real-time trajectory prediction and obstacle avoidance of dynamic obstacles such as pets and moving furniture need to solve core problems such as multi-modal perception data fusion, dynamic intention reasoning, and low-latency decision-making. The system needs to accurately capture the motion state and sudden behaviors of obstacles, such as a pet suddenly turning, to ensure the real-time nature and safety of decision-making.

[0003] Currently, some solutions adopt a lightweight method based on single-modal threshold segmentation and linear extrapolation. The moving target area is extracted through a background subtraction algorithm, and the historical trajectory sequence is constructed using the centroid coordinates of the target to directly extrapolate the future position. For multi-obstacle scenarios, simple clustering based on distance thresholds is used to classify adjacent pixel regions as the same obstacle, and finally, an obstacle avoidance instruction is generated in combination with the movement speed of the robot.

[0004] However, this solution has obvious shortcomings: poor adaptability to dynamic environments. The background subtraction algorithm is sensitive to light changes and shadow interference, and it is easy to misjudge stationary objects as obstacles. The ability to distinguish multiple obstacles is weak. Clustering based on pixel distance is difficult to distinguish dense moving targets and is often misjudged as a single obstacle, leading to risks of missed detection or collisions. Summary of the Invention

[0005] This application provides an optimization method and system for room obstacle detection based on a cross-modal attention mechanism to solve the problems of low prediction accuracy and weak multi-obstacle discrimination ability in the prior art.

[0006] In a first aspect, this application provides an optimization method for room obstacle detection based on a cross-modal attention mechanism, including:

[0007] Obtain the motion trajectory data of the mobile device, calculate the compensation coefficient of the device's moving direction and the viewing angle jitter amplitude based on the motion trajectory data, and generate a visual coordinate system that matches the physical space;

[0008] Activate the cross-modal acquisition mode of the multi-spectral sensor in the visual coordinate system, and divide the visible light and near-infrared dual-band overlapping perception area according to the moving forward area corresponding to the mobile device, where the area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection, and the area covered by the near-infrared band is used to detect the temperature anomaly area related to biological signs;

[0009] Input the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, and generate an obstacle topology map;

[0010] In the obstacle topology map, extract the connected regions that conform to the periodic biological motion law in the temperature anomaly region as living obstacles, and determine the collision risk boundary of static obstacles according to the projection offset of the geometric deformation feature between consecutive frames;

[0011] Generate the path planning parameters of the mobile device based on the periodic motion law of the living obstacle and the offset trend of the collision risk boundary.

[0012] Optionally, input the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into a cross-modal attention network, and generate dynamic allocation parameters based on the moving speed and direction change amount of the mobile device through the cross-modal attention network;

[0013] Calculate the weight distribution ratio of the visible light data stream and the near-infrared data stream through the dynamic allocation parameters, and perform weighted fusion on the geometric deformation feature and the temperature anomaly region according to the weight distribution ratio;

[0014] Perform cross-modal interaction on the weighted geometric deformation feature and the temperature anomaly region. Generate a deformation-thermal correlation mark by overlapping and comparing the geometric deformation contour and the temperature anomaly region in the spatial dimension, and construct a dynamic displacement chain by combining the relative offset of the deformation contour and the temperature anomaly region in consecutive frames with the displacement correction factor of the compensation coefficient in the time dimension;

[0015] Generate an obstacle topology map that fuses biological features and deformations based on the dynamic displacement chain and the deformation-thermal correlation mark.

[0016] Optionally, when the direction deviation between the displacement direction of the geometric deformation contour in the dynamic displacement chain and the migration direction of the temperature anomaly region is less than a preset threshold, extract synchronous offset trajectory data from the dynamic displacement chain to construct a dynamic trajectory layer of the living obstacle;

[0017] Through the spatial superposition relationship between the geometric deformation contour and the temperature anomaly region, mark the projection boundary of the geometric deformation contour coinciding with the outer edge boundary of the temperature anomaly region as the spatial distribution layer of the static obstacle;

[0018] Based on the fracture characteristics of the geometric deformation contour in the deformation heat correlation markers and the periodic intensity fluctuations in the temperature anomaly regions, when the fracture position matches the fluctuation period phase, an association response layer of potential interaction behaviors is generated through the coverage relationship between the extended region of the fracture boundary and the thermal fluctuation region;

[0019] Superimpose the spatial distribution layer, the dynamic trajectory layer, and the association response layer in the order of priority to form an obstacle topology map that integrates geometric deformation features and biothermal features.

[0020] Optionally, in the visual coordinate system, determine the moving forward region according to the moving direction of the mobile device, form a fan-shaped coverage region in the moving forward region with the current position of the device as the vertex along the moving direction, and adjust the radius length of the fan-shaped coverage region according to the moving speed;

[0021] Activate the dual-band synchronous acquisition mode of the visible light sensor and the near-infrared sensor within the adjusted fan-shaped coverage region, so that the physical acquisition ranges of both completely cover the fan-shaped coverage region, where the visible light sensor focuses on the continuous frame contour changes of the floor projection plane, and the near-infrared sensor focuses on the thermal intensity gradient changes between adjacent pixels;

[0022] Map the continuous frame contour changes collected by the visible light sensor into geometric deformation features, calculate the position coordinates of the deformation boundary through the displacement difference of the contour boundaries between adjacent frames, and at the same time map the thermal intensity gradient changes collected by the near-infrared sensor into temperature anomaly regions, and extract the boundary coordinates of the temperature anomaly regions through the spatial distribution of thermal intensity mutation points;

[0023] Based on the position coordinates and the boundary coordinates, divide the visible light and near-infrared overlapping perception regions within the adjusted fan-shaped coverage region.

[0024] Optionally, extract periodic parameters based on the periodic motion law of the living obstacle and calculate the predicted range of its motion trajectory, and generate a first avoidance direction adjustment amount according to the geometric relationship between the predicted range of the motion trajectory and the current position of the mobile device;

[0025] Calculate the angle between the projection offset rate of the geometric deformation feature and the moving direction based on the offset trend of the collision risk boundary. When the projection offset rate exceeds the preset threshold, determine the dynamic safety distance through the product of the angle and the moving speed and generate a second avoidance direction adjustment amount;

[0026] Perform a direction conflict detection on the first avoidance direction adjustment amount and the second avoidance direction adjustment amount. If the direction angle is less than the preset angle, superimpose and generate a composite avoidance parameter. If the direction angle is greater than or equal to the preset angle, select the direction with a smaller dynamic safety distance as the dominant avoidance direction;

[0027] Generate the path planning parameters of the mobile device based on the composite avoidance parameters or the dominant avoidance direction and in combination with the moving speed.

[0028] Optionally, extract the periodic intensity fluctuation parameters of the temperature anomaly region from the obstacle topology map, calculate the fluctuation period through the intensity change curve of the temperature anomaly region in consecutive frames, and when the fluctuation period matches the preset biological motion period threshold range, mark the corresponding temperature anomaly region as a candidate region for living obstacles;

[0029] Within the candidate region, analyze the continuity of the moving trajectory of the temperature anomaly region, and verify the biological motion law through the displacement direction and speed fluctuation pattern of the centroid of the temperature anomaly region in consecutive frames. When the random change rate of the displacement direction exceeds the preset threshold and the fluctuation amplitude of the speed conforms to the biological acceleration characteristics, confirm the candidate region as a living obstacle;

[0030] Synchronously based on the projection offset of the geometric deformation feature in consecutive frames, calculate the included angle relationship between the deformation contour moving trajectory of the static obstacle and the moving direction of the mobile device. When the included angle between the projection offset direction and the device moving direction is less than the preset acute angle threshold, mark the region corresponding to the deformation contour as the collision risk boundary of the static obstacle.

[0031] Optionally, generate the expected steering angle of the mobile device according to the included angle between the avoidance direction in the composite avoidance parameters or the dominant avoidance direction and the current moving direction, and at the same time adjust the moving speed based on the speed attenuation coefficient in the composite avoidance parameters to generate the target moving speed in the avoidance state;

[0032] Associate the expected steering angle and the target moving speed in terms of direction speed components, and combine them to generate the steering curvature parameter;

[0033] Fuse the steering curvature parameter and the moving speed to generate the path planning parameters including speed adjustment instructions and steering control instructions;

[0034] The calculating the compensation coefficient of the moving direction of the device and the amplitude of the view jitter based on the motion trajectory data includes:

[0035] Decompose the moving direction of the device into a horizontal component and a vertical component based on the motion trajectory data, generate a horizontal jitter compensation factor according to the projection value and direction angle deviation of the moving speed of the horizontal component, and generate a vertical jitter compensation factor according to the projection value and direction angle deviation of the moving speed of the vertical component;

[0036] Generate the compensation coefficient through the non-linear fusion of the horizontal jitter compensation factor and the vertical jitter compensation factor.

[0037] Second aspect, the present application provides an optimized system for room obstacle detection based on a cross-modal attention mechanism, including:

[0038] An acquisition module that acquires the motion trajectory data of a mobile device, calculates a compensation coefficient for the device's moving direction and the amplitude of view jitter based on the motion trajectory data, and generates a visual coordinate system that matches the physical space;

[0039] A collection module that activates the cross-modal acquisition mode of a multispectral sensor in the visual coordinate system, divides the visible light and near-infrared dual-band overlapping perception regions according to the forward moving area corresponding to the mobile device, where the visible light band coverage area is used to detect the geometric deformation characteristics of ground projections, and the near-infrared band coverage area is used to detect temperature anomaly regions related to biological signs;

[0040] An adjustment module that inputs the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly region into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network to generate an obstacle topology map;

[0041] An extraction module that extracts the connected regions that conform to the periodic biological motion law in the temperature anomaly region in the obstacle topology map as living obstacles, and determines the collision risk boundary of static obstacles according to the projection offset amount of the geometric deformation characteristics between consecutive frames;

[0042] A generation module that generates the path planning parameters of the mobile device based on the periodic motion law of the living obstacles and the offset trend of the collision risk boundary.

[0043] Third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an optimized method for room obstacle detection based on a cross-modal attention mechanism as described in the first aspect above.

[0044] Fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it implements an optimized method for room obstacle detection based on a cross-modal attention mechanism as described in the first aspect.

[0045] In the embodiments of the present application, the motion trajectory data of a mobile device is obtained, a compensation coefficient for the device moving direction and the viewing angle jitter amplitude is calculated based on the motion trajectory data, and a visual coordinate system matching the physical space is generated; in the visual coordinate system, a cross-modal acquisition mode of a multispectral sensor is activated, and a visible light and near-infrared dual-band overlapping perception area is divided according to the mobile forward area corresponding to the mobile device, wherein the area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection object, and the area covered by the near-infrared band is used to detect the temperature abnormal area related to biological signs; the compensation coefficient, the geometric deformation characteristics and the temperature abnormal area are input into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, so as to generate an obstacle topology map; in the obstacle topology map, a connected domain conforming to the periodic biological motion law in the temperature abnormal area is extracted as a living obstacle, and the collision risk boundary of the static obstacle is determined according to the projection offset amount of the geometric deformation characteristics between consecutive frames; based on the periodic motion law of the living obstacle and the offset trend of the collision risk boundary, path planning parameters of the mobile device are generated.

[0046] The embodiments of the present application have the following beneficial effects:

[0047] By fusing the motion trajectory compensation of the mobile device and the multispectral cross-modal perception, a visual coordinate system matching the physical space is dynamically constructed. The geometric deformation characteristics (such as shape and contour) of the ground projection object are detected by the visible light band, and the temperature abnormality of biological signs is captured by the near-infrared band, realizing multi-dimensional perception of environmental obstacles; further combined with the cross-modal attention network, the weights of visible light and near-infrared are adaptively allocated according to the mobile state of the device, and a topology map including static and living obstacles is generated, in which living targets (such as humans and animals) are accurately identified through the periodic biological motion law, and the collision risk boundary of the static obstacle is deduced through the projection offset amount of the geometric deformation characteristics between consecutive frames; finally, by comprehensively considering the dynamic behaviors and risk trends of the two types of obstacles, real-time and safe path planning parameters are generated, significantly improving the ability of the mobile device to distinguish between biological and static obstacles, the dynamic obstacle avoidance response speed and the environmental adaptability in complex scenarios, and being applicable to application scenarios such as service robots and autonomous driving that need to balance safety and efficiency.

[0048] Furthermore, by dynamically allocating parameters to adaptively adjust the weights of visible light and near-infrared, the real-time performance of static obstacle detection in high-speed moving scenarios (relying on geometric deformation) and the accuracy of living body recognition in low-speed scenarios (relying on temperature anomalies) are improved; the deformation thermal correlation marking in the spatial dimension enhances the ability to distinguish the mixed area of living organisms and static obstacles (such as separating the overlapping features when a human body approaches an obstacle), and the dynamic displacement chain in the time dimension combines with the displacement correction factor to eliminate the interference of device jitter and accurately calculate the movement trajectory and risk boundary of the obstacle; the finally generated topological map significantly reduces the misjudgment rate (such as high-temperature non-biological interference and missed detection of static obstacles) by fusing spatio-temporal two-dimensional information, providing a high-precision and robust environmental perception basis for obstacle avoidance decisions in complex dynamic scenarios.

[0049] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0051] Figure 1 FIG. shows a flowchart of an optimized method for room obstacle detection based on a cross-modal attention mechanism provided by the present application;

[0052] Figure 2 FIG. shows a schematic structural diagram of an optimized system for room obstacle detection based on a cross-modal attention mechanism provided by the present application;

[0053] Figure 3 FIG. shows a schematic structural diagram of a computing device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.

[0055] In some of the processes described in the specification, claims, and the above-mentioned drawings of this application, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0056] Researchers have found that there are significant defects in the obstacle detection methods of existing mobile devices in complex scenarios: traditional single-modal sensors (such as pure vision or pure infrared) are difficult to take into account both the geometric features of static obstacles and the biological signs of living targets, resulting in missed detections or false detections; the perspective jitter and motion trajectory deviation during device movement are likely to cause coordinate system misalignment, reducing the accuracy of multi-modal data fusion; the ability to distinguish the attributes of obstacles (such as living / static) in real time and predict risks in a dynamic environment is insufficient, affecting the reliability of path planning. Based on this, an optimized method for room obstacle detection based on a cross-modal attention mechanism is provided. This method can construct a high-precision visual coordinate system through motion trajectory compensation and multi-spectral collaborative perception, dynamically allocate the modal weights of visible light and near-infrared according to the device motion state, and generate an obstacle topology map that combines biological signs and collision risks using spatio-temporal two-dimensional features (geometric deformation contours and temperature anomaly regions), and finally realizes the periodic motion tracking of living obstacles and the real-time obstacle avoidance decision-making of static obstacles.

[0057] The technical solution of this application is applicable to scenarios such as indoor navigation of service robots, complex road condition perception of autonomous driving vehicles, and human body dynamic detection in security monitoring. It is particularly suitable for application environments where there are coexisting biological and static obstacles, frequent device movement jitter, and drastic changes in environmental light or temperature.

[0058] Next, the technical solutions in the embodiments of this application will be described clearly and completely with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.

[0059] Figure 1 A flowchart of an optimized method for room obstacle detection based on a cross-modal attention mechanism is provided for the embodiments of this application, as Figure 1 shown, this method includes:

[0060] 101. Obtain the motion trajectory data of the mobile device, calculate the compensation coefficients of the device movement direction and the viewing angle jitter amplitude based on the motion trajectory data, and generate a visual coordinate system that matches the physical space;

[0061] In this step, the motion trajectory data refers to a continuous sequence of spatio-temporal coordinates collected by the accelerometer, gyroscope, and GPS sensor of the mobile device. For example, 10 sets of data including position , speed and acceleration are collected per second, and are used to describe the device movement direction and the viewing angle jitter amplitude. The device movement direction includes the azimuth deviation , and the viewing angle jitter amplitude includes the pitch angle fluctuation . The compensation coefficients are dynamic correction parameters calculated based on the trajectory data. For example, the horizontal compensation and the vertical compensation are generated after fusing the sensor noise through Kalman filtering, where the noise variance is used to offset the influence of the device jitter on the visual coordinate system. The visual coordinate system is a reference system that maps the physical space to the local coordinates of the device. Its origin is the centroid of the device, and the Z-axis points to the direction of gravity . The matching accuracy with the physical space is controlled within the error range meters through coordinate transformation. The coordinate transformation is implemented using the Euler angle rotation matrix:

[0062] .

[0063] In the embodiment of the present application, first, the motion trajectory data is collected through the inertial measurement unit and GPS module of the mobile device, including acceleration angular velocity and displacement increment and other parameters. The Kalman filtering algorithm is used to smooth the noise data, and the device movement direction and the viewing angle jitter amplitude are calculated based on the kinematic model. The device movement direction is obtained through the heading angle calculation formula, where the heading angle:

[0064]

[0065] The viewing angle jitter amplitude is obtained by integrating the angular velocity to get the pitch angle fluctuation:

[0066]

[0067] Secondly, the device attitude data is converted into the visual coordinate system in the physical space through the coordinate transformation algorithm, where the compensation coefficients:

[0068]

[0069] To offset the perspective shift caused by jitter. Finally, the compensated visual coordinate system is written to memory, with the origin of the coordinate system being the current position of the device The X-axis and Y-axis are aligned with the north-south and east-west directions of the physical space respectively for multispectral sensor calibration

[0070] In the autonomous navigation system of a certain intelligent sweeping robot, the acceleration of the device is collected in real time through a built-in nine-axis inertial measurement unit sensor angular velocity and the magnetometer data B to construct a three-dimensional motion trajectory model. When the robot moves on an uneven ground, an improved Kalman filtering algorithm is used to denoise the original data, and the filtering equation is

[0071]

[0072] And combined with the quaternion attitude solution method , calculate the instantaneous offset of the pitch angle and roll angle of the device. And through the dynamic weighted fusion algorithm , convert the offset into a compensation coefficient for perspective jitter, and establish a visual coordinate system with the centroid of the device as the origin. This coordinate system is rigidly registered with the physical space grid map through SLAM point cloud data, and the registration error satisfies , achieving centimeter-level spatial alignment

[0073] 102. Activate the cross-modal acquisition mode of the multispectral sensor in the visual coordinate system, and divide the visible light and near-infrared dual-band overlapping perception area according to the moving forward area corresponding to the mobile device. The area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection, and the area covered by the near-infrared band is used to detect the temperature anomaly area related to biological signs

[0074] In this step, the cross-modal acquisition mode of the multispectral sensor is a cooperative working mode that synchronously activates the visible light sensor and the near-infrared sensor. The working wavelength of the visible light sensor is 400 - 700 nm, the working wavelength of the near-infrared sensor is 700 - 1000 nm, and the sampling interval is 33 milliseconds. The forward moving area is a fan-shaped detection area centered on the device's moving direction, with a horizontal opening angle of 60 degrees and a depth of 10 meters. The geometric deformation features of the ground projection objects are extracted from the visible light band coverage area through the Canny edge detection algorithm, such as the trapezoidal distortion caused by the tilt of a rectangular object. The calculation formula for the geometric deformation feature is that the distortion rate is equal to Δ aspect ratio divided by the original aspect ratio. The temperature anomaly area is identified from the near-infrared band coverage area through the thermal radiation intensity threshold, and the threshold is set to be greater than or equal to 35 degrees Celsius or the temperature difference from the environment is greater than or equal to 3 degrees Celsius. The temperature anomaly area is the area marked as potential biological signs, with a confidence level of not less than 85%.

[0075] In the embodiment of the present application, first, based on the visual coordinate system generated in step 101, the synchronous acquisition mode of the multispectral sensor is activated. The working wavelength of the visible light camera is 400 - 700 nm, the working wavelength of the near-infrared sensor is 800 - 1400 nm, and the two are spatially aligned in the forward fan-shaped area along the device's moving direction, with a field of view angle of 120 degrees. The forward area is divided into a visible light band coverage area and a near-infrared band coverage area through the region segmentation algorithm, where the visible light area covers the central 60-degree field of view angle, and the near-infrared area covers 30-degree field of view angles on both sides. The geometric deformation features of the ground projection objects are extracted from the visible light area through the background difference method, including parameters such as edge curvature and aspect ratio. The temperature anomaly area is used to identify biological heat sources through temperature threshold segmentation, and the temperature threshold is set to be greater than 35 degrees Celsius as abnormal. The dual-band data is aligned in the overlapping area through the spatio-temporal registration module, and this module is based on the feature point matching algorithm, and finally marked as the cross-modal fusion area.

[0076] For example, after completing the coordinate system calibration, the sweeping robot activates the multispectral fusion camera mounted on the top and switches to the cross-modal synchronous acquisition mode. According to the current moving speed v of the robot equal to 0.3 m / s and the direction angle θ, the forward 60-degree field of view angle is divided into a visible light and near-infrared dual-band overlapping sensing area. The working wavelength of the visible light is 400 - 700 nm, and the ground texture is captured at a frame rate of 30 fps. The geometric deformation of projections such as carpet edges and wires is detected through SIFT feature matching, and the deformation threshold is set to 2 mm. The working wavelength of the near-infrared is 850 - 1700 nm, and an uncooled microbolometer is used to scan the temperature distribution within 1.5 meters in front with an accuracy of thermal sensitivity of 0.05 degrees Celsius, and the biological heat source area that is 0.5 degrees Celsius higher than the ambient temperature, such as the feet of pets or humans, is identified.

[0077] 103. Input the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into the cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, and generate an obstacle topology map;

[0078] In this step, the cross-modal attention network is a dual-modal neural network composed of a visible light branch ResNet-18 and a near-infrared branch MobileNet-V2. The moving speed and direction change amount are obtained by calculating the first derivative of the trajectory data. For example, the speed v = 1.2 m / s and the direction change rate ω = 0.1 rad / s. The attention weight distribution dynamically adjusts the contribution ratio of visible light and near-infrared according to the moving speed. When the speed v > 2 m / s, the visible light weight = 0.8. The obstacle topology map is a grid map that marks the position, type, and safety distance of obstacles, and the resolution δ = 0.05 m / grid.

[0079] In the embodiment of the present application, first, input the compensation coefficient calculated in step 101 , the geometric deformation feature G of the ground projection object in the visible light band extracted in step 102, and the temperature anomaly region T detected in the near-infrared band into the cross-modal attention network. Through the multi-head attention mechanism, this network maps the geometric deformation feature and the temperature anomaly region into high-dimensional feature vectors and . According to the real-time speed and direction change amount of the mobile device, dynamically calculate the attention weight distribution ratios and of the two modalities. The specific implementation is as follows:

[0080] When the speed increases or the direction changes violently, that is, when the condition is satisfied, the network increases the weight of the near-infrared modality, which is increased from the baseline value of 0.4 to 0.7, to give priority to responding to living obstacles; when moving at a low speed and stably, that is, the visible light modality weight dominates, strengthening the detection of static obstacles.

[0081] The network realizes dynamic weight adjustment through the following formula:

[0082]

[0083] where is the sigmoid function, and are the speed and direction change sensitivity coefficients, and is the bias term.

[0084] Finally, the network outputs an obstacle topological map that fuses bimodal features. , where the nodes in the figure represent the positions of obstacles, and the edge weights represent the probability of collision risk between obstacles. For dynamic obstacles, the velocity vector and the predicted trajectory are additionally marked. The compensation coefficient is input into the pre-trained cross-modal attention network together with the geometric deformation feature vector G and the coordinates T of the temperature anomaly region. This network contains two layers of bidirectional LSTM modules. When the robot's acceleration is detected, the weight of the near-infrared modality is dynamically increased to 0.7; when the robot is in a uniform motion state, the weight of the visible light modality is restored to 0.6. The network outputs an obstacle probability map , which generates a topological grid map containing the height and the heat source intensity attributes after morphological closing operation. Among them, dynamic obstacles are marked as red grids, and static obstacles are marked as blue grids.

[0085] 104. In the obstacle topological map, extract the connected regions that conform to the periodic biological motion law in the temperature anomaly region as living obstacles, and determine the collision risk boundary of static obstacles according to the projection offset of the geometric deformation feature between consecutive frames;

[0086] In this step, the periodic biological motion law is the living motion feature extracted by near-infrared data spectrum analysis, where the fast Fourier transform is used to detect the feature frequency, and the typical step frequency range is The connected region is defined as the smallest bounding rectangle composed of continuous pixels in the temperature anomaly region, which needs to meet the conditions of area and aspect ratio . The projection offset is calculated by the optical flow method, which represents the displacement of the static obstacle in the continuous visual coordinate system, and the calculation formula is . Among them, represents the coordinate of the feature point at time. The collision risk boundary is dynamically adjusted according to the projection offset, and its inflation radius The calculation formula of The safety factor takes a value of 1.5.

[0087] In the embodiment of the present application, first, 8-neighborhood clustering analysis is performed on the temperature anomaly region, and the motion periodicity feature is extracted by discrete Fourier transform. For the detected spectral peak , it is determined as a biological motion feature when is satisfied. At the same time, an area threshold is set for connected component filtering. For geometric deformation features, a feature point projection calculation method based on the homography matrix is adopted, and the displacement is calculated by the formula . When exceeds the threshold , it is determined as a static obstacle, and its collision risk boundary is generated by the convex hull algorithm. The boundary dilation amount adopts an adaptive algorithm, and the calculation formula is .

[0088] In specific implementation, multi-frame analysis is performed on the temperature anomaly area: the continuous periodic temperature fluctuation area is extracted by the background difference method, and its frequency feature needs to satisfy to correspond to the biological signs. Combining the optical flow method to calculate the included angle between the regional movement vector and the robot's movement trajectory. When , it is determined as a living obstacle approaching positively, and a three-level avoidance strategy is triggered. For static obstacles, the Lucas-Kanade dense optical flow algorithm is used to calculate the projection offset of geometric feature points within 5 consecutive frames. When satisfies the radial distribution condition in the polar coordinate system: , its collision risk boundary is determined as a fan-shaped area centered on the obstacle with a radius , where is the radial-tangential displacement ratio threshold.

[0089] 105. Based on the periodic motion law of the living obstacle and the offset trend of the collision risk boundary, generate the path planning parameters of the mobile device.

[0090] In this step, the periodic motion law of the living obstacle is used to predict its future position, such as the Kalman filter prediction error of ±0.1 m. The offset trend of the collision risk boundary is a parameter obtained by linear regression analysis of the boundary expansion rate, such as an increase of 0.05 m per second. The path planning parameters include the obstacle avoidance priority (living weight 0.9), the path curvature limit (radius ≥ 0.8 m), and the speed adjustment coefficient (the speed decays to 50% when approaching the risk boundary), generating a safe path, and the path length increases by no more than 20%.

[0091] In the embodiments of the present application, first, according to the periodic motion law of the living obstacle extracted in step 104 (such as the pedestrian walking frequency of 1.2 Hz and the moving direction due north) and the offset trend of the collision risk boundary of the static obstacle (such as the boundary expanding westward at a rate of 0.1 m / s), the dynamic window algorithm is used to calculate the feasible path of the mobile device. Specifically, the future trajectory of the living obstacle is predicted by Kalman filtering to generate a dynamic avoidance area centered on the living body (radius 1.5 m); at the same time, according to the offset trend of the static boundary, a buffer zone with real-time expansion and contraction is generated around the obstacle (such as width 0.5 m). After fusing the constraints of the dynamic avoidance area and the buffer zone, based on the kinematic parameters of the mobile device (maximum steering angle of 30 degrees, acceleration upper limit of 1 m / s²), path planning parameters are generated, and the parameters synchronously control the exposure duration ratio of the visible light and near-infrared bands in the forward area of the mobile device before movement, and trigger the cross-modal attention network to perform incremental iterative updates on the trajectory prediction module of the connected domain of the living obstacle. It includes the target heading angle (such as adjusted to 25 degrees east of northeast), the speed curve (such as decreasing from 2 m / s to 0.8 m / s), and the steering angle sequence, and is sent to the drive controller for execution through the serial communication protocol.

[0092] Based on the motion period function of the living obstacle , where determined by the heat source fluctuation frequency and the risk boundary gradient data of the static obstacle, the path planning module adopts an improved dynamic window algorithm (DWA). A time-varying repulsive force field is applied to the living obstacle, and the repulsive force coefficient decays with the inverse square of the distance; for the static obstacle, a direction constraint term is introduced to limit the feasible speed of the robot in the tangent direction of the risk boundary inside. Finally, the planning parameters including the maximum safe speed , the steering angle and the emergency braking probability are output, and the hub motor is controlled through the ROS system to realize the S-shaped obstacle avoidance trajectory, and the measured obstacle avoidance success rate is improved.

[0093] In summary, through steps 101 to 105, the mobile device realizes high-precision autonomous obstacle avoidance and navigation in a complex environment. By fusing visible light and near-infrared multi-modal perception data, the system can simultaneously identify the geometric deformation characteristics of static obstacles and the biometric signals of dynamic living bodies, and construct an obstacle topology map containing spatio-temporal information. The innovative cross-modal attention mechanism can dynamically adjust the perception weight according to the motion state of the device, ensuring stable environmental modeling ability even during fast movement or perspective jitter. The finally generated path planning parameters comprehensively consider the periodic motion law of living obstacles and the collision risk boundary of static obstacles, enabling the device to have predictive obstacle avoidance ability and significantly improving the navigation reliability in challenging scenarios such as low light and complex backgrounds.

[0094] To solve the problem of cross-modal data fusion for obstacle detection in mobile devices in complex environments, in some embodiments, in step 103, inputting the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the change in the moving speed and direction of the mobile device through the cross-modal attention network to generate an obstacle topology map, including:

[0095] 201. Input the compensation coefficient, the geometric deformation feature, and the temperature anomaly region into the cross-modal attention network, and generate dynamic allocation parameters through the cross-modal attention network based on the change in the moving speed and direction of the mobile device;

[0096] In step 201, the compensation coefficient refers to a dynamic correction parameter calculated based on the device's movement trajectory (for example, horizontal compensation 0.7, vertical compensation 0.5), which is used to offset the influence of perspective jitter on the visual coordinate system. The geometric deformation feature is the ground object shape distortion parameter extracted from visible light data (such as trapezoidal distortion rate = 0.15). The temperature anomaly region is the region where the thermal radiation intensity detected by the near-infrared sensor exceeds the threshold (such as temperature difference ≥ 3°C). The cross-modal attention network is a dual-modal neural network composed of a visible light branch (ResNet-18) and a near-infrared branch (MobileNet-V2). The change in the moving speed and direction of the mobile device is a parameter calculated by the first derivative of the trajectory data (for example, speed = 1.5 m / s, direction change rate = 0.2 rad / s). The dynamic allocation parameter is a weight adjustment parameter generated by the cross-modal attention network according to the speed and direction change (for example, when the speed > 1 m / s, the visible light weight increases by 5%).

[0097] In the embodiments of the present application, first, input the compensation coefficient (such as K = 0.85) calculated in step 101, the geometric deformation features (such as edge curvature, aspect ratio) of the visible light band extracted in step 102, and the temperature anomaly region (such as heat source contour) of the near-infrared band into the cross-modal attention network. This network adopts a multi-head attention mechanism, and respectively performs high-dimensional encoding on the geometric deformation feature and the temperature anomaly region through a convolutional neural network (CNN) to generate feature vectors of the visible light and near-infrared modalities; subsequently, based on the real-time moving speed (such as 2 m / s) and direction change amount (such as deflecting 15 degrees per second) of the mobile device, calculate the initial attention scores of the two modalities through a fully connected layer, and dynamically weight-adjust the scores using the compensation coefficient as a weight factor, and finally output the dynamic allocation parameters of the visible light and near-infrared modalities (such as visible light weight 0.7, near-infrared weight 0.3). This process realizes the adaptive allocation of modal weights through three-stage processing of feature embedding, dynamic parameter calculation, and compensation weighting.

[0098] 202. Calculate the weight distribution ratio of the visible light data stream and the near-infrared data stream based on the dynamic allocation parameter, and perform weighted fusion on the geometric deformation feature and the temperature anomaly region according to the weight distribution ratio;

[0099] In step 202, the dynamic allocation parameter is an intermediate variable that controls the fusion ratio of visible light and near-infrared data (such as visible light weight 0.7, near-infrared weight 0.3). The visible light data stream is a continuous image sequence collected by a visible light sensor (resolution 1920×1080, 30fps). The near-infrared data stream is a thermal radiation intensity matrix output by a near-infrared sensor (resolution 640×480, 30fps). The weight distribution ratio is the contribution ratio of the visible light and near-infrared modalities determined according to the dynamic allocation parameter (such as 7:3). The weighted fusion is a process of superimposing the geometric deformation feature (visible light modality) and the temperature anomaly region (near-infrared modality) according to the weight ratio (for example, the fusion formula: 0.7×geometric feature + 0.3×temperature feature).

[0100] In the embodiment of the present application, first, based on the dynamic allocation parameter (such as visible light weight 0.7, near-infrared weight 0.3), the weight distribution ratio (visible light ratio ≈ 0.67, near-infrared ratio ≈ 0.33) is calculated by normalizing through the Softmax function. Subsequently, spatial alignment processing is performed on the visible light and near-infrared data: an affine transformation is used to map the two to the same visual coordinate system to eliminate the perspective shift caused by device movement; then, the geometric deformation feature (such as the edge gradient matrix) and the temperature anomaly region (such as the heat source probability map) are weighted and superimposed according to the weight ratio to generate a fused feature map (such as visible light feature × 0.67 + near-infrared feature × 0.33), and the noise of the fusion result is suppressed by Gaussian filtering to eliminate discrete anomaly points. This process realizes the effective integration of cross-modal data through weight allocation, spatial alignment, and feature fusion.

[0101] 203. Perform cross-modal interaction on the weighted geometric deformation feature and the temperature anomaly region, generate a deformation heat correlation mark by overlapping and comparing the geometric deformation contour and the temperature anomaly region in the spatial dimension, and construct a dynamic displacement chain by combining the relative offset of the deformation contour and the temperature anomaly region in consecutive frames with the displacement correction factor of the compensation coefficient in the time dimension;

[0102] In step 203, cross-modal interaction refers to the joint analysis of visible light and near-infrared data in the spatial and temporal dimensions. The geometric deformation contour is the contour of ground objects extracted by edge detection (such as the sequence of polygon vertex coordinates). The temperature anomaly region is the calibrated high-temperature connected domain in the thermal image (area ≥ 0.5 m²). The deformation-thermal correlation mark is the mark of the overlap between the geometric contour and the temperature region in the spatial dimension (for example, the region with a contour overlap rate > 80% is marked as highly correlated). Consecutive frames refer to adjacent visual coordinate system data in the time series (interval 33 ms). The relative offset is the displacement difference between the deformation contour and the temperature region in consecutive frames (such as a lateral offset of 0.2 meters). The displacement correction factor is a dynamic correction parameter calculated based on the compensation coefficient (for example, the horizontal displacement compensation factor is 0.8). The dynamic displacement chain is the displacement trajectory constructed by the offset amounts and correction factors of multiple frames in the time dimension (for example, the trajectory sequence: [0.1 m, 0.15 m, 0.2 m]).

[0103] In the embodiment of the present application, first, in the spatial dimension, the sliding window algorithm is used to scan the fused feature map to locate the overlapping region between the geometric deformation contour and the temperature anomaly region (such as the overlapping area ratio > 50%), and assign a deformation-thermal correlation mark to the overlapping region (the binary mask is marked as 1, and the non-overlapping region is 0). In the time dimension, the optical flow method (Lucas-Kanade algorithm) is used to track the displacement vectors of the deformation contour and the temperature anomaly region in consecutive frames (such as ), and the displacement amount is corrected in combination with the compensation coefficient (such as K = 0.9) ( ), and finally the corrected displacement vectors are stored in the time series to construct a dynamic displacement chain (such as [(0.18, 0.09), (0.20, 0.10)]). Through the joint analysis of the spatial overlap mark and the time displacement chain, the spatio-temporal consistency modeling of cross-modal interaction is realized.

[0104] 204. Generate an obstacle topology map that fuses biometric features and deformations based on the dynamic displacement chain and the deformation-thermal correlation mark.

[0105] In step 204, the dynamic displacement chain is a trajectory chain that describes the movement trend of obstacles in the time dimension (such as a displacement increment of 0.1 meter per second). The deformation-thermal correlation mark is the joint feature identifier of geometric deformation and temperature anomaly in the spatial dimension (for example, the confidence level of the overlapping region ≥ 90%). Biometric features refer to the living body attributes identified through temperature anomalies and periodic motion laws (such as the human walking frequency of 1.2 Hz). The obstacle topology map is a grid map that fuses deformation, temperature, and motion features (resolution 0.1 meter / grid), and marks the positions and risk levels of static obstacles (gray) and living body obstacles (red).

[0106] In the embodiments of the present application, first, the center point of the region marked with deformation heat correlation is used as a topological node (such as coordinates (x = 3.5, y = 2.1)), and the motion correlation between nodes (such as cosine similarity) is calculated based on the dynamic displacement chain, and this is used as the edge weight (such as similarity 0.8 corresponding to weight 0.2). Subsequently, a topological model of nodes and edges is built through a graph neural network (GNN), and a biological feature (such as a temperature value of 42 °C) and a deformation feature (such as a curvature of 0.6) are bound to each node, and finally an obstacle topological map is generated. The nodes in the map represent the positions and attributes of obstacles, and the edge weights represent the collision risk correlation intensity (in the range of 0 - 1). Through node generation, motion correlation analysis, and attribute binding, this process realizes the structured expression of multi-dimensional obstacle information.

[0107] In summary, in steps 201 to 204, multi-modal intelligent perception and dynamic obstacle modeling of the mobile device in a complex environment are realized. Through the adaptive weight allocation mechanism of the cross-modal attention network, the system can optimize the fusion strategy of visible light and near-infrared data in real time according to the device motion state, ensuring the best environmental perception effect in different motion scenarios. The innovative deformation heat correlation marking technology realizes the precise spatial alignment of geometric features and biological features, and the construction of the dynamic displacement chain effectively integrates the motion trajectory information in the time dimension. The finally generated obstacle topological map not only contains the accurate geometric contours of static obstacles but also completely retains the motion characteristics of dynamic biological targets, providing the mobile device with environmental cognitive capabilities with both spatial accuracy and temporal continuity.

[0108] To solve problems such as multi-modal data conflict, motion interference, and lack of biological features existing in traditional obstacle detection methods in complex dynamic scenarios, in some embodiments, generating an obstacle dynamic topological map that fuses biological features and deformations in step 204 includes:

[0109] 301. When the direction deviation between the displacement direction of the geometric deformation contour in the dynamic displacement chain and the migration direction of the temperature anomaly region is less than a preset threshold, extract synchronous offset trajectory data from the dynamic displacement chain to construct a dynamic trajectory layer of a living obstacle;

[0110] In step 301, the dynamic displacement chain is a sequence of obstacle motion trajectories constructed by multi-frame displacement correction factors, such as a displacement increment of 0.1 meters per second. The geometric deformation contour is the edge contour of ground objects extracted from visible light data, such as a rectangular distortion contour composed of a set of polygon vertex coordinates. The migration direction of the temperature anomaly region is the moving direction of the thermal radiation region in consecutive frames, such as a moving path with an azimuth angle of 30°. The direction deviation refers to the angular difference between the displacement direction of the geometric deformation contour and the migration direction of the temperature region. Specifically, when the deviation angle ≤ 10 degrees, a threshold judgment is triggered. The preset threshold is a critical parameter for judging direction consistency. For example, the direction deviation angle threshold is 15 degrees. The synchronized offset trajectory data is a set of displacement trajectories extracted when the direction deviation is below the threshold, such as a trajectory chain containing the coordinates [0.1m, 0.2m, …, 1.0m] for 10 consecutive frames. The living obstacle is a dynamic object with periodic biological signs, such as a moving target with a pedestrian step frequency feature of 1.2 Hz. The dynamic trajectory layer is a rasterized map layer that marks the movement path of the living obstacle, and the raster granularity is 0.1 meter per pixel.

[0111] In the embodiment of the present application, first, the direction deviation between the displacement direction of the geometric deformation contour (such as the vector ΔV = (0.2, 0.1)) and the migration direction of the temperature anomaly region (such as the vector ΔU = (0.18, 0.09)) in the dynamic displacement chain is calculated by the vector angle cosine algorithm. The formula is: direction deviation = 1 - (ΔV·ΔU) / (||ΔV||×||ΔU||). When the direction deviation is less than the preset threshold (such as 0.1), it is determined that the two motions are synchronized. Subsequently, the time series clustering algorithm (such as DBSCAN) is used to extract the synchronized offset trajectory data from the dynamic displacement chain, specifically including the trajectory starting point, the moving direction angle, and the average speed. Finally, a smooth dynamic trajectory layer of the living obstacle is generated by B-spline curve fitting, and the trajectory data is stored in the spatio-temporal database according to the time stamp.

[0112] 302. Through the spatial superposition relationship between the geometric deformation contour and the temperature anomaly region, when the projection boundary of the geometric deformation contour coincides with the outer edge boundary of the temperature anomaly region, it is marked as the spatial distribution layer of static obstacles;

[0113] In step 302, the spatial superposition relationship refers to the overlapping degree between the geometric deformation contour and the temperature anomaly region in the visual coordinate system. For example, 80% of the area coincides in the horizontal direction. The projection boundary of the geometric deformation contour is the set of two-dimensional projection outer contour vertices of ground objects in the visible light band, such as the four corner coordinates of a rectangular object. The outer edge boundary of the temperature anomaly region refers to the contour polygon of the high-temperature region in thermal imaging, such as the side length data of an irregular shape generated by the thermal radiation intensity threshold. The static obstacle is an object without biological characteristics and with a fixed position, such as a roadblock stone with a diameter of 0.5 meters. The spatial distribution layer is a raster layer that marks the position and contour of the static obstacle, and the raster is filled with gray coding.

[0114] In the embodiment of the present application, first, the intersection over union (IoU) algorithm is used to calculate the overlap degree between the projected boundary of the geometric deformation contour (such as the rectangular box coordinates (x1, y1, x2, y2)) and the outer edge boundary of the temperature anomaly region (such as the polygon vertex set). When the IoU value is greater than the threshold (such as 0.7), it is determined that the two are spatially coincident, and the coincident region is marked as the spatial distribution layer of the static obstacle through rasterization processing. Specifically, the morphological dilation algorithm is used to expand the boundary of the coincident region by 0.2 m to generate a safety buffer zone, and the spatial distribution layer is stored in a binary matrix (the obstacle area is 1, and the idle area is 0).

[0115] 303. Based on the fracture characteristics of the geometric deformation contour and the periodic intensity fluctuation of the temperature anomaly region in the deformation heat correlation mark, when the fracture position matches the fluctuation period phase, an association response layer of potential interaction behavior is generated through the coverage relationship between the extended region of the fracture boundary and the thermal fluctuation region;

[0116] In step 303, the deformation heat correlation mark is an overlapping identifier of the geometric contour and the temperature anomaly region in space. For example, a red semi-transparent overlay is set in the overlapping region. The fracture characteristic is the discontinuous edge fracture in the geometric deformation contour caused by occlusion or damage, such as a contour notch with a length of 0.3 m. The periodic intensity fluctuation refers to the regular change of the thermal radiation intensity in the temperature region. The fracture position is the set of physical coordinates of the contour fracture points. The fluctuation period phase is the time stage mark of the thermal intensity fluctuation. For example, when t = 0.3T is divided within the period T, it is the rising edge stage. The extended region of the fracture boundary is a circular buffer zone expanded 0.5 m centered on the fracture point, which is used to detect potential interaction regions. The thermal fluctuation region is a sub-region with periodic changes in the temperature anomaly region, such as the core area of thermal radiation change with a radius of 0.8 m. The coverage relationship is the determination of the position overlap ratio between the extended region and the fluctuation region. The potential interaction behavior refers to the approaching event of a living obstacle to a static obstacle during movement. The association response layer is a yellow grid layer marking the interaction risk area.

[0117] In the embodiment of the present application, first, the fracture characteristics (such as the fracture length > 0.5 m) of the geometric deformation contour in the deformation heat correlation mark are extracted through the Canny edge detection algorithm, and at the same time, the Fourier transform is used to analyze the intensity fluctuation period (such as the main frequency of 1.2 Hz) of the temperature anomaly region. When the time stamp of the fracture position and the phase difference of the fluctuation period are less than the preset value (such as 10% of the period duration), it is determined that the two match. Subsequently, the region growing algorithm is used to generate an extended region by expanding 0.3 m from the fracture boundary, and the coverage area ratio between it and the thermal fluctuation region is calculated. If the coverage ratio exceeds the threshold (such as 60%), an interaction behavior mark (such as "suspected living body contact") is assigned to this region, and finally, a thermal distribution map of the association response layer is generated.

[0118] 304. Superimpose the spatial distribution layer, the dynamic trajectory layer, and the association response layer in the order of priority to form an obstacle topological map that integrates geometric deformation features and biological heat features.

[0119] In step 304, the spatial distribution layer is a gray grid background layer that calibrates the spatial positions of static obstacles. The dynamic trajectory layer is a superimposed layer that marks the moving paths of living bodies with red linear shapes, and the line width is 3 pixels. The association response layer is a semi-transparent yellow layer that highlights potential interaction areas, and its coverage priority is higher than that of the static layer. The priority order defines the display weight rules for layer superimposition. Integrating geometric deformation features and biological heat features refers to combining the contour distortion rate of visible light (deformation rate ≥ 15%) and the body temperature fluctuation period of near-infrared (period matching error ≤ 0.1 s). The obstacle topological map is the final grid map that integrates the information of the three layers, with an output resolution of 1024 × 768 pixels, containing the coordinates, types, and risk level parameters of various types of obstacles.

[0120] In the embodiments of the present application, first, according to the preset priority rules (the spatial distribution layer is the bottom layer, the dynamic trajectory layer is the middle layer, and the association response layer is the top layer), the three-layer data is superimposed through an image fusion algorithm. Specifically, the binary matrix of static obstacles in the spatial distribution layer (the obstacle area is 1, and the idle area is 0) is used as the base map, and a rasterized rendering is used to generate the bottom layer framework; subsequently, the living body movement trajectory data in the dynamic trajectory layer (such as a B-spline curve path) is superimposed on the middle layer through a semi-transparent color band (such as a 50% transparency in the blue channel) to distinguish dynamic and static obstacles; finally, the heat distribution map in the association response layer (such as a pseudo-color mapping: red indicates high risk, and yellow indicates early warning) is covered on the top layer through an Alpha blending algorithm to highlight potential interaction behavior areas. During the fusion process, coordinate alignment and pixel-level synthesis of multi-layer data are performed based on the OpenGL rendering engine, and finally, an obstacle topological map that includes geometric deformation features (such as contour positions, broken boundaries) and biological heat features (such as temperature intensity, periodic fluctuations) is generated, and it supports real-time query of the attribute data of each layer through an interactive interface (such as clicking on a node to obtain the temperature value, deformation curvature, and movement trajectory history).

[0121] In summary, through steps 301 to 304, the intelligent upgrade of obstacle dynamic modeling and behavior prediction in complex environments is realized. Through multi-modal data fusion and spatio-temporal feature analysis, the system establishes a three-dimensional obstacle representation system that includes three-dimensional spatial distribution, dynamic movement trajectories, and potential interaction behaviors. This technology significantly improves the accuracy of obstacle recognition and the depth of environmental understanding in complex scenarios, especially demonstrating excellent perception capabilities in complex scenarios where dynamic targets interact with static environments.

[0122] To solve the problems of mismatched multi-modal perception data acquisition range and insufficient dynamic target detection accuracy of mobile devices in complex environments, in some embodiments, in step 102, the cross-modal acquisition mode of the multi-spectral sensor is activated in the visual coordinate system, and the visible light and near-infrared dual-band overlapping perception regions are divided according to the forward moving region corresponding to the mobile device, where the region covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection, and the region covered by the near-infrared band is used to detect the temperature anomaly region related to biological signs, including:

[0123] 401. In the visual coordinate system, determine the forward moving region according to the moving direction of the mobile device, and form a fan-shaped coverage region with the current position of the device as the vertex along the moving direction in the forward moving region, and adjust the radius length of the fan-shaped coverage region according to the moving speed;

[0124] In step 401, the visual coordinate system refers to a local coordinate reference system with the current position of the device as the origin and the Z-axis pointing in the moving direction, such as a three-dimensional space coordinate system (error ±0.1m) updated in real time by the SLAM algorithm. The forward moving region is the detection region directly in front of the moving direction of the device, with a horizontal opening angle of 60°, and the depth range is dynamically adjusted according to the moving speed. The fan-shaped coverage region is a fan-shaped space with the device position as the vertex and extending along the moving direction, and its radius length is linearly related to the moving speed (for example, when the speed is 1m / s, the radius is 5m, and the radius increases by 2m for every 0.5m / s increase in speed). The moving speed is a real-time speed parameter calculated by fusing the accelerometer and GPS, such as the current speed is 2.3m / s.

[0125] In the embodiments of the present application, first, based on the visual coordinate system (the origin is the current position of the device, and the X / Y axes are aligned with the physical space directions) generated in step 101, the forward moving region is determined through the heading angle of the mobile device (such as θ = 30°). With the current position of the device as the vertex, a fan-shaped coverage region is extended along the moving direction (θ ± 60°), and its initial radius is set by a preset value (such as 5m). Subsequently, according to the real-time speed of the mobile device (such as v = 2m / s), the fan radius is dynamically adjusted through a linear interpolation formula: radius R = R0×(1 + 0.2v) (R0 is the initial radius, and the coefficient 0.2 is the speed scaling factor). For example, when the speed increases to 3m / s, the radius expands to 5×(1 + 0.2×3) = 8m. This adjustment mechanism ensures that a farther region is covered during high-speed movement to sense obstacles in advance.

[0126] 402. In the adjusted fan-shaped coverage region, activate the dual-band synchronous acquisition mode of the visible light sensor and the near-infrared sensor, so that their physical acquisition ranges completely cover the fan-shaped coverage region, where the visible light sensor focuses on the continuous frame contour changes of the floor projection plane, and the near-infrared sensor focuses on the thermal intensity gradient changes between adjacent pixels;

[0127] In step 402, the adjusted sector coverage area refers to a sector detection range whose radius changes dynamically with speed (e.g., when the speed = 3 m / s, the radius = 11 m). The visible light sensor is a device for collecting images in the 400 - 700 nm band (resolution 1920×1080, frame rate 30 Hz). The near-infrared sensor is a device for collecting the thermal radiation intensity in the 700 - 1000 nm band (resolution 640×480, frame rate 30 Hz). The dual-band synchronous acquisition mode is a data synchronization mechanism for aligning the timestamps of the visible light and near-infrared sensors (synchronization error ≤ 1 ms). The physical acquisition range is the physical space area actually covered by the sensor (e.g., the visible light horizontal viewing angle is 70°, and the near-infrared viewing angle is 65°). The continuous frame contour change of the floor projection plane is a sequence of the projected shapes of ground objects extracted from visible light images (e.g., the rectangular contour shifts by 0.2 m between two frames). The thermal intensity gradient change is the difference in thermal radiation between adjacent pixels in the near-infrared data (e.g., a gradient ≥ 3 °C / pixel is marked as a mutation point).

[0128] In the embodiment of the present application, within the sector coverage area generated in step 401, the acquisition timings of the visible light sensor (400 - 700 nm) and the near-infrared sensor (800 - 1400 nm) are synchronized through a hardware trigger signal to ensure that their physical fields of view completely cover the sector area. The visible light sensor collects continuous images of the floor projection plane at a frame rate of 30 fps, and extracts the contour change through the background difference algorithm (e.g., edge displacement = 0.2 m); the near-infrared sensor collects thermal radiation data at the same frame rate, and calculates the thermal intensity gradient between adjacent pixels through the Sobel operator (e.g., a gradient value > 10 °C / pixel is determined as abnormal). The data of the two sensors are aligned by timestamps and mapped into a sector area grid map in the same visual coordinate system.

[0129] 403. Map the continuous frame contour change collected by the visible light sensor into geometric deformation features, calculate the position coordinates of the deformation boundary through the displacement difference between the contour boundaries of adjacent frames, and at the same time map the thermal intensity gradient change collected by the near-infrared sensor into a temperature anomaly area, and extract the boundary coordinates of the temperature anomaly area through the spatial distribution of thermal intensity mutation points;

[0130] In step 403, the continuous frame contour change refers to the shape displacement of the ground object projection in the visible light image sequence (for example, the aspect ratio change rate of the trapezoidal contour between two frames is 0.1). The geometric deformation feature is the deformation parameter calculated from the contour vertex coordinates (for example, the average vertex displacement difference = 0.15 m). The position coordinates of the deformation boundary are the set of contour vertices of the geometric deformation region (such as the rectangle vertex coordinates [x1, y1], [x2, y2],...). The thermal intensity mutation point refers to the pixel point coordinates in the near-infrared data where the thermal gradient exceeds the threshold (such as ≥5 °C / pixel). The boundary coordinates of the temperature anomaly region are the set of polygon vertices generated by clustering the mutation points (such as the clustering radius is 0.3 m and the minimum number of points is 10).

[0131] In the embodiment of the present application, first, for the continuous frame contour data of the visible light sensor, the displacement difference of the contour boundary between adjacent frames is calculated by the optical flow method (Lucas-Kanade algorithm) (such as ), and the displacement difference is mapped to the visual coordinate system based on the homography matrix to generate the absolute position coordinates of the geometric deformation boundary (such as (x = 3.2, y = 4.5)). Secondly, for the thermal intensity gradient data of the near-infrared sensor, the region growing algorithm is used to start from the gradient mutation point (such as the gradient > 8 °C / pixel) and expand outward to the pixels with the gradient value lower than the threshold (such as 2 °C / pixel) to generate the closed polygon boundary of the temperature anomaly region, and its boundary coordinates are calculated by the centroid coordinate method (such as the vertex set [(3.0, 4.3), (3.5, 4.8)]).

[0132] 404. Based on the position coordinates and the boundary coordinates, divide the visible light and near-infrared overlapping perception regions within the adjusted sector coverage area.

[0133] In step 404, the position coordinates are the contour vertex data of the geometric deformation region in the visible light mode (such as the polygon vertex sequence). The boundary coordinates are the contour vertex data of the temperature anomaly region in the near-infrared mode (such as the outer rectangle vertex of the thermal image). The visible light and near-infrared overlapping perception region is the spatial intersection region of the coverage ranges of the two sensors (such as the overlapping area ratio ≥85%), and the coordinate alignment algorithm (such as ICP registration) is used to ensure the spatial consistency between the modes, with the error ≤0.05 m.

[0134] In the embodiments of the present application, first, the geometric deformation boundary position coordinates of the visible light sensor and the temperature anomaly region boundary coordinates of the near-infrared sensor are mapped to the grid map of the same sector coverage area, and the overlapping area between the two is calculated by the polygon intersection algorithm (such as the Weiler-Atherton algorithm). If the proportion of the intersection area in the area of any one region exceeds the threshold (such as 30%), the overlapping area is marked as the overlapping perception area of visible light and near-infrared, and a priority label is assigned (such as priority 1 is the dual-modal confirmation area). Finally, the boundary coordinate set (such as the polygon vertex sequence) of the overlapping area and its priority attribute are output for subsequent cross-modal fusion processing.

[0135] In summary, through steps 401 to 404, the intelligent dynamic optimization of the forward environmental perception area of the mobile device and the precise alignment of multi-modal data are realized. This technology breaks through the spatial alignment problem in traditional multi-sensor data fusion, enabling pixel-level position association between geometric contour changes and thermal anomaly regions, and providing an accurate spatial benchmark for subsequent cross-modal feature fusion. This method of dynamic perception area adjustment and precise data alignment significantly improves the capture ability and perception accuracy of the mobile device for environmental features in complex motion states.

[0136] To solve the problems of inaccurate prediction of the movement of living obstacles and conflict in obstacle avoidance decisions by the mobile device in a dynamic environment, in some embodiments, in step 105, generating the path planning parameters of the mobile device based on the periodic movement law of the living obstacle and the offset trend of the collision risk boundary includes:

[0137] 501. Extract periodic parameters based on the periodic movement law of the living obstacle and calculate the prediction range of its movement trajectory, and generate a first avoidance direction adjustment amount according to the geometric relationship between the prediction range of the movement trajectory and the current position of the mobile device;

[0138] In step 501, the periodic movement law of the living obstacle refers to the biometric parameters extracted by near-infrared data spectrum analysis, such as a walking frequency of 1.2 Hz or a breathing frequency of 0.3 Hz. The periodic parameters include the movement period (such as period T = 0.83 s) and the amplitude (such as displacement amplitude ±0.5 m). The prediction range of the movement trajectory is the future position distribution area calculated based on the periodic parameters. For example, the fan-shaped area (radius 2 m, opening angle 90°) where the living body may move within the next 3 seconds is predicted by Kalman filtering. The geometric relationship is the relative position parameter between the current position of the mobile device and the prediction range. For example, the distance between the device and the boundary of the prediction area is 1.5 m and the azimuth deviation is 30°. The first avoidance direction adjustment amount is the path correction parameter generated according to the geometric relationship. For example, it deflects 15° to the left to avoid the prediction area.

[0139] In the embodiments of the present application, first, the periodic motion parameters of the living obstacle are extracted through the Fast Fourier Transform (FFT) (such as the pedestrian step frequency of 1.2 Hz and the limb swing amplitude of 0.5 m), and the motion trajectory range within the next 3 seconds is predicted based on the Kalman filtering algorithm to generate a set of trajectory points (such as ). Subsequently, the shortest distance between the current position of the mobile device and each point of the predicted trajectory is calculated through the Euclidean distance formula. If the shortest distance is less than the safety threshold (such as 1.5 m), a first avoidance direction adjustment amount is generated according to the direction of the line connecting the device position and the nearest point on the trajectory. For example, when the direction of the line connecting the device and the nearest point on the trajectory is the northeast direction, the heading angle correction amount is calculated through the arctangent function , and is mapped into the device control instruction to adjust the moving direction to avoid overlapping with the trajectory of the living obstacle.

[0140] 502. Calculate the angle between the projection offset rate of the geometric deformation feature and the moving direction based on the offset trend of the collision risk boundary. When the projection offset rate exceeds the preset threshold, determine the dynamic safety distance through the product of the angle and the moving speed and generate a second avoidance direction adjustment amount;

[0141] In step 502, the offset trend of the collision risk boundary refers to the linear regression result of the expansion rate of the static obstacle contour. For example, the boundary expands outward by 0.1 m per second. The projection offset rate is the displacement speed of the geometric deformation feature in the visual coordinate system. For example, the contour vertex moving rate is 0.3 m / s. The angle is the angle between the projection offset rate direction and the device moving direction. For example, the offset direction and the device forward direction form an angle of 45°. The preset threshold is the minimum rate for determining the offset risk. For example, avoidance is triggered when the rate ≥ 0.2 m / s. The dynamic safety distance is the buffer distance calculated through the product of the angle and the moving speed. For example, when the angle is 45° and the speed is 1 m / s, the safety distance = 1 m × sin(45°) ≈ 0.7 m. The second avoidance direction adjustment amount is the path correction parameter generated according to the dynamic safety distance. For example, deflect 20° to the right to maintain the safety distance.

[0142] In the embodiments of the present application, first, the projection offset rate of the geometric deformation feature (such as the collapsed wall contour) is calculated through the optical flow method (such as ), and the angle between its offset direction and the device moving direction is calculated using the vector dot product formula:

[0143]

[0144] where V1 is the device moving direction vector and V2 is the projection offset direction vector. When the offset rate exceeds the preset threshold (such as 0.2 m / s), the dynamic safety distance S is calculated through the formula Calculate (where K is the empirical coefficient 0.5 and v is the device speed). If S is less than the preset safety value (such as 1 m), then generate a second avoidance direction adjustment amount. Specifically, by scaling the included angle α proportionally (such as =α / 2) to determine the course angle correction direction. For example, when α = 60°, generate a correction instruction to turn left by 30° to make the device deviate from the geometric deformation area.

[0145] 503. Perform a direction conflict detection on the first avoidance direction adjustment amount and the second avoidance direction adjustment amount. If the direction included angle is less than the preset angle, then superimpose and generate a composite avoidance parameter. If the direction included angle is greater than or equal to the preset angle, then select the direction with a smaller dynamic safety distance as the dominant avoidance direction;

[0146] In step 503, the direction conflict detection is a process of determining whether the included angle between the first and second avoidance direction adjustment amounts will cause a path conflict. For example, the included angle between the two directions is 30°. The preset angle is the critical value for determining direction conflict. For example, the included angle threshold is set to 45°. The composite avoidance parameter is the vector superposition result of the two when the direction included angle is less than the threshold. For example, the first adjustment amount of turning left by 15° and the second adjustment amount of turning right by 20° are combined to be a 5° right turn. The dominant avoidance direction is to select the direction with lower risk when the direction included angle is greater than or equal to the threshold. For example, if the included angle between the two directions is 60° and the safety distance of the second direction is smaller (0.5 m < 0.7 m), then the second avoidance direction is preferentially adopted.

[0147] In the embodiment of the present application, first calculate the first avoidance direction adjustment amount and the second adjustment amount of the direction vector included angle: If it satisfies: Then through weighted average: Generate a composite avoidance parameter; if it satisfies: Then compare the dynamic safety distances and , select the smaller value direction as the dominant avoidance direction. For example, when (avoiding birds) and (avoiding swinging wires), if , then preferentially select as the dominant direction.

[0148] 504. Based on the composite avoidance parameter or the dominant avoidance direction, and in combination with the moving speed, generate the path planning parameter of the mobile device.

[0149] In step 504, the composite avoidance parameter is the weighted vector sum of two adjustment amounts. For example, the combined direction deflects by 10° and the speed decreases to 0.8 m / s. The dominant avoidance direction is the direct application of a single adjustment amount. For example, select a right turn of 20° and maintain a speed of 1 m / s. The path planning parameters include the direction deflection angle (such as 10°), the speed attenuation coefficient (such as decreasing to 70% of the original speed), and the path curvature limit (such as the radius of curvature ≥ 1 m).

[0150] In the embodiments of the present application, first, the dynamic window algorithm (DWA) is used to fuse the avoidance direction and speed constraints: According to a candidate heading angle set is generated (such as Δθ ± 10°), and the feasible speed range is calculated in combination with the maximum acceleration of the device (such as 1 m / s²) ( ). Subsequently, a B-spline curve is used to fit the optimal combination of the heading angle and speed to generate a sequence of path points (such as ), and the optimal path is selected based on a cost function (the farther away from the obstacle and the more stable the speed, the lower the cost). Finally, the target heading angle, speed curve, and steering command are output, and are sent to the actuator through ROS (Robot Operating System) to drive the device to travel along the planned path while maintaining a dynamic safety distance from the obstacle.

[0151] The following is a specific example:

[0152] In the dynamic obstacle avoidance scenario of a home service robot, the system realizes intelligent avoidance decision-making through multi-modal perception: When it detects that a pet dog (a living obstacle) swings left and right periodically at a speed of 0.8 m / s (the period is 1.2 s ± 0.15 s), its motion law parameters are extracted and the future 3-second trajectory prediction range (sector area ± 45°) is calculated to generate the first avoidance direction adjustment amount (deflect 25° to the right); at the same time, it is monitored that the geometric deformation characteristics of the left sofa leg (a static obstacle) shift outward at a rate of 0.3 m / s due to the risk of collision by the robot's side brush (the angle with the current moving direction is 50°), triggering the second avoidance direction adjustment amount (deflect 15° to the left). After the direction conflict detection, it is found that the angle between the two adjustment amounts is 40° < the preset threshold of 45°, so they are superimposed to generate the composite avoidance parameter (the combined direction deflects 10° to the right), and the safety distance (0.5 m) is dynamically calculated in combination with the current moving speed of 0.6 m / s. Finally, the path planning parameters including decelerating to 0.4 m / s and correcting the heading by 10° are output. This solution enables the robot to avoid living pets (prediction trajectory matching degree) and dynamic deformation obstacles (anti-collision success rate) synchronously within 1.2 seconds, reducing path detours compared to a single avoidance strategy.

[0153] In summary, through steps 501 to 504, intelligent obstacle avoidance and adaptive path planning of the mobile device in a complex dynamic environment are achieved. By integrating the motion trajectory prediction of living obstacles and the collision risk analysis of static obstacles, the system establishes a multi-dimensional collaborative decision-making mechanism, which can intelligently balance the requirements of dynamic and static obstacle avoidance. When a direction conflict is detected, the system automatically selects the optimal obstacle avoidance strategy based on the safety distance priority to ensure the best passage efficiency while guaranteeing safety. This technology significantly improves the autonomous navigation ability of the device in complex scenarios with both dynamic and static obstacles, and realizes an integrated intelligent decision-making of environmental perception, risk assessment, and path planning.

[0154] To solve the problems of inaccurate recognition of living obstacles and difficult assessment of the collision risk of static obstacles in a dynamic environment, in some embodiments, in step 104, in the obstacle topology map, extracting a connected domain that conforms to the periodic biological motion law in the temperature anomaly region as a living obstacle, and determining the collision risk boundary of the static obstacle according to the projection offset of the geometric deformation feature between consecutive frames, includes:

[0155] 601. Extract the periodic intensity fluctuation parameter of the temperature anomaly region from the obstacle topology map, calculate the fluctuation period through the intensity change curve of the temperature anomaly region in consecutive frames, and when the fluctuation period matches the preset biological motion period threshold range, mark the corresponding temperature anomaly region as a candidate region for the living obstacle;

[0156] In step 601, the periodic intensity fluctuation parameter of the temperature anomaly region refers to the regular parameter of the thermal radiation intensity changing with time extracted from the near-infrared sensor data, such as a periodic fluctuation of 1.2 times per second ± 0.1 Hz. The intensity change curve of the temperature anomaly region in consecutive frames is an intensity-time series generated from multi-frame thermal imaging data (such as a sampling rate of 30 Hz), such as 30 thermal intensity sampling points per second. The fluctuation period is the period value corresponding to the main frequency of the Fourier transform of the intensity change curve (such as the period T = 0.83 seconds). The preset biological motion period threshold range is the typical motion frequency interval of the living obstacle (such as the human walking frequency of 0.8 - 2.0 Hz, the animal breathing frequency of 0.2 - 0.5 Hz). The candidate region for the living obstacle is the temperature anomaly region whose fluctuation period conforms to the biological threshold, such as marked as a red highlighted block (confidence level ≥ 90%).

[0157] In the embodiment of the present application, first, time series sampling is performed on the intensity data of the temperature anomaly region in the obstacle topology map to generate an intensity change curve of consecutive frames (for example, the interval between each frame is 100 ms), and periodic fluctuation parameters (such as the main frequency of 1.2 Hz and the amplitude of ±5 °C) are extracted through fast Fourier transform (FFT). Subsequently, the calculated fluctuation period is matched with the preset biological motion period threshold range (such as the human walking frequency of 0.8 - 2 Hz and the animal breathing frequency of 0.3 - 1 Hz). If the main frequency falls within the threshold range, a living obstacle candidate label is assigned to the temperature anomaly region. For example, when the main frequency of temperature fluctuation in a certain region is 1.5 Hz, it is determined that it conforms to the human walking characteristics and is marked as a candidate region. During this process, the continuous frame data is segmented through a sliding window algorithm to eliminate noise interference (such as environmental temperature drift) to ensure the robustness of periodic parameter extraction.

[0158] 602. In the candidate region, analyze the continuity of the movement trajectory of the temperature anomaly region, and verify the biological motion law through the displacement direction and speed fluctuation mode of the centroid of the temperature anomaly region in consecutive frames. When the random change rate of the displacement direction exceeds the preset threshold and the fluctuation amplitude of the speed conforms to the biological acceleration characteristics, confirm that the candidate region is a living obstacle;

[0159] In step 602, the candidate region refers to the suspected living region preliminarily screened in step 601 (such as a circular region with a radius of 0.5 meters). The continuity of the movement trajectory of the temperature anomaly region is calculated by the motion smoothness index (such as the standard deviation of displacement direction change ≤ 10°) of the displacement sequence of the centroid of the region in consecutive frames (such as the centroid coordinates [x1, y1] → [x2, y2] → [x3, y3]). The displacement direction of the centroid is the azimuth angle of the line connecting the centroids between adjacent frames (such as the direction angle changes from 30° to 35° between two frames, and the change rate is 5° / frame). The speed fluctuation mode is the standard deviation of the centroid movement speed (such as the fluctuation amplitude of the speed sequence [0.8 m / s, 1.0 m / s, 0.9 m / s] is ±0.1 m / s). The verification of the biological motion law is through the comprehensive determination of the randomness of the displacement direction (such as the direction change rate ≥ 5° / frame) and the speed fluctuation amplitude (such as conforming to the human walking acceleration of ±0.3 m / s²). The random change rate of the displacement direction is the variance of the direction angle change per unit time (such as triggering the living body determination when the variance ≥ 15°). The biological acceleration characteristic is the unique speed change mode of living body movement (such as the acceleration peak conforms to a sine curve, R² ≥ 0.8).

[0160] In the embodiment of the present application, for the candidate region marked in step 601, first, the centroid position of the temperature anomaly region in consecutive frames is tracked through the optical flow method (Lucas - Kanade algorithm), and the displacement vector is calculated (such as )and instantaneous velocity (such as v = 0.3 m / s). Subsequently, the random change rate of the displacement direction (such as the standard deviation of the direction angle σ = 25° within 10 frames) and the velocity fluctuation amplitude (such as the root mean square acceleration of 0.1 m / s²) are statistically calculated and compared with the biological motion feature library: If the direction change rate exceeds the preset threshold (such as σ > 20°) and the acceleration conforms to the biological characteristics (such as the human walking acceleration of 0.1 - 0.5 m / s²), it is determined as a living obstacle. During this process, the centroid trajectory is smoothed through Kalman filtering to eliminate the error caused by sensor jitter, and finally the topological map label is updated to confirm the candidate area as a living obstacle.

[0161] 603. Synchronize the projection offset of the geometric deformation feature in consecutive frames, calculate the included angle relationship between the deformation contour movement trajectory of the static obstacle and the movement direction of the mobile device. When the included angle between the projection offset direction and the device movement direction is less than the preset acute angle threshold, mark the area corresponding to the deformation contour as the collision risk boundary of the static obstacle.

[0162] In step 603, the projection offset of the geometric deformation feature refers to the position difference of the static obstacle in consecutive frames in the visual coordinate system (such as an offset of 0.2 meters between two frames). The deformation contour movement trajectory of the static obstacle is a displacement sequence constructed by the multi-frame projection offset (such as the trajectory points [0.1 m, 0.3 m, 0.5 m]). The movement direction of the mobile device is the azimuth angle of the device's current movement (such as 0° in the due north direction). The included angle relationship is the included angle between the deformation contour movement trajectory direction and the device movement direction (such as the trajectory direction is 30°, the device direction is 45°, and the included angle is 15°). The preset acute angle threshold is the critical angle for determining the conflict between the trajectory and the device direction (such as 30°). The projection offset direction refers to the vector direction of the deformation contour movement trajectory (such as the azimuth angle of 30°). The collision risk boundary is the safety buffer area of the static obstacle marked according to the included angle threshold (such as generating a red warning boundary with an expanded contour of 0.5 meters when the included angle < 30°).

[0163] In the embodiment of the present application, first, the projection offset of the geometric deformation feature in consecutive frames is calculated through the homography matrix (such as ), extract its movement trajectory direction vector (such as ), and obtain the device's own movement direction vector based on the IMU data of the mobile device (such as ). Calculate the included angle between the two through the vector included angle formula: , when θ is less than the preset acute angle threshold (such as 30°), it is determined that the deformation contour offset direction is close to coinciding with the device movement direction, and there is a collision risk. At this time, the morphological dilation algorithm is used to expand the deformation contour by 0.5 m to generate the collision risk boundary, and it is marked as the high-risk area of the static obstacle in the topological map. This process improves the adaptability to different scenarios through dynamic threshold adjustment (such as increasing the dilation distance according to the device speed).

[0164] The following is a specific example:

[0165] In the multi-obstacle recognition scenario of a domestic service robot, the system achieves precise classification through thermal-deformation multimodal analysis: First, extract the temperature anomaly area (such as a hot spot of 39.5°C ± 0.8°C) from the topological map, and analyze its intensity fluctuation curve through Fourier transform. When a periodic fluctuation of 1.2Hz ± 0.3Hz (matching the movement characteristics of feline animals) is detected, it is marked as a live candidate area; then analyze the movement trajectory of the centroid of the hot spot in the candidate area. If there is a direction mutation > 45° / 0.1s and a speed fluctuation of 0.3 - 0.7m / s (meeting the start-stop characteristics of pets) within 8 consecutive frames (133ms), it is confirmed as a live obstacle (confidence level 94%). At the same time, calculate the projection offset of the furniture deformation contour (such as a coffee table leg tilted by 5°) in the moving direction of the robot (0.5m / s, due east) through RGB-D data. When the westward movement rate of the deformation area is detected to be 0.05m / s (angle < 15°), it is marked as a static collision risk boundary and a virtual anti-collision area with an outward expansion of 8cm is generated. This solution improves the accuracy of live recognition and reduces the misjudgment rate of static obstacles. In the composite scenario of a pet suddenly dashing out (speed 0.8m / s) and furniture displacement (0.1m / s), the system can complete the classification response within 80ms.

[0166] In summary, through steps 601 to 603, the mobile device realizes the intelligent recognition and risk assessment of dynamic and static obstacles in a complex environment. The system establishes a dual verification mechanism based on biological movement characteristics and geometric deformation characteristics through the fusion analysis of multimodal perception data, and can accurately distinguish between live obstacles and static obstacles. This technology significantly improves the accuracy and reliability of obstacle recognition. Especially in complex scenarios with a mixture of dynamic and static obstacles, it can provide the mobile device with accurate environmental awareness and safety warning capabilities, effectively ensuring the safety and stability of autonomous navigation.

[0167] To solve the problems of inaccurate steering control and insufficient motion jitter compensation during the dynamic obstacle avoidance process of a mobile device, in some embodiments, generating the path planning parameters of the mobile device based on the composite avoidance parameter or the dominant avoidance direction and combining with the moving speed in step 504 includes:

[0168] 701. Generate the expected steering angle of the mobile device according to the angle between the avoidance direction in the composite avoidance parameter or the dominant avoidance direction and the current moving direction, and at the same time adjust the moving speed based on the speed decay coefficient in the composite avoidance parameter to generate the target moving speed in the avoidance state;

[0169] In step 701, the composite avoidance parameter is the path correction parameter synthesized after the direction conflict detection in step 503, for example, a vector including a steering angle of 10° and a speed attenuation coefficient of 0.7. The dominant avoidance direction is a single avoidance direction selected after the conflict detection, for example, turning right by 20°. The angle between the avoidance direction and the current moving direction is the deflection angle of the correction direction from the original moving direction of the device (e.g., the angle = 15°). The expected steering angle is the device steering instruction parameter generated according to the angle (e.g., turning left by 15°). The speed attenuation coefficient is the deceleration ratio calculated according to the dynamic safety distance (e.g., the original speed of 1 m / s is reduced to 0.7 m / s). The target moving speed is the adjusted real-time speed in the avoidance state (e.g., 0.7 m / s ± 0.1 m / s).

[0170] In the embodiment of the present application, first, the angle between the composite avoidance direction (or the dominant avoidance direction) and the current moving direction is calculated through the vector angle formula:

[0171]

[0172] where is the avoidance direction vector, is the current moving direction vector. According to the angle the expected steering angle is generated: And the current speed is adjusted through the speed attenuation coefficient (e.g., ) to generate the target moving speed: For example, when the current speed is and the angle , the expected steering angle is , and the target speed is reduced to . This process ensures that the avoidance action is smooth and conforms to the kinematic constraints of the device by dynamically adjusting the steering and speed.

[0173] 702. Perform direction-speed component association on the expected steering angle and the target moving speed, and combine them to generate a steering curvature parameter;

[0174] In step 702, the expected steering angle is the input parameter for device steering control (e.g., the steering angle of the steering gear is 15°). The target moving speed is the adjusted real-time speed value (e.g., 0.7 m / s). The direction-speed component association is to decompose the speed into components in the steering direction (e.g., the tangential speed component along the steering angle of 15° is 0.68 m / s, and the normal component is 0.18 m / s). The steering curvature parameter is the theoretical turning radius calculated according to the steering angle and the speed (e.g., the curvature radius R = 1 m / s² / (0.7 m / s × tan15°) ≈ 3.8 m).

[0175] In the embodiment of the present application, first, the expected steering angle Associated with the target moving speed Calculate the steering curvature parameter through the curvature formula: where is the control period (such as seconds). For example, when (converted to radians: 0.42 rad), then: .

[0176] Subsequently, truncate the curvature exceeding the threshold through the curvature limit condition (such as ) to ensure that the steering action is within the allowable range of the device's mechanical structure. This process generates curvature parameters under physical constraints, providing a guarantee of geometric feasibility for path planning.

[0177] 703. Integrate the steering curvature parameter and the moving speed to generate path planning parameters including speed adjustment instructions and steering control instructions;

[0178] In step 703, the steering curvature parameter is a geometric constraint parameter for path planning (such as the curvature radius ≥ 1.5 m). The moving speed is the current actual speed of the device (such as 0.7 m / s). The speed adjustment instruction is a pulse width modulation (PWM) signal for controlling the motor output (such as a 30% reduction in the duty cycle). The steering control instruction is a control command for driving the steering mechanism (such as a steering servo PWM signal pulse width of 1500 μs). The path planning parameter is a set including speed and steering instructions (such as JSON format instruction: {"speed": 0.7, "steer_angle": 15}).

[0179] In the embodiment of the present application, first, based on the steering curvature parameter C, generate a sequence of path points through a B-spline curve (such as an arc path corresponding to the curvature radius R = 1 / C), and combine it with the target moving speed to generate a speed curve (such as the tangential speed along the path). Subsequently, fuse the steering and speed constraints through the dynamic window algorithm (DWA) to screen the optimal combination of path points and speed, and finally generate path planning parameters, including a sequence of target heading angles (such as ), a sequence of speed instructions (such as ), and a sequence of steering angles (such as ). These parameters are sent to the drive controller through ROS (Robot Operating System) to drive the device to perform avoidance actions according to the planned path.

[0180] 704. Decompose the device's moving direction into horizontal and vertical components based on the motion trajectory data, generate a horizontal jitter compensation factor according to the projection value and direction angle deviation of the moving speed of the horizontal component, and generate a vertical jitter compensation factor according to the projection value and direction angle deviation of the moving speed of the vertical component;

[0181] In step 704, the motion trajectory data is a sequence of spatio-temporal coordinates during the movement of the device (such as 10 sets of data per second). The horizontal component is the projection of the device's movement direction on the X-axis in a two-dimensional plane (such as velocity x = 0.8 m / s, direction angle θ = 30°). The vertical component is the projection of the device's movement direction on the Y-axis in a two-dimensional plane (such as velocity y = 0.6 m / s). The projected value of the movement speed is the magnitude of the component of the velocity in the horizontal or vertical direction (such as horizontal projection 0.8 m / s, vertical projection 0.6 m / s). The direction angle deviation is the deflection angle between the actual movement direction of the device and the target direction (such as deviation angle = 5°). The horizontal jitter compensation factor is a correction parameter calculated based on the horizontal projection speed and the angle deviation (such as horizontal compensation factor = 0.8 × cos5° ≈ 0.796). The vertical jitter compensation factor is a correction parameter for the vertical projection speed and the angle deviation (such as vertical compensation factor = 0.6 × sin5° ≈ 0.052).

[0182] In the embodiment of the present application, first, the device motion trajectory data is kinematically decomposed into a horizontal component (X-axis velocity , direction angle deviation ) and a vertical component ( -axis velocity , direction angle deviation ).

[0183] The horizontal jitter compensation factor is calculated by the formula: , and the vertical jitter compensation factor is calculated by the formula: .

[0184] For example: when the horizontal velocity and the direction deviation (converted to radians: 0.087 rad): ;

[0185] When the vertical velocity and (converted to radians: 0.052 rad): , this process provides a quantitative basis for jitter compensation by quantifying the coupled influence of the direction deviation and the velocity.

[0186] 705. Generate a compensation coefficient through the non-linear fusion of the horizontal jitter compensation factor and the vertical jitter compensation factor.

[0187] In step 705, the horizontal jitter compensation factor is a dynamic parameter (such as 0.796) that offsets the horizontal jitter of the device. The vertical jitter compensation factor is a dynamic parameter (such as 0.052) that offsets the vertical jitter. Nonlinear fusion combines the horizontal and vertical compensation factors through a weighted sum of squares formula (such as compensation coefficient = ). The compensation coefficient is the comprehensive correction parameter finally used to stabilize the visual coordinate system (such as horizontal compensation 0.78, vertical compensation 0.12).

[0188] In the embodiment of the present application, first, a nonlinear weighted fusion algorithm is used to combine the horizontal jitter compensation factor with the vertical jitter compensation factor into a global compensation coefficient , where = 0.6, = 0.4 are weight coefficients, and λ = 0.2 is an interaction term coefficient. For example, when = 0.92, = 0.94, = 0.6×0.92 + 0.4×0.94 + 0.2×0.92×0.94 ≈ 0.93. This coefficient is used to adjust the control gain of the device actuator (such as motor torque, steering servo response speed) to suppress the trajectory deviation caused by motion jitter.

[0189] In summary, through steps 701 to 705, an integrated solution for intelligent motion planning and stability control of a mobile device during dynamic obstacle avoidance is achieved. By establishing a collaborative optimization mechanism for steering angle and speed adjustment, the system can generate a smooth obstacle avoidance trajectory that takes into account both safety and motion efficiency in real time. This technology not only realizes the precise coordination of direction control and speed control to ensure the smooth execution of obstacle avoidance actions, but also can adaptively adjust the compensation parameters according to the real-time motion state, significantly improving the motion stability and control accuracy in complex obstacle avoidance scenarios, providing a reliable motion control guarantee for the autonomous navigation of mobile devices.

[0190] Figure 2 FIG. shows a schematic structural diagram of an optimized system for room obstacle detection based on a cross-modal attention mechanism provided by an embodiment of the present application. As Figure 2 shown, the system includes:

[0191] An acquisition module 21 that acquires the motion trajectory data of the mobile device, calculates the compensation coefficient of the device moving direction and the viewing angle jitter amplitude based on the motion trajectory data, and generates a visual coordinate system that matches the physical space;

[0192] The acquisition module 22 activates the cross-modal acquisition mode of the multispectral sensor in the visual coordinate system, and divides the visible light and near-infrared dual-band overlapping perception area according to the moving forward area corresponding to the mobile device. The area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection object, and the area covered by the near-infrared band is used to detect the temperature anomaly area related to biological signs;

[0193] The adjustment module 23 inputs the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly area into the cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, so as to generate an obstacle topology map;

[0194] The extraction module 24 extracts the connected domains that conform to the periodic biological movement law in the temperature anomaly area in the obstacle topology map as living obstacles, and determines the collision risk boundary of the static obstacle according to the projection offset amount of the geometric deformation characteristics between consecutive frames;

[0195] The generation module 25 generates the path planning parameters of the mobile device based on the periodic movement law of the living obstacle and the offset trend of the collision risk boundary.

[0196] Figure 2 The described room obstacle detection optimization system based on cross-modal attention mechanism can execute Figure 1 The described room obstacle detection optimization method based on cross-modal attention mechanism in the illustrated embodiment, its implementation principle and technical effects will not be elaborated. For the room obstacle detection optimization system based on cross-modal attention mechanism in the above embodiment, the specific ways for each module and unit to execute operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0197] In a possible design, Figure 2 The room obstacle detection optimization system based on cross-modal attention mechanism in the illustrated embodiment can be implemented as a computing device, such as Figure 3 shown, this computing device may include a storage component 31 and a processing component 32;

[0198] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32.

[0199] The processing component 32 is used for the Figure 1 room obstacle detection optimization method based on cross-modal attention mechanism in the above

[0200] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.

[0201] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0202] Of course, the computing device may also necessarily include other components, such as input / output interfaces, display components, communication components, etc.

[0203] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module may be an output device, an input device, etc.

[0204] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.

[0205] Among them, the computing device may be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device may refer to a cloud server, and the above processing component, storage component, etc. may be basic server resources leased or purchased from a cloud computing platform.

[0206] The embodiment of the present application also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 An optimization method for room obstacle detection based on a cross-modal attention mechanism shown in the embodiment.

[0207] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0208] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0209] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0210] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An optimized method for room obstacle detection based on a cross-modal attention mechanism, characterized in that Including: Obtaining the motion trajectory data of a mobile device, calculating a compensation coefficient for the device's moving direction and the perspective jitter amplitude based on the motion trajectory data, and generating a visual coordinate system matching the physical space; Activating a cross-modal acquisition mode of a multispectral sensor in the visual coordinate system, dividing a visible light and near-infrared dual-band overlapping perception area according to the moving forward area corresponding to the mobile device, wherein the area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection object, and the area covered by the near-infrared band is used to detect the temperature anomaly area related to biological signs; Inputting the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly area into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, so as to generate an obstacle topology map; In the obstacle topology map, extracting the connected domain in the temperature anomaly area that conforms to the periodic biological motion law as a living obstacle, and determining the collision risk boundary of the static obstacle according to the projection offset amount of the geometric deformation characteristics between consecutive frames; Generating path planning parameters for the mobile device based on the periodic motion law of the living obstacle and the offset trend of the collision risk boundary; 2. The method according to claim 1, wherein Inputting the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly area into a cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities according to the moving speed and direction change amount of the mobile device through the cross-modal attention network, so as to generate an obstacle topology map, including: Inputting the compensation coefficient, the geometric deformation characteristics, and the temperature anomaly area into a cross-modal attention network, and generating a dynamic allocation parameter based on the moving speed and direction change amount of the mobile device through the cross-modal attention network; Calculating the weight distribution ratio of the visible light data stream and the near-infrared data stream through the dynamic allocation parameter, and performing weighted fusion on the geometric deformation characteristics and the temperature anomaly area according to the weight distribution ratio; Performing cross-modal interaction on the weighted geometric deformation characteristics and the temperature anomaly area, generating a deformation heat correlation mark by overlapping and comparing the geometric deformation contour and the temperature anomaly area in the spatial dimension, and constructing a dynamic displacement chain by combining the relative offset of the deformation contour and the temperature anomaly area in consecutive frames with the displacement correction factor of the compensation coefficient in the time dimension; Generating an obstacle topology map integrating biological characteristics and deformation based on the dynamic displacement chain and the deformation heat correlation mark; 3. The method according to claim 2, characterized in that, Generating a dynamic topology map of an obstacle integrating biological characteristics and deformation based on the dynamic displacement chain and the deformation heat correlation mark, including: When the direction deviation between the displacement direction of the geometric deformation contour and the migration direction of the temperature anomaly area in the dynamic displacement chain is less than a preset threshold, extracting synchronous offset trajectory data from the dynamic displacement chain to construct a dynamic trajectory layer of the living obstacle; Based on the spatial superposition relationship between the geometric deformation contour and the temperature anomaly region, when the projection boundary of the geometric deformation contour coincides with the outer boundary of the temperature anomaly region, it is marked as the spatial distribution layer of static obstacles; Based on the fracture characteristics of the geometric deformation contour and the periodic intensity fluctuation of the temperature anomaly region in the deformation-thermal correlation marking, when the fracture position matches the phase of the fluctuation period, the correlation response layer of potential interaction behaviors is generated through the coverage relationship between the extended region of the fracture boundary and the thermal fluctuation region; The spatial distribution layer, the dynamic trajectory layer and the correlation response layer are superimposed in the order of priority to form an obstacle topology map integrating geometric deformation features and biothermal features.

4. The method according to claim 1, wherein In the visual coordinate system, activate the cross-modal acquisition mode of the multispectral sensor. According to the moving forward region corresponding to the mobile device, divide the visible light and near-infrared dual-band overlapping perception regions, where the region covered by the visible light band is used to detect the geometric deformation features of the ground projection, and the region covered by the near-infrared band is used to detect the temperature anomaly region related to biological signs, including: In the visual coordinate system, determine the moving forward region according to the moving direction of the mobile device, and in the moving forward region, form a fan-shaped coverage region with the current position of the device as the vertex along the moving direction, and adjust the radius length of the fan-shaped coverage region according to the moving speed; In the adjusted fan-shaped coverage region, activate the dual-band synchronous acquisition mode of the visible light sensor and the near-infrared sensor, so that the physical acquisition ranges of the two completely cover the fan-shaped coverage region, where the visible light sensor focuses on the continuous frame contour change of the floor projection surface, and the near-infrared sensor focuses on the thermal intensity gradient change between adjacent pixels; Map the continuous frame contour change collected by the visible light sensor into geometric deformation features, calculate the position coordinates of the deformation boundary through the displacement difference of the contour boundaries of adjacent frames, and at the same time map the thermal intensity gradient change collected by the near-infrared sensor into the temperature anomaly region, and extract the boundary coordinates of the temperature anomaly region through the spatial distribution of the thermal intensity mutation points; Based on the position coordinates and the boundary coordinates, divide the visible light and near-infrared overlapping perception regions in the adjusted fan-shaped coverage region.

5. The method according to claim 1, wherein, Based on the periodic motion law of the living obstacle and the offset trend of the collision risk boundary, generate the path planning parameters of the mobile device, including: Extract the periodic parameters based on the periodic motion law of the living obstacle and calculate the prediction range of its motion trajectory, and generate the first avoidance direction adjustment amount according to the geometric relationship between the prediction range of the motion trajectory and the current position of the mobile device; Calculate the angle between the projection offset rate of the geometric deformation feature and the moving direction based on the offset trend of the collision risk boundary. When the projection offset rate exceeds the preset threshold, determine the dynamic safety distance through the product of the angle and the moving speed and generate the second avoidance direction adjustment amount; Perform direction conflict detection on the first avoidance direction adjustment amount and the second avoidance direction adjustment amount. If the direction angle is less than the preset angle, generate a composite avoidance parameter by superposition. If the direction angle is greater than or equal to the preset angle, select the direction with a smaller dynamic safety distance as the dominant avoidance direction; Based on the composite avoidance parameter or the dominant avoidance direction, and in combination with the moving speed, generate the path planning parameter of the mobile device.

6. The method according to claim 1, wherein In the obstacle topology map, extract the connected domain that conforms to the periodic biological motion law in the temperature anomaly region as a living obstacle, and determine the collision risk boundary of the static obstacle according to the projection offset amount of the geometric deformation feature between consecutive frames, including: Extract the periodic intensity fluctuation parameter of the temperature anomaly region from the obstacle topology map, calculate the fluctuation period through the intensity change curve of the temperature anomaly region in consecutive frames. When the fluctuation period matches the preset biological motion period threshold range, mark the corresponding temperature anomaly region as the candidate region of the living obstacle; In the candidate region, analyze the continuity of the movement trajectory of the temperature anomaly region, and verify the biological motion law through the displacement direction and speed fluctuation mode of the centroid of the temperature anomaly region in consecutive frames. When the random change rate of the displacement direction exceeds the preset threshold and the fluctuation amplitude of the speed conforms to the biological acceleration feature, confirm that the candidate region is a living obstacle; Synchronously based on the projection offset amount of the geometric deformation feature in consecutive frames, calculate the included angle relationship between the deformation contour movement trajectory of the static obstacle and the movement direction of the mobile device. When the included angle between the projection offset direction and the device movement direction is less than the preset acute angle threshold, mark the region corresponding to the deformation contour as the collision risk boundary of the static obstacle.

7. The method according to claim 5, wherein Based on the composite avoidance parameter or the dominant avoidance direction, and in combination with the moving speed, generate the path planning parameter of the mobile device, including: According to the included angle between the avoidance direction in the composite avoidance parameter or the dominant avoidance direction and the current movement direction, generate the expected steering angle of the mobile device. At the same time, adjust the moving speed based on the speed attenuation coefficient in the composite avoidance parameter to generate the target moving speed in the avoidance state; Perform direction-velocity component association on the expected steering angle and the target moving speed, and combine them to generate the steering curvature parameter; Fuse the steering curvature parameter and the moving speed to generate the path planning parameter including the speed adjustment instruction and the steering control instruction; The calculating the compensation coefficient of the device movement direction and the view jitter amplitude based on the movement trajectory data includes: Decompose the device movement direction into a horizontal component and a vertical component based on the movement trajectory data, generate a horizontal jitter compensation factor according to the moving speed projection value and the direction angle deviation of the horizontal component, and generate a vertical jitter compensation factor according to the moving speed projection value and the direction angle deviation of the vertical component; Generate the compensation coefficient through the non-linear fusion of the horizontal jitter compensation factor and the vertical jitter compensation factor.

8. An optimized system for room obstacle detection based on a cross-modal attention mechanism, characterized in that, Including: An acquisition module, which acquires the movement trajectory data of the mobile device, calculates the compensation coefficient of the device movement direction and the view jitter amplitude based on the movement trajectory data, and generates a visual coordinate system that matches the physical space; The acquisition module activates the cross-modal acquisition mode of the multi-spectral sensor in the visual coordinate system, and divides the visible light and near-infrared dual-band overlapping perception area according to the moving forward area corresponding to the mobile device, where the area covered by the visible light band is used to detect the geometric deformation characteristics of the ground projection object, and the area covered by the near-infrared band is used to detect the temperature anomaly area related to the biological signs; The adjustment module inputs the compensation coefficient, the geometric deformation characteristics and the temperature anomaly area into the cross-modal attention network, so as to adjust the attention weight distribution between the visible light and near-infrared modalities through the cross-modal attention network according to the moving speed and direction change amount of the mobile device, so as to generate an obstacle topology map; The extraction module extracts the connected domain that conforms to the periodic biological movement law in the temperature anomaly area in the obstacle topology map as a living obstacle, and determines the collision risk boundary of the static obstacle according to the projection offset amount of the geometric deformation characteristics between consecutive frames; The generation module generates the path planning parameters of the mobile device based on the periodic movement law of the living obstacle and the offset trend of the collision risk boundary.

9. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an optimization method for room obstacle detection based on cross-modal attention mechanism according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that, A computer program is stored, and when the computer program is executed by a computer, it implements an optimization method for room obstacle detection based on cross-modal attention mechanism according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Obstacle speed detection method and device, computer equipment and storage medium

    CN117930220A

  • Multi-sensor fusion obstacle detection method

    CN117968860A