Indoor robot positioning method and system
By combining communication signals and multi-source data fusion positioning methods, a hierarchical optimization expression is constructed, achieving high-precision robot positioning in complex indoor environments. This solves the problem of insufficient positioning accuracy in existing technologies and meets the application needs of point-to-point services and cleaning robots.
Patent Information
- Application Number
- CN202511628912.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-03
AI Technical Summary
Existing indoor robot positioning methods lack accuracy in complex environments, making it difficult to meet the positioning needs of point-to-point service robots and cleaning robots, especially in situations with changing lighting conditions or occlusion, where positioning accuracy and robustness are insufficient.
By combining communication signals and multi-source data fusion for localization, the communication connection signals between the target robot and indoor detection equipment are captured, and combined with indoor layout and multimodal image data, a hierarchical optimization expression is constructed to achieve accurate determination of the initial coarse path and fine path.
A balance between robot positioning accuracy and real-time performance was achieved in complex indoor environments. The problems of low positioning accuracy of single signals, large signal fluctuations, and poor robustness of attitude measurement caused by lighting and occlusion were solved, meeting the needs of point-to-point service robot docking and cleaning robot traversal.
Smart Images

Figure CN121453028A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent positioning technology, and in particular to an indoor robot positioning method and system. Background Technology
[0002] Indoor robot localization is one of the core technologies for achieving autonomous navigation and task execution, especially in applications such as point-to-point service robots (e.g., food delivery and item delivery robots) and cleaning robots (e.g., automated sweeping and disinfection robots), where the requirements for positioning accuracy, path continuity, and environmental adaptability are even more stringent. Existing technologies for indoor robot localization mainly include wireless signal localization (such as WiFi fingerprinting and UWB), visual localization (such as SLAM), and multi-sensor fusion localization, but these methods have several shortcomings: Wireless signal positioning relies on signal strength and time difference, making it susceptible to indoor obstruction, multipath effects, and device deployment density. In complex layouts, its accuracy fluctuates significantly. Furthermore, it fails to adequately utilize multi-dimensional characteristics of communication signals, such as device identification and signal attenuation patterns from multiple directions, hindering the construction of high-precision initial paths. This can easily lead to task failure or inefficiency for point-to-point service robots that need to accurately reach service locations or cleaning robots that need to traverse cleaning areas.
[0003] Visual positioning: Limited by lighting conditions and scene texture richness, positioning drift is prone to occur in low light and textureless areas. Furthermore, relying solely on visual information is not robust enough to meet the continuous and reliable positioning requirements of cleaning robots in dim corners and service robots in complex lighting scenarios.
[0004] Multi-sensor fusion positioning: Although it integrates information from multiple sensors, the insufficient depth in the collaborative utilization of multimodal data leads to inadequate path construction accuracy, making it difficult to meet positioning accuracy requirements in complex indoor environments. For example, point-to-point service robots are prone to taking detours when delivering across areas due to insufficient path accuracy, and cleaning robots are prone to missing cleaning areas due to positioning deviations in complex furniture layouts.
[0005] The aforementioned shortcomings limit the positioning performance of indoor robots in complex scenarios. Therefore, there is an urgent need for an indoor robot positioning method to address these technical challenges and meet the practical application needs of point-to-point service robots and cleaning robots. Summary of the Invention
[0006] This invention provides an indoor robot positioning method and system to solve the aforementioned technical problems.
[0007] This invention provides an indoor robot positioning method, comprising: Step 1: Capture the communication connection signal between the communication component and the indoor detection device during the movement of the target robot, and determine the initial coarse path of the target robot by combining the deployment location of the indoor detection device and the indoor layout. The initial coarse path consists of several coarse position units, and the coarse position units include fine units, coarse frame units, and fuzzy units. The position accuracy of the fine units is greater than that of the coarse frame units, and the position accuracy of the coarse frame units is greater than that of the fuzzy units. The communication connection signal includes the unique identifier of the indoor detection device that the communication component connects to each time, the corresponding connection sensing signal strength, the signal attenuation characteristics under unobstructed conditions and under obstructed conditions in each specified direction. Step 2: Capture multimodal images of the target robot's motion behavior at each moment of movement based on the imaging device and motion posture data measured by the target robot's own sensors during the movement process, and obtain the fused posture at the same moment of movement; Step 3: Based on the fused pose of two adjacent coarse positions, the positional distance between the two adjacent coarse positions, and the motion pattern of the target robot, construct a global optimization expression for the coarse frame unit, a local optimization expression for the fine frame unit, and a correction expression for the fuzzy frame unit to obtain and output the fine path.
[0008] Preferably, determining the initial coarse path of the target robot includes: Collect the signal strength of the remaining indoor detection devices (excluding the indoor detection devices that have established a stable communication connection with the communication component) and the communication component, and classify the remaining indoor detection devices into stable devices that continuously detect signals and unstable devices that do not continuously detect signals. A dynamic network connection diagram is constructed by combining the deployment location of each indoor detection device, the indoor layout, the unique identifier of the connected device, the sensing signal strength, the signal attenuation characteristics, and the correlation between the signal strength of other stable devices and other unstable devices. Based on the dynamic network connection graph, the coarse positions are initially determined time-by-time. Multiple time-by-time coarse positions are then concatenated in a time sequence to form an initial coarse path composed of several coarse position units of varying sizes.
[0009] Preferably, before initially determining the coarse position based on the dynamic network connection graph time-by-time, the method further includes: The sensor signal strength of the communication components and various indoor detection devices of the target robot during its movement, the dynamic occlusion sequence of the indoor layout, and the motion acceleration trend of the communication components are collected. The intensity of the sensed signal is matched with the signal attenuation feature library of each indoor detection device in the time and frequency domain to generate a matching feature set containing the time domain stability and frequency domain distribution characteristics of signal attenuation. At the same time, combined with the motion acceleration trend and dynamic occlusion sequence, the orientation-motion-occlusion multi-dimensional constraint space of the target robot is constructed. Based on the multi-dimensional constraint space, matching feature set, and preset unit construction type determination system, the unit construction type at the current movement moment is determined, and the unit construction type is attached to the dynamic network connection graph in the form of dynamic attributes. The determination system integrates the second derivative change rate of signal intensity, the Markov transition probability of occlusion state, and the entropy value feature of motion acceleration as determination conditions for refined, coarse-framed, and fuzzy construction types, respectively.
[0010] Preferably, the coarse position is initially determined time-by-time based on the dynamic network connection graph of the additional construction type, including: Based on the dynamic network connection graph with additional dynamic attributes, a real-time mapping function of construction type-location contribution is constructed, and the topological adjacency matrix of indoor layout is introduced as a regularization term to establish a multi-objective optimization model. Based on the multi-objective optimization model, the coarse position of the target robot is determined time-by-time by employing high-precision localization based on sparse Bayesian learning for refined construction type units, range shrinkage based on adaptive particle filtering for coarse-frame construction type units, and temporal correlation prediction based on attention mechanism for fuzzy construction type units. The real-time mapping function weights the refined construction type by an exponential function of signal confidence, the coarse-frame construction type by a logarithmic function of spatial coverage, and the fuzzy construction type by a Gaussian function of temporal correlation.
[0011] Preferably, step 2 includes: The system captures multimodal images of the target robot's motion behavior at each moment of movement, captured by an imaging device, as well as multi-source motion posture data measured by the target robot's multi-axis inertial measurement unit and tactile sensor array. Cross-modal feature fusion is performed on the multimodal images of the motion behavior to extract an image feature set containing appearance semantic features, three-dimensional structural features, and dynamic thermal distribution features. Simultaneously, spatiotemporal calibration is performed on the multi-source motion posture data to obtain a sensor feature set containing acceleration time-domain spectrum, angular velocity frequency-domain features, and tactile pressure distribution matrix. Based on the multi-source motion posture data, the motion state of the target robot is obtained through machine learning classification. A dynamic fusion weight network for image feature set and sensor feature set is constructed. The dynamic fusion weight network assigns attention weights to image feature set according to motion mode confidence and fusion weights to sensor feature set according to consistency entropy value of multi-source data for different motion states. Based on the dynamic fusion weight network, multi-scale spatiotemporal alignment fusion of image feature sets and sensor feature sets is performed to obtain the fused pose at the same movement moment. The fused pose includes a visual-motion coupled pose matrix, action semantic labels, and fusion confidence of multi-source data.
[0012] Preferably, the construction includes a global optimization expression for the coarse-frame units, a local optimization expression for the fine-frame units, and a correction expression for the fuzzy-frame units, including: Multimodal feature decomposition is performed on the fused poses of two adjacent coarse positions to obtain a set of pose features including a six-dimensional pose vector, an action semantic probability distribution, and a multi-source data reliability matrix of the fused poses. The system simultaneously collects the positional distance between two adjacent coarse positions, resolves it into Euclidean distance and indoor topological correlation distance, and extracts the temporal gradient and spatial distribution entropy of the distance. At the same time, it identifies the motion pattern of the target robot. With the global spatial deviation norm of the attitude feature set, the product of Euclidean distance and dynamic weight of motion pattern, and the robust regularization term of the multi-source data reliability matrix as the core, an adaptive constraint factor based on motion pattern is introduced to form a global optimization expression with multiple constraints coupled. The local detail deviation of the pose feature set, the intersection of indoor topological association distance and dynamic weight of motion mode, and the semantic obstacle matrix of indoor layout are fused together. At the same time, the action semantic probability distribution of the fused pose is embedded as a weighting coefficient to form a local optimization expression that is strongly correlated with the indoor environment layout. Based on the first and second expressions of the nearest non-fuzzy unit of the corresponding fuzzy unit, and combined with the change vector of the adjacent coarse position of the corresponding fuzzy unit, the correction expression is obtained.
[0013] Preferably, the fine path is obtained and output, including: Multi-objective joint optimization is performed on the global optimization expression, local optimization expression and correction expression. A dynamic regularization term based on the gradient of the temporal change of location distance is introduced to achieve the coordinated convergence of coarse-framed units, fuzzy units and fine-framed units. The fine path is obtained based on the collaborative convergence results.
[0014] This invention provides an indoor robot positioning system, comprising: The coarse path determination module is used to capture the communication connection signals between the communication component and the indoor detection device during the movement of the target robot, and determine the initial coarse path of the target robot by combining the deployment location of the indoor detection device and the indoor layout. The initial coarse path is composed of several coarse position units, and the coarse position units include fine units, coarse frame units, and fuzzy units. The position accuracy of the fine units is greater than that of the coarse frame units, and the position accuracy of the coarse frame units is greater than that of the fuzzy units. The communication connection signals include the unique identifier of the indoor detection device that the communication component connects to each time, the corresponding connection sensing signal strength, the signal attenuation characteristics based on the unobstructed condition in each specified direction, and the signal attenuation characteristics based on the occluded condition. The fusion module is used to capture multimodal images of the target robot's actions at each moment of movement captured by the camera and motion posture data measured by the target robot's own sensors during the movement process, so as to obtain the fused posture at the same moment of movement. The expression refinement module is used to construct a global optimization expression for the coarse frame unit, a local optimization expression for the fine frame unit, and a correction expression for the fuzzy frame unit based on the fused pose of two adjacent coarse frames, the positional distance between the two adjacent coarse frames, and the motion pattern of the target robot, so as to obtain the fine path and output it.
[0015] Compared with the prior art, the beneficial effects of this application are as follows: By employing a three-level positioning logic—coarse positioning via communication signals, attitude acquisition through multi-source data fusion, and hierarchical path optimization—a balance between robot positioning accuracy and real-time performance in complex indoor environments is achieved. The initial coarse path is correlated with the indoor layout through multi-dimensional communication signals, addressing the issues of low positioning accuracy and large signal fluctuations caused by single signals. The attitude fusion, achieved through dynamic fusion of multimodal images and sensor data, addresses the problems of poor robustness in attitude measurement and weak light drift in visual positioning caused by lighting and occlusion. The hierarchical optimization expression is differentiated for different accuracy requirements, solving the problem of a lack of specificity in path optimization and meeting the core requirements of point-to-point service robot docking and cleaning robot traversal.
[0016] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.
[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of an indoor robot positioning method according to an embodiment of the present invention; Figure 2 This is a structural diagram of an indoor robot positioning system according to an embodiment of the present invention. Detailed Implementation
[0019] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0020] This invention provides an indoor robot positioning method, such as... Figure 1 As shown, it includes: Step 1: Capture the communication connection signal between the communication component and the indoor detection device during the movement of the target robot, and determine the initial coarse path of the target robot by combining the deployment location of the indoor detection device and the indoor layout. The initial coarse path consists of several coarse position units, and the coarse position units include fine units, coarse frame units, and fuzzy units. The position accuracy of the fine units is greater than that of the coarse frame units, and the position accuracy of the coarse frame units is greater than that of the fuzzy units. The communication connection signal includes the unique identifier of the indoor detection device that the communication component connects to each time, the corresponding connection sensing signal strength, the signal attenuation characteristics under unobstructed conditions and under obstructed conditions in each specified direction. Step 2: Capture multimodal images of the target robot's motion behavior at each moment of movement based on the imaging device and motion posture data measured by the target robot's own sensors during the movement process, and obtain the fused posture at the same moment of movement; Step 3: Based on the fused pose of two adjacent coarse positions, the positional distance between the two adjacent coarse positions, and the motion pattern of the target robot, construct a global optimization expression for the coarse frame unit, a local optimization expression for the fine frame unit, and a correction expression for the fuzzy frame unit to obtain and output the fine path.
[0021] In this embodiment, the target robot refers to an autonomous mobile robot that needs to achieve indoor positioning, including point-to-point service robots and cleaning robots.
[0022] In this embodiment, the communication component is integrated into the target robot. This hardware module, used to establish a communication connection with indoor detection equipment and transmit signals, needs to support signal strength acquisition, device identification, and signal attenuation characteristic recording. Specifically, an ESP32-C3 wireless communication module is selected, supporting WiFi 802.11b / g / n and Bluetooth 5.0 protocols. The signal acquisition frequency is 10Hz, and it can acquire the received signal strength indication value of the connected device in real time. The module has a built-in unique device identifier (MAC address) to achieve unique connection identification with the indoor detection equipment.
[0023] In this embodiment, the indoor detection device is deployed at a fixed location indoors. It is used to establish a connection with the target robot's communication components and provide a location reference. It must have stable communication capabilities and known deployment coordinates. Specifically, a TP-Link TL-WDR5620 wireless router is selected as the indoor detection device, supporting a WiFi signal coverage radius of 15m. During deployment, its coordinates in the indoor coordinate system are determined using a laser rangefinder. For example, device A is deployed at (10.0m, 5.0m, 2.5m), and device B is deployed at (20.0m, 5.0m, 2.5m). Each device has a built-in unique identifier (SSID bound to MAC address) that can be recognized by the robot's communication components.
[0024] In this embodiment, the communication connection signal is the signal data transmitted when the communication component establishes a connection with the indoor detection device. It includes three core pieces of information: the unique identifier of the indoor detection device that the communication component connects to each time, the corresponding sensing signal strength, and the signal attenuation characteristics based on unobstructed / obstructed conditions in each specified direction. Among them, the signal attenuation characteristics are obtained by testing the signal attenuation values of the communication component and the indoor detection device in four specified directions (east, south, west, and north) in an unobstructed area (such as an open conference room) and an obstructed area (such as an office separated by one brick wall). For example, when there is no obstruction, the signal attenuation is 8dB per 10m; when there is one brick wall, the signal attenuation is 25dB per 10m, and a signal attenuation feature library is constructed.
[0025] In this embodiment, the initial coarse path is a preliminary target robot movement path determined based on communication connection signals, the deployment location of indoor detection equipment, and the indoor layout. It is formed by concatenating several coarse position units in a time sequence, with a lower accuracy than the final fine path, serving as a basic framework for subsequent path refinement. The coarse position units are the basic position units constituting the initial coarse path, categorized into three types: refined units, coarse-framed units, and fuzzy units, with their positional accuracy decreasing sequentially. This is used to manage positioning accuracy differently based on the complexity of the indoor environment. Specifically: a refined unit is defined as a square area with a side length of 0.5m, used in areas with high accuracy requirements, such as delivery robot docking points and key cleaning areas for cleaning robots; a coarse-framed unit is defined as a square area with a side length of 2m, used in areas with medium accuracy requirements, such as indoor corridors and open office areas; and a fuzzy unit is defined as a square area with a side length of 5m, used in areas with weak signals and complex environments, such as storage rooms and multi-obstructed passageways, where the approximate location is determined through temporal correlation prediction.
[0026] In this embodiment, signal capture is initiated by activating the target robot, whose communication component scans indoor detection devices in real time, captures communication connection signals, and records the unique identification of the connected device, the strength of the sensed signal, and the signal attenuation characteristics in the corresponding direction once every 0.5 seconds. For example, at time t1, the connected device A has the identification TP-Link_88A1, RSSI-65dBm, and is unobstructed to the east. The attenuation characteristics are consistent with the unobstructed database data.
[0027] Data association involves associating the captured communication connection signals with the pre-stored deployment locations of indoor detection devices. For example, given a layout diagram of device A (10.0, 5.0), including a corridor width of 1.5m and wall positions of x=8m and x=22m, based on RSSI-65dBm and unobstructed attenuation characteristics, the distance between the robot and device A is calculated to be approximately 8m. Considering the unobstructed wall in the x-direction, it is preliminarily determined that the robot is located within the area of x=10.0±8m and y=5.0±0.75m.
[0028] Coarse position cell division is based on the current area accuracy requirements, such as medium accuracy for a corridor. The initially determined area is divided into coarse frame cells, such as x=12.0-14.0m and y=4.5-6.5m, as coarse position cells at time t1.
[0029] Path concatenation involves obtaining multiple coarse location units based on a time series. For example, at time t2, the coarse-framed units x=14.0-16.0m and y=4.5-6.5m are concatenated to form an initial coarse path. For instance, in a 1000-square-meter office building on the 3rd floor, containing 10 offices and 1 corridor, 8 indoor detection devices (routers) are deployed. The target robot is a delivery robot that needs to deliver goods from the break room (x=5.0m, y=5.0m) to office 302 (x=35.0m, y=15.0m). The robot communication component captures a signal of RSSI -58dBm from device C (x=3.0m, y=5.0m) near the tea room. The signal is unobstructed, and the calculated distance is approximately 5m. Combined with the location of the tea room door (x=4.0m, y=4.0-6.0m), the coarse position unit at time t1 is determined to be a refined unit with a position of x=4.5-5.0m, y=4.5-5.0m. Since the tea room is a stopping point, high accuracy is required. After moving to the corridor, the robot captures a signal of RSSI -72dBm from device D (x=15.0m, y=5.0m). The signal is unobstructed, and the coarse position unit at time t2 is determined to be a coarse frame unit with a position of x=13.0-15.0m, y=4.5-6.5m. The coarse position units from time t1 to tn are then concatenated to form an initial coarse path from the tea room to the corridor.
[0030] In this embodiment, the imaging device is integrated on the target robot. It is a multimodal imaging device used to capture multimodal images of the robot's actions and behaviors. For example, a visible light camera of model OV5640, an infrared thermal imaging module of model MLX90640, and a depth imaging module are selected and installed on the robot's head. The height is 1.2m and the horizontal viewing angle is 120°. Image data is collected synchronously.
[0031] In this embodiment, the self-sensor is a sensor mounted on the target robot to measure motion posture data. A 6-axis IMU of model MPU6050 is selected and installed in the center of the robot chassis; a tactile sensor array of model TTP223 is selected and installed on the bottom edge of the robot to detect contact with the ground or obstacles.
[0032] In this embodiment, the fused posture is robot posture information obtained by fusing multimodal images of action behavior and motion posture data at the same moment of movement. It includes pose, action semantics and fusion confidence, and is used for subsequent path optimization.
[0033] In this embodiment, the motion pattern refers to the target robot's movement pattern, including four categories: linear movement, turning, stopping, and obstacle avoidance. This is identified by fusing action semantic tags from the posture with motion posture data. Specifically, based on action semantic tags from the fused posture, such as linear movement and right turn, and combined with IMU acceleration and angular velocity thresholds, such as x-axis acceleration fluctuations during linear movement... angular velocity along the z-axis during turning The motion pattern is identified using a support vector machine classifier, with an accuracy rate of no less than 95%.
[0034] In this embodiment, the global optimization expression is a path optimization mathematical model constructed for the coarse-frame unit, which is used to correct path deviations from a global perspective and ensure path continuity.
[0035] In this embodiment, the local optimization expression is a path optimization mathematical model constructed for the refined unit, which is used to correct local position deviations and improve positioning accuracy.
[0036] In this embodiment, the correction expression is a path optimization mathematical model constructed for the fuzzy unit. The position is corrected based on the optimization results of the neighboring unfuzzy units to ensure path integrity.
[0037] The beneficial effects of the above technical solution are as follows: through a three-level positioning logic of coarse positioning using communication signals, attitude acquisition through multi-source data fusion, and hierarchical path optimization, a balance between robot positioning accuracy and real-time performance in complex indoor environments is achieved: the initial coarse path is associated with the indoor layout through multi-dimensional communication signals, solving the problem of low positioning accuracy of a single signal; the attitude fusion is achieved by dynamically fusing multi-modal images and sensor data, solving the problem of poor robustness of attitude measurement caused by lighting and occlusion; the hierarchical optimization expression is optimized differently for different accuracy requirement units, solving the problem of lack of specificity in path optimization, and meeting the core requirements of point-to-point service robot docking and cleaning robot traversal.
[0038] This invention provides an indoor robot localization method, which determines the initial coarse path of the target robot, including: Collect the signal strength of the remaining indoor detection devices (excluding the indoor detection devices that have established a stable communication connection with the communication component) and the communication component, and classify the remaining indoor detection devices into stable devices that continuously detect signals and unstable devices that do not continuously detect signals. A dynamic network connection diagram is constructed by combining the deployment location of each indoor detection device, the indoor layout, the unique identifier of the connected device, the sensing signal strength, the signal attenuation characteristics, and the correlation between the signal strength of other stable devices and other unstable devices. Based on the dynamic network connection graph, the coarse positions are initially determined time-by-time. Multiple time-by-time coarse positions are then concatenated in a time sequence to form an initial coarse path composed of several coarse position units of varying sizes.
[0039] In this embodiment, the indoor detection device that is currently establishing a stable communication connection with the communication component refers to the indoor detection device that is currently establishing a stable connection with the target robot's communication component, and is the main reference device for coarse position calculation.
[0040] The remaining indoor detection devices are those whose communication components can scan for signals but have not established a stable connection, in addition to the currently connected devices. They are used to assist in correcting coarse positions.
[0041] The other devices are those with a signal continuous detection time >3s and RSSI value fluctuation <3dBm. They have high signal stability and high auxiliary correction weight. Otherwise, they are considered as unstable devices.
[0042] In this embodiment, the dynamic network connection graph is a dynamic network model constructed using indoor detection devices as nodes and the signal strength and connection status of communication components and devices as edge weights, combined with the indoor layout. This model is used to intuitively reflect the communication relationships between the robot and the devices. Specifically: Node definition: Taking indoor detection devices as nodes, it includes currently connected device A, stable device E, and unstable device F. Node attributes include device coordinates: A(10.0,5.0), E(15.0,10.0), F(20.0,8.0), and device type: currently connected / stable / unstable.
[0043] Edge weight definition: Edge weight = (RSSI normalized value × 0.6) + (connection stability × 0.4), where RSSI normalized value = (RSSI measured value + 90) / 30, the purpose is to map -90~-60dBm to 0~1, and connection stability = continuous probe time / 5, the purpose is to map 0~5s to 0~1.
[0044] Graph Construction: Using the Graphviz tool, nodes and edges are drawn on the background of the indoor layout diagram to form a dynamic network connection graph, which is updated once every 1 second. The currently connected devices are represented by red circles, stable devices by blue squares, and unstable devices by gray triangles. The thickness of the edges corresponds to the edge weights, with larger weights resulting in thicker edges.
[0045] In this embodiment, the coarse position determination at each time step is based on the node coordinates and edge weights of the dynamic network connection graph. The weighted triangulation method is used to calculate the coarse position at each time step. For example, at time t2, with devices A (weight 0.7) and E (weight 0.56) as references, the position (13.5m, 6.8m) is calculated and divided into coarse frame units x=13.0-15.0m, y=6.5-7.1m.
[0046] The beneficial effects of the above technical solution are as follows: by classifying other devices, constructing dynamic network connection diagrams, and refining the logic of coarse position concatenation at each time step, the problem of insufficient signal utilization and single position judgment in the initial coarse path construction is solved: the other devices are divided into stable and unstable types to realize differentiated utilization of signal data, the dynamic network connection diagram intuitively reflects the communication relationship, improves the accuracy of coarse position calculation, and provides a more reliable basic framework for subsequent fine path optimization.
[0047] This invention provides an indoor robot positioning method, which, before initially determining the coarse position based on the dynamic network connection graph time-by-time, further includes: The sensor signal strength of the communication components and various indoor detection devices of the target robot during its movement, the dynamic occlusion sequence of the indoor layout, and the motion acceleration trend of the communication components are collected. The intensity of the sensed signal is matched with the signal attenuation feature library of each indoor detection device in the time and frequency domain to generate a matching feature set containing the time domain stability and frequency domain distribution characteristics of signal attenuation. At the same time, combined with the motion acceleration trend and dynamic occlusion sequence, the orientation-motion-occlusion multi-dimensional constraint space of the target robot is constructed. Based on the multi-dimensional constraint space, matching feature set, and preset unit construction type determination system, the unit construction type at the current movement moment is determined, and the unit construction type is attached to the dynamic network connection graph in the form of dynamic attributes. The determination system integrates the second derivative change rate of signal intensity, the Markov transition probability of occlusion state, and the entropy value feature of motion acceleration as determination conditions for refined, coarse-framed, and fuzzy construction types, respectively.
[0048] In this embodiment, the dynamic occlusion sequence is a sequence of changes in occlusion objects such as walls, furniture, and human bodies in the indoor layout over time. The location, type, and duration of occlusion are recorded to analyze the dynamic impact of signal attenuation. Specifically, indoor infrared sensors detect occlusion objects in real time, recording occlusion information every 0.5 seconds. For example, at time t1, there is a human body occluding at x=12.0-13.0m and y=5.0-6.0m for 1 second, forming a dynamic occlusion sequence. For instance, there is human body occlusion at time t1, and no occlusion from time t2 to t5.
[0049] In this embodiment, the acceleration trend of the communication component is the acceleration change trend of the communication component as the robot moves. It is calculated using acceleration data measured by the IMU to reflect the smoothness of the robot's movement and is used to help determine the unit type. For example, if the acceleration fluctuation is large, the environment is complex and tends to be a fuzzy unit.
[0050] In this embodiment, time-frequency domain matching involves matching the time-domain waveform and frequency-domain features of the measured induced signal intensity with data in the signal attenuation feature library, calculating the matching similarity, and considering a similarity > 0.8 as a successful match, thereby generating a matching feature set that includes time-domain stability and frequency-domain distribution features.
[0051] In this embodiment, the orientation-motion-occlusion multi-dimensional constraint space is a spatial model constructed using robot orientation, motion acceleration trend, and occlusion state as three-dimensional constraint dimensions. This model is used to narrow down the range of coarse position units. For example, the orientation dimension (x=12.0-14.0m, y=4.5-6.5m based on preliminary judgment of communication signals) and the motion acceleration trend dimension (fluctuation) are used. Using three dimensions as axes (stability), occlusion (human occlusion at time t1, no occlusion from time t2 to t5), a constraint space is constructed. Each dimension has a set range threshold, such as a directional range of ±0.5m and acceleration fluctuations. The duration of the occlusion is less than 1 second.
[0052] In this embodiment, motion acceleration trend acquisition: the IMU acquires acceleration data every 0.05 seconds and calculates the acceleration fluctuation value within 1 second, such as the x-axis acceleration fluctuation at time t1-t5.
[0053] In this embodiment, the determination system rules are as follows: Refinement unit determination criterion: Rate of change of the second derivative of signal strength Signal stability is considered achieved when the Markov transition probability (unobstructed → unobstructed) is greater than 0.9, and motion acceleration entropy is considered to be stable. Criteria for coarsening elements: The rate of change of the second derivative of signal strength is <0.1 dB / s², 0.7 Markov transition probability in occlusion state 0.9, 0.1 The entropy value of motion acceleration is <0.2; Blurring unit determination criterion: rate of change of the second derivative of signal strength Markov transition probability of occlusion state < 0.7, motion acceleration entropy value 0.2. It should be noted that the rate of change of the second derivative of signal strength reflects the rate of signal change, the Markov transition probability reflects the continuity of the occlusion state, and the acceleration entropy value reflects the randomness of motion.
[0054] Type determination: Calculate the current parameter: rate of change of the second derivative of signal strength. The Markov transition probability of the occlusion state is 0.95 and the motion acceleration entropy value is 0.08, which meets the criteria for fine element determination, and the element construction type is determined to be fine element.
[0055] Attribute appending involves adding refined units as dynamic attributes to the node attributes of the corresponding coarse-position units in the dynamic network connection graph, such as refining the node annotation of the coarse-position unit at time t5.
[0056] The beneficial effects of the above technical solution are as follows: Through the logic of multi-dimensional data acquisition, time-frequency domain matching, constraint space construction, and type determination, dynamic and accurate determination of coarse position unit type is achieved: time-frequency domain matching makes full use of signal attenuation feature library to improve signal analysis depth; multi-dimensional constraint space integrates orientation, motion, and occlusion information to narrow the determination range; the determination system is based on quantitative parameters to avoid subjective judgment error and reduce the position error of the initial coarse path.
[0057] This invention provides an indoor robot localization method, which preliminarily determines a coarse position time-by-time based on a dynamic network connection graph with additional construction types, including: Based on the dynamic network connection graph with additional dynamic attributes, a real-time mapping function of construction type-location contribution is constructed, and the topological adjacency matrix of indoor layout is introduced as a regularization term to establish a multi-objective optimization model. Based on the multi-objective optimization model, the coarse position of the target robot is determined time-by-time by employing high-precision localization based on sparse Bayesian learning for refined construction type units, range shrinkage based on adaptive particle filtering for coarse-frame construction type units, and temporal correlation prediction based on attention mechanism for fuzzy construction type units. The real-time mapping function weights the refined construction type by an exponential function of signal confidence, the coarse-frame construction type by a logarithmic function of spatial coverage, and the fuzzy construction type by a Gaussian function of temporal correlation.
[0058] In this embodiment, the real-time mapping function is constructed by setting a contribution function based on the construction type. The contribution of the refined unit, f1: f1 = exp(0.5) S), where S is the signal confidence level and its value ranges from 0 to 1.
[0059] The contribution of the coarsened unit f2: f2 = ln(2 C), where C is the spatial coverage, with a value ranging from 0.5 to 1.
[0060] The contribution of the fuzzy unit f3: Where T represents temporal correlation, with a value ranging from 0 to 1. It is 0.5. The value of 0.2 was set in advance in the experiment.
[0061] Construction of topological adjacency matrix: Define indoor areas as nodes such as corridor C1, offices O1 and O2, and storage room S1, and construct a 3×3 matrix. Taking C1, O1, and S1 as an example: .
[0062] Multi-objective optimization model establishment: A model is constructed with location accuracy, spatial rationality, and temporal continuity as optimization objectives. ,in, The weights are 0.5, 0.3, and 0.2, respectively, and are pre-defined. f represents the contribution of the real-time mapping function. This represents the distance error between the coarse position and the reference device; Madj is an element of the topological adjacency matrix, where 1 indicates reasonable and 0 indicates unreasonable. This is the distance from the previous coarse position. If there is no previous position, the default distance is 0.
[0063] In this embodiment, differential positioning solution and coarse position determination are performed as follows: Refined cell solution: The sparse Bayesian learning algorithm is used to construct a sparse model by inputting the measured RSSI value (-65dBm), signal attenuation feature library data, and real-time mapping function contribution f1=1.608, and estimating the coarse position such as (14.0m, 5.0m).
[0064] Coarse-frame element solution: An adaptive particle filter algorithm is used. The number of particles is set to 50 when the signal is stable and 100 when the signal fluctuates. The dynamic network connection graph data and the topological adjacency matrix are input. The position range is narrowed by particle resampling, such as from x=13.0-15.0m to x=13.5-14.5m, and y=4.5-6.5m to y=4.8-5.2m, to determine the coarse position (14.0m, 5.0m).
[0065] Fuzzy unit solution: A temporal correlation prediction algorithm with attention mechanism is used to assign attention weights to the coarse positions of the previous 5 time steps, such as t1(20.0m,10.0m), t2(21.0m,10.5m), t3(22.0m,11.0m), t4(23.0m,11.5m), and t5(24.0m,12.0m) (the weight of recent time steps is higher, such as t5 weight 0.3 and t4 weight 0.25), and predict the current coarse position (25.0m,12.5m).
[0066] The beneficial effects of the above technical solution are as follows: through real-time mapping function, topological adjacency matrix, multi-objective optimization model, and differentiated solution logic, accurate positioning of different types of coarse location units can be achieved: real-time mapping function quantifies the contribution of different types of units, improves the rationality of weight allocation, topological adjacency matrix constrains the space rationality, avoids position determination from deviating from the actual environment, and differentiated algorithm selects the appropriate method for different accuracy requirements, taking into account both accuracy and efficiency, and improving the overall accuracy of the initial coarse path.
[0067] This invention provides an indoor robot positioning method, step 2 of which includes: The system captures multimodal images of the target robot's motion behavior at each moment of movement, captured by an imaging device, as well as multi-source motion posture data measured by the target robot's multi-axis inertial measurement unit and tactile sensor array. Cross-modal feature fusion is performed on the multimodal images of the motion behavior to extract an image feature set containing appearance semantic features, three-dimensional structural features, and dynamic thermal distribution features. Simultaneously, spatiotemporal calibration is performed on the multi-source motion posture data to obtain a sensor feature set containing acceleration time-domain spectrum, angular velocity frequency-domain features, and tactile pressure distribution matrix. Based on the multi-source motion posture data, the motion state of the target robot is obtained through machine learning classification. A dynamic fusion weight network for image feature set and sensor feature set is constructed. The dynamic fusion weight network assigns attention weights to image feature set according to motion mode confidence and fusion weights to sensor feature set according to consistency entropy value of multi-source data for different motion states. Based on the dynamic fusion weight network, multi-scale spatiotemporal alignment fusion of image feature sets and sensor feature sets is performed to obtain the fused pose at the same movement moment. The fused pose includes a visual-motion coupled pose matrix, action semantic labels, and fusion confidence of multi-source data.
[0068] Multimodal images of robot motion behavior are images captured by imaging devices that include visible light textures and infrared thermal distribution, such as images of the robot's posture when moving in a straight line and images of heading changes when turning.
[0069] In this embodiment, the visible light image uses CNN to extract appearance semantic features, such as semantic labels for the robot's linear movement, with a feature dimension of 256. Infrared image: U-Net is used to extract dynamic features of heat distribution, such as the location of the heat-generating area of the motor remains unchanged, with a feature dimension of 128; Depth image: PointNet is used to extract 3D structural features, such as the robot being 1.5m away from the wall with no change, and the feature dimension is 128; Feature fusion: The three types of features are concatenated (256+128+128=512 dimensions). An attention mechanism is used to assign a weight of 0.5 to the semantic features, 0.25 to the heat distribution features, and 0.25 to the three-dimensional structural features to obtain the image feature set.
[0070] Spatiotemporal calibration: Timestamp alignment: Unify the timestamps of IMU and tactile sensor data to 0.1s, such as taking the average acceleration and angular velocity of the IMU at 0.1s; Spatial coordinate calibration: Converting the IMU's local coordinate system acceleration data into global coordinate system data, achieved using the identity matrix at a heading angle of 0°; Feature extraction: Fourier transform is performed on the calibrated IMU data to obtain the acceleration time-domain spectrum and angular velocity frequency-domain features. Mean filtering is performed on the tactile data to obtain the pressure distribution matrix (8×1 dimension), forming a sensor feature set.
[0071] In this embodiment, motion state classification is based on acceleration and angular velocity thresholds from sensor feature sets, such as linear motion: x-axis acceleration fluctuations. For turning: z-axis angular velocity < 5° / s; z-axis angular velocity > 5° / s. The motion state is identified by a random forest classifier, such as the current movement being in a straight line.
[0072] Dynamic fusion weight network construction: A two-layer fully connected neural network is used. The input is the one-hot encoding of motion state features, such as linear movement [1,0,0,0], and the output is the image feature set weight w_img and the sensor feature set weight w_sen (w_img+w_sen=1), as follows: Linear movement: w_img=0.6 (depending on visual texture), w_sen=0.4 (depending on IMU stationary data); Turning: w_img=0.4 (visual features are easily blurred due to turning), w_sen=0.6 (depends on IMU angular velocity); Dock: w_img=0.7 (depending on visual positioning points), w_sen=0.3 (motion data stable); Obstacle avoidance: w_img=0.5 (visual recognition of obstacles), w_sen=0.5 (IMU detection of acceleration changes).
[0073] In this embodiment, spatiotemporal alignment is achieved by aligning the timestamps and spatial coordinates of the image feature set with those of the sensor feature set, such as using the center of the robot chassis as the origin and the global coordinate system as the reference.
[0074] Feature fusion: The image feature set and the sensor feature set are weighted and fused by moving in a straight line with w_img=0.6 and w_sen=0.4 to obtain the fused feature vector.
[0075] Fusion pose output: Based on the fusion feature vector, calculate the visual-motion coupled pose matrix, such as (15.5m, 5.0m, 0.1m, heading angle 0°), and combine it with the action semantic label: linear movement with a confidence of 0.95, and the multi-source data fusion confidence of 0.92 to form the fusion pose.
[0076] The beneficial effects of the above technical solution are as follows: by using the logic of multimodal data acquisition, cross-modal feature fusion, dynamic weight allocation, and multi-scale spatiotemporal alignment, it solves the problem of insufficient robustness of multi-source data fusion: multimodal shooting equipment and multi-axis sensors ensure data comprehensiveness, cross-modal feature fusion and spatiotemporal calibration eliminate data deviation, and the dynamic fusion weight network adaptively adjusts the weights according to the motion state, improving the fusion reliability in different scenarios, providing high-precision attitude support for path optimization, and improving the accuracy and continuity of subsequent fine paths.
[0077] This invention provides an indoor robot localization method, which constructs a global optimization expression for coarse-frame units, a local optimization expression for fine-frame units, and a correction expression for fuzzy-frame units, including: Multimodal feature decomposition is performed on the fused poses of two adjacent coarse positions to obtain a set of pose features including a six-dimensional pose vector, an action semantic probability distribution, and a multi-source data reliability matrix of the fused poses. The system simultaneously collects the positional distance between two adjacent coarse positions, resolves it into Euclidean distance and indoor topological correlation distance, and extracts the temporal gradient and spatial distribution entropy of the distance. At the same time, it identifies the motion pattern of the target robot. With the global spatial deviation norm of the attitude feature set, the product of Euclidean distance and dynamic weight of motion pattern, and the robust regularization term of the multi-source data reliability matrix as the core, an adaptive constraint factor based on motion pattern is introduced to form a global optimization expression with multiple constraints coupled. The local detail deviation of the pose feature set, the intersection of indoor topological association distance and dynamic weight of motion mode, and the semantic obstacle matrix of indoor layout are fused together. At the same time, the action semantic probability distribution of the fused pose is embedded as a weighting coefficient to form a local optimization expression that is strongly correlated with the indoor environment layout. Based on the first and second expressions of the nearest non-fuzzy unit of the corresponding fuzzy unit, and combined with the change vector of the adjacent coarse position of the corresponding fuzzy unit, the correction expression is obtained.
[0078] In this embodiment, multimodal feature decomposition decomposes the multidimensional features of the fused pose, such as pose, action semantics, and fusion confidence, into independent feature subsets.
[0079] The pose feature set is a set of features obtained after multimodal feature decomposition, which includes a six-dimensional vector of pose (x, y, z, heading angle, pitch angle, roll angle) and a probability distribution of action semantics, reflecting the core information of the fused pose.
[0080] The multi-source data reliability matrix is a matrix that quantifies the reliability of image data and sensor data. The matrix elements are reliability values, which are used to weight the credibility of different data sources in the optimization expression.
[0081] The indoor topology association distance is the shortest path distance between two adjacent coarse locations within the indoor topology, reflecting the actual reachability of the location and used for local optimization.
[0082] The time-domain gradient of distance is the rate of change of the distance between two adjacent coarse positions per unit time, reflecting the robot's movement speed, and is used to dynamically adjust and optimize the regularization term.
[0083] The spatial distribution entropy of distance is the randomness of the spatial distribution of the distance between two adjacent coarse locations. The smaller the entropy value, the more uniform the distribution, and it is used to judge the complexity of the spatial environment.
[0084] The adaptive constraint factor is a constraint coefficient that is dynamically adjusted based on the robot's motion pattern. The constraint strength varies depending on the motion pattern.
[0085] The robust regularization term is used in the optimization expression to suppress the impact of data noise. It is calculated based on the reliability matrix of multi-source data. The higher the reliability, the smaller the weight of the regularization term.
[0086] The semantic obstacle matrix is a matrix that reflects the location and type of indoor obstacles. A matrix element of 1 indicates the presence of an obstacle, and 0 indicates the absence of an obstacle. It is used to constrain the rationality of the path in local optimization.
[0087] The first expression is the global optimization expression corresponding to the nearest coarse-framed unit of the fuzzy unit, and the second expression is the local optimization expression corresponding to the nearest fine-framed unit of the fuzzy unit. Both are used to correct the position of the fuzzy unit.
[0088] The change vector of adjacent coarse positions is the displacement vector of two adjacent coarse positions in the x and y directions. It is used to reflect the position change trend and help correct the position of the fuzzy unit.
[0089] Specifically: Multimodal eigenvalue decomposition and reliability matrix construction: Multimodal feature decomposition: Decomposing the fused pose at two adjacent coarse positions, such as time t5 (28.0m, 14.5m) and time t6 (28.5m, 15.0m), yields: The pose six-dimensional vector is: t5(28.0,14.5,0.1,0°,0°,0°), t6(28.5,15.0,0.1,0°,0°,0°); Action semantic probability distribution: t5 docking 0.95, linear movement 0.05, t6 docking 0.98, linear movement 0.02; forming a set of posture features.
[0090] Construction of multi-source data reliability matrix: Based on fusion confidence (t5=0.92, t6=0.95) and data noise level (image data noise 0.05, sensor data noise 0.08), reliability values are calculated. Image data reliability = 1 - noise = 0.95; Sensor data reliability = 1 - noise = 0.92; construct matrix It is a diagonal matrix because the image and sensor data are independent.
[0091] In this embodiment, location distance resolution: Euclidean distance d_geo: calculates the straight-line distance between timest5 and t6; Indoor topological association distance d_top: At times t5-t6, the location is in office 302 with no obstacles, and the topological association distance = Euclidean distance; Time-domain gradient The time interval between t5 and t6 is 1 second. ; Spatial distribution entropy H_d: The spatial distribution within the office is uniform, H_d=0.2.
[0092] Motion pattern recognition: Action semantic probability distribution based on fused pose with docking > 0.95 and acceleration fluctuations The motion pattern is identified as docking using an SVM classifier.
[0093] In this embodiment, the global optimization expression ,in, λ1 is the global spatial deviation norm; λ1 is the mode constraint factor, which takes a value of 0.9 when in docking mode. It is pre-set and can be directly called. d_geo is the Euclidean distance metric, with a value ranging from 0 to 1. , w_mode is the dynamic weight of the motion mode; λ2 Tr is the robust regularization term, and λ2 is the regularization coefficient. Tr1 is the trace of the reliability matrix. The dynamic weights of the motion modes are: 0.6 for straight-line movement, 0.8 for turning, 0.9 for stopping, and 0.7 for obstacle avoidance.
[0094] In this embodiment, the local optimization expression ,in, For local detail deviations, for example, a deviation of 0.5m in the x-direction and 0.5m in the y-direction between time t5 and t6, the maximum value of 0.5m is taken, d_top w_mode c represents the weighted cross term, where c = 0.1. d_top represents the topological metric distance, ranging from 0 to 1. p_sem represents the weighted probability distribution of the action semantics, such as an action semantic docking probability of 0.95, which is pre-set and can be used directly. w_mode represents the value corresponding to the motion mode, which varies depending on the mode and is pre-set. Tr2 is the value obtained based on the semantic obstacle matrix. For example, the matrix corresponding to an office without obstacles is [0,0], in which case Tr2[0,0] = 0. It should be noted that in the semantic obstacle matrix, 1 represents an obstacle (such as a wall or furniture), and 0 represents no obstacle.
[0095] In this embodiment, for example, the nearest non-fuzzy unit of the fuzzy unit at time t4 is the coarse-frame unit at time t3 and the fine-frame unit at time t5. In this case, the correction expression... ,in, The weights for J1 and J2 are 0.4, and γ is 0.2, representing the weights of the change vector. For the change vector The magnitude of the change vector between adjacent coarse positions. The change vector between time t3 (26.0m, 13.0m) and time t4 (22.5m, 12.5m) is (-3.5m, -0.5m).
[0096] The beneficial effects of the above technical solution are as follows: by constructing logic through multimodal feature decomposition, location distance analysis, and hierarchical optimization expressions, the problem of lack of specificity in path optimization is solved: the global optimization expression ensures the continuity of the global path for coarse-framed units, the local optimization expression improves the local positioning accuracy for fine-framed units, and the correction expression ensures the integrity of the path by utilizing neighboring unit information for fuzzy units, thus providing accurate optimization model support for fine path acquisition.
[0097] This invention provides an indoor robot localization method, which obtains and outputs a fine path, including: Multi-objective joint optimization is performed on the global optimization expression, local optimization expression and correction expression. A dynamic regularization term based on the gradient of the temporal change of location distance is introduced to achieve the coordinated convergence of coarse-framed units, fuzzy units and fine-framed units. The fine path is obtained based on the collaborative convergence results.
[0098] In this embodiment, the algorithm selected is the particle swarm optimization algorithm because it has a fast convergence speed, is easy to implement, and is suitable for real-time positioning scenarios. Parameter settings: number of particles 50, inertia weight 0.7, learning factor c1=c2=1.494, maximum number of iterations 20, convergence error threshold 0.01.
[0099] Calculation of dynamic regularization term: based on the gradient of temporal variation of location distance For example, if the speed is 0.707 m / s between time t5 and t6, construct the regularization term: Where λ5 = 0.05 is the regularization coefficient; The regularization term is added to the objective function of multi-objective optimization to balance the flexibility and stability of path adjustment. When R_d is large, the constraints should be relaxed appropriately. When R_d is small, strengthen the constraints, i.e.: At this point, R_d = 0.025. At this point, R_d = 0.01.
[0100] Objective function construction: J1, J2, J3 are combined with the dynamic regularization term to construct a multi-objective optimization function. Where w01, w02, and w03 are the weights of the coarse-frame unit, fine-frame unit, and fuzzy-frame unit, respectively: 0.3, 0.5, and 0.2; PSO algorithm optimization: Initialize the particle swarm, iteratively update the particle positions and velocities, and gradually decrease J_total until the convergence condition is met. For example, after 15 iterations, J_total = 0.95405 → 0.9400, the deviation is <0.01, and the algorithm converges.
[0101] Convergence criteria: When the deviations of the optimization results of J1, J2, and J3 from the optimal values are all <0.01, and the optimized positions of each coarse position unit are spatially continuous (i.e., the overlap area of adjacent units is >30%), and the indoor topological adjacency matrix is met (e.g., corridor units and office units are adjacent), then collaborative convergence is determined.
[0102] Fine path generation: Based on the converged optimization values of each coarse position unit, such as time t5 (28.2m, 14.8m) and time t6 (28.5m, 15.0m), a fine path is formed by concatenating them in time series to meet the positioning requirements of the robot for point-to-point service and cleaning.
[0103] The beneficial effects of the above technical solution are as follows: by using the logic of multi-objective joint optimization, dynamic regularization, and collaborative convergence, the problem of local optima and global discontinuity in path optimization is solved. Multi-objective joint optimization takes into account the needs of different precision units at the same time. The dynamic regularization term adaptively adjusts the constraint strength according to the movement speed. Collaborative convergence ensures the continuity of the path space and the rationality of the topology. Ultimately, the positioning accuracy of fine paths is improved, providing precise position support for robot autonomous navigation and task execution.
[0104] This invention provides an indoor robot positioning system, such as... Figure 2 As shown, it includes: The coarse path determination module is used to capture the communication connection signals between the communication component and the indoor detection device during the movement of the target robot, and determine the initial coarse path of the target robot by combining the deployment location of the indoor detection device and the indoor layout. The initial coarse path is composed of several coarse position units, and the coarse position units include fine units, coarse frame units, and fuzzy units. The position accuracy of the fine units is greater than that of the coarse frame units, and the position accuracy of the coarse frame units is greater than that of the fuzzy units. The communication connection signals include the unique identifier of the indoor detection device that the communication component connects to each time, the corresponding connection sensing signal strength, the signal attenuation characteristics based on the unobstructed condition in each specified direction, and the signal attenuation characteristics based on the occluded condition. The fusion module is used to capture multimodal images of the target robot's actions at each moment of movement captured by the camera and motion posture data measured by the target robot's own sensors during the movement process, so as to obtain the fused posture at the same moment of movement. The expression refinement module is used to construct a global optimization expression for the coarse frame unit, a local optimization expression for the fine frame unit, and a correction expression for the fuzzy frame unit based on the fused pose of two adjacent coarse frames, the positional distance between the two adjacent coarse frames, and the motion pattern of the target robot, so as to obtain the fine path and output it.
[0105] The beneficial effects of the above technical solution are as follows: through a three-level positioning logic of coarse positioning using communication signals, attitude acquisition through multi-source data fusion, and hierarchical path optimization, a balance between robot positioning accuracy and real-time performance in complex indoor environments is achieved: the initial coarse path is associated with the indoor layout through multi-dimensional communication signals, solving the problem of low positioning accuracy of a single signal; the attitude fusion is achieved by dynamically fusing multi-modal images and sensor data, solving the problem of poor robustness of attitude measurement caused by lighting and occlusion; the hierarchical optimization expression is optimized differently for different accuracy requirement units, solving the problem of lack of specificity in path optimization, and meeting the core requirements of point-to-point service robot docking and cleaning robot traversal.
[0106] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. An indoor robot positioning method, characterized in that, include: Step 1: Capture the communication connection signal between the communication component and the indoor detection device during the movement of the target robot, and determine the initial coarse path of the target robot by combining the deployment location of the indoor detection device and the indoor layout. The initial coarse path consists of several coarse position units, and the coarse position units include fine units, coarse frame units, and fuzzy units. The position accuracy of the fine units is greater than that of the coarse frame units, and the position accuracy of the coarse frame units is greater than that of the fuzzy units. The communication connection signal includes the unique identifier of the indoor detection device that the communication component connects to each time, the corresponding connection sensing signal strength, the signal attenuation characteristics based on the unobstructed situation in each specified direction, and the signal attenuation characteristics based on the obstructed situation. Step 2: Capture multimodal images of the target robot's motion behavior at each moment of movement based on the imaging device and motion posture data measured by the target robot's own sensors during the movement process, and obtain the fused posture at the same moment of movement; Step 3: Based on the fused pose of two adjacent coarse positions, the positional distance between the two adjacent coarse positions, and the motion pattern of the target robot, construct a global optimization expression for the coarse frame unit, a local optimization expression for the fine frame unit, and a correction expression for the fuzzy frame unit to obtain and output the fine path.
2. The indoor robot positioning method according to claim 1, characterized in that, Determining the initial coarse path of the target robot includes: Collect the signal strength of the remaining indoor detection devices (excluding the indoor detection devices that have established a stable communication connection with the communication component) and the communication component, and classify the remaining indoor detection devices into stable devices that continuously detect signals and unstable devices that do not continuously detect signals. A dynamic network connection diagram is constructed by combining the deployment location of each indoor detection device, the indoor layout, the unique identifier of the connected device, the sensing signal strength, the signal attenuation characteristics, and the correlation between the signal strength of other stable devices and other unstable devices. Based on the dynamic network connection graph, the coarse positions are initially determined time-by-time. Multiple time-by-time coarse positions are then concatenated in a time sequence to form an initial coarse path composed of several coarse position units of varying sizes.
3. The indoor robot positioning method according to claim 2, characterized in that, Before initially determining the coarse position based on the dynamic network connection graph at each time step, the process also includes: The sensor signal strength of the communication components and various indoor detection devices of the target robot during its movement, the dynamic occlusion sequence of the indoor layout, and the motion acceleration trend of the communication components are collected. The intensity of the sensed signal is matched with the signal attenuation feature library of each indoor detection device in the time and frequency domain to generate a matching feature set containing the time domain stability and frequency domain distribution characteristics of signal attenuation. At the same time, combined with the motion acceleration trend and dynamic occlusion sequence, the orientation-motion-occlusion multi-dimensional constraint space of the target robot is constructed. Based on the multi-dimensional constraint space, matching feature set, and preset unit construction type determination system, the unit construction type at the current movement moment is determined, and the unit construction type is attached to the dynamic network connection graph in the form of dynamic attributes. The determination system integrates the second derivative change rate of signal intensity, the Markov transition probability of occlusion state, and the entropy value feature of motion acceleration as determination conditions for refined, coarse-framed, and fuzzy construction types, respectively.
4. The indoor robot positioning method according to claim 3, characterized in that, Based on the additional construction type, the coarse position of the dynamic network connection graph is initially determined time-by-time, including: Based on the dynamic network connection graph with additional dynamic attributes, a real-time mapping function of construction type-location contribution is constructed, and the topological adjacency matrix of indoor layout is introduced as a regularization term to establish a multi-objective optimization model. Based on the multi-objective optimization model, the coarse position of the target robot is determined time-by-time by employing high-precision localization based on sparse Bayesian learning for refined construction type units, range shrinkage based on adaptive particle filtering for coarse-frame construction type units, and temporal correlation prediction based on attention mechanism for fuzzy construction type units. The real-time mapping function weights the refined construction type by an exponential function of signal confidence, the coarse-frame construction type by a logarithmic function of spatial coverage, and the fuzzy construction type by a Gaussian function of temporal correlation.
5. The indoor robot positioning method according to claim 1, characterized in that, Step 2 includes: The system captures multimodal images of the target robot's motion behavior at each moment of movement, captured by an imaging device, as well as multi-source motion posture data measured by the target robot's multi-axis inertial measurement unit and tactile sensor array. Cross-modal feature fusion is performed on the multimodal images of the motion behavior to extract an image feature set containing appearance semantic features, three-dimensional structural features, and dynamic thermal distribution features. Simultaneously, spatiotemporal calibration is performed on the multi-source motion posture data to obtain a sensor feature set containing acceleration time-domain spectrum, angular velocity frequency-domain features, and tactile pressure distribution matrix. Based on the multi-source motion posture data, the motion state of the target robot is obtained through machine learning classification. A dynamic fusion weight network for image feature set and sensor feature set is constructed. The dynamic fusion weight network assigns attention weights to image feature set according to motion mode confidence and fusion weights to sensor feature set according to consistency entropy value of multi-source data for different motion states. Based on the dynamic fusion weight network, multi-scale spatiotemporal alignment fusion of image feature sets and sensor feature sets is performed to obtain the fused pose at the same movement moment. The fused pose includes a visual-motion coupled pose matrix, action semantic labels, and fusion confidence of multi-source data.
6. The indoor robot positioning method according to claim 1, characterized in that, Construct global optimization expressions for coarse-frame units, local optimization expressions for fine-frame units, and correction expressions for fuzzy-frame units, including: Multimodal feature decomposition is performed on the fused poses of two adjacent coarse positions to obtain a set of pose features including a six-dimensional pose vector, an action semantic probability distribution, and a multi-source data reliability matrix of the fused poses. The system simultaneously collects the positional distance between two adjacent coarse positions, resolves it into Euclidean distance and indoor topological correlation distance, and extracts the temporal gradient and spatial distribution entropy of the distance. At the same time, it identifies the motion pattern of the target robot. With the global spatial deviation norm of the attitude feature set, the product of Euclidean distance and dynamic weight of motion pattern, and the robust regularization term of the multi-source data reliability matrix as the core, an adaptive constraint factor based on motion pattern is introduced to form a global optimization expression with multiple constraints coupled. The local detail deviations of the pose feature set, the intersection of indoor topological association distance and dynamic weight of motion mode, and the semantic obstacle matrix of indoor layout are fused together. At the same time, the action semantic probability distribution of the fused pose is embedded as a weighting coefficient to form a local optimization expression that is strongly correlated with the indoor environment layout. Based on the first and second expressions of the nearest non-fuzzy unit of the corresponding fuzzy unit, and combined with the change vector of the adjacent coarse position of the corresponding fuzzy unit, the correction expression is obtained.
7. The indoor robot positioning method according to claim 6, characterized in that, Obtain and output the fine path, including: Multi-objective joint optimization is performed on the global optimization expression, local optimization expression and correction expression. A dynamic regularization term based on the gradient of the temporal change of location distance is introduced to achieve the coordinated convergence of coarse-framed units, fuzzy units and fine-framed units. The fine path is obtained based on the collaborative convergence results.
8. An indoor robot positioning system, characterized in that, include: The coarse path determination module is used to capture the communication connection signals between the communication component and the indoor detection device during the movement of the target robot, and determine the initial coarse path of the target robot by combining the deployment location of the indoor detection device and the indoor layout. The initial coarse path is composed of several coarse position units, and the coarse position units include fine units, coarse frame units, and fuzzy units. The position accuracy of the fine units is greater than that of the coarse frame units, and the position accuracy of the coarse frame units is greater than that of the fuzzy units. The communication connection signals include the unique identifier of the indoor detection device that the communication component connects to each time, the corresponding connection sensing signal strength, the signal attenuation characteristics based on the unobstructed condition in each specified direction, and the signal attenuation characteristics based on the occluded condition. The fusion module is used to capture multimodal images of the target robot's actions at each moment of movement captured by the camera and motion posture data measured by the target robot's own sensors during the movement process, so as to obtain the fused posture at the same moment of movement. The expression refinement module is used to construct a global optimization expression for the coarse frame unit, a local optimization expression for the fine frame unit, and a correction expression for the fuzzy frame unit based on the fused pose of two adjacent coarse frames, the positional distance between the two adjacent coarse frames, and the motion pattern of the target robot, so as to obtain the fine path and output it.