Robot control method, robot, robot controller and storage medium

By fusing multi-source information from lidar, ultra-wideband positioning modules, and inertial measurement units, and combining it with voice-interactive control, the problems of inaccurate robot positioning and low human-machine interaction efficiency in metal-intensive workshops have been solved, achieving high-precision and efficient human-machine collaborative operation.

CN121756355APending Publication Date: 2026-03-31SUNWODA MOBILITY ENERGY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-13
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In industrial sites with dense metal equipment and high noise levels, existing technologies for industrial robots suffer from insufficient positioning accuracy, easy drift in positioning results, and low efficiency in human-machine interaction. This makes it particularly difficult to meet the needs of high-precision and low-intervention continuous automated production under complex working conditions.

Method used

The robot acquires environmental map information via LiDAR, spatial ranging information via UWB positioning module, and robot pose status information via inertial measurement unit. Multi-source positioning information is fused and combined with voice interaction commands to control the robot to perform actions.

Benefits of technology

It improves the positioning accuracy and stability of robots in complex environments, enhances autonomous operation capabilities and the convenience of human-machine interaction, and meets the requirements of high-precision and low-intervention continuous automated production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121756355A_ABST
    Figure CN121756355A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot control, and discloses a robot control method, a robot, a robot controller and a storage medium, and the method comprises the steps: obtaining environment map information through a laser radar as first positioning information, acquiring space ranging information between the robot and the target base station through an ultra-wideband positioning module as second positioning information, and acquiring information representing the pose state of the robot through an inertial measurement unit as third positioning information; performing fusion processing on the first positioning information, the second positioning information and the third positioning information to obtain a current positioning result of the robot; and in response to the voice interaction instruction, processing the voice interaction instruction to obtain a target instruction, and controlling the robot to execute a target action according to the target instruction and the positioning result. According to the method, the autonomous operation capability of the robot in a complex environment and the convenience and intelligence level of man-machine interaction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robot control technology, and in particular to a robot control method, a robot, a robot controller, and a storage medium. Background Technology

[0002] With the development of intelligent manufacturing and flexible production, industrial robots have been widely deployed in workshop environments with dense metal equipment for material handling, precision assembly, and flexible production line reconfiguration. Current industrial robot positioning largely relies on single-type sensors, such as LiDAR-based environmental mapping and positioning, or ultra-wideband (UWB) ranging and positioning; in some scenarios, wireless positioning technologies such as Wi-Fi and Bluetooth are also used to achieve area recognition and trajectory tracking. Regarding human-machine interaction, industrial robots are typically controlled via button panels, remote controls, or teach pendants. Some companies are beginning to explore the introduction of consumer-grade voice control systems that strive for natural interaction, aiming to improve operational efficiency and ease of use in complex working conditions.

[0003] However, in typical industrial settings characterized by dense metal equipment and high noise levels, the aforementioned technical approaches reveal significant limitations. Firstly, lidar is susceptible to strong reflections from metal equipment and shelves, resulting in prominent false echoes and occlusion in point cloud data, making it difficult to reliably meet the ±1cm precision assembly positioning requirements. Secondly, UWB exhibits significant multipath effects in metallic environments, with ranging errors typically reaching 5–10cm. This, coupled with signal drift issues in highly reflective scenarios for wireless technologies like Wi-Fi and Bluetooth, leads to easily misaligned robot positioning results, necessitating frequent manual calibration and impacting cycle time and production capacity. Thirdly, industrial noise levels of 70–90dB cause consumer-grade voice systems to have a misrecognition rate exceeding 30%, forcing operators to rely on buttons, remote controls, or teach pendants for control. When handling objects with both hands, the current operation must be interrupted to switch control modes, creating a clear contradiction between human-machine interaction efficiency and the need for precise positioning, making it difficult to support high-precision, low-intervention continuous automated production. Summary of the Invention

[0004] In view of this, embodiments of this application provide a robot control method, a robot, a robot controller, and a storage medium, which can effectively solve problems such as insufficient positioning accuracy of industrial robots in metal-intensive workshops, easy drift of positioning results, and low human-machine interaction efficiency in high-noise environments.

[0005] In a first aspect, embodiments of this application provide a robot control method applied to a robot equipped with a lidar, an ultra-wideband positioning module, and an inertial measurement unit, comprising: The robot obtains environmental map information through the lidar as the first positioning information, obtains spatial ranging information between the robot and the target base station through the ultra-wideband positioning module as the second positioning information, and obtains information characterizing the robot's pose state through the inertial measurement unit as the third positioning information. The first positioning information, the second positioning information, and the third positioning information are fused together to obtain the current positioning result of the robot. In response to a voice interaction command, the voice interaction command is processed to obtain a target command, and the robot is controlled to perform a target action based on the target command and the positioning result.

[0006] In some embodiments, the step of acquiring environmental map information via the lidar as first positioning information, acquiring spatial ranging information between the robot and the target base station via the ultra-wideband positioning module as second positioning information, and acquiring information characterizing the robot's pose state via the inertial measurement unit as third positioning information includes: The environmental point cloud data collected by the lidar is acquired, and environmental map information is constructed based on the environmental point cloud data as the first positioning information; The ultra-wideband positioning module receives wireless ranging signals and measures the spatial distance information between the robot and the target base station based on the wireless ranging signals, which is used as the second positioning information; wherein, the ultra-wideband positioning module adjusts the carrier signal frequency at preset intervals during the process of receiving the wireless ranging signals. The robot's initial pose information is obtained based on the motion state data collected by the inertial measurement unit. Compensation processing is performed on the initial pose information to obtain corrected pose information, which is used as the third positioning information.

[0007] In some embodiments, acquiring the environmental point cloud data collected by the lidar and constructing environmental map information based on the environmental point cloud data includes: Extract the reflection intensity value corresponding to each spatial sampling point in the environmental point cloud data; The reflection intensity value is compared with a preset intensity threshold. Point cloud data with a reflection intensity value less than or equal to the preset intensity threshold are retained, and point cloud data with a reflection intensity value greater than the preset intensity threshold are discarded to obtain the target point cloud data. Based on the target point cloud data, localization and mapping operations are performed to generate current environment map information.

[0008] In some embodiments, receiving wireless ranging signals through the ultra-wideband positioning module and measuring spatial ranging information between the robot and the target base station based on the wireless ranging signals includes: The ultra-wideband positioning module receives wireless ranging signals and performs differential processing on the wireless ranging signals to obtain differential ranging signals. The differential ranging signal is constructed into an input feature tensor for signal analysis and input into a preset convolutional neural network model to perform type identification on the differential ranging signal in order to obtain a direct signal that represents the direct path between the robot and the target base station. The spatial ranging information between the robot and the target base station is determined based on the direct signal.

[0009] In some embodiments, obtaining the robot's initial pose state information based on the motion state data collected by the inertial measurement unit, and performing compensation processing on the initial pose state information to obtain corrected pose state information, includes: Acquire the acceleration and angular velocity data output by the inertial measurement unit; Complementary filtering is performed based on the acceleration data and the angular velocity data to generate the initial pose state information used to characterize the robot's posture state; When the initial pose state information meets the preset tilt threshold condition, Kalman filtering is performed based on the initial pose state information and the spatial ranging information to generate the corrected pose state information.

[0010] In some embodiments, fusing the first positioning information, the second positioning information, and the third positioning information to obtain the robot's current positioning result includes: The lidar positioning information is determined based on the first positioning information; Ultra-wideband positioning information is determined based on the second positioning information; Inertial positioning information is determined based on the third positioning information; Obtain the weighting coefficients corresponding to the lidar positioning information, the ultra-wideband positioning information, and the inertial positioning information, respectively; The lidar positioning information, the ultra-wideband positioning information, and the inertial positioning information are weighted and fused with their respective weight coefficients to obtain a weighted fusion result, and the robot's current positioning result is generated based on the weighted fusion result.

[0011] In some embodiments, the step of processing the voice interaction command in response to obtain a target command, and controlling the robot to perform a target action based on the target command and the positioning result, includes: Receive the voice signal corresponding to the voice interaction command, and perform noise reduction and recognition processing on the voice signal to generate a voice recognition result; Based on the speech recognition results, instruction association processing is performed to generate instruction candidate information; The speech recognition result and the instruction candidate information are subjected to fuzzy matching processing to obtain the target instruction; Based on the target instruction and the positioning result, control instructions are generated to control the robot to perform the target action.

[0012] In some embodiments, the method further includes: After receiving the completion information of the robot's target action, corresponding execution status information is generated; The execution status information is converted into voice broadcast content, and the task execution status is fed back through the voice output unit.

[0013] Secondly, embodiments of this application provide a robot, including: a robot body, a lidar module, an ultra-wideband positioning module, an inertial measurement module, a control module, and a voice interaction module disposed on the robot body; The lidar module is used to acquire environmental map information as the first positioning information; The ultra-wideband positioning module is used to acquire the spatial ranging information between the robot and the target base station as the second positioning information; The inertial measurement module is used to acquire information characterizing the robot's pose state as third positioning information; The voice interaction module is used to receive voice interaction commands; The control module is used to fuse the first positioning information, the second positioning information and the third positioning information to obtain the current positioning result of the robot; The control module is further configured to respond to the voice interaction command, process the voice interaction command to obtain a target command, and control the robot body to perform a target action based on the target command and the positioning result.

[0014] Thirdly, embodiments of this application provide a robot controller, which stores a computer program and executes the computer program to implement the robot control method of the first aspect described above.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium, wherein when the computer program is executed on a processor, it implements the robot control method of the first aspect described above.

[0016] The embodiments of this application have the following beneficial effects: First, environmental map information is acquired using LiDAR as the first positioning information; second, spatial ranging information between the robot and the target base station is acquired using an ultra-wideband positioning module as the second positioning information; and third, information characterizing the robot's pose state is acquired using an inertial measurement unit (IMU) as the third positioning information. These three types of positioning information are fused to obtain the robot's current positioning result. Furthermore, in response to voice interaction commands, the voice interaction commands are processed to obtain the target command, and the robot is controlled to execute the target action based on the target command and the positioning result. This method, by fusing environmental map, spatial ranging information, and pose state information, achieves joint estimation of the robot's position and attitude, improving the accuracy and stability of the positioning result. Simultaneously, the introduction of a voice interaction command-based control mechanism enables the robot to execute actions matching the target command based on the obtained accurate positioning result, enhancing the robot's autonomous operation capability in complex environments and the convenience and intelligence of human-machine interaction. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart of a robot control method according to an embodiment of this application is shown; Figure 2 A schematic diagram of the robot architecture in the robot control method of this application is shown; Figure 3 Another flowchart of the robot control method according to an embodiment of this application is shown; Figure 4 A schematic diagram of the operation of the lidar in the robot control method of this application is shown; Figure 5 This paper shows a schematic diagram of the operation of the ultra-wideband positioning module in the robot control method of this application embodiment; Figure 6 Another flowchart of the robot control method according to an embodiment of this application is shown; Figure 7 This paper shows a schematic diagram of the robot receiving voice signals in the robot control method of an embodiment of this application; Figure 8 A schematic diagram of a robot control method according to an embodiment of this application is shown. Detailed Implementation

[0019] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.

[0020] The components of the embodiments of this application described and illustrated in the accompanying drawings can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of this application provided in the drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0021] In the following text, the terms "comprising," "having," and their cognates, which may be used in various embodiments of this application, are intended only to indicate a particular feature, number, step, operation, element, component, or combination thereof, and should not be construed as primarily excluding the presence of one or more other features, numbers, steps, operations, elements, components, or combinations thereof, or adding the possibility of one or more combinations thereof. Furthermore, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0022] Unless otherwise specified, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments of this application pertain. Terms (such as those defined in commonly used dictionaries) shall be interpreted as having the same meaning as in their contextual meaning in the relevant technical field and shall not be construed as having an idealized or overly formal meaning, unless clearly defined in the various embodiments of this application.

[0023] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0024] Considering the problems of insufficient positioning accuracy, easy drift of positioning results, and low human-machine interaction efficiency in high-noise environments of industrial robots in metal-intensive workshops, a robot control method is proposed. This method utilizes LiDAR to acquire environmental map information, an ultra-wideband positioning module (UWB module) to acquire spatial ranging information between the robot and the target base station, and an inertial measurement unit (IMU module) to acquire information representing the robot's pose state. The multi-source positioning information is fused to obtain the robot's current accurate positioning result. When responding to voice interaction commands, the voice interaction commands are processed to obtain the target command. Based on the target command and the positioning result, the robot is controlled to execute the target action.

[0025] The robot control method will be explained below with reference to some specific embodiments.

[0026] Figure 1 A flowchart of a robot control method according to an embodiment of this application is shown. Exemplarily, this robot control method is applied to a robot equipped with a lidar, an ultra-wideband positioning module, and an inertial measurement unit, and includes the following steps: In step S100, environmental map information is acquired through LiDAR as the first positioning information, spatial ranging information between the robot and the target base station is acquired through the ultra-wideband positioning module as the second positioning information, and information characterizing the robot's pose state is acquired through the inertial measurement unit as the third positioning information.

[0027] The first positioning information refers to the environmental map information constructed based on the scanning of the robot's surroundings by LiDAR, used to represent the distribution of obstacles, passable areas, and the spatial structure corresponding to the robot's coordinate system; the second positioning information refers to the spatial ranging information obtained from the ranging process between the ultra-wideband positioning module and the target base station deployed on the ceiling of the workshop, used to characterize the spatial distance between the robot and the target base station; the third positioning information refers to the pose state information formed after the acceleration and angular velocity data output by the inertial measurement unit are processed by attitude calculation and compensation, used to characterize the robot's attitude and position in three-dimensional space. For example, Figure 2 As shown, the lidar, anti-metal UWB antenna, six-axis IMU, and robot main control box (also known as robot controller) installed on the robot body work synchronously to collect point cloud data, ranging signals, and motion state data from the serial port (RS232), SPI interface, and I2C interface, respectively. The pre-processed environmental map information, spatial ranging information, and pose state information are used as the first positioning information, the second positioning information, and the third positioning information, respectively.

[0028] In one alternative embodiment, such as Figure 3 As shown, step S100 includes the following sub-steps: S110: Acquire environmental point cloud data collected by lidar, and construct environmental map information based on the environmental point cloud data as the first positioning information.

[0029] Environmental point cloud data refers to the set of multiple points obtained by industrial-grade LiDAR when scanning the robot's surrounding environment. Each point contains at least two dimensions: spatial coordinates and reflection intensity value. The reflection intensity value is a unitless normalized grayscale value output by the device, typically ranging from 0 to 4095, used to distinguish highly reflective metallic objects from ordinary target surfaces. For example, the LiDAR's scanning frequency can be any value between 5 Hz and 50 Hz, and the horizontal scanning range can be any angle between 180° and 360°. Specific parameters are configured according to the dynamic requirements and perception coverage needs of the application scenario: a higher scanning frequency can be selected in scenarios requiring rapid response to moving obstacles; in environments with complex spatial structures or multiple interference sources, an omnidirectional scanning mode is used to achieve all-round environmental perception. Environmental map information refers to the map data obtained after filtering the environmental point cloud data by reflection intensity, coordinate transformation, and rasterization, used to describe the distribution of obstacles and passable areas within the robot's working area.

[0030] For example, such as Figure 4 As shown, the lidar is installed on the top of the robot body at a height of about 1.5m above the ground. It continuously scans with the laser scanning plane parallel to the ground (e.g., 360° scan, 10Hz frequency). After each frame of scanning, it reads the current environmental point cloud data, performs threshold segmentation and weight calculation on the reflection intensity value of each sampling point in the point cloud, and projects the filtered target point cloud onto a two-dimensional grid coordinate system to generate the current environmental map information. This environmental map information is then marked as the first positioning information.

[0031] In one optional implementation, step S110 includes the following sub-steps: Extract the reflection intensity values ​​corresponding to each spatial sampling point in the environmental point cloud data. Compare the reflection intensity values ​​with a preset intensity threshold, retaining point cloud data with reflection intensity values ​​less than or equal to the preset threshold and discarding point cloud data with reflection intensity values ​​greater than the preset threshold to obtain the target point cloud data. Perform localization and mapping operations based on the target point cloud data to generate the current environmental map information.

[0032] The reflection intensity threshold is an upper limit set for highly reflective objects such as metals. In this embodiment, it can be set to 2000. When the reflection intensity of a spatial sampling point in the point cloud exceeds this threshold, it is determined to be a high-reflection interference point. The target point cloud data refers to the set of valid point clouds retained after threshold filtering. The point cloud weights can be calculated by weighting the observations according to a Gaussian distribution. The weight calculation expression is as follows: ; Where z is the current point cloud observation value, For predicted values, This represents the observation variance.

[0033] Exemplary approach: First, the reflection intensity values ​​corresponding to each sampling point are extracted from the environmental point cloud data. Each reflection intensity value is compared with a preset intensity threshold of 2000. Point cloud data with reflection intensity values ​​less than or equal to the threshold are retained, while those with reflection intensity values ​​greater than the threshold are discarded to obtain target point cloud data. Then, for the observation values ​​of each point in the target point cloud data, the corresponding weights are calculated using the aforementioned Gaussian weighting formula. Next, using the robot's body coordinate system as a reference, the target point cloud data is mapped to a preset two-dimensional map coordinate system and rasterized at a fixed resolution (e.g., 0.05m / pixel). Particle filter SLAM or Gmapping algorithms are used to complete pose estimation and map updates, generating current environmental map information. Finally, the generated environmental map information is used as the first localization information.

[0034] S120 receives wireless ranging signals through an ultra-wideband positioning module, performs dynamic frequency adjustment processing on the wireless ranging signals, and measures the spatial ranging information between the robot and the target base station based on the frequency-adjusted wireless ranging signals as the second positioning information.

[0035] Among them, the wireless ranging signal refers to the ranging-specific signal transmitted by the target base station to the robot's ultra-wideband positioning module and received by the module within the 3.1-10.6 GHz frequency hopping range, such as... Figure 5 As shown; frequency dynamic adjustment processing refers to compensating for and frequency hopping the carrier frequency based on the relative speed between the robot and the base station (where the distance between the base stations is 5-8m), in order to reduce the ranging error caused by multipath effects and frequency shift in the metallic environment. The frequency shift can be calculated according to the Doppler frequency shift formula:

[0036] Where v is the relative velocity and c is the speed of light. The carrier frequency is used; spatial ranging information refers to the distance between the robot and the target base station and its related confidence parameters estimated based on the signal samples of the direct path.

[0037] Exemplary, the ultra-wideband positioning module adopts a dual-antenna differential receiving structure (50cm spacing), dynamically adjusting the operating frequency three times per second. After completing the dynamic frequency adjustment, it performs differential operations and constructs feature tensors on the received ranging signals, and inputs the feature tensors into a pre-trained convolutional neural network model to classify direct signals and multipath signals. Finally, it uses the signal samples determined to be direct paths to calculate spatial ranging information, and uses this ranging information as the second positioning information.

[0038] In one optional implementation, step S120 includes the following sub-steps: The robot receives wireless ranging signals via an ultra-wideband positioning module and performs differential processing on these signals to obtain differential ranging signals. These differential ranging signals are then used to construct input feature tensors for signal analysis and fed into a pre-defined convolutional neural network model. The model performs type identification on the differential ranging signals to obtain direct signal samples representing the direct path between the robot and the target base station. Based on these direct signal samples, the spatial ranging information between the robot and the target base station is determined.

[0039] Among them, differential ranging signal refers to the signal sequence obtained by performing differential operation on the original wireless ranging signal in the time dimension, which is used to reduce fixed bias and slow-varying interference; input feature tensor is a two-dimensional data structure formed by rearranging differential ranging signal according to sampling frequency and sample length. In this embodiment, a time-series signal with a length of 128 points and a sampling frequency of 200Hz can be combined into a [128,1] feature tensor; convolutional neural network model (CNN) includes multi-layer convolution, pooling and fully connected structures, used for feature extraction and binary classification of input feature tensor; direct signal sample is the ranging sample that the CNN model classifies as Loss of Sight (LoS, Line-of-Sight, direct path); direct path probability It can be estimated using logistic regression, for example: ; Where S is the signal strength and t is the timestamp. , is the weight parameter, and b is the bias term.

[0040] Exemplarily, firstly, a wireless ranging signal is received through an ultra-wideband positioning module, and differential processing is performed on the signal to obtain a differential ranging signal. Subsequently, the differential ranging signal is reorganized in chronological order into a temporal intensity sequence of length 128, and an input feature tensor of [128,1] is constructed and input into a pre-defined convolutional neural network model. Inside the CNN model, the signal is processed sequentially as follows: First convolutional layer: kernel size 3×1, number of kernels 32, stride 1, activation function ReLU; First pooling layer: using max pooling, pooling window size 2; Second convolutional layer: kernel size 3×1, number of kernels 64, activation function ReLU; Second pooling layer: pooling window size 2; Fully connected layer: output dimension 128, activation function ReLU; Classification output layer: using the Softmax activation function, outputting two probability values, corresponding to the direct signal (LoS) and multipath signal (NLoS), respectively. This yields the probability that a sample belongs to a direct signal or a multipath signal; when When the percentage is greater than a preset threshold (≥92%), the corresponding sample is determined to be a direct signal sample. Based on the flight time information corresponding to these direct signal samples, the spatial ranging information between the robot and the target base station is calculated, and the ranging result and its confidence level are combined to form the second positioning information.

[0041] S130: Obtain the robot's initial pose state information based on the motion state data collected by the inertial measurement unit, perform compensation processing on the initial pose state information to obtain the corrected pose state information, which is used as the third positioning information.

[0042] The motion state data refers to the three-axis acceleration and three-axis angular velocity data output by the six-axis IMU sensor during robot motion. The initial pose state information refers to the preliminary posture and position estimation results of the robot calculated based on acceleration and angular velocity data using algorithms such as complementary filtering. The corrected pose state information refers to a more accurate pose estimation obtained by updating the initial pose state information using Kalman filtering in conjunction with external space ranging information. Exemplarily, the IMU module collects acceleration and angular velocity data, calculates and fuses the attitude angle using complementary filtering, and determines whether the current posture significantly deviates from the horizontal state based on a preset tilt threshold. When the tilt angle exceeds, for example, 5°, the Kalman filtering update process is triggered, incorporating UWB ranging information as an observation into the state estimation to obtain the corrected pose state information, which is then used as the third positioning information in subsequent positioning fusion.

[0043] In one optional implementation, step S130 includes the following sub-steps: The system acquires acceleration and angular velocity data output from the inertial measurement unit (IMU). Complementary filtering is performed based on the acceleration and angular velocity data to generate initial pose state information characterizing the robot's attitude. If the initial pose state information meets a preset tilt threshold condition, Kalman filtering is performed based on the initial pose state information and spatial ranging information to generate corrected pose state information.

[0044] Among them, acceleration data is a three-axis signal used to reflect the change of linear acceleration of the robot, and angular velocity data is a three-axis signal used to reflect the angular motion state of the robot; complementary filtering is used to fuse the angle obtained by gyroscope integration and the attitude angle obtained by acceleration calculation, and its fusion formula can be expressed as: ; in, The attitude angle is obtained by integrating the angular velocity. The attitude angles obtained from acceleration calculations. To integrate weighting factors, Inclination angle Defined as pitch angle With roll angle The maximum value among absolute values, i.e. ; The Kalman filter update process can be described by the following formula: ; ; ; Where K is the Kalman gain, P is the state covariance matrix, H is the observation matrix, R is the observation noise covariance matrix, x is the state vector, and z is the observation vector.

[0045] Exemplary approach: First, acquire the triaxial acceleration and triaxial angular velocity data output by the inertial measurement unit, and calculate the fused attitude angles based on the aforementioned complementary filtering formula to obtain initial pose state information characterizing the robot's attitude state; then, based on the pitch angle in the initial pose state information... With roll angle Calculate the tilt angle ,when When the value is greater than 5, the robot is determined to be in a posture state that needs correction. At this time, the spatial ranging information at the corresponding moment is obtained. The initial pose state information and the spatial ranging information are combined and the Kalman gain is calculated according to the above Kalman filter formula. The pose state vector is then updated to obtain the corrected pose state information. Finally, the corrected pose state information is used as the third localization information.

[0046] Step S200: The first positioning information, the second positioning information and the third positioning information are fused to obtain the robot's current positioning result.

[0047] Demonstratively, after independently processing the data from the LiDAR, UWB, and inertial measurement unit (IMU), the three types of positioning information are aligned with a unified timestamp to construct corresponding LiDAR positioning information, UWB positioning information, and IMU positioning information. A preset weighted fusion strategy is then used to calculate the positioning result representing the robot's current two-dimensional planar pose. For example, in a battery cell workshop application scenario, after fusing the above three types of information, the original positioning error of approximately 5cm can be compressed to ±1cm. Specifically, the measured error at the battery cell stacking station can be controlled to approximately 0.8cm, thus meeting the positioning requirements for precision material docking and assembly.

[0048] In one alternative embodiment, such as Figure 6 As shown, step S200 includes the following sub-steps: S210, determine the lidar positioning information based on the first positioning information.

[0049] Among them, the lidar positioning information refers to the pose estimation result of the robot on the global map determined based on the first positioning information (environmental map information); for example, based on the current environmental map information output by lidar SLAM and the particle filter pose estimation result, the two-dimensional coordinates and orientation angle of the robot in the global map coordinate system are extracted from the first positioning information, and the pose estimation is recorded as lidar positioning information.

[0050] S220, determine ultra-wideband positioning information based on the second positioning information.

[0051] Among them, ultra-wideband positioning information refers to the robot position constraints calculated in the base station coordinate system based on the second positioning information (spatial ranging information); for example, based on the spatial ranging information between the robot and the target base station output by the UWB module, under the known base station installation position coordinates, multiple ranging values ​​are converted into the robot's position constraints in the workshop plane, and ultra-wideband positioning information is formed. S230, determine inertial positioning information based on third positioning information.

[0052] Among them, inertial positioning information refers to the short-time integral pose estimation derived based on the third positioning information (corrected pose state information); for example, based on the acceleration and angular velocity output by the IMU module, the corrected pose state information obtained by complementary filtering and Kalman filtering is used to extract the displacement and attitude changes of the robot within a short time window, and the estimation results are organized into inertial positioning information. S240, respectively obtain the weight coefficients corresponding to the lidar positioning information, ultra-wideband positioning information and inertial positioning information.

[0053] The weighting coefficients are the weighting parameters assigned to the lidar positioning information, ultra-wideband positioning information, and inertial positioning information respectively in the weighted fusion calculation. They reflect the reliability of each type of information under the current environment and operating conditions. For example, weighting coefficients are configured for the lidar positioning information, ultra-wideband positioning information, and inertial positioning information respectively, based on the observation accuracy and historical statistical errors of each sensor under the current operating conditions. It can also be dynamically adjusted according to environmental changes; S250 performs weighted fusion calculations on the lidar positioning information, ultra-wideband positioning information, and inertial positioning information with their respective weight coefficients to obtain a weighted fusion result, and generates a positioning result to characterize the robot's current two-dimensional planar pose based on the weighted fusion result.

[0054] As an example, the three types of localization information are weighted and fused according to a weighted fusion formula to obtain a localization result that characterizes the robot's current two-dimensional planar pose. For example, the following weighted fusion expression can be used:

[0055] in, This indicates the location result after fusion. These represent the pose estimates obtained from lidar positioning information, ultra-wideband positioning information, and inertial positioning information, respectively. These are the corresponding weighting coefficients.

[0056] For example, in a battery cell workshop with dense metal shelves, when the LiDAR field of view is obstructed and the point cloud quality deteriorates, the weight of UWB ranging information and IMU inertial information is increased and the weight of LiDAR positioning information is decreased by statistically analyzing historical errors. This maintains high positioning stability under multipath effects and local obstruction conditions. Conversely, when the robot moves to an open area and the SLAM matching quality is high, the weight of LiDAR positioning information can be increased to make full use of high-resolution map constraints.

[0057] Step S300: In response to the voice interaction command, process the voice interaction command to obtain the target command, and control the robot to execute the target action based on the target command and the positioning result, such as... Figure 6 As shown.

[0058] Among them, voice interaction commands refer to control requests issued by operators via voice within the robot's working area, such as "transfer the battery cell at workstation B12 to production line 3," etc. Target commands refer to structured control commands that can directly drive the robot to execute, obtained after noise reduction, recognition, association, and matching processing of the voice interaction commands. The positioning result refers to the data used to characterize the robot's current two-dimensional planar pose obtained through the fusion of the aforementioned multi-source sensors. Exemplarily, the microphone array on top of the robot collects the operator's voice signal. After spectral subtraction noise reduction and speech recognition are performed by the speech processing unit, a speech recognition result is generated. Subsequently, the speech processing unit combines a preset command library and recent task log information to perform command association, generating several candidate commands, and selecting the target command through fuzzy matching based on edit distance. Finally, the robot's main control system generates corresponding control commands based on the target command and the current positioning result, driving the robot to perform corresponding handling, alignment, or placement actions.

[0059] In one optional implementation, step S300 includes the following sub-steps: S310 receives the voice signal corresponding to the voice interaction command, performs noise reduction and recognition processing on the voice signal, and generates a voice recognition result.

[0060] Here, the speech signal refers to the raw audio data collected by a microphone array, containing speech components and workshop environmental noise. Noise reduction processing refers to the process of attenuating the speech signal's spectrum and updating noise estimation to address background interference such as mechanical noise and wind noise. The speech recognition result refers to the text sequence obtained after acoustic modeling and decoding the denoised speech signal. For example, as... Figure 7 As shown, six omnidirectional microphones (30cm apart, 360° coverage) are mounted in a circular array on the top of the robot, and two directional microphones (10° front focus) are mounted on the front, sharing the same mounting bracket as the omnidirectional microphones. A 3M filter with a 0.1mm aperture is placed in front of the microphone array to physically attenuate mechanical noise. The voice signal is sent to the voice processing unit (preferably an NVIDIA Jetson AGX Xavier processor integrated into the robot controller, connected to the microphone array via an I2S interface) through an I2S interface. First, spectral subtraction noise reduction is performed, using the formula... Each frame of spectrum is processed, where This represents the speech spectrum of the current frame. Represents the noise estimation spectrum. Here is the spectral reduction coefficient. Noise estimation can be recursively derived using the following formula:

[0061] in, This is the noise update factor (e.g., 0.7). This is a noise estimate for the previous frame.

[0062] The speech data, after spectral subtraction, is input into a lightweight Wenet model (Transformer architecture) for speech recognition, and the corresponding text sequence is output as the speech recognition result. For example, in a battery cell workshop with a noise environment of 70–90 dB, after joint processing by physical filters and spectral subtraction, the speech processing unit can output stable speech recognition results while ensuring a recognition latency of less than 100 ms, which can then be used for subsequent command association and matching. In other implementations, the speech recognition model can be trained offline using knowledge distillation, and its loss function can be expressed as: ; in, For large models, soft tags, Output for small model This is the temperature coefficient.

[0063] S320 performs instruction association processing based on speech recognition results to generate instruction candidate information.

[0064] Among them, instruction association processing refers to, given the speech recognition result, combining the current work scenario, task history and environmental context information, filtering and generating several candidate instructions that are highly relevant to the current context from a preset instruction library; instruction candidate information refers to a candidate set consisting of several candidate instructions and corresponding context labels, which is used in the subsequent fuzzy matching step to select the final target instruction.

[0065] For example, after receiving the speech recognition result, the system activates a set of basic instructions related to the current process (e.g., "cell stacking operation area") from a preset task library, such as "material picking", "barcode scanning", "inspection", and "placement". Then, the system searches the operation history of the last 10 minutes in the task log, counts the frequently occurring instruction sequences (e.g., executing "barcode scanning-placement" multiple times in a row), and adds these high-frequency instructions as priority candidates to the instruction candidate information. At the same time, the system generates contextual information by combining the current material identifier, the material ID of the last operation, and the record of the most recent manual operation completion. For example, it adds instructions such as "continue to move type B cells" and "continue to process material A031" to the candidate set.

[0066] For example, when the relative distance d(i,j) between two employees or between an employee and materials is less than 2m, a command association mechanism based on the current speech recognition result can be triggered according to the sensor perception result, so as to reduce the verbal length of complex commands and improve the efficiency of task issuance.

[0067] S330 performs fuzzy matching processing on the speech recognition result and the command candidate information to obtain the target command.

[0068] Fuzzy matching refers to measuring the similarity between the speech recognition result and each candidate instruction by calculating the edit distance when there are recognition errors or incomplete descriptions, and selecting the candidate instruction with the smallest distance that meets a preset threshold as the target instruction. The target instruction is the structured instruction that is most similar to the speech recognition result and can be executed in the current context, as determined after fuzzy matching. For example, the system calculates the edit distance d(i,j) between the speech recognition result and each candidate instruction based on the Levenshtein distance, and its recursive relationship can be expressed as: ; Where d(i,j) represents the edit distance between the recognized text prefix of length i and the candidate instruction prefix of length j. This represents the cost incurred for characters being the same or different at the current position (0 for the same, 1 for different). When the edit distance corresponding to a candidate instruction is less than a preset threshold (e.g., 2), the candidate instruction is considered to be the target instruction. If multiple candidate instructions simultaneously meet the threshold condition, the one with the smallest edit distance and consistent with the current context (e.g., the ID of the most recent material transported) can be selected as the final target instruction.

[0069] For example, when the operator says "Continue processing the previous one", the system calculates the edit distance between the voice recognition result and the candidate instruction set consisting of the task log and the current material information. It also combines the records of material operations in the last 5 minutes (e.g., material ID=A031) and the current process (e.g., warehouse entry verification) to generate candidate instructions such as "Continue to verify material A031", "Continue to scan material A031", and "Place A031 to the current workstation". The final target instruction is obtained by comprehensively judging the edit distance and context consistency.

[0070] S340 generates control commands based on target instructions and positioning results to control the robot to perform target actions.

[0071] Among them, control instructions refer to execution commands that can be directly issued to the robot motion control system after combining the semantic content of the target instructions and the robot's current positioning results. Target actions refer to the specific operations that the robot performs in physical space according to the control instructions, such as moving, turning, grasping, and placing.

[0072] For example, when the target instruction is "place A031 to the designated workstation on production line 3", the robot's main control system first reads the current positioning results, including the robot's position coordinates and heading angle in the workshop plane coordinate system. Combined with the target coordinates of the designated workstation on production line 3 in the workshop map, it calls the path planning module to generate a motion trajectory from the current position to the target position. Subsequently, it calculates the corresponding speed command and attitude adjustment parameters based on the generated trajectory, encapsulates them into control commands, and sends them to the robot motion control unit through the internal bus to drive the robot to move along the planned path and complete the placement operation at the target position. After the task is completed, the voice processing unit can also broadcast execution status information such as "transfer completed" and "placement completed" through the speaker.

[0073] In other implementations, the generation of control commands can also be combined with obstacle avoidance strategies to perform collision detection on the environmental map information output by the LiDAR and the current positioning results, and automatically avoid static shelves and dynamic pedestrians in the planned trajectory to ensure the robot's operational safety during the execution of target actions.

[0074] In one optional implementation, the robot control method further includes the following sub-steps: After receiving the completion information of the robot's target action, S400 generates the corresponding execution status information.

[0075] Among them, completion information refers to the task status data reported by the robot motion control unit or the host control system when the target action is completed, including at least the task identifier, start time, end time and execution result mark; execution status information refers to the task status description data formed after the completion information is structured and organized, which is used to represent the execution status of the task corresponding to the current voice command.

[0076] For example, after the robot completes the task of "transferring type B battery cells to workstation 3 on production line" according to the target instruction, the robot's main control system encapsulates information such as task completion flag, actual completion time, and whether there are any abnormal shutdowns or obstacle avoidance events into completion information and reports it to the voice processing unit or task management module. The task management module generates execution status information based on the completion information, such as constructing structured text such as "Task number T001 has been completed, material type B battery cells have been delivered to the designated workstation on production line 3, and there were no abnormalities in the execution process", for subsequent voice broadcasting and human-machine feedback.

[0077] For example, in a battery cell workshop scenario, the system can store the execution status information along with the aforementioned task logs for subsequent voice command association of "the task we just completed" and historical task tracing.

[0078] The S500 converts execution status information into voice broadcast content and provides feedback on task execution status through the voice output unit.

[0079] The voice broadcast content refers to text content or voice data generated based on execution status information that is suitable for voice synthesis and broadcasting, used to prompt the operator with the task execution results in natural language form; the voice output unit can be a speaker integrated into the robot body or other audio output devices connected to the voice processing unit.

[0080] For example, after receiving the execution status information, the voice processing unit concatenates the task number, material information, target workstation, and execution result marker into a broadcast text according to a preset template, such as "The handling task has been completed, the current material has arrived at the target workstation, please continue to the next operation", and calls the voice synthesis module on the same hardware platform as the aforementioned voice recognition to generate voice data; then, it broadcasts the information in real time through a speaker electrically connected to the voice processing unit, so that the operator can know the task execution status without looking at the screen in a high-noise environment.

[0081] For example, during the entire voice interaction process, the overall delay from receiving the execution completion signal to broadcasting the execution result can be controlled at the 100ms level. It shares the same processing platform with the aforementioned voice recognition link, thereby ensuring that the human-computer interaction in the industrial field has good real-time performance and continuity.

[0082] In other implementations, the voice broadcast content can also automatically adjust the volume or the number of broadcasts based on the current workshop noise level.

[0083] Figure 8 A schematic diagram of a robot according to an embodiment of this application is shown. Exemplarily, the robot 10 includes: a robot body 100, and a lidar module 110, an ultra-wideband positioning module 120, an inertial measurement module 130, a control module 140, and a voice interaction module 150 disposed on the robot body; The lidar module 110 is used to acquire environmental map information as the first positioning information; The ultra-wideband positioning module 120 is used to acquire spatial ranging information between the robot and the target base station as second positioning information. The inertial measurement module 130 is used to acquire information characterizing the robot's pose state as third positioning information; The voice interaction module 150 is used to receive voice interaction commands; The control module 140 is used to fuse the first positioning information, the second positioning information and the third positioning information to obtain the current positioning result of the robot; The control module 140 is further configured to respond to the voice interaction command, process the voice interaction command to obtain a target command, and control the robot body to perform a target action based on the target command and the positioning result.

[0084] It is understood that the system in this embodiment corresponds to the method in the above embodiments, and the options in the above embodiments are also applicable to this embodiment, so they will not be described again here.

[0085] This application also provides a robot, exemplary in that the robot includes a processor and a memory, wherein the memory stores a computer program, and the processor, by running the computer program, causes the robot to perform the functions of the various modules in the above-described method or system.

[0086] The processor can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, including at least one of a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Network Processor (NP), Digital Signal Processor (DSP), Application-Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0087] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), etc. The memory is used to store computer programs, and the processor can execute the computer programs accordingly after receiving execution instructions.

[0088] This application also provides a computer-readable storage medium for storing the computer program used in the robot described above. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0089] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can also be implemented in other ways. The system embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that, as an alternative implementation, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0090] In addition, the functional modules or units in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0091] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a smartphone, personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0092] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A robot control method, applied to a robot equipped with a lidar, an ultra-wideband positioning module, and an inertial measurement unit, characterized in that, The method includes: The robot obtains environmental map information through the lidar as the first positioning information, obtains spatial ranging information between the robot and the target base station through the ultra-wideband positioning module as the second positioning information, and obtains information characterizing the robot's pose state through the inertial measurement unit as the third positioning information. The first positioning information, the second positioning information, and the third positioning information are fused together to obtain the current positioning result of the robot. In response to a voice interaction command, the voice interaction command is processed to obtain a target command, and the robot is controlled to perform a target action based on the target command and the positioning result.

2. The robot control method according to claim 1, characterized in that, The process of acquiring environmental map information via the lidar as first positioning information, acquiring spatial ranging information between the robot and the target base station via the ultra-wideband positioning module as second positioning information, and acquiring information characterizing the robot's pose state via the inertial measurement unit as third positioning information includes: The environmental point cloud data collected by the lidar is acquired, and environmental map information is constructed based on the environmental point cloud data as the first positioning information; The ultra-wideband positioning module receives wireless ranging signals and measures the spatial distance information between the robot and the target base station based on the wireless ranging signals, which is used as the second positioning information; wherein, the ultra-wideband positioning module adjusts the carrier signal frequency at preset intervals during the process of receiving the wireless ranging signals. The robot's initial pose information is obtained based on the motion state data collected by the inertial measurement unit. Compensation processing is performed on the initial pose information to obtain corrected pose information, which is used as the third positioning information.

3. The robot control method according to claim 2, characterized in that, The step of acquiring the environmental point cloud data collected by the lidar and constructing environmental map information based on the environmental point cloud data includes: Extract the reflection intensity value corresponding to each spatial sampling point in the environmental point cloud data; The reflection intensity value is compared with a preset intensity threshold. Point cloud data with a reflection intensity value less than or equal to the preset intensity threshold are retained, and point cloud data with a reflection intensity value greater than the preset intensity threshold are discarded to obtain the target point cloud data. Based on the target point cloud data, localization and mapping operations are performed to generate current environment map information.

4. The robot control method according to claim 2, characterized in that, The step of receiving wireless ranging signals through the ultra-wideband positioning module and measuring spatial ranging information between the robot and the target base station based on the wireless ranging signals includes: The ultra-wideband positioning module receives wireless ranging signals and performs differential processing on the wireless ranging signals to obtain differential ranging signals. The differential ranging signal is constructed into an input feature tensor for signal analysis and input into a preset convolutional neural network model to perform type identification on the differential ranging signal in order to obtain a direct signal that represents the direct path between the robot and the target base station. The spatial ranging information between the robot and the target base station is determined based on the direct signal.

5. The robot control method according to claim 2, characterized in that, The initial pose state information of the robot is obtained based on the motion state data collected by the inertial measurement unit, and compensation processing is performed on the initial pose state information to obtain the corrected pose state information, including: Acquire the acceleration and angular velocity data output by the inertial measurement unit; Complementary filtering is performed based on the acceleration data and the angular velocity data to generate the initial pose state information used to characterize the robot's posture state; When the initial pose state information meets the preset tilt threshold condition, Kalman filtering is performed based on the initial pose state information and the spatial ranging information to generate the corrected pose state information.

6. The robot control method according to claim 1, characterized in that, The step of fusing the first positioning information, the second positioning information, and the third positioning information to obtain the robot's current positioning result includes: The lidar positioning information is determined based on the first positioning information; Ultra-wideband positioning information is determined based on the second positioning information; Inertial positioning information is determined based on the third positioning information; Obtain the weighting coefficients corresponding to the lidar positioning information, the ultra-wideband positioning information, and the inertial positioning information, respectively; The lidar positioning information, the ultra-wideband positioning information, and the inertial positioning information are weighted and fused with their respective weight coefficients to obtain a weighted fusion result, and the robot's current positioning result is generated based on the weighted fusion result.

7. The robot control method according to claim 1, characterized in that, The process of responding to a voice interaction command, processing the voice interaction command to obtain a target command, and controlling the robot to perform a target action based on the target command and the positioning result includes: Receive the voice signal corresponding to the voice interaction command, and perform noise reduction and recognition processing on the voice signal to generate a voice recognition result; Based on the speech recognition results, instruction association processing is performed to generate instruction candidate information; The speech recognition result and the instruction candidate information are subjected to fuzzy matching processing to obtain the target instruction; Based on the target instruction and the positioning result, control instructions are generated to control the robot to perform the target action.

8. The robot control method according to claim 7, characterized in that, The method further includes: After receiving the completion information of the robot's target action, corresponding execution status information is generated; The execution status information is converted into voice broadcast content, and the task execution status is fed back through the voice output unit.

9. A robot, characterized in that, include: The robot body, a lidar module, an ultra-wideband positioning module, an inertial measurement module, a control module, and a voice interaction module are mounted on the robot body; The lidar module is used to acquire environmental map information as the first positioning information; The ultra-wideband positioning module is used to acquire the spatial ranging information between the robot and the target base station as the second positioning information; The inertial measurement module is used to acquire information characterizing the robot's pose state as third positioning information; The voice interaction module is used to receive voice interaction commands; The control module is used to fuse the first positioning information, the second positioning information and the third positioning information to obtain the current positioning result of the robot; The control module is further configured to respond to the voice interaction command, process the voice interaction command to obtain a target command, and control the robot body to perform a target action based on the target command and the positioning result.

10. A robot controller, characterized in that, The robot controller stores a computer program, which is used to execute the computer program to implement the robot control method according to any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed on a processor, implements the robot control method according to any one of claims 1-8.