Control system of remote robot
Through the multimodal fusion system of data acquisition equipment, controller and user-side, the problem of multi-sensor data fusion and motion control accuracy in complex assembly tasks is solved, and the high-precision operation of remote robots and user-friendly feedback is achieved, which improves the efficiency and accuracy of remote assembly.
Patent Information
- Application Number
- CN202510561441.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional robot control systems have shortcomings in multi-sensor data fusion, motion control accuracy and human robot interaction under complex assembly tasks and high-precision requirements, especially in remote operations, it is difficult to efficiently and accurately map human operations to robot end effectors.
The combined system of data acquisition equipment, controller and user-side is adopted to extract video and audio features through SIFT and MFCC algorithms, combine adaptive weighting fusion strategies, dynamically adjust weights, realize multimodal data fusion, control remote robots for assembly, and provide multi-dimensional feedback information.
It significantly improves the control accuracy and operating experience of the remote robot, ensures smooth and continuous motion control, provides comprehensive and intuitive user feedback, and improves the efficiency and accuracy of remote assembly tasks.
Smart Images

Figure CN120244973A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a control system for a remote robot. Background Art
[0002] With the continuous development of automation and remote control technologies, robots have been widely used in various industrial and service scenarios. Although traditional robot control systems can achieve basic operation tasks, there are still some problems when facing complex assembly tasks and high-precision requirements, especially in the fusion of multi-sensor data, the accuracy of motion control, and the interaction between robots and humans. In remote operation tasks, how to efficiently and accurately map human operations to the end effector of the robot and improve the flexibility and accuracy of the robot remains an urgent problem to be solved.
[0003] Based on this, the present invention proposes a control system for a remote robot to solve the above technical problems. Summary of the Invention
[0004] The present invention describes a control system for a remote robot, which can effectively improve the control accuracy of the remote robot.
[0005] The present invention provides a control system for a remote robot, including a data acquisition device, a controller, a remote robot, and a user terminal;
[0006] The controller is respectively communicatively connected to the data acquisition device, the remote robot, and the user terminal;
[0007] The data acquisition device is used to collect video data and audio data of the user within the same preset time period;
[0008] The controller is used to perform the following operations: receiving the video data and the audio data sent by the data acquisition device; sequentially performing feature extraction on the video data and the audio data to obtain video features and audio features; performing feature fusion on the video features and the audio features according to the quality dynamic weight to obtain fusion features; determining the motion posture based on the fusion features; controlling the remote robot to perform remote assembly according to the motion posture; updating the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset duration, receiving the video data and the audio data sent by the data acquisition device again;
[0009] The user terminal is used to receive the visual data, force sense data, and tactile data of the remote robot during the remote assembly process sent by the controller;
[0010] The remote robot is used to receive the motion posture sent by the controller and perform remote assembly according to the motion posture.
[0011] According to the control system of the remote robot provided by the present invention, the system includes a data acquisition device, a controller, a remote robot, and a user terminal. The controller establishes a two-way communication connection with the data acquisition device, the remote robot, and the user terminal through a high-speed communication link. The data acquisition device accurately acquires the video data and audio data of the user at a high sampling rate within the same preset time period (this preset time period will change dynamically as the cycle progresses). The controller is used to perform the following operations: receiving the video data and audio data sent by the data acquisition device, extracting key frames and matching feature points from the video data by using algorithms such as Scale-Invariant Feature Transform (SIFT) algorithm, and performing spectrum analysis and feature vector extraction on the audio data by combining with Mel Frequency Cepstral Coefficient (MFCC) algorithm, so as to generate the audio features for accurately representing the user's voice commands and the video features representing the user's action commands respectively. Subsequently, based on the adaptive weighted fusion strategy, according to the quality evaluation values determined by the quantization indexes such as the clarity of the video data and the frame rate stability, and the quality evaluation values determined by the quantization indexes such as the signal-to-noise ratio and the harmonic distortion rate of the audio data, the quality dynamic weights are dynamically adjusted, and the video features and the audio features are deeply fused to obtain the fusion features that can comprehensively and accurately represent the user's true intention. Based on the fusion features, the matching motion posture is determined, and the remote robot is accurately controlled to carry out remote assembly operations by virtue of this posture. At the same time, the quality of the video data and the audio data is continuously evaluated dynamically, and based on these quality data, the quality dynamic weights are intelligently updated to ensure that the fusion features always have high accuracy, thereby greatly improving the control accuracy of the remote robot. Whenever the preset duration ends, the interruption mechanism will be automatically triggered, and "acquiring the video data and audio data of the user within the same preset time period" will be executed again to ensure the smooth and uninterrupted continuous operation of the motion control of the remote robot. The user terminal is used to receive the visual data, force sense data, and tactile data of the remote robot during the remote assembly process sent by the controller; the visualization and feedback processing software equipped on the user terminal analyzes and presents these data to provide comprehensive and intuitive feedback information for the user. The remote robot is used to receive the motion posture sent by the controller and perform remote assembly according to the motion posture. Through the above configuration method, the present invention can effectively improve the control accuracy of the remote robot. Description of the Drawings
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 The figure shows a schematic block diagram of a control system of a remote robot according to an embodiment.
[0014] Reference numerals:
[0015] 11 - Data acquisition device;
[0016] 12 - Controller;
[0017] 13 - Remote robot;
[0018] 14 - User terminal. Detailed implementation manners
[0019] The solution provided by the present invention will be described below with reference to the accompanying drawings.
[0020] Figure 1 The figure shows a schematic block diagram of a control system of a remote robot 13 according to an embodiment. As Figure 1 shown, it includes a data acquisition device 11, a controller 12, a remote robot 13 and a user terminal 14;
[0021] The controller 12 is communicatively connected to the data acquisition device 11, the remote robot 13 and the user terminal 14 respectively;
[0022] The data acquisition device 11 is used to collect video data and audio data of the user within the same preset time period;
[0023] The controller 12 is used to perform the following operations: receive the video data and audio data sent by the data acquisition device 11; sequentially perform feature extraction on the video data and audio data to obtain video features and audio features; perform feature fusion on the video features and audio features according to the quality dynamic weight to obtain fusion features; determine the motion posture based on the fusion features; control the remote robot 13 to perform remote assembly according to the motion posture; update the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset time period, receive the video data and audio data sent by the data acquisition device 11 again;
[0024] The user terminal 14 is used to receive the visual data, force sense data and tactile data of the remote robot 13 during the remote assembly process sent by the controller 12;
[0025] The remote robot 13 is used to receive the motion posture sent by the controller 12 and perform remote assembly according to the motion posture.
[0026] In this embodiment, the system includes a data acquisition device 11, a controller 12, a remote robot 13, and a client 14. The controller 12 establishes a two-way communication connection with the data acquisition device 11, the remote robot 13, and the client 14 through a high-speed communication link. The data acquisition device 11 accurately acquires the user's video data and audio data at a high sampling rate within the same preset time period (this preset time period will change dynamically as the loop progresses). The controller 12 is used to perform the following operations: receive the video data and audio data sent by the data acquisition device 11, extract key frames and match feature points from the video data using an algorithm such as Scale-Invariant Feature Transform (SIFT), and perform spectral analysis and feature vector extraction on the audio data in combination with the Mel Frequency Cepstral Coefficient (MFCC) algorithm, thereby generating audio features that accurately represent the user's voice commands and video features that represent the user's action commands respectively. Subsequently, based on an adaptive weighted fusion strategy, according to the quality evaluation values determined by quantization indicators such as the clarity of the video data and the frame rate stability, and the quality evaluation values determined by quantization indicators such as the signal-to-noise ratio and harmonic distortion rate of the audio data, dynamically adjust the quality dynamic weights, and deeply fuse the video features and audio features to obtain fusion features that can comprehensively and accurately represent the user's true intention. Based on this fusion feature, determine the matching motion posture, and precisely control the remote robot 13 to carry out remote assembly operations with this posture. At the same time, continuously dynamically evaluate the quality of the video data and audio data, and based on these quality data, intelligently update the quality dynamic weights to ensure that the fusion features always have high accuracy, thereby greatly improving the control accuracy of the remote robot 13. Whenever the preset duration ends, an interruption mechanism will be automatically triggered, and "acquire the user's video data and audio data within the same preset time period" will be executed again to ensure that the motion control of the remote robot 13 proceeds smoothly and continuously without interruption. The client 14 is used to receive the visual data, force perception data, and tactile data of the remote robot 13 during the remote assembly process sent by the controller 12; the visualization and feedback processing software equipped on the client 14 analyzes and presents these data to provide comprehensive and intuitive feedback information to the user. The remote robot 13 is used to receive the motion posture sent by the controller 12 and perform remote assembly according to the motion posture. Through the above configuration method, the present invention can effectively improve the control accuracy of the remote robot 13.
[0027] In an embodiment of the present invention, the data acquisition device 11 includes a microphone array and a camera;
[0028] The microphone array is used to collect audio data;
[0029] The camera is used to collect video data.
[0030] In this embodiment, the data acquisition device 11 adopts a combined configuration of a microphone array and a camera. Among them, the microphone array consists of multiple high-sensitivity omnidirectional microphones, which directionally collect user voices through beamforming technology, effectively suppressing environmental noise interference and achieving high-quality acquisition of audio data; the camera is equipped with a high-resolution image sensor and an autofocus lens, supporting high-frame-rate video shooting and wide-angle vision, and can clearly capture user action details to ensure the integrity and clarity of video data. The two work together to provide accurate and reliable raw data for the system.
[0031] In an embodiment of the present invention, the user terminal 14 includes a display screen, a tactile sensor, and a force sensor;
[0032] The display screen is used to display visual data to the user;
[0033] The tactile sensor is used to transmit tactile data to the user;
[0034] The force sensor is used to transmit force data to the user.
[0035] In this embodiment, the user terminal 14 constructs a multi-modal interaction device system, which consists of a display screen, a tactile sensor, and a force sensor, aiming to provide users with a comprehensive and intuitive operation feedback experience.
[0036] Specifically, the display screen adopts high-resolution and high-refresh-rate display technology, and can present the visual data of the remote robot 13 during the remote assembly process to the user with clear and smooth pictures. Whether it is the fine details of parts or the overall assembly scene, users can obtain accurate visual information through the display screen to make accurate judgments and decisions. The tactile sensor uses advanced microelectromechanical system (MEMS) technology, which can accurately sense the surface texture, temperature and other tactile characteristics of the objects contacted by the remote robot 13, and transmit these tactile data to the user's hand in the form of vibration, pressure, etc., enabling the user to feel as if they are touching the objects in the remote environment, enhancing the realism and immersion of the operation. The force sensor is based on a high-precision force sensing principle, which measures the magnitude and direction of the force received by the remote robot 13 during operation in real time, and converts the force data into feedback force that users can perceive. When the remote robot 13 applies or receives an external force, the force sensor will make the user's hand feel the corresponding resistance or thrust, helping the user better control the force and movement of the remote robot 13, and improving the accuracy and stability of the operation. Through the coordinated work of the display screen, the tactile sensor and the force sensor, the user terminal 14 provides users with multi-dimensional feedback information of vision, touch and force, enabling users to more realistically feel the operation environment when remotely operating the remote robot 13 for assembly tasks, and significantly improving the operation experience and control accuracy.
[0037] In an embodiment of the present invention, after the controller 12 determines the motion posture based on the fusion features, the controller 12 is further configured to execute:
[0038] The motion posture is constrained by the angle constraint equations output by the remote robot 13 to obtain the constrained motion posture, and the constrained motion posture is used as the final control data of the remote robot 13.
[0039] In this embodiment, after completing the step of determining the motion posture based on the fusion features, there are subsequent key steps: substituting the determined motion posture into the angle constraint equations output by the remote robot 13 for constraint calculation. This set of equations integrates various factors such as the physical limitations of the robot joints, dynamic characteristics, and task requirements. After being processed by this set of equations, the constrained motion posture can be obtained, and finally, this constrained motion posture is used as the final control data of the remote robot 13. This operation ensures that the motion posture of the remote robot 13 not only conforms to the user's intention but also can be executed safely, stably, and efficiently within the framework of the robot's own hardware conditions and the actual requirements of the task, effectively avoiding problems such as mechanical failures and motion out-of-control caused by unreasonable postures, and greatly improving the reliability and accuracy of the control of the remote robot 13.
[0040] In an embodiment of the present invention, the angle constraint equations output by the remote robot 13 are determined by the calculation logic of the controller 12 executing the following formula:
[0041]
[0042] In the formula, θ opt is the optimized joint angle vector, λ1 is the first weight coefficient, λ2 is the second weight coefficient, λ3 is the third weight coefficient, F(θ) is the forward kinematic function of the robot, P * is the desired position and attitude vector of the end effector, is the dynamic constraint term, C limit (θ) is the joint angle limit constraint term, θ0 is the joint angle vector at the previous moment, θ = [θ1, θ2, …, θ n T is the current joint angle vector, and α1 is the fourth weight coefficient.
[0043] In this embodiment, F(θ) is the forward kinematic function of the robot, P * is the desired position and attitude vector of the end effector, ||F(θ) - P * || 2 is the kinematic constraint to ensure that the end effector reaches the desired position and attitude, is the dynamic constraint term, which is determined by the robot dynamics equation Obtained in combination with restrictions such as the maximum torque of the joint, To ensure that the movement is within the dynamic capabilities, C limit C(θ) is the joint angle limit constraint term, C limit C(θ) = max(0, θ - θ max ) + max(0, θ min - θ), ensuring that the joint angle is within the physically feasible range, θ max is, θ min are respectively the maximum and minimum value vectors of the joint angle, λ1 is the first weight coefficient, λ2 is the second weight coefficient, and λ3 is the third weight coefficient used to measure the importance of each constraint. The objective function J(θ) encourages the difference between the current joint angle and the angle at the previous moment not to be too large, and limits the speed of change of the joint angle to achieve smoother movement.
[0044] In an embodiment of the present invention, when the controller 12 executes the update of the quality dynamic weight based on the quality of the video data and the quality of the audio data, it is specifically used to execute the calculation logic of the following formula:
[0045]
[0046] In the formula, D is the updated quality dynamic weight, Q is the quality dynamic weight before update, Q V is the quality of the video data, Q A is the quality of the audio data.
[0047] In an embodiment of the present invention, after the controller 12 executes the remote assembly of the remote robot 13 according to the motion posture, the controller 12 is further used to execute:
[0048] Based on the current operating state of the remote robot 13, determine the feedback ratios of the visual data, force sense data, and tactile data;
[0049] And send the feedback ratios to the user's client 14;
[0050] Among them, the operating state includes free space motion operation, fine operation, and contact operation.
[0051] In this embodiment, when driving the remote robot 13 to perform a remote assembly task according to the motion posture, the controller 12 will also perform the following steps: Based on the specific operating state of the current remote robot 13, using a predefined algorithm model, accurately calculate the feedback ratios of visual data, force sense data, and tactile data. The operating states are mainly divided into three modes: free space motion operation, fine operation, and contact operation. In the free space motion operation mode, through the previous calibration and optimization strategy of the present invention, the feedback ratio of visual information is set to 0.7, the feedback ratio of force sense data is 0.2, and the feedback ratio of tactile data is 0.1. This ratio allocation aims to highlight the importance of visual guidance in free space while taking into account force sense and tactile information to provide a more comprehensive perception. In the fine operation mode, the feedback ratios of visual, force sense, and tactile data are precisely adjusted to 0.4, 0.4, and 0.2. This balanced allocation helps the user to rely on visual observation of operation details, precise control of force by force sense, and perception of minute contact changes by touch when performing fine operations. In the contact operation state, the ratios of the three are 0.1, 0.6, and 0.3 respectively. This setting emphasizes the dominant role of force sense and tactile data in contact operations because at this time, force sense feedback is crucial for judging the magnitude and direction of the contact force, and tactile data can provide subtle information about the characteristics of the contact surface, while the relative importance of visual information decreases. After determining the above feedback ratios, the present invention will transmit them to the user's client 14 in a timely manner through an efficient communication protocol. Through this mechanism of dynamically adjusting the data feedback ratio according to different operating states, the user can more realistically and subtly perceive various key information in the actual operation environment, thereby significantly improving the control accuracy of the remote robot 13 and ensuring the efficient and accurate completion of the remote assembly task.
[0052] In an embodiment of the present invention, the preset duration executed by the controller 12 is determined according to the integrity of the human skeleton sequence information in the video data.
[0053] In this embodiment, when collecting video data, if situations such as an object blocking the picture or a part of the user's body accidentally entering the video picture occur, it will cause a deviation in the motion posture determined based on the fusion features. In view of this, the present invention needs to extend the preset duration to provide the user with sufficient time to adjust their own posture or the surrounding environment, so as to ensure the efficient and accurate progress of the remote assembly task. Specifically, there is a clear relationship between the integrity of the human skeleton sequence information and the preset duration, that is, the lower the integrity of the human skeleton sequence information, the longer the preset duration extended by the present invention, so as to ensure that the user has enough time to optimize the operation and improve the completion quality of the remote assembly task to the greatest extent.
[0054] In an embodiment of the present invention, the preset duration is determined by the controller 12 executing the calculation logic of the following formula:
[0055]
[0056] Wherein, T is the preset duration, k is the proportionality constant, ε is the denominator adjustment coefficient, α2 is the exponential parameter, β is the weight coefficient of the integrity change rate, is the integrity change rate, γ is the weight coefficient of the historical average integrity, is the average integrity of the human skeleton sequence information within a past period of time, δ is the weight coefficient of the environmental interference factor, and N is the environmental interference index.
[0057] In this embodiment, k is the proportionality constant, which is used to adjust the overall scale of the preset duration and can be determined according to the characteristics and experience of the actual present invention. ε is the denominator adjustment coefficient to avoid the situation where the denominator is zero. α2 is the exponential parameter, which is used to adjust the influence degree of the integrity I on the preset duration T. β is the weight coefficient of the integrity change rate, which reflects the influence degree of the integrity change rate on the preset duration, is the integrity change rate, that is, the change amount of the integrity per unit time. It can reflect the dynamic change situation of the integrity. For example, if the integrity rises rapidly within a short period of time, it indicates that the information quality is improving rapidly, and the preset duration may need to be shortened accordingly. γ is the weight coefficient of the historical average integrity, which reflects the influence of the historical integrity on the current preset duration. is the average integrity of the human skeleton sequence information within a past period of time. Considering the historical integrity can make the determination of the preset duration more stable and avoid large fluctuations in the preset duration caused by accidental fluctuations in the current integrity. N is the environmental interference index, which is used to measure the interference degree of the current environment on the acquisition of the human skeleton sequence information in the audio data and can be comprehensively evaluated by factors such as the noise level and signal strength.
[0058] The specific embodiments of the present invention have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A control system for a remote robot, characterized in that, It includes a data acquisition device (11), a controller (12), a remote robot (13) and a user terminal (14); The controller (12) is communicatively connected to the data acquisition device (11), the remote robot (13) and the user terminal (14) respectively; The data acquisition device (11) is used to acquire video data and audio data of the user within the same preset time period; The controller (12) is used to perform the following operations: receive the video data and the audio data sent by the data acquisition device (11); sequentially perform feature extraction on the video data and the audio data to obtain video features and audio features; Fuse the video features and the audio features according to the quality dynamic weight to obtain fused features; Based on the fused features, determine the motion posture; Control the remote robot (13) to perform remote assembly according to the motion posture; update the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset time period, receive the video data and the audio data sent by the data acquisition device (11) again; The user terminal (14) is used to receive visual data, force sense data, and tactile data of the remote robot (13) during the remote assembly process sent by the controller (12); The remote robot (13) is used to receive the motion posture sent by the controller (12) and perform remote assembly according to the motion posture.
2. The system according to claim 1, characterized in that, The data acquisition device (11) includes a microphone array and a camera; The microphone array is used to acquire the audio data; The camera is used to acquire the video data.
3. The system according to claim 1, wherein The user terminal (14) includes a display screen, a tactile sensor, and a force sense sensor; The display screen is used to display the visual data to the user; The tactile sensor is used to transmit the tactile data to the user; The force sense sensor is used to transmit the force sense data to the user.
4. The system according to claim 1, wherein After the controller (12) performs determining the motion posture based on the fused features, the controller (12) is further used to perform: Constrain the motion posture through the angle constraint equations output by the remote robot (13) to obtain the constrained motion posture, and use the constrained motion posture as the final control data of the remote robot (13).
5. The system according to claim 4, characterized in that, The angle constraint equations output by the remote robot (13) are determined by the calculation logic of the following formula executed by the controller (12): Where, θ opt is the optimized joint angle vector, λ1 is the first weight coefficient, λ2 is the second weight coefficient, λ3 is the third weight coefficient, F(θ) is the forward kinematic function of the robot, P * is the desired position and attitude vector of the end effector, is the dynamic constraint term, C limit (θ) is the joint angle limit constraint term, θ0 is the joint angle vector at the previous moment, θ = [θ1, θ2, …, θ n T is the current joint angle vector, and α1 is the fourth weight coefficient. 6. The system according to claim 1, wherein When the controller (12) updates the quality dynamic weight based on the quality of the video data and the quality of the audio data, it is specifically used to execute the calculation logic of the following formula: Where D is the updated quality dynamic weight, Q is the quality dynamic weight before update, and Q V is the quality of video data, and Q A is the quality of audio data.
7. The system according to claim 1, wherein After the controller (12) performs controlling the remote robot (13) to perform remote assembly according to the motion posture, the controller (12) is further used to perform: Based on the current operation state of the remote robot (13), determine the feedback ratios of the visual data, the force sense data, and the tactile data; And send the feedback ratios to the user terminal (14) of the user; Wherein, the operation state includes free space motion operation, fine operation, and contact operation.
8. The system according to claim 1, wherein The preset duration executed by the controller (12) is determined according to the integrity of the human skeleton sequence information in the video data.
9. The system according to claim 8, wherein The preset duration is determined by the controller (12) executing the calculation logic of the following formula: where T is the preset duration, k is a proportionality constant, ε is a denominator adjustment coefficient, α2 is an exponential parameter, β is a weight coefficient of the integrity change rate, is the integrity change rate, γ is a weight coefficient of the historical average integrity, I is the average integrity of the human skeleton sequence information in the past period of time, δ is a weight coefficient of the environmental interference factor, and N is an environmental interference index.