Remote robot control method, device and equipment and medium

By acquiring user video and audio data and utilizing feature extraction and adaptive weighted fusion strategies, a fusion feature control system is generated to control a remote robot. This solves the problems of insufficient accuracy and interaction in traditional robot control methods for complex assembly tasks, and achieves efficient and accurate remote robot control.

CN120395831APending Publication Date: 2025-08-01SHENZHEN POLYTECHNIC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510561576.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Traditional robot control methods are insufficient for complex assembly tasks and high-precision requirements, especially in terms of multi-sensor data fusion, motion control accuracy, and robot-human interaction. They are difficult to efficiently and accurately map human operations to the robot's end effector.

Method used

By acquiring users' video and audio data, features are extracted using the SIFT and MFCC algorithms. Combined with an adaptive weighted fusion strategy, quality weights are dynamically adjusted to generate fused features to control remote robots. Visual, force, and tactile data are transmitted in real time via a 5G network to provide comprehensive feedback.

Benefits of technology

It improves the control precision and reliability of remote robots, ensures smooth and continuous motion control, and enhances the efficiency and accuracy of remote assembly tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120395831A_ABST
    Figure CN120395831A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a remote robot control method, device and equipment and a medium. The method comprises the following steps: acquiring video data and audio data of a user in the same preset time period; performing feature extraction on the video data and the audio data in sequence to obtain video features and audio features; performing feature fusion on the video features and the audio features according to the quality dynamic weight to obtain fusion features; determining a motion posture based on the fusion feature; controlling the remote robot to perform remote assembly according to the motion posture, and sending visual data, force sense data and touch sense data of the remotely assembled remote robot to a user side of the user; and updating the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset time length, re-executing the step of obtaining the video data and the audio data of the user in the same preset time period. Through the configuration mode, the control precision of the remote robot can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a control method, device, equipment and medium for a remote robot. Background Art

[0002] With the continuous development of automation and remote control technologies, robots have been widely used in various industrial and service scenarios. Although traditional robot control methods can achieve basic operation tasks, there are still some problems when facing complex assembly tasks and high-precision requirements, especially in the fusion of multi-sensor data, the accuracy of motion control, and the interaction between robots and humans. In remote operation tasks, how to efficiently and accurately map human operations to the end effector of the robot and improve the flexibility and accuracy of the robot is still an urgent problem to be solved.

[0003] Based on this, the present invention proposes a control method and device for a remote robot to solve the above technical problems. Summary of the Invention

[0004] The present invention describes a control method, device, equipment and medium for a remote robot, which can effectively improve the control accuracy of the remote robot.

[0005] According to a first aspect, the present invention provides a control method for a remote robot, including:

[0006] Obtaining video data and audio data of a user within the same preset time period;

[0007] Successively performing feature extraction on the video data and the audio data to obtain video features and audio features;

[0008] Performing feature fusion on the video features and the audio features according to a quality dynamic weight to obtain fusion features;

[0009] Determining a motion posture based on the fusion features;

[0010] Controlling a remote robot to perform remote assembly according to the motion posture, and sending visual data, force sense data, and tactile data of the remote robot during remote assembly to the user terminal of the user;

[0011] Updating the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset duration, re-executing the step of "obtaining video data and audio data of a user within the same preset time period".

[0012] According to a second aspect, the present invention provides a control device for a remote robot, including:

[0013] An acquisition unit, configured to acquire video data and audio data of a user within the same preset time period;

[0014] A first data processing unit, configured to sequentially perform feature extraction on the video data and the audio data to obtain video features and audio features;

[0015] A second data processing unit, configured to perform feature fusion on the video features and the audio features according to a quality dynamic weight to obtain fusion features;

[0016] A third data processing unit, configured to determine a motion posture based on the fusion features;

[0017] A fourth data processing unit, configured to control a remote robot to perform remote assembly according to the motion posture, and send visual data, force perception data, and tactile data of the remote robot during remote assembly to the user's client;

[0018] A fifth data processing unit, configured to update the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset duration, re - execute the step of "acquiring video data and audio data of the user within the same preset time period".

[0019] According to a third aspect, the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method of the first aspect is implemented.

[0020] According to a fourth aspect, the present invention provides a computer - readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method of the first aspect.

[0021] According to the control method, device, equipment and medium of the remote robot provided by the present invention, with the help of the sensor array and data acquisition module, the user's video data and audio data are accurately acquired at a high sampling rate within the same preset time period. Then, the video data is subjected to key frame extraction and feature point matching using algorithms such as the scale-invariant feature transform (SIFT), and the audio data is subjected to spectrum analysis and feature vector extraction using the Mel-frequency cepstral coefficient (MFCC) algorithm, thereby generating audio features for accurately representing the user's voice commands and video features for representing the user's action commands. Subsequently, based on an adaptive weighted fusion strategy, the quality weights are dynamically adjusted based on the quality assessment values determined by quantitative indicators such as the clarity and frame rate stability of the video data, and the quality assessment values determined by quantitative indicators such as the signal-to-noise ratio and harmonic distortion rate of the audio data, and the video features and audio features are deeply fused to obtain a fusion feature that can comprehensively and accurately represent the user's true intentions. Based on this fusion feature, a matching motion posture is determined, and this posture is used to accurately control the remote robot to carry out remote assembly operations. In the remote assembly process, the visual data collected by the visual sensors (such as industrial cameras) carried by the remote robot, the force data obtained by the six-dimensional force sensor, and the tactile data perceived by the tactile sensor array will be transmitted to the user's client in real time and synchronously through a high-speed, low-latency wireless communication link (such as a 5G network). The visualization and feedback processing software equipped on the user end parses and presents these data to provide users with comprehensive and intuitive feedback information. At the same time, the quality of the video data and audio data is continuously dynamically evaluated, and based on these quality data, the quality dynamic weight is intelligently updated to ensure that the fusion feature always has high accuracy, thereby greatly improving the control accuracy of the remote robot. Whenever the preset time period ends, the interruption mechanism will be automatically triggered, and the "obtaining the user's video data and audio data within the same preset time period" will be executed again to ensure that the motion control of the remote robot is smooth and uninterrupted. Through the above configuration method, the present invention can effectively improve the control accuracy of the remote robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 A schematic flow chart of a remote robot control method according to one embodiment is shown;

[0024] Figure 2A schematic block diagram of a control device for a remote robot according to an embodiment is shown. Detailed implementation

[0025] The solution provided by the present invention will be described below with reference to the accompanying drawings.

[0026] Figure 1 A schematic flowchart of a control method for a remote robot according to an embodiment is shown. It can be understood that this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities. As Figure 1 shown, the method includes:

[0027] Step 100: Obtain video data and audio data of the user within the same preset time period;

[0028] Step 102: Extract features from the video data and audio data in sequence to obtain video features and audio features;

[0029] Step 104: Perform feature fusion on the video features and audio features according to the quality dynamic weight to obtain fusion features;

[0030] Step 106: Determine the motion posture based on the fusion features;

[0031] Step 108: Control the remote robot to perform remote assembly according to the motion posture, and send the visual data, force sense data, and tactile data of the remote robot during remote assembly to the user's client;

[0032] Step 110: Update the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset duration, re - execute the step of "obtaining video data and audio data of the user within the same preset time period".

[0033] In this embodiment, with the aid of a sensor array and a data acquisition module, within the same preset time period (this preset time period will change dynamically as the cycle progresses), video data and audio data of the user are accurately acquired at a high sampling rate. Subsequently, algorithms such as Scale-Invariant Feature Transform (SIFT) are used to extract key frames and match feature points from the video data, and the Mel Frequency Cepstral Coefficient (MFCC) algorithm is combined to perform spectral analysis and extract feature vectors from the audio data, thereby generating audio features that accurately represent the user's voice commands and video features that represent the user's motion commands respectively. Then, based on an adaptive weighted fusion strategy, according to the quality evaluation values determined by quantization metrics such as the clarity of the video data and the frame rate stability, and the quality evaluation values determined by quantization metrics such as the signal-to-noise ratio and harmonic distortion rate of the audio data, the quality dynamic weights are dynamically adjusted, and the video features and audio features are deeply fused to obtain fusion features that can comprehensively and accurately represent the user's true intention. Based on this fusion feature, a matching motion posture is determined, and with this posture, the remote robot is precisely controlled to carry out remote assembly operations. During the remote assembly process, the visual data collected by the visual sensor (such as an industrial camera) carried by the remote robot, the force perception data obtained by the six-axis force sensor, and the tactile data sensed by the tactile sensor array will be transmitted to the user's client in real-time and synchronously through a high-speed and low-latency wireless communication link (such as a 5G network). The visualization and feedback processing software equipped on the client parses and presents these data, providing comprehensive and intuitive feedback information to the user. At the same time, the quality of the video data and audio data is continuously evaluated dynamically, and based on these quality data, the quality dynamic weights are intelligently updated to ensure that the fusion features always have high accuracy, thereby greatly improving the control accuracy of the remote robot. Whenever the preset duration ends, an interruption mechanism is automatically triggered, and "acquire the video data and audio data of the user within the same preset time period" is executed again to ensure that the motion control of the remote robot continues smoothly and without interruption. Through the above configuration method, the present invention can effectively improve the control accuracy of the remote robot.

[0034] In an embodiment of the present invention, after determining the motion posture based on the fusion feature, it further includes:

[0035] The motion posture is constrained by the output angle constraint equations of the remote robot to obtain the constrained motion posture, and the constrained motion posture is used as the final control data of the remote robot.

[0036] In this embodiment, after completing the step of determining the motion posture based on the fusion features, there are subsequent key steps: substituting the determined motion posture into the remote robot output angle constraint equations for constraint operations. This system of equations integrates various factors such as the physical limitations of the robot joints, dynamic characteristics, and task requirements. After being processed by this system of equations, the constrained motion posture can be obtained, and finally, this constrained motion posture is used as the final control data for the remote robot. This operation ensures that the motion posture of the remote robot not only conforms to the user's intention but also can be executed safely, stably, and efficiently within the framework of the robot's own hardware conditions and the actual requirements of the task, effectively avoiding problems such as mechanical failures and motion out-of-control caused by unreasonable postures, and greatly improving the reliability and accuracy of remote robot control.

[0037] In an embodiment of the present invention, the remote robot output angle constraint equations are determined according to the following formula:

[0038]

[0039]

[0040] In the formula, θ opt is the optimized joint angle vector, λ1 is the first weight coefficient, λ2 is the second weight coefficient, λ3 is the third weight coefficient, F(θ) is the forward kinematic function of the robot, P * is the desired position and attitude vector of the end effector, is the dynamic constraint term, C limit (θ) is the joint angle limit constraint term, θ0 is the joint angle vector at the previous moment, θ = [θ1, θ2, …, θ n T is the current joint angle vector, and α1 is the fourth weight coefficient.

[0041] In this embodiment, F(θ) is the forward kinematic function of the robot, P * is the desired position and attitude vector of the end effector, ||F(θ) - P * || 2 is the kinematic constraint to ensure that the end effector reaches the desired position and attitude, is the dynamic constraint term, which is obtained by combining the robot dynamics equation with limitations such as the maximum joint torque, used to ensure that the motion is within the dynamic capacity range, C limit (θ) is the joint angle limit constraint term, C limit (θ) = max(0, θ - θ max ) + max(0, θ min - θ), to ensure that the joint angles are within the physically feasible range, θ​max is θ min are the maximum and minimum vectors of joint angles respectively, λ1 is the first weight coefficient, λ2 is the second weight coefficient, and λ3 is the third weight coefficient used to measure the importance of each constraint. The objective function J(θ) encourages the difference between the current joint angle and the angle at the previous moment not to be too large, and limits the speed of change of the joint angle to achieve smoother motion.

[0042] In an embodiment of the present invention, the quality dynamic weight is updated by the following formula:

[0043]

[0044] In the formula, D is the updated quality dynamic weight, Q is the quality dynamic weight before update, Q V is the quality of video data, Q A is the quality of audio data.

[0045] In an embodiment of the present invention, after controlling the remote robot for remote assembly according to the motion posture and sending the visual data, force sense data, and tactile data of the remote robot during remote assembly to the user's client, it further includes;

[0046] Based on the current operating state of the remote robot, determine the feedback ratio of the visual data, force sense data, and tactile data;

[0047] And send the feedback ratio to the user's client;

[0048] Among them, the operating state includes free space motion operation, fine operation, and contact operation.

[0049] In this embodiment, when driving a remote robot to perform a remote assembly task according to the motion posture and transmitting the visual data, force perception data, and tactile data of the remote robot during the assembly process to the user's client, the present invention further performs the following key steps: Based on the specific operating state of the current remote robot, using a predefined algorithm model, accurately calculate the feedback ratios of the visual data, force perception data, and tactile data. The operating states are mainly divided into three modes: free-space motion operation, fine operation, and contact operation. In the free-space motion operation mode, through the previous calibration and optimization strategy of the present invention, the feedback ratio of visual information is set to 0.7, the feedback ratio of force perception data is 0.2, and the feedback ratio of tactile data is 0.1. This ratio distribution aims to highlight the importance of visual guidance in free space while taking into account force perception and tactile information to provide a more comprehensive perception. In the fine operation mode, the feedback ratios of visual, force perception, and tactile data are precisely adjusted to 0.4, 0.4, and 0.2. This balanced distribution helps the user to rely on visual observation of operation details, precise control of force by force perception, and perception of minute contact changes by touch when performing fine operations. In the contact operation state, the ratios of the three are 0.1, 0.6, and 0.3 respectively. This setting emphasizes the leading role of force perception and tactile data in contact operations because at this time, force perception feedback is crucial for judging the magnitude and direction of the contact force, and tactile data can provide subtle information about the characteristics of the contact surface, while the relative importance of visual information decreases. After the present invention determines the above feedback ratios, it will transmit them to the user's client in a timely manner through an efficient communication protocol. Through this mechanism of dynamically adjusting the data feedback ratio according to different operating states, the user can more realistically and delicately perceive various key information in the actual operation environment, thereby significantly improving the control accuracy of the remote robot and ensuring the efficient and precise completion of the remote assembly task.

[0050] In an embodiment of the present invention, the preset duration is determined according to the integrity of the human skeleton sequence information in the video data.

[0051] In this embodiment, when collecting video data, if situations such as object occlusion of the picture or accidental intrusion of some areas of the user's body into the video picture occur, it will cause deviations in the motion posture determined based on the fusion features. In view of this, the present invention needs to extend the preset duration to provide the user with sufficient time to adjust their own posture or the surrounding environment, so as to ensure the efficient and precise progress of the remote assembly task. Specifically, there is a clear relationship between the integrity of the human skeleton sequence information and the preset duration, that is, the lower the integrity of the human skeleton sequence information, the longer the preset duration extended by the present invention, so as to ensure that the user has enough time to optimize the operation and improve the completion quality of the remote assembly task to the greatest extent.

[0052] In an embodiment of the present invention, the preset duration is determined by the following formula:

[0053]

[0054] In the formula, T is the preset duration, k is the proportionality constant, ε is the denominator adjustment coefficient, α2 is the exponential parameter, β is the weight coefficient of the integrity change rate, is the integrity change rate, γ is the weight coefficient of the historical average integrity, is the average integrity of the human skeleton sequence information within a past period of time, δ is the weight coefficient of the environmental interference factor, and N is the environmental interference index.

[0055] In this embodiment, k is the proportionality constant, which is used to adjust the overall scale of the preset duration and can be determined according to the characteristics and experience of the actual present invention. ε is the denominator adjustment coefficient to avoid the situation of a zero denominator. α2 is the exponential parameter, which is used to adjust the influence degree of the integrity I on the preset duration T. β is the weight coefficient of the integrity change rate, which reflects the influence degree of the integrity change rate on the preset duration, is the integrity change rate, that is, the change amount of the integrity per unit time. It can reflect the dynamic change situation of the integrity. For example, if the integrity rises rapidly within a short period of time, it indicates that the information quality is improving rapidly, and the preset duration may need to be shortened accordingly. γ is the weight coefficient of the historical average integrity, which reflects the influence of the historical integrity on the current preset duration. is the average integrity of the human skeleton sequence information within a past period of time. Considering the historical integrity can make the determination of the preset duration more stable and avoid large fluctuations in the preset duration due to accidental fluctuations in the current integrity. N is the environmental interference index, which is used to measure the interference degree of the current environment on the acquisition of the human skeleton sequence information in the audio data and can be comprehensively evaluated through factors such as the noise level and signal strength.

[0056] The above describes specific embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0057] According to an embodiment of another aspect, the present invention provides a control device for a remote robot. Figure 2 The schematic block diagram of the control device for the remote robot according to an embodiment is shown. It can be understood that the device can be implemented by any device, equipment, platform, and device cluster with computing and processing capabilities. Such as Figure 2As shown in the figure, the device includes: an acquisition unit 200, a first data processing unit 202, a second data processing unit 204, a third data processing unit 206, a fourth data processing unit 208, and a fifth data processing unit 210. The main functions of each component unit are as follows:

[0058] The acquisition unit is configured to acquire video data and audio data of the user within the same preset time period;

[0059] The first data processing unit is configured to sequentially perform feature extraction on the video data and the audio data to obtain video features and audio features;

[0060] The second data processing unit is configured to perform feature fusion on the video features and the audio features according to the quality dynamic weight to obtain fusion features;

[0061] The third data processing unit is configured to determine a motion posture based on the fusion features;

[0062] The fourth data processing unit is configured to control a remote robot to perform remote assembly according to the motion posture, and send visual data, force sense data, and tactile data of the remote robot during remote assembly to the user's client;

[0063] The fifth data processing unit is configured to update the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset duration, re-execute the step of "acquiring video data and audio data of the user within the same preset time period".

[0064] As a preferred implementation manner, after determining the motion posture based on the fusion features, it further includes:

[0065] Constraining the motion posture through an angle constraint equation set output by the remote robot to obtain a constrained motion posture, and using the constrained motion posture as the final control data of the remote robot.

[0066] As a preferred implementation manner, the angle constraint equation set output by the remote robot is determined according to the following formula:

[0067]

[0068] In the formula, θ opt is the optimized joint angle vector, λ1 is the first weight coefficient, λ2 is the second weight coefficient, λ3 is the third weight coefficient, F(θ) is the forward kinematics function of the robot, P * is the expected position and attitude vector of the end effector, is the dynamic constraint term, C limit(θ) is the joint angle limit constraint term, θ0 is the joint angle vector at the previous moment, θ = [θ1, θ2, …, θ n T is the current joint angle vector, and α1 is the fourth weight coefficient.

[0069] As a preferred implementation manner, the quality dynamic weight is updated by the following formula:

[0070]

[0071] In the formula, D is the updated quality dynamic weight, Q is the quality dynamic weight before update, Q V is the quality of video data, and Q A is the quality of audio data.

[0072] As a preferred implementation manner, after controlling the remote robot to perform remote assembly according to the motion posture and sending the visual data, force sense data, and tactile data of the remote robot during remote assembly to the user's client, it further includes;

[0073] Based on the current operation state of the remote robot, determine the feedback ratios of the visual data, force sense data, and tactile data;

[0074] And send the feedback ratios to the user's client;

[0075] Wherein, the operation state includes free space motion operation, fine operation, and contact operation.

[0076] As a preferred implementation manner, the preset duration is determined according to the integrity of the human skeleton sequence information in the video data.

[0077] As a preferred implementation manner, the preset duration is determined by the following formula:

[0078]

[0079] In the formula, T is the preset duration, k is the proportionality constant, ε is the denominator adjustment coefficient, α2 is the exponential parameter, β is the weight coefficient of the integrity change rate, is the integrity change rate, γ is the weight coefficient of the historical average integrity, is the average integrity of the human skeleton sequence information in the past period of time, δ is the weight coefficient of the environmental interference factor, and N is the environmental interference index.

[0080] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, it causes the computer to execute the method described in combination with Figure 1 what is described.

[0081] According to an embodiment of still another aspect, an electronic device is further provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method combined with Figure 1 is implemented.

[0082] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment.

[0083] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0084] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention shall be included in the protection scope of the present invention.

Claims

1. A control method for a remote robot, characterized in that, Including: Obtain the video data and audio data of the user within the same preset time period; Successively perform feature extraction on the video data and the audio data to obtain video features and audio features; Fuse the video features and the audio features according to the quality dynamic weight to obtain fused features; Determine the motion posture based on the fused features; Control the remote robot to perform remote assembly according to the motion posture, and send the visual data, force sense data, and tactile data of the remote robot during remote assembly to the user's client; Update the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset duration, re-execute the step of "obtaining the video data and audio data of the user within the same preset time period".

2. The method according to claim 1, characterized in that After determining the motion posture based on the fused features, it further includes: Constrain the motion posture through the angle constraint equation set output by the remote robot to obtain the constrained motion posture, and use the constrained motion posture as the final control data of the remote robot.

3. The method according to claim 2, wherein The angle constraint equation set output by the remote robot is determined according to the following formula: where θ opt is the optimized joint angle vector, λ1 is the first weight coefficient, λ2 is the second weight coefficient, λ3 is the third weight coefficient, F(θ) is the forward kinematic function of the robot, P * is the desired position and attitude vector of the end effector, is the dynamic constraint term, C limit (θ) is the joint angle limit constraint term, θ0 is the joint angle vector at the previous moment, θ = [θ1, θ2, …, θ n T is the current joint angle vector, and α1 is the fourth weight coefficient.​ 4. The method according to claim 1, wherein The quality dynamic weight is updated through the following formula: Where D is the updated quality dynamic weight, Q is the quality dynamic weight before update, and Q V is the quality of video data, and Q A is the quality of audio data.

5. The method according to claim 1, characterized in that, After controlling the remote robot to perform remote assembly according to the motion posture and sending the visual data, force sense data, and tactile data of the remote robot during remote assembly to the user's client, it further includes; Determine the feedback ratio of the visual data, force sense data, and tactile data based on the current operation state of the remote robot; And send the feedback ratio to the user's client; Wherein, the operation state includes free space motion operation, fine operation, and contact operation.

6. The method according to claim 1, wherein The preset duration is determined according to the integrity of the human skeleton sequence information in the video data.

7. The method according to claim 5, wherein The preset duration is determined according to the following formula: Wherein, T is the preset duration, k is a proportionality constant, ε is a denominator adjustment coefficient, α2 is an exponential parameter, β is a weight coefficient of the integrity change rate, is the integrity change rate, γ is a weight coefficient of the historical average integrity, is the average integrity of the human body skeleton sequence information in the past period of time, δ is a weight coefficient of environmental interference factors, and N is an environmental interference index.

8. A control device for a remote robot, characterized in that, Including: An acquisition unit configured to obtain the video data and audio data of the user within the same preset time period; A first data processing unit configured to successively perform feature extraction on the video data and the audio data to obtain video features and audio features; A second data processing unit configured to fuse the video features and the audio features according to the quality dynamic weight to obtain fused features; A third data processing unit configured to determine the motion posture based on the fused features; A fourth data processing unit configured to control the remote robot to perform remote assembly according to the motion posture, and send the visual data, force sense data, and tactile data of the remote robot during remote assembly to the user's client; A fifth data processing unit configured to update the quality dynamic weight based on the quality of the video data and the quality of the audio data, and after a preset duration, re-execute the step of "obtaining the video data and audio data of the user within the same preset time period".

9. An electronic device, characterized in that, Including a memory and a processor, where a computer program is stored in the memory, and when the processor executes the computer program, the method described in any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which, when executed on a computer, causes the computer to execute the method according to any one of claims 1 to 7.