A body intelligent AI dolly robot and a control method thereof
Patent Information
- Application Number
- CN202610953726.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-09-22
AI Technical Summary
专业运镜手法如推、拉、摇、移、跟、环绕是提升视频观感与艺术效果的关键,但通常依赖专业摄影师、辅助器材与长期经验积累,普通用户难以快速实现
1、采用轮式移动底盘与多自由度双臂机器人结合的具身智能结构,可模拟专业摄影师的复杂运镜手法,实现推、拉、摇、移、环绕、俯仰专业镜头效果,显著提升视频拍摄质量。双臂协同控制相比单臂机器人具有更大的运动空间和更灵活的姿态调整能力,能够完成更加复杂的复合运镜动作;
Smart Images

Figure CN122802777A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of camera movement robot technology, specifically to an embodied intelligent AI camera movement robot and its control method. Background Technology
[0002] With the rapid development of short videos, self-media, and the cultural tourism promotion industry, users' demand for high-quality, professional video content continues to rise. Professional camera techniques such as push, pull, pan, tilt, tracking, and circling are key to enhancing the viewing experience and artistic effect of videos, but they usually rely on professional photographers, auxiliary equipment, and long-term experience, which are difficult for ordinary users to achieve quickly.
[0003] Existing shooting aids suffer from numerous technical shortcomings: handheld gimbals can only achieve basic image stabilization and simple angle adjustments, unable to autonomously plan complex camera movement paths, and are prone to fatigue when held for extended periods, making it difficult to complete long-distance, high-difficulty camera movements; track-based or fixed shooting robots have limited movement patterns, are restricted by track length and installation location, have limited coverage, lack environmental perception and dynamic adjustment capabilities, and cannot adapt to complex and ever-changing outdoor scenes; some shooting devices with AI functions rely heavily on monocular vision recognition, which can only achieve basic target tracking, lacking embodied intelligence in environmental understanding, autonomous decision-making, and body coordination control capabilities, and cannot generate personalized and professional camera movement solutions based on scene, lighting, weather, and user style requirements.
[0004] Meanwhile, existing devices generally lack a lightweight interactive mode that integrates payment, function activation, style selection, path preview, and final video transmission. Their deployment and use in cultural and tourism scenic spots, street shooting, and short drama creation scenarios are costly and complex to operate. For example, professional film and television shooting robots are expensive and usually require professional operation, making them unsuitable for ordinary users; while consumer-grade shooting equipment has limited functionality and cannot meet users' needs for professional camera movement effects.
[0005] Currently, some technological attempts involving intelligent shooting robots have emerged in the industry. For example, existing single-arm shooting robots can achieve basic rotation and pitch movements, but their limited space prevents them from performing complex composite camera movements. Some mobile shooting robots can only perform simple follow-up shooting and lack professional camera movement planning capabilities. Furthermore, some AI shooting software can only process videos in post-production and cannot achieve professional camera movement control during the shooting stage. In summary, existing technologies cannot meet the needs of mass-market, low-cost, and high-quality professional video shooting. Therefore, we propose an embodied intelligent AI camera movement robot and its control method. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a body-worn intelligent AI camera-moving robot and its control method, thus solving the problems mentioned below.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A body-embody intelligent AI camera-moving robot, comprising: The wheeled mobile chassis is equipped with a built-in lidar, depth camera, inertial measurement unit and drive mechanism, and is used to achieve autonomous movement, obstacle avoidance, positioning and attitude stabilization in real-world scenarios. The dual-arm robot system has its main body fixed above the wheeled mobile chassis and includes a left robotic arm and a right robotic arm. Each robotic arm has no less than 3 degrees of freedom. The ends of the two arms are equipped with adaptive flexible cooperative grippers for adjustable gripping and fixing of smartphones. The multimodal perception module includes an ambient light sensor, a temperature and humidity sensor, a gesture recognition camera, and a microphone array, which are used to collect light intensity, environmental meteorological information, spatial scene information, and user gestures and voice interaction commands in real time. The core control and AI processing module has a built-in embodied intelligence model that integrates computer vision, natural language understanding and motion planning algorithms. It is electrically connected to the wheeled mobile chassis, the dual-arm robot system and the multimodal perception module, respectively. It is used to infer and generate camera movement control paths based on user input information and environmental perception data, and to coordinate the control of the chassis and dual arms to perform camera movement actions. The user interaction and payment module includes a display unit and a QR code interaction unit, which are used to display QR codes to enable payment, function activation, camera style selection, path preview and wireless transmission of video files. The voice interaction unit includes a microphone and a speaker, used to recognize user voice control commands and play prompts at the start and end of shooting. Furthermore, in the present invention, each robotic arm in the dual-arm robot system has a 6-degree-of-freedom structure and can work in conjunction with the wheeled mobile chassis to complete at least one composite camera movement action among pitch, roll, circling, telescopic, push, pull, pan, and shift.
[0008] Furthermore, the core control and AI processing module of the present invention is configured to: receive scene text descriptions and camera movement style descriptions input by the user through a QR code interactive interface; integrate real-time lighting, weather, and spatial layout information collected by a multimodal perception module; generate multiple differentiated camera movement control paths through an embodied intelligent large model, each path including the driving trajectory, driving speed, double arm joint movement sequence, and focus and angle control parameters of the mobile phone camera; and push the multiple camera movement paths to the user terminal or device display unit in preview form for selection.
[0009] Furthermore, the embodied intelligent large model in the core control and AI processing module is trained using a large amount of film and television camera movement clips, professional shooting parameters, scene tags, style tags and lighting conditions data, and is constructed using supervised learning and reinforcement learning methods.
[0010] Furthermore, the camera movement path generated by the embodied intelligent large model of the present invention simultaneously includes chassis kinematic constraints and dual-arm workspace constraints, and the execution deviation is compensated at the millisecond level through the real-time feedback closed loop of chassis IMU and dual-arm joint encoder.
[0011] Furthermore, the gesture recognition camera and voice interaction unit are configured to: after the user selects a camera movement path, recognize a preset gesture command or voice command, and trigger the robot to perform automated camera movement and shooting according to the selected path.
[0012] Furthermore, the user interaction and payment module is also configured as follows: When the user selects "Modify description and reshoot", the current user session state is retained, and the core control and AI processing module is triggered to regenerate the camera movement path based on the modified description. The reshoot does not trigger a new payment request.
[0013] Furthermore, the present invention also includes a solar-assisted power supply unit, which is installed above the wheeled mobile chassis and adopts a flexible solar panel design to supplement power to the robot's various modules in outdoor scenarios.
[0014] A control method for an embodied intelligent AI camera-moving robot includes the following steps: S1. The user scans the QR code displayed on the device to enter the interactive service interface, completes online payment, and starts the shooting function; S2. Users can input a description of the shooting scene and their camera movement style requirements through the interactive interface or by voice input. S3. The robot automatically collects information on current ambient lighting, weather, spatial dimensions, and obstacle distribution through a multimodal perception module; S4. The core control and AI processing module combines user descriptions and environmental information to generate at least two alternative camera movement paths through an embodied intelligent big model, and pushes them to the user for preview and selection. S5. The user confirms the selected target camera movement path through voice or gesture commands; S6. The voice interaction unit has a built-in wake word detection model. When it recognizes any word in the preset wake word set issued by the user, it generates a trigger signal. The core control and AI processing module responds to the trigger signal, starts the camera movement path execution program, synchronously controls the wheeled mobile chassis and the dual-arm robot to perform camera movement actions, and automatically controls the mobile phone to record video. S7. During filming, the robot dynamically adjusts the camera movement path in real time according to changes in the environment to ensure safe movement and stable footage; S8. After the camera movement path is completed, the robot automatically stops recording and plays an end prompt sound, then transfers the video file to the user's mobile phone; S9. Users can choose to save the video or modify the description and replan the shooting based on the final result.
[0015] Furthermore, in step S7 of the present invention, when the road surface is detected to be tilted or uneven, the robot adjusts the target state of the robotic arm in real time according to the acquired pose information to ensure the stability of the camera movement center during the shooting process.
[0016] This invention provides a holographic intelligent AI camera-moving robot and its control method. Compared with the prior art, it has the following advantages: 1. Employing a embodied intelligent structure combining a wheeled mobile chassis with a multi-degree-of-freedom dual-arm robot, it can simulate the complex camera movements of professional photographers, achieving professional shot effects such as push, pull, pan, tilt, circle, and pitch, significantly improving video shooting quality. Compared to single-arm robots, dual-arm collaborative control offers greater motion space and more flexible posture adjustment capabilities, enabling the completion of more complex compound camera movements; 2. Based on a well-trained embodied intelligent big model, multiple camera movement schemes are automatically generated by combining multi-dimensional information such as scene, style, lighting, and weather, achieving a high degree of personalization and intelligence. The camera movement path generated by the AI model takes into account the physical constraints of the chassis and arms, and can compensate for execution deviations in real time to ensure smooth camera movement. 3. The entire process—payment, startup, style selection, preview, and image transfer—is completed via QR code, coupled with voice and gesture controls, making it simple and convenient. Ordinary users can quickly complete the shooting without professional skills. No dedicated app download is required, lowering the barrier to entry for users. 4. It features a closed-loop mechanism for recording completion notification, saving, and reshooting, supports user session state retention, and allows users to modify descriptions and reshoot without repeated payments. It is highly practical and offers a superior user experience. Users can adjust the camera movement style multiple times based on the final video quality until they achieve a satisfactory result. 5. It has the ability to adapt to all environments and can dynamically adjust the camera movement path according to real-time environmental perception information to achieve obstacle avoidance and image stabilization optimization. When the road surface is detected to be tilted or uneven, the robot can fine-tune the posture of the robotic arm in real time to ensure the stability of the camera movement center. 6. It can be widely deployed in cultural and tourism scenic spots, commercial streets, self-media creation, and short drama shooting scenes. It can not only provide tourists with a convenient check-in experience, but also meet the needs of content creators for efficient and low-cost shooting. It has broad market application prospects. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall structure of the embodied intelligent AI camera-moving robot of the present invention; Figure 2This is a block diagram showing the module connection of the embodied intelligent AI camera movement robot system of the present invention; Figure 3 This is a schematic diagram of the control method for the embodied intelligent AI camera robot of the present invention.
[0018] In the diagram: 1. Wheeled mobile chassis; 2. Dual-arm robot system; 3. Multi-mode perception module; 4. Core control and AI processing module; 5. User interaction and payment module; 6. Voice interaction unit. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1: like Figure 1 , Figure 2 As shown, this embodiment provides a body-worn intelligent AI camera robot, including a wheeled mobile chassis 1, a LiDAR installed at the front of the chassis, and an integrated main control unit and power supply module. A dual-arm robot system 2 is fixed on the top of the chassis, each robotic arm has a 6-DOF structure, and a flexible end gripper holds a smartphone together. The robot is equipped with a panoramic camera multi-mode perception module 3 and a microphone array as a voice interaction unit 6. A touch screen and a QR code display area are set on the front as a user interaction and payment module 5. The robot's operating time can be increased by adding a solar panel.
[0021] The wheeled mobile chassis 1 adopts a four-wheel drive omnidirectional drive structure, adaptable to various road surfaces such as scenic stone slab roads, cement roads, and indoor floors. The chassis integrates a LiDAR, depth camera, inertial measurement unit (IMU), and odometer, working in conjunction with a SLAM navigation module to complete environmental map construction, real-time positioning, and dynamic obstacle avoidance. The chassis can achieve omnidirectional movement including forward, backward, lateral translation, and rotation, with a maximum speed of 1.5 m / s and a positioning accuracy of ±5 cm.
[0022] The dual-arm robot system 2 features a symmetrical 6-DOF robotic arm design. Each joint integrates a high-precision servo motor, reducer, and encoder, achieving a repeatability accuracy of ±0.1mm. The end effectors of both arms utilize adaptive flexible grippers capable of holding smartphones ranging from 4.7 inches to 6.7 inches, with automatically adjustable gripping force to prevent damage. The coordinated movement of the two arms enables a variety of professional camera movements, including 360° surround shooting, high and low angle tilting shooting, forward and backward push-pull shooting, and left and right panning shooting.
[0023] The multimodal sensing module integrates an ambient light sensor, a temperature and humidity sensor, a gesture recognition camera, and a microphone array. The ambient light sensor has a detection range of 0-100,000 lux, accurately identifying different lighting conditions such as sunny days, cloudy days, and dusk. The gesture recognition camera uses a 3D depth camera and can recognize various preset gestures such as waving, making a heart shape with your hands, and OK. The microphone array uses a 4-microphone array design, supporting voice recognition within a 3-meter range with an accuracy rate of over 95%.
[0024] The core control and AI processing module 4 utilizes high-performance edge computing devices and incorporates a large-scale embodied intelligence model. The model training dataset contains over 100,000 professional film and television camera movement clips, covering various styles including traditional Chinese, modern, suspenseful, and aesthetically pleasing works, as well as indoor, outdoor, daytime, and nighttime scenes. Based on user-input natural language descriptions, the model can generate three differentiated camera movement paths within one second and optimize and adjust them in real time.
[0025] The user interaction and payment module 5 features a 10.1-inch high-definition touchscreen display with a resolution of 1920×1200. The QR code interaction unit supports mainstream payment methods such as WeChat and Alipay, automatically redirecting to the interactive interface upon payment completion. Video file transfer supports both Wi-Fi direct connection and cloud short link, with transfer speeds up to 10MB / s.
[0026] The solar-assisted power supply unit uses high-efficiency flexible solar panels with a conversion efficiency of up to 22%. Under standard lighting conditions, the charging power can reach 50W, which can effectively extend the robot's continuous working time by 2-3 hours.
[0027] Example 2: When using this robot in cultural and tourism scenic spots, users can scan the QR code on the display screen with their mobile phones to enter the mini-program and complete a single payment. The user inputs the shooting scene as "ancient garden style" and the style as "beautiful and slow surround". The robot detects that the current environment is one with sufficient natural light and diffused light through environmental sensors, and generates three camera movement paths through the AI model: Path A is a low-angle slow surround, with the chassis moving along an arc with a radius of 2 meters at a speed of 0.3m / s, and the arms keeping the camera at a height of 1.2 meters, always pointing at the user; Path B involves advancing forward from the front and then moving laterally to follow up. The chassis first advances forward 2 meters at a speed of 0.2 m / s, and then moves laterally to the left 3 meters at a speed of 0.4 m / s. Path C involves taking a high-angle shot and then gradually pulling away, raising the camera to a height of 2.5 meters with both arms for the overhead shot, and then slowly moving the chassis back while lowering the camera height with both arms.
[0028] The user selects path A and gives the voice command "Start Shooting." The robot moves autonomously along the path and controls the phone to record. During shooting, if slight undulations are detected on the road surface, the robot adjusts its robotic arm posture in real time to compensate for chassis vibrations and ensure stable footage. Upon completion, a notification sound is emitted, and the video is transmitted directly to the user's phone via Wi-Fi. The service is completed after the user confirms and saves the video. If the user is not satisfied with the final result, they can select "Edit Description" and add "Closer Shot to the Person." The system retains the current session state, regenerates the camera path, and shoots again without triggering a new payment request.
[0029] Example 3: Self-media users can use this robot to shoot short chase scenes. After scanning the QR code to start the robot, they can use voice input to select a dimly lit indoor environment, a tense chase style, and rapid panning and sudden stops. The robot detects that the ambient light is weak and there are many obstacles indoors. The AI model generates two high-dynamic camera movement schemes: Path 1 combines S-shaped movement of the chassis with rapid panning of both arms to simulate a chase perspective; Path 2 combines low-angle tracking shots with sudden stop close-ups to highlight the tense atmosphere.
[0030] After the user confirms path 1 with a gesture, the robot starts recording. It then moves rapidly along the preset path while simultaneously controlling the phone to record in high frame rate mode. During recording, if an obstacle is detected, the robot automatically adjusts its camera movement to avoid it while maintaining a consistent camera rhythm. Once recording is complete, the robot automatically transmits the video to the user's phone, allowing the user to choose to save the video or edit the description and re-record.
[0031] The above embodiments are only used to illustrate the core technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or some of the technical features can be replaced. Such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0032] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0033] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their likenesses.
Claims
1. A body-embody intelligent AI camera-moving robot, characterized in that, include: The wheeled mobile chassis is equipped with a built-in lidar, depth camera, inertial measurement unit and drive mechanism, and is used to achieve autonomous movement, obstacle avoidance, positioning and attitude stabilization in real-world scenarios. The dual-arm robot system has its main body fixed above the wheeled mobile chassis and includes a left robotic arm and a right robotic arm. Each robotic arm has no less than 3 degrees of freedom. The ends of the two arms are equipped with adaptive flexible cooperative grippers for adjustable gripping and fixing of smartphones. The multimodal perception module includes an ambient light sensor, a temperature and humidity sensor, a gesture recognition camera, and a microphone array, which are used to collect light intensity, environmental meteorological information, spatial scene information, and user gestures and voice interaction commands in real time. The core control and AI processing module has a built-in embodied intelligence model that integrates computer vision, natural language understanding and motion planning algorithms. It is electrically connected to the wheeled mobile chassis, the dual-arm robot system and the multimodal perception module, respectively. It is used to infer and generate camera movement control paths based on user input information and environmental perception data, and to coordinate the control of the chassis and dual arms to perform camera movement actions. The user interaction and payment module includes a display unit and a QR code interaction unit, which are used to display QR codes to enable payment, function activation, camera style selection, path preview and wireless transmission of video files. The voice interaction unit includes a microphone and a speaker, used to recognize user voice control commands and play prompts at the start and end of shooting.
2. The embodied intelligent AI camera-moving robot and its control method according to claim 1, characterized in that: Each robotic arm in the dual-arm robot system has a 6-degree-of-freedom structure and can work in conjunction with a wheeled mobile chassis to complete at least one composite camera movement action among pitch, roll, orbit, telescopic, push, pull, pan, and tilt.
3. The embodied intelligent AI camera-moving robot according to claim 2, characterized in that: The core control and AI processing module is configured to: receive scene text descriptions and camera movement style descriptions input by the user through the QR code interactive interface; integrate real-time lighting, weather and spatial layout information collected by the multimodal perception module; and generate multiple differentiated camera movement control paths through the embodied intelligent big model. Each path includes the driving trajectory, driving speed, double arm joint movement sequence of the wheeled mobile chassis, and the focus and angle control parameters of the mobile phone camera. Multiple camera movement paths are pushed to the user terminal or device display unit in preview form for selection.
4. The embodied intelligent AI camera-moving robot according to claim 3, characterized in that: The embodied intelligent large model in the core control and AI processing module is trained using a large amount of film and television camera movement clips, professional shooting parameters, scene tags, style tags and lighting conditions data, and is constructed using supervised learning and reinforcement learning methods.
5. The embodied intelligent AI camera-moving robot according to claim 4, characterized in that: The camera movement path generated by the embodied intelligent large model includes both chassis kinematic constraints and dual-arm workspace constraints, and the execution deviation is compensated at the millisecond level through a real-time feedback closed loop of chassis IMU and dual-arm joint encoder.
6. The embodied intelligent AI camera-moving robot according to claim 5, characterized in that: The gesture recognition camera and voice interaction unit are configured to: after the user selects a camera movement path, recognize preset gesture commands or voice commands, and trigger the robot to perform automated camera movement and shooting according to the selected path.
7. The embodied intelligent AI camera-moving robot according to claim 6, characterized in that: The user interaction and payment module is also configured as follows: When the user selects "Modify description and reshoot", the current user session state is retained, and the core control and AI processing module is triggered to regenerate the camera movement path based on the modified description. The reshoot does not trigger a new payment request.
8. The embodied intelligent AI camera-moving robot according to claim 7, characterized in that: It also includes a solar-assisted power supply unit, which is installed above the wheeled mobile chassis and uses a flexible solar panel design to supplement power to the robot's various modules in outdoor scenarios.
9. The control method for a body-worn intelligent AI camera-moving robot according to claim 8, characterized in that: Includes the following steps: S1. The user scans the QR code displayed on the device to enter the interactive service interface, completes online payment, and starts the shooting function; S2. Users can input a description of the shooting scene and their camera movement style requirements through the interactive interface or by voice input. S3. The robot automatically collects information on current ambient lighting, weather, spatial dimensions, and obstacle distribution through a multimodal perception module; S4. The core control and AI processing module combines user descriptions and environmental information to generate at least two alternative camera movement paths through an embodied intelligent big model, and pushes them to the user for preview and selection. S5. The user confirms the selected target camera movement path through voice or gesture commands; S6. The voice interaction unit has a built-in wake word detection model. When it recognizes any word in the preset wake word set issued by the user, it generates a trigger signal. The core control and AI processing module responds to the trigger signal, starts the camera movement path execution program, synchronously controls the wheeled mobile chassis and the dual-arm robot to perform camera movement actions, and automatically controls the mobile phone to record video. S7. During filming, the robot dynamically adjusts the camera movement path in real time according to changes in the environment to ensure safe movement and stable footage; S8. After the camera movement path is completed, the robot automatically stops recording and plays an end prompt sound, then transfers the video file to the user's mobile phone; S9. Users can choose to save the video or modify the description and replan the shooting based on the final result.
10. The control method for a body-worn intelligent AI camera-moving robot according to claim 9, characterized in that: In step S7, when the road surface is detected to be tilted or uneven, the robot adjusts the target state of the robotic arm in real time based on the acquired pose information to ensure the stability of the camera movement center during the shooting process.