Method for vehicle remote control

By using remote vehicle control methods, drones can follow vehicles in their shooting mode and edge devices can analyze image data to generate vehicle control commands. This solves the problem that users cannot operate drones and control vehicles at the same time, thus improving the self-driving tour experience and safety.

CN121500829APending Publication Date: 2026-02-10CHANGCHUN AUTOMOTIVE TEST CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511524988.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

When users take aerial photos using the drones equipped in their vehicles during self-driving tours, they cannot simultaneously operate the drones and control the vehicles manually, resulting in a poor user experience, especially in terms of meeting safety and personalization needs.

Method used

A method for remote vehicle control is provided, which uses a drone to follow the vehicle and capture images in two modes: autonomous and manual. In autonomous mode, the drone flies automatically and collects image data, and the edge device generates vehicle control commands to achieve remote control. In manual mode, the driver manually operates the drone, and the edge device analyzes the image data to generate vehicle control commands.

Benefits of technology

This allows users to capture images of vehicles in motion while manually operating the drone, meeting their personalized needs, enhancing the user experience, and ensuring driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121500829A_ABST
    Figure CN121500829A_ABST
Patent Text Reader

Abstract

The invention provides a method for vehicle remote control, and the method comprises the steps: further selecting an unmanned aerial vehicle following shooting mode after a driver inputs an unmanned aerial vehicle following shooting instruction, and enabling an unmanned aerial vehicle to automatically follow a target vehicle to fly according to a preset aerial shooting strategy in an autonomous vehicle following shooting mode, vehicle image data are collected in real time; in the manual vehicle following shooting mode, a driver manually controls the unmanned aerial vehicle to fly, the unmanned aerial vehicle collects vehicle image data in real time and transmits the vehicle image data to the target vehicle airport, the target vehicle airport transmits the vehicle image data to the edge device, and the edge device generates a vehicle control instruction based on the vehicle image data. And the vehicle control instruction is sent to the target vehicle machine, and the target vehicle machine controls the target vehicle according to the vehicle control instruction, so that a user can shoot a picture in the running process of the vehicle while manually operating the unmanned aerial vehicle, more individual requirements of the user are met, and the use experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of vehicle remote control technology, and in particular to a method for vehicle remote control. BACKGROUND

[0002] With the development of civilian unmanned aerial vehicle technology, more and more users begin to purchase unmanned aerial vehicles for aerial photography. Compared with general cameras, unmanned aerial vehicles have a wider and freer view, and can take pictures at angles that cannot be achieved by conventional shooting. At present, some car manufacturers start to cooperate with unmanned aerial vehicle manufacturers to equip unmanned aerial vehicles on some new energy vehicle models in order to cater to the needs of consumers and expand the market. Users can view the image pictures taken by the unmanned aerial vehicle in real time through the vehicle machine, and the vehicle machine also has an automatic control program of the unmanned aerial vehicle, which can control the unmanned aerial vehicle to follow the vehicle and take pictures in a preset manner. Users who cannot operate the unmanned aerial vehicle skillfully can also obtain image pictures with aesthetic sense. Users can also manually control the unmanned aerial vehicle through a remote control device, thereby greatly improving the shooting experience of users during self-driving.

[0003] However, there are still some problems when users use the unmanned aerial vehicle equipped in the vehicle for aerial photography. For example, although the vehicle machine has a plurality of preset unmanned aerial vehicle automatic aerial photography schemes to meet the aerial photography needs of novice users, users who can operate the unmanned aerial vehicle skillfully tend to operate the unmanned aerial vehicle autonomously to obtain more unique aerial photography angles and pictures through different panning ways. However, when operating the unmanned aerial vehicle autonomously, the user cannot control the vehicle. For users who drive alone, they cannot take pictures of the vehicle driving while manually operating the unmanned aerial vehicle due to safety considerations, and the user experience needs to be further improved. SUMMARY

[0004] The present application aims to provide a method for vehicle remote control to solve the above technical problems existing in the prior art.

[0005] To achieve the above application purpose, the present application provides a method for vehicle remote control, which comprises: S101, a driving personnel inputs a unmanned aerial vehicle following vehicle shooting instruction, and a target vehicle machine releases the unmanned aerial vehicle after receiving the unmanned aerial vehicle following vehicle shooting instruction; S102, the unmanned aerial vehicle flies in a cruising state and follows the target vehicle after taking off; S103, the driving personnel selects a unmanned aerial vehicle following vehicle shooting mode, the unmanned aerial vehicle following vehicle shooting mode comprising an autonomous following vehicle shooting mode and a manual following vehicle shooting mode, step S104 is executed when the driving personnel selects the autonomous following vehicle shooting mode, and step S105 is executed when the driving personnel selects the manual following vehicle shooting mode. S104. The drone automatically follows the target vehicle according to the preset aerial photography strategy and collects vehicle image data in real time, and transmits the vehicle image data to the target vehicle's in-vehicle system. S105. The driver manually controls the drone to fly, while the drone collects vehicle image data in real time and transmits the vehicle image data to the target vehicle's infotainment system. The target vehicle's infotainment system transmits the vehicle image data to the edge device. The edge device generates vehicle control commands based on the vehicle image data and sends them to the target vehicle's infotainment system. The target vehicle's infotainment system controls the target vehicle according to the vehicle control commands.

[0006] Furthermore, step S103 specifically includes the following operations: S201. The vehicle's infotainment system acquires the real-time location of the target vehicle. S202. The target vehicle's in-vehicle infotainment system acquires real-time electronic fence data and determines the electronic fence area based on the real-time electronic fence data. S203. Determine whether the target vehicle is in the electronic fence area based on the real-time location of the target vehicle, or whether the distance between the target vehicle and the electronic fence area is less than the preset distance threshold and has a continuous shortening trend. S204. If the target vehicle is within the electronic fence area, or the distance between the target vehicle and the electronic fence area is less than the preset distance threshold and shows a continuous shortening trend, then the driver is prohibited from selecting the manual following and shooting mode.

[0007] Furthermore, step S104 specifically includes the following operations: S301. The target vehicle's infotainment system displays selectable preset aerial photography strategies, and the driver selects the preset aerial photography strategy through the target vehicle's infotainment system. S302, The vehicle's infotainment system sends the preset aerial photography strategy selected by the driver to the drone; The S303 drone automatically follows the target vehicle according to a preset aerial photography strategy, collects vehicle image data in real time during flight, and transmits the vehicle image data to the target vehicle's onboard unit.

[0008] Furthermore, the target vehicle's in-vehicle infotainment system displays selectable aerial photography strategies, specifically including the following operations: S401. The vehicle's infotainment system obtains the current time and the real-time location of the target vehicle. S402. Obtain popular aerial photography vehicle videos through the API interface of a third-party platform, input the popular aerial photography vehicle videos into the aerial photography strategy recognition model for processing, and obtain the aerial photography strategy recognition results. The aerial photography strategy recognition results include, but are not limited to, aerial photography location, aerial photography time, aerial photography trajectory, and camera movement method. S403. Based on the current time and the real-time location of the target vehicle, match each aerial photography strategy identification result with the current time and determine the first weight coefficient of each aerial photography strategy identification result based on the matching results. S404. Determine the second weighting coefficient of the aerial photography strategy recognition result based on the popularity of the popular aerial photography vehicle videos corresponding to the recognition results of each aerial photography strategy. S405. Generate optional aerial photography strategies based on the aerial photography strategy recognition results, and determine the sorting method of optional aerial photography strategies for visualization display based on the first weight coefficient and the second weight coefficient of the aerial photography strategy recognition results. S406. When the target vehicle's infotainment system displays the available aerial photography strategies, it sorts the available aerial photography strategies according to the sorting method determined in step S405.

[0009] Furthermore, the target vehicle's in-vehicle infotainment system displays selectable aerial photography strategies, specifically including the following operations: S501, the target vehicle's infotainment system acquires the driver's historical selection of aerial photography strategies; S502. Perform cluster analysis on historical aerial photography strategies to obtain strategy preference information; S503. Match the strategy preference information with the current available aerial photography strategies, and determine the third weighting coefficient of the available aerial photography strategies based on the degree of matching. S504. Adjust the ranking of the selectable aerial photography strategies after sorting in step S406 according to the third weighting coefficient.

[0010] Furthermore, step S105 specifically includes the following operations: S601. Determine the appropriate driving speed for the target vehicle based on the manual control parameters of the drone set by the operator. S602, The vehicle's infotainment system will transmit the vehicle image data to the edge device according to the driving speed; S603. The edge device inputs vehicle image data into the scene target detection model for processing to obtain scene target detection results. The scene target detection results include, but are not limited to, roads, vehicles, pedestrians, and obstacles. S604: The edge device continuously generates vehicle control commands based on real-time scene target detection results and adapted driving speed, and transmits the vehicle control commands to the target vehicle.

[0011] Furthermore, step S105 also includes the following operations: S701, The driver issues a voice wake-up command to the target vehicle's infotainment system; S702. After the target vehicle's infotainment system receives the voice wake-up command, it prompts the driver to input a voice command. S703: The driver sends a speed limit voice command to the target vehicle's infotainment system. The target vehicle's infotainment system receives and recognizes the speed limit voice command and obtains the speed limit command. S704. The vehicle's infotainment system transmits the speed limit command to the edge device. When the edge device detects the speed limit command, it continuously generates vehicle control commands based on the real-time scene target detection results and the speed limit command.

[0012] Furthermore, the vehicle image data is input into the scene target detection model for processing, specifically including the following operations: S801. The vehicle image data is split into multiple time-series continuous multi-view images, and a top view of the driving scene at different times is generated based on the multi-view images. S802. Extract the time features of the top view of the driving scene at different times, and further fuse the information of different spatial ranges in the top view of the driving scene to obtain the spatiotemporal features of the top view of the scene. S803. Decode the spatiotemporal features of the scene top view to obtain the position and classification information of different targets in the current scene.

[0013] Furthermore, step S802 specifically includes the following operations: S901. For the current driving scene top view, obtain the previous n frames of historical driving scene top views, and align the previous n frames of historical driving scene top views with the current driving scene top view. S902. The aligned top view of the driving scene is processed through multiple 3D CNN models and average pooling layers to obtain a global spatiotemporal semantic context representation. After compressing the channel dimensions, a rasterized top view of the driving scene is obtained. S903. Input the rasterized top view of the driving scene into the spatial convolutional layer for processing, and obtain residual features by passing the processing result through residual connections. S904. Input the residual features into the spatial pooling layer and output the first feature map at different scales. At the same time, perform pooling and upsampling operations on the residual features to obtain the second feature map. Concatenate the first feature map and the second feature map to obtain the concatenated feature. S905. Dimensionally reduce the splicing features to obtain the spatiotemporal features of the scene top view.

[0014] Furthermore, step S803 specifically includes the following operations: S1001. Input the spatiotemporal features of the scene top view into the target segmentation head for processing, segment the scene targets from the scene top view, assign a unique identifier to each segmented scene target, and obtain the classification information of the scene targets. S1002. Determine the geometric center of each scene target obtained from the segmentation in the scene top view, and obtain the position information of the scene target in the scene top view.

[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention provides a method for remote vehicle control. When using a drone equipped in a vehicle, the driver can choose between an autonomous vehicle-following shooting mode or a manual vehicle-following shooting mode according to their needs. In the manual vehicle-following shooting mode, the user can manually operate the drone to obtain a more unique and free aerial shooting perspective and ideal footage. On the other hand, based on the vehicle image data collected by the drone, the target vehicle's onboard unit transmits it to an edge device for analysis and processing, thereby generating vehicle control commands. By sending vehicle control commands to the target vehicle's onboard unit, remote control of the target vehicle is achieved. This allows the user to capture footage of their own vehicle while manually operating the drone, thus meeting more personalized user needs and improving the user experience. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only preferred embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the overall structure of a method for remote vehicle control provided in an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the process for the electronic fence restriction of drone following and shooting modes provided in an embodiment of the present invention.

[0019] Figure 3 This is a schematic diagram of the autonomous vehicle-following shooting mode provided in an embodiment of the present invention.

[0020] Figure 4 This is a schematic diagram of the process for visually displaying selectable aerial photography strategies provided in an embodiment of the present invention.

[0021] Figure 5 This is a schematic diagram of the process for optimizing the ranking of optional aerial photography strategies based on historical selection strategies, provided in an embodiment of the present invention.

[0022] Figure 6 This is a schematic diagram of the target vehicle speed control process adapted to the control parameters of the unmanned aerial vehicle provided in the embodiment of the present invention.

[0023] Figure 7This is a schematic diagram of the target vehicle speed control process based on voice commands provided in an embodiment of the present invention.

[0024] Figure 8 This is a schematic diagram of the scene target detection model processing flow provided in the embodiment of the present invention.

[0025] Figure 9 This is a schematic diagram of the spatiotemporal feature extraction process of a scene top view provided in an embodiment of the present invention.

[0026] Figure 10 This is a schematic diagram of the spatiotemporal feature decoding process of a scene top view provided in an embodiment of the present invention. Detailed Implementation

[0027] The principles and features of the present invention are described below with reference to the accompanying drawings. The listed embodiments are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0028] Reference Figure 1 This embodiment provides a method for remote vehicle control, the method comprising: S101. The driver inputs the drone's command to follow the vehicle and film. After receiving the command, the vehicle's infotainment system releases the drone.

[0029] In this embodiment, the target vehicle is the vehicle driven by the driver, and the vehicle infotainment system is the intelligent vehicle infotainment system installed in the target vehicle.

[0030] For example, when not in use, the drone is stored in a cabin on the roof of the target vehicle. This cabin may have an electrically operated hatch that opens when the drone needs to be launched and closes when the drone is stored. The vehicle's infotainment system can directly control the opening and closing of the hatch. When the driver inputs a drone follow-the-vehicle filming command through the vehicle's infotainment system, the system opens the hatch and sends a takeoff command to the drone.

[0031] S102. After takeoff, the drone will fly in cruise mode and follow the target vehicle.

[0032] In this step, after the drone takes off, it first maintains a preset altitude and keeps the distance between it and the target vehicle relatively stable, following the target vehicle as it moves, in order to wait for further instructions from the vehicle's infotainment system.

[0033] S103. The driver selects the drone following vehicle shooting mode. The drone following vehicle shooting mode includes autonomous following vehicle shooting mode and manual following vehicle shooting mode. When the driver selects the autonomous following vehicle shooting mode, step S104 is executed. When the driver selects the manual following vehicle shooting mode, step S105 is executed.

[0034] For example, in this step, the driver selects the drone following mode through the vehicle's infotainment system. This operation can be performed after the drone takes off, or it can be performed immediately after the driver inputs the drone following command. In the latter case, the drone directly enters the drone following mode selected by the driver after takeoff.

[0035] S104. The drone automatically follows the target vehicle according to the preset aerial photography strategy and collects vehicle image data in real time, transmitting the vehicle image data to the target vehicle's in-vehicle system.

[0036] S105. The driver manually controls the drone to fly, while the drone collects vehicle image data in real time and transmits the vehicle image data to the target vehicle's infotainment system. The target vehicle's infotainment system transmits the vehicle image data to the edge device. The edge device generates vehicle control commands based on the vehicle image data and sends them to the target vehicle's infotainment system. The target vehicle's infotainment system controls the target vehicle according to the vehicle control commands.

[0037] In this embodiment, when the driver selects the manual following and shooting mode, the driver manually controls the drone's flight and shooting via the drone remote control device. At this time, the target vehicle's driving control is achieved by the vehicle's onboard unit executing vehicle control commands issued by the edge device. The vehicle control commands generated by the edge device are generated after analyzing the vehicle image data collected by the drone. The vehicle image data collected by the drone during aerial photography often includes not only the target vehicle itself but also roads, other vehicles, pedestrians, obstacles, and other targets. The edge device identifies these targets in the vehicle image data and the target vehicle's direction of travel, generating corresponding control commands to help the target vehicle travel along the lane and avoid pedestrians, obstacles, and other vehicles.

[0038] In this embodiment, for users who are not proficient in operating drones but want to capture aerial videos or photos during a road trip, the autonomous vehicle-following shooting mode can be selected. This allows the drone to fly and capture images according to a preset aerial photography strategy, eliminating the need for manual operation and making it simple and easy to use. For users with more personalized aerial photography needs, and who do not have other drivers to replace them or whose companions are unable to drive, the manual vehicle-following shooting mode can be selected. In this mode, the driver can manually control the drone, while the vehicle's driving is remotely controlled by an edge device that analyzes the vehicle image data transmitted by the drone. This satisfies the user's personalized needs and ensures driving safety.

[0039] For example, when a driver manually controls a drone, the footage captured by the drone may not continuously include the target vehicle and road. To better implement the method provided in this embodiment without affecting the user experience, the drone can be equipped with two sets of lenses. One set is a main lens, responsible for capturing the images the user wants to capture; its captured footage may not include the target vehicle and road. The other set is a secondary lens, responsible for continuously tracking the target vehicle, ensuring that the captured footage includes the target vehicle and road, thereby continuously providing edge devices with raw data to support the generation of control commands. When the driver selects the autonomous following shooting mode, the secondary lens is in standby mode. In some specific embodiments, the secondary lens can be activated only when the target vehicle and road are no longer visible in the footage captured by the main lens.

[0040] As an optional implementation method, refer to Figure 2 Step S103 specifically includes the following operations: S201, The vehicle's infotainment system obtains the real-time location of the target vehicle.

[0041] For example, the vehicle's infotainment system can locate the vehicle's real-time position using the vehicle's built-in GPS or BeiDou positioning device.

[0042] S202. The target vehicle's in-vehicle infotainment system acquires real-time electronic fence data and determines the electronic fence area based on the real-time electronic fence data.

[0043] S203. Determine whether the target vehicle is within the electronic fence area based on its real-time location, or whether the distance between the target vehicle and the electronic fence area is less than a preset distance threshold and shows a continuous shortening trend.

[0044] S204. If the target vehicle is within the electronic fence area, or the distance between the target vehicle and the electronic fence area is less than the preset distance threshold and shows a continuous shortening trend, then the driver is prohibited from selecting the manual following and shooting mode.

[0045] In manual vehicle-following shooting mode, the target vehicle's movement is taken over by the edge device, and the driver is primarily responsible for operating the drone. This mode is only suitable for sparsely populated outdoor environments with simple road traffic conditions. In urban areas or rural towns, due to the variable road traffic conditions, manual vehicle-following shooting mode needs to be disabled for traffic safety reasons. In this implementation, multiple electronic fence zones are pre-defined on the map. Within or near these electronic fence zones, manual vehicle-following shooting mode cannot be used to ensure driving safety. The extent of the electronic fence zones is determined based on electronic fence data, which can be updated in real time.

[0046] In some further implementations, different drone following modes are disabled in different electronic fence areas. For example, electronic fence area A only restricts users from enabling the manual following mode, while electronic fence area B restricts users from enabling all drone following modes.

[0047] As another possible implementation, refer to Figure 3 Step S104 specifically includes the following operations: S301. The target vehicle's infotainment system displays selectable preset aerial photography strategies, which the driver can choose through the vehicle's infotainment system.

[0048] S302, The vehicle's infotainment system sends the preset aerial photography strategy selected by the driver to the drone.

[0049] The S303 drone automatically follows the target vehicle according to a preset aerial photography strategy, collects vehicle image data in real time during flight, and transmits the vehicle image data to the target vehicle's onboard unit.

[0050] In this embodiment, the preset aerial photography strategy can be stored in a strategy library locally on the target vehicle's infotainment system; or it can be stored in a cloud-based strategy library. When the driver selects the autonomous following shooting mode, the target vehicle's infotainment system retrieves the latest preset aerial photography strategy from the cloud-based strategy library and displays it to the driver on the infotainment screen for selection. The preset aerial photography strategy controls the drone's automatic flight path, flight speed, and camera movement around the target vehicle. When the driver selects a preset aerial photography strategy, the target vehicle's infotainment system controls the drone to automatically fly according to the selected strategy and capture vehicle image data. The vehicle image data captured by the drone is transmitted back to the target vehicle's infotainment system for further processing, including but not limited to editing, saving, and sharing the vehicle image data. In other words, after the drone automatically collects vehicle image data, the target vehicle's infotainment system can also edit it with one click to generate a finished video, allowing users to quickly share it on social media and further enhance the user experience.

[0051] As a further possible implementation, refer to Figure 4 The target vehicle's in-vehicle infotainment system displays selectable aerial photography strategies, including the following operations: S401, The vehicle's infotainment system obtains the current time and the real-time location of the target vehicle.

[0052] S402. Obtain popular aerial photography vehicle videos through the API interface of a third-party platform, input the popular aerial photography vehicle videos into the aerial photography strategy recognition model for processing, and obtain the aerial photography strategy recognition results. The aerial photography strategy recognition results include, but are not limited to, aerial photography location, aerial photography time, aerial photography trajectory, and camera movement method.

[0053] In this embodiment, the aerial photography strategy recognition model is a neural network model that has been pre-trained using labeled sample aerial photography vehicle videos and sample aerial photography strategy recognition results.

[0054] S403. Based on the current time and the real-time location of the target vehicle, match each aerial photography strategy identification result with the current time and determine the first weight coefficient of each aerial photography strategy identification result based on the matching results.

[0055] S404. Determine the second weighting coefficient of the aerial photography strategy recognition result based on the popularity of the popular aerial vehicle videos corresponding to each aerial photography strategy recognition result.

[0056] S405. Generate optional aerial photography strategies based on the aerial photography strategy recognition results, and determine the sorting method of optional aerial photography strategies for visualization display based on the first weight coefficient and the second weight coefficient of the aerial photography strategy recognition results.

[0057] S406. When the target vehicle's infotainment system displays the available aerial photography strategies, it sorts the available aerial photography strategies according to the sorting method determined in step S405.

[0058] In this implementation, the first weight coefficient of each aerial photography strategy identification result reflects the degree of matching between its corresponding aerial vehicle video and the current time and the real-time location of the target vehicle. It is understandable that the closer the shooting time and location of popular aerial vehicle videos obtained from third-party platforms match the current time and the real-time location of the target vehicle, the easier it is for the vehicle image data collected according to its corresponding aerial photography strategy to achieve the user's desired "output" effect. The popularity of popular aerial vehicle videos largely reflects the degree of public acceptance of the underlying aerial photography strategy. This implementation first identifies the aerial photography strategy based on popular online aerial vehicle videos, then determines its first and second weight coefficients based on time, location, and popularity, and finally prioritizes pushing aerial photography strategies with high matching and popularity to users based on the first and second weight coefficients. This improves user satisfaction with vehicle image data collected by drones in autonomous vehicle-following shooting mode, thereby further enhancing the user experience.

[0059] As a further implementation method, refer to Figure 5 The target vehicle's in-vehicle infotainment system displays selectable aerial photography strategies, which include the following operations: S501. The vehicle's infotainment system acquires the driver's historical aerial photography strategy selections. These historical aerial photography strategies are preset strategies previously selected by the driver.

[0060] S502. Perform cluster analysis on historical aerial photography strategies to obtain strategy preference information.

[0061] S503. Match the strategy preference information with the current available aerial photography strategies, and determine the third weight coefficient of the available aerial photography strategies based on the degree of matching.

[0062] S504. Adjust the ranking of the selectable aerial photography strategies after sorting in step S406 according to the third weighting coefficient.

[0063] In this implementation, by recording the driver's historical aerial photography strategies and performing cluster analysis on all historical aerial photography strategies, strategies with high commonality in aerial photography trajectories / camera movement are grouped into the same cluster. Then, by comparing the number of historical aerial photography strategies in each cluster, the cluster containing the most strategies is determined. By analyzing the cluster center of the cluster containing the most strategies, the user's preferred aerial photography strategy is determined, obtaining strategy preference information. This strategy preference information is then matched with the currently available aerial photography strategies. Based on the matching results, a third weight coefficient for the available aerial photography strategies is determined. The available aerial photography strategies are further ranked and optimized based on this third weight coefficient. This approach, in addition to considering time, location, and popularity, further considers user preferences to prioritize and push preset aerial photography strategies that may better suit the user's preferences, thereby further improving the user experience.

[0064] As another possible implementation method, refer to Figure 6 Step S105 specifically includes the following operations: S601. Determine the appropriate driving speed for the target vehicle based on the operator's manual control parameters for the drone.

[0065] In this embodiment, the manual control parameters include at least the flight speed of the drone.

[0066] S602, the target vehicle's infotainment system will transmit the vehicle's speed and image data to the edge device.

[0067] S603. The edge device inputs vehicle image data into the scene target detection model for processing to obtain scene target detection results, including but not limited to roads, vehicles, pedestrians, and obstacles.

[0068] S604: The edge device continuously generates vehicle control commands based on real-time scene target detection results and adapted driving speed, and transmits the vehicle control commands to the target vehicle.

[0069] In this implementation, in the autonomous vehicle-following shooting mode, when the edge device generates control commands, it considers the flight speed of the drone and the scene target detection results when controlling the vehicle speed, thereby ensuring driving safety. While ensuring that the vehicle speed does not exceed the preset threshold, it matches the flight speed of the drone and avoids collisions with pedestrians, vehicles and obstacles on the road, thus achieving dual protection of aerial shooting effect and driving safety.

[0070] As a further possible implementation, refer to Figure 7 Step S105 also includes the following operations: S701, The driver issues a voice wake-up command to the target vehicle's infotainment system.

[0071] S702. After receiving the voice wake-up command, the vehicle's infotainment system prompts the driver to input a voice command.

[0072] S703: The driver issues a speed limit voice command to the target vehicle's infotainment system. The target vehicle's infotainment system receives and recognizes the speed limit voice command and obtains the speed limit instruction.

[0073] S704. The vehicle's infotainment system transmits the speed limit command to the edge device. When the edge device detects the speed limit command, it continuously generates vehicle control commands based on the real-time scene target detection results and the speed limit command.

[0074] In this implementation, the driver can control the speed of the target vehicle by giving voice commands. As long as the driver's intended speed does not exceed the preset safe speed threshold, the edge device prioritizes the driver's needs over the drone's flight speed when controlling the target vehicle's movement speed. In manual following and shooting mode, the driver can focus on controlling the drone while conveniently controlling the vehicle's movement speed to capture the desired footage, thus further enhancing the user experience.

[0075] As a further possible implementation, refer to Figure 8 The vehicle image data is input into the scene target detection model for processing, which specifically includes the following operations: S801. The vehicle image data is split into multiple time-series continuous multi-view images, and a top view of the driving scene at different times is generated based on the multi-view images.

[0076] In this step, firstly, for any given moment, the multi-view image is downsampled using a convolutional neural network as the backbone to obtain its image and depth features. The outer product of these features is then calculated to obtain the 3D features of the multi-view image in the UAV camera coordinate system. Through a mapping transformation, the 3D features of the multi-view image in the UAV camera coordinate system are converted into image features in the top-view coordinate system, denoted as the top-view image features. The top-view image features are then normalized to avoid distortion caused by noise or outliers. Finally, all top-view image features are kept at the same scale. Dimensionality reduction is then used to obtain the top-view view of the driving scene at each moment.

[0077] S802. Extract the time features of the top view of the driving scene at different times, and further fuse the information of different spatial ranges in the top view of the driving scene to obtain the spatiotemporal features of the top view of the scene.

[0078] S803. Decode the spatiotemporal features of the scene top view to obtain the position and classification information of different targets in the current scene.

[0079] Specifically, refer to Figure 9 Step S802 specifically includes the following operations: S901. For the current driving scene top view, obtain the previous n frames of historical driving scene top views, and align the previous n frames of historical driving scene top views with the current driving scene top view.

[0080] S902. The aligned top view of the driving scene is processed through multiple 3D CNN models and average pooling layers to obtain a global spatiotemporal semantic context representation. After compressing the channel dimensions, a rasterized top view of the driving scene is obtained.

[0081] S903. Input the rasterized top view of the driving scene into the spatial convolutional layer for processing, and obtain residual features by passing the processing result through residual connections.

[0082] S904. Input the residual features into the spatial pooling layer and output the first feature map at different scales. At the same time, perform pooling and upsampling operations on the residual features to obtain the second feature map. Concatenate the first feature map and the second feature map to obtain the concatenated feature.

[0083] S905. Dimensionally reduce the splicing features to obtain the spatiotemporal features of the scene top view.

[0084] In this implementation, as the vehicle moves, the surrounding scene continuously changes over time, and there is a clear correlation between the temporally consecutive top-down views of the scene. Therefore, after aligning the current driving scene top-down view with the driving scene top-down views of the previous n frames, the global spatiotemporal semantic context information is first captured using a 3D CNN model and an average pooling layer. Then, channel dimension compression is performed using convolutional layers to achieve continuous contextual semantic compensation in the temporal dimension. Next, spatial convolutional layers compensate for long-distance features in the driving scene top-down view, and spatial pooling layers use different dilation rates to process residual features, obtaining first feature maps of different scales. The first feature maps are then concatenated with second feature maps to obtain concatenated features of different levels. Finally, by reducing the dimensionality of the concatenated features and extracting higher-level semantic information, the spatiotemporal features of the scene top-down view are obtained.

[0085] Reference Figure 10 Step S803 specifically includes the following operations: S1001. Input the spatiotemporal features of the scene top view into the target segmentation head for processing, segment the scene targets from the scene top view, assign a unique identifier to each segmented scene target, and obtain the classification information of the scene targets.

[0086] S1002. Determine the geometric center of each scene target obtained from the segmentation in the scene top view, and obtain the position information of the scene target in the scene top view.

[0087] This implementation generates top-down views of the driving scene at different times by using multi-view images captured by drones to expand the perception range of the vehicle's driving environment. It further extracts the temporal features of the top-down views and fuses information from different spatial ranges in the top-down views to obtain the spatiotemporal features of the top-down views. Finally, by decoding the spatiotemporal features of the top-down views, it obtains the position and classification information of different targets in the scene where the target vehicle is located. This implementation can capture the key dynamic features of each target in the scene where the target vehicle is located, providing better data support for edge devices to analyze and generate control commands, and meeting the environmental perception requirements of autonomous driving.

[0088] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for remote vehicle control, characterized in that, The method includes: S101. The driver inputs the drone's command to follow the vehicle and film. After the target vehicle's infotainment system receives the command, it releases the drone. S102. After the drone takes off, it flies in cruise mode and follows the target vehicle. S103. The driver selects the drone following vehicle shooting mode. The drone following vehicle shooting mode includes autonomous following vehicle shooting mode and manual following vehicle shooting mode. When the driver selects the autonomous following vehicle shooting mode, step S104 is executed. When the driver selects the manual following vehicle shooting mode, step S105 is executed. S104. The drone automatically follows the target vehicle according to the preset aerial photography strategy and collects vehicle image data in real time, and transmits the vehicle image data to the vehicle's infotainment system. S105. The driver manually controls the drone to fly, while the drone collects vehicle image data in real time and transmits the vehicle image data to the target vehicle's infotainment system. The target vehicle's infotainment system transmits the vehicle image data to the edge device. The edge device generates vehicle control commands based on the vehicle image data and sends them to the target vehicle's infotainment system. The target vehicle's infotainment system controls the target vehicle according to the vehicle control commands.

2. The method for remote vehicle control according to claim 1, characterized in that, Step S103 specifically includes the following operations: S201. The vehicle's infotainment system acquires the real-time location of the target vehicle. S202. The target vehicle's in-vehicle infotainment system acquires real-time electronic fence data and determines the electronic fence area based on the real-time electronic fence data. S203. Determine whether the target vehicle is in the electronic fence area based on the real-time location of the target vehicle, or whether the distance between the target vehicle and the electronic fence area is less than the preset distance threshold and has a continuous shortening trend. S204. If the target vehicle is within the electronic fence area, or the distance between the target vehicle and the electronic fence area is less than the preset distance threshold and shows a continuous shortening trend, then the driver is prohibited from selecting the manual following and shooting mode.

3. The method for remote vehicle control according to claim 1, characterized in that, Step S104 specifically includes the following operations: S301. The target vehicle's infotainment system displays selectable preset aerial photography strategies, and the driver selects the preset aerial photography strategy through the target vehicle's infotainment system. S302, The vehicle's infotainment system sends the preset aerial photography strategy selected by the driver to the drone; The S303 drone automatically follows the target vehicle according to a preset aerial photography strategy, collects vehicle image data in real time during flight, and transmits the vehicle image data to the target vehicle's onboard unit.

4. The method for remote vehicle control according to claim 3, characterized in that, The target vehicle's in-vehicle infotainment system displays selectable aerial photography strategies, including the following operations: S401. The vehicle's infotainment system obtains the current time and the real-time location of the target vehicle. S402. Obtain popular aerial photography vehicle videos through the API interface of a third-party platform, input the popular aerial photography vehicle videos into the aerial photography strategy recognition model for processing, and obtain the aerial photography strategy recognition results. The aerial photography strategy recognition results include, but are not limited to, aerial photography location, aerial photography time, aerial photography trajectory, and camera movement method. S403. Based on the current time and the real-time location of the target vehicle, match each aerial photography strategy identification result with the current time and determine the first weight coefficient of each aerial photography strategy identification result based on the matching results. S404. Determine the second weighting coefficient of the aerial photography strategy recognition result based on the popularity of the popular aerial photography vehicle videos corresponding to the recognition results of each aerial photography strategy. S405. Generate selectable aerial photography strategies based on the aerial photography strategy recognition results, and determine the sorting method of the selectable aerial photography strategies for visualization display based on the first weight coefficient and the second weight coefficient of the aerial photography strategy recognition results. S406. When the target vehicle's infotainment system displays the available aerial photography strategies, it sorts the available aerial photography strategies according to the sorting method determined in step S405.

5. A method for remote vehicle control according to claim 4, characterized in that, The target vehicle's in-vehicle infotainment system displays selectable aerial photography strategies, which include the following operations: S501, the target vehicle's infotainment system acquires the driver's historical selection of aerial photography strategies; S502. Perform cluster analysis on historical aerial photography strategies to obtain strategy preference information; S503. Match the strategy preference information with the current available aerial photography strategies, and determine the third weighting coefficient of the available aerial photography strategies based on the degree of matching. S504. Adjust the ranking of the selectable aerial photography strategies after sorting in step S406 according to the third weighting coefficient.

6. The method for remote vehicle control according to claim 1, characterized in that, Step S105 specifically includes the following operations: S601. Determine the appropriate driving speed for the target vehicle based on the manual control parameters of the drone set by the operator. S602, The vehicle's infotainment system will transmit the vehicle image data to the edge device according to the driving speed; S603. The edge device inputs vehicle image data into the scene target detection model for processing to obtain scene target detection results. The scene target detection results include, but are not limited to, roads, vehicles, pedestrians, and obstacles. S604: The edge device continuously generates vehicle control commands based on real-time scene target detection results and adapted driving speed, and transmits the vehicle control commands to the target vehicle.

7. A method for remote vehicle control according to claim 6, characterized in that, Step S105 also includes the following operations: S701, The driver issues a voice wake-up command to the target vehicle's infotainment system; S702. After the target vehicle's infotainment system receives the voice wake-up command, it prompts the driver to input a voice command. S703: The driver sends a speed limit voice command to the target vehicle's infotainment system. The target vehicle's infotainment system receives and recognizes the speed limit voice command and obtains the speed limit command. S704. The vehicle's infotainment system transmits the speed limit command to the edge device. When the edge device detects the speed limit command, it continuously generates vehicle control commands based on the real-time scene target detection results and the speed limit command.

8. A method for remote vehicle control according to claim 6, characterized in that, The vehicle image data is input into the scene object detection model for processing, which includes the following operations: S801. The vehicle image data is split into multiple time-series continuous multi-view images, and a top view of the driving scene at different times is generated based on the multi-view images. S802. Extract the time features of the top view of the driving scene at different times, and further fuse the information of different spatial ranges in the top view of the driving scene to obtain the spatiotemporal features of the top view of the scene. S803. Decode the spatiotemporal features of the scene top view to obtain the position and classification information of different targets in the current scene.

9. A method for remote vehicle control according to claim 8, characterized in that, Step S802 specifically includes the following operations: S901. For the current driving scene top view, obtain the previous n frames of historical driving scene top views, and align the previous n frames of historical driving scene top views with the current driving scene top view. S902. The aligned top view of the driving scene is processed through multiple 3D CNN models and average pooling layers to obtain a global spatiotemporal semantic context representation. After compressing the channel dimensions, a rasterized top view of the driving scene is obtained. S903. Input the rasterized top view of the driving scene into the spatial convolutional layer for processing, and obtain residual features by passing the processing result through residual connections. S904. Input the residual features into the spatial pooling layer and output the first feature map at different scales. At the same time, perform pooling and upsampling operations on the residual features to obtain the second feature map. Concatenate the first feature map and the second feature map to obtain the concatenated feature. S905. Dimensionally reduce the splicing features to obtain the spatiotemporal features of the scene top view.

10. A method for remote vehicle control according to claim 8, characterized in that, Step S803 specifically includes the following operations: S1001. Input the spatiotemporal features of the scene top view into the target segmentation head for processing, segment the scene targets from the scene top view, assign a unique identifier to each segmented scene target, and obtain the classification information of the scene targets. S1002. Determine the geometric center of each scene target obtained from the segmentation in the scene top view, and obtain the position information of the scene target in the scene top view.