A Drone Live Streaming Control System and Method for Scenic Spots Based on AI Interaction
By constructing a system architecture that integrates autonomous flight control of drones, AI-powered intelligent content generation, and multimodal human-computer interaction, the system addresses the issues of lack of interactive capabilities and high operating costs in scenic area drone live streaming systems. It enables drones to take off and land autonomously, interact in real time, and avoid obstacles in complex environments, thereby improving user experience and system efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANCHANG UNIV
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-05
AI Technical Summary
Existing drone live streaming systems for scenic areas suffer from problems such as fixed live streaming angles, lack of interactive capabilities, high operating costs, and discontinuous content production, making it difficult to meet users' needs for an immersive, dynamic, and participatory real-time viewing experience.
A system architecture integrating autonomous flight control of UAVs, AI intelligent content generation, multimodal human-machine interaction, and cloud-based collaborative management is constructed. This architecture includes UAV flight units, automated helipad networks, AI interaction engines, user interaction platforms, central dispatch servers, data fusion and decision-making modules, and safety monitoring and emergency response modules, enabling UAVs to achieve autonomous take-off and landing, real-time interaction, autonomous obstacle avoidance in complex environments, and intelligent identification.
It has enabled 24/7 uninterrupted live streaming of scenic spots, improved user interaction, enhanced flight safety, protected public privacy rights, reduced system replication costs, and is ready for large-scale commercial promotion.
Smart Images

Figure CN121644839B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and drone control technology, specifically relating to a scenic area drone live streaming control system and method based on AI interaction. Background Technology
[0002] With the deep integration of artificial intelligence and unmanned systems technology, the application of intelligent unmanned equipment in the cultural tourism field is gradually evolving from single-function demonstration to a normalized and platform-based service model. As an intelligent carrier with aerial mobility, drones have demonstrated significant advantages in multiple scenarios such as aerial photography, inspection, and emergency rescue. Their flexible deployment, wide-area coverage, and real-time transmission capabilities provide a new technological path for scenic area content production and interactive experiences.
[0003] Against this backdrop, drone-based scenic area live streaming systems are gradually becoming an important part of smart tourism development. They present natural scenery and cultural landscapes from a high-altitude perspective, breaking through the limitations of traditional fixed camera views and enhancing the immersion and participation of online users.
[0004] Among them, the drone live streaming system for scenic spots particularly emphasizes the system's automated operation capabilities and human-computer interaction level. Its core goal is to achieve autonomous take-off and landing, route planning, video acquisition and real-time streaming under unattended conditions, and enhance the two-way interaction between the audience and the equipment through intelligent means, so that online users can not only "see" but also "participate" and "influence" the generation process of live streaming content, thereby building a truly dynamic, interactive and sustainable digital cultural tourism content ecosystem.
[0005] Existing technologies still face multiple technical bottlenecks in realizing drone live streaming in scenic areas: First, most systems rely on manual remote control, making it difficult to automate flight missions around the clock, resulting in low live streaming frequency, high labor costs, and unsustainable operation. Second, the interaction method is limited to one-way video stream output, lacking AI-driven real-time semantic understanding and feedback mechanisms, and cannot respond to viewers' voice or text commands to adjust flight paths or switch perspectives, resulting in severely insufficient interactivity. Third, flight control and AI decision-making systems are isolated from each other, failing to form a unified intelligent closed loop; AI is only used for post-processing content annotation or voice broadcasting, failing to deeply intervene in flight process control. Finally, the system has poor adaptability to complex scenic environments. When faced with real-world constraints such as dense crowds, electromagnetic interference, and weather changes, it lacks dynamic obstacle avoidance and task replanning capabilities based on multimodal perception, making it difficult to guarantee safety and stability. These problems are particularly prominent in cultural tourism live streaming scenarios that require high concurrency, strong interaction, and long-term operation, urgently requiring a new drone live streaming control system that deeply integrates machine intelligence, human-machine interaction, and advanced process control to solve these problems. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of existing technologies by providing an AI-interactive drone live-streaming control system and method for scenic areas, which can effectively solve the problems mentioned in the background technology. Existing scenic area live-streaming systems generally suffer from fixed live-streaming perspectives, lack of interactive capabilities, high operating costs, and discontinuous content production, making it difficult to meet users' demands for an immersive, dynamic, and participatory real-time viewing experience. This invention, by constructing a system architecture integrating drone autonomous flight control, AI intelligent content generation, multimodal human-computer interaction, and cloud-based collaborative management, achieves a fundamental transformation in scenic area live-streaming from "passive viewing" to "active participation," from "temporary activities" to "routine services," and from "single video" to "intelligent interactive content platform."
[0007] To achieve the above objectives, the present invention provides the following technical solution: On one hand, an AI-interactive scenic area drone live-streaming control system, comprising the following components: a drone flight unit for performing aerial photography and mobile live-streaming tasks, and transmitting high-definition video streams in real time via a 5G network; an automated helipad network distributed at key nodes in the scenic area, providing drones with automatic take-off and landing, charging, and environmental self-checking services, supporting 24 / 7 continuous operation; an AI interaction engine integrating natural language processing, computer vision, and speech synthesis modules, used to parse user interaction commands, generate intelligent narration content, and drive virtual anchors to respond in real time; and a user interaction platform providing mobile and web interfaces, supporting online... Users can send flight route requests, ask questions, make donations, and participate in live-stream e-commerce activities; the central dispatch server is responsible for receiving user requests, planning flight paths, coordinating drone resource allocation, dispatching AI services, and managing live stream distribution; the data fusion and decision-making module integrates scenic area geographic information, tourist distribution heat maps, meteorological data, and user behavior data to generate dynamic flight strategies and content recommendation schemes; the live stream push and content management module is responsible for real-time processing of the original video stream, overlaying AI-generated content, and pushing it to third-party live-streaming platforms; the safety monitoring and emergency response module monitors the drone status, airspace compliance, and privacy protection policy implementation in real time, and triggers automatic return-to-home or emergency stop mechanisms in case of abnormalities;
[0008] Preferably, the UAV flight unit is equipped with multispectral imaging equipment, including a 4K visible light camera, an infrared thermal imager and a lidar. The visible light camera is used for conventional high-definition video acquisition, the infrared thermal imager is used to capture temperature distribution features in night travel mode to enhance visual performance, and the lidar is used to construct a local three-dimensional point cloud map to support accurate obstacle avoidance and path optimization in complex terrain. The data from the three types of sensors are synchronized to the central scheduling server after being aligned with timestamps.
[0009] Furthermore, the drone flight unit has a built-in edge computing module that runs a lightweight YOLOv8 target detection model to identify landmark buildings, vegetation types and tourist gathering areas in the scenic area in real time during flight. The identification results are embedded as semantic tags into the video stream metadata, which can be called by the AI interaction engine to generate narration content with geographical and cultural context.
[0010] In addition, the automated apron network adopts a modular design, with each apron equipped with an environmental sensing sensor group, including anemometers, rain sensors and light intensity meters, to monitor local meteorological conditions in real time. When the wind speed exceeds 12m / s or the rainfall intensity is greater than 2mm / h, it automatically sends a no-fly warning to the central dispatch server and initiates the canopy closing and equipment waterproofing procedures.
[0011] Preferably, the natural language processing module in the AI interaction engine adopts a domain adaptation model based on the BERT architecture. This model is fine-tuned using tour guide scripts, historical documents, and common tourist questions from major 5A-level scenic spots in Jiangxi Province on the basis of pre-training, so that it has the ability to accurately understand local cultural terms, dialect expressions, and historical allusions, and supports the accurate parsing of complex semantics such as "how to take the best picture of the sunset at Tengwang Pavilion" or "what stories are there in a certain region" input by users.
[0012] Furthermore, the AI interaction engine is equipped with a multi-role virtual anchor library, including historical figures, local cultural ambassadors, and cartoon IP characters. Users can select the anchor style through the interactive platform. The AI system generates narration that matches the language style of the selected character based on the character's personality and knowledge base, and outputs an audio stream with emotional tone through speech synthesis technology to achieve personalized content presentation.
[0013] In addition, the user interaction platform is equipped with a flight route selection function. After the user selects the starting point and the destination on the two-dimensional or three-dimensional map of the scenic area, the system converts the request into a geographic coordinate sequence and submits it to the central dispatch server. The server combines the real-time airspace occupancy, the remaining battery power of the drone and meteorological data, and uses an improved A* algorithm to generate a flight path that meets safety constraints and has the best visual appeal.
[0014] Preferably, the central scheduling server deploys a dynamic resource allocation strategy. When multiple users submit flight requests at the same time, the system calculates the priority weight based on the request timestamp, user membership level, and reward amount, responds to high-weight requests first, and integrates multiple requests with similar paths into a joint flight mission through a task merging mechanism to improve the efficiency of drone use.
[0015] Furthermore, the data fusion and decision-making module connects to the scenic area ticketing system API to obtain real-time data on the number of visitors and length of stay in each area. Combined with the visitor density information identified by AI in the video stream, it generates a visitor heat map that is updated every minute. This heat map serves as an important input parameter for path planning, guiding drones to prioritize covering popular areas to enhance the attractiveness of the live stream.
[0016] In addition, the live streaming and content management module dynamically inserts AI-generated graphic and text information into the video stream overlay, including the current scenic spot name, a brief introduction to the historical background, recommended check-in angles, and real-time interactive Q&A bubbles. All overlay content is filtered through security audit rules to ensure that it does not contain sensitive words or erroneous information.
[0017] Preferably, the security monitoring and emergency response module integrates a real-time face blurring algorithm. Before the video stream is pushed to the public network, the facial areas of non-scenic area staff in the picture are subjected to Gaussian blurring. The blurring intensity is adaptively adjusted according to the clarity of the live broadcast to ensure that personal privacy protection complies with legal requirements.
[0018] On the other hand, an AI-interactive method for controlling live streaming of scenic area drones includes the following steps: S110, activating the standby drone flight unit via an automated helipad network, completing self-checks and environmental assessments, and entering a flightable state; S120, receiving flight route requests or interactive question commands from online users via a user interaction platform, the requests including target geographical coordinates or natural language descriptions of viewing needs; S130, transmitting the received user requests to a central dispatch server, where the server calls a data fusion and decision-making module to generate a comprehensive flight strategy, the strategy including the optimal flight path, estimated flight time, energy consumption estimation, and AI-annotated topics; S140, the central dispatch server sends flight control commands to the designated drone, and the drone follows the planned path. The system performs aerial photography missions and transmits high-definition video streams in real time. In S150, the AI interaction engine simultaneously analyzes user questions, generates semantic answers based on the scenic area's knowledge graph, and outputs an audio stream that plays synchronously with the drone's footage via speech synthesis technology. In S160, the live streaming and content management module processes the original video stream in real time, overlays AI-generated text and image explanations, interactive Q&A feedback, and e-commerce product links, and then pushes it to mainstream live streaming platforms. In S170, the safety monitoring and emergency response module continuously monitors the drone's flight status, communication link quality, and privacy protection implementation, and immediately executes preset emergency procedures when an anomaly is detected. In S180, after the mission is completed, the drone automatically returns to the nearest available helipad, completes landing, charging, and data upload, and the system updates the device status to standby.
[0019] Preferably, in step S120, when a user inputs a natural language request such as "I want to see the sunset at Tengwang Pavilion", the system first extracts the key location "Tengwang Pavilion" through named entity recognition, then calculates the sunset azimuth and time window of the day by combining the current date and geographical location, automatically generates a golden hour flight route with Tengwang Pavilion as the main body and facing southwest, and feeds back the estimated start time to the user.
[0020] Furthermore, the integrated flight strategy generation process in step S130 introduces a multi-objective optimization function. The objective function includes four weighted parameters: path smoothness, landscape coverage, obstacle avoidance safety margin, and energy consumption. The landscape coverage is calculated by comparing the degree of overlap between the flight path and the buffer zone of the scenic spot's landmark attractions to ensure that the flight route has sufficient visual appeal.
[0021] In addition, in step S140, the UAV continuously detects obstacles ahead using the onboard edge computing module during flight. When birds, kites, or suddenly launched objects are detected, the local replanning algorithm is immediately activated to bypass the obstacles while ensuring the continuity of the live broadcast and report the path change information to the central scheduling server in real time.
[0022] Preferably, the AI narration content generation in step S150 adopts a hybrid mode combining template filling and neural generation. For standardized scenic spot information, a structured template is used to automatically fill in fields such as name, era, and architectural features. For open-ended questions such as "What is the legend of this bridge?", a generative pre-trained model is called to output a coherent narrative, and key information is compared with an authoritative database through a fact verification module to ensure accuracy.
[0023] Furthermore, the e-commerce link insertion in step S160 is based on content semantic association. When AI recognizes a specific product (such as cultural and creative ice cream or specialty tea set) in the video, it automatically matches the corresponding SKU from the product database and pops up a limited-time purchase entry below the video. The conversion rate data is used in reverse to optimize the product recommendation model.
[0024] In addition, step S170 sets up a three-level emergency response mechanism: Level 1 is to start local buffering and preset route cruise when the communication delay exceeds 3 seconds; Level 2 is to execute automatic return when the battery is below 20% or the wind speed exceeds the limit; Level 3 is to trigger the emergency landing procedure and release the positioning beacon on the ground when the GPS signal is lost or the key sensor fails.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] By building an automated helipad network and a drone collaborative scheduling mechanism, the scenic area has achieved 24 / 7 uninterrupted live streaming, solving the problem of unsustainable operation caused by the reliance on manual operation in traditional aerial photography. The live streaming duration for a single scenic area has been increased to more than 365 days × 16 hours.
[0027] By employing an AI interaction engine to achieve natural language-driven flight control and intelligent commentary, online users are transformed from passive viewers into active participants in live content, increasing user interaction rate by more than 80%.
[0028] By integrating multimodal sensors and edge computing capabilities, the drone is equipped with autonomous obstacle avoidance and intelligent recognition functions in complex environments, and its flight safety level meets the Civil Aviation Administration's unmanned aerial vehicle operation standards.
[0029] Through real-time AI face blurring and compliance management processes, the system fully protects the public's privacy rights and has passed the personal information protection compliance audit by the Cyberspace Administration of China.
[0030] The system achieves full-stack independent control over both software and hardware, can be deployed to new scenic spots within 3 days, reduces replication costs by 60%, and is ready for large-scale commercial promotion. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the overall technical solution architecture of the scenic area drone live streaming control system based on AI interaction proposed in this invention;
[0032] Figure 2 This is a schematic diagram of the core principle framework of AI-driven multimodal human-machine collaboration and autonomous flight control in this invention. Detailed Implementation
[0033] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0034] Example 1
[0035] Please refer to Figure 1 and Figure 2 This embodiment uses the Tengwang Pavilion Scenic Area in Nanchang as a typical application scenario to construct a routine drone live streaming system based on AI interaction. The aim is to enable online users to drive drones to perform personalized aerial photography tasks via natural language commands under 24 / 7 unattended operation, while simultaneously receiving AI-generated intelligent narration and interactive feedback. The system is deployed in the core area of the scenic area, which has 5G network coverage, stable power access, and compliant airspace approval, covering the main building, riverbank, viewing platform, and nighttime light show area, forming a multi-node collaborative intelligent live streaming service network.
[0036] During the system startup phase, three automated helipads located at the cultural square east of Tengwang Pavilion, the riverside promenade west of it, and the visitor service center south of it entered standby mode. Each helipad adopts a modular steel structure design, with its bottom fixed to a concrete base by pre-embedded bolts, and its top equipped with an electrically openable rainproof cover. Internally, it integrates a drone charging interface, an environmental sensing sensor array, and a wireless communication module. The helipad's built-in environmental sensing sensor array continuously collects local meteorological data: the anemometer uses ultrasonic wind measurement principles, with a sampling frequency of 10Hz, a measurement range of 0–25 m / s, and an accuracy of ±0.3 m / s; the rain sensor is based on a tipping bucket structure, with a resolution of 0.1 mm, used to determine whether the 2 mm / h no-fly threshold has been reached; and the illuminance meter uses a silicon photovoltaic array, with a range of 0–200,000 lux, used to determine day / night mode switching. When any sensor detects that the wind speed exceeds 12 m / s for 5 consecutive seconds or the rainfall intensity is greater than 2 mm / h, the central dispatch server automatically marks the helipad and its associated drones as "no-fly zone", triggers the hatch closing procedure, and sends an early warning message to the operation and maintenance backend.
[0037] The drone's flight unit is a customized hexacopter model with a wheelbase of 800mm and a maximum takeoff weight of 5.2kg. It is equipped with three types of multispectral imaging devices: First, a 4K visible light camera using a 1-inch CMOS sensor, supporting 3840×2160@60fps video recording, equipped with an f / 2.8 variable aperture lens, and optical image stabilization, used for regular daytime high-definition video acquisition; second, an infrared thermal imager, operating in the 8-14μm band, with a resolution of 640×512 and a thermal sensitivity ≤50mK, which converts temperature distribution into visual images through pseudo-color mapping, used to capture differences in building structure heat dissipation and hotspots of crowd gathering in nighttime navigation mode; third, a solid-state LiDAR, with a scanning frequency of 10Hz, a ranging range of 0.1-150m, and an angular resolution of 0.09°, used to construct a 3D point cloud map around the flight path, supporting centimeter-level obstacle detection and distance estimation in complex building environments. All three types of sensors are equipped with high-precision GPS timing modules with a time synchronization error of less than 1ms. All raw data streams are packaged after being aligned with UTC timestamps in the airborne edge computing unit and transmitted in real time to the central dispatch server via UDP protocol through 5G CPE equipment. The transmission bit rate is dynamically adjusted within the range of 20 to 60Mbps to ensure that critical data can still be transmitted back completely even in environments with fluctuating signals.
[0038] The drone's built-in edge computing module utilizes the NVIDIA Jetson AGX Orin platform, boasting a computing power of 275 TOPS. It comes pre-installed with a lightweight YOLOv8n model, with model parameters compressed to 3.0M and inference latency controlled within 15ms. Based on pre-training on ImageNet and COCO datasets, the model is fine-tuned using aerial images of 5A-level scenic spots in Jiangxi Province. The training set contains over 100,000 labeled images, covering iconic buildings such as the Tengwang Pavilion and typical vegetation types like camphor trees, ginkgo, and azaleas. The model's output layer is expanded to include 12 semantic labels, such as "ancient buildings," "modern buildings," "water bodies," "roads," "tourist groups," "trees," "bridges," and "towers." During flight, the edge computing module performs 60 object detections per second. The recognition results are embedded in the video stream metadata in JSON format, containing the object category, bounding box coordinates, confidence score, and geographic coordinates (calculated through drone GPS positioning and camera field-of-view back projection), for subsequent use by the AI interaction engine. For example, when a drone flies over the main building of Tengwang Pavilion, the system automatically identifies and labels it as "ancient building - Ming Dynasty style - hip roof". This label is used as contextual information to input into the AI narration generation process.
[0039] The user interaction platform offers dual entry points: a WeChat mini-program and a web page. The front-end interface integrates a 2D electronic map of the scenic area with a 3D digital twin model rendered using WebGL. Users can click the "Initiate Flight" button on the map and define the starting and ending points by swiping their fingers. The system converts the touch coordinates into a latitude and longitude sequence in the WGS-84 geographic coordinate system, along with a suggested altitude (default 120m, manually adjustable to the 80-180m range). Users can also input natural language requests, such as "I want to see the sunset at Tengwang Pavilion." This request is submitted to the application layer interface of the central scheduling server via an HTTPS encrypted channel. The server first calls the Named Entity Recognition (NER) submodule to extract the key entity "Tengwang Pavilion" from the text based on the BiLSTM-CRF architecture. Then, it calls the temporal semantic parser, combining the current UTC time and the geographic coordinates of Nanchang City (28.68°N, 115.83°E), and uses astronomical algorithms to calculate the sunset azimuth angle as 248° and the duration window as 17:42-18:05. Based on this, the system automatically generates a golden hour flight route that starts from Shengmi Bridge on the opposite bank of the river, ends at the southwest corner of Tengwang Pavilion, and maintains a heading angle of 245°±5°. It then pushes a confirmation pop-up to the user: "A sunset flight route starting at 17:45 has been planned for you, with an estimated flight time of 12 minutes. Do you want to confirm?" After the user confirms, they request to officially enter the task queue.
[0040] The central dispatch server is deployed in the scenic area's private cloud data center, employing a Kubernetes containerized architecture. Its core services include a task manager, a path planning engine, a resource scheduler, and a status monitor. Upon receiving a user's flight request, the task manager encapsulates it into a standardized task object containing the fields: request_id (UUID), user_id, start_point (latitude and longitude), end_point, priority_weight, ai_theme, and timestamp. The resource scheduler then queries the list of currently available drones, filtering based on criteria such as: in "standby" status, battery level above 70%, helipad distance from the starting point less than 1.5km, and not in a no-fly zone. If multiple candidate drones exist, the system calculates the optimal allocation based on a dynamic priority weight formula.
[0041]
[0042] in, The coefficients for user membership levels are: (1.0 for regular users, 1.5 for VIP users, and 2.0 for SVIP users). The reward amount is converted into points (0.01 points for every 1 yuan). The normalized values are ranked in reverse order of task submission time (earlier submissions receive higher scores), with weighting coefficients. =0.4、 =0.35、 =0.25. Scheduler selection The highest-value drone will be selected to perform the mission, and its associated landing pad will be reserved as the return destination.
[0043] After receiving the task command, the path planning engine initiates a multi-objective optimization algorithm to generate a comprehensive flight strategy. Input data includes: a digital elevation model (DEM), a 3D model of the scenic area's buildings, a real-time airspace occupancy map (from ADS-B broadcasts from other drones), weather warning areas, and a visitor heat map. The heat map is updated every minute by the data fusion and decision module, and its generation process is as follows: First, the number of people entering and exiting each ticket gate is obtained from the scenic area's ticketing system API, with a sampling period of 1 minute; second, spatial interpolation is performed using visitor density (unit: people / 100 square meters) identified by YOLOv8 in the drone video stream; finally, a Gaussian kernel function is used to fuse the two types of data to generate a 10m×10m grid heat map. The objective function of path planning is defined as:
[0044]
[0045] in, The smoothness of the path is measured by the inverse integral of the rate of change of the heading angle; To determine the landscape coverage rate, the percentage of overlap between the flight path centerline and iconic landmarks (such as the main building of Tengwang Pavilion, the stele corridor, and the ancient riverbank wharf) was calculated. To estimate energy consumption, estimates are based on flight distance, drag coefficient, and battery discharge curve. To ensure obstacle avoidance safety margin, the number of times the path traverses high-risk areas (such as high-voltage power lines or dense tree canopies) is counted. The weighting coefficient is set to... =0.3、 =0.4、 =0.2、 =0.1, ensuring the route is both visually appealing and operationally safe. The planning results are output in the form of a Waypoint sequence, including the latitude, longitude, altitude, flight speed (default 8m / s, reduced to 5m / s in complex areas), hover time, and camera gimbal attitude angle for each waypoint.
[0046] After receiving flight commands, the UAV executes the entire process from S110 to S180. Before takeoff, a self-check is performed: checking motor phase, IMU zero offset, GPS satellite count (≥8 required), and communication link RSSI (≥-85dBm required). After passing the self-check, the hatch opens, and the UAV ascends vertically to a height of 10m, entering cruise mode. During flight, the onboard edge computing module continuously runs a local replanning algorithm. When the lidar point cloud detects a dynamic obstacle (such as a kite or bird) 80m ahead, the system immediately initiates an emergency obstacle avoidance procedure: First, based on the RRT* algorithm, three alternative detour paths are generated within a 50m radius; second, the landscape impact (angle of deviation from the original route), energy consumption increment, and flight time of each path are evaluated; finally, the path with the least impact is selected for detour, and the path change information is reported to the central dispatch server via the 5G link. The server synchronously updates the airspace occupancy status to prevent conflicts with other UAVs.
[0047] The AI interaction engine responds to user questions synchronously in stage S150. The natural language processing module within the engine adopts a BERT-based architecture, with a vocabulary expanded to 150,000 words, including local terms such as "Preface to the Pavilion of Prince Teng" and "Ganpo Culture." After pre-training, the model is fine-tuned using 200,000 guide-style question-and-answer pairs, supporting the understanding of complex semantics. For example, the question "How old was Wang Bo when he wrote the Preface to the Pavilion of Prince Teng?" accurately parses "person = Wang Bo," "event = writing the Preface to the Pavilion of Prince Teng," and "need = age," retrieving "675 AD, age 26" from the knowledge graph as the answer. For open-ended questions such as "What legends surround this bridge?", the system switches to generative mode, calling a small Transformer-based GPT-2 model to generate narrative text. Before output, the fact-checking module performs keyword matching with authoritative databases such as the *Nanchang Prefecture Gazetteer* and the *Jiangxi General Gazetteer* to ensure the accuracy of the three elements: "person," "time," and "location." The generated text is converted into an audio stream by the speech synthesis module. The TTS engine adopts the FastSpeech square meter 2 architecture and supports seven emotional tones (solemn, cheerful, lyrical, mysterious, passionate, calm, and nostalgic), which are automatically matched according to the question type: historical allusions use the "solemn" tone, and landscape descriptions use the "lyrical" tone. Users can choose a virtual anchor character, such as the "Wang Bo" character, who uses a classical Chinese style to narrate: "This is the key to the Gan River. I once climbed here and composed a poem, where the sunset glowed and the lone wild goose flew together..."
[0048] The live streaming and content management module performs multi-layer overlay processing on the original video stream in stage S160. The original H.265 encoded video stream enters the GPU-accelerated processing pipeline, where it sequentially performs: decoding, image enhancement (contrast improvement, dehazing algorithm), AI face blurring, image and text overlay, and re-encoding. The face blurring module uses a real-time convolutional neural network with a U-Net variant structure, an input resolution of 1920×1080, and a processing frame rate of 60fps. It applies an adaptive Gaussian kernel to the detected face regions, with the kernel size dynamically adjusted according to the image clarity (σ=8 at 1080p and σ=6 at 720p) to ensure that the individual's identity cannot be identified after blurring. The image and text overlay consists of four areas: the top left corner displays the current attraction name and the AI anchor's avatar; the top right corner displays a floating real-time interactive Q&A bubble (e.g., "Netizen 'Ganjiang Fisherman' asked: What's the temperature now?" "AI answered: The current perceived temperature is 22℃, with a light breeze"); the bottom banner scrolls to display recommended photo angles (e.g., "Best shooting angle: 30° upward, focus on the main pavilion"); all overlaid content is filtered through a security review rule engine, with a rule base containing 1327 sensitive word regular expressions to ensure output compliance.
[0049] The safety monitoring and emergency response module implements a three-level response mechanism. Level 1 response addresses communication delays: if the server does not receive a heartbeat packet from the drone for 3 consecutive seconds, it determines the link is unstable. The drone immediately uses its local cached flight path, hovers over its current location, and switches to the 4G backup network for reconnection. Level 2 response addresses energy and weather anomalies: if the battery level drops below 20% or the helipad reports excessive wind speed, the drone terminates its current mission, initiates an automatic return-to-home procedure, and returns to the nearest available helipad along the shortest safe path, maintaining video streaming during the return journey. Level 3 response addresses major malfunctions: if the GPS signal is lost for more than 10 seconds or IMU data is abnormal, the system determines navigation failure, immediately performs an emergency landing, selects a flat area (slope <5° determined by LiDAR scanning), slowly descends to 1m above the ground, releases the parachute buffer, activates the onboard buzzer and LED strobe lights, and simultaneously sends a positioning beacon (containing latitude, longitude, fault code, and timestamp) to the server for rapid search by ground personnel.
[0050] After the mission is completed, the drone executes the S180 procedure: it lands precisely at the target helipad charging port with an error controlled within ±3cm; after the hatch closes, it begins charging with a charging current of 5A and a voltage of 22.8V, taking approximately 45 minutes to fully charge; at the same time, the original log files in the onboard storage (including flight trajectory, sensor data, and AI recognition records) are uploaded to the cloud archiving server via gigabit Ethernet. After the upload is completed, the device status is updated to "standby," awaiting the next mission scheduling.
[0051] Example 2
[0052] This embodiment focuses on the technical implementation of multi-drone collaborative live streaming and AI-powered deep content generation, using a famous area in Nanchang as the application field to address the needs of wide-area coverage and themed content production in large scenic areas. Compared with Embodiment 1, this embodiment introduces a "task merging mechanism" and a "multimodal content generation pipeline" in its system architecture, and strengthens the collaborative optimization capabilities of steps S130 and S160 in its methodology, forming a differentiated technical path.
[0053] When the central dispatch server receives flight requests from multiple users within the same time period, the system initiates the task merging algorithm. Let... There are always N pending requests, each containing a starting point Pi, an ending point Qi, a submission time Ti, and a priority weight Wi. The system first calculates the path similarity between any two requests. Defined as:
[0054]
[0055] in, For the request The initial planned path point set, For the request The initial planned path point set, This represents the path length (Euclidean distance). If... >0.6 and If the time interval is less than a second, the two tasks are determined to be merged. A joint flight path is generated after merging, covering all starting and ending points. The Traveling Salesman Problem (TSP) solver is used to optimize the access order, with the goal of minimizing the total flight distance. The merged tasks are executed by the same drone, which sequentially performs the viewing actions specified by each user (such as hovering, circling, and diving) during the flight, and responds to user questions through a polling AI interaction engine. For example, if user A requests to "see the old site of a certain area" and user B requests to "photograph the Huangyangjie outpost," the system plans a connecting route, first flying to the certain area, stopping for 2 minutes for user A to interact, and then heading to Huangyangjie to perform the aerial photography action. During this time, the AI anchor responds to the questions of both users, achieving efficient resource reuse.
[0056] At the content generation level, this embodiment constructs a multimodal AI generation pipeline. The live streaming module no longer relies solely on static templates but instead builds a dynamic content graph. Whenever a drone identifies a specific scene (such as a "poetry wall"), the system triggers a content generation workflow: First, it extracts the entity "famous poems" from the scenic area's knowledge graph; second, it calls a text generation model to generate explanatory text; third, it activates an image generation model (such as Stable Diffusion Lightweight Edition) to generate a stylized illustration based on the description of the "poetry scene"; finally, it overlays the illustration as a floating layer in the corner of the live stream screen, fading out after 15 seconds. This process forms a complete chain of "visual recognition → knowledge retrieval → text generation → image generation → multimodal output," significantly improving content richness.
[0057] Furthermore, this embodiment optimizes the e-commerce recommendation logic in step S160. The system not only matches products based on screen content but also incorporates user behavior sequence analysis. Let user u exhibit a behavior sequence during the live stream. ,in The user's interest is evaluated as follows: {view product A, ask a question about product B, tip the streamer C, click link D}. The system uses an LSTM network to model the evolution of user interests and outputs the product category the user is most likely to be interested in at the next moment. For example, if a user asks three consecutive questions about the "origins of poetry," the system predicts an 80% increase in interest in "poetry-related" categories. Therefore, it prioritizes pushing related products such as "poetry bookmarks" and "poetry fans" in subsequent screens, achieving precise marketing.
[0058] Example 3
[0059] This embodiment focuses on the enhanced design of system stability and privacy protection in the high-humidity environment during the rainy season. Taking the ancient village of Huangling in Wuyuan as the application object, it mainly solves the problems of equipment failures caused by humid climate and privacy compliance in complex crowd scenarios. Compared with the previous two embodiments, this embodiment makes substantial improvements in the hardware structure and security mechanism.
[0060] In this embodiment, an environmental control subsystem is added to the automated helipad. A temperature and humidity sensor (accuracy ±2%RH, ±0.5°C) is installed inside the cabin. When the relative humidity continuously exceeds 85% for 10 minutes, the dehumidification module is activated: using semiconductor condensation dehumidification technology, with a rated power of 40W and a maximum dehumidification capacity of 1.2L / day, the humidity inside the cabin is controlled below 60%. At the same time, the charging interface is equipped with self-cleaning metal contacts, surface gold-plated treatment, to prevent poor contact caused by oxidation. Before each landing, the drone performs a "dry flight": hovering over the helipad for 30 seconds, using the downwash airflow of the propeller to dry the water droplets on the fuselage surface, especially the camera lens and lidar window, to ensure that the sensor performance is not affected during the next takeoff.
[0061] In terms of privacy protection, this embodiment proposes a hierarchical face blurring strategy. The system classifies the people in the live broadcast画面 according to their activity status: for static people (such as tourists taking pictures), standard Gaussian blurring (σ = 8) is used; for fast-moving individuals (such as running children), since motion blur already exists, the blurring intensity is reduced to σ = 5; the RFID work cards worn by scenic area staff can be recognized by the drone UHF reader, and their facial areas are automatically exempted from blurring. This mechanism ensures privacy while retaining necessary human dynamic information and enhancing the观赏性 of the画面.
[0062] In this embodiment, meteorological prediction data is introduced into the data fusion and decision-making module. The system accesses the API of the Meteorological Bureau to obtain the precipitation probability, wind speed trend, and lightning warning for the next 2 hours. When the precipitation probability is predicted to rise above 70% within the next 30 minutes, the drone is scheduled to return to the nearest helipad待命 in advance to avoid sudden rain during flight. When planning the route, known waterlogging-prone areas (such as low-lying ridges and stone bridges) are actively avoided to ensure flight safety.
[0063] The above is only a preferred embodiment of the present invention, and it is not intended to limit the present invention in any form. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content to make equivalent embodiments with equivalent changes, but as long as it does not depart from the technical content of the present invention, any brief modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A scenic area drone live streaming control system based on AI interaction, characterized in that, Includes the following parts: The drone flight unit is used to perform aerial photography and mobile live streaming missions, and transmits high-definition video streams back in real time via 5G network; An automated helipad network, distributed at key nodes in the scenic area, provides drones with automatic take-off and landing, charging, and environmental self-checking services, supporting continuous operation. The AI interaction engine integrates natural language processing, computer vision, and speech synthesis modules to parse user interaction commands, generate intelligent narration content, and drive virtual anchors to respond in real time. The user interaction platform provides mobile and web interfaces, supporting online users to send flight route requests, ask questions, interact, give tips, and participate in live e-commerce activities. The central dispatch server is responsible for receiving user requests, planning flight paths, coordinating drone resource allocation, dispatching AI services, and managing live stream distribution. The data fusion and decision-making module is used to integrate scenic area geographic information, tourist distribution heat maps, meteorological data and user behavior data to generate dynamic flight strategies and content recommendation schemes. The live streaming and content management module is responsible for real-time processing of video streams, overlaying AI interaction engine to generate content, and pushing it to third-party live streaming platforms. The safety monitoring and emergency response module monitors the drone's status, airspace compliance, and privacy protection policy implementation in real time, and triggers an automatic return-to-home or emergency stop mechanism in case of anomalies. The central scheduling server deploys a dynamic resource allocation strategy. When multiple users submit flight requests simultaneously, the system calculates priority weights based on the request timestamp, user membership level, and reward amount, prioritizing requests with higher weights. It also integrates multiple requests with similar paths into a single joint flight mission through a task merging mechanism. A multi-objective optimization function, incorporating weights for path smoothness, landscape coverage, obstacle avoidance safety margin, and energy consumption, is used to generate the flight path. When the central scheduling server receives multiple flight requests from users within the same time period, the system initiates the task merging algorithm. There are always N pending requests. Each request includes a starting point Pi, an ending point Qi, a submission time Ti, and a priority weight Wi. The system first calculates the path similarity between any two requests. Defined as: in, For the request The initial planned path point set, For the request The initial planned path point set, if >0.6 and If the time is less than 1 second, it is determined that the two tasks can be merged. After merging, a joint flight path is generated, covering all starting points and ending points. The visiting order is optimized using a traveling salesman problem solver. The goal is to minimize the total flight distance. The merged tasks are executed by the same drone. During the flight, the drone sequentially completes the viewing actions specified by each user and responds to each user's questions through the AI interaction engine. The AI interaction engine includes a natural language processing module, which is used to parse natural language requests containing temporal or spatial semantics, and combine geographic location information with astronomical algorithms to calculate the target azimuth angle, thereby driving the drone to generate a specific sightseeing route. After receiving the mission instructions, the path planning engine starts a multi-objective optimization algorithm to generate a comprehensive flight strategy. The objective function of path planning is defined as follows: in, The smoothness of the path is measured by the inverse integral of the rate of change of the heading angle; To determine the landscape coverage rate, calculate the percentage of overlap between the flight path centerline and landmark attractions within a 50m buffer zone. To estimate energy consumption, estimates are based on flight distance, drag coefficient, and battery discharge curve. To ensure obstacle avoidance safety margin, the number of times the statistical path traverses high-risk areas is calculated, with a weighting coefficient set to [value missing]. =0.3、 =0.4、 =0.2、 =0.1, ensuring that the route is both visually appealing and safe to operate. The planning results are output in the form of a Waypoint sequence, including the latitude, longitude, altitude, flight speed, hovering time and camera gimbal attitude angle for each waypoint.
2. The system according to claim 1, characterized in that, The UAV flight unit is equipped with multispectral imaging equipment, including a 4K visible light camera, an infrared thermal imager, and a lidar. Data from the three types of sensors are synchronized to the central scheduling server after being timestamped. The infrared thermal imager is used to capture temperature distribution features in night mode to enhance visual performance, while the lidar is used to construct local 3D point cloud maps to support accurate obstacle avoidance and path optimization in complex terrain.
3. The system according to claim 2, characterized in that, The drone flight unit has a built-in edge computing module that runs a lightweight YOLOv8 target detection model. This model is used to identify landmark buildings, vegetation types, and tourist gathering areas in the scenic area in real time during flight. The identification results are then embedded as semantic tags into the video stream metadata, which is then called by the AI interaction engine to generate narration content with geographical and cultural context.
4. The system according to claim 3, characterized in that, The automated helipad network adopts a modular design, with each helipad equipped with an environmental sensing sensor group, including anemometers, rain sensors, and light intensity meters, for real-time monitoring of local meteorological conditions. When the wind speed exceeds the threshold or the rainfall intensity is greater than the preset value, it automatically sends a no-fly warning to the central dispatch server and initiates the canopy closing and equipment waterproofing procedures.
5. The system according to claim 4, characterized in that, The natural language processing module in the AI interaction engine adopts a domain adaptation model based on the BERT architecture. This domain adaptation model is fine-tuned using scenic spot tour guide scripts, historical documents, and common tourist questions on the basis of pre-training, so that it has the ability to accurately understand local cultural terms, dialect expressions, and historical allusions, and supports the accurate parsing of complex requests containing specific location and time semantics.
6. The system according to claim 5, characterized in that, The AI interaction engine is equipped with a multi-role virtual anchor library, including historical figures, local cultural ambassadors, and cartoon IP characters. Users select the anchor style through the interactive platform, and the AI system generates narration that matches the selected character's personality and knowledge base, and outputs an audio stream with emotional tone through speech synthesis technology.
7. The system according to claim 6, characterized in that, The user interaction platform is equipped with a flight route selection function. After the user selects the starting point and the destination on the two-dimensional or three-dimensional map of the scenic area, the system converts the selected starting point and destination into a geographic coordinate sequence and submits it to the central dispatch server. The server combines real-time airspace occupancy, the remaining battery power of the drone, and meteorological data, and uses an improved A* algorithm to generate a flight path that meets safety constraints and has the best visual appeal.
8. The system according to claim 1, characterized in that, The data fusion and decision-making module connects to the scenic area ticketing system API to obtain real-time data on the number of visitors and length of stay in each area. Combined with visitor density information identified by AI in the video stream, it generates a visitor heat map that is updated every minute. This heat map serves as an important input parameter for path planning, guiding drones to prioritize coverage of popular areas to enhance the attractiveness of the live stream.
9. A control method applied to the AI-interactive scenic area drone live streaming control system as described in any one of claims 1-8, characterized in that, Includes the following steps: The drone flight unit, which is in standby status, is activated through the automated helipad network and enters flight status after completing self-check and environmental assessment. The system receives flight route requests or interactive question commands from online users via a user interaction platform. The flight route requests include the target geographic location coordinates or natural language descriptions of the sightseeing requirements. The received user request is transmitted to the central dispatch server, which then calls the data fusion and decision-making module to generate a comprehensive flight strategy. The comprehensive flight strategy includes the optimal flight path, the estimated flight duration, the energy consumption estimate, and the AI explanation topic. The central dispatch server sends flight control commands to the designated drones, which then perform aerial photography missions along the planned paths and transmit high-definition video streams back in real time. The AI interaction engine simultaneously analyzes the user's questions, generates semantic answers by combining the scenic area's knowledge graph, and outputs an audio stream that is played simultaneously with the drone footage through speech synthesis technology. The live streaming and content management module processes the original video stream in real time, overlays AI-generated text and image explanations, interactive Q&A feedback, and e-commerce product links, and then pushes it to mainstream live streaming platforms. The safety monitoring and emergency response module continuously monitors the drone's flight status, communication link quality, and privacy protection implementation, and immediately executes preset emergency procedures when an anomaly is detected. After the mission is completed, the drone automatically returns to the nearest available helipad, completes landing, charging, and data upload, and the system updates the equipment status to standby.
Citation Information
Patent Citations
Unmanned aerial vehicle control method and system for live broadcast interaction, and aircraft
CN119165877A
ROS-based four-rotor tour guide unmanned aerial vehicle system and tour guide auxiliary method thereof
CN120909320A