Scenic spot unmanned aerial vehicle live broadcast control system and method based on AI interaction
By constructing a system architecture that integrates autonomous flight control of drones and AI-powered intelligent content generation, the problems of lack of interactivity and high operating costs in scenic area drone live streaming systems have been solved, enabling continuous and highly interactive live streaming services and improving user experience and business revenue.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-04
- Publication Date
- 2026-03-10
AI Technical Summary
Existing drone live streaming systems for scenic areas suffer from problems such as fixed live streaming angles, lack of interactive capabilities, high operating costs, and discontinuous content production, making it difficult to meet users' needs for an immersive, dynamic, and participatory real-time viewing experience.
A system architecture integrating UAV autonomous flight control, AI intelligent content generation, multimodal human-machine interaction, and cloud-based collaborative management is constructed, including UAV flight units, automated helipad networks, AI interaction engines, user interaction platforms, central dispatch servers, data fusion and decision-making modules, and safety monitoring and emergency response modules, to achieve autonomous take-off and landing, real-time interaction, obstacle avoidance in complex environments, and mission replanning for UAVs.
It has enabled 24/7 uninterrupted live streaming of the scenic area, which has increased user interaction, reduced operating costs, enhanced system security and privacy protection capabilities, opened up new channels for non-ticket revenue, and is ready for large-scale commercial promotion.
Smart Images

Figure CN121644839A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of artificial intelligence and unmanned aerial vehicle control technology, and particularly relates to a scenic spot unmanned aerial vehicle live broadcast control system and method based on AI interaction. BACKGROUND
[0002] With the deep integration of artificial intelligence and unmanned system technology, the application of intelligent unmanned equipment in the cultural tourism field is gradually evolving from single-function demonstration to normal and platform service mode. As an intelligent carrier with air mobility, unmanned aerial vehicles have shown significant advantages in aerial photography, inspection, emergency rescue and other scenarios. Their flexible deployment, wide coverage and real-time transmission capabilities provide a new technical path for scenic spot content production and interactive experience.
[0003] Under this background, the scenic spot live broadcast system based on unmanned aerial vehicles has gradually become an important part of smart tourism construction. It presents natural scenery and cultural landscape from a high-altitude perspective, breaks through the limitations of traditional fixed camera views, and improves the immersion and participation of online users.
[0004] Among them, the unmanned aerial vehicle live broadcast system for scenic spots places particular emphasis on the automation of the system and the level of human-computer interaction. Its core goal is to achieve autonomous take-off and landing, flight planning, video acquisition and real-time streaming under unmanned conditions, and to enhance the two-way interaction between the audience and the equipment through intelligent means. Online users can not only "see", but also "participate" and "influence" the generation process of live broadcast content, thus building a truly dynamic, interactive and sustainable digital tourism content ecosystem.
[0005] The existing technology still faces multiple technical bottlenecks in implementing the scenic spot unmanned aerial vehicle live broadcast function. First, most systems rely on manual remote control operations, making it difficult to achieve all-weather and all-time automated execution of flight tasks, resulting in low live broadcast frequency, high labor costs and unsustainable operation. Second, the interactive mode is limited to one-way video stream output, lacking real-time semantic understanding and feedback mechanisms driven by AI, and unable to respond to audience voice or text instructions for flight path adjustment or view switching, with serious lack of interactivity. Third, the flight control and AI decision-making systems are separate from each other, without forming a unified intelligent closed loop. AI is only used for post-content labeling or voice broadcasting, without deep involvement in flight process regulation. Finally, the system has poor adaptability to complex scenic environments. In the face of real-world constraints such as crowded areas, electromagnetic interference and weather changes, it lacks dynamic obstacle avoidance and task re-planning capabilities based on multi-modal perception, making it difficult to ensure safety and stability. These problems are particularly prominent in tourism live broadcast scenarios that require high concurrency, strong interaction and long-term operation, and a new unmanned aerial vehicle live broadcast control system that deeply integrates machine intelligence, human-computer interaction and advanced process control is urgently needed to solve them. SUMMARY
[0006] The present application aims to make up for the deficiencies of the prior art, and provides a scenic spot unmanned aerial vehicle live broadcast control system and method based on AI interaction, which can effectively solve the problems in the background art. The existing scenic spot live broadcast system generally has the problems of fixed live broadcast angle, lack of interaction ability, high operation cost, discontinuous content output, etc., and it is difficult to meet the user's demand for immersive, dynamic and participatory real-time viewing experience. The present application builds a system architecture integrating unmanned aerial vehicle autonomous flight control, AI intelligent content generation, multi-modal human-computer interaction and cloud collaborative management, and realizes the fundamental change of scenic spot live broadcast from "passive watching" to "active participation", from "temporary activity" to "normal service", and from "single video" to "intelligent interactive content platform".
[0007] To achieve the above-mentioned purpose, the present application provides the following technical scheme: on the one hand, a scenic spot unmanned aerial vehicle live broadcast control system based on AI interaction, the system comprises the following components: unmanned aerial vehicle flight unit, used for executing aerial shooting and mobile live broadcast tasks, and transmitting high-definition video stream in real time through 5G network; automatic landing apron network, distributed in key nodes of the scenic spot, providing automatic take-off and landing, charging and environment self-checking services for the unmanned aerial vehicle, supporting 7*24 hours continuous operation; AI interaction engine, integrating natural language processing, computer vision and speech synthesis modules, used for analyzing user interaction instructions, generating intelligent commentary content and driving virtual anchors to respond in real time; user interaction platform, providing mobile terminal and Web terminal interface, supporting online users to send flight route requests, ask questions, reward and participate in live broadcast e-commerce activities; central dispatch server, responsible for receiving user requests, planning flight paths, coordinating unmanned aerial vehicle resource allocation, dispatching AI services and managing live broadcast stream distribution; data fusion and decision module, used for integrating scenic spot geographic information, tourist distribution heat map, meteorological data and user behavior data, generating dynamic flight strategy and content recommendation scheme; live broadcast push stream and content management module, responsible for real-time processing of original video stream, superimposing AI generated content and pushing to third-party live broadcast platform; safety monitoring and emergency response module, real-time monitoring of unmanned aerial vehicle state, airspace compliance and privacy protection strategy execution, and triggering automatic return or emergency stop mechanism in abnormal situation; Preferably, the unmanned aerial vehicle flight unit is equipped with multi-spectral imaging equipment, including a 4K visible light camera, an infrared thermal imager and a laser radar, wherein the visible light camera is used for regular high-definition video acquisition, the infrared thermal imager is used for capturing temperature distribution characteristics in night tour mode to enhance visual expressiveness, and the laser radar is used for constructing a local three-dimensional point cloud map to support precise obstacle avoidance and path optimization in complex terrain, and the data of the three types of sensors are transmitted to the central dispatch server synchronously after time stamp alignment; Further, the unmanned aerial vehicle flight unit is built-in with an edge computing module, which runs a lightweight YOLOv8 target detection model to identify scenic landmark buildings, vegetation types and tourist gathering areas in real time during flight, and embeds the identification results as semantic tags into video stream metadata for the AI interaction engine to call to generate commentary content with geographical and cultural context. In addition, the automated apron network adopts a modular design, and each apron is equipped with an environmental perception sensor group, including an anemometer, a rain sensor and an illumination intensity meter, for real-time monitoring of local weather conditions. When the wind speed exceeds 12 m / s or the rainfall intensity is greater than 2 mm / h, an automatic no-fly warning is sent to the central dispatch server, and the hatch closing and equipment waterproofing program is started. Preferably, the natural language processing module in the AI interaction engine uses a domain adaptation model based on the BERT architecture, which is fine-tuned using the tour guide words, historical documents and tourist common question corpus of the main 5A-level scenic spots in Jiangxi Province based on pre-training, so that it has accurate understanding ability for local cultural terms, dialect expressions and historical anecdotes, and supports accurate analysis of complex semantics such as "How to take the best sunset at the Pavilion of Prince Teng" or "What's the story of a certain area" input by the user. Further, the AI interaction engine is configured with a multi-role virtual anchor library, including historical figures, local cultural spokespersons and cartoon IP images. Users can select the anchor style through the interactive platform, and the AI system generates commentary words in the language style of the selected role according to its personality setting and knowledge base, and outputs audio streams with emotional intonation through voice synthesis technology, realizing personalized content presentation. In addition, the user interaction platform is provided with a flight route point selection function. After the user selects the starting point and ending point on the two-dimensional or three-dimensional map of the scenic spot, the system converts the request into a geographic coordinate sequence and submits it to the central dispatch server. The server combines real-time airspace occupation, remaining power of the unmanned aerial vehicle and weather data to generate a flight path that meets safety constraints and has the best visual observation. Preferably, the central dispatch server deploys a dynamic resource allocation strategy. When multiple users submit flight requests at the same time, the system calculates the priority weight according to the request timestamp, user membership level and reward amount, responds to high-weight requests first, and integrates multiple requests with similar paths into a joint flight task through a task merging mechanism to improve the efficiency of the unmanned aerial vehicle. Further, the data fusion and decision module accesses the scenic spot ticket system API to obtain real-time data of the number of visitors and stay time in each area, and generates a minute-level updated visitor heat distribution map based on the AI-identified visitor density information in the video stream. The heat map is an important input parameter for path planning, guiding the unmanned aerial vehicle to preferentially cover high-popularity areas to improve the attraction of live streaming. In addition, the live streaming and content management module dynamically inserts AI-generated text information in the video stream overlay layer, including the current scenic spot name, historical background introduction, recommended check-in angle, and real-time interactive Q&A bubbles. All overlay content is filtered through security audit rules to ensure that it does not contain sensitive words or incorrect information. Preferably, the security monitoring and emergency response module integrates a real-time face blurring algorithm. Before the video stream is pushed to the public network, the face area of non-scenic staff in the picture is processed with Gaussian blur. The blur intensity is adjusted adaptively according to the live streaming clarity to ensure that personal privacy protection meets legal requirements. On the other hand, an AI interaction-based scenic area unmanned aerial vehicle live broadcast control method, the specific steps of which are as follows: S110, start the unmanned aerial vehicle flight unit in standby state through the automatic landing pad network, complete self-checking and environment evaluation, and enter the flyable state; S120, receive the flight route request or interactive question instruction of the online user via the user interaction platform, the request including the target geographic location coordinates or the natural language description of the viewing demand; S130, transmit the received user request to the central dispatch server, and generate a comprehensive flight strategy by calling the data fusion and decision module, the strategy including the optimal flight path, the estimated flight time, the energy consumption estimation, and the AI commentary theme; S140, the central dispatch server sends the flight control instruction to the specified unmanned aerial vehicle, the unmanned aerial vehicle executes the aerial photography task along the planned path and returns the high-definition video stream in real time; S150, the AI interaction engine synchronously analyzes the user question content, generates a semantic answer combined with the scenic area knowledge graph, and outputs an audio stream through voice synthesis technology for synchronous playback with the unmanned aerial vehicle shooting picture; S160, the live streaming and content management module processes the original video stream in real time, superimposes the AI-generated text commentary, interactive Q&A feedback, and e-commerce commodity link information, and pushes it to the mainstream live broadcast platform; S170, the security monitoring and emergency response module continuously monitors the unmanned aerial vehicle flight state, communication link quality, and privacy protection execution, and immediately executes the preset emergency program when an abnormality is found; S180, after the task is completed, the unmanned aerial vehicle automatically returns to the nearest idle landing pad, completes landing, charging, and data uploading, and the system updates the device state to standby; Preferably, in the S120 step, when the user inputs a natural language request such as "I want to see the sunset at the Pavilion of Prince Teng", the system first extracts the key point "Pavilion of Prince Teng" through named entity recognition, then calculates the sunset azimuth and time window based on the current date and geographic location, and automatically generates a flight route with the Pavilion of Prince Teng as the main body and facing the southwest direction during the golden time period, and feeds back the estimated start time to the user. Further, the comprehensive flight strategy generation process in the S130 step introduces a multi-objective optimization function, the objective function includes path smoothness, landscape coverage, obstacle avoidance safety margin and energy consumption four weight parameters, wherein the landscape coverage is calculated by comparing the overlap degree of the flight path and the buffer area of the landmark scenic spot, to ensure that the flight route has enough visual appeal; In addition, in the S140 step, the unmanned aerial vehicle continuously detects the front obstacles during flight using the on-board edge computing module, when birds, kites or suddenly rising objects are identified, the local re-planning algorithm is started immediately, the obstacles are bypassed under the premise of ensuring the continuity of live broadcast and the path change information is reported to the central dispatch server in real time; Preferably, the AI commentary content generation in the S150 step adopts a hybrid mode combining template filling and neural generation, for standardized scenic spot information, structured templates are used to automatically fill in fields such as name, age, architectural features, for open-ended questions such as "what is the legend of this bridge", a generative pre-training model is called to output coherent narratives, and the key information is compared with authoritative databases through a fact checking module to ensure accuracy; Further, the e-commerce link insertion in the S160 step is based on content semantic association, when the AI identifies specific goods (such as cultural and creative ice cream, special tea set) appearing in the picture, the corresponding SKU is automatically matched from the goods database, and a time-limited purchase portal is popped up below the video, the conversion rate data is used to optimize the goods recommendation model in reverse; In addition, the S170 step sets up a three-level emergency response mechanism: level one is to start local caching and preset route cruising when communication delay exceeds 3 seconds; level two is to execute automatic return when power is less than 20% or wind speed exceeds the standard; level three is to trigger the emergency landing program and release the positioning beacon on the ground when the GPS signal is lost or the key sensor fails; Compared with the prior art, the present application has the following beneficial effects: By constructing an automated airport network and a UAV cooperative scheduling mechanism, uninterrupted live broadcast in scenic spots 7x24 hours is realized, solving the problem of unsustainable operation caused by traditional aerial photography relying on manual operation, and the live broadcast time in a single scenic spot is increased to more than 365 daysx16 hours; The AI interaction engine is used to realize natural language driven flight control and intelligent commentary, so that online users change from passive viewers to active participants in live broadcast content, and the user interaction rate is increased by more than 80%; The integration of multi-modal sensors and edge computing capabilities enables the UAV to have autonomous obstacle avoidance and intelligent identification functions in complex environments, and the flight safety level reaches the Civil Aviation Administration Unmanned Aerial Vehicle Operation Standard; A diversified revenue model of "live broadcast + e-commerce + tipping" is constructed, a new channel for non-ticket revenue is opened up for scenic spots, and the monthly online revenue of the pilot scenic spot increases by 400,000 yuan; Through real-time AI-powered facial blurring and compliance management processes, the system comprehensively protects the public's privacy rights and has passed the personal information protection compliance audit by the Cyberspace Administration of China. The system achieves full-stack independent control over both software and hardware, can be deployed to new scenic spots within 3 days, reduces replication costs by 60%, and is ready for large-scale commercial promotion. Attached Figure Description
[0008] Figure 1 This is a schematic diagram of the overall technical architecture of the AI-interactive scenic area drone live streaming control system proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of AI-driven multimodal human-machine collaboration and autonomous flight control in this invention. Detailed Implementation
[0009] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structure, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0010] Example 1 Please refer to Figure 1 and Figure 2 This embodiment uses the Tengwang Pavilion Scenic Area in Nanchang as a typical application scenario to construct a routine drone live streaming system based on AI interaction. The aim is to enable online users to drive drones to perform personalized aerial photography tasks via natural language commands under 24 / 7 unattended operation, while simultaneously receiving AI-generated intelligent narration and interactive feedback. The system is deployed in the core area of the scenic area, which has 5G network coverage, stable power access, and compliant airspace approval, covering the main building, riverbank, viewing platform, and nighttime light show area, forming a multi-node collaborative intelligent live streaming service network.
[0011] During the system startup phase, the three automated landing pads located in the cultural square east of the Pavilion of Prince Teng, the riverside promenade west of the pavilion, and the tourist service center south of the pavilion enter standby state. Each landing pad is designed with modular steel structure, fixed to the concrete base by pre-buried bolts at the bottom, and equipped with a rainproof cabin cover that can be opened and closed electrically at the top. Inside the landing pad, there is an integrated group of unmanned aerial vehicle charging interfaces, environmental perception sensors, and wireless communication modules. The environmental perception sensor group inside the landing pad continuously collects local meteorological data: the anemometer uses ultrasonic wind measurement principle, with a sampling frequency of 10 Hz, a measurement range of 0~25 m / s, and an accuracy of ±0.3 m / s; the rain sensor is based on a tipping bucket structure, with a resolution of 0.1 mm, used to determine whether the no-fly threshold of 2 mm / h is reached; the light intensity meter uses a silicon photocell array, with a range of 0~200,000 lux, used to determine the day-night mode switching. When any sensor detects that the wind speed exceeds 12 m / s for more than 5 seconds or the rainfall intensity is greater than 2 mm / h, the central dispatching server automatically marks the landing pad and its associated unmanned aerial vehicle as "no-fly state", triggers the cabin cover closing program, and sends warning information to the operation and maintenance background.
[0012] The unmanned aerial vehicle flight unit uses a customized hexacopter model, with a body axis distance of 800 mm and a maximum take-off weight of 5.2 kg. It carries three types of multi-spectral imaging equipment: the first is a 4K visible light camera with a 1-inch CMOS sensor, supporting 3840×2160@60fps video recording, equipped with a f / 2.8 variable aperture lens, and with optical image stabilization function, used for regular daytime high-definition video acquisition; the second is an infrared thermal imager with a working waveband of 8~14 μm and a resolution of 640×512, with a thermal sensitivity of ≤50 mK, which converts temperature distribution into visual images through pseudo-color mapping, used for capturing building structure heat dissipation differences and crowd gathering hotspots in night tour mode; the third is a solid-state laser radar with a scanning frequency of 10 Hz, a ranging range of 0.1~150 m, and an angular resolution of 0.09°, used for constructing a three-dimensional point cloud map of the surrounding flight path, supporting obstacle detection and distance estimation with centimeter-level precision in complex buildings. All three types of sensors are equipped with high-precision GPS timing modules, with a time synchronization error of less than 1 ms. After all raw data streams are aligned by UTC timestamp in the on-board edge computing unit, they are packaged and transmitted in real time to the central dispatching server through a 5G CPE device using UDP protocol, with a dynamic transmission code rate range of 20~60 Mbps, ensuring that key data can still be completely returned in a fluctuating signal environment.
[0013] The UAV built-in edge computing module adopts NVIDIA Jetson AGX Orin platform, with a computing power of 275 TOPS, and pre-installs a lightweight YOLOv8n model with a model parameter compression of 3.0M and an inference delay controlled within 15ms. Based on the pre-training of the ImageNet and COCO datasets, the model is fine-tuned using real aerial images of 5A-level scenic spots in Jiangxi Province. The training set contains more than 100,000 labeled images, covering landmark buildings such as the Pavilion of Prince Teng, as well as typical vegetation species such as camphor trees, ginkgo trees, and azaleas. The model output layer is expanded to include 12 semantic labels, including "ancient architecture", "modern architecture", "water body", "road", "tourist group", "tree", "bridge", and "tower". During flight, the edge computing module performs 60 target detections per second, and the recognition results are embedded in the video stream metadata packet in JSON format, including target category, bounding box coordinates, confidence score, and geographic coordinates (calculated by UAV GPS positioning and camera field of view angle back projection), for subsequent AI interaction engine call. For example, when the UAV flies over the main building of the Pavilion of Prince Teng, the system automatically identifies and labels "ancient architecture - Ming Dynasty style - hip and gable roof", which is input as context information into the AI commentary generation process.
[0014] The user interaction platform provides a WeChat mini-program and a Web page dual entrance, and the front-end interface integrates a two-dimensional electronic map of the scenic spot and a WebGL rendered three-dimensional digital twin model. Users can click the "Launch Flight" button on the map, and draw the starting and ending points by sliding their fingers. The system converts the touch coordinates into latitude and longitude sequences in the WGS-84 geographic coordinate system, and adds an altitude suggestion value (default 120m, can be manually adjusted to 80-180m interval). Users can also input natural language requests, such as "I want to see the sunset at the Pavilion of Prince Teng". The request is submitted to the application layer interface of the central dispatch server through the HTTPS encrypted channel. The server first calls the Named Entity Recognition (NER) submodule to extract the key entity "Pavilion of Prince Teng" from the text based on the BiLSTM-CRF architecture; then calls the time semantic parser, combines the current UTC time and the geographic coordinates of Nanchang City (North Latitude 28.68°, East Longitude 115.83°), and calculates the sunset azimuth as 248° and the duration window as 17:42-18:05 through astronomical algorithms. The system automatically generates a flight route starting from the Shengmi Bridge across the river and ending at the southwest corner of the Pavilion of Prince Teng, with a heading angle of 245°±5°, and pushes a confirmation popup to the user: "A sunset flight route starting at 17:45 has been planned for you, with an estimated flight time of 12 minutes, do you confirm?" After the user confirms, the request enters the task queue.
[0015] The central dispatch server is deployed in the scenic spot private cloud data center, using Kubernetes container architecture, the core services include task manager, path planning engine, resource scheduler and state monitor. When receiving a user flight request, the task manager encapsulates it as a standardized task object, including fields: request_id (UUID), user_id, start_point (latitude and longitude), end_point, priority_weight, ai_theme, timestamp. The resource scheduler then queries the current available drone list, screening conditions include: in "standby" state, battery level higher than 70%, the distance between the starting point and the parking apron is less than 1.5 km, not in the no-fly zone. If there are multiple candidate drones, the system calculates the optimal allocation according to the dynamic priority weight formula: wherein, is the user membership level coefficient (ordinary user 1.0, VIP user 1.5, SVIP user 2.0), is the reward amount conversion score (1 yuan corresponds to 0.01 points), is the task submission time inverse ranking normalization value (the earlier the submission, the higher the score), weight coefficient = 0.4, = 0.35, = 0.25. The scheduler selects the drone with the highest value to execute the task, and reserves its associated parking apron as the return endpoint.
[0016] The path planning engine receives the task instruction and starts the multi-objective optimization algorithm to generate a comprehensive flight strategy. The input data includes: digital elevation model (DEM), scenic building 3D model, real-time airspace occupation map (from other drones' ADS-B broadcast), weather warning area, and tourist heat distribution map. The heat map is updated every minute by the data fusion and decision module, and its generation process is as follows: first, get the number of people entering and leaving each ticket gate from the scenic ticket system API, with a sampling period of 1 minute; second, perform spatial interpolation on the tourist density (unit: people / 100 square meters) identified by YOLOv8 in the drone video stream; finally, use the Gaussian kernel function to fuse the two types of data to generate a 10m x 10m resolution raster heat map. The objective function of path planning is defined as: wherein, is the path smoothness, measured by the integral inverse of the rate of change of heading angle; is the landscape coverage rate, calculated as the overlapping length ratio of the 50m buffer area of the flight path center line and the landmark scenic spots (such as the main building of the Pavilion of Prince Teng, the Stele Corridor, and the ancient wharf on the riverbank). To estimate energy consumption, the flight distance, wind resistance coefficient, and battery discharge curve are estimated. To avoid obstacles and ensure safety margin, the number of times the path crosses high-risk areas (such as high-voltage lines and dense tree canopies) is counted. The weight coefficients are set as = 0.3, = 0.4, = 0.2, = 0.1 to ensure that the route is both visually appealing and safe to operate. The planning results are output in the form of a waypoint sequence, including the latitude, longitude, and altitude of each waypoint, flight speed (default 8 m / s, reduced to 5 m / s in complex areas), hovering time, and camera gimbal attitude angle.
[0017] After receiving the flight instructions, the UAV performs the full process operation from S110 to S180. Before takeoff, self-checking is performed: check the motor phase, IMU zero offset, GPS star number (require ≥ 8), and communication link RSSI (require ≥ -85 dBm). After passing the self-checking, the hatch is opened, and the UAV vertically rises to 10 m in height and enters the cruise mode. During flight, the onboard edge computing module continuously runs the local re-planning algorithm. When the laser radar point cloud detects dynamic obstacles (such as kites or birds) 80 m ahead, the system immediately starts the emergency obstacle avoidance process: first, generate three alternative detour paths within a radius of 50 m based on the RRT* algorithm; second, evaluate the landscape influence degree (deviation from the original route angle), energy consumption increment, and flight time of each path; finally, select the path with the smallest influence to execute the detour, and report the path change information to the central dispatch server through the 5G link, and the server synchronously updates the airspace occupation state to prevent conflicts with other UAVs.
[0018] The AI interaction engine responds to user questions synchronously at stage S150. The natural language processing module in the engine uses the BERT-base architecture, and the word table is expanded to 150,000, including local proprietary words such as "Tengwang Pavilion Preface" and "Ganpo Culture". The model is fine-tuned after pre-training using 200,000 tour guide word questions and answers, supporting the understanding of complex semantics, such as "How old was Wang Bo when he wrote the Tengwang Pavilion Preface?" The model can accurately analyze "person = Wang Bo", "event = writing the Tengwang Pavilion Preface", and "demand = age", and retrieve "AD 675, 26 years old" as the answer from the knowledge graph. For open-ended questions such as "What legends are there about this bridge?", the system switches to a generative mode, calling a small GPT-2 model based on Transformer to generate narrative text. The output is checked by a fact-checking module against authoritative databases such as "Nanchang Prefecture Annals" and "Jiangxi Annals" to ensure the accuracy of "person", "time", and "location". The generated text is converted into an audio stream by a speech synthesis module. The TTS engine uses the FastSpeech square meter 2 architecture, supporting seven emotional tones (solemn, cheerful, lyrical, mysterious, passionate, calm, and nostalgic), and automatically matches the tone based on the type of question: "solemn" for historical anecdotes and "lyrical" for landscape descriptions. Users can choose a virtual anchor role, such as "Wang Bo" in a classical style of ancient Chinese language, who explains, "This is the key to the Ganjiang River. In the past, I climbed the hill and composed a poem, with the setting sun and the lone duck flying together..."
[0019] The live streaming and content management module performs multi-layer superimposition processing on the original video stream at stage S160. The original H.265 encoded video stream enters the GPU acceleration processing pipeline and performs decoding, image enhancement (contrast enhancement, dehazing algorithm), AI face blurring, graphic overlay, and re-encoding in sequence. The face blurring module uses a real-time convolutional neural network with a U-Net variant structure, an input resolution of 1920x1080, and a processing frame rate of 60fps. An adaptive Gaussian kernel is applied to the detected face area, with the kernel size dynamically adjusted based on the picture clarity (σ=8 for 1080p and σ=6 for 720p), ensuring that the blurred face cannot be identified. The graphic overlay layer includes four areas: the upper left corner displays the current scenic spot name and AI anchor avatar; the upper right corner displays real-time interactive question bubbles (e.g., "Netizen 'Ganjiang Fisherman' asks: What is the current temperature?" "AI answers: The current body temperature is 22°C, with a gentle breeze"); the lower horizontal banner scrolls to display recommended shooting angles (e.g., "Best shooting angle: 30° upward, focus on the main pavilion"); and the bottom permanent e-commerce portal automatically matches the SKU number TGG-SG001 in the product database when the AI recognizes "Tengwang Pavilion Cultural and Creative Ice Cream" in the picture, and pops up a limited-time purchase button that links to the WeChat mini-program mall. All superimposed content is filtered by a security audit rule engine, which includes 1,327 sensitive word regular expressions to ensure compliance.
[0020] The safety monitoring and emergency response module implements a three-level response mechanism. Level 1 response addresses communication delays: if the server does not receive a heartbeat packet from the drone for 3 consecutive seconds, it determines the link is unstable. The drone immediately uses its local cached flight path, hovers over its current location, and switches to the 4G backup network for reconnection. Level 2 response addresses energy and weather anomalies: if the battery level drops below 20% or the helipad reports excessive wind speed, the drone terminates its current mission, initiates an automatic return-to-home procedure, and returns to the nearest available helipad along the shortest safe path, maintaining video streaming during the return journey. Level 3 response addresses major malfunctions: if the GPS signal is lost for more than 10 seconds or IMU data is abnormal, the system determines navigation failure, immediately performs an emergency landing, selects a flat area (slope <5° determined by LiDAR scanning), slowly descends to 1m above the ground, releases the parachute buffer, activates the onboard buzzer and LED strobe lights, and simultaneously sends a positioning beacon (containing latitude, longitude, fault code, and timestamp) to the server for rapid search by ground personnel.
[0021] After the mission is completed, the drone executes the S180 procedure: it lands precisely at the target helipad charging port with an error controlled within ±3cm; after the hatch closes, it begins charging with a charging current of 5A and a voltage of 22.8V, taking approximately 45 minutes to fully charge; at the same time, the original log files (including flight trajectory, sensor data, and AI recognition records) in the onboard storage are uploaded to the cloud archiving server via gigabit Ethernet. After the upload is completed, the device status is updated to "standby," awaiting the next mission scheduling.
[0022] Example 2 This embodiment focuses on the technical implementation of multi-drone collaborative live streaming and AI-powered deep content generation, using a famous area in Nanchang as the application field to address the needs of wide-area coverage and thematic content production in large scenic areas. Compared with Embodiment 1, this embodiment introduces a "task merging mechanism" and a "multimodal content generation pipeline" in its system architecture, and strengthens the collaborative optimization capabilities of steps S130 and S160 in its methodology, forming a differentiated technical path.
[0023] When the central dispatch server receives flight requests from multiple users within the same time period, the system initiates a task merging algorithm. Let... There are N pending requests at any given time. Each request includes a starting point Pi, an ending point Qi, a submission time Ti, and a priority weight Wi. The system first calculates the path similarity between any two requests. Defined as: in, For the request Preliminary planned path point set, For the request Preliminary planned path point set, represents the path length (Euclidean distance). If > 0.6 and seconds, it is determined that the two tasks can be merged. After merging, a joint flight path is generated, covering all starting points and ending points, and the traveling salesman problem (TSP) solver is used to optimize the access order, with the goal of minimizing the total flight distance. The merged tasks are executed by the same unmanned aerial vehicle, which sequentially completes the viewing actions (such as hovering, circling, and diving) specified by each user during the flight, and polls the AI interaction engine to respond to questions from each user. For example, user A requests to "see the old site in a certain area", and user B requests to "take pictures at the Yellow Border Whistle", the system plans a serial route, first flies to a certain area, stays for 2 minutes for user A to interact, then goes to the Yellow Border, executes the aerial photography action, and the AI anchor responds to the questions of the two users during the period, realizing efficient reuse of resources.
[0024] At the content generation level, the embodiment constructs a multi-modal AI generation pipeline. The live streaming module no longer relies solely on static templates, but constructs a dynamic content graph. Whenever the unmanned aerial vehicle recognizes a specific scene (such as "poetry wall"), the system triggers the content generation workflow: first, extract the entity "famous poetry" from the scenic spot knowledge graph; second, call the text generation model to generate a commentary; third, start the image generation model (such as the lightweight version of Stable Diffusion) to generate a stylized illustration according to the description "poetry scene"; finally, superimpose the illustration as a floating layer in the corner of the live streaming screen, fade out after 15 seconds. This process forms a complete chain of "visual recognition -> knowledge retrieval -> text generation -> image generation -> multi-modal output", significantly improving the richness of the content.
[0025] In addition, the embodiment optimizes the e-commerce recommendation logic in step S160. The system not only matches goods based on screen content, but also introduces user behavior sequence analysis. Let user u have a behavior sequence during the live streaming, where ∈{watch product A, square meters ask about B, square meters reward anchor C, square meters click link D}. The system uses an LSTM network to model the evolution of user interest, outputting the most likely product category of interest at the next moment. For example, if a user asks three consecutive questions about "the source of poetry", the system predicts that the interest degree of "poetry-related" categories will increase by 80%, and accordingly, related goods such as "poetry bookmarks" and "poetry fans" will be preferentially pushed in subsequent screens, achieving precise marketing.
[0026] Embodiment Three This embodiment is designed to enhance the stability of the system in high-humidity environments during the rainy season and to protect privacy. It takes the ancient village of Luoling in Wuyuan as the application object, and focuses on solving the problems of device failure caused by humid climate and privacy compliance in complex crowd scenarios. Compared with the previous two embodiments, this embodiment makes substantial improvements in hardware structure and security mechanisms.
[0027] The automatic apron is additionally provided with an environmental control subsystem in the embodiment. A temperature and humidity sensor (precision ±2% RH, ±0.5°C) is installed inside the cabin, and when the relative humidity exceeds 85% for 10 minutes, a dehumidification module is started: a semiconductor condensation dehumidification technology is used, with a rated power of 40 W and a maximum dehumidification capacity of 1.2 L / day, to control the humidity in the cabin to below 60%. At the same time, the charging interface is provided with self-cleaning metal contacts, with a gold plating treatment on the surface to prevent oxidation from causing poor contact. The UAV performs a "dry flight" before each landing: it hovers 30 seconds above the apron, using the wash air flow generated by the propellers to dry the water droplets on the body surface, especially the camera lens and the laser radar window, to ensure that the sensor performance is not affected during the next take-off.
[0028] In terms of privacy protection, the embodiment proposes a hierarchical face blurring strategy. The system classifies the characters in the live picture according to their activity states: a standard Gaussian blur (σ=8) is used for stationary crowds (such as tourists taking pictures); the blurring strength is reduced to σ=5 for fast-moving individuals (such as running children) because motion blur already exists; and the face area of the staff wearing RFID badges can be automatically exempted from blurring when the face area is identified by the UAV UHF reader. This mechanism ensures privacy while retaining necessary dynamic information of the characters, improving the picture's aesthetic value.
[0029] The data fusion and decision module introduces meteorological prediction data in the embodiment. The system accesses the meteorological bureau API to obtain the precipitation probability, wind speed trend and lightning warning in the next 2 hours. When the precipitation probability is predicted to rise above 70% in the next 30 minutes, the UAV is scheduled to return to the nearest apron to wait for orders, avoiding encountering sudden rainfall during flight. Known areas prone to water accumulation (such as low-lying dikes and stone bridges) are actively avoided during path planning to ensure flight safety.
[0030] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the above disclosed technical content without departing from the scope of the technical solution of the present application, and equivalent embodiments with equivalent changes and modifications are equivalent to the above embodiments. Any modification, change, equivalent change and modification of the above embodiments made according to the technical essence of the present application without departing from the scope of the technical solution of the present application are still within the scope of the technical solution of the present application.
Claims
1. An AI interaction-based scenic area unmanned aerial vehicle live broadcast control system, characterized in that , including the following parts: A drone flight unit for performing aerial shooting and live streaming tasks and transmitting high-definition video streams in real time through a 5G network; An automated apron network distributed at key points in the scenic area to provide automatic take-off and landing, charging, and environmental self-checking services for drones, supporting continuous operation; An AI interaction engine integrating natural language processing, computer vision, and speech synthesis modules to analyze user interaction instructions, generate intelligent commentary content, and drive virtual anchors to respond in real time; A user interaction platform providing mobile and web interfaces to support online users in sending flight route requests, asking questions, rewarding, and participating in live streaming e-commerce activities; A central dispatch server responsible for receiving user requests, planning flight paths, coordinating drone resource allocation, dispatching AI services, and managing live streaming distribution; A data fusion and decision-making module for integrating scenic area geographic information, tourist distribution heat maps, weather data, and user behavior data to generate dynamic flight strategies and content recommendation schemes; A live streaming and content management module responsible for real-time processing of video streams, superimposing AI interaction engine-generated content, and pushing it to third-party live streaming platforms; A safety monitoring and emergency response module that monitors drone status, airspace compliance, and privacy protection policy implementation in real time and triggers automatic return or emergency stop mechanisms in the event of an anomaly.
2. The system of claim 1, wherein The drone flight unit is equipped with multispectral imaging equipment, including a 4K visible light camera, an infrared thermal imager, and a laser radar. The data from the three types of sensors are synchronized and transmitted to the central dispatch server after timestamp alignment. The infrared thermal imager is used to capture temperature distribution characteristics in night tour mode to enhance visual expressiveness, and the laser radar is used to construct a local three-dimensional point cloud map to support precise obstacle avoidance and path optimization in complex terrain.
3. The system of claim 2, wherein The drone flight unit has an embedded edge computing module running a lightweight YOLOv8 target detection model to identify scenic landmark buildings, vegetation types, and tourist gathering areas in real time during flight, and embed the recognition results as semantic tags in video stream metadata for the AI interaction engine to generate commentary content with geographical and cultural context.
4. The system of claim 3, wherein The automated apron network adopts a modular design, with each apron equipped with an environmental perception sensor group, including an anemometer, a rain sensor, and an illuminance meter, to monitor local weather conditions in real time. When the wind speed exceeds a threshold or the rainfall intensity is greater than a preset value, the system automatically sends a no-fly warning to the central dispatch server and starts the hatch closing and equipment waterproofing procedures.
5. The system of claim 4, wherein The natural language processing module in the AI interaction engine uses a domain adaptation model based on the BERT architecture. This model is fine-tuned using scenic tour guides, historical documents, and tourist common question corpora on the basis of pre-training, enabling it to accurately understand local cultural terms, dialect expressions, and historical anecdotes, and support precise analysis of complex requests containing specific location and time semantics.
6. The system of claim 5, wherein The AI interaction engine is configured with a multi-role virtual anchor library, including historical figure images, local culture spokespersons and cartoon IP images. Users can select an anchor style through an interactive platform. The AI system generates commentary words in line with the language style of the selected role according to the personality setting and knowledge base of the role, and outputs audio streams with emotional intonation through voice synthesis technology.
7. The system of claim 6, wherein The user interaction platform is provided with a flight route point selection function. After the user selects a starting point and an ending point on a two-dimensional or three-dimensional map of the scenic spot, the system converts the selected starting point and ending point into a geographic coordinate sequence and submits it to the central dispatch server. The server combines real-time airspace occupation, remaining power of the unmanned aerial vehicle and meteorological data, and generates a flight path that meets the safety constraints and has the optimal visual observation by using an improved A* algorithm.
8. The system of claim 7, wherein The central dispatch server deploys a dynamic resource allocation strategy. When multiple users submit flight requests at the same time, the system calculates the priority weight according to the request timestamp, user membership level and reward amount, responds to high-weight requests first, and integrates multiple requests with similar paths into a joint flight task through a task merging mechanism to improve the efficiency of the unmanned aerial vehicle.
9. The system of claim 8, wherein The data fusion and decision module accesses the API of the scenic spot ticket system to obtain real-time data of the number of visitors and the length of stay in each area, combines the visitor density information recognized by AI in the video stream, and generates a minute-level updated visitor heat distribution map. The heat map is an important input parameter for path planning, guiding the unmanned aerial vehicle to preferentially cover high-popularity areas to improve the attraction of live streaming.
10. A control method applied to the AI interaction-based scenic spot unmanned aerial vehicle live broadcast control system of any one of claims 1-9, characterized in that The method comprises the following steps: Starting the unmanned aerial vehicle flight unit in standby state through the automatic airport network, completing self-checking and environment evaluation, and entering flight state; Receiving flight route requests or interactive question instructions of online users through the user interaction platform, wherein the flight route request includes target geographic location coordinates or natural language description of viewing demand; Transmitting the received user request to the central dispatch server, and generating a comprehensive flight strategy by the server calling the data fusion and decision module, wherein the comprehensive flight strategy includes the optimal flight path, the estimated flight time, the energy consumption estimation and the AI commentary theme; The central dispatch server sends flight control instructions to the specified unmanned aerial vehicle, and the unmanned aerial vehicle executes aerial photography task along the planned path and returns high-definition video stream in real time; The AI interaction engine synchronously analyzes the user question content, generates a semantic answer combining the scenic spot knowledge graph, and outputs an audio stream through voice synthesis technology for synchronous playback with the unmanned aerial vehicle shooting picture; The live streaming and content management module processes the original video stream in real time, superimposes the AI-generated text commentary, interactive question and answer feedback and e-commerce commodity link information, and pushes it to the mainstream live streaming platform; The safety monitoring and emergency response module continuously monitors the flight state of the unmanned aerial vehicle, the quality of the communication link and the execution of privacy protection, and immediately executes the preset emergency program when an abnormality is found; After the task is completed, the unmanned aerial vehicle automatically returns to the nearest idle airport, completes landing, charging and data uploading, and the system updates the device state to standby.
Citation Information
Patent Citations
Unmanned aerial vehicle control method and system for live broadcast interaction, and aircraft
CN119165877A
Automatic route planning method of unmanned aerial vehicle for electric power inspection
CN120333453A
Scenic spot immersive viewing and tour guide system based on unmanned aerial vehicle and XR head-mounted display
CN120653111A
ROS-based four-rotor tour guide unmanned aerial vehicle system and tour guide auxiliary method thereof
CN120909320A
Defect diagnosis system for power line unmanned aerial vehicle inspection
CN121256486A
Cited By
Critical heat flux density prediction method based on micro-nano hierarchical structure and machine learning
CN121646356A
Critical heat flux prediction method based on micro-nano hierarchical structure and machine learning
CN121646356B