AI-driven multi-mode collaborative perception interaction system and mixed reality terminal
Through the AI-driven multimodal collaborative perception interaction system, combined with multiple sensors and deep learning models, the problems of low detection accuracy, high false alarm rate and limited detection range in drone rescue are solved, and high-precision and panoramic rescue support is achieved.
Patent Information
- Application Number
- CN202510351435.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing drone rescue system has low detection accuracy, high false alarm rate, limited detection range, severe environmental noise interference, and cannot effectively identify signs of human life in high temperature and complex environments.
Using AI-driven multimodal collaborative perception interaction system, combining visual, auditory, motion and environmental modal sensors, cross-modal information fusion and interaction is carried out through deep learning models, and a three-dimensional model is built to provide immediate feedback and adaptive interaction.
It improves detection accuracy and identification accuracy in complex environments, reduces false alarm rates, expands the detection range, and provides panoramic stereoscopic environmental observation and integrated rescue solutions.
Smart Images

Figure CN120406724A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and specifically to an AI-driven multi-modal collaborative perception and interaction system and a mixed reality terminal. Background Art
[0002] With the continuous improvement of unmanned aerial vehicle technology, unmanned aerial vehicles have been widely used in search and rescue. The existing technology uses unmanned aerial vehicles combined with infrared imaging, biosensors, and acoustic wave sensors. The unmanned aerial vehicle flies and scans over the disaster area. The infrared camera captures the thermal imaging pictures of the ground or ruins in real time. Rescue workers analyze the heat source distribution in the received images to identify possible human or animal life signs. The biosensor detects the carbon dioxide concentration in the air, the gas components generated by breathing, or other vital sign signals in real time. These data are transmitted wirelessly to the ground station. Rescue workers analyze the data to determine whether there are life signs and determine their approximate locations. The acoustic wave sensor listens to the sounds in the environment in real time, such as calls for help, knocking sounds, or movement sounds. Through sound source localization technology, the unmanned aerial vehicle can determine the source location of the sound and transmit the information to the rescue workers.
[0003] However, the existing technology still has the following problems:
[0004] 1) Unmanned Aerial Vehicle + Infrared Imaging
[0005] Great influence from environmental interference: Infrared imaging relies on the detection of thermal radiation. However, in high-temperature environments (such as fire scenes) or direct sunlight, there are many heat source interferences, which may lead to false alarms. Weather conditions such as rain, fog, or thick smoke will also affect the clarity and accuracy of infrared imaging.
[0006] High false alarm rate: Infrared imaging may misjudge animals, heating mechanical equipment, or other heat sources as humans.
[0007] 2) Unmanned Aerial Vehicle + Biosensor
[0008] Weak signal: Biosensors rely on detecting weak vital sign signals (such as breathing and heartbeat). However, in complex environments, these signals may be masked by noise, resulting in detection failure.
[0009] Environmental interference: In chemical leakage or fire scenes, harmful gases or high temperatures in the air may interfere with the normal operation of the sensor, reducing the detection accuracy.
[0010] Limited detection range: The detection range of biosensors is usually small, and the unmanned aerial vehicle needs to fly at a low altitude, which may increase the operation difficulty and risk.
[0011] 3) Unmanned Aerial Vehicle + Acoustic Wave Sensor
[0012] Environmental noise interference: At the disaster scene, environmental noises (such as wind sounds, mechanical sounds, and rescue equipment sounds) may cover up the calls for help or knocking sounds of survivors, resulting in detection failure.
[0013] Limited detection distance: The effective detection distance of the acoustic wave sensor is short. Especially in an open or noisy environment, the sound signal attenuates quickly, making it difficult to accurately locate.
[0014] Survivor status limitation: If the survivor is unable to make a sound (such as being unconscious, injured, or trapped in a soundproof environment), acoustic wave detection will not be effective.
[0015] Therefore, an AI-driven multi-modal collaborative perception interaction system and a mixed reality terminal are proposed to assist drones in search and rescue operations. Summary of the Invention
[0016] Aiming at the deficiencies of the prior art, the present invention provides an AI-driven multi-modal collaborative perception interaction system and a mixed reality terminal, which have the advantages of establishing a three-dimensional model of the disaster area, multi-modal collaborative perception, and improving search and rescue efficiency, and solve the problems of information flattening, high operation threshold, and fuzzy positioning in traditional drone rescue.
[0017] To achieve the above object, the present invention provides the following technical solutions: An AI-driven multi-modal collaborative perception interaction system, comprising:
[0018] Multi-modal perception module: Used to obtain the content collected by multiple sensors, cameras, and voice input devices carried by the drone, construct an information modal network, covering the physical information collected by the sensors and the digital information obtained by the cameras and language input devices;
[0019] Multi-modal information processing module: Used to denoise, align, and extract features from the collected multi-modal information, and preprocess the multi-modal information data;
[0020] Collaborative perception and fusion module: Achieve cross-modal alignment through timestamp synchronization and spatial coordinate mapping, use a deep learning model to obtain cross-modal joint features, and identify scene information and personnel information;
[0021] AI-driven module: Used to process the information input across modalities, generate a unified recognition and understanding of different information, optimize the interaction strategy in a dynamic environment, and co-train a multi-modal model across devices;
[0022] Interaction and feedback module: Used to construct an interaction interface through voice dialogue, gesture recognition, and VP visualization, and provide an instant response through tactile vibration, visual cues, or voice synthesis.
[0023] Furthermore, the information modality network includes a visual modality, an auditory modality, a motion modality, and an environmental modality: The visual modality captures on-site information through a drone camera;
[0024] The auditory modality acquires voice information in the venue through a voice input device;
[0025] The motion modality captures the object posture through an accelerometer and a gyroscope;
[0026] The environmental modality obtains the temperature, humidity, and gas composition on-site through sensors.
[0027] Furthermore, the multi-modal information processing module includes time synchronization, spatial alignment, and feature extraction:
[0028] Time synchronization unifies the clock through the PS / PTP protocol and aligns the voice and action video using dynamic time warping;
[0029] Spatial alignment performs joint calibration of the camera and LiDAR, establishes a coordinate mapping using a checkerboard calibration board, and constructs a 3D environment model shared by multiple sensors;
[0030] Feature extraction uses YOLO object detection, ResNet image classification, and Mask R-CNN instance segmentation for visual feature extraction, MFCC acoustic feature extraction and Whisper speech-to-text for speech feature extraction, and pressure distribution matrix coding and vibration spectrum analysis for tactile feature extraction.
[0031] Furthermore, the collaborative perception and fusion module performs cross-modal joint feature extraction through the Transformer framework of the deep learning model, uses a neural network to learn graph-structured data, and extracts and discovers the features and patterns in the graph-structured data.
[0032] Furthermore, the AI-driven module also includes a multi-modal large model, a reinforcement learning unit, and a federated learning unit:
[0033] The multi-modal large model trains the model with positive and negative sample pairs, makes cross-modal samples with similar semantics close in the embedding space, constructs a pre-training task using the natural association between modalities, converts data of different modalities into sequence inputs, models cross-modal interactions through the self-attention mechanism, designs a specific embedding layer for each modality, and then inputs it into the shared Transformer backbone network;
[0034] The reinforcement learning unit optimizes the interaction strategy in a dynamic environment;
[0035] The federated learning unit collaboratively trains a multi-modal model across devices while protecting privacy.
[0036] Further, the interaction and feedback module includes multimodal input, multimodal output, and adaptive interaction:
[0037] The multimodal input converts the collected voice into text, judges the user's emotion through intonation and speech rate, captures the gestures of the person, simulates physical interaction through vibration or resistance, and captures subtle tactile signals;
[0038] The multimodal output converts text into natural speech, superimposes virtual information onto the real scene, adjusts the UI layout according to the user's behavior, transmits information through device vibration, and simulates physical interaction;
[0039] The adaptive interaction includes context awareness and dynamic adjustment:
[0040] The context awareness adjusts the interaction mode according to the environment;
[0041] The dynamic adjustment selects the best feedback method according to the user's preference and switches to the backup plan when some modalities fail.
[0042] An AI-driven multimodal collaborative perception interaction mixed reality terminal, including a mixed reality terminal constructed by an AI-driven multimodal collaborative perception interaction system, and the construction of the mixed reality terminal includes the following steps:
[0043] 1) Use RealityCapture to construct a three-dimensional model of the search and rescue scene, obtain the factors extracted by the AI-driven multimodal collaborative perception interaction system, match the same feature points in different images through the feature matching algorithm, establish the corresponding relationship between the images, and calculate the position of each feature point in the three-dimensional space through feature matching to generate a sparse point cloud;
[0044] 2) Generate a dense point cloud through the dense matching algorithm. After the point cloud is generated, the system will convert the point cloud into a three-dimensional mesh model through the triangulation algorithm and import the three-dimensional model of the disaster area as the basic scene;
[0045] 3) Achieve precise positioning of the VisionPro wearer through SLAM technology, calculate the field of view angle in combination with gyroscope data, receive the data pushed by the server in real time, and use ray detection to judge whether the position of the living body is within the user's field of view;
[0046] 4) If it is within the field of view, use the particle system of Unity to render the dynamic point cloud effect; if it is not within the field of view, only display the marker points at the corresponding positions on the three-dimensional map, and optimize the visual effect through the URP rendering pipeline of Unity to ensure the real-time rendering performance of the point cloud and the marker.
[0047] Further, the specific steps for the RealityCapture to perform automated three-dimensional modeling are as follows:
[0048] 1) Data preprocessing: The system classifies and organizes the images, eliminates blurred, duplicate, or low-quality frames, and creates structured data tags based on time, geographical location, or shooting angle.
[0049] 2) Image alignment and point cloud generation: Using the MVS multi-view stereo vision algorithm, the system performs feature point matching and spatial alignment on the images to construct a sparse point cloud covering the disaster area.
[0050] 3) Generate a dense point cloud with millimeter-level accuracy through dense matching to restore details such as terrain and building debris.
[0051] 4) Meshing and texture mapping: Convert the point cloud into a triangular mesh model, and accurately map the color information of the original photo to the mesh surface through the "UV unwrapping" technology to generate a 3D model with real textures.
[0052] Furthermore, the mixed reality terminal captures the user's gesture actions through the built-in camera or sensor, uses ARFoundation to perform real-time analysis and classification of the user's gestures. After recognizing a specific gesture, it converts the gesture information into corresponding drone control commands, and the commands are transmitted to the drone end through Socket communication technology. After receiving the commands, the drone performs corresponding actions.
[0053] The camera carried by the drone will capture the surrounding environment in real time and transmit the video stream to the mixed reality terminal through 3DWebView. The user can view the images captured by the drone from the first-person perspective through the immersive display interface of the mixed reality terminal.
[0054] Compared with the prior art, the technical solution of this application has the following beneficial effects:
[0055] 1. The present invention optimizes the visual effect through the URP rendering pipeline of Unity to ensure the real-time rendering performance of the point cloud and the markers. The entire system adopts a hierarchical architecture, and data interaction between modules is carried out through standardized interfaces, ensuring the scalability and maintainability of the system.
[0056] 2. The present invention quickly constructs a three-dimensional full-scene dynamic model of the disaster area, providing panoramic three-dimensional environmental observation and various data supports for decision-making. The overall solution takes lightweight interaction, holographic perception, and real-time modeling as core advantages, solves the pain points such as information flattening, high operation threshold, and fuzzy positioning in traditional rescue, and provides an integrated intelligent solution for rescue under collapsed objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a framework diagram of the AI-driven multi-modal collaborative perception and interaction system of the present invention;
[0058] Figure 2Flowchart of the AI - driven multi - modal collaborative perception and interaction mixed reality terminal of the present invention. Detailed implementation manners
[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0060] Please refer to Figure 1 , a kind of AI - driven multi - modal collaborative perception and interaction system in this embodiment includes:
[0061] Multi - modal perception module: used to obtain the content collected by multiple sensors, cameras and voice input devices carried by the drone, construct an information modal network, covering the physical information collected by the sensors and the digital information obtained by the cameras and language input devices;
[0062] Multi - modal information processing module: used to denoise, align and extract features from the collected multi - modal information, and pre - process the multi - modal information data;
[0063] Collaborative perception and fusion module: realizes cross - modal alignment through timestamp synchronization and spatial coordinate mapping, uses a deep learning model to obtain cross - modal joint features, and identifies scene information and personnel information;
[0064] AI - driven module: used to process the information input across modalities, generate a unified recognition and understanding of different information, optimize the interaction strategy in a dynamic environment, and co - train a multi - modal model across devices;
[0065] Interaction and feedback module: used to construct an interaction interface through voice dialogue, gesture recognition and VP visualization, and provide an instant response through tactile vibration, visual cues or speech synthesis.
[0066] In this embodiment, the visual modality of the information modal network is the on - site information collected by the drone camera, the auditory modality is the voice information in the venue obtained by the voice input device, the motion modality captures the object posture through the accelerometer and gyroscope, and the environmental modality obtains the temperature, humidity and gas composition of the site through the sensor.
[0067] In this embodiment, the multimodal information processing module synchronizes clocks through the PS / PTP protocol, aligns speech and action videos using dynamic time warping, establishes a coordinate mapping using a checkerboard calibration board through the joint calibration of the camera and LiDAR, constructs a 3D environment model shared by multiple sensors, extracts visual features using YOLO object detection, ResNet image classification, and Mask R-CNN instance segmentation, extracts speech features using MFCC acoustic feature extraction and Whisper speech-to-text, and extracts tactile features using pressure distribution matrix encoding and vibration spectrum analysis.
[0068] In this embodiment, the collaborative perception and fusion module performs cross-modal joint feature extraction through the deep learning model Transformer framework. Transformer is an Encoder-Decoder architecture. The middle part of Transformer can be divided into two parts: the encoding component and the decoding component. Among the decoders of the same layer of the Transformer model, the encoding component consists of multiple layers of encoders, and the decoding component also consists of the same number of layers of decoders.
[0069] In this embodiment, the AI-driven module trains the model with positive and negative sample pairs to make cross-modal samples with similar semantics close in the embedding space, constructs a pre-training task using the natural association between modalities, converts data of different modalities into sequence inputs, models cross-modal interactions through the self-attention mechanism, designs specific embedding layers for each modality, and then inputs them into the shared Transformer backbone network to optimize the interaction strategy in a dynamic environment and co-train a multimodal model across devices while protecting privacy.
[0070] In this embodiment, the interaction and feedback module converts the collected speech into text, judges the user's emotion by intonation and speech rate, captures the person's gestures, simulates physical interactions through vibration or resistance, captures subtle tactile signals, converts the text into natural speech, superimposes virtual information on the real scene, adjusts the UI layout according to the user's behavior, and transmits information through device vibration to simulate physical interactions.
[0071] Please refer to Figure 2 , an AI-driven multimodal collaborative perception and interaction mixed reality terminal, including a mixed reality terminal constructed by an AI-driven multimodal collaborative perception and interaction system. The construction of the mixed reality terminal includes the following steps:
[0072] 1) Use RealityCapture to construct a three-dimensional model of the search and rescue scene, obtain the factors extracted by the AI-driven multimodal collaborative perception and interaction system, match the same feature points in different images through a feature matching algorithm, establish the corresponding relationship between the images, and calculate the position of each feature point in the three-dimensional space through feature matching to generate a sparse point cloud;
[0073] 2) Through the dense matching algorithm, a dense point cloud is generated. After the point cloud is generated, the system will convert the point cloud into a three-dimensional mesh model through the triangulation algorithm and import the three-dimensional model of the disaster area as the basic scene;
[0074] 3) The precise positioning of the VisionPro wearer is achieved through SLAM technology. The field of view angle is calculated in combination with gyroscope data, and the data pushed by the server is received in real time. Ray detection is used to determine whether the position of the living body is within the user's field of view;
[0075] 4) If it is within the field of view, the dynamic point cloud effect is rendered using the particle system of Unity; if it is not within the field of view, only a marker point is displayed at the corresponding position on the three-dimensional map, and the visual effect is optimized through Unity's URP rendering pipeline to ensure the real-time rendering performance of the point cloud and the marker.
[0076] In this embodiment, the specific steps for RealityCapture to perform automated 3D modeling are as follows:
[0077] 1) Data preprocessing. The system classifies and organizes the images, eliminates blurred, duplicate or low-quality frames, and establishes structured data tags according to time, geographical location or shooting angle;
[0078] 2) Image alignment and point cloud generation. Using the MVS multi-view stereo vision algorithm, the images are subjected to feature point matching and spatial alignment to construct a sparse point cloud covering the disaster area;
[0079] 3) A dense point cloud with millimeter-level accuracy is generated through dense matching to restore details such as terrain and building debris;
[0080] 4) Meshing and texture mapping. The point cloud is converted into a triangular mesh model, and the color information of the original photo is accurately mapped to the mesh surface through the "UV unwrapping" technology to generate a three-dimensional model with real textures.
[0081] In summary, the present invention optimizes the visual effect through Unity's URP rendering pipeline to ensure the real-time rendering performance of the point cloud and the marker. The entire system adopts a hierarchical architecture, and data interaction between modules is carried out through standardized interfaces, ensuring the scalability and maintainability of the system, quickly constructing a three-dimensional full-scene dynamic model of the disaster area scene, providing panoramic three-dimensional environment observation and various data supports for decision-making. The overall solution takes lightweight interaction, holographic perception and real-time modeling as its core advantages, solves the pain points such as information flatness, high operation threshold and fuzzy positioning in traditional rescue, and provides an integrated intelligent solution for rescue under collapsed objects.
[0082] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising said element.
[0083] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An AI-driven multi-modal collaborative perception and interaction system, characterized in that, Including: Multi-modal perception module: It is used to obtain the content collected by multiple sensors, cameras, and voice input devices carried by the drone, construct an information modality network, covering the physical information collected by sensors and the digital information obtained by cameras and language input devices; Multi-modal information processing module: It is used to denoise, align, and extract features from the collected multi-modal information, and preprocess the multi-modal information data; Collaborative perception and fusion module: It realizes cross-modal alignment through timestamp synchronization and spatial coordinate mapping, uses a deep learning model to obtain cross-modal joint features, and identifies scene information and personnel information; AI-driven module: It is used to process the information input across modalities, generate a unified recognition and understanding of different information, optimize the interaction strategy in a dynamic environment, and co-train the multi-modal model across devices; Interaction and feedback module: It is used to construct an interaction interface through voice dialogue, gesture recognition, and VP visualization, and provide an immediate response through tactile vibration, visual cues, or speech synthesis.
2. An AI-driven multi-modal collaborative perception and interaction system according to claim 1, characterized in that, The information modality network includes visual modality, auditory modality, motion modality, and environmental modality: The visual modality is the on-site information collected by the drone camera; The auditory modality is the voice information in the venue obtained through the voice input device; The motion modality captures the object posture through the accelerometer and gyroscope; The environmental modality obtains the temperature, humidity, and gas composition on-site through sensors.
3. An AI-driven multi-modal collaborative perception and interaction system according to claim 1, wherein, The multi-modal information processing module includes time synchronization, spatial alignment, and feature extraction: Time synchronization unifies the clock through the PS / PTP protocol, and uses dynamic time warping to align speech and action videos; Spatial alignment is achieved through the joint calibration of the camera and LiDAR, uses a checkerboard calibration board to establish a coordinate mapping, and constructs a 3D environment model shared by multiple sensors; Feature extraction uses YOLO object detection, ResNet image classification, and Mask R-CNN instance segmentation for visual feature extraction, uses MFCC acoustic feature extraction and Whisper speech-to-text to achieve speech feature extraction, and uses pressure distribution matrix coding and vibration spectrum analysis for tactile feature extraction.
4. An AI-driven multi-modal collaborative perception and interaction system according to claim 3, characterized in that, The collaborative perception and fusion module performs cross-modal joint feature extraction through the Transformer framework of the deep learning model, uses a neural network to learn graph-structured data, and extracts and discovers the features and patterns in the graph-structured data.
5. An AI-driven multi-modal collaborative perception and interaction system according to claim 4, wherein The AI-driven module also includes a multi-modal large model, a reinforcement learning unit, and a federated learning unit: The multi-modal large model trains the model through positive and negative sample pairs, makes cross-modal samples with similar semantics close in the embedding space, constructs a pre-training task using the natural association between modalities, converts data of different modalities into sequence inputs, models cross-modal interactions through the self-attention mechanism, designs a specific embedding layer for each modality, and then inputs it into the shared Transformer backbone network; The reinforcement learning unit optimizes the interaction strategy in a dynamic environment; The federated learning unit co-trains the multi-modal model across devices while protecting privacy.
6. An AI-driven multi-modal collaborative perception and interaction system according to claim 1, characterized in that, The interaction and feedback module includes multi-modal input, multi-modal output, and adaptive interaction: Multimodal input converts the collected voice into text, judges the user's emotion by intonation and speech rate, captures the gestures of personnel, simulates physical interaction through vibration or resistance, and captures subtle tactile signals; Multimodal output converts text into natural speech, superimposes virtual information onto the real scene, adjusts the UI layout according to the user's behavior, transmits information through device vibration, and simulates physical interaction; Adaptive interaction includes context awareness and dynamic adjustment: Context awareness adjusts the interaction mode according to the environment; Dynamic adjustment selects the best feedback method according to the user's preference and switches to an alternative solution when some modalities fail.
7. An AI-driven multi-modal collaborative perception and interaction mixed reality terminal, including a mixed reality terminal constructed by using the AI-driven multi-modal collaborative perception and interaction system described in any one of claims 1-6, characterized in that, The construction of the mixed reality terminal includes the following steps: 1) Use RealityCapture to build a three-dimensional model of the search and rescue scene, obtain the factors extracted by the AI-driven multimodal collaborative perception and interaction system, match the same feature points in different images through the feature matching algorithm, establish the corresponding relationship between the images, and calculate the position of each feature point in the three-dimensional space through feature matching to generate a sparse point cloud; 2) Generate a dense point cloud through the dense matching algorithm. After the point cloud is generated, the system will convert the point cloud into a three-dimensional mesh model through the triangulation algorithm and import the three-dimensional model of the disaster area as the basic scene; 3) Achieve precise positioning of the VisionPro wearer through SLAM technology, calculate the field of view angle in combination with gyroscope data, receive the data pushed by the server in real time, and use ray detection to judge whether the position of the living body is within the user's field of view; 4) If it is within the field of view, use the particle system of Unity to render the dynamic point cloud effect; if it is not within the field of view, only display the marker points at the corresponding positions on the three-dimensional map, and optimize the visual effect through the URP rendering pipeline of Unity to ensure the real-time rendering performance of the point cloud and the marker.
8. An AI-driven multi-modal collaborative perception and interaction hybrid reality terminal according to claim 7, characterized in that The specific steps for RealityCapture to perform automated 3D modeling are as follows: 1) Data preprocessing, the system classifies and organizes the images, eliminates blurred, repeated or low-quality frames, and establishes structured data labels according to time, geographical location or shooting angle; 2) Image alignment and point cloud generation, use the MVS multi-view stereo vision algorithm to match the feature points of the images and perform spatial alignment to construct a sparse point cloud covering the disaster area; 3) Generate a dense point cloud with millimeter-level accuracy through dense matching to restore details such as terrain and building debris; 4) Meshing and texture mapping, convert the point cloud into a triangular mesh model, and accurately map the color information of the original photo to the mesh surface through the "UV unwrapping" technology to generate a three-dimensional model with real texture.
9. An AI-driven multi-modal collaborative perception and interaction mixed reality terminal according to claim 7, characterized in that The mixed reality terminal captures the user's gesture actions through the built-in camera or sensor, uses ARFoundation to perform real-time analysis and classification of the user's gestures, and after recognizing a specific gesture, converts the gesture information into the corresponding drone control instructions. The instructions are transmitted to the drone end through Socket communication technology, and the drone executes the corresponding actions after receiving the instructions; The camera mounted on the drone will capture the surrounding environment in real time and transmit the video stream to the mixed reality terminal through 3DWebView. Users can view the images captured by the drone from a first-person perspective through the immersive display interface of the mixed reality terminal.