Shared unmanned aerial vehicle travel shooting interaction method and system
By combining multi-sensor and deep learning algorithms with a multimodal fusion network, the technical challenge of judging user emotions in complex travel photography scenarios has been solved, enabling real-time and accurate emotion recognition and personalized services on the drone platform, adapting to resource constraints and protecting user privacy.
Patent Information
- Application Number
- CN202511133146.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-14
AI Technical Summary
In complex travel photography scenarios, contextual information such as user posture, angle, and distance changes rapidly, leading to numerous technical challenges in emotion recognition. These challenges include drastic changes in outdoor lighting, ever-changing user posture and angle, and limited computing resources of drones, making it difficult to achieve accurate and real-time emotion recognition.
By acquiring user posture and angle information through multiple sensors, extracting facial features using deep learning algorithms, and improving robustness through image enhancement technology, a multimodal fusion network is used for feature representation. User emotions are classified using an emotion discrimination model, and the model is optimized by combining incremental learning and federated learning to adapt to the resource constraints of the drone platform.
It achieves real-time and accurate emotion recognition on drone platforms in complex environments, protects user privacy, provides personalized shooting services, and improves user experience and system performance.
Smart Images

Figure CN120954070A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of drone interaction, and more particularly to a shared drone travel photography interaction method and system. Background Technology
[0002] In complex travel photography scenarios, the user's posture, angle, and distance, among other contextual information, change rapidly, posing numerous technical challenges to emotion recognition. First, drastic changes in outdoor lighting make facial feature extraction difficult and introduce noise interference, affecting the accuracy of emotion recognition. Second, the highly variable postures and angles of users make single-modal emotion recognition methods inadequate, necessitating the integration of multimodal information for joint judgment. Third, the limited computing platform resources and power of drones make it difficult to support computationally intensive emotion recognition models, requiring a balance between accuracy and real-time performance. Finally, effectively organizing, managing, and utilizing massive amounts of user emotion data for model training and optimization is a pressing issue. These intertwined technical challenges greatly complicate emotion recognition technology in complex scenarios, requiring a systematic analysis of the root causes and the comprehensive application of multiple technologies to build a robust, efficient, and intelligent emotion recognition system that provides more personalized and attentive photography services for travel photographers. Summary of the Invention
[0003] This invention provides a shared drone travel photography interaction method, mainly including: The system acquires user posture and angle information from multiple sensors and preprocesses it; it then uses deep learning algorithms to extract facial features and improves the robustness of the extraction through image enhancement techniques; combining the user's posture, angle, and facial features, it obtains a comprehensive feature representation through feature fusion; it uses this comprehensive feature representation for user identification and behavior analysis to construct a comprehensive user profile; it acquires the user's voice and body language information, and inputs it along with the facial feature vector into a multimodal fusion network to achieve joint representation of different modal features; finally, it inputs the multimodal fused feature vector into an emotion discrimination model to obtain the user's emotion classification result.
[0004] Furthermore, the process of acquiring user posture and angle information through multiple sensors and performing preprocessing includes: using a Kalman filter algorithm to remove noise and outliers to ensure the accuracy of posture and angle data.
[0005] Furthermore, the method of extracting facial features from users using deep learning algorithms and improving the robustness of the extraction through image enhancement techniques includes: using methods such as histogram equalization and contrast enhancement to reduce noise interference caused by changes in outdoor ambient lighting.
[0006] Furthermore, the process of combining the user's posture, angle, and facial features to obtain a comprehensive feature representation through feature fusion includes: using principal component analysis (PCA) for dimensionality reduction, removing redundant and irrelevant features, and obtaining a compact and discriminative feature vector.
[0007] Furthermore, the method of using comprehensive feature representation for user identification and behavior analysis to construct a comprehensive user profile includes: employing support vector machine (SVM) and logistic regression (LR) algorithms to achieve accurate identification of user identity and behavioral intent.
[0008] Furthermore, the acquisition of the user's voice and body posture information, along with the facial feature vector, is input into a multimodal fusion network to achieve joint representation of different modal features, including: using a multi-head attention mechanism to calculate the importance weights of different modal features.
[0009] Furthermore, the step of inputting the feature vector after multimodal fusion into the emotion discrimination model to obtain the classification result of user emotion includes: using a bidirectional long short-term memory network (Bi-LSTM) and an attention mechanism to capture the dynamic change features of user emotion, and optimizing the model through a cross-entropy loss function.
[0010] A shared drone travel photography interactive system, based on the aforementioned shared drone travel photography interactive method, is characterized by comprising: a data acquisition and preprocessing module, a facial feature extraction and enhancement module, a feature fusion module, a user identification and behavior analysis module, a multimodal information fusion module, and an emotion discrimination module; The data acquisition and preprocessing module is used to acquire the user's posture and angle information through multiple sensors and perform preprocessing. The facial feature extraction and enhancement module is used to extract the user's facial features using deep learning algorithms and to improve the robustness of the extraction through image enhancement technology. The feature fusion module is used to combine the user's posture, angle, and facial features to obtain a comprehensive feature representation through feature fusion. The user identification and behavior analysis module is used to identify and analyze users using comprehensive feature representations, and to build a comprehensive user profile. The multimodal information fusion module is used to acquire the user's voice and body posture information, which, together with the facial feature vector, is input into the multimodal fusion network to achieve joint representation of different modal features; The emotion discrimination module is used to input the feature vector after multimodal fusion into the emotion discrimination model to obtain the classification result of the user's emotion.
[0011] A computer device includes: a memory and a processor; the memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the shared drone travel photography interaction method.
[0012] A computer-readable storage medium having a computer program stored thereon, characterized in that: when the computer program is executed by a processor, it implements the steps of the shared drone travel photography interaction method.
[0013] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a shared drone travel photography interaction method and system. The method acquires user posture, facial, and voice information through multiple sensors, extracts features using deep learning algorithms, and utilizes a multimodal fusion network to achieve joint representation of features from different modalities. This invention introduces an attention mechanism to capture dynamic changes in emotion and compresses the model through knowledge distillation to adapt it to the drone platform. To address data privacy issues, this invention employs federated learning for distributed training. Furthermore, this invention introduces an incremental learning mechanism, enabling the model to continuously learn from new data, and constructs a monitoring mechanism to optimize model performance in real time. This method effectively integrates multimodal information, achieving real-time and accurate emotion recognition on the drone platform while protecting privacy, providing new technical support for human-computer interaction. Attached Figure Description
[0014] Figure 1 This is a flowchart of a shared drone travel photography interactive method and system according to the present invention. Detailed Implementation
[0015] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] like Figure 1 This embodiment of a shared drone travel photography interaction method may specifically include: S101. The user's posture and angle information are acquired through multiple sensors, and facial features are extracted using deep learning algorithms. Image enhancement techniques are employed to overcome noise interference caused by changes in outdoor lighting conditions, improving the robustness of facial feature extraction. For the extracted facial features, a convolutional neural network is used for representation learning to obtain compact and discriminative feature vectors.
[0017] User pose and angle information is acquired through multiple sensors. The acquired information is preprocessed, and a Kalman filter algorithm is used to remove noise and outliers, resulting in accurate pose and angle data. Image enhancement techniques are employed to preprocess the acquired user facial images, using methods such as histogram equalization and contrast enhancement to reduce noise interference caused by changes in outdoor lighting conditions and improve image quality. The preprocessed user facial images are then input into a facial feature extraction model based on a convolutional neural network (CNN). Through multi-layer convolution and pooling operations, key facial features are extracted. The extracted facial features are then subjected to dimensionality reduction using principal component analysis (PCA) to remove redundant and irrelevant features, resulting in compact and discriminative feature vectors. The pose and angle data are fused with the facial feature vectors to construct a comprehensive user feature representation. An autoencoder is used for unsupervised learning of the comprehensive features. By minimizing the reconstruction error, a more accurate and robust user feature representation is obtained. The optimized user feature representation is then input into a support vector machine (SVM) classifier to achieve user identification. Simultaneously, a behavioral analysis model based on Hidden Markov Models (HMM) is used to model and predict user behavior patterns. By correlating user identification results with behavioral analysis results, a comprehensive user profile is constructed, providing a foundation for subsequent personalized services and interactions. User interaction data is continuously collected, and the model is updated and optimized online to continuously improve the system's identification accuracy and predictive capabilities.
[0018] Specifically, in modern technology applications, acquiring user posture and angle information through multiple sensors is a common practice, especially in the fields of virtual reality and augmented reality. For example, in VR games, sensors can capture the player's body movements for precise character control. However, sensor data often contains noise and outliers, making Kalman filtering algorithms particularly important. Kalman filtering is an effective method for estimating linear dynamic systems. It continuously corrects the estimation of the system state through prediction and update steps, thereby removing noise and improving data accuracy. For example, in autonomous vehicles, Kalman filtering is used to integrate data from GPS and inertial navigation systems to provide more accurate position and velocity estimates. When processing user facial images, changes in outdoor lighting often affect image quality. Image enhancement techniques, such as histogram equalization and contrast enhancement, can effectively improve the visual appeal and quality of images. Histogram equalization adjusts the brightness distribution of an image, ensuring clear visibility in different brightness areas. This is particularly important in surveillance video analysis, helping to improve the accuracy of facial recognition. Furthermore, the preprocessed image is input into a Convolutional Neural Network (CNN)-based model for facial feature extraction. CNNs, through their multi-layered structure, can capture hierarchical features in images, which are highly discriminative for recognizing different faces. For example, in airport security systems, CNNs can help quickly and accurately identify the facial features of different passengers, improving security efficiency. The extracted facial features are then subjected to dimensionality reduction using Principal Component Analysis (PCA). PCA extracts the main feature components from the data, removing noise and redundant information, making the feature vectors more compact and discriminative. In the financial security field, PCA is often used to identify and verify customer identities to prevent identity theft. After fusing pose and angle data with facial feature vectors, a comprehensive feature representation of the user can be constructed. This fused data can be further optimized through autoencoders. By learning to reconstruct the input data, autoencoders can discover deeper feature relationships within the data, resulting in more accurate and robust feature representations. In smart home systems, this method can help the system better understand users' behavioral habits and lifestyles, thereby providing more personalized services. Finally, the optimized feature representation is input into a Support Vector Machine (SVM) classifier to achieve accurate user identification. SVM distinguishes different categories of data by constructing an optimal decision boundary, which is particularly important in access control systems, ensuring that only authorized individuals can enter specific secure areas. Simultaneously, combining this with Hidden Markov Models (HMMs) to analyze user behavior patterns can further enhance the system's security and intelligence. For example, in medical monitoring, analyzing patient behavior patterns can promptly detect abnormal behavior and provide timely medical assistance.The integrated application of these technologies can not only improve the accuracy of user identification and behavior analysis, but also provide solid technical support for subsequent personalized services and interactions. Continuous optimization and updating of these models will further enhance the overall performance of the system and the user experience.
[0019] S102. Based on the user's posture and angle data, combined with facial features, a comprehensive feature representation is obtained through feature fusion. The user recognition results and behavior analysis results are correlated to construct a comprehensive user profile, obtain interaction data, and update and optimize the model online to improve the system's recognition accuracy and prediction capabilities, thus providing a foundation for personalized services and interactions.
[0020] To acquire user pose, angle, and facial features, depth cameras such as Kinect or RealSense can be used to collect the user's skeletal joint positions and facial keypoint coordinates. Feature extraction algorithms such as PCA or LDA are then used to reduce the dimensionality of this data, resulting in a compact feature representation vector. Based on the comprehensive feature representation, a Support Vector Machine (SVM) algorithm is employed for user identification. A soft-margin SVM with a linear kernel function can be selected, and the penalty coefficient C is chosen through cross-validation. The SMO algorithm is used for training to obtain a user identity classification model. Based on user identity information, the user's historical clicks, browsing, and search records are retrieved from the log database. This behavioral data is cleaned and feature-engineered to construct a behavioral feature vector. Logistic Regression (LR) is used to model the behavioral features and determine the user's current behavioral intent, such as purchase intention. The LR model's parameters are optimized using gradient descent. The user identification results from the SVM are correlated with the behavioral intent analysis results from LR to construct structured user profile data. This profile data includes multiple dimensions such as user identity information, behavioral preferences, and intent. During user interaction, real-time logs of user clicks, browsing, and searches are acquired and updated in the behavioral log database. Simultaneously, incremental log data is used for feature extraction and serves as new sample data. The parameters of the LR model are then updated online using the gradient descent algorithm, enabling continuous optimization of the behavioral intent analysis model. The model's performance is evaluated by calculating metrics such as precision, recall, and AUC for user identification and behavioral intent analysis. Continuous iterative optimization of the model enhances the overall system's intelligence level, providing data support for personalized recommendations and intelligent interactions.
[0021] Specifically, in modern intelligent systems, depth cameras such as Kinect or RealSense capture the positions of users' skeletal joints and the coordinates of facial key points, forming the basis for accurate user recognition and interaction. For example, in smart home systems, cameras can monitor the activities of elderly people and children in real time. By analyzing their postures and facial expressions, the system can determine their emotions and needs, automatically adjusting environmental settings, such as temperature and lighting, to provide a more comfortable living environment. Using feature extraction algorithms such as Principal Component Analysis (PCA) or Linear Discriminant Analysis (LDA) to reduce the dimensionality of the collected data is crucial for extracting the most representative features, thereby reducing computational load and increasing processing speed. In the retail industry, by analyzing customers' walking paths and dwell points in stores, PCA can help extract key factors influencing customer purchasing decisions, such as store layout and merchandise placement, and then optimize these factors to increase sales. Support Vector Machine (SVM) algorithms are used for user recognition, constructing an optimal decision boundary to distinguish feature vectors from different users. In financial services, banks can use SVM to identify and verify customer identities, ensuring that only authorized users can conduct transactions, thus preventing fraud. Choosing a soft-margin SVM with a linear kernel function and selecting the penalty coefficient C through cross-validation effectively balances model complexity and training data fit, avoiding overfitting. Combining user historical behavior records, such as clicks, browsing, and searches, allows for the construction of more detailed user profiles. In e-commerce, analyzing users' browsing and purchase history enables precise recommendations of new products or promotions they may be interested in, thereby improving user satisfaction and purchase conversion rates. Logistic Regression (LR) is used to analyze user behavioral intent, and gradient descent optimizes model parameters, resulting in more accurate predictions. Correlating the user identification results of SVM with the behavioral intent analysis results of LR constructs structured user profile data that includes multiple dimensions such as user identity information, behavioral preferences, and intent. This comprehensive information application is particularly important in personalized recommendation systems. For example, in video streaming services, the system can recommend movies or TV shows that match a user's tastes in real time based on their historical viewing behavior and preferences. By acquiring user interaction logs in real time and updating them to a behavioral log database, the system can continuously learn and adapt to changes in user behavior. On online education platforms, by analyzing students' learning behavior and progress, educational software can adjust course difficulty and recommend suitable learning materials to help students learn more effectively. Finally, by calculating evaluation metrics such as precision and recall for user identification and AUC for behavioral intent analysis, the model's performance can be comprehensively evaluated. These metrics help the technical team identify the model's strengths and weaknesses, guiding subsequent optimization efforts and continuously improving the system's intelligence and user experience.
[0022] S103. Acquire the user's voice and body language information, and input them along with facial feature vectors into a multimodal fusion network to achieve joint representation of features from different modalities. The multimodal fusion network employs an attention mechanism and a cross-modal interaction layer to perform weighted fusion and complementary enhancement of features from different modalities. The fused feature vector is input into a gated recurrent unit network to model temporal information and capture the dynamic changes in the user's emotions. Based on the modeled emotional feature sequence, the weights of features at different times are dynamically adjusted through the attention mechanism to focus on key segments of the user's emotions.
[0023] After acquiring the user's voice, body posture, and facial feature information, the raw data undergoes preprocessing, including data cleaning, noise removal, and feature normalization, to ensure data quality and consistency. Then, for feature data of different modalities, independent convolutional neural networks (CNNs) are used for feature extraction and representation learning, yielding voice feature vectors, body posture feature vectors, and facial feature vectors. The extracted feature vectors from different modalities are input into a multimodal fusion network, where a multi-head attention mechanism is used to calculate the importance weights of the features from different modalities. Attention scores are generated by performing a dot product between the feature vectors and the learnable query vector, and then normalized using a softmax function to obtain the weight distribution. The feature vectors from different modalities are weighted and summed to achieve feature fusion. A cross-modal interaction layer is introduced into the multimodal fusion network, using feature concatenation and gating mechanisms to achieve complementary enhancement of features from different modalities. Feature vectors from different modalities are concatenated together, and then nonlinear transformations and information transfer are performed on the features through gating units (such as gated recurrent units, GRUs), capturing the interactions and dependencies between modalities and enhancing the representational power of the features. The fused multimodal feature vectors are input into a Bidirectional Long Short-Term Memory (Bi-LSTM) network. Leveraging the gating structure and bidirectional information transfer of LSTM, the network captures the long- and short-term dependencies and dynamic changes in user emotions. By setting appropriate hidden layer sizes and numbers, and employing gradient pruning and regularization techniques, the model's temporal modeling ability and generalization performance are improved. An attention mechanism is applied to the output of the Bi-LSTM, learning the importance weights of features at different time points through self-attention. The output sequence of the Bi-LSTM is multiplied by the learnable query vector to generate an attention score, which is then normalized using a softmax function to obtain the attention weight distribution. The output sequence of the Bi-LSTM is weighted and summed according to the attention weights to obtain the focused emotion feature representation. Finally, the attention-focused emotion feature vector is input into a fully connected layer, where a series of nonlinear transformations and feature transformations enhance the discriminativeness and separability of the features. A softmax activation function is applied to the output of the fully connected layer, transforming the output into a probability distribution of emotion categories. Using the cross-entropy loss function as the optimization objective, the model is trained and updated end-to-end through the backpropagation algorithm and gradient descent optimizer, ultimately achieving multi-class recognition of user emotions.
[0024] Specifically, in modern emotion recognition systems, the first step is to preprocess the user's speech, body posture, and facial features to ensure data quality and consistency. For example, when processing speech data, background noise and echo are typically removed, which helps improve the accuracy of speech recognition. In the preprocessing of body posture data, captured movements may need to be smoothed to eliminate jitter caused by sensor errors. Facial data preprocessing may include illumination correction and contrast adjustment of images to ensure that facial expressions can be clearly captured and recognized. After data preprocessing, data from various modalities are fed into independent convolutional neural networks for feature extraction. Taking facial features as an example, convolutional neural networks can extract key facial features, such as the position and shape of the eyes, mouth, and nose, which are crucial for recognizing user expressions. Speech feature vector extraction focuses on identifying elements such as tone, rhythm, and volume from the speech signal; these features can reflect the user's emotional state. The extracted feature vectors are then input into a multimodal fusion network, where a multi-head attention mechanism is used to calculate the importance weights of features from different modalities. For example, when a user speaks with a calm expression but reveals tension or excitement in their voice, a multi-head attention mechanism can more accurately identify the user's emotions by increasing the weight of speech features and decreasing the weight of facial features. During multimodal fusion, the cross-modal interaction layer further enhances the complementarity of different modal features through feature concatenation and gating mechanisms. For instance, concatenating facial and body feature vectors and processing them through a gated recurrent unit can effectively capture the interaction between the user's gestures and facial expressions while speaking; this interaction is key to understanding the user's emotions. The fused multimodal feature vector is then input into a bidirectional long short-term memory network (BiLSTM), which can capture the long- and short-term dependencies and dynamic changes in emotions. For example, in a conversation, a user's emotions may gradually shift from calm to excitement; BiLSTM, with its bidirectional structure, can simultaneously consider past and future emotional changes, thus more accurately capturing the overall trend of emotions. Applying a self-attention mechanism to the output of BiLSTM allows for further analysis of the importance of emotional features at different time points. In this way, the model can focus on those moments most crucial for emotion recognition, such as the emotional peak when a user is discussing a specific topic. Finally, through a series of nonlinear transformations in the fully connected layers, the attention-focused emotion feature vector is transformed into a probability distribution of emotion categories. This step ensures that the model can effectively distinguish different emotional states, such as happiness, sadness, or anger, from the comprehensive feature representation. The model is optimized using the cross-entropy loss function, ensuring both the accuracy of emotion recognition and the model's generalization ability.
[0025] S104. Input the feature sequence processed by the attention mechanism into the emotion discrimination model to obtain the classification result of the user's emotion. To address the resource limitations of the drone platform, a knowledge distillation method is used to compress the emotion discrimination model. During knowledge distillation, the output probability distribution of the large teacher model and the intermediate layer features are used to guide the training of the small student model, achieving effective knowledge transfer. By adjusting the temperature parameter and the weights of the loss function, the model performance and computational complexity are balanced.
[0026] This process involves acquiring multimodal data related to user emotions, including text, speech, and video, and preprocessing the data such as denoising and normalization. Convolutional Neural Networks (CNNs) are used to extract features from video frames, and Long Short-Term Memory (LSTM) networks are used to extract features from speech and text data, resulting in feature vectors for each modality. The extracted multimodal feature vectors are then input into an attention mechanism module, where features from different modalities are weighted and fused using attention weights to highlight key features related to emotional expression. These attention weights can be learned by training a feedforward neural network. The fused multimodal feature vectors are then input into a pre-trained large-scale emotion discrimination model, such as BERT or XLNet, to obtain the probability distribution of user emotions. The pre-trained model can be trained on a large-scale emotion-labeled dataset to learn rich emotion representation capabilities. To adapt to the resource constraints of the drone platform, a lightweight student model, such as MobileNet or ShuffleNet, is designed. Knowledge distillation is used to guide the training of the student model using the output probability distribution of the teacher model and intermediate layer features, achieving model compression. During knowledge distillation, a temperature parameter is introduced to control the softening degree of the teacher model's output probability distribution. Higher temperature parameters result in a smoother probability distribution, which helps the student model learn more knowledge. By adjusting the temperature parameters and weights in the distillation loss function, a trade-off between performance and computational complexity is struck to obtain an emotion discrimination model that meets the resource constraints of drones. The compressed emotion discrimination model is deployed on a drone platform, utilizing the drone's onboard cameras, microphones, and other sensors to collect multimodal user data in real time. After preprocessing and feature extraction, the collected data is input into the deployed emotion discrimination model for inference, determining the user's emotional state in real time and taking corresponding interaction strategies based on the emotional state, such as playing soothing music or adjusting flight attitude, to improve the user experience. Simultaneously, user feedback data, such as user emotion annotations and interactive behaviors, is collected on the drone platform to further fine-tune the emotion discrimination model, continuously improving its performance and generalization ability, and achieving continuous optimization and iteration of the model.
[0027] Specifically, in modern emotion recognition systems, the first step is to preprocess the user's multimodal data, including text, speech, and video data. For example, for video data, denoising is a necessary step, which can be achieved by using filters to remove random noise from the image, while normalization ensures that video frames have a uniform brightness and contrast standard before being input into the model. Such preprocessing helps the subsequent feature extraction process to be more accurate, because clear images and uniform standards can reduce errors during model training. Next, convolutional neural networks (CNNs) are used to extract features from the video frames. CNNs can effectively identify key visual elements in the video, such as the user's facial expressions and gestures, which are crucial for emotion recognition. For example, when a user displays a happy expression, a CNN can identify the corners of the mouth and the shape of the eyes when smiling. Similarly, long short-term memory networks (LSTMs) are used to process speech and text data because they are good at handling sequential data and capturing dynamic changes over time, such as pitch variations in speech and the frequency of emotional words in text. The feature vectors extracted from each modality are then input into an attention mechanism module, a step that involves weighted fusion by calculating the importance weights of features from different modalities. For example, if a video shows a user crying while the voice is calm, the attention mechanism might give higher weight to video features because the video provides more direct evidence of emotion. This weighted fusion helps the model more accurately identify the user's true emotional state. The fused feature vector is then fed into a pre-trained large-scale emotion discrimination model, such as BERT or XLNet. These models, due to their depth and complexity, are able to understand and process complex emotional expressions. For example, the BERT model can analyze implicit emotions in text through its multi-layered self-attention mechanism, providing in-depth insights into the user's emotional state. To adapt to resource-constrained drone platforms, lightweight student models, such as MobileNet, are designed and trained via knowledge distillation. In this process, the output of the teacher model (such as BERT) and intermediate layer features are used to guide the training of the student model. By introducing a temperature parameter to control the softening of the teacher model's output probability distribution, the student model can learn more subtle differences in emotional expression from the teacher model. For example, setting a higher temperature parameter can make the teacher model's output probability distribution smoother, helping the student model capture more nuanced emotion classification boundaries. Finally, the compressed emotion discrimination model was deployed on a drone platform, using its onboard camera and microphone to collect multimodal data from users in real time. By analyzing users' emotional states in real time, the drone can adjust its interaction strategies accordingly, such as playing soothing music when it detects a user's low mood, or performing light and agile flight maneuvers when the user is happy, thereby enhancing the user's interactive experience.
[0028] S105. Design a distributed data storage and processing architecture, employing a combination of a distributed file system and a non-relational database to uniformly organize and manage heterogeneous sentiment data. Based on this distributed data architecture, a federated learning approach is used for model training and optimization. During federated learning, each node trains its model based on local data, sharing only model parameters rather than the original data, while the central server aggregates and updates the global model. Differential privacy and secure aggregation technologies are used to protect data privacy and model security. The sentiment discrimination model is trained in parallel on different nodes to achieve collaborative model optimization.
[0029] Based on the characteristics of heterogeneous sentiment data, non-relational databases such as MongoDB and distributed file systems such as HDFS are used for unified storage and management, achieving efficient data organization and access. Necessary preprocessing of the heterogeneous sentiment data is performed, including data cleaning, feature extraction, and standardization, to prepare for subsequent model training. Distributed computing frameworks such as Spark can be used to process large-scale data in parallel. Federated learning frameworks, such as FATE or PySyft, are used to train deep learning models such as CNN and LSTM in parallel on various nodes. Collaborative optimization is achieved by sharing encrypted model parameters instead of raw data, protecting data privacy while improving model performance. Homomorphic encryption is used during parameter aggregation to ensure that the central server cannot access the nodes' privacy information. The trained local models are uploaded to the central server, and ensemble learning algorithms such as weighted averaging and voting are used to fuse the prediction results of multiple models to form the final sentiment discrimination model. Weights can be dynamically adjusted based on factors such as the amount of data on each node and model performance. The optimized sentiment discrimination model is deployed for real-time distributed inference on new sentiment data. Based on the heterogeneous characteristics of the data, such as text, voice, and images, corresponding scheduling strategies are designed to dynamically allocate data to different computing nodes for parallel processing, thereby improving the system's throughput and response speed. The inference results return the emotion category and its confidence level, which can be used for applications such as emotion statistical analysis and abnormal emotion detection and early warning. Simultaneously, the inference data and results are recorded in a distributed database to continuously optimize the model and achieve adaptive evolution of emotion discrimination.
[0030] Specifically, in modern emotion recognition systems, using non-relational databases like MongoDB and distributed file systems like HDFS is a suitable choice for processing and storing large amounts of heterogeneous data, such as text, speech, and video. For example, MongoDB can flexibly handle structured and unstructured data, making it suitable for storing text and user behavior data, while HDFS can effectively store and process large-scale video and audio files. This data storage method not only improves data access efficiency but also facilitates subsequent data processing and analysis. In the data preprocessing stage, using distributed computing frameworks like Spark can process large amounts of data in parallel, significantly improving processing speed. For example, frame extraction and preliminary facial recognition can be performed on video data, noise reduction and feature extraction can be performed on speech data, and part-of-speech tagging and sentiment analysis can be performed on text data. These operations can be distributed across multiple nodes on Spark and executed in parallel, greatly shortening data processing time. Using federated learning frameworks, such as FATE or PySyft, deep learning models, such as CNN and LSTM, can be trained in parallel on various nodes without sharing the original data. This method ensures data privacy through encryption technology. For example, each participating node only needs to upload encrypted model parameters to the central server for aggregation, instead of the raw data. This ensures data privacy while leveraging data from all nodes to improve the overall performance of the model. During the model aggregation phase, homomorphic encryption is used to ensure that the central server cannot obtain detailed data information from any node during the aggregation process. For example, the encrypted model parameters uploaded by each node can be calculated and updated without decryption, so even if the central server is attacked, attackers cannot obtain any user data. Finally, the optimized sentiment discrimination model is deployed on the server for real-time sentiment inference. Based on the heterogeneous nature of the data, corresponding scheduling strategies are designed to dynamically allocate computing tasks to different computing nodes. For example, for sentiment analysis tasks on video data, priority can be given to nodes with more GPU resources to speed up video processing; while text and voice data, due to different processing requirements, can be assigned to nodes specifically designed for these tasks. This dynamic scheduling strategy effectively improves the overall response speed and throughput of the system. The sentiment categories and their confidence levels in the inference results can be used for further sentiment statistical analysis and abnormal sentiment detection and early warning. For example, if the system detects that a user has been displaying sadness for several consecutive days, the system can automatically trigger an alarm to notify relevant personnel for attention and intervention. At the same time, this inference data and results will also be recorded in a distributed database for continuous model optimization and adaptive evolution.
[0031] S106. An incremental learning mechanism is introduced to enable the model to continuously learn new emotion data. During incremental learning, a combination of elastic weight merging and knowledge distillation is used to effectively integrate information from new data while retaining previously learned knowledge. A monitoring mechanism for emotion discrimination is constructed to monitor the quality of the discrimination results in real time. When the discrimination accuracy falls below a preset threshold, the model's retraining and optimization process is automatically triggered. User feedback is collected, and samples with incorrect discrimination are analyzed to continuously iterate and optimize the emotion discrimination model. By dynamically adjusting the learning rate and regularization strength, the model's adaptability to both new and old data is balanced.
[0032] Acquire incremental sentiment data, perform preprocessing such as cleaning and labeling, and extract textual features. Employ an elastic weight merging method to combine the feature weights of the new data with the weights of the original sentiment discrimination model, resulting in a merged model weight. Use the merged model as the teacher model to train a student model adapted to the new data using knowledge distillation techniques. The student model can use lightweight LSTM or GRU networks, implemented in TensorFlow or PyTorch frameworks. Utilize the student model to construct a sentiment discrimination monitoring mechanism, performing real-time sentiment discrimination on the new data and calculating the discrimination accuracy. If the accuracy falls below a preset threshold (e.g., 8), trigger the model retraining process. Obtain user feedback on the sentiment discrimination results, label incorrectly judged samples, and merge them with the incremental data to form a new training dataset. Using the new training dataset, dynamically adjust the learning rate and regularization strength using methods such as grid search to fine-tune and optimize the student model. Evaluate the performance of the optimized student model on the test set, calculating evaluation metrics such as accuracy, precision, and recall. If the expected results are met (e.g., an accuracy improvement of more than 2%), the sentiment discrimination model of the online service will be updated to the optimized student model. User feedback and new sentiment data will be continuously collected, and an incremental learning process will be triggered regularly (e.g., weekly) to continuously optimize and iterate the sentiment discrimination model to adapt to the ever-changing data distribution and user needs.
[0033] Specifically, in modern emotion recognition systems, processing and updating models to adapt to new data is a key challenge. First, incremental emotion data is acquired, which may originate from social media, customer service records, or online interactive platforms. Cleaning and labeling this data are necessary steps to ensure its quality and usability. For example, cleaning removes irrelevant information such as URLs and advertisements, while labeling may involve emotion classification, such as categorizing text as "happy" or "sad." After data preprocessing, feature extraction follows, typically involving natural language processing techniques for text data, such as using TF-IDF or Word2Vec models to transform text data into machine-processable numerical features. These features are then used to train the emotion discrimination model. Regarding model updates, a flexible weight merging approach is an effective strategy. This method allows the feature weights of old and new data to be dynamically adjusted based on their importance, enabling the model to better adapt to new data distributions. For example, if new data shows a change in the expression of "anxiety," the model can more accurately identify this change by adjusting the weights of relevant features. The merged model is then used as a teacher model, and knowledge distillation techniques are employed to train a new student model. Knowledge distillation is a model compression technique that allows a small model (student) to learn the output of a large model (teacher), enabling the small model to achieve near-large model performance. Lightweight network architectures such as LSTM or GRU can be used in this process; these architectures are suitable for processing sequence data, such as text and speech data, and are computationally efficient. After model deployment, real-time monitoring of the model's emotion discrimination accuracy becomes possible. If the accuracy falls below a preset threshold, such as below 80%, a retraining process is triggered. This mechanism ensures that the model continuously provides high-quality emotion discrimination services. To further optimize the model, collecting user feedback on the emotion discrimination results is crucial. This feedback helps identify samples where the model misclassifies emotions, which can then be used for retraining, continuously improving the model's performance. By dynamically adjusting the learning rate and regularization strength, such as using grid search methods, the student model can be fine-tuned on new training datasets. Finally, the performance improvement can be verified by evaluating the optimized student model on a test set. If the test results show an accuracy improvement of more than 2%, this indicates that the model update is successful, and the model can be deployed to the online service. In addition, continuously collecting user feedback and new sentiment data, and regularly triggering incremental learning processes are key strategies to ensure that the sentiment discrimination model can adapt to the ever-changing data distribution and user needs.
[0034] S107. By obtaining user feedback on the emotion discrimination results, the samples with incorrect discrimination are labeled and merged with incremental data to form a new training dataset. The student model is fine-tuned and optimized based on the new training dataset to obtain the optimized emotion discrimination model. User feedback and new emotion data are continuously collected, and the incremental learning process is triggered regularly to continuously iterate and optimize the model to adapt to the changing data distribution and user needs.
[0035] The process involves acquiring user feedback data on emotion assessment results, preprocessing and filtering the feedback data to remove invalid and duplicate feedback, manually annotating the filtered feedback data to obtain a high-quality emotion-annotated dataset, merging the emotion-annotated dataset with incremental emotion data to form a new emotion training dataset, and fine-tuning the original LSTM emotion assessment model using gradient descent with appropriate settings for hyperparameters such as learning rate and batch size to obtain an optimized emotion assessment model. The process continues to collect user feedback data and new emotion data, triggering an incremental learning process when the data volume reaches a preset threshold or at fixed time intervals. In the incremental learning process, newly collected user feedback data is preprocessed, filtered, and annotated with emotion to obtain incremental emotion-annotated data. This incremental emotion-annotated data is merged with the original training dataset, and knowledge distillation is used to balance the weights of the old and new data to form a new emotion training dataset for the next round of model optimization. Through continuous incremental learning and iterative model optimization, the emotion assessment model adapts to changing data distributions and user needs, improving the accuracy of emotion assessment and user satisfaction.
[0036] To enhance the visitor experience and ensure a safe and orderly tour route, this invention provides a method that combines ground markings with drone voice-controlled guidance to help visitors move from one location (point A) to another (point B). The specific implementation is as follows: Ground markings: Lay out clear footprint patterns or other forms of directional signs between key points in the scenic area. These signs clearly indicate the direction that tourists should follow. Drone voice-controlled guidance: When tourists arrive at point A, the drone pre-positioned there takes off, playing a voice message through its built-in speaker to inform tourists of the next destination (point B) and instruct them to follow the footprints on the ground. Simultaneously, the drone maintains a certain altitude, continuously monitoring the position of the tourist group and, if necessary, providing audio reminders to tourists to maintain formation or pay attention to their footing.
[0037] To enhance interactivity and fun, this invention also includes a special shooting trigger mechanism: the drone will only automatically activate its camera to take a picture when a visitor makes a specific gesture (such as "OK" or a thumbs-up). This design not only captures the visitor's most natural facial expressions but also adds to the enjoyment of the experience, specifically including: Gesture recognition: The drone is equipped with a high-resolution camera and advanced image processing algorithms, which can monitor tourists' movements in real time. Once a gesture that meets preset conditions is detected, such as an "OK" gesture or a thumbs-up, the camera will be activated immediately. Feedback confirmation: After successfully recognizing the gesture, the drone can also confirm to tourists that filming is about to begin by flashing LED lights or making a sound again, making the whole process more user-friendly and intuitive.
[0038] Considering that some scenic spots may offer personalized souvenir customization services, such as photo printing, this invention further integrates an order recognition function. When multiple tourists appear in the same scene simultaneously, the drone can accurately identify individuals who have pre-ordered such services online, specifically including: Facial recognition: A facial recognition model trained using deep learning technology enables drones to quickly locate target customers in crowds. It maintains high recognition accuracy even in complex background environments. Order matching: Once the specific tourist's identity is confirmed, the drone can provide them with exclusive service options based on the records in the background database, such as directly sending the photos taken to the tourist's mobile phone, or arranging subsequent physical delivery processes.
[0039] Specifically, in modern emotion recognition systems, user feedback data on emotion discrimination results is a key resource for improving model performance. First, this feedback data undergoes preprocessing and filtering to remove invalid and duplicate feedback. For example, if a user submits the same emotion feedback multiple times, the system will only retain the first submission. This step ensures data quality and uniqueness, laying a solid foundation for subsequent data processing. Next, the filtered feedback data undergoes manual emotion labeling. This process may involve multiple labeling experts classifying the emotions in the text content, such as labeling the text as "anger" or "joy." This method yields high-quality emotion-labeled datasets, which directly impact the effectiveness and accuracy of model training. The emotion-labeled dataset is then merged with incremental emotion data to form a new emotion training dataset. For example, if the original training set mainly contains emotional expressions from social media, while the new feedback dataset includes emotional expressions from customer service conversations, merging these two datasets allows the model to more comprehensively understand and recognize emotional expressions from different sources. The next crucial step is to fine-tune the existing LSTM emotion discrimination model using gradient descent. During this process, different learning rates and batch sizes can be set to observe changes in model performance. For example, smaller batch sizes may allow for more frequent weight updates during model training, helping the model capture more subtle emotional changes, but they can also lead to increased fluctuations in the training process. Continuously collecting user feedback data on emotion discrimination results and new emotion data is crucial for ensuring continuous model optimization. When the data volume reaches a preset threshold or at fixed time intervals, the system will automatically trigger an incremental learning process. This periodic model update mechanism helps the model adapt to possible changes in user emotional expression, such as the emergence of new popular phrases or the impact of social events. In the incremental learning process, newly collected user feedback data, after preprocessing, filtering, and emotion labeling, forms incremental emotion-labeled data. This data is then merged with the original training dataset, and knowledge distillation is used to balance the weights of the old and new data. This technique allows a small model (student) to learn the output of a large model (teacher), enabling the small model to achieve performance close to that of the large model while better adapting to new data distributions. Through this series of steps, the emotion discrimination model can continuously adapt to changing data distributions and user needs, improving the accuracy of emotion discrimination and user satisfaction. This continuous process of model iteration and optimization ensures that the emotion recognition system can play its maximum role in practical applications, helping businesses and organizations to better understand user emotions and thus provide more humanized services.
[0040] The above description is merely a specific implementation of this specification. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the scope of protection of this specification is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this specification, and these modifications or substitutions should all be covered within the scope of protection of this specification.
Claims
1. A shared drone travel photography interactive method, characterized in that, include: The user's posture and angle information is acquired through multiple sensors and preprocessed. Deep learning algorithms are used to extract facial features from users, and image enhancement techniques are used to improve the robustness of the extraction. By combining the user's posture, angle, and facial features, a comprehensive feature representation is obtained through feature fusion. Utilize comprehensive feature representations for user identification and behavior analysis to construct comprehensive user profiles; The user's voice and body posture information are acquired and input into a multimodal fusion network along with facial feature vectors to achieve joint representation of features from different modalities. The feature vector after multimodal fusion is input into the emotion discrimination model to obtain the classification result of the user's emotion.
2. The shared drone travel photography interaction method as described in claim 1, characterized in that, The process of acquiring and preprocessing the user's posture and angle information through multiple sensors includes: Kalman filtering is used to remove noise and outliers, ensuring the accuracy of attitude and angle data.
3. The shared drone travel photography interaction method as described in claim 1, characterized in that, The process of extracting facial features from users using deep learning algorithms and improving the robustness of the extraction through image enhancement techniques includes: Histogram equalization and contrast enhancement methods are used to reduce noise interference caused by changes in outdoor ambient light.
4. The shared drone travel photography interaction method as described in claim 1, characterized in that, The process of combining the user's posture, angle, and facial features to obtain a comprehensive feature representation through feature fusion includes: Principal component analysis (PCA) is used for dimensionality reduction to remove redundant and irrelevant features, resulting in compact and discriminative feature vectors.
5. The shared drone travel photography interaction method as described in claim 1, characterized in that, The method of using comprehensive feature representation for user identification and behavior analysis to construct a comprehensive user profile includes: The Support Vector Machine (SVM) and Logistic Regression (LR) algorithms are used to accurately identify user identity and behavioral intent.
6. The shared drone travel photography interaction method as described in claim 1, characterized in that, The process of acquiring the user's voice and body language information, and inputting it along with facial feature vectors into a multimodal fusion network to achieve joint representation of different modal features includes: Multi-head attention is used to calculate the importance weights of features in different modalities.
7. The shared drone travel photography interaction method as described in claim 1, characterized in that, The process of inputting the multimodal fused feature vector into the emotion discrimination model to obtain the user's emotion classification result includes: We employ a bidirectional long short-term memory network (Bi-LSTM) and an attention mechanism to capture the dynamic changes in user emotions, and optimize the model using a cross-entropy loss function.
8. A shared drone travel photography interactive system, based on the shared drone travel photography interactive method described in any one of claims 1-7, characterized in that: It includes a data acquisition and preprocessing module, a facial feature extraction and enhancement module, a feature fusion module, a user identification and behavior analysis module, a multimodal information fusion module, and an emotion discrimination module; The data acquisition and preprocessing module is used to acquire the user's posture and angle information through multiple sensors and perform preprocessing. The facial feature extraction and enhancement module is used to extract the user's facial features using deep learning algorithms and to improve the robustness of the extraction through image enhancement technology. The feature fusion module is used to combine the user's posture, angle, and facial features to obtain a comprehensive feature representation through feature fusion. The user identification and behavior analysis module is used to identify and analyze users using comprehensive feature representations, and to build a comprehensive user profile. The multimodal information fusion module is used to acquire the user's voice and body posture information, which, together with the facial feature vector, is input into the multimodal fusion network to achieve joint representation of different modal features; The emotion discrimination module is used to input the feature vector after multimodal fusion into the emotion discrimination model to obtain the classification result of the user's emotion.
9. A computer device, comprising: Memory and processor; The memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the shared drone travel photography interaction method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the shared drone travel photography interaction method as described in any one of claims 1 to 7.