Voice interaction method and device of eVTOL, equipment and medium

By collecting user voice data in real time in eVTOL and combining it with multimodal analysis of image and location information, the problem of limited user interaction in existing technologies is solved, enabling more comprehensive voice responses and information provision, thus improving user experience and security.

CN121963704APending Publication Date: 2026-05-01GUANGDONG GAOYU TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG GAOYU TECHNOLOGY CO LTD
Filing Date
2025-12-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing eVTOL voice interaction technology cannot meet the diverse needs of users and the complexity of flight scenarios. Users are unable to ask open-ended questions or perform complex information queries, resulting in a reduced user experience.

Method used

By collecting user voice information in real time in the flight equipment, detecting whether it contains the target question, acquiring the image of the environment below using the downward acquisition device, combining the multimodal analysis model with the position input of the flight equipment, outputting the target description and performing audio conversion to respond with voice.

Benefits of technology

It enables comprehensive analysis of user voice information to provide accurate voice responses, thereby improving user experience and flight safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963704A_ABST
    Figure CN121963704A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice interaction, in particular to an eVTOL voice interaction method and device, equipment and a medium. According to the method, user voice information is collected in real time in flight equipment, whether the user voice information contains a target problem or not is detected, if it is detected that the user voice information contains the target problem, downward collection equipment on the flight equipment is used for collecting the environment below, a target image is obtained, and audio conversion is conducted on target description; and obtaining a reply voice, and playing the reply voice. User voice is collected in real time through flight equipment, after a target problem in the user voice is detected, a target image of the lower environment is obtained through downward collection equipment, a multi-modal analysis model is input in combination with the current position of the flight equipment, and target description is output and converted into reply voice to be played. Therefore, the semantic result of the user voice information is comprehensively analyzed, and voice reply is accurately carried out according to the semantic result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice interaction technology, and in particular to a voice interaction method, apparatus, device and medium for eVTOL. Background Technology

[0002] With the increasing severity of urban traffic congestion and the growing demand for efficient and environmentally friendly travel, electric vertical takeoff and landing (eVTOL) aircraft, as an emerging mode of transportation, are gradually becoming an important solution for future urban air mobility. In the use of eVTOL, a good human-machine interaction experience is crucial for improving user satisfaction and flight safety. Existing eVTOL voice interaction technologies have many shortcomings and cannot meet the diverse needs of users and the complexity of flight scenarios.

[0003] Currently, most eVTOL systems use simple voice command response technology, allowing users to interact with the aircraft only through preset, fixed commands such as "take off," "land," "turn left," and "turn right." This technology offers very limited interaction options, preventing users from asking open-ended questions or performing complex information queries. When users want to understand the environmental information below the flight path or inquire about facilities near their destination, the technology fails to provide an effective response, severely restricting the user's ability to obtain information and degrading the user experience.

[0004] Therefore, how to comprehensively analyze the semantic results of user voice information in order to accurately respond to users' voice messages has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of this application provide a voice interaction method, apparatus, device, and medium for eVTOL to solve the problem of how to comprehensively analyze the semantic results of user voice information in order to accurately provide voice responses based on the semantic results.

[0006] In a first aspect, embodiments of this application provide a voice interaction method for eVTOL, including: The system collects user voice information in real time within the flight equipment and detects whether the user voice information contains the target question. If the user's voice information contains the target question, the downward acquisition device on the flight equipment is used to acquire the image of the environment below to obtain the target image. The current position of the flight equipment is obtained, and the current position and the target image are input into a trained multimodal analysis model to output the corresponding target description. The target description is converted into audio to obtain a response voice, which is then played.

[0007] Secondly, an embodiment of this application provides a voice interaction device for eVTOL, comprising: The information detection module is used to collect user voice information in real time in the flight equipment and detect whether the user voice information contains the target question; The image acquisition module is used to acquire images of the environment below using the downward acquisition device on the flight equipment if the user's voice information contains the target question, thereby obtaining a target image. The model analysis module is used to obtain the current position of the flight equipment, input the current position and the target image into the trained multimodal analysis model, and output the corresponding target description; An audio conversion module is used to convert the target description into audio, obtain a response voice, and play the response voice.

[0008] Thirdly, embodiments of this application provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the eVTOL voice interaction method as described in the first aspect.

[0009] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the eVTOL voice interaction method as described in the first aspect.

[0010] The beneficial effects of the embodiments in this application compared with the prior art are: This application involves real-time acquisition of user voice information within a flight-enabled device. The device detects whether the user voice information contains a target question. If a target question is detected, a downward-facing acquisition device on the flight-enabled device captures images of the surrounding environment to obtain the target image. The current position of the flight-enabled device is then acquired. The current position and the target image are input into a trained multimodal analysis model, which outputs a corresponding target description. This target description is then converted into audio to obtain a response voice, which is then played. By acquiring user voice information in real-time and detecting a target question, the downward-facing acquisition device obtains a target image of the surrounding environment. This image is then combined with the current position of the flight-enabled device and input into a multimodal analysis model to output a target description, which is then converted into a response voice for playback. This comprehensive analysis of the semantic results of the user voice information allows for accurate voice responses based on these results. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram of an application environment for a voice interaction method for eVTOL provided in Embodiment 1 of this application; Figure 2 This is a flowchart illustrating a voice interaction method for eVTOL provided in Embodiment 2 of this application; Figure 3 This is a flowchart illustrating a voice interaction method for eVTOL provided in Embodiment 3 of this application; Figure 4 This is a flowchart illustrating a voice interaction method for eVTOL provided in Embodiment 4 of this application; Figure 5 This is a flowchart illustrating a voice interaction method for eVTOL provided in Embodiment 5 of this application; Figure 6 This is a schematic diagram of the structure of an eVTOL voice interaction device provided in Embodiment Six of this application; Figure 7 This is a schematic diagram of the structure of a computer device provided in Embodiment 7 of this application. Detailed Implementation

[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0014] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0015] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0016] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0017] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0018] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0019] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0020] To illustrate the technical solution of this application, specific embodiments are described below.

[0021] The first embodiment of this application provides a voice interaction method for eVTOL, which can be applied to applications such as... Figure 1 In this application environment, the client and server connect and communicate. Users can provide eVTOL voice interaction conditions, requirements, and operation instructions through the client. The server is used to control the eVTOL voice interaction method according to the control instructions sent by the client.

[0022] The client side includes, but is not limited to, PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server side can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0023] See Figure 2 This is a flowchart illustrating a voice interaction method for eVTOL provided in Embodiment 2 of this application. The above-described voice interaction method for eVTOL can be applied to... Figure 1 The server-side component.

[0024] like Figure 2 As shown, the voice interaction method of this eVTOL may include the following steps: Step S201: Collect user voice information in real time in the flight equipment and detect whether the user voice information contains the target question.

[0025] Optionally, detecting whether the user's voice information contains the target question may include the following steps: The user's voice information is input into a preset text conversion model to obtain the text information corresponding to the user's voice. The text information is input into a preset classification model for analysis to obtain text semantics, and the text semantics are analyzed to determine whether they contain the target question.

[0026] Optionally, after detecting whether the user's voice information contains the target question, the following steps may be included: If the user's voice information does not contain the target question, then the user's voice information is analyzed using a preset language response model, and a response is output. The reply content is converted into audio to obtain a reply voice, and the reply voice is played.

[0027] The role of the text-to-text (TTC) model is to convert speech signals into text information that computers can understand and process. In practical applications, speech recognition technologies are used, such as end-to-end speech recognition models based on deep learning, like DeepSpeech and Wav2Vec. These models, trained on large amounts of speech data, can accurately identify the text content within speech. For example, if a user says, "Are there any rivers nearby?", the flying device collects this speech, inputs it into the TTC model, and the model outputs the text information "Are there any rivers nearby?".

[0028] The primary task of classification models is to perform in-depth semantic analysis of textual information. They classify and understand input text based on pre-defined rules or patterns trained on large amounts of data. For example, machine learning-based text classification algorithms such as Support Vector Machines (SVM) and Naive Bayes classifiers, or more advanced deep learning-based pre-trained language models such as BERT, can be used. These models can extract key information from the text, understand its true meaning, and thus determine whether it contains the target question. For instance, inputting the question "Are there any rivers nearby?" into a classification model will, if the target question is about nearby natural water bodies, allow the model to identify that the sentence semantically contains the target question.

[0029] If the user's voice message does not contain the target question, a pre-defined language response model is used to analyze the user's voice message and output a response. When the classification model determines that the user's voice message does not contain the target question, the flight device will call the pre-defined language response model. This model can be rule-based or machine learning-based. Rule-based models generate corresponding responses based on a series of pre-written rules and the input question. Machine learning-based models are trained on a large amount of question-and-answer data to learn the mapping relationship between questions and responses, thereby generating appropriate response content. For example, if a user says "How's the weather today?", while the target question is about the environment below, the classification model will determine that the voice message does not contain the target question. In this case, the language response model will analyze the sentence and, based on its own knowledge base or pre-trained content, output a response such as "I can't get real-time weather information; you can check a weather forecast app."

[0030] The reply is converted into audio, resulting in a spoken response. This step transforms the text-based reply into a speech signal, allowing the user to receive feedback from the device through hearing. Audio conversion typically uses Text-to-Speech (TTS) technology, which generates a corresponding speech waveform based on the input text. Modern TTS technology can achieve very natural and fluent speech synthesis, and can even simulate different timbres and intonations. For example, the reply "I can't get real-time weather information, you can check a weather forecast app" can be converted into speech using TTS technology, and the flight device's voice playback module will play this speech for the user to hear.

[0031] Step S202: If the user voice information contains the target question, the downward acquisition device on the flight equipment is used to acquire the environment below to obtain the target image.

[0032] The target problem is related to the specific conditions of the environment below the flight equipment, such as terrain, the presence of specific objects, etc. Downward acquisition equipment refers to the downward acquisition devices equipped on the flight equipment. Downward acquisition devices can be cameras, which have specific optical and imaging capabilities to adapt to different flight environments and lighting conditions. For example, a camera with a wide-angle lens can be used to obtain a wider downward field of view, or a high-resolution camera can be equipped to clearly capture details of the downward environment.

[0033] Downward acquisition equipment captures images of the environment directly below the flying equipment, capturing the scene below in real time and converting it into digital image signals to provide raw data for subsequent analysis and processing.

[0034] The data acquisition process is real-time; once a target problem is detected, the camera immediately begins working, quickly capturing images of the surrounding environment. This ensures that the acquired images reflect the actual situation at the current moment, providing users with the latest and most accurate information.

[0035] The camera focuses on the environment below through an optical lens, uses an image sensor to convert light signals into electrical signals, and then, after a series of processing and encoding steps, finally obtains a digital image of the target that can be stored and transmitted. This image contains various features and details of the environment below, such as the color of the land and the shape of objects.

[0036] Step S203: Obtain the current position of the flight equipment, input the current position and the target image into the trained multimodal analysis model, and output the corresponding target description.

[0037] The precise location of the flight equipment is obtained by using the Global Positioning System (GPS). Furthermore, other positioning technologies, such as Inertial Navigation Systems (INS), can be combined to improve the accuracy and reliability of positioning, especially when GPS signals are interfered with or blocked.

[0038] The current location information of the flight equipment can help multimodal analysis models better understand the actual scene represented by the target image. For example, the environmental characteristics of different geographical locations may vary greatly. Knowing the location information, the model can perform more accurate analysis of the image based on common geographical features, topography, and other characteristics of the region.

[0039] A multimodal analysis model refers to a model that simultaneously processes multiple data types, in this case, location information and image information. This integrated processing approach fully leverages the advantages of different data sources, providing more comprehensive and accurate analysis results than a single data source. Multimodal analysis models are trained using a large amount of labeled data. This labeled data includes images from different locations and their corresponding environmental descriptions. During training, the model learns the correlation between location information and image features, and how to generate accurate target descriptions based on this information. For example, the model learns that in an image of a specific geographic location, a particular combination of colors and shapes may represent specific land features (such as mountains, farmland, etc.).

[0040] When the current location of the flight equipment and the target image are input into the trained model, the multimodal analysis model first extracts various features from the image, such as color, shape, and texture, while combining them with the geographical background knowledge provided by the location information. Then, the model uses this information to reason and analyze, identify various objects and scene elements in the image, and generate corresponding descriptions.

[0041] The target description is a detailed description of the environment below the flying equipment. It can include terrain features (such as mountains, plains, rivers, etc.), land feature types (such as buildings, forests, farmland, etc.), object characteristics (such as size, color, shape, etc.), and other relevant information. For example, the model might output "Below the current location is a green farmland with a winding stream in the middle, and several small buildings are distributed around the farmland."

[0042] The accuracy of the objective description depends on the performance of the multimodal analysis model and the quality of the input data. An accurate description provides users with valuable information, helping them better understand the environment and make appropriate decisions.

[0043] Step S204: Perform audio conversion on the target description to obtain the response voice, and play the response voice.

[0044] Audio conversion primarily relies on Text-to-Speech (TTS) technology. The core of TTS is converting input text information into natural and fluent speech signals. It comprises three main steps: text analysis, prosodic processing, and speech synthesis. The text analysis stage performs grammatical and semantic analysis on the input target description, determining the pronunciation and intonation of each word. Prosodic processing adds appropriate pauses, stresses, and other prosodic features to the text based on language habits and context. The speech synthesis stage generates the corresponding speech waveform based on the preceding processing results. In practical applications, there are various TTS implementation methods. Rule-based methods can be used to convert text into speech according to predefined speech rules, but this method has poor flexibility and relatively low naturalness of the speech. Currently, methods based on statistical models (such as Hidden Markov Models) and deep learning models (such as Tacotron and Transformer-TS) are more commonly used. These methods can learn from large amounts of speech data to generate more natural and fluent speech effects.

[0045] After audio conversion, the system encodes and processes the digitally represented voice signal to produce a playable response voice file. This file typically uses common audio formats such as MP3 and WAV for easy playback within the flight equipment's audio playback module.

[0046] The flight equipment is equipped with a dedicated audio playback module, such as a speaker. This module is responsible for receiving and playing the converted response audio. It needs to have good sound quality and volume adjustment functions to ensure that users can clearly hear the response audio even under various environmental noise conditions.

[0047] To enhance the user experience, optimizations can be made to the playback of reply voice messages based on actual conditions. For example, the volume can be automatically adjusted according to the noise level of the flight equipment, and a prompt tone can be added before the voice message plays to remind the user that they are about to receive a reply.

[0048] This application involves real-time acquisition of user voice information within a flight-enabled device. The device detects whether the user voice information contains a target question. If a target question is detected, a downward-facing acquisition device on the flight-enabled device captures images of the surrounding environment to obtain the target image. The current position of the flight-enabled device is then acquired. The current position and the target image are input into a trained multimodal analysis model, which outputs a corresponding target description. This target description is then converted into audio to obtain a response voice, which is then played. By acquiring user voice information in real-time and detecting a target question, the downward-facing acquisition device obtains a target image of the surrounding environment. This image is then combined with the current position of the flight-enabled device and input into a multimodal analysis model to output a target description, which is then converted into a response voice for playback. This comprehensive analysis of the semantic results of the user voice information allows for accurate voice responses based on these results.

[0049] See Figure 3 This is a flowchart illustrating a voice interaction method for eVTOL provided in Embodiment 3 of this application. Figure 3 As shown, step S203, which involves obtaining the current position of the flight equipment, inputting the current position and the target image into a trained multimodal analysis model, and outputting a corresponding target description, may include the following steps: Step S301: Input the target image into a preset image recognition model for recognition to obtain an image description; Step S302: Convert the image description and the current position into text to obtain the converted content. Input the converted content into a preset semantic response model and output the corresponding target description.

[0050] The pre-defined image recognition model can be based on deep learning techniques, such as convolutional neural networks (CNNs). It is trained on a large amount of image data to learn the feature patterns of different objects and scenes. For example, during training, the model learns the characteristic representations of objects such as trees, rivers, and buildings in images, including their color, shape, and texture. When the target image is input into the image recognition model, the model extracts and analyzes its features. It identifies various objects and scene elements in the image and determines their categories and locations. For example, the image model might identify a green area as a forest, a blue line as a river, and regularly arranged squares as buildings.

[0051] Based on the recognition results, the image recognition model generates a corresponding image description. This image description includes information about the main objects and scenes in the image, such as "the image contains a forest, a river, and several buildings." The level of detail in the description depends on the model's performance and the quality of the training data.

[0052] The image description obtained in step S301 is integrated with the current location information of the flight equipment. The current location information is usually represented in the form of latitude and longitude or an address. During integration, it is converted into a natural language description and combined with the image description. For example, if the current location is "longitude 120.123, latitude 30.456", it will be converted to "at the location of longitude 120.123, latitude 30.456", and then combined with the image description to form "at the location of longitude 120.123, latitude 30.456, the image contains a forest, a river, and several buildings." This is the converted content.

[0053] The primary function of the pre-defined semantic response model is to perform semantic understanding and processing of the input transformed content, and to generate a more detailed, accurate, and natural target description based on its knowledge base and training data. This model can understand the meaning of the input content and, combined with contextual information and common sense, conduct a deeper analysis and description of the environmental situation.

[0054] The semantic response model outputs a corresponding target description based on the transformed input content. This description is more detailed and specific, such as "Below the location at longitude 120.123, latitude 30.456, there is a vast forest with lush trees, a clear river meandering through it, and several uniquely styled buildings scattered along its banks." The target description takes into account location information and various elements in the image, providing users with more comprehensive and vivid environmental information.

[0055] This application embodiment provides users with a detailed and accurate environmental description by combining location and image information.

[0056] See Figure 4 This is a flowchart illustrating a voice interaction method for eVTOL provided in Embodiment 4 of this application. Figure 4 As shown, after inputting the target image into a preset image recognition model for recognition and obtaining an image description in step S301, the following steps may be included: Step S401: Store the image description and the corresponding target image to obtain an image repository.

[0057] Step S402: Return to the step of storing the image description and the corresponding target image to obtain an image repository, until the preset storage conditions are met to obtain an updated image repository.

[0058] Step S403: Traverse the updated image repository, extract image location information from the image description, and sort the images in the updated image repository according to the image location information to obtain a sorted image repository to be recommended.

[0059] Step S404: Generate recommended content for the image library to be recommended based on the current location.

[0060] After identifying the target image and obtaining its description, the image description and the corresponding target image are stored to form an image repository. The main function of this repository is to preserve environmental information acquired during flight for subsequent analysis, retrieval, and use. The image repository provides a rich data foundation for subsequent processing, such as training more accurate models and monitoring environmental changes.

[0061] A database can be used for storage, where image descriptions are stored as text in database records, while the target image is stored as a file on the server's disk or cloud storage, and the storage path of the image file is associated with the database record. This storage method facilitates data management and retrieval.

[0062] This step is a cyclical process, meaning that the storage operation is repeated every time a new image description and target image are obtained, continuously expanding the image repository. Preset storage conditions can take various forms, such as the number of stored images reaching a certain threshold (e.g., stopping after storing 1000 images), the storage time reaching a certain set value (e.g., stopping storage after the flight ends), or the database space reaching its limit. When these conditions are met, the loop stops, and the image repository becomes the updated image repository, containing image information acquired at multiple stages of the flight.

[0063] Iterate through all image descriptions in the updated image repository and extract location-related information. This location information may have already been included in the image descriptions during previous processing.

[0064] Based on the extracted image location information, the images in the updated image repository are sorted. The sorting rule can be based on distance from the current location, with images closer to the current location appearing first. This sorted image repository becomes the sorted image library for recommendation, facilitating subsequent recommendations based on the current location.

[0065] Using the current location of the flight equipment as a reference, image information related to the current location is filtered from the sorted image library to be recommended, and recommended content is generated. Since the image library to be recommended has been sorted by location, images and their descriptions that are closer to the current location can be selected first.

[0066] Recommended content can be a brief description of the selected images, such as "Near your current location, there is a beautiful forest and a clear river," or it can include image thumbnails or links to allow users to view further details. Recommended content provides users with environmental information about their current location, helping them better understand the flight area.

[0067] See Figure 5 This is a flowchart illustrating a voice interaction method for eVTOL provided in Embodiment 5 of this application. Figure 5 As shown, after performing audio conversion on the target description to obtain the response speech and playing the response speech in step S204, the following steps may also be included: Step S501: Acquire facial images of passengers in the flight equipment; perform emotion analysis on the facial images according to a preset emotion classification library to obtain the emotion type of the facial images. Step S502: Determine an emotion coping strategy based on the emotion type.

[0068] Optionally, if the emotion type is a preset abnormal emotion, then the tone corresponding to the abnormal emotion is selected as the playback mode of the reply voice.

[0069] The flight equipment is equipped with image acquisition devices such as cameras, which will activate at appropriate times (such as after playing a response voice) to capture images of the passenger's face. The captured facial images need to clearly and completely capture the passenger's facial features in order to conduct accurate emotion analysis later.

[0070] The predefined emotion classification database is a predefined database containing various emotion types and their corresponding characteristics. Emotion types can include happiness, sadness, anger, surprise, fear, calmness, etc. Each emotion has its unique facial features; for example, happiness may involve an upturned mouth and squinting eyes, while sadness may involve a furrowed brow and downturned mouth. These features are quantified, encoded, and stored in the emotion classification database as a reference standard for emotion analysis.

[0071] The acquired passenger facial images are input into an emotion analysis system, which extracts key facial features such as the shape and position of the eyes, eyebrows, and mouth. These extracted features are then compared and matched with features in a pre-defined emotion classification database. By calculating a similarity score, the system determines the emotion type that best matches the facial image, thus identifying the corresponding emotion type.

[0072] The system offers different response strategies for different emotional states. If the analysis indicates the passenger is in a happy emotional state, the response could be to continue providing positive and interesting information, such as sharing anecdotes about the flight route, to further enhance the passenger's pleasant experience. When a passenger exhibits sadness, the system can play soothing and comforting music or offer encouraging and warm words, such as, "Don't be sad, you can relax during the flight and look at the beautiful scenery outside the window." For anger, the system needs to promptly reassure the passenger, expressing apologies, inquiring about any unmet needs, and resolving any potential issues as soon as possible. If the passenger is surprised, the system can provide more detailed information about what caused the surprise, satisfying the passenger's curiosity. For example, if the surprise is due to seeing a unique view outside the window, the system can explain the characteristics and formation of that view. When fear is detected, the system can play relaxing and calming music while providing safety reminders and reassurances. For passengers in a calm emotional state, the system can continue to provide normal flight information and services, maintaining good communication and interaction.

[0073] If the analyzed emotion type falls under a preset abnormal emotion category, the playback method of the response message will be adjusted. Specifically, the tone of the response message will be selected to match the passenger's emotional state, thereby enhancing communication effectiveness and passenger experience. Once the emotional response plan is determined, the relevant systems on the flight equipment will operate accordingly. If it is a voice message, it will be played through the audio playback module; if it is music, an appropriate track will be selected; if interaction with passengers is required, it will be done through the crew or a smart voice assistant.

[0074] This application addresses different emotions with appropriate solutions to further enhance passenger enjoyment.

[0075] Corresponding to the eVTOL voice interaction method in the above embodiments, Figure 6 This paper shows a structural block diagram of the eVTOL voice interaction device provided in Embodiment Six of this application. The above-described eVTOL voice interaction device can be applied to... Figure 1 For ease of explanation, only the parts of the server-side shown are relevant to the embodiments of this application.

[0076] See Figure 6 The eVTOL voice interaction device includes: The information detection module 61 is used to collect user voice information in real time in the flight equipment and detect whether the user voice information contains the target question; The image acquisition module 62 is used to acquire the target image by using the downward acquisition device on the flight equipment to acquire the environment below if the user's voice information contains the target question; Model analysis module 63 is used to obtain the current position of the flight equipment, input the current position and the target image into the trained multimodal analysis model, and output the corresponding target description; The audio conversion module 64 is used to convert the target description into audio to obtain a response voice and play the response voice.

[0077] Optionally, the information detection module 61 includes: The conversion unit is used to input the user's voice information into a preset text conversion model to obtain text information corresponding to the user's voice. The semantic analysis unit is used to input the text information into a preset classification model for analysis, obtain the text semantics, and analyze whether the text semantics contain the target question.

[0078] Optionally, the voice interaction device includes: The voice analysis module is used to analyze the user voice information using a preset language response model and output response content if the user voice information does not contain the target question after detecting whether the user voice information contains the target question. An audio conversion module is used to convert the reply content into audio, obtain reply speech, and play the reply speech.

[0079] Optionally, the model analysis module 63 includes: The image description unit is used to input the target image into a preset image recognition model for recognition and obtain an image description; The description output unit is used to convert the image description and the current position into text to obtain the converted content, input the converted content into a preset semantic response model, and output the corresponding target description.

[0080] Optionally, the voice interaction device includes: An image storage module is used to store the image description and the corresponding target image after the target image is input into a preset image recognition model for recognition and image description is obtained, thereby obtaining an image storage library. The execution module is returned to the step of storing the image description and the corresponding target image to obtain an image repository, until the preset storage conditions are met and an updated image repository is obtained. The image sorting module is used to traverse the updated image repository, extract image location information from the image description, and sort the images in the updated image repository according to the image location information to obtain a sorted image library to be recommended.

[0081] Optionally, the voice interaction device includes: The emotion analysis module is used to convert the target description into audio to obtain a response voice, play the response voice, acquire the facial image of the passenger in the flight equipment, and perform emotion analysis on the facial image according to a preset emotion classification library to obtain the emotion type of the facial image. The solution determination module is used to determine an emotion coping strategy based on the emotion type.

[0082] Optionally, the sentiment analysis module includes: The voice playback unit is used to select the tone corresponding to the abnormal emotion as the playback mode of the reply voice if the emotion type is a preset abnormal emotion.

[0083] It should be noted that the information interaction and execution process between the above modules, units, and sub-units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0084] Figure 7 This is a schematic diagram of the structure of a computer device provided in Embodiment Seven of this application. Figure 7 As shown, the computer device of this embodiment includes: at least one processor ( Figure 7 Only one is shown in the diagram), a memory, and a computer program stored in the memory and executable on at least one processor, wherein the processor executes the computer program to implement the steps in any of the above-described eVTOL voice interaction method embodiments.

[0085] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 7 The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.

[0086] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0087] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of the computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.

[0088] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0089] The implementation of all or part of the processes in the methods of the above embodiments can also be accomplished by a computer program product. When the computer program product is run on a computer device, it enables the computer device to execute the steps in the above method embodiments.

[0090] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0091] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0092] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0093] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0094] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A voice interaction method for eVTOL, characterized in that, The voice interaction method includes: The system collects user voice information in real time within the flight equipment and detects whether the user voice information contains the target question. If the user's voice information contains the target question, the downward acquisition device on the flight equipment is used to acquire the image of the environment below to obtain the target image. The current position of the flight equipment is obtained, and the current position and the target image are input into a trained multimodal analysis model to output the corresponding target description. The target description is converted into audio to obtain a response voice, which is then played.

2. The eVTOL voice interaction method according to claim 1, characterized in that, The detection of whether the user's voice information contains the target question includes: The user's voice information is input into a preset text conversion model to obtain the text information corresponding to the user's voice. The text information is input into a preset classification model for analysis to obtain text semantics, and the text semantics are analyzed to determine whether they contain the target question.

3. The eVTOL voice interaction method according to claim 1, characterized in that, After detecting whether the user's voice information contains the target question, the process includes: If the user's voice information does not contain the target question, then the user's voice information is analyzed using a preset language response model, and a response is output. The reply content is converted into audio to obtain a reply voice, and the reply voice is played.

4. The eVTOL voice interaction method according to claim 1, characterized in that, The process of obtaining the current position of the flight equipment, inputting the current position and the target image into a trained multimodal analysis model, and outputting a corresponding target description includes: The target image is input into a preset image recognition model for recognition to obtain an image description; The image description and the current location are converted into text to obtain the converted content. The converted content is then input into a preset semantic response model to output the corresponding target description.

5. The eVTOL voice interaction method according to claim 4, characterized in that, After inputting the target image into a preset image recognition model for recognition and obtaining an image description, the process further includes: The image description and the corresponding target image are stored to obtain an image repository; Return to the step of storing the image description and the corresponding target image to obtain an image repository, until the preset storage conditions are met, and obtain an updated image repository; Traverse the updated image repository, extract the image location information from the image description, and sort the images in the updated image repository according to the image location information to obtain the sorted image repository to be recommended. Based on the current location, generate recommended content for the image library to be recommended.

6. The eVTOL voice interaction method according to claim 1, characterized in that, After performing audio conversion on the target description to obtain the response speech and playing the response speech, the method further includes: Acquire facial images of passengers in flight equipment, perform emotion analysis on the facial images according to a preset emotion classification library, and obtain the emotion type of the facial images; Based on the described emotion type, determine the emotion coping strategy.

7. The eVTOL voice interaction method according to claim 6, characterized in that, The step of determining an emotion coping strategy based on the emotion type includes: If the emotion type is a preset abnormal emotion, then the tone corresponding to the abnormal emotion is selected as the playback mode of the reply voice.

8. A voice interaction device for eVTOL, characterized in that, The eVTOL voice interaction device includes: The information detection module is used to collect user voice information in real time in the flight equipment and detect whether the user voice information contains the target question; The image acquisition module is used to acquire images of the environment below using the downward acquisition device on the flight equipment if the user's voice information contains the target question, thereby obtaining a target image. The model analysis module is used to obtain the current position of the flight equipment, input the current position and the target image into the trained multimodal analysis model, and output the corresponding target description; An audio conversion module is used to convert the target description into audio, obtain a response voice, and play the response voice.

9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the eVTOL voice interaction method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the eVTOL voice interaction method as described in any one of claims 1 to 7.