Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

45 results about "Speech generation" patented technology

Speech generation. Speech generation and recognition are used to communicate between humans and machines. Rather than using your hands and eyes, you use your mouth and ears. This is very convenient when your hands and eyes should be doing something else, such as: driving a car, performing surgery, or (unfortunately) firing your weapons at the enemy.

Food fraud detection device using terahertz spectroscopy

This invention provides a food fraud detection device using terahertz spectroscopy that achieves a compact and stable structural configuration. [Solution] The food fraud detection device 100 using terahertz spectroscopy according to the present invention comprises a base housing 20 and an upright housing 21 attached to the base housing 20. A display unit 1 is provided in the upright housing 21, and a user input unit 2 is provided in the base housing 20. The main body of the device 100 is equipped with a wireless connection antenna 7, a voice generation unit 8, a voice input port 81, first to fourth connection ports 31, 32, 33, 34, an external sensor connection port 4, a server connection port 5, and a port 6 for opening the device. An optical sensor 10, an ultrasonic sensor 11, and a chemical sensor 12 are arranged on the base housing 20 together with corresponding control switches 13, 14, 15.
Owner:シャイマ アブデルラオフ モハメド アブデルモフセン +2

Information processing device, information processing method, and information processing program

The present invention provides an information processing device that enables a more natural improvement in the accuracy of identifying past related utterances made by a user or identifying their relationship with other users. [Solution] The system includes: a speech acquisition unit 1 that acquires user utterances; a prompt generation unit 2 that generates prompts for a large-scale language model to generate tag information that includes at least one of the following based on the acquired utterances: the relationships between the constituent elements of the utterance, the attributes to which the utterance belongs, or a summary of the utterance; an interface 3 that transmits the generated prompts to the large-scale language model and receives response information that includes tag information generated by the large-scale language model in response to the transmitted prompts; and a recording unit 5 that records the acquired utterances and the received tag information in association.
Owner:PIONEER IP

Method and system for providing video conference service including artificial intelligence-based speech interpretation function

PCT designated stageWO2026146907A1Speech translationSpeech sound
Disclosed are a method and system for providing a speech conference service including an artificial intelligence-based speech interpretation function. According to one embodiment, the method for providing a video conference service may comprise the steps of: setting two or more interpretation bots participating as virtual participants in a video conference service; interpreting a speech of a first language input by participants of the video conference service into a speech of at least one other language different from the first language through processes of speech recognition, translation, and synthetic speech generation between the two or more interpretation bots and artificial intelligence; and providing the interpreted speech.
Owner:LINE PLUS

Voice generation method and apparatus, product, device, and medium

PCT designated stageWO2026108241A1Speech synthesisSpeech soundTarget text
A voice generation method and apparatus, a product, a device, and a medium, which are applied to the technical field of voice generation. The method comprises: using a quantizer to discretize a voice feature vector of an original voice signal, obtaining a discrete symbol representation corresponding to the original voice signal (S11); extracting a text feature corresponding to a target text (S12); and inputting the text feature and the discrete symbol representation into a voice generation model, so that the voice generation model uses the text feature as a condition, and generates a target voice on the basis of the discrete symbol representation (S13). The quantizer is obtained by training a self-organizing map network by means of voice signal training samples. The method may restore an original voice feature more accurately, and improve the quality of the generated voice.
Owner:SHANGHAI SOULGATE TECH CO LTD

A large model-based speech generation method, device and medium

PendingCN122369426AMultiplexingAcoustics
The application discloses a large model-based speech generation method and device and medium, and belongs to the technical field of data processing. The method comprises the following steps: dividing target language text information into multiple speech generation units through a text division model; in response to the confidence that the target language text information is divided into multiple speech generation units through the text division model being lower than a first threshold value, dividing the target language text information into multiple speech generation units according to text structure features and parameter distribution features; in response to the speech generation units meeting a multiplexing condition, obtaining a first speech output result according to the speech generation units; in response to the speech generation units not meeting the multiplexing condition, calling a speech generation processing object to obtain a second speech output result, and combining the first speech output result and the second speech output result to generate a target speech result. The application can improve the overall processing efficiency, resource utilization rate and result multiplexing capability under a parameterized speech task.
Owner:SHANGHAI ZHONGAN XINKE INFORMATION TECH SERVICES CO LTD

A speech generation method, system, device and medium

PendingCN122313960ASpeech soundAudio frequency
This specification provides a speech generation method, system, apparatus, and medium through several embodiments. The method includes: determining a target intent based on a user audio stream and environmental data; extracting acoustic features from the user audio stream and visual features from the user video stream; determining the user's emotion distribution based on the acoustic and visual features; and generating speech playback parameters and a target speech based on the emotion distribution and the target intent; the speech playback parameters are the playback parameters for the target speech.
Owner:HANSONG NANJING TECH LTD

Electronic device, method, and non-transitory computer-readable storage medium for determining uplink transmission power

This electronic device may comprise at least one processor. The at least one processor is configured to: while an application for providing an interpretation service for a call, which is performed by using a communication circuit, is executed, transmit uplink voice data of a first language generated on the basis of a user's utterance; determine a transmission time interval for transmitting uplink voice data of a second language on the basis of generating the uplink voice data of the second language from the utterance of the first language; determine transmission power for transmitting the uplink voice data of the second language on the basis of the transmission time interval for transmitting the uplink voice data of the second language; and transmit the uplink voice data of the second language on the basis of the transmission power.
Owner:SAMSUNG ELECTRONICS CO LTD

A pet companion voice generation method and device based on voiceprint modeling, equipment and storage medium

PendingCN122337180APersonalizationEngineering
This invention relates to a method, apparatus, device, and storage medium for generating pet companion voice based on voiceprint modeling. The method includes: acquiring the voiceprint features of a target user and constructing a personalized timbre model based on a quality score; acquiring pet behavior data collected by a pet companion terminal, identifying the pet's behavior categories and the intensity of the behavior; determining a basic tone template based on the behavior categories, and dynamically adjusting the speech parameters in the basic tone template according to a preset mapping relationship based on the behavior intensity to obtain target tone parameters; and synthesizing target companion voice based on the target tone parameters, the personalized timbre model, and target text matching the behavior category, and sending it to the pet companion terminal for playback. This invention achieves emotional and adaptive automated pet companionship, and ensures modeling accuracy and usage security through a recording quality scoring mechanism and a sensitive content filtering mechanism.
Owner:QIERLING BEIJING HEALTH TECH CO LTD

An end-to-end speech dialogue system and method based on cross-modal retrieval enhancement

PendingCN122116896ADigital data information retrievalSpeech recognitionCognitive capabilityPhonological memory
The application discloses an end-to-end voice dialogue system and method based on cross-modal retrieval enhancement, comprising a voice input and feature extraction module, a cross-modal retrieval enhancement module, an end-to-end voice generation module and a voice memory module. Through the system and method process setting, direct mapping and generation from voice to voice can be realized, error accumulation problems caused by modular architecture can be eliminated, and the overall robustness of the system is improved. Through the end-to-end model, the training and deployment process is simplified, the scalability and real-time interaction capability of the system are improved, and the system can be better applied to intelligent customer service and other application scenarios with extremely high requirements for accuracy and efficiency. Through the efficient cross-modal retrieval enhancement generation mode, relevant field professional knowledge can be retrieved from a text knowledge base in real time according to received audio, the accuracy and timeliness of the response information are ensured, direct, efficient and accurate retrieval from voice to text knowledge is realized, and the cognitive ability of the system in the vertical field is enhanced.
Owner:CHINA UNICOM WO MUSIC & CULTURE CO LTD

Speech generation method and apparatus, computer device and readable storage medium

The application relates to a speech generation method and device, computer equipment and a readable storage medium. The method comprises the following steps: generating a diffusion sequence according to an original word element corresponding to a to-be-processed object and a target mask quantity; inputting the diffusion sequence into a target network model which has been trained to obtain a first candidate audio word element corresponding to each mask label and a first confidence; determining a to-be-updated mask label in the mask label according to a current inference time step, the first candidate audio word element and the first confidence, updating the to-be-updated mask label, and generating an updated diffusion sequence; using the updated diffusion sequence to return to the step of obtaining the first candidate audio word element corresponding to each mask label and the first confidence until a preset inference time step is reached, generating a target audio according to the first candidate audio word element of each current mask label, thereby effectively reducing the error propagation risk and improving the reliability and accuracy of audio generation.
Owner:BAIRONG ZHIXIN (BEIJING) TECH CO LTD

Speech generation method and device based on semantic adjustment, equipment and medium

PendingCN122313946ASemantic representationSpeech reconstruction
This invention relates to the field of speech synthesis technology, and discloses a speech generation method, apparatus, device, and medium based on semantic adjustment. The method includes: encoding text content and style cues to obtain text semantic representation and style semantic representation; performing semantic consistency analysis based on the two to obtain semantic mismatch results; determining dynamic guidance strength based on the semantic mismatch results, and performing autoregressive generation processing in conjunction with alternative style cues to obtain an acoustically labeled sequence; and reconstructing the acoustically labeled sequence to obtain the target speech. This invention can be applied to business scenarios such as fintech and healthcare. By obtaining semantic mismatch results through semantic consistency analysis and determining dynamic guidance strength based on the semantic mismatch results, the speech generation process can be adjusted according to changes in semantic consistency, thereby reducing the deviation caused by inconsistencies between semantic and emotional expression and improving the expressive coordination of the target speech.
Owner:PING AN TECH (SHENZHEN) CO LTD

An interaction method, apparatus, product, electronic device, and medium

The present disclosure provides an interaction method, device, product, electronic equipment and medium, relating to the technical field of artificial intelligence. The method comprises: receiving input information of a user, generating a control result according to the associated information of the input information; injecting the control result into the decoding process of a speech interaction model used to generate target speech as a constraint condition of the model output; generating target speech matched with the control result based on the speech interaction model, forming a mandatory constraint on the generation direction by injecting the control result through the model, enabling the model to follow the control result when generating target speech, actively regulating the speech generation process, and making the generated target speech matched with the interaction demand of the input information, thereby improving the pertinence and rationality of human-computer interaction.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1

Method for generating a virtual voice

The present application relates to the technical field of speech processing, and more particularly to a virtual speech generation method, comprising: step S1, generating a virtual speech, detecting a sound wave shape of the virtual speech, and calculating a similarity according to the sound wave shape by a central control module; step S2, rating and determining whether the generation of the virtual speech is qualified; step S3, the central control module determines whether to update a test statement when determining that the generation of the virtual speech does not meet the standard, or adjusts the frequency and amplitude of the reverse-phase sound wave in the noise reduction process to a corresponding value; and step S4, the central control module determines whether the generation of the virtual speech is qualified again when determining that the generation of the virtual speech meets the standard, the present application avoids the phenomenon of missing words in the generated virtual speech, improves the quality of the generated virtual speech, and improves the efficiency of virtual speech generation while ensuring the quality of the virtual speech.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Imaging display device and method

An imaging display device and method. Imaging information corresponding to speech information can be obtained by means of speech, and screen imaging is then controlled on the basis of the imaging information; moreover, an imaging display device is provided to refract imaging content by means of a beam-splitting film after the screen imaging, so as to display an image on a regular hexahedron for viewing. Since a formed image is generated on the basis of the speech of a user, the current requirements of the user can be met, and imaging content is enriched.
Owner:BEIJING YIZHI YIDE CULTURE TECHNOLOGY CO LTD

Methods and systems of text-conditioned audio-visual speech generation with multi-modal latent diffusion models

Methods, systems, and computer programs are presented for audio-visual speech generation with multi-modal latent diffusion models. One method includes encoding raw audio signals and video frames into respective latent spaces using audio and visual autoencoders. A text transcript is processed into phoneme sequences using a text transcript processor. The audio and visual latent spaces are conditioned using the text transcript and a conditioning variable. Joint distributions of the visual and audio latent spaces, text transcript, and conditioning variable are learned using a multi-modal latent diffusion model. The model adds noise to the latent audio-visual representations and predicts the noise through denoising neural networks. An inverted diffusion process is utilized to generate diverse speech content and speaker characteristics, resulting in realistic audio-visual speech. The technology presented provides a novel approach to conditional speech generation with potential applications in speech synthesis, voice conversion, and speech recognition.
Owner:TENSORTYPE INC

A humanized voice intervention method for inhalation error of an oral-nasal aerosol dispenser

PendingCN122337174AGuaranteed therapeutic effectRealize humanized voice interventionAerosol drug deliveryPatient suctioning
This invention discloses a humanized voice intervention method for inhalation errors in nasal and oral aerosol delivery devices, aiming to address the lack of humanization in existing voice intervention methods, which leads to patient misunderstanding and resistance. The method includes: establishing a patient-doctor corpus by combining expert experience, literature, and guidelines; fine-tuning a language parsing model to obtain an error-sensitive parsing model; using the error-sensitive parsing model to obtain a total intervention feature library and an urgency level sub-library; establishing a two-level query mechanism for patient inhalation errors to query intervention text in the urgency level sub-library; and fine-tuning the voice generation model based on the aforementioned patient-doctor corpus, incorporating new tokens, and evolving through prompting learning into a patient-friendly intervention voice generation model capable of generating intervention voices that are appropriate for the urgency level. The intervention voices generated by this invention can form a humanized voice intervention that approximates the tone of clinical medical staff, and can adapt the tone according to the urgency level of the error, effectively improving the patient's understanding and acceptance of the intervention suggestions.
Owner:CHILDRENS HOSPITAL OF CHONGQING MEDICAL UNIV

Speech generation method and device based on double-layer style modulation, equipment and medium

The application relates to the technical field of speech synthesis, and discloses a speech generation method and device based on double-layer style modulation, equipment and a medium, which comprises the following steps: receiving an input text and a style description text, standardizing the input text, converting phonemes, and extracting prosodic features; generating phoneme embedding vectors and performing context coding; combining the prosodic features to perform gate modulation to obtain local prosodic features; performing semantic coding on the style description text and decomposing the style description text to obtain local style components and global style components; performing fusion modulation to obtain style modulation features; performing time length prediction and sequence expansion on the style modulation features to obtain expanded features; mapping and decoding the expanded features to obtain acoustic features; and generating a speech waveform based on the acoustic features. The application can be applied to business scenarios such as financial technology and medical health, rhythm features and paralanguage style modeling are separated, fusion modulation and feature constraint processing are performed, and therefore the controllability and stability of speech generation are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Streaming audio generation system

Providing a user-friendly streaming audio generation system. [Solution] The streaming speech generation system 10 includes a speech recognition unit 20 that inputs the speaker's voice as text data to a generation AI 30, a generation AI 30 that generates a response sentence to the voice recognized by the speech recognition unit 20 as text data, a splitting unit 35 that receives the response sentence (text data) output by the generation AI 30 as streaming data, and sequentially divides it into short sentences and outputs them in real time, and a speech synthesis unit 40 that outputs the short sentences output by the splitting unit 35 as speech.
Owner:YANO YUAI OFFICE CO LTD +1

Digital human emotional speech generation method based on emotional semantic modeling

PendingCN122369428AImplementing fine-grained emotion modelingAccurately identify differential contributionsGraph neural networksSemantic feature
This invention relates to the field of artificial intelligence technology, specifically to a digital human emotional speech generation method based on emotion semantic modeling. The method includes multimodal data preprocessing, multimodal semantic feature extraction, emotion semantic unit partitioning, graph-based emotion interaction, speech expression parameter generation, emotional speech generation, and digital human emotion collaborative driving. By constructing emotion semantic units and introducing an emotion contribution evaluation mechanism, fine-grained emotion modeling of the speech semantic structure is achieved, accurately identifying the differential contributions of different semantic segments in a sentence to the overall emotional expression, thereby significantly improving the accuracy and controllability of emotional speech generation. Furthermore, by constructing a dual-graph structure of emotion interaction graph and emotion evolution graph, and combining graph neural networks to jointly model the interaction relationships and evolutionary processes between emotion semantic units, a structured expression of complex dynamic emotional changes is achieved, enhancing the generated speech's ability to exhibit emotional coherence and naturalness.
Owner:BEIJING XILIANLIAN TECHNOLOGY CO LTD

Electronic device for supporting performance of function corresponding to utterance of user by using artificial intelligence model, and operation method thereof

Provided is an electronic device comprising: a microphone; a communication circuit for communicating with an external electronic device; a memory storing instructions; and a processor. The instructions, when executed by the processor, may cause the electronic device to: acquire an utterance of a user through the microphone or the communication circuit; receive, from at least one external electronic device through the communication circuit, information about at least one control command for controlling the at least one external electronic device; generate a first prompt to be input into a first large language model (LLM), on the basis of the received information and the utterance; identify, from an output value that has been output from the first LLM as a result of inputting the first prompt into the first LLM, an update target in the at least one external electronic device and update content of a second prompt to be input into a second LLM related to the update target; and transmit the update content through the communication circuit to an external electronic device configured to generate the second prompt.
Owner:SAMSUNG ELECTRONICS CO LTD

A script semantic-driven multi-digital human collaborative generation method

This invention relates to the field of artificial intelligence technology, specifically to a script semantic-driven multi-digital human collaborative generation method. The method includes script input, script semantic feature encoding, script semantic dual-graph construction, dynamic-static graph joint optimization based on a constraint-driven mechanism, script temporal semantic hierarchical modeling, multi-digital human role behavior strategy generation, retrieval-enhanced digital human speech generation processing, triple consistency constraint fusion, and multimodal scene generation. By constructing a script semantic dual-graph structure containing role nodes, scene nodes, and event nodes, and introducing a joint modeling and optimization mechanism of a global static relationship graph and a local dynamic interaction graph, the accuracy and expressive power of complex script semantic modeling are improved. Furthermore, by introducing a multi-digital human behavior generation mechanism based on role relationship constraints, and combining retrieval-enhanced speech generation with semantic, emotional, and temporal triple consistency constraints, the consistency and naturalness of multi-digital human collaborative expression are enhanced.
Owner:BEIJING XILIANLIAN TECHNOLOGY CO LTD

A controllable speech generation method and system based on emotion trajectory reasoning

This invention discloses a controllable speech generation method and system based on emotion trajectory reasoning. The method includes: an instruction understanding and emotion reasoning step: performing joint semantic parsing on the input speech text content and user instruction information, performing semantic reasoning on the explicit and implicit emotional needs in the user instruction information, constructing a multi-stage emotion trajectory structure of the target speech, and generating prosodic control factors and pronunciation control information corresponding to each emotion stage to form stage-level speech control factors; a controllable speech generation step: inputting the stage-level speech control factors and the speech text content into a conditional speech generation model, generating multiple candidate speech under the constraints of the stage-level speech control factors; and an optimization and verification step: performing multi-dimensional comprehensive evaluation on the multiple candidate speech, optimizing the generation process based on the evaluation results, and selecting the optimal candidate speech as the final output speech.
Owner:SUN YAT SEN UNIV +1

Method and apparatus for constructing text-to-speech model, electronic device, readable medium, and program product

PCT designated stageWO2026138018A1Speech inputAcoustics
A method and apparatus for constructing a text-to-speech model, an electronic device, a readable medium, and a program product. The method comprises: inputting preset training speech into a preset vector quantizer, so as to obtain a training semantic dispersion feature of the training speech, the training semantic dispersion feature comprising a language style of the training speech (101); acquiring training text corresponding to the training speech, and using the training text and the training semantic dispersion feature to train a preset autoregressive speech model, so as to obtain a semantic dispersion feature generation model (102); acquiring a training Mel-frequency spectrogram corresponding to the training semantic dispersion feature (103); using the training semantic dispersion feature and the training Mel-frequency spectrogram to train a preset optimal transport conditional flow matching model, so as to obtain a Mel-frequency spectrogram generation model (104); and on the basis of the Mel-frequency spectrogram generation model and the semantic dispersion feature generation model, constructing a text-to-speech model (105).
Owner:CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD

A method and device for distributed voice interaction in AI-driven offline merchant operations

A method and apparatus for distributed voice interaction in offline AI-driven operations are disclosed. The method includes: acquiring voice data from offline operational scenarios and extracting several candidate voices from the voice data through voiceprint recognition; performing order-specific keyword recognition on the candidate voices to determine the target interactive voice and generating a voiceprint mask based on the target interactive voice; performing noise filtering processing on the target interactive voice using the voiceprint mask to obtain standard voice data; generating response text or business instructions based on the standard voice data and the business model, and sending the response text or business instructions to the corresponding target terminal. The method demonstrates high accuracy and good adaptability in voice instruction recognition for order-taking scenarios in noisy offline commercial environments. Furthermore, by combining cloud-based business models and role-based permission mapping, it achieves accurate distribution and response of instructions, meeting the AI-driven operational interaction needs of offline merchants across multiple roles and scenarios.
Owner:SHENZHEN IBOX INFORMATION TECH CO LTD

Speech generation method based on large model and method for training deep learning model

The disclosure provides a large model-based speech generation method and a method for training a deep learning model, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, speech technology, video production, intelligent customer service, intelligent agents and the like. The large model-based speech generation method comprises: processing input information for a target speech by using a large model to obtain a target knowledge graph, wherein a node in the target knowledge graph represents a character object in the input information, and attribute information of the node represents a character state of the character object; performing information extraction on attribute information of a plurality of nodes related to the character object to obtain state change information, wherein the state change information describes a character state change of the character object; and performing a speech generation task according to the input information and the state change information to generate a target speech.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

External electronic device

The disclosure provides a voice communication system, including a caller voice receiving unit, a voiceprint recognition unit, a voiceprint training unit, a text input unit, a language translation model, a voice generation unit, a receiver voice receiving unit, a voice recognition unit, and a text recording unit. The caller voice receiving unit, the voiceprint recognition unit, and the voiceprint training unit are configured to train a voice synthesis model by using a voice sample of a caller. The text input unit, the language translation model, and the voice generation unit are configured to generate a synthesized voice by using caller text and the voice synthesis model. The receiver voice receiving unit and the voice recognition unit are configured to convert a reply voice into reply text. The text recording unit is configured to record the caller text and the reply text. The disclosure also provides a voice communication method.
Owner:ASUSTEK COMPUTER INC

A speech generation method and a speech generation device based on timbre features

ActiveCN120600002BSpeech soundTimbre
This application provides a speech generation method and apparatus based on timbre features. The speech generation method includes: acquiring a target parsing model; inputting descriptive text into the target parsing model to obtain a target sound feature vector corresponding to the descriptive text; wherein the descriptive text includes a target timbre text description; inputting the target sound feature vector into a target fusion model to generate a target timbre vector; and inputting the target timbre vector and the text to be converted into a speech generation model to obtain target speech that conforms to the descriptive text. Through this method and apparatus, target speech that meets the timbre features required by the user is generated, improving the accuracy of speech generation under a specified timbre and satisfying the user's demand for speech generation with diverse timbres and high-precision emotional expression.
Owner:SHANGHAI XIYU JIZHI TECH CO LTD

A method and system for dynamic emotion-based speech generation based on stream matching and multimodal context

This invention discloses a dynamic emotion-based speech generation method and system based on stream matching and multimodal context, belonging to the field of intelligent speech technology. The method includes: acquiring the text to be synthesized in the current round and context information from N previous rounds of dialogue; obtaining semantic representation vectors and emotion representation vectors using a causal cross-attention mechanism; obtaining the discrete coordinates of each phoneme in the text to be synthesized in the current round in a three-dimensional continuous space based on the emotion representation vectors; generating a phoneme-level VAD emotion trajectory sequence based on the discrete coordinates of adjacent phonemes; inputting the phoneme-level VAD emotion trajectory sequence and the semantic representation vector into a constructed vector field to obtain an emotional acoustic representation; and inputting the emotional acoustic representation into a residual vector quantization decoder to output a speech waveform. This invention eliminates acoustic artifacts, achieves self-consistent emotional logic in multi-turn dialogues, and restores highly realistic micro-emotional details such as breathing and vibrato, making it suitable for real-time interaction scenarios such as intelligent assistants and virtual humans.
Owner:YUANYU HUANYU ARTIFICIAL INTELLIGENCE TECHNOLOGY (SUZHOU) CO LTD

Large model-based full-modal unmanned aerial vehicle post-flood disaster intelligent patrol system and method

The application relates to the technical field of intelligent inspection, and particularly discloses a full-mode unmanned aerial vehicle post-flood disaster intelligent inspection system and method based on a large model. The system adopts a VTOL fixed-wing unmanned aerial vehicle platform, integrates various sensors such as visible light, infrared thermal imaging, LiDAR and a microphone array, realizes all-around and multi-dimensional information collection of the disaster area environment, can obtain richer and more reliable data under day and night, various weather and shielding conditions, can realize unified and deep fusion analysis and understanding of multi-source heterogeneous data through an advanced large multi-modal model MLLM, automatically completes complex tasks such as flood range plotting, infrastructure damage assessment and especially detection of trapped personnel in combination with visual and auditory clues, greatly reduces the dependence on manual interpretation, and simultaneously realizes unprecedented real-time two-way voice interaction between the system and ground personnel by using the natural language processing and voice generation capability of the MLLM, and enhances the human-machine cooperation efficiency and decision support level.
Owner:NORTH CHINA UNIV OF WATER RESOURCES & ELECTRIC POWER