Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

677 results about "Environmental sounds" patented technology

Long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation

The invention discloses a long video multi-modal understanding and question-answering method and system based on large model and retrieval enhancement generation. The method comprises the following steps: 1) a multi-modal feature extraction module; 2) a multi-modal synchronization and alignment mechanism; 3) constructing a structured memory pool; 4) querying a drive generation mechanism; 5) incremental updating and memory compression strategy; and 6) unifying the multi-modal representation space. The invention provides a long video multi-mode understanding method fusing a large language model and retrieval enhancement generation, and aims to break through the limitation of a traditional method in the aspects of single-mode processing and semantic fragmentation. According to the method, video image features are extracted through a visual model (such as YOLO and ViT), voice transcription and environment voice description are obtained in combination with an audio model (such as Whisper and Qwen-Audio), and unified coding of vision, voice and audio in a long video is achieved. Then, a structured memory pool is constructed through semantic consistency segmentation and timestamp alignment technologies to store time slice data of different modalities.
Owner:GUANGZHOU BINGO SOFTWARE +1

Voiceprint recognition method, model construction method and system for authority verification

The invention discloses a voiceprint recognition method and system for authority verification and a model construction method and system, and relates to the technical field of user identification, and the method comprises the steps: sending a random verification instruction; acquiring response voice of the person, extracting entrance identity acoustic features, comparing the entrance identity acoustic features with the identity reference model to obtain a comparison result, and determining whether the person is authorized to enter a predetermined operation scene or not based on the comparison result; recording an entry state during authorization; in the predetermined operation scene, continuously capturing environment sound, and extracting a target voice segment; extracting real-time identity acoustic features and real-time state acoustic features from the target voice segment; comparing the real-time identity acoustic features with an identity reference model; acquiring real-time context information; determining a corresponding predefined behavior rule according to the role and the real-time context information; and based on a predefined behavior rule, executing dynamic verification to obtain a dynamic verification result, and generating output information. According to the invention, the accuracy of voiceprint recognition can be effectively improved.
Owner:WUHAN LINK SOFTWARE CO LTD

Multi-scene adaptive noise reduction method and system for audio terminal equipment

The invention relates to the technical field of audio noise reduction, in particular to a multi-scene adaptive noise reduction method and system for audio terminal equipment. The method comprises the following steps: identifying a real-time environmental sound signal; environment noise source identification is carried out on the real-time environment sound signal, noise distribution perception modeling is carried out, and an environment noise distribution field is constructed; performing current scene noise change analysis on the environment noise distribution field, performing dynamic noise distribution rendering, and constructing a dynamic noise scene model; performing hedging suppression parameter calculation based on the dynamic noise scene model, performing adaptive noise suppression processing, and constructing an adaptive noise suppression strategy; acquiring an audio input signal of a user; and performing local audio gain processing on an audio input signal of a user, and performing multi-band decomposition to generate an audio frequency parameter of each band. The stability and the accuracy of the noise reduction effect are improved, so that the practicability and the adaptability of the audio terminal equipment are enhanced.
Owner:HANK ELECTRONICS

Photovoltaic equipment fault detection method and system based on voiceprint recognition

The invention relates to the technical field of voiceprint recognition, in particular to a photovoltaic equipment fault detection method and system based on voiceprint recognition, and the method comprises the following steps: collecting the sound of the surrounding environment of a current photovoltaic power station, and building scene associated equipment spectrum peak features; according to the method, classification of environment sound is introduced in the sound acquisition stage, and scene association is performed on the operation sound of the photovoltaic equipment and the surrounding background sound in combination with Mel-frequency cepstrum coefficient extraction and acoustic scene recognition, so that a more targeted reference framework is provided for subsequent analysis. Modeling of multi-source sound features is achieved by synchronously collecting operation sound of an inverter and a cooling fan in equipment and extracting spectrum peak frequency points and amplitudes. In the signal processing process, multiple sub-bands are divided on the basis of scene correlation spectrum peak frequency points, so that each sub-band is aligned with the acoustic characteristics of equipment, the structural precision of short-time energy envelope extraction is improved, and it is ensured that the cross-correlation analysis result is more reliable.
Owner:WUWEI SHENNENG NORTH ENERGY DEV CO LTD

Method of indoor old man falling detection system based on mixed mode

The invention relates to the technical field of indoor old people safety monitoring, and discloses an indoor old people tumble detection system based on a hybrid mode, which comprises a video acquisition module, an audio acquisition module, an identity recognition module, a tumble detection module, an emotion recognition module and an alarm module. According to the method of the indoor old people tumble detection system based on the mixed mode, video and audio data are creatively fused, multi-dimensional monitoring of the state of the old people, capturing of action postures of the old people by the video acquisition module and obtaining of voice and environment sound information by the audio acquisition module are achieved, and the two complements each other; abnormal sound in the audio, such as collision sound of falling and shouting sound, can assist in judgment, the multi-dimensional data acquisition and analysis mode greatly reduces misjudgment and missed judgment caused by single data, the accuracy and reliability of detection are remarkably improved, and the safety of old people is comprehensively guaranteed.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

Immersive active control wing-mounted flight simulation system

The invention relates to the technical field of wing-mounted flight simulation, in particular to an immersive active control wing-mounted flight simulation system. The system comprises a physical simulation platform used for bearing an experiencer and providing immersive sensory feedback; the motion capture subsystem is used for collecting motion data of an experiencer in real time; the wing-mounted flight aerodynamic modeling subsystem is used for calculating aerodynamic characteristics in the wing-mounted flight process; the dynamic wind feeling simulation subsystem is used for dynamically adjusting the intensity and direction of a wind field and simulating wind pressure distribution in different flight states; the dynamic stress control subsystem is used for applying dynamic force feedback to the experiencer and simulating stress change in the flight process; the three-dimensional visual subsystem is used for simulating a visual picture in real time; and the sound subsystem is used for comprehensively simulating the environment sound effect in the flight process. According to the invention, the immersion and reality of wing-mounted flight simulation are greatly improved, and the safety and reliability are ensured.
Owner:SHANGHAI JIYE ENTERPRISE MANAGEMENT PARTNERSHIP (LLP)

Ambient sound feature-based authenticity analysis method, apparatus and device, and medium

The invention relates to the technical field of voice processing, can be applied to business scenes of financial science and technology, medical treatment and health and the like, and discloses an authenticity analysis method, device and equipment based on environmental sound characteristics and a medium. Performing voice separation processing on the original voice data to generate environment voice data and pure voice data, analyzing the environment voice data and retrieving knowledge information associated with the environment voice data, and identifying the content of the pure voice data to generate dialogue text data, and inputting the environment sound features, the knowledge information, the dialogue text data and the user declaration information into an analysis model, and outputting a authenticity analysis result. According to the method, the environment sound data and the voice content are separated, the available features of the environment sound data and the voice content are extracted respectively, and the background knowledge and the user declaration information are combined to perform fusion reasoning in the unified analysis model, so that the accuracy of authenticity judgment and the adaptability to complex scenes can be improved.
Owner:PING AN TECH (BEIJING) CO LTD

Campus safety monitoring method and system based on multi-mode perception

The invention provides a campus safety monitoring method and system based on multi-modal perception, and relates to the technical field of campus safety, and the method comprises the steps: collecting multi-modal data in real time through a front-end perception unit, the multi-modal data comprising human body motion data, environmental sound data and personnel physiological data; independently analyzing the multi-modal data, and respectively generating a first early warning signal, a second early warning signal and a third early warning signal; the three early warning signals are uploaded to a cloud analysis platform for fusion analysis, so that the cloud analysis platform determines a comprehensive alarm level based on the signal intensity, the superposition weight and the time-space dynamic threshold of the three early warning signals; the space-time dynamic threshold value is dynamically adjusted based on the incident time and the incident place; early warning information is directionally distributed to a preset terminal according to the comprehensive alarm level and the incident site, and emergency response operation is triggered. According to the method provided by the invention, the campus spoofing event can be detected under the condition that the privacy of students is not influenced, and the detection result is more accurate.
Owner:LANGFANG XIHANG TECHNOLOGY CO LTD

Method and system for triggering intelligent conversation through audio-video reality

A method and a system for triggering a smart dialogue with audio-video reality are provided, in which a reality image interface is opened in a user device, in which a camera is activated to obtain an environment image, and a microphone is activated to obtain an environment audio. Afterwards, the server may receive the location information and the reality image request from the user device, and may include identifying the environment image to obtain the environment object, and identifying the environment audio to obtain the environment sound. Then, after an intelligent dialogue link point displayed on the reality image interface is triggered, an intelligent dialogue program is started, an intelligent dialogue interface can be started in the intelligent dialogue program, and after the chat robot is imported, dialogue content can be generated by running a natural language model according to the position information, the environment object and / or the environment sound.
Owner:PEIXI TECHNOLOGY CO LTD

Upmix processing method, system and device for stereo audio and video, and storage medium

The invention provides an up-mixing processing method, system and device for a stereo audio video, and a storage medium, and relates to the technical field of audio and video processing, and the method comprises the steps: obtaining a to-be-processed stereo audio video, carrying out the multi-stream separation of the to-be-processed stereo audio video, obtaining a plurality of independent sound source components, the to-be-processed stereo audio video comprises a to-be-processed stereo audio or a to-be-processed stereo video; performing dry and wet sound processing on the to-be-processed stereo sound video to obtain an environment sound component; performing frequency band extension on the environment sound component to obtain a target environment sound component; and carrying out space rendering on each independent sound source component and the target environment sound component by adopting a vector-based amplitude translation method, and carrying out multichannel output fusion to obtain a multichannel audio / video. According to the method, the sound source is accurately positioned and efficiently separated, the dynamic sense and the space sense of the sound source object are enhanced, and the audio and video upmix processing quality and the user experience are improved.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Remote monitoring nursing method, device and equipment for old people and storage medium

The invention relates to the technical field of medical monitoring, and particularly discloses a remote monitoring nursing method, device and equipment for old people and a storage medium, and the method comprises the steps: obtaining wrist acceleration information and environment sound information; acquiring a wrist acceleration frequency spectrum and an environment sound frequency spectrum; calculating the ratio of the noise intensity of the environment sound information at each frequency to the environment noise reference intensity to obtain a noise intensity proportion; acquiring a weighted wrist acceleration frequency spectrum; splicing to form a fusion feature vector; and adopting a preset health risk identification model to carry out risk level evaluation so as to judge the current health state of the old people, and generating alarm information when the risk level exceeds a preset threshold value. According to the method, monitoring and abnormal early warning of the health state of the old people are realized, the motion information and the environment sound information are comprehensively utilized, the interference of sound noise of household appliances on the acceleration information can be effectively inhibited, and the accuracy of health risk identification is improved.
Owner:ZHEJIANG EAST VOCATIONAL TECH COLLEGE

Systems and methods for dictation with a digital stethoscope

The present description relates to methods and systems for a medical dictation. In one example, a stethoscope includes a first microphone positioned to capture physiological sounds of a patient, a second microphone positioned to capture ambient sounds, one or more processors, and memory storing instructions executable by the one or more processors to: during a stethoscope mode, obtain a first signal from the first microphone and a second signal from the second microphone, process the first signal to capture a physiological sound signal, including performing noise cancellation on the first signal based on the second signal, and transmit the physiological sound signal to an external computing device and / or a speaker of the stethoscope; and during a dictation mode, obtain the second signal from the second microphone, process the second signal to capture a voice signal, and transmit the voice signal to the external computing device.
Owner:EKO HEALTH INC

Multi-modal information fusion emotion detection method based on visible light and voiceprint

The invention relates to the technical field of artificial intelligence, and particularly provides a multi-modal information fusion emotion detection method based on visible light and voiceprint. The method comprises the following steps: respectively acquiring video images and environment sounds of old people in a monitoring environment through a visible light image module and an audio voiceprint acquisition module; a facial expression detection module is used for recognizing a human face in the video image, and a visual anomaly signal is output; recognizing an audio voiceprint signal in the environmental sound according to a voiceprint detection module, and outputting an audio abnormal signal; according to the method, the detection accuracy is effectively improved, the missing report is reduced, the method is suitable for home and old-age care institution environments, and the method is of great significance to guarantee the safety of old people and alleviate serious consequences caused by falling down.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Gas sensor for measuring gas concentration

The present disclosure provides a gas sensor, including a substrate, a first housing fixed on the substrate and enclosed with the substrate to form a first chamber, and a first infrared transmitter and a first acoustic sensor connected to the substrate. The first acoustic sensor and the first infrared transmitter are housed in the first chamber, and the first housing is provided with a first venthole. The gas sensor also includes an environmental detection assembly connected to the substrate and located outside the first housing, and a differential processor connected to the substrate. The differential processor of the present disclosure can eliminate the ambient sound signal and the vibration signal in the first detection signal according to the second detection signal. Eliminate the strong interference of noise and vibration in the external environment, and improve the accuracy of the gas concentration detection of the gas sensor.
Owner:AAC ACOUSTIC TECH (SHENZHEN) CO LTD

Sound output control method based on multiple game levels

The invention relates to a sound output control method based on multiple game levels, and relates to the field of sound output control of cloud games, which comprises the following steps: establishing an entry library, selecting specific entries from the entry library, randomly generating levels by the game according to the selected entries, and extracting a basic music library corresponding to the specific entries; randomly selecting a fixed number of basic music from the extracted basic music library, combining the basic music into background music, environment sound effect, atmosphere sound effect and customs clearance music by using a random algorithm, and generating a music playing strategy according to the selected vocabulary entry; a large batch of random music is intelligently generated through machine learning to meet the requirements of multiple game levels, and the workload of game music making is greatly reduced; music used in the game is split into basic music capable of being freely combined according to types and use occasions, and the requirement of the game for the music is met on the premise that resources and cost are saved.
Owner:BEIJING CHUANDU HAPPY TECHNOLOGY CO LTD

Head-mounted earmuff device

The invention discloses a head-mounted earmuff device comprising: two earmuff bodies and a headband body. Each earmuff body has an earpad portion and a housing portion. The earpad portion has an accommodating hole. The housing portion is pivotably connected to the earpad portion. The housing portion covers the outer side surface of the earpad portion to isolate environmental sounds. The housing portion is pivoted to leave away from the outer side surface of the earpad portion to allow reception of the environmental sounds. The headband body has a mounting member and two assembling portions, and a pivoting end of the assembling portion is pivotably connected to the earpad portion. A sliding end of the assembling portion is slidably connected to a mounting section of the mounting member, and a dimension of a head-mounted space is adjustable by sliding and fixing the two assembling portions with respect to the mounting member.
Owner:WU JING FENG

Noise reduction regulation and control method for earphone and noise reduction earphone

The invention belongs to the technical field of earphone noise reduction, and provides a noise reduction regulation and control method for an earphone and a noise reduction earphone. The method comprises the following steps: acquiring a first environment sound signal of an environment area, and extracting human voice features from the first environment sound signal to construct a human voice feature template library; in response to a noise reduction regulation and control instruction, collecting a second environment sound signal of the environment area, converting the time domain signal into a frequency domain signal through Fourier transform, and extracting a frequency spectrum feature from the frequency domain signal; performing matching calculation on the spectrum features and a human voice feature template library to determine environment voice signals, human voice signals and noise signals; an active noise reduction algorithm is adopted to generate offset sound waves opposite to the noise signals in phase, the human voice signals are amplified, and the offset sound waves and the amplified human voice signals are output through a loudspeaker of the earphone. According to the invention, effective noise filtering and clear transparent transmission of specific human voice are realized, smooth communication can be realized without taking off the earphone, and the use experience is improved.
Owner:GUANGZHOU SOUNDBOX ACOUSTIC TECH

Vehicle door control method and device, electronic equipment and vehicle

The invention discloses a vehicle door control method and device, electronic equipment and a vehicle, and belongs to the technical field of vehicles. The method comprises the steps that when the vehicle senses an external knocking action, the environment sound intensity is detected; the response feasibility of the knocking signal is determined according to the environment sound intensity to control the vehicle door to be opened. According to the method, a user can unlock and open the vehicle door only through simple knocking without traditional key or key operation, the method is particularly suitable for the situation that a heavy object is held by hand or the key cannot be easily taken out, whether the knocking signal is responded or not is determined according to the environment sound intensity, false triggering can be reduced, high accuracy of knocking to open the door in the complex environment is achieved, and the user experience is improved. And the trouble that the user can take effect and is invalid when knocking to open the vehicle door is reduced.
Owner:BYD CO LTD

Systems and methods for automated movie generation and editing

A system and method to generate a video is provided. The method may include generating, based on a user input including a description of a desired video, a structured script including one or more of scene descriptions, dialogue, or explicit shot-level information. The method also includes generating, based on the structured script, a sequence of video frames representing one or more scenes. The method further includes generating, based on the structured script and the sequence of video frames, an audio track including one or more of ambient sounds, sound effects, or music. The generated audio track being temporally synchronized with the sequence of video frames. The method also includes combining the sequence of video frames with the audio track to generate a synchronized video output representing the desired video.
Owner:META PLATFORMS INC

Apparatus and a method for providing a customizable and interactive ambient sound experience

An apparatus for providing a customizable and interactive ambient sound experience, comprising a computing device configured to populate a user interface (UI) data structure presenting selectable sound categories by generating configurable audio outputs, each having audio parameters, for each selectable sound category, initializing audio settings linked audio parameters, and populate the UI data structure using configurable audio outputs and audio settings, transmit the populated UI data structure to a downstream device, adjust, in response to a user input targeting a first set of configurable audio outputs, at least one audio setting to modify the audio parameters linked to the first set of configurable audio outputs, and output, at the downstream device, the first set of configurable audio outputs by overlaying the first set of configurable audio outputs with a second set of configurable audio outputs to create a composite audio output, and outputting the composite audio output.
Owner:POCKET BARD LLC

Audio Device with Ambient Sound Source Selection and Natural Language Control

Various implementations include approaches for device control and / or sound source selection in audio devices. In some implementations, an audio device includes: an electro-acoustic transducer for providing an audio output; a set of microphones for detecting ambient sounds; and a processor coupled with the electro-acoustic transducer and the set of microphones, the processor configured to: evaluate microphone signals from the set of microphones to identify classes of sound sources in the ambient sounds; and adjust output of at least one class of the ambient sounds relative to another class of ambient sounds based on a user input.
Owner:BOSE CORP

Debris flow early warning system and method based on Transform fused sound and image

The invention relates to a debris flow early warning system and method based on Transform fusion sound and image, the system comprises a front end acquisition module, an edge calculation module and a remote alarm module, the system can synchronously acquire visible light image and environment sound information of a monitoring area, and uploads the acquired data to a cloud server in real time through a communication module; the cloud model is based on a Transform structure, carries out deep fusion analysis on sound and image features, and combines synthetic data and a data enhancement strategy, thereby effectively improving the recognition precision of the model for weak symptoms before debris flow occurrence and the overall generalization ability of the system. The system can be widely applied to debris flow high-incidence areas, meets the actual requirements of high-frequency monitoring, intelligent identification and rapid early warning of geological disasters in a complex environment, solves the problems of single mode, low identification precision, difficult arrangement and the like of an existing system, and improves the intelligent level and response efficiency of the geological disaster monitoring system.
Owner:DANMO INTELLIGENT TECH (HANGZHOU) CO LTD +1

Sound environment analysis and monitoring method based on artificial intelligence

The invention discloses a sound environment analysis and monitoring method based on artificial intelligence, and belongs to the field of artificial intelligence. Sound feature extraction and parameter analysis; converting sound into text; converting the sound event into a text description form by using a multi-mode big language model with a sound event analysis function and outputting text information; timestamps are added to the output text information, and the text information is classified, sorted, recorded and stored for a user to trace back; determining a key sound event by identifying a dialogue scene; after the key sound event is triggered, the device reminds the user according to a specific form. Any equipment does not need to be implanted, and postoperative risks and maintenance cost do not exist; a user can perceive environment sound and understand dialogue content without listening by himself or herself, and a hearing aid is not needed; sign language actions do not need to be captured through a camera, and the problem of dialect sign language recognition errors does not exist; the user does not need to stare at the character display device in real time during use, and the user is reminded in a specific mode after the key sound event is triggered.
Owner:SHENZHEN TECH UNIV

Remote recognition system based on spatial respiratory tract abnormal sound

The invention discloses a remote recognition system based on spatial respiratory tract abnormal sound, and relates to the technical field of sound recognition. The multi-mode sensor array is used for collecting breathing sound in a non-contact mode and obtaining sound source space information of the breathing sound, the cross infection risk is eliminated, and the limitation that a traditional stethoscope can only collect sound signals is broken through. The spatial audio processing module performs preset audio processing on the breathing sound to obtain a breathing sound signal, and obtains sound source control characteristics based on the sound source spatial information; the feature extraction module extracts time-frequency joint features and nonlinear dynamic features from the breath sound signals, and performs feature fusion on the time-frequency joint features and the nonlinear dynamic features and sound source control features to obtain multi-dimensional feature vectors; and the abnormal sound recognition model performs abnormal sound type recognition and spatial positioning on the multi-dimensional feature vector. Therefore, the environmental sound interference can be effectively solved, the type and position of the breathing sound can be intelligently identified, detected and positioned, the identification accuracy of the breathing sound is effectively improved, and meanwhile, the dependence on medical staff is also reduced.
Owner:JIANG SU ZHI ZI NA MI KE JI YOU XIAN GONG SI

Systems and methods for adaptive additive sound

A method for adaptive additive sound includes receiving ambient sound data corresponding to ambient sound in a first zone acquired by a microphone in the first zone, analyzing the ambient sound data from the first zone, generating audio signal data for the second zone based at least in part on the ambient sound data from the first zone, and transmitting the audio signal data for the second zone to a speaker in the second zone. The first zone is separate from the second zone within a space.
Owner:GOOGLE LLC

Stretchable ear clamping type Bluetooth earphone

The invention relates to the technical field of Bluetooth earphones, in particular to a stretchable ear clamping type Bluetooth earphone which is placed in an earphone cabin and comprises an ear hook, an auxiliary ear shell, a battery, an adjusting assembly, a rubber sleeve, an extension assembly, a main ear shell, a sound production unit and a sealing film, the auxiliary ear shell is connected to the lower left end of the ear hook, and the battery is connected to an inner cavity of the auxiliary ear shell; the adjusting assembly is connected to the ear hook, the rubber sleeve is arranged on the peripheral face of the ear hook in a sleeving mode, the extension assembly is connected to the lower right end of the ear hook, the main ear shell is connected to the lower end of the extension assembly, the sound production unit is connected with an inner cavity of the main ear shell, and the sealing film is connected between the sound production unit and the main ear shell. And a sealing cavity is formed between the ear hook and the rubber sleeve. By arranging the movable main ear shell, the main ear shell is close to the ear canal, so that the interference degree of environmental sound is reduced, and in the process that the main ear shell is close to the ear canal, a vacuum space is formed in the main ear shell to wrap the sounding unit, so that the purpose of sound insulation is achieved.
Owner:赖俊辉

Activity charting when using personal artificial intelligence assistants including differentiating a patient from a different person based on audio associated with toiletting

The present disclosure provides for the division of environmental sounds from speech sounds to extract and analyze the behaviors and activities occurring in the environment; thus expanding and improving the functionality of AI assistant devices. These environmental sounds can be securely uploaded to an electronic chart, and may be used to aid in the treatment of existing conditions or the prophylaxis / mitigation of conditions not yet experienced by a patient under observation. Accordingly, the present disclosure provides for improved functionality in assistant devices and devices linked to the AI assistant devices.
Owner:RESMED CORPORATION +2

Intelligent health monitoring alarm system

The invention discloses an intelligent health monitoring alarm system, and the system comprises an environment behavior monitoring module which is used for collecting environment and behavior data related to daily activities of a user, and the environment and behavior data comprises water data, power data, gas data, image data, environment sound data and refrigerator door state change data; wherein the environment sound data particularly refers to a playing sound source signal which is acquired by an indoor microphone and is generated by a television or audio playing equipment; the data processing center is used for receiving and storing the environment and behavior data; and the user habit learning module is used for establishing and dynamically updating a personalized daily behavior baseline model of the user through a machine learning algorithm based on historical data. According to the invention, through multi-dimensional non-contact monitoring, the influence of contact equipment on the self-respect of the elderly can be avoided, and more comprehensive and intelligent health monitoring alarm of the elderly can be monitored at the same time.
Owner:DALIAN MEDICAL UNIVERSITY

Behavior prediction method and device based on feature fusion

The invention relates to the technical field of video monitoring processing. The invention provides a behavior prediction method and device based on feature fusion. The method comprises the steps of collecting video data of an airport target monitoring area to form a video frame sequence, synchronously collecting environment sound of the airport target monitoring area to form an environment sound time sequence, performing multi-behavior type labeling, forming a positive sample and a negative sample, and obtaining a training data set; establishing a behavior prediction model based on a hybrid architecture of a CNN algorithm and an RNN algorithm, and performing feature fusion and model optimization when the training data set is used to train the behavior prediction model to obtain an optimized behavior prediction model; and carrying out feature fusion on a to-be-predicted video and the collected environment sound, and inputting the fused features into the optimized behavior prediction model to obtain a behavior prediction result. The risk behavior of the target area can be efficiently and accurately predicted.
Owner:CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD