Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24875 results about "Audio frequency" patented technology

An audio frequency (abbreviation: AF) or audible frequency is a periodic vibration whose frequency is in the band audible to the average human. The SI unit of audio frequency is the hertz (Hz). It is the property of sound that most determines pitch.

Headset antenna and connector for the same

ActiveUS20090033574A1Increase the equivalent impedanceImprove rendering capabilitiesAntenna supports/mountingsAntenna adaptation in movable bodiesHeadphonesAudio frequency
A headset antenna and a connector for the same are provided. The headset antenna includes an audio signal line, an antenna and a high impedance element in specified application frequency ranges. The audio signal line is adapted for transmitting an audio signal and the antenna is adapted for receiving an RF signal. The high impedance element is disposed on a transmission path of the audio signal and generates a high impedance at a specified frequency band of the RF signal, so that the audio signal line is equivalent to an open circuit and the antenna obtains a better receiving capability.
Owner:HTC CORP

Multimodal intelligent agent system for dynamic environmental monitoring and human-centered support

A multimodal intelligent agent system for dynamic environmental monitoring and user-centered support, consisting of: a multimodal sensor module configured to continuously acquire environmental and behavioral data from multiple input modalities, including at least one visual sensor, at least one acoustic sensor, at least one environmental conditions sensor, and at least one proximity or motion detection sensor, each generating modality-specific data streams representing visual images, audio waveforms, physical environmental parameters, and motion signatures within a monitored environment; a data preprocessing and fusion subsystem that is operationally coupled with the multimodal sensor module and configured to normalize, temporally align, and transform the modality-specific data streams into high-dimensional feature embeddings using a variety of encoders, wherein the visual encoder uses convolutional or vision transformer architectures, the audio encoder uses a spectral-temporal feature extractor, and the sensor encoder transforms raw analog data into context vectors suitable for multimodal alignment; a multimodal processing unit consisting of a transformer-based large language model (LLM) trained on paired multimodal datasets and configured to perform semantic fusion, context abstraction, and inference across the aforementioned aligned multimodal feature embeddings to generate a contextual understanding of environmental and behavioral states; an adaptive agent controller coupled to the multimodal inference processing unit and configured to instantiate, manage, and terminate a variety of task-specific intelligent agents, each agent being a software unit configured to perform a specialized function selected from meeting summarization, behavioral analysis, misplaced object detection, or environmental anomaly identification, with the agents dynamically interacting with the inference engine to retrieve contextually relevant multimodal embeddings for task execution; a personalization and adaptive learning subsystem consisting of a user preference database and a neural memory structure configured to update and refine model parameters based on user-specific interaction history, thereby enabling personalized output generation, prioritization of recommendations, and long-term behavioral adaptation; and An output generation interface is operationally connected to the adaptive agent controller and configured to produce multimodal output in textual, visual, and auditory form. The interface is capable of displaying human-readable summaries, notifications, and visual reconstructions of identified entities or environmental states.
Owner:GOUNDER MOHAN SELLAPPA DR BENGALURU +3

Large-scene monitoring video abnormal event early warning method based on multi-modal large model

The invention relates to the technical field of abnormal event early warning, and provides a large-scene monitoring video abnormal event early warning method based on a multi-mode large model. According to the invention, the problems of delay, low accuracy and limited coverage range of abnormal event early warning of large-scene monitoring videos in the prior art are solved. According to the main scheme, multiple paths of high-resolution monitoring videos are spliced and preprocessed to generate a panoramic video; synchronously acquiring and preprocessing audio and sensor data to construct a multi-modal data set; video key frames are extracted by adopting a traditional small model, and the video key frames and multi-modal data are jointly input into a multi-modal large model based on a Transform architecture for deep feature fusion; abnormal events such as tumble, congestion and fight are identified based on the fusion features; triggering an early warning mechanism to send event type and position information in real time; and storing the full-dimensional data of the abnormal event for tracing analysis. The real-time processing performance is optimized through edge calculation, the complex scene understanding ability is enhanced in combination with a multi-modal large model, and the detection precision and the response speed are remarkably improved.
Owner:PEKING UNIV (TIANJIN BINHAI) NEW GENERATION INFORMATION TECH RES INST +1

Compensation for nonuniform delayed group communications

A method for synchronizing audio reproduction in collocated end devices is presented. Each of the devices auto-correlates using noise to determine a threshold prior to the antenna receiving an audio signal. When the devices receive a common audio signal, they provide audio outputs. Each device cross-correlates its audio output with the audio outputs of the other devices. The timing of the audio output of each device is then adjusted such that the audio outputs of all of the devices align temporally with the lagging or leading device.
Owner:MOTOROLA SOLUTIONS INC

Illegal content auditing method and device based on multi-modal data, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical treatment and health and the like, and discloses a violation content auditing method, device and equipment based on multi-modal data and a medium. Inputting the visual semantic features and the composite audio features into a multi-modal model, generating fusion features through model alignment and fusion, and analyzing the fusion features based on a knowledge base to judge whether illegal content fragments exist in the multi-modal data, and when the illegal content fragments exist, positioning the illegal content fragments in the multi-modal data and generating an auditing report. According to the method, the visual semantic features and the composite audio features are fused, cross-modal compliance analysis is realized in combination with the knowledge base, frame-level or time-axis-level positioning is performed on the illegal content segments, and the auditing report containing the evidence is generated, so that the problems of insufficient single-modal detection accuracy and poor positioning capability are solved, and the auditing accuracy is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Intelligent conference summary automatic generation method based on voice recognition and large model

The invention discloses an intelligent conference summary automatic generation method based on voice recognition and a large model. The method comprises the following steps: S1, executing voice activity detection operation on an audio data stream; s2, extracting embedding vectors of continuous and effective voice segments, and generating a voice segment set to which a spokesman belongs; s3, inputting the voice fragment set to which the spokesman belongs into an improved Whisper model, fusing a Speaker-Aware attention mechanism and a connection time sequence classification auxiliary path, and outputting a conference transcription text sequence set; s4, inputting the processed structured dialogue format into a GPT-4 large language model, and generating a conference semantic representation sequence; s5, generating a conference summary first draft text according to a preset summary generation template; and S6, performing formatting output operation on the conference summary first draft text. The conference semantic elements can be automatically extracted, the structured summary text can be generated, and the method is suitable for efficient conference recording and task tracking in government affair office, enterprise collaboration, academic discussion and other scenes.
Owner:JIANGSU GUOHUACHENJIAGANG POWER GENERATION CO LTD

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Humanoid robot body intelligent cooperative control system based on multi-mode perception fusion

The invention relates to the technical field of robot cooperative control, in particular to a humanoid robot body intelligent cooperative control system based on multi-modal sensing fusion, and the system comprises a multi-modal sensing fusion module which recognizes a target boundary position, analyzes a pressure change, and combines with a posture to extract audio features to generate an environment sensing graph; the dynamic time sequence adjustment module optimizes an action rhythm adjustment detail generation execution plan, the cross-modal behavior correction module corrects an offset optimization track generation coordination sequence, the task priority distribution module analyzes task distribution to generate an execution list, and the time sequence conflict correction module optimizes a path adjustment conflict generation coordination path. According to the method, an environment perception graph is constructed through matching and fusion of multi-source perception data, the action sequence and interval are dynamically adjusted to optimize an execution chain, trajectory offset is corrected to improve action precision, nearest response and load balancing are achieved through real-time task allocation, conflict blocking is reduced through path rearrangement and time sequence coordination, and a perception decision execution closed loop is formed; and the identification precision and the cooperation efficiency are improved.
Owner:SHANGHAI DIJIETONG DIGITAL TECH CO LTD

Multi-modal depression recognition system based on MFE-CCAGNN model

The invention belongs to the field of artificial intelligence, and provides a multi-modal depression recognition system based on an MFE-CCANNN model, which comprises a data acquisition unit, a data preprocessing unit and an MFE-CCANNN model unit. The data acquisition unit synchronously acquires multi-mode data such as videos, audios, texts and fNIRS when a subject performs the same interview task. The data preprocessing unit comprises a video preprocessing unit, an audio preprocessing unit, a text preprocessing unit and an fNIRS preprocessing unit. The MFE-CCARNN model unit comprises a video, audio, text and fNIRS neural signal feature extraction module, a multi-modal feature fusion module and a classification module, and depression recognition and classification result output are achieved. The system supports four-level depression degree discrimination, is high in recognition precision, portable in deployment, high in interpretability and the like, and is suitable for psychological health screening and clinical auxiliary evaluation scenes.
Owner:TONGJI UNIV

Wind turbine generator voiceprint fault recognition method

The invention provides a wind turbine generator voiceprint fault recognition method, and relates to the technical field of wind turbine generator state monitoring and fault diagnosis, and the method comprises the steps: carrying out the noise reduction of an original audio signal through variational mode decomposition, screening a target mode of which the frequency, energy and kurtosis accord with features, and reconstructing the signal; extracting a Mel frequency cepstrum coefficient and a sensing noise robust coefficient, and generating multi-dimensional voiceprint data in combination with statistical characteristics such as a frequency spectrum gravity center, a spectrum entropy, energy, kurtosis and a zero-crossing rate; constructing a support set based on the prototype network, realizing small sample fault classification by calculating the Euclidean distance between the feature vector and the prototype vector, and outputting a preliminary result; judging whether the voiceprint is abnormal according to a preset threshold value, if so, storing the voiceprint into a dynamic abnormal voiceprint knowledge base; frequently occurring abnormal samples are manually labeled and added into a support set, the prototype network is retrained to update the model, and continuous optimization of the fault recognition capability is achieved.
Owner:CGN (SHANXI) NEW ENERGY INVESTMENT CO LTD

Plug-and-play wireless high-definition audio and video transmission method and system

The invention relates to the technical field of wireless high-definition audio and video transmission, and discloses a plug-and-play wireless high-definition audio and video transmission method and system.The method comprises the steps that foreground space texture features and audio and voice segments are obtained, and an initial feature set containing space and semantic features is obtained through semantic correlation analysis; a unified representation vector is output through multi-modal feature fusion, and a dynamic importance score is obtained by determining time sequence consistency, calculating an importance weight and performing normalization; and when the score exceeds a threshold value, segmenting the video frame in real time to determine an attention focus area and optimize a boundary, thereby generating an attention weight matrix, preferentially allocating bandwidth and forming a partition differentiation compression result. And in combination with a network bandwidth state, protection is enhanced for a key stream, the priority and the bit rate are dynamically adjusted, and high-quality audios and videos are output through decoding and recombination. According to the invention, the bandwidth allocation and compression strategy can be dynamically optimized, the transmission quality of a key area is guaranteed when the bandwidth fluctuates, and the audio and video transmission efficiency and experience are improved.
Owner:深圳市翼联网络通讯有限公司

Systems and methods for generating an equal-loudness contour response using an auricular device

A system may include a storage device, configured to store computer-executable instructions. A system may include an ear-bud configured to be positioned within an ear canal of a user, the ear-bud comprising: a speaker, a microphone; and one or more processors in communication with the storage device, wherein the computer-executable instructions, when executed by the one or more processors, cause the one or more processors to: obtain a user hearing profile, obtain an equal-loudness hearing profile, receive audio data from the microphone, and generate a second audio data based on a first sound-pressure level, a second sound-pressure level, a first frequency; and cause the speaker to emit the second audio data within the ear canal of the user, such that the user perceives the audio data as if the user has normal hearing.
Owner:MASIMO CORP

Emotion prediction and disease derivation method and system based on multi-modal fusion

The invention discloses an emotion prediction and disease derivation method and system based on multi-modal fusion, and the system comprises a data collection and preprocessing module, an emotion fusion module, an abnormal condition detection and cloud uploading module, and a disease possibility derivation module. The data acquisition and preprocessing module comprises a video part, a text part and an audio part, and the video part comprises face emotion recognition and prediction and human motion recognition and prediction; the text part comprises text content emotion recognition and prediction; the audio part comprises voice-to-text and voice tone emotion recognition and prediction, the system comprehensively captures an emotion state by fusing multi-mode information such as video, text and voice, and the accuracy and prediction capability of emotion recognition are improved; and by predicting the future emotion trend, the abnormal condition is warned in advance, and the response timeliness is improved.
Owner:JIANGSU UNIV OF SCI & TECH IND TECH RES INST OF ZHANGJIAGANG

Deepfake detection

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.
Owner:PINDROP SECURITY INC

Multi-modal emotion recognition method and system based on cross-modal alignment and matching enhancement

The invention discloses an emotion recognition method and system based on cross-modal alignment and matching enhancement. According to the method, firstly, feature extraction is carried out on text, audio and video modalities in a data set, and then a text and audio cross-modal emotion alignment module and a text and video cross-modal emotion alignment module are constructed respectively, so that cross-modal semantic alignment is realized. Constructing an emotion label matching module based on an alignment result, generating modal pairs with similar emotions but different labels by using a difficult negative sample mining strategy, and paying attention to cross-modal emotion consistency through a dichotomy task guide model; performing modal feature fusion on the three modals through a six-layer attention crossing mechanism, finally splicing feature vectors, inputting the spliced feature vectors into a long-sequence context fusion modeling module for deep modal fusion, and capturing cross-modal interaction information; and the fused features are sent to an emotion classification module, and a final emotion category recognition result is output.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-modal emotion fusion analysis method and system

The invention discloses a multi-modal emotion fusion analysis method and system, and the method comprises the steps: carrying out the feature extraction of multi-modal emotion data modal by modal through a feature extraction module, and generating a text original feature, an audio original feature and a visual original feature; performing cross-modal alignment interactive fusion on the original text features, the original audio features and the original visual features based on a unified semantic alignment module, and constructing collaborative fusion features; performing mode and channel double-layer dynamic fusion optimization by adopting a dynamic fusion regulation and control module according to the text original feature, the audio original feature, the visual original feature and the collaborative fusion feature, and determining a unified fusion feature; performing hierarchical residual semantic gating enhancement based on the unified fusion features according to a high-order semantic abstraction module to generate semantic enhancement features; and inputting the semantic enhancement features into an emotion prediction module, and outputting an emotion analysis result. Based on the above scheme, a more stable, accurate and reliable emotion recognition result can be provided.
Owner:GUANGDONG UNIV OF TECH

Digital human interaction method and device based on multi-modal sentiment analysis and medium

The invention discloses a digital human interaction method and device based on multi-modal sentiment analysis and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: collecting multi-modal data of a user in real time through a multi-source sensor device; the multi-modal data comprises face video stream data, voice audio stream data and text dialogue data; calling data analysis engines corresponding to different modalities, and extracting corresponding modal feature sequences; according to the current interaction scene, the modal feature sequence and the historical dialogue context features are fused, and a comprehensive emotion evaluation result is generated; outputting a corresponding multi-modal response data packet based on the modal feature sequence through an interactive response engine corresponding to a comprehensive emotion evaluation result; and executing the multi-modal response data packet. And after feature fusion is carried out in combination with the current interaction scene, the generated response can more accurately fit the current emotion demand and communication context of the user, so that the digital human can be more easily fused into various scenes needing emotion interaction.
Owner:INSPUR ZHUOSHU BIG DATA IND DEV CO LTD

System and method for contextual analysis and metadata database generation for user-specific speech patterns

A system for contextual analysis and metadata database generation for user-specific speech patterns is disclosed. The system accesses a speech signal of a user and identifies the user based on the voice print associated with the user. The system splits the speech signal into a first set of audio frames, where each audio frame comprises an utterance of one or more words. The system determines a context associated with each word. In response, the system detects a context change between a first text and a second text. The system generates a contextually split set of frames by splitting the speech signal into a second set of audio frames according to the detected context changes.
Owner:BANK OF AMERICA CORP

Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement

The invention relates to the technical field of artificial intelligence, and discloses a Bluetooth communication intelligent speech translation method and system based on multi-mode enhancement, and the method comprises the steps: collecting a multi-channel audio signal through a built-in multi-microphone array of a Bluetooth device, carrying out the dynamic direction self-adaptive beam forming of the multi-channel audio signal, and carrying out the self-adaptive beam forming of the multi-channel audio signal; extracting a Mel spectrogram feature of the direction enhancement signal, identifying lip regions of a plurality of candidate speakers in each frame of real-time speaking video captured by a camera, performing time sequence convolution on the lip regions to obtain a lip movement time sequence embedded vector, calculating a correlation score with the Mel spectrogram feature, separating the direction enhancement signal, and obtaining a lip movement time sequence embedded vector; and performing text transcription and conversion on the high-confidence separation voice to obtain a translation language text, and sending the synthesized target translation voice to a preset mobile terminal through the Bluetooth device to obtain a target translation result. According to the method, the real-time performance and accuracy of speech translation are improved in a multi-person scene, far-field speech, noise interference and accent difference.
Owner:SHENZHEN DIE MICRO SEMICON CO LTD

Intelligent predictive maintenance system for audio equipment fault

The invention discloses an intelligent predictive maintenance system for an audio equipment fault, and the system comprises a data collection module which collects the internal sensor data during the operation of audio equipment, and outputs an audio signal feature parameter, an external environment parameter, and a historical operation log; the feature preprocessing module is used for carrying out standardization and noise reduction processing on the multi-source data based on a multi-modal feature fusion algorithm; the health state evaluation module outputs a health state evaluation result of the audio equipment in real time through a hybrid analysis model combining a convolutional neural network, a long and short-term memory network and an attention mechanism; the fault risk prediction module is used for performing real-time fault risk prediction according to the evaluation result and generating predictive maintenance decision parameters; and the maintenance decision and early warning module is used for outputting fault early warning information according to the prediction parameters and automatically generating maintenance operation suggestions when the early warning level reaches a preset condition. According to the invention, the operation reliability of the audio equipment can be effectively improved, and intelligent prediction and advanced maintenance of faults are realized.
Owner:SHENZHEN JIEYU INFORMATION TECH CO LTD

Multidirectional frame audio stream transmission method, device, equipment and medium

The invention discloses a multidirectional frame audio stream transmission method, device, equipment and medium, and the method is realized through cooperation of a transmitting end and a receiving end: the transmitting end cuts an original audio stream into independent audio frames, gives priority identifiers to the independent audio frames, and determines redundant coding parameters and transmission paths for different priority frames in combination with a predefined static strategy; generating a data packet containing an original data block and a redundant data block, and sending the data packet through at least one network path; and a receiving end caches the multi-path data packet, recovers lost data by using redundant data blocks to recombine a complete audio frame, and executes error concealment processing on the frame which cannot be recombined to generate a replacement frame. According to the method, based on a multi-path parallel transmission, forward error correction (FEC) redundancy mechanism and a cost-aware static scheduling strategy, lossless forwarding and instantaneous recovery of audio streams are realized on the premise of not waiting for network feedback and avoiding inter-frame dependence, and high-quality real-time audio transmission service can still be provided in a complex network environment.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Physiological monitoring soundbar

A soundbar for medical monitoring which may comprise a speaker, a sensor, and a hardware processor. The speaker can be configured to emit audio signals. The sensor can be configured to obtain sensor data relating to a physiology of a subject. The sensor can include a camera and the sensor data can include image data. The hardware processor can be configured to access the sensor data and determine a health status of the subject based on at least the sensor data.
Owner:MASIMO CORP

Task execution strategy generation and adjustment method and device, equipment and medium

PendingCN120951241ABiological modelsLive feedbackEngineering
The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as personal intelligence, financial science and technology, medical health and the like, and discloses a task execution strategy generation and adjustment method, device, equipment and medium. Encoding the visual information, the audio information and the language instruction information to obtain a visual feature, an audio feature and a language feature, fusing the visual feature, the audio feature and the language feature to generate a comprehensive feature, generating an initial task execution strategy based on the comprehensive feature and executing a corresponding action, and in the execution process, according to real-time feedback information of the environment, executing the corresponding action according to the initial task execution strategy. And dynamically adjusting the initial task execution strategy by adopting a reinforcement learning model to obtain an updated task execution strategy. Through multi-modal information fusion and reinforcement learning dynamic adjustment, optimization and flexible updating of a task execution strategy in a complex environment are realized, and the autonomous decision-making capability of the intelligent equipment is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Video translation method and system based on artificial intelligence

The invention discloses a video translation method and system based on artificial intelligence. The method relates to the technical field of video translation and comprises the following steps of original sound track extraction, target AI speaker adaptation, AI dubbing generation and mouth shape synchronization and video synthesis. According to the method, independent audio and video streams are obtained by adopting an audio and video separation technology, and multiple original sound tracks are extracted through a voice separation model; matching or generating an adaptive target AI speaker module in a preset tone library; converting the original language voice into a text, translating the text into a target language text, and synthesizing an AI dubbing audio track in combination with a target AI speaker module; and finally, the independent video stream and the multi-AI dubbing audio track are input into the mouth shape synchronization model to output a translated video, so that the timbre fitting degree, the voice quality and the voice consistency of the same speaker of AI dubbing are improved, and meanwhile, the resource utilization rate of video translation and the processing efficiency under batch tasks are improved. The problem that in the prior art, video translation is low in quality and efficiency is solved.
Owner:BEIJING DEEP LOGIC INTELLIGENT TECHNOLOGY CO LTD

Digital human interaction system based on web terminal

The invention discloses a digital human interaction system based on a web end, relates to the technical field of digital human interaction, and aims to solve the problem of accumulated dislocation of browser end digital population animation and actual audio playback caused by multiple clocks and buffer scheduling. The system comprises a visual position cooperative control module, a session initialization module, an audio track and rendering canvas binding module, a multi-domain alignment time base cluster establishment module, construction of a time base cluster containing a system reference sub-time base and a content logic sub-time base, potential candidate anchor point generation module, an optimal anchor point selection module, synchronization error calculation, and judgment of a synchronization steady state, a fine adjustment state or a lost state. An optimal anchor point is screened to adjust the animation, and a visual effect studio dynamic maintenance module and an enhancement generation module assist in out-of-step processing and parameter optimization; through cooperation of multiple modules, accurate synchronization of audio and digital human animation is realized.
Owner:NANJING SUPERMIND INFORMATION TECHNOLOGY CO LTD

Low-delay audio input switching method and system, storage medium and equipment

The invention relates to the technical field of audio control, and discloses a low-delay audio input switching method and system, a storage medium and equipment, and the method comprises the steps: carrying out the parallel pre-initialization of a plurality of pieces of audio input equipment, and building an independent parallel audio data cache for each piece of equipment; monitoring the state of each audio input device and the audio stream quality in real time, and judging whether switching is triggered or not based on a multi-factor decision model; after the switching decision is triggered, seamless audio data stream switching is executed, and format unification processing and cross fade-in and fade-out transition are included; a unified equipment operation interface is provided through the hardware abstraction layer, and system resources are optimized and managed; a delay sensing closed-loop control mechanism is constructed, processing delay of each link is monitored in real time, a caching strategy, a processing algorithm and resource allocation parameters are dynamically adjusted, self-adaptive balance of low delay and high tone quality is achieved, and through the method, quick and smooth switching of audio input equipment is achieved, delay is remarkably reduced, and real-time audio experience is improved.
Owner:LINKPLAY TECHNOLOGY INC NANJING

Multi-modal emotion recognition method and device, electronic equipment and storage medium

The invention discloses a multi-modal emotion recognition method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring text, video and audio data of a user and respectively performing feature extraction to obtain text features, audio features and facial features; the three features are input into a pre-trained multi-modal emotion recognition model, the multi-modal emotion recognition model comprises a first fusion module, a second fusion module, a third fusion module, a fourth fusion module and a classification module, the audio features and the text features are fused through the first fusion module, and audio text features are obtained; fusing the facial features and the text features by using a second fusion module to obtain facial text features; performing feature enhancement on the text features by using a third fusion module to obtain enhanced text features; fusing the audio text features, the face text features and the enhanced text features by using a fourth fusion module to obtain multi-modal features; and classifying the multi-modal features by using a classification module to obtain a sentiment classification result of the user.
Owner:AGRICULTURAL BANK OF CHINA

Bird identification method and device based on sound-image multi-modal fusion

The invention discloses a bird identification method based on sound-image multi-modal fusion. The bird identification method comprises the following steps: S1, carrying out standardized frame-level preprocessing on bird audio signals; s2, acoustic features are extracted and enhanced, and an acoustic high-level feature vector which highlights birdsong discrimination information and suppresses environmental noise is obtained; s3, visual image standardization preprocessing; s4, performing visual feature extraction and multi-scale fusion to obtain a visual high-level feature vector which enhances correspondence to the bird key form area and inhibits background interference; s5, performing dynamic weighted fusion on the decision-making layer to obtain a bird existence probability; and S6, comparing the bird existence probability with a preset threshold value of the corresponding bird, and judging whether the bird exists or not and the type of the existing bird. Through cross-modal feature enhancement and adaptive fusion, the precision, robustness and real-time performance of bird recognition in a complex orchard environment are significantly improved, and a core technical support is provided for green intelligent bird repelling.
Owner:NANJING FORESTRY UNIV

Multi-modal large model optimization method and device for complex task and medium

The embodiment of the invention discloses a multi-modal large model optimization method and device for a complex task and a medium, and relates to the technical field of multi-modal large models.The method comprises the steps that multi-modal input data are received, modal specific feature extraction is carried out on the multi-modal input data, and initial feature representation of each modal is generated, the multi-modal input data comprises image data, text data and audio data; generating a dynamic parameter adjustment strategy based on a pre-acquired task type and input data complexity, and adjusting calculation parameters of a multi-modal encoder in the multi-modal large model according to the dynamic parameter adjustment strategy to determine a corresponding optimization modal processing sub-module; and performing cross-modal fusion by optimizing the modal processing sub-module and the initial feature expression to generate a fused multi-modal feature expression, and performing semantic analysis and reasoning on the multi-modal feature expression by using a pre-trained multi-modal large model to generate a task output result.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Abnormal blood sampling test data evaluation processing method and system of blood sampling system

The invention provides an abnormal blood sampling inspection data evaluation processing method and system of a blood sampling system. Wherein the diffusion track dynamic state of an anticoagulant-blood contact surface in a blood sampling tube is tracked in real time, dielectric response waveform distortion characteristics on the two sides are synchronously collected through a capacitance sensor array, and diffusion form parameters are generated; curvature sudden change is analyzed to trigger an acoustic fluid device to modulate audio frequency and flow velocity, and a composite disturbance field is formed; performing capacitance axial scanning on the disturbed mixed fluid, extracting phase angle offset and gradient abrupt change point coordinates, and constructing a dielectric anomaly three-dimensional map; matching a standard dielectric model, positioning a coordinate set of a concentration gradient imbalance area and a coordinate set of an eddy current attenuation area, and generating an anticoagulant abnormal index and a mixed defect matrix in combination with phase angle offset difference; and establishing a data anomaly degree evaluation function based on the anomaly index and the defect matrix, dynamically associating a clinical error threshold, and outputting a test evaluation result. According to the invention, anticoagulant diffusion can be monitored in real time, and the mixing quality error can be accurately evaluated.
Owner:BEIJING CANCER HOSPITAL PEKING UNIV CANCER HOSPITAL