Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3 results about "Multimedia signal processing" patented technology

Audio and video optimization processing method and system for handheld terminal

ActiveCN121037602BSelective content distributionReal time analysisMultimedia signal processing
The application discloses an audio and video optimization processing method and system for a handheld terminal, and relates to the technical field of multimedia signal processing, which comprises the following steps: independently of an audio coding process, a lightweight neural network model is used to perform real-time analysis on original audio data in parallel, and a structured audio semantic descriptor is output; in response to a selected current video coding strategy mode, a set of operation parameters of a video encoder is dynamically reconstructed to code a synchronously collected video frame; if the current video coding strategy mode is a speech active mode, the coding frame rate is increased and a region of interest coding for a face region is started; if the current video coding strategy mode is a music dominant mode, the coding resolution is increased; and if the current video coding strategy mode is a silent listening mode, the coding frame rate and resolution are reduced. Through dynamic video coding strategy adjustment based on audio semantics, the application realizes accurate balance of video quality, smoothness and resource consumption.
Owner:SHANGHAI DAQI INFORMATION TECH CO LTD

Video conference interaction equipment with self-adaptive energy-saving control

The invention relates to the technical field of video conferences and multimedia signal processing, and discloses video conference interaction equipment with self-adaptive energy-saving control, which comprises a central processing unit, an eye movement tracking sensor, a video processing unit and a display driving controller, the energy consumption arbitration controller module calculates a real-time comprehensive priority weight according to the semantic metadata and a fixation point coordinate, and for a video stream of which the real-time comprehensive priority weight is lower than a preset threshold value, the virtualized coding and decoding interface layer module intercepts a predictive coding frame and injects the predictive coding frame into a virtual null-hop unit constructed into a full-skip mode; triggering the video processing unit to execute zero copy address mapping operation, bypass inverse quantization and inverse transformation operation; the display driving controller synchronously reduces the refresh rate and the backlight current of the corresponding display area, and the power consumption of equipment is reduced while the decoding context continuity is maintained through cooperative control of virtualized null-hop injection and partition display.
Owner:HUASHENG INTELLIGENT TECH (GUANGZHOU) CO LTD

An audio-video segmentation method, system, and medium based on global temporal mixing and multi-scale audio input.

PendingCN122317335AFrame sequenceMultimedia signal processing
This invention discloses an audio-video segmentation method, system, and medium based on global temporal mixing and multi-scale audio input, belonging to the field of multimedia signal processing and information fusion technology. First, video data is divided into audio data and a continuous sequence of video frames, and acoustic and video features are extracted. Then, visual and acoustic features are input into a global temporal audio-video mixer to perform cross-modal global temporal fusion, obtaining non-homogeneous acoustic query features and visual features. Finally, a mask prediction decoder is used to achieve accurate matching between acoustic query features and visual features. This invention breaks through the dimensionality limitations of traditional fusion methods by designing a GTAVM module, significantly improving the accuracy of target segmentation; simultaneously, by introducing the MSAI mechanism, the model's adaptability to complex sound environments is enhanced.
Owner:BEIHANG UNIV