Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1720 results about "Frame sequence" patented technology

Intelligent campus safety early warning method and system based on edge computing and big data

The invention provides a smart campus safety early warning method and system based on edge computing and big data. The method comprises the following steps: collecting multiple paths of video data streams in real time through an edge computing node, performing frame sequence segmentation and spatial-temporal feature extraction, generating an initial behavior feature set, performing multi-dimensional correlation analysis on the initial behavior feature set based on a preset behavior semantic tag, extracting a spatial-temporal behavior feature vector corresponding to a target monitoring scene, and obtaining a spatial-temporal behavior feature vector; and transmitting the time-space behavior feature vector to a central server, inputting the time-space behavior feature vector into a target behavior recognition model, generating a behavior semantic description sequence corresponding to the video data stream, determining a real-time behavior monitoring result according to a matching result of the behavior semantic description sequence and a preset abnormal behavior rule base, receiving the real-time behavior monitoring result through an edge computing node, and sending the real-time behavior monitoring result to the central server. And adaptive adjustment is carried out on acquisition parameters of the video data stream based on a dynamic priority strategy. According to the invention, the real-time performance, the accuracy and the system sustainability of campus behavior monitoring can be improved.
Owner:GUANGDONG SANZHU TECH CO LTD

Network edge monitoring and early warning method based on video image AI analysis

The invention discloses a network edge monitoring and early warning method based on video image AI analysis, and the method comprises the following steps: S1, obtaining video data, processing the video data, and generating an image frame sequence; s2, analyzing an image frame sequence, and extracting space and time features; s3, taking the target as a hypergraph node, constructing a hyperedge based on space and time features, and dynamically adjusting hyperedge connection by using a genetic algorithm; s4, constructing a multi-layer hypergraph Transform network, extracting multi-scale spatial-temporal characteristics, and calculating a semantic relationship between nodes; s5, setting a butterfly optimization algorithm initial population, and dynamically optimizing model parameters through global and local search; s6, constructing an anomaly detection model, identifying an abnormal behavior, and feeding back a result to optimize model parameters and hyperedge selection; and S7, deploying the model at an edge node, triggering early warning when an abnormal behavior is detected, and pushing information to a management platform. According to the invention, through video image AI analysis, accurate detection and real-time early warning of abnormal behaviors in a network edge scene are realized.
Owner:SHAANXI VIDEO BIG DATA CONSTR & OPERATION CO LTD

Audio and video dual-mode emotion recognition method and system based on adapter fusion

The invention relates to the technical field of artificial intelligence and emotion calculation, in particular to an audio and video dual-mode emotion recognition method and system based on adapter fusion. The method comprises the following steps: acquiring a video frame sequence and an audio signal, and preprocessing the video frame sequence and the audio signal; constructing an emotion recognition model; based on a bimodal feature extraction module, a space adapter and a global adapter are embedded in sequence, and corresponding modal enhanced space features and global features are obtained in sequence; generating intermediate representations of the corresponding modes based on the global features, and performing feature fusion according to the intermediate representations to obtain fusion features of the corresponding modes; the fusion features are spliced, time sequence features are extracted, and final features are obtained; inputting the final features into a classifier to obtain a predicted emotion category, training an emotion recognition model by adopting a loss function, and determining an optimal emotion recognition model; and inputting a to-be-recognized video frame sequence and an audio signal into the emotion recognition model, and outputting a recognition result.
Owner:NANJING MEDICAL UNIV

Real-time video analysis method based on deep learning

The invention relates to the technical field of computer vision, and discloses a real-time video analysis method based on deep learning. The method comprises the following steps: acquiring a real-time video stream through image acquisition equipment, and performing frame segmentation processing to generate a continuous video frame sequence; and extracting features of the video frame sequence by using a pre-trained convolutional neural network to obtain a multi-dimensional feature vector, inputting the multi-dimensional feature vector into the time sequence analysis model to calculate dynamic relevance, and outputting an inter-frame movement track and object behavior features. And constructing a scene understanding map containing a spatial position and a time evolution relationship according to the above-mentioned data, and carrying out abnormal event detection and generating event marking data based on the map. And performing semantic analysis on the event marking data, determining an abnormal event type and a confidence score, triggering a real-time alarm signal according to a result, and updating a historical event database. In the analysis process, the resource occupancy rate of the system is continuously monitored, the calculation precision is dynamically adjusted, a degradation processing mechanism is started when a preset threshold value is exceeded, and key area analysis is preferentially guaranteed.
Owner:HANGZHOU SIYUAN INFORMATION TECH CO LTD

Image target tracking method and device based on deep learning

The invention relates to the technical field of image processing, and provides an image target tracking method and device based on deep learning, and the method comprises the steps: obtaining an original frame sequence of a target video, carrying out the adaptive frame screening of the original frame sequence, and obtaining a plurality of feature frame sequences, key frames and common frames, performing feature extraction and feature mapping on all the feature frame sequences to obtain target trajectory nodes, performing trajectory optimization calculation on the key frames and the common frames by using all the target trajectory nodes to obtain key frame trajectories and common frame trajectories, and inputting the key frame trajectories and the common frame trajectories into a preset deep learning model to perform trajectory reconstruction optimization so as to obtain a target trajectory; and obtaining a tracking target trajectory. By performing adaptive frame screening on the original frame sequence and generating the feature frame sequence, the key frame and the common frame, the data processing efficiency and the information extraction accuracy are improved, the accurate generation of the target motion track is realized, and the problems of reduced tracking accuracy and delayed system response in a complex scene are solved.
Owner:HOHEM TECHNOLOGY CO LTD

Video semantic segmentation method based on time sequence cross attention mechanism

The invention discloses a video semantic segmentation method based on a time sequence cross attention mechanism, and belongs to the field of computer vision and the field of material detection.The video semantic segmentation method comprises the steps that firstly, a video used for training is preprocessed, a frame sequence is extracted, and then a coding-decoding network for multi-level feature extraction and fusion is constructed; according to the method, feature extraction is enhanced through a time sequence cross attention module, network parameters are optimized through weighted IoU loss and binary cross entropy BCE loss, then frame-by-frame prediction segmentation is carried out on a target video by using a trained model, and a multi-classification segmentation result is exported. According to the method, a time sequence cross attention mechanism is integrated into the SAMUNet network, the segmentation precision is effectively improved for image data with time sequences, the time cost and the labor cost of material video processing are greatly reduced, the method can be widely applied to the field of industrial detection, and the product quality and the production efficiency are improved.
Owner:ZHEJIANG UNIV

Game scene optimization method and system based on dynamic rendering

The invention provides a game scene optimization method and system based on dynamic rendering, and belongs to the technical field of game development.The game scene optimization method comprises the steps that firstly, a real-time rendering state data set of a current game scene, rendering resource occupation information containing scene elements and picture frame rendering time consumption parameters are obtained; then, a pre-trained rendering demand analysis model is called for analysis, a rendering demand feature set of the scene elements is obtained, the rendering demand feature set comprises element visual importance features and rendering resource sensitivity features, and a scene element dynamic optimization strategy is generated based on the rendering demand feature set; the scene element dynamic optimization strategy comprises an element detail level adjustment rule and a rendering resource allocation priority parameter, executing a rendering parameter adjustment operation on a target scene element according to the scene element dynamic optimization strategy, generating adjusted scene rendering configuration data, and finally inputting the scene rendering configuration data into a game rendering pipeline for real-time rendering processing. And outputting the optimized game scene picture frame sequence. Therefore, the game picture quality and the operation performance can be improved.
Owner:CHENGDU FUQIAN TECH CO LTD +1

Video stream processing method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of medical health, financial science and technology and the like, and discloses a video stream processing method which comprises the steps of collecting current environment parameters, generating a mode switching instruction and determining a target processing mode; obtaining multi-dimensional context awareness data, and selecting a target detection model; key area coordinate parameters in the video frame sequence are extracted, and grading resolution parameters are determined; and based on the target processing mode, the target detection model and the grading resolution parameter, constructing a video processing strategy matrix, executing the video processing strategy matrix to perform coding processing on the video stream, and generating a target coding video stream. According to the invention, through intelligent mode switching based on the current environment parameters, dynamic adjustment of the video processing mode is realized, and the adaptive capacity of the system in a complex environment is improved; through target detection model selection in combination with multi-dimensional context awareness data, the detection precision is optimized, and the reliability of visual analysis is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Task instruction generation method and device based on cross-modal fusion, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a task instruction generation method, device and equipment based on cross-modal fusion, and a medium, and the method comprises the steps: carrying out the decoding and noise reduction of an input video, generating a frame sequence, and recognizing a plurality of key frames based on the inter-frame similarity; extracting spatial features of the key frames to form a sequence, and generating video spatial-temporal features in combination with time features; performing semantic preprocessing on the input text to obtain text semantic features, and acquiring motion sensor signals to obtain motion features; fusing the video spatio-temporal features, the text semantic features and the action features to generate fused features; and generating a perception vector based on the fusion feature and outputting a task instruction. According to the method, multi-modal fusion is realized through key frame extraction and space-time fusion mechanisms in combination with text semantic features and action features, and the perception expression ability and the task instruction generation accuracy are improved by using time sequence information and multi-source perception input of the video.
Owner:PING AN TECH (SHENZHEN) CO LTD

Visual call information processing method and system based on 5G

The invention relates to the field of data processing, and provides a 5G-based video call information processing method and system, and the method comprises the steps: continuously obtaining a real-time video frame sequence and 5G network environment perception data in a video call scene, carrying out the multi-dimensional state mapping processing of the 5G network environment perception data, constructing a network transmission adaption model, and carrying out the real-time video frame sequence and 5G network environment perception data. Generating a video coding control instruction based on the network transmission adaptation model, performing content-aware coding conversion on the real-time video frame sequence, and outputting a coding optimization stream; in the transmission process of the coding optimization stream, link state fluctuation information is obtained through a 5G network feedback channel, transmission strategy dynamic calibration is performed on the coding optimization stream according to the link state fluctuation information, and a calibration transmission stream is obtained; and carrying out decoding time sequence alignment processing on the calibration transport stream, generating a visual call output sequence which is synchronous with the time of the original video stream unit, and pushing the visual call output sequence to a receiving end presentation device.
Owner:CHENGDU IKE IND CO LTD

Non-contact heart rate detection method, system and device based on visual Transform and multi-scale feature aggregation and medium

The invention discloses a non-contact heart rate detection method, system and device based on visual Transform and multi-scale feature aggregation and a medium. The method comprises the steps that a face visible light video is obtained, the unified video frame rate is re-sampled through the frame rate, face key point positioning and region division are carried out on each frame of image, and a face visible light video sequence is obtained; performing time migration operation on the video sequence to obtain a difference frame sequence, performing channel fusion on the video sequence and a difference frame, and performing down-sampling on a spatial dimension to obtain a low-resolution feature tensor; constructing a non-contact heart rate extraction model, inputting the low-resolution feature tensor into the non-contact heart rate extraction model, and outputting a heart rate value of the target; the non-contact heart rate extraction model comprises a feature enhancement module, a multi-scale mask feature aggregation module, a Transform time sequence modeling module and an rPPG predictor, the system, the device and the medium are used for achieving the non-contact heart rate detection method based on visual Transform and multi-scale feature aggregation, and the precision of rPPG signal extraction is improved.
Owner:NORTHWEST UNIV

Video monitoring abnormal behavior identification method and system based on edge AI

The invention discloses a video monitoring abnormal behavior identification method and system based on edge AI, and relates to the technical field of intelligent video analysis, and the method comprises the steps: carrying out the frame segmentation processing of a video monitoring data stream, obtaining a video frame sequence, building a target trajectory prediction model, predicting the position region of a current frame, and generating target position prediction data; performing multi-scale feature extraction on the video frame sequence, performing fusion matching on a target feature vector and target position prediction data, and constructing an enhanced feature matrix; carrying out weight distribution on the key behavior characteristics by adopting an attention mechanism, setting a dynamic threshold adjustment mechanism, and dynamically adjusting an abnormal behavior judgment threshold according to the personnel density; and inputting the adjusted feature data into an abnormal behavior classifier for identification and judgment, outputting an abnormal behavior identification result, and generating an abnormal event report. According to the method, the abnormal behavior detection accuracy in a complex monitoring scene is improved, the false report and missing report rate is reduced, and the millisecond-level real-time response capability is realized.
Owner:NANJING CHAOS INTERNET OF THINGS TECH CO LTD

Video stream analysis method and system based on unmanned aerial vehicle inspection

The invention provides a video stream analysis method and system based on unmanned aerial vehicle routing inspection, and relates to the technical field of unmanned aerial vehicles. Firstly, original video stream data of a target area is collected by using a camera device carried by an unmanned aerial vehicle, and then feature extraction is performed on the original video stream data; the method comprises the steps of obtaining a space correlation feature set and a time dynamic feature set of a video frame sequence, then carrying out anomaly detection on the two feature sets based on a preset anomaly detection model, generating an anomaly distribution feature set of a target area, generating a dynamic path optimization instruction according to the anomaly distribution feature set and real-time flight parameters of an unmanned aerial vehicle, and optimizing the dynamic path according to the dynamic path optimization instruction. The method is used for adjusting the inspection path of the unmanned aerial vehicle so that the unmanned aerial vehicle can focus on covering the abnormal area, and calling the incremental learning module to update the parameters of the anomaly detection model based on the anomaly distribution feature set so as to ensure that the anomaly detection model can adapt to the change of the target area and improve the anomaly detection accuracy. Therefore, more efficient, comprehensive and accurate inspection of the target area is realized.
Owner:DEYANG JINGKAI ZHIHANG TECH CO LTD

Video image color correction method and system based on artificial intelligence

The invention discloses a video image color correction method and system based on artificial intelligence, and relates to the technical field of color correction, and the method comprises the steps: collecting video data, carrying out the image size normalization and denoising processing, and obtaining the preprocessed video frame sequence data; and extracting reflectivity data and illumination data by using a Retinex-Net deep learning model, calculating an illumination adjustment coefficient, carrying out pixel correction, and carrying out color temperature correction on the pixel data of the video frame subjected to pixel correction according to the illumination adjustment coefficient. According to the method, the preliminary color mapping matrix is generated through the color histogram matching method, the color distribution of the image can be optimized, the image is enabled to better conform to the color features of the target reference frame, the color temperature adjustment is performed on the image according to the illumination adjustment coefficient and the illumination threshold, the cold and warm tones of the video are enabled to be accurately adjusted, and the image quality is improved. Through successively carrying out pixel correction, color correction and local area correction, a stable color adjustment frame is provided, so that the video styles are consistent.
Owner:WUHAN HENGJI INTELLIGENT CLOUD NETWORK TECHNOLOGY CO LTD

Network camera monitoring identification method and system based on artificial intelligence

The invention provides a network camera monitoring identification method and system based on artificial intelligence, and the method comprises the steps: firstly obtaining a monitoring data stream which is outputted by a network camera and comprises a video frame sequence and a corresponding time sequence metadata sequence, and then carrying out the spatial-temporal context coding processing of the monitoring data stream; the method comprises the following steps: generating a context feature cube containing spatial position information and time evolution information, then executing a normal behavior mode learning operation based on the context feature cube, and generating a reference feature library containing typical scene feature templates and feature evolution rule description; and dynamically matching and comparing the context feature cube of the current time period with the reference feature library, calculating a feature matching deviation value and generating an abnormal confidence score, and finally generating a monitoring early warning instruction containing abnormal occurrence time, a space coordinate range and a confidence level identifier according to the abnormal confidence score and corresponding space-time position information. And the accuracy and the early warning effect of monitoring and identification of the network camera are effectively improved.
Owner:SICHUAN XINSAIHU INTERNET OF THINGS TECHNOLOGY CO LTD

Knowledge-intensive visual question and answer automatic data generation method and device

The invention relates to a knowledge-intensive visual question and answer automatic data generation method and device, and the method comprises the steps: constructing an original visual data set containing the professional knowledge of a target domain according to a static image, a video stream and multimedia content; extracting a representative frame sequence, converting the audio information into text information, and extracting character information in the static image to construct a structured visual instance database; according to the prompt text meeting the preset professional depth condition, establishing a three-level prompt system containing domain knowledge, an evaluation standard and a generation specification; generating a corresponding visual question and answer pair data set according to the dynamic cooperation of the main agent and the domain expert agent; generating a multi-agent quality evaluation system according to the quality evaluation result; and designing a difficulty grading mechanism according to the negative example sample. According to the method, the professionality, the accuracy and the diversity of the visual question and answer data are remarkably improved, and reliable data support is provided for training and evaluation of a multi-modal large model.
Owner:TSINGHUA UNIVERSITY

Video generation method and device based on action coherence, equipment and medium

The invention relates to the technical field of data analysis, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses a video generation method, device, equipment and medium based on action coherence, the method comprises the following steps: obtaining a video action material set, carrying out key frame analysis on the video action material set to obtain a key video frame sequence, extracting an inter-frame residual vector between consecutive frames in the key video frame sequence, adjusting a preset initial diffusion model by using the inter-frame residual vector to obtain an optimized diffusion model, performing cosine scaling on the inter-frame residual vector by using the optimized diffusion model to obtain a scaled residual vector, and obtaining a video adjustment text, and carrying out noise addition and splicing on the key video frame sequence by using the video adjustment text and the zoom residual vector to obtain a noise video, and carrying out noise reduction on the noise video to obtain a target action video. According to the method and the device, the action coherence in the customized generated video can be effectively improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Aircraft target tracking method and system based on compensation prediction

The invention discloses an aircraft target tracking method and system based on compensation prediction, which are used for improving the target tracking precision in an image transmission delay scene. The method comprises the following steps: firstly, acquiring an image frame sequence of a target aircraft by using an airborne monocular camera, extracting a target center coordinate through a small target detection algorithm, and constructing a position sequence; the method comprises the following steps: extracting current high-frequency I MU data aiming at the condition that an image frame has transmission delay, inputting the current high-frequency I MU data into an LSTM-DKF model constructed by fusing LSTM and a delay Kalman filter, and predicting and generating a process noise and observation noise covariance matrix; and initializing a delay Kalman filter by using the matrix, and recursively predicting the target position during the delay period. And when the delayed image frame is received, backtracking and updating the state of the filter, recurring to the current moment again, and outputting the compensated target position. And finally, pixel deviation is calculated according to the compensation position, an aircraft tracking control instruction is generated, and high-precision target tracking is realized.
Owner:GUANGDONG UNIV OF TECH

Video generation method and apparatus, device, and medium

Embodiments of the present application provide a video generation method and apparatus, a device, and a medium. The method can be applied to the technical field of video content generation, and is used for improving the video generation quality. The method comprises: acquiring a sample video frame sequence from a sample video, determining a first step count and sample original noise, and performing data noise addition processing on the sample video frame sequence to obtain video input data; inputting the video input data, a sample text encoded feature corresponding to sample description text, and first embedding information corresponding to the first step count into an initial generation model; and performing noise prediction on the sample video frame sequence by means of M spatiotemporal residual components and M spatiotemporal attention components in the initial generation model to obtain sample predicted noise, correcting a network parameter in the initial generation model on the basis of the sample original noise and the sample predicted noise, and determining the initial generation model comprising the corrected network parameter as a video generation model.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Model training method and device, electronic equipment and storage medium

The invention discloses a model training method and device, electronic equipment and a storage medium. Comprising the following steps: sampling an original video frame sequence based on a target training temporal-spatial resolution to obtain target training data; inputting the target training data into a teacher model, generating a target token set through dynamic token selection, and generating teacher training features through forward propagation; performing multi-scale cutting on the target token set according to a target self-attention weight of the teacher model to generate at least three student training masks with different token numbers; inputting the target training data and the different student training masks into a student model for forward propagation, and generating student training features; and performing alignment distillation on the student training features and the teacher training features to obtain a target student model. The defect of a video understanding model in downstream flexible reasoning is overcome, and dynamic token selection and multi-scale mask training under high temporal-spatial resolution are utilized, so that the model can obtain better performance under various downstream calculated amount limits.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Sign language translation method and system based on pre-training diffusion large language model

The invention provides a sign language translation method and system based on a pre-training diffusion large language model, and belongs to the field of sign language video translation. The method comprises the following steps: preprocessing a video containing sign language actions to obtain a sign language video frame sequence, inputting the sign language video frame sequence into a visual feature extraction network to extract features, and fusing to obtain a time sequence visual fusion feature sequence; giving a text cue word of a sign language translation task, constructing an initial mask sequence for a target translation position, taking the text cue word, the time sequence visual fusion feature sequence and the initial mask sequence as guide conditions, injecting the guide conditions into a diffusion language model, iteratively denoising and predicting lexical elements of a masked position in combination with a diffusion mask mechanism, and obtaining the sign language translation task. A natural language translation sequence is obtained, and sign language translation is completed; wherein when the diffusion language model is trained, through an internal feature alignment mechanism, the guiding effect of guiding conditions on text generation is optimized, so that the accuracy, coherence and robustness of long text translation are improved, and the actual requirements of a barrier-free public service scene are better met.
Owner:ZHEJIANG UNIV

System and method for identifying subjects and items in an area of real space and calibrating presentation data for the subjects

The technology disclosed relates to a system and methods for providing presentation data to a subject in an area of real space, including obtaining respective sequences of frames of corresponding fields of view in an area of real space; detecting a subject in the area of real space; analyzing a sequence of frames in the respective sequences of space; calibrating presentation data; and triggering a presentation of the calibrated presentation data to the detected subject. Analysis of the sequence of frames includes identifying objects in the area of real space, identifying subject data of the detected subject with respect to the identified objects, and identifying a connection between the identified subject data of the detected subject and a particular identified object. The calibration of the presentation data is dependent on the identified connection between the identified subject data of the detected subject and the particular identified object.
Owner:STANDARD COGNITION CORP

Remote sensing video segmentation method and segmentation system based on text guidance

The invention discloses a remote sensing video segmentation method and segmentation system based on text guidance, belongs to the crossing field of remote sensing image processing and computer vision, and relates to a remote sensing video segmentation method and segmentation system. The invention aims to solve the problems that the existing remote sensing video segmentation technology is poor in flexibility, cannot interact with natural languages, is insufficient in generalization ability for new categories or complex targets, and cannot meet the requirement of quickly and accurately extracting semantic information in a dynamic remote sensing scene. The method comprises the following steps: 1, acquiring a video frame sequence, and acquiring a key frame based on the video frame sequence; 2, obtaining an initial segmentation mask; 3, obtaining an optimized mask; 4, calculating a minimum bounding rectangle of the optimized mask, obtaining a bounding box of the minimum bounding rectangle, and obtaining an expanded bounding box; and 5, inputting the expanded bounding box and the video frame sequence obtained in the step 1 into an improved SAM2 video segmentation model, and outputting a frame-by-frame segmentation result of the region of interest by the improved SAM2 video segmentation model.
Owner:HARBIN INST OF TECH

Unsupervised region-growing network for object segmentation in atmospheric turbulence

An unsupervised region-growing network (RGN) is trained to perform object segmentation on video data degraded by atmospheric turbulence. The method includes obtaining input data containing turbulence-degraded video, extracting a video frame sequence, and training the RGN using a selected algorithm incorporating a region-growing algorithm and a grouping loss function. A bidirectional optical flow sequence is computed for multiple reference frames within the video sequence. Pixel-level masks are generated for detected moving objects, followed by applying the region-growing algorithm to create coarse masks. A grouping loss function refines these masks to ensure consistency across consecutive frames. The trained RGN outputs refined masks as object segmentation data for the received video, improving segmentation accuracy in turbulent environments. This approach enables robust object detection and segmentation without requiring prior video restoration, maintaining fidelity to the original turbulence-distorted input.
Owner:CLEMSON UNIV RES FOUND +2

Video editing method based on grid layout alternate diffusion and multi-attention control

The invention relates to the technical field of video analysis, in particular to a video editing method based on grid layout alternate diffusion and multi-attention control, and the method comprises the steps: segmenting an original video frame sequence into a plurality of grids, each grid comprising a plurality of pixel space video frames which are continuously arranged, and forming grid data; mapping the gridding data to a low-dimensional submerged space through an encoder, and generating initial submerged space feature data; the initial submerged space feature data are edited, the editing process comprises a diffusion process and a sampling process, the diffusion process is based on a pre-trained stable diffusion model, and a time attention module is embedded in the diffusion process; in the sampling process, executing an odd-even time step alternate replacement strategy on the grid layout to promote cross-grid global consistency, and dynamically fusing attention maps of a reconstruction branch and an editing branch according to a timestamp threshold to generate de-noised data; and decoding, splitting and recombining the de-noised data through a decoder to generate an edited continuous video frame sequence.
Owner:ANHUI UNIV

Method and system for quickly generating movie and television animation scene in combination with AI algorithm

The invention relates to the technical field of generative artificial intelligence, and particularly provides a film and television animation scene rapid generation method and system combined with an AI algorithm, and the method comprises the steps: firstly receiving scene description text input, carrying out the semantic analysis, obtaining a scene element set and element association features, and then calling a pre-trained scene resource matching model, matching a scene resource combination in a multi-modal resource library according to an analysis result, then performing spatio-temporal layout optimization on the resource combination, generating a layout adjustment parameter and a dynamic adjustment parameter, constructing a three-dimensional space topological structure according to the layout adjustment parameter and the dynamic adjustment parameter, and rendering to generate a target animation scene frame sequence, and finally, synchronously calibrating the target animation scene frame sequence and a preset plot time axis, and outputting a complete animation scene flow, thereby improving the efficiency and accuracy of movie and television animation scene generation by means of an AI algorithm, and realizing rapid and intelligent scene generation highly conforming to the plot.
Owner:CHENGDU LIFANG VISION TECHNOLOGY CO LTD

Multi-mode audio and video synchronous processing method and system based on distributed architecture

The invention relates to the technical field of computers, and discloses a multi-mode audio and video synchronous processing method and system based on a distributed architecture. The method comprises the following steps: generating a unique logic sequence identifier consisting of a device code, a media type and a frame number for each frame of audio and video data; distributing the frame carrying the identifier to a distributed node for processing and adding a local timestamp; and the sink node reconstructs an original frame sequence according to the identifier, calculates a synchronous offset by combining high-precision clock calibration, and realizes dynamic reordering and output through a double-buffer structure and a self-adaptive drift compensation algorithm. The system comprises a source end acquisition module, an identifier generation module, a task scheduling module, a distributed processing cluster module, a convergence synchronization module and a synchronous output module. According to the scheme, the system effectively guarantees the audio and video synchronization precision in a high-concurrency heterogeneous environment, the synchronization error is controlled within 20 milliseconds, and the playing quality and the user experience of scenes such as live broadcast and cloud games are remarkably improved.
Owner:SHENZHEN HAIWEI HENGTAI INTELLIGENT TECH CO LTD

Multi-modal fusion green port digital management and control system and method

The invention discloses a multi-modal fusion green port digital management and control system and method, and relates to the technical field of port management and control, and the method comprises the following steps: collecting the signal-to-noise ratio data of a radio channel of a video link, and generating a port area electromagnetic interference thermodynamic diagram through spatial interpolation; and dynamically adjusting a synchronous clock of each camera according to the electromagnetic interference thermodynamic diagram, mapping the frame triggering offset into a pixel compensation matrix, and performing sub-pixel-level space-time correction on the image frame to obtain a corrected video frame sequence. According to the method, target credibility judgment and path optimization under multi-modal perception are realized through construction of an interference thermodynamic diagram, space-time correction, artifact recognition and point cloud fusion, an interference model and a recognition threshold are dynamically regulated and controlled through closed-loop feedback, and an end-to-end self-adaptive port digital management and control method is constructed. The artifact identification accuracy and scheduling stability are significantly improved, and the anti-interference and green efficient operation capabilities of the port management and control system are enhanced.
Owner:TIANJIN RES INST FOR WATER TRANSPORT ENG M O T

Intelligent traceless erasing method and system for video picture characters

The invention provides an intelligent traceless erasing method and system for video picture characters, and relates to the technical field of video stream processing, and the method comprises the steps: carrying out the space-time dual-domain feature extraction of a preprocessed first video frame sequence, and obtaining a first space-time feature; and establishing a space-time double-domain attention model according to the first space-time feature, the space-time double-domain attention model calculating fusion space domain and time domain features through cross attention, dynamically distributing fusion weights based on the character movement speed, obtaining a second video frame sequence of the to-be-processed video stream, preprocessing the second video frame sequence, and extracting a second space-time feature. Texture and motion characteristics of front and back frames are synchronously referenced through a space-time double-domain attention model, so that the time sequence coherence of a repaired area is improved, a time sequence consistency compensation mechanism driven by an SSIM is combined, the flicker frequency is reduced, and the time-space domain weight is adaptively adjusted according to the character speed through a dynamic weight distribution strategy, so that the time sequence coherence of the repaired area is improved. And stable erasing and rapid moving of characters are realized.
Owner:XIAN LINGXIANG BIRD CULTURE COMM CO LTD

Large-scene multi-modal analysis system and method based on unmanned aerial vehicle

The invention relates to the technical field of aerial video analysis, and discloses a large-scene multi-modal analysis system and method based on an unmanned aerial vehicle. According to the method, an original video frame sequence is obtained from aerial photography equipment, a scene change period and an object motion trend are extracted by calculating inter-frame motion feature vectors, and whether a scene is in a stable or dynamic state is judged by means of a state classification model. Calculating a tracking target position according to preset parameters in a dynamic state, and adjusting and analyzing a target area in combination with a motion trend; generating a camera parameter adjustment instruction according to the deviation value of the target area and the actual object position, and outputting the camera parameter adjustment instruction to an execution mechanism; acquiring a complexity value through the scene complexity feedback signal, updating an analysis strategy by adopting an adaptive algorithm when the complexity value exceeds a preset range, and outputting an adjustment signal; the output time sequence of control and strategy adjustment signals is optimized, meanwhile, the frame sequence is monitored in real time, the parameter change trend is analyzed, the preset parameters are dynamically updated to meet the user requirements, and efficient video analysis processing is achieved.
Owner:深圳市新创中天信息科技发展有限公司