Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3782 results about "Video based" patented technology

Underground mine operation state analysis system and method based on video monitoring data

The invention discloses an underground mine operation state analysis system and method based on video monitoring data, and the system comprises a data collection module which is used for collecting mine video and environment parameter data through a distributed sensor network, and generating a multi-dimensional data fusion set based on a space-time label technology; the edge analysis module is used for extracting feature parameters through a convolutional neural network algorithm based on the multi-dimensional data fusion set and generating a mine operation state recognition result; the fence construction module is used for constructing a three-dimensional digital model and a dynamic safety boundary based on the mine operation state recognition result to form a real-time monitoring reference framework; and the decision execution module is used for performing hierarchical risk assessment on the monitoring data in the security boundary based on the real-time monitoring reference framework, and generating a security early warning and disposal scheme with a tracing identifier. Each piece of early warning and disposal information is attached with a unique tracing identification code, so that follow-up event backtracking analysis is facilitated, and the risk management and control capability is continuously improved.
Owner:河北省水文工程地质勘查院(河北省遥感中心) +3

Action localization method, device, electronic equipment, and computer-readable storage medium

An action localization method, device, electronic equipment, and computer-readable storage medium are provided. The action localization method includes: identifying at least one target video segment containing a target object in a video; acquiring a first action recognition result of at least one image frame in the at least one target video segment and a second action recognition result of the target video segment; and acquiring an action localization result of the video based on the first action recognition result and the second action recognition result.
Owner:SAMSUNG ELECTRONICS CO LTD

Multi-modal diffusion-based long video role scene decoupling generation method and system

The invention discloses a long video role scene decoupling generation method and system based on multi-modal diffusion, and relates to the technical field of image processing, and the method comprises the steps: S1, synthesizing the advanced features of a role and a scene through a SigLIP encoder and a DINOv2 encoder; s2, performing cross-modal feature fusion on the advanced features to obtain joint features, and compressing the joint features to obtain compact vectors; s3, generating text features according to the text prompt; s4, potential codes are generated from an input video through a causal 3D convolution encoder, the potential codes pass through a linear projection matrix and then are spliced with a memory state for dimension reduction, and a segmented potential vector sequence is obtained; s5, performing decoupling perception generation on the segmented potential vector sequence through an improved 3D-UNet, and performing deconvolution up-sampling reconstruction after deterministic sampling to obtain an RGB video segmented sequence; according to the method, the key problems of rough dynamic control, limited generation length and over-high resource consumption in long video generation are solved, and the quality and efficiency of the generated video are remarkably improved.
Owner:湖南马栏山视频先进技术研究院有限公司

Video analysis-based multi-scene operator violation behavior identification method and system

The invention discloses a video analysis-based multi-scene operator violation behavior identification method and system, and belongs to the technical field of intelligent operation safety monitoring and artificial intelligence identification, and the method comprises the steps: collecting a real-time video stream of a multi-scene operation site; recognizing a continuous action time sequence in the real-time video stream by using an action recognition depth model; constructing the continuous action time sequence into an action behavior sequence; the action behavior sequence is constructed into a directed behavior graph with time, space and action labels, the directed behavior graph is compared with a directed behavior graph corresponding to the standard action behavior sequence, and illegal behaviors are recognized; and carrying out multi-mode early warning on the identified illegal behaviors. According to the method, the bottleneck that the traditional image recognition technology is weak in action sequence semantic understanding and poor in environmental adaptability is broken through, and accurate recognition and real-time early warning of illegal behaviors in multi-scene operation are achieved.
Owner:CHENGDU HANGTIAN PHOTOELECTRIC TECH

Virtual fitting method and device, storage medium and electronic equipment

The invention discloses a virtual fitting method and device, a storage medium and electronic equipment, and is applied to the technical field of computer image processing, and the method comprises the steps: processing a target video, and obtaining the human body posture data of each video frame in the target video; constructing a 3D human body model of each video frame by using the human body posture data of each video frame; for the 3D human body model of each video frame, acquiring clothing rendering information of the 3D human body model, and rendering virtual clothing on the 3D human body model based on a dynamic fitting algorithm and the clothing rendering information to obtain a virtual rendering video frame of the video frame; and generating an AR fitting video based on each virtual rendering video frame, and displaying the AR fitting video to the user. Therefore, the selected costume can be tried on in the AR form, the fitting effect of the costume can be displayed for the user, the problem that the size is improper due to the fact that the costume purchased online cannot be tried on is solved, the probability of refunding and changing goods is reduced, and good shopping experience is provided for the user.
Owner:小芒电子商务有限责任公司

Short video network public opinion information identification method based on image processing technology

The invention discloses a short video network public opinion information identification method based on an image processing technology, and relates to the technical field of artificial intelligence and image processing, and the method comprises the following steps: S001, through obtaining image frames, audio tracks and time sequence information of a short video, constructing a multi-modal fusion model, extracting continuous image frames with suspicious identity features, and carrying out the recognition of the short video network public opinion information; generating a forgery risk area distribution map; and S002, performing semantic consistency verification according to the counterfeit risk region distribution map, and extracting space and time anomaly features existing among facial micro-expressions, pronunciation actions and background semantics in the image frame. According to the method, a multi-modal model is constructed by fusing image, audio and time information, fine abnormal features, traceability forgery starting points and propagation paths in a deep forgery video are identified, and an identification strategy and a public opinion response mechanism are dynamically adjusted, so that accurate identification, adaptive processing and closed-loop control of short video public opinion risks are realized; and the identification accuracy and the treatment efficiency are improved.
Owner:TIBET UNIV

Semantic analysis of video data for event detection and validation

Methods, systems, and computer programs are presented to perform semantic analysis of video data for event detection and validation. One method includes an operation for training a machine-learning model with text-image pairs to create a semantic model. The semantic model generates text embeddings based on text describing events and image embeddings from video frames. The system calculates similarity values between text and image embeddings to determine event occurrences when the similarity exceeds a predetermined threshold. The system enables precise event detection by analyzing object relationships and attributes within video frames, reducing false alarms and enhancing monitoring efficiency. Further, these techniques can be used to detect any event describable by human text or by some image examples.
Owner:FIRST-CITIZENS BANK & TRUST CO

Automatic lecturer video generation method based on AI speech synthesis and animation driving

The invention discloses a lecturer video automatic generation method based on AI speech synthesis and animation driving. The method comprises the following steps: performing structured analysis on a PPT or a text script through an improved interior point method and an incremental shortest path algorithm; performing semantic grouping by applying a full-dynamic parallel single-link clustering algorithm and generating an enhanced script with an expressive mark; a CosyVoice technology is combined with a low-rank approximation method to generate a high-quality voice data stream; establishing a mapping relation between contents and action expressions through semantic analysis, and generating a complete action expression instruction set; and driving the digital human model by using the msueTalk technology, and generating a final lecturer teaching video through a parallel rendering algorithm. According to the invention, the method achieves the efficient and automatic generation of the education video, remarkably improves the content production efficiency, reduces the production cost, and guarantees the specialty and expressive force of the teaching video.
Owner:SHENZHEN XUEYOU TECHNOLOGY CO LTD

Panoramic video three-dimensional point cloud sparse reconstruction method and device integrating key frame screening

The invention provides a panoramic video three-dimensional point cloud sparse reconstruction method and device integrated with key frame screening, and the method comprises the steps: obtaining a panoramic video of a target scene, and carrying out the video quality evaluation of the panoramic video, and obtaining a video evaluation result; determining a target resolution of downsampling according to the video evaluation result; carrying out downsampling processing on the panoramic video based on the target resolution to obtain a downsampled second video; for the second video, dynamically adjusting a frame extraction interval based on a preset key frame selection mechanism, and extracting a frame index of the key frame from the second video according to the frame extraction interval to obtain a plurality of frame indexes; and obtaining a plurality of corresponding key frames in the panoramic video based on the plurality of frame indexes, and obtaining the three-dimensional scene sparse point cloud of the target scene based on the plurality of key frames, thereby solving the technical problem of low processing efficiency in the prior art.
Owner:TIANJIN FIRE SCI & TECH RES INST OF MEM

Non-inductive identity recognition and real-time tracking method based on personnel in video

The invention discloses a non-inductive identity recognition and real-time tracking method based on personnel in a video. The method comprises the following steps: S1, acquiring video stream data and extracting a key frame image; s2, key point coordinates are extracted, and a time sequence skeleton sequence is constructed; s3, decomposing gait, action and attitude features to generate standardized features; s4, inputting the standardized features into an improved multi-scale image convolutional neural network, extracting identity representation features, and generating a target identity code; s5, matching identity codes based on similarity measurement, calculating a consistency score, and establishing an identity tracking trajectory; s6, tracking is carried out in combination with identity codes and consistency scores, identity mapping is dynamically updated, and non-inductive recognition is achieved; and S7, updating a track in real time, adjusting a tracking state, and ensuring identity stability. According to the method, the improved graph convolutional network and the consistency measurement technology are combined, the identity recognition precision and tracking stability are improved, and cross-scene non-inductive identity recognition is achieved.
Owner:BEIJING LIYANG ZHIGUANG TECH CO LTD

Commodity video intelligent generation method based on multi-modal analysis and dynamic narrative architecture

The invention relates to the technical field of video generation, in particular to an intelligent commodity video generation method based on multi-modal analysis and a dynamic narrative architecture, which comprises the following steps: inputting original commodity data and a user behavior log, and outputting a dynamic selling point weight vector through a selling point value evaluation engine; s2, inputting the dynamic selling point weight vector output in S1 into a narrative flow generator, and executing the following steps: intercepting first K core selling points according to the weight vector to form a narrative trunk chain; based on the historical preference data of the user, inserting an emotion enhancement node to generate a branch enhancement narrative flow; generating a multi-version narrative flow instruction set in combination with the real-time network bandwidth data; retrieving a preset video clip library according to the instruction set to generate an initial video sequence; detecting semantic fault regions between adjacent segments to generate compensation animation parameters; and injecting compensation animation parameters and outputting a continuous narrative video stream. According to the invention, the jamming feeling and the frame skipping phenomenon caused by the change of a narrative structure or different sources of video clips are reduced, and the overall watching experience of a user is improved.
Owner:SHENZHEN YINGMENG INTELLIGENT TECHNOLOGY CO LTD

Safety production behavior monitoring method and system based on AI video analysis

The invention provides a safety production behavior monitoring method and system based on AI video analysis. The method comprises the steps of collecting a real-time video data stream of a production area; inputting each frame of video image in the real-time video data stream into a target detection model for target detection to obtain a personnel target output by the target detection model and a target position coordinate of the personnel target in each frame of video image; cutting out a local image area of the personnel target in each frame of video image based on the target position coordinate, and performing feature recognition based on the local image area to obtain personnel features; performing comparison on the basis of the personnel characteristics and the personnel standard behavior characteristics to obtain personnel behavior states, and performing track association on the basis of the personnel behavior states corresponding to the continuous multi-frame video images to obtain personnel behavior tracks; and performing safety production behavior monitoring based on the personnel behavior state and the personnel behavior track, and generating an abnormal behavior early warning signal. According to the method and the device, the real-time performance and the accuracy of safety monitoring in a production scene are improved.
Owner:SHENZHEN YINXING INTELLIGENT DATA CO LTD

Video quality detection method based on multi-scene self-adaption

The invention provides a video quality detection method based on multi-scene self-adaption. The method comprises the following steps: selecting an area in a video picture, extracting multi-dimensional features of the area, determining the type of a scene where the video picture is located, and dynamically adjusting video detection parameters of the video picture; detecting the quality of the video picture based on the video detection parameters; intercepting an abnormal video picture, and adding the abnormal video picture into the constructed equipment fault template library M; and performing local model retraining on the video quality detection model based on the equipment fault template library M, and detecting the quality of a to-be-detected video picture based on the video quality detection model. On the basis of deep analysis of different scene features, the scene type of the video picture is accurately judged through multi-dimensional feature calculation, and then the video quality detection parameters are dynamically adjusted according to the scene type, so that the detection system can automatically adapt to the optimal detection parameters according to scene changes, misjudgment and missed judgment caused by the scene changes are avoided, and the detection efficiency is improved. And the accuracy of video quality detection in different scenes is remarkably improved.
Owner:CHINA SHIPBUILDING LINGJIU HIGH TECH (WUHAN) CO LTD +1

Intelligent driving behavior identification method and system based on video analysis

The invention provides a driving behavior intelligent identification method and system based on video analysis, and the method comprises the steps: obtaining a driver face video stream and a road environment video stream collected by a vehicle-mounted camera, and reading the driving information recorded by a whole vehicle communication network; recognizing an eyelid closing state, a sight line direction and a head posture in the driver face video stream based on a posture recognition model, and performing fatigue distraction analysis to obtain driver state information; performing motion trail analysis on the road environment video stream and the driving information, and performing driving risk assessment in combination with the driver state information to obtain driving assessment information; and performing early warning construction according to the driving evaluation information, generating early warning prompt information, and synchronously writing the early warning prompt information, the driving evaluation information and the driver state information into a safety data protection memory. The fatigue and distraction states of the driver can be recognized more accurately, and the accuracy of state judgment is improved.
Owner:SHENZHEN ZHIJU CLOUD SERVICE TECH CO LTD

Interactive automatic explanation method for converting traditional video into artificial intelligence digital human

The invention provides an interactive automatic explanation method for converting a traditional video into an artificial intelligence digital human, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining original video data and an audio track, and carrying out the semantic analysis of the audio track, and obtaining multi-mode deconstruction data; generating an explanation script for each time period of the video based on the explanation text, and performing timestamp labeling on the visual elements to form a time sequence synchronization data structure; in the playing process, a virtual image generator is driven to synthesize digital human dynamic expression output in real time according to the current playing time point; after a user interruption request is received, semantic matching is carried out on a query intention in the explanation script, a target explanation fragment and visual elements are positioned, and complementary explanation content is generated; and driving the virtual image generator to synthesize dynamic output synchronized with the supplementary explanation, and after interaction is completed, recovering playing or skipping to a specified time point according to a user instruction. According to the invention, the conversion from the traditional video to the interactive intelligent explanation video is realized, and the watching experience and learning efficiency of the user are improved.
Owner:BEIJING MENGKE TECH CO LTD

Virtual fitting video generation method and system based on multi-view face fixation

The invention discloses a virtual fitting video generation method and system based on multi-view face fixation, and relates to the field of generative artificial intelligence and computer vision, and the virtual fitting video generation method based on explicit geometric constraints comprises the following steps: S1, obtaining an original fitting image and an action cue word, and constructing a multi-source input data set; s2, segmenting the original fitting image to obtain multi-view modeling, and generating a multi-angle face image and a clothing texture feature parameter; s3, analyzing the action cue word to generate a target posture sequence, and generating head and tail frame virtual fitting images; s4, performing pairing analysis on the virtual fitting images of the head frame and the tail frame to generate attitude transition parameters; s5, performing dynamic texture repair on the image to generate a transition frame image sequence; and S6, generating a virtual fitting video and outputting the virtual fitting video. According to the method, accurate segmentation and multi-angle modeling of the face area and the clothing area are realized through image segmentation and a rigid geometric transformation algorithm.
Owner:QINGDAO UNIV

Al-based video content analysis method and system

The invention relates to the technical field of content recognition, in particular to an Al-based video content analysis method and system, and the method comprises the following steps: carrying out the image segmentation learning based on a video sequence through employing a U-Net convolutional neural network, analyzing scenes and elements in a video frame, recognizing and isolating key visual elements in a video through network learning, and carrying out the recognition of the key visual elements in the video. Comprising objects and figures. According to the method, image segmentation learning is carried out by adopting the U-Net convolutional neural network, key visual elements in the video can be identified and isolated more accurately, a clearer basis is provided for follow-up scene change and key event tracking, video content analysis is carried out by applying the graph neural network, the identification capability of dynamic scenes and events is enhanced, and the identification efficiency is improved. Deeper structured understanding is provided for video content indexing, the video quality is analyzed through a structural similarity index evaluation method, and the reason of visual quality reduction can be accurately recognized and improved.
Owner:CHENDA (GUANGZHOU) NETWORK TECH CO LTD +1

Flow monitoring method and system based on video image analysis and processing

The invention relates to the technical field of water flow monitoring, and particularly discloses a flow monitoring method and system based on video image analysis and processing. The method comprises the following steps: carrying out communication acquisition of image monitoring and water level monitoring at a plurality of monitoring acquisition sites; performing video frame extraction and image preprocessing; feature point detection and displacement analysis; converting the water surface flow velocity into a plurality of station water surface flow velocities, and performing fitting correction; and calculating station flow data of the plurality of monitoring acquisition stations. Water surface image data are obtained through image monitoring and water level monitoring communication collection of a plurality of monitoring collection sites in combination with video frame extraction and image preprocessing, the site pixel flow velocity is calculated through feature point detection and displacement analysis, the site pixel flow velocity is converted into corrected water surface flow velocity through a space projection model, and site flow data are calculated. Non-contact full-flow monitoring is achieved, the potential safety hazard that equipment needs to be arranged in water in a traditional method is eliminated, the maintenance difficulty is remarkably reduced, and the method is particularly suitable for flow monitoring in flood periods and dangerous water areas.
Owner:JIANGXI SHANLIU HUILIAN TECHNOLOGY CO LTD

Key frame extraction method and device based on dynamic reinforcement learning, equipment and medium

The invention relates to the technical field of computer vision, can be applied to the medical field and the financial science and technology field, and discloses a key frame extraction method, device and equipment based on dynamic reinforcement learning and a medium, which are applied to electronic application in a high-frequency transaction abnormal behavior monitoring scene or can be applied to a medical operation key frame extraction scene. The method comprises the steps of obtaining an original video stream and performing preprocessing to generate a standardized video frame; performing feature extraction and feature splicing on the standardized video frame, and performing time sequence modeling on the generated pair frame-level mixed feature vector to generate video-level time sequence representation; generating an enhancement action instruction based on the video-level time sequence representation through the strategy network, and performing enhancement processing on the standardized video frame according to the enhancement action instruction to generate an enhanced video frame; performing optimization processing on the strategy network according to the enhanced video frame to generate an updated strategy network; and performing feature extraction and optimization on the enhanced video frame to generate a target key frame. According to the invention, the key frame extraction precision is improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Method, system and equipment for processing abnormal black screen during vehicle-machine interconnection

The invention provides an abnormal blank screen processing method, system and equipment during vehicle-machine interconnection, and the method comprises the steps: predicting the current abnormal blank screen risk probability based on vehicle-machine interconnection parameters, carrying out the abnormal blank screen detection through employing a preset multilayer abnormal blank screen detection algorithm, and when the abnormal blank screen risk probability reaches or exceeds a preset probability threshold value, carrying out the abnormal blank screen detection. And when the abnormal blank screen is detected, controlling the display screen to display a preset picture or video, or when the abnormal blank screen is detected, determining an abnormal blank screen type, and controlling the display screen to display the preset picture or video based on the abnormal blank screen type. According to the method provided by the invention, the technical defects of lack of blank screen risk prediction and active prevention and control capabilities and lack of prevention, control and recovery of various types of blank screens when the existing vehicle-mounted system and the terminal equipment are interconnected are effectively solved.
Owner:CHENGDU DESAY SV KAWA TECHNOLOGY CO LTD

Video data processing method and device, equipment and medium

The invention relates to the field of video processing, in particular to a video data processing method and device, equipment and a medium. In a background replacement link, based on precise operation of video processing requirements, an adaptive algorithm can be selected according to scene characteristics, and errors are preliminarily reduced. And subsequently, error compensation is carried out on the generated intermediate video data, the error region is corrected in a targeted manner, the edge is filled and optimized by utilizing the edge pixel characteristics of the foreground region, and iteration processing is carried out until the error is lower than a preset value, so that the accuracy of background replacement is greatly improved. And finally, the target video data is injected into the virtual camera of the cloud mobile phone, the method is applied to a mobile terminal scene, the background and the foreground are naturally fused under high-precision scenes such as live broadcast and virtual conferences, the image flaws are remarkably reduced, the strict requirements of a user on the video image quality are met with a high-precision background replacement effect, and the overall user experience is improved.
Owner:启朔(深圳)科技有限公司

Long video target inference segmentation method based on context mark prompt

The invention belongs to the technical field of image segmentation, and discloses a context mark prompt-based long video target reasoning segmentation method, which comprises a pre-training image encoder, a multi-layer perceptron mapping module, a multi-modal feature fusion module, a large language model and a mask spreading device. The method comprises the following steps: sampling support frames from equally divided video clips, and processing the support frames and key frames together through a pre-trained image encoder and a multi-layer perceptron mapping module to obtain corresponding visual features; the multi-modal feature fusion module injects the visual features of the reference expression and the support frame into potential queries through a plurality of fusion modules to generate enriched potential queries; the enriched potential queries guide the large language model to generate key frames and full video level lt; sEGgt, SEGgt; and finally, the SAM2-based mask spreading device is used for accurately decoding and continuously and consistently spreading the SAM2-based mask spreading device in all frames. According to the method, the problems of long-distance dependence modeling and consistency tracking are solved through context mark prompt and a multi-modal feature fusion module.
Owner:DALIAN UNIV OF TECH

Text-guided video generation

A method, apparatus, non-transitory computer readable medium, and system for video generation include obtaining an input image having an element depicted in a first view angle, generating a synthetic image depicting the element of the input image from a second view angle different from the first view angle, generating an intermediate image by interpolating based on the synthetic image, and generating a video based on the synthetic image and the intermediate image, where the video depicts the element of the input image from a changing view angle.
Owner:ADOBE INC

System and Method for Event-Driven Video Synthesis Using Textual Descriptions

A video generation framework that is controllable, unsupervised and based on events (CUBE) includes an event camera, which captures changes in light intensity at each pixel of a scene asynchronously and generates event camera data. A text-to-image diffusion model that is conditioned on textual descriptions integrates the event camera data to control video synthesis. Further, an edge extraction module translates event data into a format usable by the text-to-image diffusion model, whereby the diffusion model synthesizes detailed and contextually accurate videos based on textual prompts. Further, an improved system (CUBE Plus) includes a content frame identification module which selectively identifies and uses only the most information-rich event segments of the event camera data to drive cross-frame attention, and an event driven attention mechanism that allows the framework to focus on event-dense moments.
Owner:THE UNIVERSITY OF HONG KONG

Video-based unsupervised visible light infrared pedestrian re-identification method

The invention belongs to the field of pedestrian re-identification, and relates to a video-based unsupervised visible light infrared pedestrian re-identification method, which comprises the following steps: acquiring query data and a data set, inputting the query data and the data set into a trained re-identification model to obtain query features and a feature set, and matching the query features with the feature set to obtain an identification result; the training process of the re-identification model comprises the following steps: acquiring visible light data SV and infrared data ST; inputting the SV and the ST into a feature extraction module to obtain visible light and infrared features FV and FT; inputting the FV and the FT into a clustering module to obtain a clustering result; inputting the clustering result into a progressive false label correction module to obtain a corrected clustering result; inputting the SV and the ST into a feature extraction module to obtain visible light and infrared features qV and qT; updating model parameters according to the qV, the qT and the corrected clustering result until a trained re-identification model is obtained; according to the method, noise samples are recovered into effective labels through intra-modal correction and inter-modal correction, and robustness is enhanced.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Video generation method and device, electronic equipment, storage medium and program product

The invention provides a video generation method and device, electronic equipment, a storage medium and a program product, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes of digital human, content generation based on artificial intelligence and the like. The method comprises the following steps: acquiring text features of a description text, image features of a virtual image in a reference image and audio features of audio, wherein the description text indicates action description information for driving the virtual image based on the audio; the role feature is bound with the audio feature to obtain a target audio feature, the role feature is associated with a corresponding virtual image in the reference image, and the role feature is used for indicating an association relationship between the target audio feature and the image feature; and generating a target video based on the text feature, the image feature and the target audio feature, wherein the target video comprises a plurality of video frames for making a sound according to the action description information based on the audio driving virtual image.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Multi-source information emotion recognition method based on common attention

The invention discloses a multi-source information emotion recognition method based on common attention, and the method comprises the steps: 1, carrying out the preprocessing of video data, separating an audio stream from a video stream, and obtaining a video frame sequence xv and audio data xa; step 2, manually extracting an MFCC acoustic feature xm from the audio data xa as a part of input; 3, based on the input of the video frame sequence xv, the audio data xa and the MFCC acoustic feature xm, modeling and coding are carried out on the video frame sequence xv, the audio data xa and the MFCC acoustic feature xm respectively, and feature representations of deeper levels are obtained; and step 4, classifying the data processed in the step 3 through a full connection layer and a Softmax function to obtain an emotion recognition classification result. According to the method, the acoustic features and the semantic features are extracted from the voice, the global features are concerned in the video, the attention weights are generated by the acoustic and visual features to act on the semantic features, and the emotion recognition effect is better by using the cooperative relationship between different information sources.
Owner:NORTHWEST UNIV

Multi-channel dynamic hypergraph sentiment analysis method and analysis network fusing time sequence consistency

The invention discloses a multi-channel dynamic hypergraph sentiment analysis method and a multi-channel dynamic hypergraph sentiment analysis network fusing time sequence consistency, belongs to the field of artificial intelligence and multi-modal sentiment calculation, and aims to solve the problems existing in the existing sentiment analysis technology. The method comprises the following steps: S1, a multi-channel feature extraction step: extracting multi-channel features of a text mode and an audio mode through a heterogeneous pre-training model; s2, a local time sequence context fusion step based on a video number: fusing short-term emotional fluctuation based on a local context mechanism of the video number, and capturing long-range dependence across time dimensions through Transform; s3, a single-modal-multi-modal hypergraph collaborative prediction step: dynamically constructing a single-modal hypergraph and a multi-modal hypergraph in a training batch, and modeling a high-order relationship by adopting spectral domain-spatial domain hybrid convolution; and S4, a multi-level multi-branch supervision step: outputting a final emotion prediction result through joint optimization of an early MLP branch and a late hypergraph branch.
Owner:HARBIN INST OF TECH

Video understanding method and device and computer program product

The embodiment of the invention provides a video understanding method and device and a computer program product, and belongs to the field of videos and big data, and the method comprises the steps: obtaining a description text, a query text token and a visual token of a target video; lLM reasoning, retrieval and region expansion are carried out on the description text to obtain candidate time regions; dense sampling, retrieval and region merging are carried out on the candidate time regions to obtain continuous time regions; determining an attention matrix according to the query text token and the visual token; performing semantic relevancy evaluation, pruning and position coding reconstruction on the continuous time region according to the attention matrix to obtain spatial-temporal characteristics; and generating an answer of the target video according to the spatial-temporal characteristics. According to the method, efficient and high-precision long video understanding is realized.
Owner:AMWAY HUASHENG DATA TECH (JIANGSU) CO LTD

Four-dimensional scene reconstruction method and apparatus, and electronic device

Embodiments of the present application disclose a four-dimensional scene reconstruction method and apparatus, and an electronic device. A specific implementation of the method includes: obtaining a multi-view video, where the multi-view video includes multi-view images, which include a video frame at an initial moment in the multi-view video; generating a three-dimensional scene model corresponding to the multi-view images; determining a deformable network corresponding to the multi-view video based on the three-dimensional scene model, the multi-view video, and camera pose information corresponding to the multi-view video; and determining a four-dimensional scene model corresponding to the multi-view video based on the three-dimensional scene model and the deformable network.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD