Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

754 results about "Video quality" patented technology

Video quality is a characteristic of a video passed through a video transmission/processing system, a formal or informal measure of perceived video degradation (typically, compared to the original video). Video processing systems may introduce some amount of distortion or artifacts in the video signal, which negatively impacts the user's perception of a system. For many stakeholders such as content providers, service providers, and network operators, the assurance of video quality is an important task.

Monitoring video enhancement method for farm

The invention belongs to the technical field of video processing, and particularly relates to a monitoring video enhancement method for a farm, which aims to solve the technical problem of low quality of an enhanced video in the prior art, and comprises the following steps: S1, processing each frame of image in a monitoring video sequence frame by frame; s2, distinguishing a target animal area from a background area, and identifying and generating an artifact mask; s3, aiming at the background area, carrying out key smoothing processing on the artifact position to inhibit the artifact; s4, for the image of the target animal area, performing adaptive nonlinear enhancement on the brightness component, and performing color correction on the chrominance component; s5, performing pixel-level fusion on the enhanced target animal area and the background area; and S6, spreading the information of the previous frame to the current frame by using the forward optical flow field, and carrying out weighted fusion on the information of the previous frame and the current frame. According to the method, the target bred animals, the background areas and the artifacts are accurately distinguished, so that refined and differentiated processing of pictures is realized.
Owner:EGG NO 1 FOOD CO LTD

Video content quality analysis and knowledge recommendation method and system based on large model

The invention belongs to the technical field of biological medicine, and particularly relates to a video content quality analysis and knowledge recommendation method and system based on a large model, and the method comprises the steps: carrying out the video obtaining through an input keyword based on a short video platform, and obtaining a video ID, a video content text and video metadata; based on the video content text, multi-dimensional text information extraction and problem diagnosis are carried out, and a structured video text is generated; performing multi-dimensional scoring on the structured video text, calculating to obtain a video quality score, and constructing an initial high-quality video candidate set; based on the initial high-quality video candidate set and the user input question, calculating a similarity score of the user input question and the initial high-quality video candidate set, and constructing a high-quality video candidate set; and in combination with the video quality score and the similarity score, carrying out weighted calculation on a comprehensive recommendation score of each video in the high-quality video candidate set, and outputting the first three recommended videos and recommendation reasons to realize precise recommendation.
Owner:湖南工商大学

Panoramic video three-dimensional point cloud sparse reconstruction method and device integrating key frame screening

The invention provides a panoramic video three-dimensional point cloud sparse reconstruction method and device integrated with key frame screening, and the method comprises the steps: obtaining a panoramic video of a target scene, and carrying out the video quality evaluation of the panoramic video, and obtaining a video evaluation result; determining a target resolution of downsampling according to the video evaluation result; carrying out downsampling processing on the panoramic video based on the target resolution to obtain a downsampled second video; for the second video, dynamically adjusting a frame extraction interval based on a preset key frame selection mechanism, and extracting a frame index of the key frame from the second video according to the frame extraction interval to obtain a plurality of frame indexes; and obtaining a plurality of corresponding key frames in the panoramic video based on the plurality of frame indexes, and obtaining the three-dimensional scene sparse point cloud of the target scene based on the plurality of key frames, thereby solving the technical problem of low processing efficiency in the prior art.
Owner:TIANJIN FIRE SCI & TECH RES INST OF MEM

Video stream adaptive low-delay real-time transmission method and system based on edge calculation

The invention discloses a video stream adaptive low-delay real-time transmission method and system based on edge calculation. The method comprises the following steps: receiving a real-time video stream from a network camera, creating a pipeline queue, adding timestamp information for each video frame, and setting a queue protection mechanism; the coded video frames are taken out from the input queue, and the frames in the video are processed through hardware acceleration decoding; the resource use condition of the system is monitored in real time; executing a self-adaptive frame skipping decision according to a performance monitoring result; timestamp generation: dynamically calculating a timestamp interval according to an actual processing frame rate; receiving the decoded original video frame and the corresponding timestamp information, accelerating decoding by using hardware, and executing a video coding operation; and packaging and transmitting the coded video data, and providing a standard protocol interface to be connected with a client for playing. According to the scheme, stable low delay and relatively low resource occupation can be kept, and meanwhile, the video quality is remarkably improved.
Owner:SICHUAN WEIBANG XINCHUANG TECH CO LTD

Video quality detection method based on multi-scene self-adaption

The invention provides a video quality detection method based on multi-scene self-adaption. The method comprises the following steps: selecting an area in a video picture, extracting multi-dimensional features of the area, determining the type of a scene where the video picture is located, and dynamically adjusting video detection parameters of the video picture; detecting the quality of the video picture based on the video detection parameters; intercepting an abnormal video picture, and adding the abnormal video picture into the constructed equipment fault template library M; and performing local model retraining on the video quality detection model based on the equipment fault template library M, and detecting the quality of a to-be-detected video picture based on the video quality detection model. On the basis of deep analysis of different scene features, the scene type of the video picture is accurately judged through multi-dimensional feature calculation, and then the video quality detection parameters are dynamically adjusted according to the scene type, so that the detection system can automatically adapt to the optimal detection parameters according to scene changes, misjudgment and missed judgment caused by the scene changes are avoided, and the detection efficiency is improved. And the accuracy of video quality detection in different scenes is remarkably improved.
Owner:CHINA SHIPBUILDING LINGJIU HIGH TECH (WUHAN) CO LTD +1

Al-based video content analysis method and system

The invention relates to the technical field of content recognition, in particular to an Al-based video content analysis method and system, and the method comprises the following steps: carrying out the image segmentation learning based on a video sequence through employing a U-Net convolutional neural network, analyzing scenes and elements in a video frame, recognizing and isolating key visual elements in a video through network learning, and carrying out the recognition of the key visual elements in the video. Comprising objects and figures. According to the method, image segmentation learning is carried out by adopting the U-Net convolutional neural network, key visual elements in the video can be identified and isolated more accurately, a clearer basis is provided for follow-up scene change and key event tracking, video content analysis is carried out by applying the graph neural network, the identification capability of dynamic scenes and events is enhanced, and the identification efficiency is improved. Deeper structured understanding is provided for video content indexing, the video quality is analyzed through a structural similarity index evaluation method, and the reason of visual quality reduction can be accurately recognized and improved.
Owner:CHENDA (GUANGZHOU) NETWORK TECH CO LTD +1

CDN-based video downlink weak network adaptive image quality optimization method and system

The invention discloses a CDN-based video downlink weak network adaptive image quality optimization method and system, and relates to the technical field of video content distribution. The method comprises the following steps: S1, generating a basic compensation amount according to frame interval deviation and frame content complexity; s2, calculating a rendering overhead index and a rendering urgency threshold according to the state of the multi-source equipment, and generating an interpolation correction coefficient for frame rendering time sequence adjustment; s3, constructing a channel damage quantization map segment trigger redundancy control and introducing an equipment load state constraint; and S4, basic waiting time is generated according to local historical features, and a final waiting time value is generated in combination with the playing buffer area state and the composite load factor. Through cooperation of frame time sequence compensation enhancement, rendering scheduling optimization and segmentation redundancy control strategies, improvement of video image quality stability and end side load adaptability in a weak network environment is realized.
Owner:GUANGZHOU GENGYUN NETWORK TECHNOLOGY CO LTD

5G high-definition video monitoring method and system for smart city and medium

The invention relates to the technical field of video monitoring communication, in particular to a 5G high-definition video monitoring method and system for a smart city, and a medium, and the method specifically comprises the steps: obtaining a 5G high-definition video based on the moving features of each feature point of a moving target between each frame and a previous frame image in a monitoring video, and the gray gradient change condition in each channel image of an RGB image of each video frame; constructing an information rich characteristic value of each video frame; calculating a bandwidth influence coefficient of each detection time period based on a data fluctuation abnormal condition of the network bandwidth data in each preset detection time period; constructing a video transmission quality coefficient by combining the features, optimizing a quantization parameter of a video coding algorithm in the current detection time period, carrying out coding transmission on the monitoring video in the current detection time period, and carrying out intelligent analysis on the transmitted monitoring video; the problem of frame loss or low video quality of video content with relatively high potential information value is reduced; and the interference on subsequent video monitoring intelligent analysis and target detection can be reduced.
Owner:HARBIN HANCAI TECH CO LTD

Intelligent network control method for low-delay video return and related equipment

The invention relates to the field of multimedia communication and network control, in particular to an intelligent network control method for low-delay video return and related equipment. The intelligent network control method comprises the following steps: acquiring network key indexes including bandwidth, delay, jitter and packet loss rate of a network link in real time, and providing real-time network environment data support for transmission strategy adjustment. According to the method, an intelligent control mechanism combining network state perception and video content feature recognition is constructed, so that the technical problem of low-delay video return in a complex network environment is effectively solved. Specifically, key indexes of a network link are collected in real time, and a lightweight CNN model is introduced to analyze the video content activeness, so that dual perception capabilities for a network environment and content features are formed, data support is provided for dynamic adjustment of coding parameters, and accurate balance between video quality and network adaptability is realized.
Owner:IFREECOMM TECH CO LTD

Audio and video low-delay return method and system in extreme environment

The invention relates to the technical field of audio and video emergency transmission, and discloses an audio and video low-delay return method and system in an extreme environment. The method comprises the following steps: acquiring original multi-modal data of audio and video acquisition equipment in a target area, and analyzing a data state of the original multi-modal data; and meanwhile, available network transmission links are monitored, and the quality is evaluated. Self-adaptive coding parameters are generated in combination with the link quality and the data state, dynamic coding is executed on original data, and a coding stream suitable for redundant transmission is formed. And carrying out cooperative distribution transmission based on the network link set and the coded stream to generate return data. According to the generation process, a device and energy management instruction is formed and issued to the acquisition and relay device. According to the method, the network state and the content characteristics are optimized in a dynamic coding link in a collaborative manner, and the audio and video quality, the real-time performance and the system energy efficiency which are transmitted back in an extreme environment are improved through a feedback closed loop for transmitting a result to equipment management.
Owner:XIAN YUNKAI INFORMATION TECHNOLOGY CO LTD

Method, device and equipment for quality evaluation and three-dimensional reconstruction of public-source geographic video data

The invention relates to a crowd-sourced geographic video data quality evaluation and three-dimensional reconstruction method, device and equipment. The method comprises the following steps: acquiring a to-be-processed video frame sequence; the to-be-processed video frame sequence is obtained by preprocessing public source geographic video data; constructing a video quality evaluation network; the video quality evaluation network comprises a pre-trained dynamic object detection module, a lens conversion detection module and a scene detection module; and processing the to-be-processed video frame sequence by using the video quality evaluation network, splicing the continuous scene frames output by the shot conversion detection module and the shot transition frames marked by the scene detection module according to timestamps to obtain a qualified video frame sequence, and performing three-dimensional reconstruction according to the qualified video frame sequence. By adopting the method, the quality and availability of the public-source geographic video data can be improved, and the precision and efficiency of three-dimensional reconstruction are improved.
Owner:TIANJIN INST OF ADVANCED TECH +2

Digital human mouth broadcast video generation method, system and device and medium

The invention discloses a method, a system and equipment for generating an oral video of a digital human, and a medium, and the method comprises the steps: obtaining oral copywriting and video material data, and analyzing and determining a timestamp of the copywriting in a video material through a multi-mode large model; converting the copywriting into audio data, preprocessing the audio data, and combining the audio data with the timestamp to generate first video data; and generating a digital human according to a user demand, and combining the digital human with the first video after image matting processing to obtain an oral playing video. According to the method, a traditional template generation mode is broken through, and customized production of the digital human broadcast video is realized through a multi-modal semantic matching and personalized digital human generation technology; and meanwhile, audio and video accurate synchronization, high-quality image matting and synthesis technologies are adopted, so that the content adaptability and the video quality are guaranteed, and the flexibility, the efficiency and the effect of digital population broadcast video production are remarkably improved.
Owner:GUANGZHOU TAIDONG TECH CO LTD

Video quality evaluation method based on spatio-temporal feature fusion

The invention belongs to the technical field of video quality evaluation, and particularly relates to a video quality evaluation method based on spatio-temporal feature fusion. Comprising the following steps: acquiring a video quality evaluation data set, preprocessing the video quality evaluation data set, and processing preprocessed video data by adopting a time feature extraction module to obtain short video features and long video features; processing the preprocessed video data by adopting a spatial feature extraction module to obtain high-level semantic features and distortion features; fusing the short video features, the long video features, the high-level semantic features and the distortion features by using a feature fusion module to obtain fused features; inputting the fusion feature into a multi-layer perceptron for processing to obtain a video quality score; calculating the total loss of the model and continuously adjusting model parameters according to the total loss of the model to obtain a trained video quality evaluation model; performing video quality evaluation by using the trained model; according to the method, the correlation between the spatial-temporal features and the inter-frame features is effectively utilized, so that the accuracy of quality evaluation is improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Image-based video generation method and device, equipment and storage medium

The invention relates to the field of artificial intelligence, financial science and technology and digital medical treatment, and discloses an image-based video generation method and device, equipment and a storage medium. The method comprises the following steps: receiving and preprocessing an input static image to generate a multi-scale feature; based on the multi-scale features, determining an optimal space-time processing path through differentiable search of a space-time architecture generator, and generating output features fused with time sequence dynamic information; generating a video frame based on the time sequence recurrent neural network and the output features; and inputting the video frames into the video frame sequence, and generating a target video through the video frame sequence. According to the method, the optimal space-time processing path can be automatically determined through differential search, manual intervention is avoided, the time sequence coherence and detail authenticity of the video are ensured by dynamically generating the output characteristics of the fusion time sequence information and generating the video, the video generation quality and efficiency are improved, and the video quality is improved. The method is suitable for high-precision video generation such as video synthesis, content creation and virtual reality in the fields of finance and medical treatment.
Owner:PING AN TECH (SHENZHEN) CO LTD

Energy-saving video adaptive bit rate optimization method based on deep reinforcement learning

The invention relates to the technical field of streaming media, in particular to an energy-saving video adaptive bit rate optimization method based on deep reinforcement learning, which introduces a video quality evaluation method of video multi-method fusion evaluation in the optimization process, effectively evaluates the perception quality of a video at different bit rates, and improves the video quality. It is ensured that the user obtains the most real picture experience in the watching process; meanwhile, an energy consumption perception model is provided by combining a neural network of an Actor-Critic architecture, and the energy consumption condition of equipment in the video playing process is estimated based on the energy consumption of video data downloading and video rendering; besides, by improving an entropy updating strategy of a near-end strategy optimization algorithm, multi-dimensional information such as network conditions and user preferences is combined more effectively, the bit rate decision of the video stream is optimized, and the video quality and the energy consumption are balanced. According to the invention, on the premise that the user experience quality is ensured, the equipment energy consumption is obviously reduced, so that the dual optimization of the video stream transmission quality and the energy consumption is realized.
Owner:HUZHOU UNIVERSITY

Three-dimensional grounded video generation

Systems and methods are disclosed related to a 3D grounded video foundation model. A video generation method and system provide 3D conditioning information to a video diffusion model to improve generated video quality (object and temporal consistency) that is grounded in three dimensions (3D). The video generation method and system also enable precise camera control, cinematic effects, and scene editing. Video output corresponding to a set of camera specifications is generated for a scene from input image(s) including one or more images of a static scene or a sequence of images (video) for a dynamic scene. The input image(s) are used to calculate a 3D cache representing the scene. The 3D cache is rendered according to the set of camera specifications to produce a frame sequence and a mask sequence that identifies missing pixels in each frame. The frame sequence is encoded and masked to generate the output video.
Owner:NVIDIA CORP

Video generation method and device, equipment and storage medium

The invention relates to a video generation method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. According to the method, for video description information input by a user, a plurality of split scripts can be generated, a plurality of split pictures are generated after the user confirms the split scripts, and a final video is generated after the user confirms the split pictures, and the video can be automatically generated only by inputting the video description information by the user, so that the video generation efficiency is improved; moreover, in the process of generating the video, the generated split script and the split picture are displayed, and the video is generated after the user confirms, so that the participation degree of the user is improved, namely, the controllability of the user on the generated video is improved, the video generation efficiency is improved, the generated video can better meet the requirements of the user, and the user experience is improved. And the quality of the generated video is improved.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Road scene multi-target tracking system based on visual perception of unmanned aerial vehicle

The invention relates to the technical field of traffic management, in particular to a road scene multi-target tracking system based on visual perception of an unmanned aerial vehicle, which realizes efficient acquisition and low-delay return of road scene videos through a high-definition camera and a 5G communication technology, and provides a reliable data basis for subsequent processing. Through data preprocessing and flight control, the video quality and the flight efficiency are improved, and the path of the unmanned aerial vehicle is dynamically adjusted according to the motion trail analysis result of the key dynamic target; a key dynamic target is obtained through a one-shot detector model and an optical flow method, and effective tracking is realized in combination with a deep learning algorithm and a multi-target tracking technology, so that the accuracy of target detection and the stability of tracking are improved; the motion trail of the key dynamic target is visually displayed by analyzing the motion trail of the key dynamic target and visualizing the analysis result of the motion trail of the key dynamic target.
Owner:合肥众安睿博智能科技有限公司

Method for improving quality of space-time domain compressed video by using dense network

The invention discloses a method for improving the quality of a space-time domain compressed video by using a dense network, and the method mainly comprises the following steps: firstly, inputting a low-quality video obtained by HEVC compression into a network, aligning a certain frame to be enhanced with three adjacent frames before and after the frame in a universal deformable convolution mode, and carrying out the reconstruction of the frame to be enhanced; extracting an initial space-time fusion feature; extracting spatial feature information containing more levels from a to-be-enhanced frame by using the dense structure; the initial space-time fusion features and the spatial features are input into a space-time information adaptive fusion module together, and optimized space-time information is obtained through a series of introduced attention mechanisms and convolution layers; and finally, sending to a reconstruction module, obtaining an optimal residual error through a network, and adding the optimal residual error with a compressed frame to obtain a reconstructed video frame. Experimental results show that the method can effectively suppress the compression effect of the video, improve the video quality and obtain a better visual effect.
Owner:SICHUAN UNIV

Video quality evaluation method with rich information representation

The invention discloses a video quality evaluation method with rich information representation, which comprises the steps of spatial domain feature extraction, motion feature extraction, time sequence modeling and quality prediction, and is characterized in that a fused feature vector is input into a time domain convolutional network for time sequence modeling, and the feature dimension is reduced to 128; and finally, directly mapping the 128-dimensional features into the quality score of the video through a full connection layer. According to the method, spatial domain feature extraction, motion feature extraction, time sequence modeling and quality prediction are set, the spatial domain feature extraction part of the model considers different features represented by video frames in RGB and YIQ color spaces so as to fully extract the features of the video frames in the spatial domain, and in addition, in order to fully understand the relation between adjacent frames, the feature extraction part of the model considers the different features represented by the video frames in the RGB and YIQ color spaces. Motion information on a time sequence is supplemented by extracting motion features of a video, so that the model has better performance, the accuracy of model prediction is improved, the overall model has better performance and accuracy in video quality evaluation, and the overall effect of the model is improved.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Video quality evaluation method and device, computer equipment and storage medium

The embodiment of the invention relates to a video quality evaluation method and device, computer equipment and a storage medium, and the method comprises the steps: carrying out the frame extraction of a target video, and obtaining an image sequence; inputting the image sequence into a trained target model, and outputting index scoring information corresponding to each video quality index of the target video through the target model to obtain multiple pieces of index scoring information; and determining video scoring information of the target video according to the multiple pieces of index scoring information. Therefore, the scoring information can be generated for the quality indexes of the multiple dimensions of the video through the target model, the video scoring information is further generated according to the scoring information of the different quality indexes, automatic scoring of the video based on the multiple quality indexes of the video is achieved, and therefore the overall quality of the video is evaluated efficiently, comprehensively and accurately.
Owner:BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD

Livestock farm safety management method and system based on video monitoring

The invention relates to the technical field of video monitoring, in particular to a livestock farm safety management method and system based on video monitoring, and the method comprises the steps: collecting a monitoring video of a livestock farm, and equally dividing each frame of video image into each CTU block; marking pixel points corresponding to the target object in each frame of video image as salient pixel points, and determining the static saliency of each CTU block; obtaining the motion pixel aggregation degree of each coordinate point in each frame of video image; determining the center-of-mass coordinate and the motion displacement of the target object in each frame of video image, and obtaining the dynamic saliency of each CTU block; and combining the static saliency and the dynamic saliency, determining a saliency weight of each CTU block, distributing a code rate for each CTU block, and carrying out coding compression on the video image for livestock farm safety management. Therefore, the monitoring video quality of the livestock farm is improved, and the safety management effect of the livestock farm is enhanced.
Owner:KAIXIN (DALIAN) INTERNET SERVICES CO LTD

Video synthesis via multimodal conditioning

A multimodal video generation framework (MMVID) that benefits from text and images provided jointly or separately as input. Quantized representations of videos are utilized with a bidirectional transformer with multiple modalities as inputs to predict a discrete video representation. A new video token trained with self-learning and an improved mask-prediction algorithm for sampling video tokens is used to improve video quality and consistency. Text augmentation is utilized to improve the robustness of the textual representation and diversity of generated videos. The framework incorporates various visual modalities, such as segmentation masks, drawings, and partially occluded images. In addition, the MMVID extracts visual information as suggested by a textual prompt.
Owner:SNAP INC

Objective video quality assessment models based on bitstream, and additional pixel domain features

Techniques are described for training and use of machine learning models to determine objective video quality scores. Video quality scores predict the quality of video content perceived by viewers. Quality scores have various uses, including the selection of encoding profiles and determination of encoding ladders. A core model and residual model may be used to determine quality scores.
Owner:AMAZON TECH INC

Intelligent image selecting and cutting system fusing visual features and quality scores

The invention relates to the technical field of industrial visual intelligence, in particular to an intelligent image selection and switching system fusing visual features and quality scores, which comprises the following steps: receiving multiple paths of video coding streams, inter-frame motion vectors and camera parameters; generating a macro block activeness distribution map based on the video coding stream and the inter-frame motion vector, and performing local window positioning and feature reconstruction on the video coding stream to generate an enhanced video vector; calculating a confidence coefficient mean value and a consistency score of the video coding stream according to the enhanced video vector, and fusing the confidence coefficient mean value and the consistency score with a channel transmission signal-to-noise ratio and a quantization noise increment to generate a video quality score; constructing a multi-criterion optimization model, and setting a feature representation vector for the multi-criterion optimization model; mapping the viewpoint weight based on the inner product of the feature representation vector, and calculating with the video quality score to generate a video switching score; and performing priority ranking based on the video switching score, and triggering a mapping switching instruction. And realizing video image scheduling by fusing the visual features and the quality score.
Owner:XINAOTE (NANJING) VIDEO TECH CO LTD

Video quality evaluation method and apparatus, device and storage medium

Embodiments of the present disclosure provide a video quality evaluation method, an apparatus, a device, and a storage medium. The method includes: performing frame sampling on a target video to obtain a plurality of video frames; cropping at least one sub-image out of each of the plurality of video frames to obtain a plurality of video sub-images; inputting the plurality of video sub-images respectively into a quality evaluation model to output quality evaluation sub-information respectively corresponding to the video sub-images, wherein the quality evaluation model comprises a self-attention network with a moving window; and fusing the quality evaluation sub-information respectively corresponding to the video sub-images to obtain quality evaluation information corresponding to the target video.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Video stream quality analysis and control system and method based on AI drive

The invention relates to a real-time video stream quality analysis and adaptive control method based on AI driving, and belongs to the technical field of artificial intelligence and video stream transmission. The method comprises the following steps: receiving a real-time video stream from a server through client equipment, and the real-time video stream comprises a dynamic adaptive streaming media and a live broadcast stream; analyzing the video stream and network conditions according to the video stream in combination with a network monitoring tool to extract video features; inputting the extracted features into a machine model based on a neural network to generate quality evaluation including network conditions, video parameters and annotations; determining whether to switch to video representation forms with different bit rates or not by analyzing a quality evaluation result; and when the switching is determined, a request for updating the playing representation form is sent to the server, so that the artificial intelligence technology is realized to optimize the video quality, reduce buffering and improve the user experience, and the method is suitable for a real-time streaming media scene.
Owner:SHANGHAI ITEST TECH CO LTD

Text generation video model training method and device, equipment and storage medium

The invention relates to a training method and device of a text generation video model, equipment and a storage medium. In the training process of the text generation video model, the video evaluation model is introduced, the target video generated by the initial text generation video model is subjected to quality evaluation, and the video quality information is generated, so that the model parameters of the initial text generation video model can be updated through the video quality information; generating a trained target text generation video model; through the mode, the video evaluation model can be used for replacing manual scoring, the manual participation degree is greatly reduced, and the labor cost and the time cost are reduced; moreover, score fluctuation caused by artificial subjective difference can be avoided, accurate video quality information is generated, and stable convergence of training is facilitated.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Screen content video quality evaluation method and device based on frequency-space complementation and semantics

The invention discloses a screen content video quality evaluation method and device based on frequency-space complementation and semanteme, and relates to the field of computer vision, and the method comprises the steps: S1, extracting a video block and a key frame of a screen content video, and inputting the key frame into a high-frequency structure texture information extraction branch to obtain high-frequency structure texture information; s2, inputting the key frame into a noise sensing module to obtain a noise sensing feature; s3, inputting the noise perception features into a self-adaptive time sequence embedding module to obtain noise and semantic information; s4, splicing the high-frequency structure texture information and the noise and semantic information, performing quality regression to obtain a quality score of a single-frame key frame, and summing and averaging to obtain a spatial domain video quality score; s5, inputting the video blocks into a Fast-VQA-based quality evaluation branch to obtain a time sequence distortion perception degradation score; and S6, dynamically fusing the spatial domain video quality score and the time sequence distortion perception degradation score to obtain a final video quality score. According to the method provided by the invention, the screen content video quality is effectively evaluated.
Owner:XIAMEN UNIV OF TECH +1

Lightweight acoustic feature extraction method for end-side audio and video quality inspection

The invention relates to a lightweight acoustic feature extraction method for end-side audio and video quality inspection. The method comprises the following steps: end-side equipment separates audio and video stream data through a double-time-sequence anchor point alignment method, and acquires independent audio data; performing grading preprocessing on the independent audio data to obtain noise-reduced independent audio data; on the basis of the segmented audio, a time domain and frequency domain collaborative extraction algorithm is adopted, multi-dimensional time domain features and frequency features sensitive to audio quality difference are obtained, and a lightweight acoustic feature vector is obtained through an incremental principal component analysis method; based on the lightweight acoustic feature vector, utilizing a dual-threshold matching judgment mode to obtain similarity between features; and on the basis of the feature similarity, through feature dimension deviation verification, obtaining a quality inspection result containing a standard grade and feature anomaly positioning. Lightweight acoustic feature extraction of end side audio and video quality inspection is realized.
Owner:SHANGHAI SIOO INFORMATION TECH CO LTD