Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1096 results about "Reference frame" patented technology

Reference frames are frames of a compressed video that are used to define future frames. As such, they are only used in inter-frame compression techniques. In older video encoding standards, such as MPEG-2, only one reference frame – the previous frame – was used for P-frames. Two reference frames (one past and one future) were used for B-frames.

EVTOL multi-camera cooperative video compression coding method based on multi-source perception and intelligent partitioning

The invention discloses an eVTOL multi-camera cooperative video compression coding method based on multi-source perception and intelligent partitioning. The method comprises the following steps: constructing a scene three-dimensional perception model by fusing multi-source data of visible light, infrared and depth sensors; the method comprises the following steps: realizing dynamic video partitioning based on motion vectors and semantic analysis, and dividing a picture into a core region, a secondary region and a background region; establishing a parallax compensation motion prediction model by adopting a cross-camera reference frame sharing mechanism; and high-fidelity compression of the key area is realized through layered entropy coding and a dynamic quantization parameter distribution strategy. And the decoding end reversely executes multi-source data fusion and partition reconstruction according to the coding metadata. The method is suitable for eVTOL multi-camera video real-time transmission scenes such as polling, surveying and mapping.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Extremely compressed video coding method based on intelligent reference frame

The invention discloses an extreme compressed video coding method based on an intelligent reference frame, comprising the following steps: S1, acquiring an original video stream, dividing the original video stream into frame groups according to a time sequence, each group comprising a current frame and a candidate reference frame; s2, carrying out region positioning on the candidate reference frame, extracting motion, edge and background regions, and generating a mapping graph; s3, calculating a quality score based on the mapping graph according to a motion vector, gradient change and a region overlapping rate; s4, selecting three frames with highest scores to form a reference set, and establishing an index structure; s5, predicting the current frame by using the reference set to generate a predicted frame and a residual error; s6, multi-path coding cost is calculated, and a path with the minimum cost is selected; and S7, entropy coding is carried out on the reference index, the motion vector, the residual error and the control parameter, a code stream is output, and calling information is recorded. According to the method, the compression ratio and prediction precision of video coding are improved, the image quality is maintained while the code rate is reduced, and the method is suitable for efficient transmission and storage of high-resolution videos.
Owner:NINGXIA ANYING INFORMATION TECHNOLOGY SERVICE CO LTD

Method and device for demonstration-based robot programming with adaptive reference frames

A method of programming an industrial robot (100), with a robot manipulator (110) and robot controller (120), comprises: recording movements of the robot manipulator during a demonstration-based robot programming session, for thereby obtaining a robot trajectory; acquiring a video of the robot programming session; and generating, on the basis of the robot trajectory and the video, a robot program (C) executable by the robot controller, wherein the robot program includes at least one motion command which is expressed in a first reference frame (0 1). The method further comprises capturing operator input data indicating a first object (151) in an image in the video. The generated robot program includes a command to identify a position of an object resembling the first object in a work area (150) of a robot which executes the robot program; and the first reference frame is defined with respect to the identified position of the object resembling the first object.
Owner:ABB (SCHWEIZ) AG

Unsupervised region-growing network for object segmentation in atmospheric turbulence

An unsupervised region-growing network (RGN) is trained to perform object segmentation on video data degraded by atmospheric turbulence. The method includes obtaining input data containing turbulence-degraded video, extracting a video frame sequence, and training the RGN using a selected algorithm incorporating a region-growing algorithm and a grouping loss function. A bidirectional optical flow sequence is computed for multiple reference frames within the video sequence. Pixel-level masks are generated for detected moving objects, followed by applying the region-growing algorithm to create coarse masks. A grouping loss function refines these masks to ensure consistency across consecutive frames. The trained RGN outputs refined masks as object segmentation data for the received video, improving segmentation accuracy in turbulent environments. This approach enables robust object detection and segmentation without requiring prior video restoration, maintaining fidelity to the original turbulence-distorted input.
Owner:CLEMSON UNIV RES FOUND +2

Skeleton sign language recognition method of double-flow space-time dynamic graph convolutional network fused with residual learning

The invention discloses a skeleton sign language recognition method of a double-flow space-time dynamic graph convolutional network fused with residual learning, and belongs to the technical field of artificial intelligence and gesture recognition. According to the method, an input gesture skeleton sequence relative to a face is divided into two data streams, namely a hand form data stream and a wrist track data stream through double-reference-system differential homeomorphic mapping; the method comprises the following steps: firstly, processing hand posture data, and capturing a hand joint spatial topological relation by combining a spatial-temporal dynamic graph convolutional network (STDGCNN) with a residual convolutional block; meanwhile, a Finsler trajectory dynamics encoder (FTDE) is adopted to carry out differential geometric modeling on the wrist trajectory, and the direction sensitivity characteristic of the trajectory is captured through a multi-scale causal convolutional network. Then, mutual enhancement of double-flow features is realized through a bidirectional cross feature enhancer (BCFE), and the problem of geometric inconsistency of a heterogeneous feature space is solved through a geometric-driven optimal transmission fusion device (Geometric-OT). The method solves the challenge that a traditional sign language recognition method processes complex space-time correlation of gesture forms and motion tracks at the same time, the technical problems of insufficient relation between hands and faces, insufficient feature expression ability and low space-time feature extraction efficiency, and the problems of geometric inconsistency, single reference system and the like. And the identification accuracy and the real-time performance are obviously improved. Experiments show that the accuracy rate of the method in complex hundreds of sign language vocabulary recognition tasks reaches 95% or above on average, the reasoning speed is only 17ms on average, high-precision real-time sign language recognition is achieved, and the method has higher robustness in complex environments such as noise and shielding.
Owner:刘良锦

Apparatus, method, and computer program for video encoding and decoding

The method includes receiving an image block unit of a frame, the image block unit including samples in a first chrominance channel, a second chrominance channel, and one luminance channel; defining a co-located reference area across a luminance component of the luminance channel, a chrominance component of the first chrominance channel, and a chrominance component of the second chrominance channel; performing motion compensation prediction of the luminance component, the first chrominance component, and the second chrominance component using the reference frame and a motion vector; obtaining a first mapping function that maps the luminance component to the first chrominance component and a second mapping function that maps the luminance component to the second chrominance component using the co-located luminance prediction and chrominance prediction; reconstructing a luminance residual; and obtaining a first chrominance prediction by using the first function and the luminance residual and obtaining a second chrominance prediction by using the second function and the luminance residual.
Owner:NOKIA TECHNOLOGIES OY

Pleno-generation face video compression framework for generative face video compression

Methods and systems implement a pleno-generation face video compression framework with bandwidth intelligence for generative models and compression. Heterogeneous-granularity facial description regularizes long-term dependencies between video frames and compensates for motion estimation errors caused by compact representations of motion information. A generative decoder reconstructs heterogeneous-granularity visual representations, providing auxiliary visual signals for attention-based recalibration of a GFVC-reconstructed face signal. A coarse-to-fine generation strategy avoids error accumulation. High efficiency for heterogeneous-granularity signal compression is achieved by two different entropy-based signal compression methods: heterogeneous-granularities feature representation from the key-reference frame as hyperpriors to optimize the entropy model for compressing heterogeneous-granularity feature from subsequent inter frames, and a feature difference operation for heterogeneous-granularities feature representation between key-reference and subsequent inter frames, such that the entropy model only compresses heterogeneous-granularities feature residual for redundancy reduction. Mixed-model dataset generation and training and model-specific dataset generation and training are also provided.
Owner:SIM IP 5 LLC

Heart motion feature extraction method based on optical flow estimation

The invention discloses a heart motion feature extraction method based on optical flow estimation, and relates to medical image processing. Preprocessing the input four-dimensional space-time cardiac magnetic resonance imaging data, namely, scaling pixel values; inputting two frames of images which are continuous in time into a feature encoder and a context encoder for feature extraction, wherein the two frames of images are divided into a reference frame and a moving frame; calculating the correlation between the feature maps through a correlation volume calculation module, constructing a correlation pyramid, and extracting the feature maps to provide matching information for subsequent optical flow estimation; the motion feature iteration enhancement module iteratively and continuously refines an optical flow estimation result through a deformable convolution and global motion aggregation (GMA) module; model parameters are optimized based on a weighted sum of luminosity consistency loss, smoothness loss, and gradient consistency loss. Edge features are adaptively captured through deformable convolution, global and local features are fused by using a GMA module, optical flow prediction errors are effectively reduced, motion estimation quality is improved, and key features of a heart edge region are maintained.
Owner:XIAMEN UNIV

Real-time monitoring method and system for milk powder stirring processing

The invention provides a real-time monitoring method and system for milk powder stirring processing, and the method comprises the steps: collecting a current frame image of a milk powder material in a stirring container, and obtaining a reference frame image of the current frame image before a preset time interval; calculating a space texture feature set, a time sequence color texture feature set and a dynamic flow field feature set of the current frame image; cascading the space texture feature set, the time sequence color texture feature set and the dynamic flow field feature set to obtain a high-dimensional state vector; calculating a mahalanobis distance between the high-dimensional state vector and a target uniform state cluster core in a pre-constructed state space; when the mahalanobis distance is smaller than a first threshold value, it is judged that the current stirring state is uniform mixing; when the mahalanobis distance is greater than a second threshold value, determining an abnormal state; when the Mahalanobis distance is between the first threshold and the second threshold, it is determined that mixing is being performed.
Owner:SHAANXI YATAI DAIRY CO LTD

ERP data intelligent supervision method based on Internet of Things

The invention discloses an ERP (Enterprise Resource Planning) data intelligent supervision method based on the Internet of Things, which relates to the technical field of enterprise informatization, and comprises the following steps: constructing a time sequence integrity observation layer in an Internet of Things data acquisition link, embedding a mirror image time anchor in each sensor sampling point, recording electromagnetic interference, time service offset and queuing delay, generating a phase fingerprint time baseline, and obtaining a phase fingerprint time baseline; constructing a unified time reference framework; operating a causal coherent analysis mechanism based on a unified time reference frame, comparing instantaneous offsets of a mirror time anchor and a real-time timestamp, identifying a phase inversion region, extracting information of a corresponding sampling node, a production line position and an energy consumption unit, generating a high-risk time slice list, and determining a time correction boundary. According to the method, a time sequence observation layer and a time reference frame are constructed, phase abnormity is identified, a data trend is reconstructed, a regulation and control instruction draft is generated, sampling time sequence recovery and scheduling convergence are realized through time inversion control, data consistency and decision accuracy are improved, and a closed-loop control system is constructed.
Owner:FUJIAN ZHILIAN ALL THINGS TECH CO LTD

Numerical simulation method for solid phase and liquid phase in rotating flow channel based on multiple reference systems

The invention relates to the field of multiphase numerical simulation, and discloses a numerical simulation method for solid and liquid phases in a rotating flow channel based on multiple reference systems, which comprises the following steps of: firstly, acquiring a rotating flow field generated by a multiple reference system method, and extracting rotating domain information as well as particle information and fluid computational domain grid information in a current time step; secondly, particles in a rotation domain and particles outside the rotation domain are determined according to the rotation domain information and the particle information; in combination with the particle information and the fluid calculation domain grid information, the final flow field speed and the reconstructed porosity of the particles in the rotation domain and the particles outside the rotation domain are calculated respectively, and then the drag force is calculated according to the final flow field speed, the reconstructed porosity and the particle information; meanwhile, according to the particle information and the rotation domain information, the Korotkoff force and the centrifugal force of the particles in the rotation domain are calculated. Finally, drag force, Korotkoff force and centrifugal force are calculated through iteration, particle information and grid information of the next time step are obtained, and therefore numerical simulation calculation of particle motion in the rotating flow channel under the fixed grid is achieved.
Owner:ZHEJIANG SCI-TECH UNIV

Video coding method and device, electronic equipment and storage medium

The invention provides a video coding method and device, electronic equipment and a storage medium, relates to the technical field of image processing, in particular to the field of video coding and the like, and can be applied to application scenes such as video live broadcast and the like. The specific implementation scheme is as follows: acquiring a reference frame which is adjacent to a current frame and has established initial ROI hierarchical division; multiplexing a motion vector generated by the reference frame in a motion compensation time domain filtering process, and determining an ROI position predicted value of the current frame; extracting feature points of the five sense organs of the current frame, and correcting an ROI position predicted value based on a feature point spatial topological relation in the initial ROI; respectively configuring differential quantization parameters for the corrected main face region, the corrected secondary face region and the corrected background region; dynamically allocating a three-layer area code rate according to a real-time network bandwidth; and outputting the coded frame of the current frame and the associated ROI level metadata. According to the scheme, the coding efficiency and quality can be improved.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Engineering project management digital model generation method and device, equipment and medium

The invention provides a method, a device, equipment and a medium for generating an engineering project management digital model, and the method comprises the steps: collecting original data related to engineering project features, carrying out the preprocessing, forming a standardized project data set, constructing a space-time reference frame based on the data set, fusing a multi-dimensional feature tensor of multiple engineering features, and carrying out the construction of a space-time reference frame. The method comprises the steps of generating a project progress node and an integrated feature tensor, fusing the project progress node and the integrated feature tensor to generate a multi-source association graph containing entity relationships and feature semantics, generating a verified engineering project management digital model according to the project progress node, the integrated feature tensor and the association graph, and dynamically updating and optimizing the model in response to change data in project execution. According to the method, the problems of data dispersion, lack of association and static stiffness of the model in a traditional method are solved, automatic construction and dynamic evolution from multi-source heterogeneous data to an intelligent decision-making model are realized, and the accuracy, real-time performance and intelligent level of project management are remarkably improved.
Owner:CISDI INFORMATION TECH CO LTD

Contact network defect identification method based on multiple modes

The invention belongs to the technical field of image processing, and particularly relates to a contact network defect identification method based on multiple modalities, which comprises the following steps of: 1, acquiring a visible light image sequence and an infrared thermal imaging image sequence of a target contact network, and carrying out time alignment and temperature calibration to obtain a unified reference frame sequence and a thermal vision consistency label set; 2, constructing a structure change graph and an overheat candidate graph in the obtained unified reference frame sequence, calling a cross-modal physical semantic autocatalysis recombination algorithm, and taking a generated thermal vision consistency label set as a constraint; and step 3, taking the output fusion evidence body as input, combining the thermal vision consistency label set and the reversible channel audit record to verify the candidate area, obtaining a defect target, and outputting a defect category and a defect position. According to the method, the precision and stability of defect detection are improved, the traceability of the result is also realized, and a reliable guarantee is provided for intelligent operation and maintenance of the electrified railway overhead line system.
Owner:CHENGDU NUOBIKAN TECH CO LTD

Video Encoding Method and Related Apparatus

A method includes performing scene detection on a current frame of picture to obtain a scene status of the current frame of picture; determining, based on the scene status, a reference frame structure corresponding to the current frame of picture, where the reference frame structure indicates a reference frame of picture of the current frame of picture and an encoding layer of the current frame of picture; and encoding the current frame of picture into a bit stream based on the reference frame structure. In a process of encoding a video, a reference frame of picture and an encoding layer of each frame of picture are adjusted in real time with reference to features such as whether scene switching occurs or whether a scene is kept stable in each frame of picture.
Owner:HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD

Single low-illumination image enhancement method for simulating multi-exposure image fusion

The invention belongs to the technical field of image data processing, and particularly discloses a single low-illumination image enhancement method for simulating multi-exposure image fusion. The method comprises the following steps: carrying out nonlinear transformation on a single low-illumination input image to generate a virtual exposure sequence containing a reference frame and a plurality of non-reference frames; inputting a virtual exposure sequence into a VMEIF-Net network, processing a reference frame and a non-reference frame through a multi-branch feature coding structure, introducing a space attention module to a non-reference frame branch to screen features consistent with the reference frame, iteratively optimizing texture details by adopting a feature feedback unit, and inhibiting artifacts through a global feature analysis unit, so as to obtain a VMEIF-Net network; and generating an enhanced image through feature fusion and decoding operation. According to the invention, through multi-branch feature extraction and feedback connection, high-level feature information is transmitted back to a feature fusion stage, so that the information transmission efficiency is improved; by introducing the global context sensing module, the receptive field range is expanded, generation of artifacts can be inhibited, and rich image detail information is kept.
Owner:WEIFANG UNIVERSITY +3

Three-dimensional reconstruction method, device and equipment of moving object and readable storage medium

The invention relates to the technical field of image processing, and discloses a three-dimensional reconstruction method, device and equipment of a moving object and a readable storage medium, and the three-dimensional reconstruction method of the moving object comprises the steps: obtaining a plurality of fringe images of a measured moving object at continuous moments, selecting one frame as a reference frame, and obtaining a plurality of fringe images of the measured moving object at continuous moments; constructing an image set together with adjacent image frames; inputting the image set into a deep learning network for pixel-level motion tracking to obtain motion position information; based on the motion information and the intensity information, constructing a phase calculation equation set to obtain a relative phase value; performing phase unwrapping processing on the relative phase value to obtain absolute phase information; and performing phase height mapping by combining system calibration parameters to obtain dynamic three-dimensional shape data. According to the method, phase distortion caused by motion interference is remarkably eliminated, the three-dimensional reconstruction precision and real-time performance in a dynamic scene are improved, and a reliable guarantee is provided for industrial detection and high-speed target measurement.
Owner:HENAN UNIVERSITY OF TECHNOLOGY

Underwater single-target tracking method based on wavelet token and space-time Transform

The invention relates to an underwater single target tracking method based on a wavelet token and a space-time Transform. The method comprises the following steps: firstly, constructing a reference frame sequence, a search frame and a previous frame historical token into a space-time input sequence, and extracting cross-frame features through a Transform encoder; then, Haar wavelet decomposition is carried out on the historical token, and a low-frequency component representing a target structure and a high-frequency component capturing motion details are separated out; then, adaptively fusing the global features and the historical components of the current search frame by using a gating mechanism, and generating a wavelet token; and finally, inputting the wavelet token and the global feature into a prediction head, and outputting a target classification confidence map and a bounding box regression map to determine the position and the scale of the target. According to the technical scheme of the invention, the interference of underwater low-illumination noise can be effectively suppressed through the wavelet token, and the space-time continuity of target motion modeling is maintained in combination with a gating strategy, so that the tracking robustness of an underwater complex scene is effectively improved.
Owner:GUILIN UNIVERSITY OF TECHNOLOGY

Adaptive transform type sets based on frame level statistics

Encoding using adaptive transform type sets based on frame level statistics includes obtaining an encoded bitstream by encoding a current block of a current frame of a current sequence of frames of an input video stream using adaptive transform type sets based on frame level statistics and outputting the encoded bitstream. Encoding the current block includes obtaining transform type statistics for previously reconstructed reference frames from the current sequence of frames, the previously reconstructed reference frames including at least one previously reconstructed reference frame, determining, in accordance with the transform type statistics, a current subset of transform types from a set of available transform types, generating encoded block data for the current block using a current transform type from the current subset of transform types, and including the encoded block data in the encoded bitstream.
Owner:GOOGLE LLC

Method and device for processing audio data

The invention relates to an audio data processing method, a neural network training method, an audio data processing device, electronic equipment and a computer readable storage medium. The method comprises the following steps: determining facial prior information of a virtual object based on reference video data; determining a target audio feature based on the target audio data; determining target expression sequence information based on the face prior information and the target audio feature; and determining target video data corresponding to the target audio data based on the target expression sequence information and a reference frame in the reference video data. According to the method and the device, the synchronization precision between the audio content and the facial component (such as the mouth) action of the virtual object can be remarkably improved, the real texture of the mouth action is enhanced, the expression difference between individuals is more accurately captured and reproduced, and the operation efficiency and the stability of the whole system are also improved while the performance of the system is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Target detection method and device based on visual large model

The invention provides a target detection method and device based on a visual large model. The method comprises the following steps: inputting a to-be-detected picture into a preset visual basic large model to obtain image features; calculating the image features and a preset positive example first feature prototype vector to generate a plurality of adaptive reference frames; performing feature conversion processing on the reference frames, and performing candidate feature similarity comparison on the reference frames and a preset positive example second feature prototype vector to obtain a first similarity score of each reference frame and each category; performing non-maximum suppression processing on each reference frame aiming at each category so as to reserve the most accurate reference frame; and taking the reference frame with the similarity score exceeding a score threshold as a detection result of the corresponding category. Through the application of the method and the device, the approximate region of the target is obtained by adopting a mode of adaptively generating the reference frame, simple and efficient target detection can be realized under the condition that only a small number of labeled samples are needed, and the actual landing requirements of detection items are better met.
Owner:GUANGZHOU YUNCONG INFORMATION TECH CO LTD +1

Video enhancement method, related device and computer program product

The invention provides a video enhancement method, related equipment and a computer program product, and relates to the technical field of image processing. The method comprises the following steps: acquiring a target frame and at least one reference frame of the target frame from a video to be enhanced; performing feature extraction on the target frame and the at least one reference frame through a feature extraction module to obtain a target frame feature corresponding to the target frame and a reference frame feature corresponding to each reference frame; performing multi-frame alignment processing on the target frame feature and each reference frame feature through a multi-frame alignment module to obtain a multi-frame alignment feature; performing image enhancement processing on the multi-frame alignment features through an image enhancement module to obtain image enhancement features; and adding the image enhancement feature and the target frame pixel by pixel to obtain an enhanced target frame. According to the embodiment of the invention, the video can be enhanced accurately and efficiently.
Owner:BEIJING SANKUAI ONLINE TECH CO LTD

Super-resolution imaging method based on focal plane splicing and adaptive fusion

The invention relates to the field of digital image processing, in particular to a super-resolution imaging method based on focal plane splicing and adaptive fusion. According to the method, sub-pixel offset among nine CCDs is preset through hardware, and nine frames of low-resolution image sequences with accurate displacement are obtained in push-broom. A central image is taken as a reference frame, high-precision mapping is realized based on hardware offset, motion estimation errors are avoided, effective pixels are screened by calculating robustness weight, an anisotropic Gaussian kernel function with a self-adaptive local structure is constructed so as to maintain image edge and detail features, and each frame is accumulated to a high-resolution grid in a weighting mode, so that a high-resolution image is obtained. And a sample compensation mechanism based on cumulative robustness is introduced, a fusion strategy is adaptively adjusted in an information insufficient area, and finally a high-resolution image is generated through normalization. The method significantly improves the imaging quality, suppresses artifacts and noise, and is suitable for the field of satellite remote sensing.
Owner:XIANGTAN UNIV

High-precision source positioning method suitable for cross-scale complex rock mass medium

The invention provides a high-precision source positioning method suitable for a cross-scale complex rock mass medium. The high-precision source positioning method comprises the following steps of construction of a homogenized reference frame, generation of a linear equation set, dynamic weight estimation, matrix truncation decomposition for enhancing stability, and output of sound source coordinates, wave velocity and trigger time. According to the method, error transmission of a single sensor is eliminated through a homogenized reference frame, noise interference is suppressed through dynamic weight estimation, the problem of ill-conditioned matrix inversion is solved in combination with matrix truncation decomposition, geometric limitation of a sensor array is broken through, and finally high-precision and high-stability cross-scale acoustic emission source positioning is achieved. Reliable technical support is provided for complex engineering environments such as a dynamic velocity field, random sensor layout and multi-scale monitoring requirements.
Owner:CHONGQING UNIV +2

Video coding method and device, video decoding method and device, electronic equipment and storage medium

The invention relates to a video coding method and device, a video decoding method and device, electronic equipment and a storage medium. The video coding method comprises the following steps: acquiring difference information between a plurality of reference frames of a current frame and the current frame; determining weights corresponding to the plurality of reference frames based on the difference information; fusing the coding information of each reference frame in the plurality of reference frames based on the weights corresponding to the plurality of reference frames to obtain fused coding information; and coding the current frame based on the fused coding information to obtain a coding frame of the current frame.
Owner:BEIJING DAJIA INTERNET INFORMATION TECH CO LTD

Occupancy coding using inter prediction with octree occupancy coding based on dynamic optimal binary coder with update on the fly (OBUF) in geometry-based point cloud compression

A G-PCC coder may determine an occupancy of a reference child node in a reference node, wherein the reference node is in a reference frame of point cloud data used for inter prediction of a current node in a current frame of the point cloud data. The G-PCC coder may further determine a context for decoding a current occupancy bit of a current child node of the current node based on the occupancy of the reference child node, and arithmetic decode the current occupancy bit using the context.
Owner:QUALCOMM INC

Audio data processing method, neural network training method, and related apparatus

The present disclosure relates to an audio data processing method, a neural network training method, an audio data processing apparatus, an electronic device, and a computer readable storage medium. The audio data processing method comprises: determining facial prior information of a virtual object on the basis of reference video data; determining a target audio feature on the basis of target audio data; determining target expression sequence information on the basis of the facial prior information and the target audio feature; and on the basis of the target expression sequence information and a reference frame in the reference video data, determining target video data corresponding to the target audio data. According to the present disclosure, the synchronization precision between audio content and a facial component (e.g., the mouth) action of a virtual object can be significantly improved, the sense of reality of the mouth action is enhanced, an expression difference between individuals is captured and reproduced more accurately, and the operation efficiency and stability of the whole system are improved while the system performance is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Video monitoring data transmission and storage method based on narrow bandwidth

The invention relates to the technical field of video compression, in particular to a narrow-bandwidth-based video monitoring data transmission and storage method, which comprises the following steps of: for an input video monitoring data frame, calculating a difference absolute value sum between a current frame and a reference frame; according to the method, the pixel difference of the current frame and the reference frame is calculated, and the average amplitude of the motion vector field is fused to form a comprehensive index of the dynamic degree of the quantized content, so that the image group structure is not fixed or periodic any more, and the dynamic degree of the quantized content can be calculated according to the score of scene activity and the change intensity. When a monitoring picture is static, a super-long image group is established to limit a compression code rate, when the picture is suddenly changed, the super-long image group is quickly switched to a short image group to ensure instant refreshing and definition of key information, meanwhile, non-uniform redistribution is performed on limited total code rate budget, bit resources can be intelligently inclined to frames with violent content change and large information amount, and the real-time refreshing and definition of the key information are ensured. Therefore, the subjective visual quality of the key dynamic moments is improved under the narrow bandwidth.
Owner:THE FIRST MONITORING AND APPLICATION CENTER CHINA EARTHQUAKE ADMINISTRATION +1

Adaptive video restoration method based on digital video technology

The invention discloses an adaptive video restoration method based on a digital video technology, and relates to the technical field of image communication. The method comprises the following steps: preprocessing an original video containing hard subtitles and automatically removing the hard subtitles to obtain a video to be repaired; screening a reference frame and a to-be-repaired frame by comparing the original video with the to-be-repaired video; determining a to-be-repaired area according to the pixel gray scale difference; extracting matched feature points, calculating a reference motion vector, estimating possible positions and motion differences of points in the to-be-repaired region in combination with object structure segmentation and region prediction, and generating a corrected motion vector; performing sub-pixel-level interpolation by using the vector, and constructing a corrected reference image; and finally, fusing the corrected image and the to-be-repaired area to generate a repaired video frame. The method can effectively improve the restoration quality of the subtitle shielded area.
Owner:EC INNOVATIONS (SHENYANG) INC