Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

898 results about "Video sequence" patented technology

Guideboard surface deformation detection method and system based on dynamic video sequence

The invention discloses a guideboard surface deformation detection method and system based on a dynamic video sequence. The method comprises the following steps: acquiring road video data with a timestamp and position information; a traffic sign is automatically identified from the road video data, and geometric anomaly detection is carried out; a time sequence tracking sequence is established for the guideboard, and whether continuous abnormity exists in the guideboard is determined through multi-frame analysis; if yes, performing local texture consistency and local contour geometric analysis on the guideboard to obtain a fusion analysis result; performing state evaluation on supporting equipment corresponding to the guideboard to obtain a state evaluation result of the guideboard supporting equipment; and determining the final surface deformation state and the corresponding grade of the guideboard based on the fusion analysis result and the state evaluation result of the guideboard supporting equipment. By implementing the method provided by the invention, automatic detection and quantitative analysis of the bending or deformation condition of the road sign can be realized without manual intervention, and early warning information can be output in time.
Owner:WINTOO INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Video object segmentation method based on query adaptive attention and discriminative memory

The invention belongs to the technical field of computer vision and digital video processing, and discloses a video object segmentation method based on query adaptive attention and discriminative memory, which comprises the following steps: step 1, constructing a video data set and preprocessing data; step 2, constructing a video object segmentation model combining multi-scale semantic feature integration and query adaptive discriminant enhancement, and training the video object segmentation model; step 3, video object segmentation reasoning; performing target segmentation reasoning on the preprocessed video sequence, outputting a frame-by-frame target segmentation mask, and generating a time sequence tracking result; 4, exporting, deploying and applying the model; through lightweight network design and a discriminative memory optimization strategy, the model is deployed to edge equipment, and real-time segmentation and visualization are realized. According to the method, on one hand, comprehensive target representation is provided on multi-scale feature extraction, and on the basis of the characteristic that only high-confidence target features are stored, error propagation is avoided, and the long-term segmentation stability of the method is ensured.
Owner:NANTONG INST OF TECH

Electric power operator behavior detection method and related equipment

The invention discloses a power worker behavior detection method and related equipment. The method comprises the following steps: acquiring first data in a scene where a worker is located through a camera in a working site; the first data comprises continuous frame images or a video sequence lasting for preset time; analyzing the first data to obtain a current first identification result of the operator; the first identification result is used for representing whether the operator lacks of wearing the protective equipment and / or carries illegal articles; determining the current posture of the operator from the first data; and fusing the first recognition result and the current posture to obtain the operation risk level of the operator. The method can be widely applied to the technical field of artificial intelligence.
Owner:FOSHAN UNIVERSITY

Long video understanding method capable of relieving time sequence illusion in video language large model

The invention provides a long video understanding method capable of relieving time sequence illusion in a video language large model. The long video understanding method is based on a static bias adaptive frame selection mechanism and a cross-modal feature fusion strategy. According to the static bias mechanism, inter-frame similarity is evaluated through a discriminator, redundant frames are identified, key frames are selected or a complete sequence is reserved, so that calculation overhead is reduced, and spatio-temporal information integrity is kept; a video frame and a text are mapped to a shared semantic space, the single-frame semantic understanding ability is enhanced, then an embedded sequence serves as a soft prompt to be input into a large language model, and a final answer is generated in an autoregression mode. According to the method, the efficiency and accuracy of long video understanding and video question and answer tasks can be remarkably improved; the problem of low training and reasoning efficiency caused by time sequence dependence redundancy and excessive computing resource consumption is effectively relieved; and through a dynamic multi-modal task processing framework and a space-time memory bank compression mechanism, the modeling capability and generalization performance of the model on a long video sequence are further improved.
Owner:LANZHOU UNIV

Video snapshot compression imaging reconstruction method and system

The invention relates to a video snapshot compression imaging reconstruction method and system. The method comprises the following steps: inputting a video frame sequence and a time-varying mask set thereof into a measurement model to obtain initial estimation; constructing a reconstruction network which comprises a feature extraction module, a gating residual network module and a video reconstruction module; the feature extraction module comprises two three-dimensional convolution layers, each three-dimensional convolution layer is connected with an activation function, and the feature extraction module extracts initial features from the initial estimation; inputting the initial features into a gating residual network module, and outputting reconstruction information features; and the video reconstruction module fuses the reconstruction information features, and performs up-sampling and detail refining to reconstruct a video sequence. According to the method, on the premise that parameters and computing power are hardly increased, ghosting and flickering are effectively restrained, the stability of long-time reconstruction is improved, and an effective scheme is provided for SCI reconstruction with the high compression ratio, the super-definition resolution ratio and the long sequence.
Owner:GUANGDONG UNIV OF TECH

Two-stage road traffic abnormal event identification method and system based on visual large model

The invention relates to a two-stage road traffic abnormal event identification method and system based on a visual large model, and the method comprises the steps: collecting the monitoring image and video data of an expressway and an urban expressway, and building a static image semantic data set and a dynamic video traffic semantic data set; utilizing the static image semantic data set to train a visual large model to obtain a first visual large model; constructing a same-preference data pair, and performing direct preference optimization training of the first visual large model by using the same-preference data pair to obtain a second visual large model; intercepting an abnormal video key frame based on the dynamic video traffic semantic data set, and performing parameter fine tuning on the second visual large model based on the abnormal video key frame to obtain a road traffic abnormal event recognition model; and collecting a monitoring image or video sequence in real time, and performing abnormal event identification by using the road traffic abnormal event identification model. Compared with the prior art, the traffic abnormal event identification method provided by the invention can effectively combine dynamic and static characteristics of data and is efficient.
Owner:TONGJI UNIV

Diffusion model video local editing method and system based on mask guidance

The invention provides a diffusion model video local editing method and system based on mask guidance. The method comprises the following steps: encoding a video sequence of a target video to obtain a first video frame submerged space feature; performing coarse-grained mask labeling on a target editing area in the target video to obtain mask information; determining a space attention weight according to the first video frame potential space feature; determining a second video frame potential space feature according to the first video frame potential space feature, the mask information and the space attention weight; determining a time attention weight according to the second video frame potential space feature; determining a third video frame potential space feature according to the mask information, the second video frame potential space feature and the time attention weight; and decoding the hidden space feature of the third video frame to generate an edited video. According to the method, frame-by-frame accurate masking is not needed, the time-space consistency of local editing of the video can be effectively enhanced, an unedited area is kept unchanged, and a high-quality and stable video editing effect is achieved.
Owner:HEFEI UNIV OF TECH

Vehicle target detection tracking and trajectory data extraction method based on deep learning

The invention discloses a vehicle target detection tracking and trajectory data extraction method based on deep learning, and relates to the technical field of unmanned aerial vehicle aerial photography. The method comprises the following steps: carrying out stable frame processing on an unmanned aerial vehicle video, extracting feature points and feature vectors through an SURF algorithm, matching and screening reliable matching pairs through an FLANN algorithm, calculating a homography matrix through an RANSAC algorithm when a condition is met, carrying out perspective transformation to eliminate jitter, and outputting a stable video sequence; vehicle target detection: introducing an AIFI module to construct an improved YOLOv5OBB model, and outputting vehicle rotation bounding box parameters and categories after training; vehicle tracking and trajectory extraction are carried out, cross-frame tracking is realized based on a DeepSORT model, original trajectory data are preprocessed, and the speed, the acceleration, the additional lane number and the ID of an adjacent vehicle are calculated. According to the method, the problems of video jitter and insufficient detection precision are effectively solved, and the accuracy and continuity of track data extraction in the highway scene are improved.
Owner:BEIJING JIAOTONG UNIV

Port forbidden area vehicle illegal parking monitoring method based on Transform and time sequence detection

The invention discloses a port forbidden area vehicle illegal parking monitoring method based on transformer and time sequence detection, and the method comprises the steps: carrying out the video collection and preprocessing, carrying out the target detection of each frame of image of a video through a detection model constructed based on a deep convolutional neural network, and carrying out the target detection through a multi-target tracking algorithm, and performing ID distribution and trajectory tracking on the same vehicle in the video sequence, analyzing the motion state of the vehicle in combination with a time sequence detection model, identifying whether a forbidden zone staying or illegal parking behavior occurs, and judging whether the vehicle is illegal. According to the scheme, CNN and Transform attention mechanisms are combined, the detection precision of vehicles in a complex port scene is improved, a multi-target tracking algorithm is utilized in vehicle trajectory matching, high-precision matching of a vehicle detection frame and a tracking trajectory is achieved in combination with a Hungary algorithm, a time sequence detection model is utilized to detect the vehicle, and the vehicle trajectory matching precision is improved. The system can recognize behavior changes of illegal parking vehicles in the time dimension, automation of the whole illegal parking judgment and warning process is achieved, and the safety response efficiency is improved.
Owner:CHINA DESIGN GROUP CO LTD +1

Interpolation filter clipping for sub-picture motion vectors

A video coding mechanism is disclosed. The mechanism includes receiving a bitstream comprising a current picture including a sub-picture coded according to inter-prediction. A motion vector for a block of the sub-picture is determined. A clipping function is applied to sample locations in a reference block to support application of an interpolation filter when the motion vector points outside of the sub-picture and when a flag is set to indicate the sub-picture is treated as a picture. The interpolation filter is applied to results of the clipping function to obtain a predicted sample value. The block is decoded based on the predicted sample value. The block is forwarded for display as part of a decoded video sequence.
Owner:HUAWEI TECH CO LTD

Video coding and decoding method, device, equipment and storage medium

The invention discloses a video coding and decoding method, device and equipment and a storage medium, and relates to the technical field of video coding and decoding, and the method comprises the steps: setting a target sequence identifier in a sequence head of a to-be-coded video frame sequence corresponding to a target video, determining a to-be-zoomed video frame, determining a target resolution of the to-be-zoomed video frame, and carrying out the zooming of the to-be-zoomed video frame; scaling each to-be-scaled video frame based on the target resolution, and adding image header data in an image header of the scaled video frame to obtain a target video frame; coding each target video frame and the original video frame in sequence to obtain a code stream; and sending the code stream to a decoding end, so that the decoding end decodes the code stream. By means of adding the image header data containing the image identifier and the target resolution in the image header, the capability of flexibly embedding frames with different resolutions in the same video sequence is realized, and the limitation of fixed resolution coding is broken.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Real-time continuous tracking method and device for optic disk

The invention provides an optic disk real-time continuous tracking method and device, and relates to the technical field of medical image processing. According to the method, the video sequence of the ophthalmologic operation is acquired and preprocessed, and then the preprocessed video sequence is subjected to multi-scale feature extraction and optic disc target detection, so that an optic disc detection target is obtained; the optic disk detection target is based on Kalman filtering prediction and appearance feature matching, an adaptive gating matching mechanism is introduced to carry out inter-frame target trajectory association, a target trajectory is maintained in combination with an appearance feature cache updating mechanism and ReID cosine similarity retrieval, and a trajectory tracking result is output; wherein the trajectory tracking result comprises a continuous tracking result of the optic disc position, the confidence coefficient, the trajectory identifier and the timestamp. According to the invention, the detection precision of the small-scale optic disc is improved, the shielding recovery and track continuity are enhanced, and high-precision, high-robustness and real-time optic disc continuous tracking is realized in complex operation scenes of strong reflection, low contrast, local shielding and the like.
Owner:XIAMEN UNIV OF TECH

Adaptive mode switching RGBT target tracking method based on space-time state feedback

The invention relates to an adaptive mode switching RGBT target tracking method based on space-time state feedback, which utilizes a light backbone network fused with an adaptive adjustment mechanism, retains inherent information of each mode, deeply mines hierarchical information between the modes, provides more discrimination features for target tracking, and improves target identifiability. A first frame illumination intensity detection module is introduced to perform illumination intensity classification on a video sequence, so that the target tracking efficiency of a special illumination scene is improved; the space-time information fed back in the tracking process is fully utilized, and the information difference between different modes is combined, so that the tracking quality is effectively evaluated; by using the adaptive mode switching network, the mode selection can be flexibly adjusted according to the tracking condition, so that the tracking quality can be improved, and the tracking efficiency can be considered.
Owner:HENAN UNIV OF SCI & TECH

Intelligent edge flame and smoke identification method based on deep learning

The invention discloses an edge flame and smoke intelligent identification method based on deep learning, and the method comprises the steps: collecting continuous video frames, carrying out the normalization, correction and noise reduction, and generating a preprocessing video sequence; constructing a background static reference frame, and carrying out pixel difference on the background static reference frame and the current frame to generate geometric refraction potential field codes; performing refraction phase mapping on the geometric refraction potential field code to form a refraction phase disturbance tensor; inputting an improved SlowFast model, dynamically adjusting three-branch sampling, and outputting a preliminary candidate region; extracting refraction, phase and energy evolution sequences, constructing a coupling sequence and correcting an identification result; and calculating a risk level, marking a high-risk area, and outputting fire early warning at edge equipment. According to the invention, through constructing the refraction potential field features and the phase disturbance features and combining the improved multi-branch SlowFast deep learning model, rapid, accurate and stable edge side intelligent identification and early warning of flames and smog are realized.
Owner:ZHONGLANG INFORMATION TECH CO LTD

Poultry counting method based on deep learning

The invention discloses a young poultry counting method based on deep learning, and relates to the field of image processing, and the method comprises the steps: information collection and feature extraction: collecting a young poultry image through a hardware platform, and carrying out the preprocessing of the image, and obtaining a continuous color image of a target image; data annotation: carrying out label annotation on the preprocessed image, carrying out accurate target detection annotation on stacked young poultry in the image to obtain an annotated data set, and providing a high-quality data set for model training; inputting the collected real-time image into a young poultry detection system constructed based on a YOLO architecture, wherein the system is used for performing target detection, target tracking and target counting on the input image; and multi-target tracking: continuously positioning spatial positions of a plurality of targets through frame-by-frame analysis of the video sequence, and maintaining a unique identity (ID) of each target. Precise counting and high-speed real-time processing of the young poultry in a high-density scene are realized.
Owner:QINGDAO XINGYI ELECTRONIC EQUIP CO LTD

Multi-agent figure video generation method and device based on potential diffusion model, equipment and storage medium

The invention discloses a multi-agent figure video generation method, device and equipment based on a potential diffusion model and a storage medium, and the method comprises the steps: processing original video data, obtaining a multi-modal input set and multi-modal features, generating state triples corresponding to multiple agents in each time step based on the multi-modal features, and generating state triples corresponding to multiple agents in each time step based on the state triples; inputting the state triad into a pre-constructed hierarchical intention decomposition model to perform intention decomposition, outputting a structure constraint sequence, inputting the structure constraint sequence into a target rendering model to perform denoising in a potential space to obtain a potential video sequence, and decoding the potential video sequence to obtain a target person video; according to the method, the multi-agent cooperation strategy is optimized through reinforcement learning, and the diffusion model is combined to generate the high-quality figure video, so that the behavior consistency, the emotion expression and the interaction coordination of the multi-agent figure video are effectively improved, and the space-time consistency and the natural fidelity degree of the multi-figure video are greatly improved.
Owner:CENT SOUTH UNIV

Bad driving behavior identification method and system, vehicle-mounted terminal equipment and storage medium

According to the bad driving behavior identification method and system, the vehicle-mounted terminal equipment and the computer readable storage medium, the continuous video sequence is acquired, the long-term time sequence features are extracted, the contact event and action stage of the hand and the target object are identified, the initial confidence coefficient is calculated, and the time-space regression processing is performed to generate the steady-state confidence coefficient sequence; the problems that in the prior art, time sequence modeling is weak, tiny target detection precision is insufficient, and the recognition result fluctuates are solved, and the method has the advantages that the long-term time sequence dependency relationship and stage change of bad driving behaviors can be effectively modeled, the recognition precision and stability are improved, and the influence of environmental interference is reduced.
Owner:HUACHEN XINYUAN CHONGQING AUTOMOBILE

Space-time consistent video depth completion method under zero sample unified diffusion framework

The invention discloses a space-time consistent video depth completion method under a zero sample unified diffusion framework. The method comprises the following steps: constructing a depth completion model comprising a variational auto-encoder, a semantic coding network and a space-time diffusion generation network; preparing training data, and generating a frame-level semantic feature vector and a conditional latent variable fusing an original depth and a relative depth for a video frame; training the space-time diffusion generation network in stages by taking the conditional latent variable sequence as input and the semantic features as conditions; in the inference stage, a video sequence to be complemented is processed through a sliding window fusion mechanism, a de-noising depth latent variable is obtained through a trained network, and finally a complemented depth sequence is output through decoding and scale recovery of a variational auto-encoder. The method has the advantages that a depth sequence with measurement consistency, structural integrity and time stability can be generated when depth completion is performed on a sparse, noisy or structurally damaged long sequence video.
Owner:浙江大学宁波国际科创中心

Aircraft maintenance simulation model training method and aircraft maintenance simulation method

The invention provides a training method of an aircraft maintenance simulation model and an aircraft maintenance simulation method, relates to the technical field of aircraft maintenance simulation, and aims to predict future evolution of aircraft maintenance so as to improve authenticity of aircraft maintenance simulation. The method comprises the following steps: acquiring training data related to maintenance simulation; training a maintenance simulation model based on the training data; the aircraft maintenance simulation model comprises a video word segmentation device, a multi-modal input encoder, a multi-modal token sequence and a multi-modal output encoder, wherein the video word segmentation device is used for encoding videos related to aircraft maintenance simulation into a video token sequence; the multi-modal input encoder is used for encoding multi-modal input data related to aircraft maintenance into a multi-modal token sequence; the potential action model is used for determining a potential action representation of a maintenance action based on the video token sequence and the multi-modal token sequence, and the dynamic prediction model is used for predicting a prediction token at the next moment based on the video token sequence, the multi-modal token sequence and the potential action representation so as to simulate a maintenance scene at the next moment.
Owner:CHINA SOUTHERN AIRLINES DIGITAL TECHNOLOGY (GUANGDONG) CO LTD

Air-ground pedestrian re-identification method combining multi-frame information and prompt learning

The invention discloses an air-ground pedestrian re-identification method combining multi-frame information and prompt learning, and the method comprises the following steps: inputting a video sequence into a trained visual encoder model, mapping the same pedestrian at different visual angles into a consistent feature space, and achieving the cross-visual-angle pedestrian re-identification and tracking; according to the visual encoder model, a random rotation transformation strategy of structure perception is introduced in the input embedding stage of a visual encoder backbone network, pedestrian vector features of each frame are rotated, visual angle rotation distortion generated by aerial shooting is simulated, and an enhanced sequence is generated. Extracting global features of the enhanced sequence and the unenhanced sequence through a backbone network; inputting the global feature into an inter-frame information attention module for time dimension attention calculation to obtain an average feature of multi-frame fusion; and then inputting the multi-frame fusion average features into a prompting and guiding visual attention module to generate a text prompt so as to guide model discriminative character features.
Owner:SOUTH CHINA UNIV OF TECH

Sub-picture layout signaling in video coding

A video coding mechanism is disclosed. The mechanism includes receiving a bitstream comprising a sub-picture partitioned from a picture and a sequence parameter set (SPS) comprising a sub-picture size and a sub-picture location. The SPS is parsed to obtain the sub-picture size of the sub-picture and the sub-picture location of the sub-picture. The sub-picture is decoded based on the sub-picture size and the sub-picture location to create a video sequence. The video sequence is forwarded for display.
Owner:HUAWEI TECH CO LTD

Video anomaly detection method and system based on local perception and global reasoning

The invention provides a video anomaly detection method and system based on local perception and global reasoning, and belongs to the technical field of video anomaly detection, and the method comprises the steps: obtaining a to-be-detected continuous video frame sequence; inputting the video frame sequence into a local anomaly sensing module and a global anomaly reasoning module in parallel; wherein the local anomaly sensing module is used for detecting appearance and motion anomalies and time-space dynamic deviation anomalies of a target level and outputting a local anomaly score; the global anomaly reasoning module is used for detecting frame-level scene semantic anomaly and event-level sequential logic anomaly and outputting a global anomaly score; and carrying out standardization, time sequence smoothing and adaptive weighted fusion on the local abnormal score and the global abnormal score to generate a final frame-level abnormal score, and judging whether the video sequence is abnormal or not according to the frame-level abnormal score. According to the method, the comprehensiveness, the accuracy and the robustness of anomaly detection in a complex real scene are effectively improved.
Owner:SHANGHAI SUNGATE NETWORK&INFO TECH CO LTD

Ball-Transform model training method and sphere target tracking method

The invention provides a Ball-Transform model training method and a sphere target tracking method, and the method comprises the steps: processing each frame of image of a video sequence through employing a target detection model, and obtaining a single-frame detection frame matrix; constructing a historical track sequence through a sliding window strategy; obtaining a current frame candidate frame set; a prediction result is generated through the three decision heads of the Ball-Transform model; updating a historical track buffer area based on a sliding window strategy, adding a current frame matching result into the buffer area, and removing historical frames exceeding the window length; executing the following decisions: if the existence probability is lower than an existence threshold, judging that the target disappears; otherwise, selecting the candidate box with the highest matching probability and the rationality score meeting the set threshold value from the candidate box set as the current frame tracking result; and outputting position coordinates and bounding box information of the target in the current frame to ensure accurate tracking of the sphere target.
Owner:恒鸿达(福建)体育科技有限公司

Systems and methods for generating video content using natural language

A computer implemented method for generating video content based on natural language input is disclosed. The method includes receiving a natural language instruction describing one or more desired characteristics of a video. A structured script file comprising at least one story beat is generated using a natural language processing engine. A storyboard comprising one or more storyboard frames is created based on the structured script file. One or more virtual components are generated based on the storyboard. An intermediate video sequence comprising a visual component and an auditory component is created using virtual components and the storyboard. The intermediate video sequence is then refined to produce a modified video sequence by applying one or more post-processing effects.
Owner:RITUAL ADS INC

Heart rate measuring method and system based on time-frequency fusion and dynamic gating attention

The invention discloses a heart rate measurement method and system based on time-frequency fusion and dynamic gating attention, and the method comprises the steps: firstly carrying out the preprocessing of an acquired face video sequence, and obtaining a stable input image sequence; thirdly, extracting a three-dimensional feature map and time sequence features which change along with time through a time-frequency convolutional network; a time-frequency fusion module is constructed, complementary time-frequency features are obtained through time-domain convolution and frequency-domain wavelet decomposition, and feature adaptive fusion is realized by using a dynamic gating mechanism; and the local time relevance of the features is enhanced through a local sliding window attention module. And inputting the three-dimensional convolution features into a prediction module, performing regression to generate a remote photoelectric volume pulse wave signal, and performing frequency domain analysis on the signal to estimate the heart rate. According to the method, pulse related signals can be stably extracted under the conditions of illumination variation and slight head movement, and the method has high heart rate estimation precision and environment robustness and can be used for non-contact vital sign monitoring application.
Owner:ANHUI NORMAL UNIV

Video abstraction method based on multi-modal fusion and dynamic time sequence modeling

The invention relates to a video abstraction method based on multi-modal fusion and dynamic time sequence modeling. The video abstraction method comprises the following steps: respectively extracting features of a video sequence and a corresponding text sequence and projecting the features; a multi-modal video abstract model comprising a dynamic time sequence module, a difference perception feature fusion module and a cross-modal gating Transform module is constructed, the dynamic time sequence module captures inter-frame time sequence dependence and motion information, the difference perception feature fusion module realizes effective fusion of cross-modal information, and the cross-modal gating Transform module performs fine-grained cross-modal interaction; and calculating the importance score of each video frame or text sentence, selecting key frames and key sentences, and generating a multi-modal video abstract. The video sequence is modeled through multi-scale convolution and motion difference perception of the dynamic time sequence module, the time sequence dependency relationship and dynamic change of different time scales in the video are effectively captured, the time sequence feature expression ability is improved, and the time sequence perception ability is enhanced in combination with relative time sequence position coding.
Owner:CHINA THREE GORGES UNIV

Generative adversarial optimization-based rPPG physiological signal reconstruction recognition system

The invention discloses an rPPG physiological signal reconstruction recognition system based on generative adversarial optimization, and relates to the technical field of data processing. The system comprises a signal preprocessing module used for extracting an initial rPPG signal from an input video sequence; and the signal reconstruction module comprises a generator and is used for receiving the initial rPPG signal and outputting a reconstructed rPPG signal. According to the method, by introducing a multi-dimensional discriminator set and physiological prior loss collaborative optimization mechanism, the precision and robustness of rPPG signal reconstruction are remarkably improved. The system can restrain signal quality from multiple angles of time domain, frequency domain and time-frequency domain, and restrain irrational fluctuation in combination with physiological laws, thereby effectively overcoming motion artifacts and illumination interference. Meanwhile, the dynamic region-of-interest selection module adaptively focuses an optimal signal region through a learnable attention mechanism, the input quality is improved from the source, and finally high-reliability estimation of the physiological parameters such as the heart rate and the respiration rate in a complex scene is achieved.
Owner:SHANGHAI LANSHENG RUIFU BIOTECHNOLOGY CO LTD

Adult joint motion range evaluation method and system based on multi-view video

The invention relates to the technical field of human body kinematics parameter measurement, in particular to an adult joint motion range evaluation method and system based on a multi-view video, and the method comprises the steps: obtaining a multi-view video sequence of a testee, and generating an observation data set containing two-dimensional key point observation, a human body segmentation mask and camera parameter information; constructing an individualized joint geometric model containing a bone segment length, a joint center, a joint principal axis and a joint angle coordinate system based on the observation data set, and generating a three-dimensional symbol distance field; constructing a factor graph and performing incremental optimization solution, and outputting a joint angle sequence and an abnormal frame set; and determining an activity range interval and carrying out anti-fact consistency check, and if the validity is not satisfied, calculating an information matrix by a Jacobian matrix to generate a supplementary collection instruction to update a result, thereby improving robustness and verifiability.
Owner:THE FIRST AFFILIATED HOSPITAL OF ZHENGZHOU UNIV

For multi-line intra prediction

A method of and an apparatus for controlling intra prediction for decoding of a video sequence are provided. The method may include based on a reference line index signaling a first reference line nearest to a coding unit, among a plurality of reference lines adjacent to the coding unit, applying intra smoothing on one or more first reference lines comprising the first reference line, among the plurality of reference lines, and based on the intra smoothing being applied on the one or more first reference lines, applying a position-dependent intra prediction combination (PDPC) on one or more third reference lines comprising the first reference line, among the plurality of reference lines, while preventing application of the PDPC on one or more fourth reference lines other than the one or more third reference lines, among the plurality of reference lines.
Owner:TENCENT AMERICA LLC

Method and system for making dome-screen film based on high-altitude aerial image

The invention belongs to the technical field of computer vision and digital media, and discloses a dome-screen film production method and system based on a high-altitude aerial image. Projecting the multi-view aerial video sequence to a unified spherical coordinate system, and constructing a panoramic space with an overlapping region; performing non-rigid geometric deformation processing on the overlapping region of the panoramic space to generate an intermediate image with aligned pixels; calculating the fusion cost of each pixel point in the overlapping region in the intermediate image, and performing path search based on the fusion cost of each pixel point to generate a dynamic suture line; generating an image mask based on the dynamic suture line, performing multi-scale frequency domain adaptive fusion on the intermediate image by using the image mask to generate a fused panoramic image, and remapping the fused panoramic image into a projection format adaptive to dome screen playing to output; and the real reducibility and the manufacturing quality of the dome film are greatly improved.
Owner:SHANGHAI HENGTIAN YIDA ADVERTISEMENT TRANSMISSION CO LTD