Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

35 results about "Fast motion" patented technology

The opposite of fast motion is slow motion. Cinematographers refer to fast motion as undercranking since it was originally achieved by cranking a handcranked camera slower than normal. Overcranking produces slow motion effects. How time-lapse works. Film is often projected at 24 frame/s, meaning 24 images appear on the screen every second.

Space-time fringe analysis method based on physical prior guidance

The invention discloses a space-time fringe analysis method based on physical prior guidance. The method comprises the following steps: firstly, acquiring a stripe sequence taking a target frame as a center, and separating background light through a background estimation network to obtain background light intensity; then, under the assumption of constant brightness and small displacement, an optical flow pixel alignment module is used for estimating inter-frame optical flow, and sub-pixel-level registration is carried out on the subsequences; and finally, inputting the alignment sequence into Attention U-Net with an attention mechanism to carry out space-time phase demodulation, and outputting a target frame phase. And then phase unwrapping is carried out, and three-dimensional reconstruction is completed in combination with system calibration parameters. In the training stage, a space-time consistency loss function and a frequency domain consistency loss function are introduced to improve the generalization ability. The method has high phase demodulation precision and time sequence stability in a low signal-to-noise ratio and fast motion scene, and is suitable for high-speed three-dimensional measurement engineering application.
Owner:NANJING UNIV OF SCI & TECH

RGB-T target tracking method and system based on space-time state evolution

The invention relates to an RGB-T target tracking method and system based on spatio-temporal state evolution, and the method comprises the steps: constructing a double-branch RGB-T target tracking model, enabling RGB and TIR branches to share the weight of a ViT encoder, employing an iterative processing frame, and transmitting spatio-temporal context information between frames through an updatable context memory Tokens. In each iteration, after RGB and TIR modal features are extracted respectively, cross-modal time sequence context modeling is carried out through a modal perception time sequence Mamba module, and the module realizes long-time target perception representation learning through a cross-modal coupling state transition mechanism and a prompt guide strategy; and carrying out intra-modal and inter-modal feature fusion in a spatial dimension through a cross-modal Mama aggregation module, and finally outputting a target position through a prediction head. According to the method, lasting cross-modal state evolution can be realized with linear complexity, the spatial-temporal characteristics of visible light and infrared light are effectively fused, and the tracking robustness and accuracy are improved in complex scenes such as rapid target movement, shielding or modal degradation.
Owner:XIAMEN UNIV OF TECH

A motion estimation based video recognition acceleration method

This invention provides a motion estimation-based method for accelerating video recognition, comprising: identifying keyframes and non-keyframes in a video sequence; extracting features from keyframes using a Bayer domain basic model to obtain perceptual features; calculating motion vectors between non-keyframes and a reference frame (the frame preceding the non-keyframe) using a fast motion estimation module, wherein the fast motion estimation module employs a pyramid block structure and performs multi-level matching search from coarse to fine under a GPU parallel architecture; deforming the features of the reference frame using the motion vectors to obtain propagation features; predicting the perceptual residual of the current frame using a perceptual residual correction network, numerically correcting the propagation features using the residual, and outputting the corrected propagation features, wherein the perceptual residual correction network is a lightweight network structure; and performing video recognition based on the corrected propagation features. This invention achieves faster and more efficient video recognition.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Handheld holder for long-focus shooting device

The application discloses a handheld gimbal capable of carrying a long-focus shooting device, which comprises a handheld part, a yaw shaft assembly and a pitch shaft assembly arranged in layers; the yaw shaft assembly comprises a yaw shaft shell, a yaw shaft driving device installed in the yaw shaft shell and a yaw shaft motor base connected with the yaw shaft driving device, and the top of the yaw shaft motor base is connected with the pitch shaft assembly; the pitch shaft assembly comprises a shaft arm fixedly connected with the yaw shaft motor base, a pitch shaft shell rotationally connected with the shaft arm, a pitch shaft driving device installed in the pitch shaft shell and fixedly connected with the shaft arm, a harmonic reducer connected with the output end of the pitch shaft driving device and fixedly connected with the pitch shaft shell; and the pitch shaft assembly further comprises a quick release part fixedly connected with the pitch shaft shell. The application has the beneficial effect that the "visual target detection / tracking" and the "high-precision execution structure of the two-axis gimbal" are integrated in a closed loop, so that the camera can keep stable and robust in tracking under the condition of rapid movement / occlusion.
Owner:DAWEI HONGYI ROBOT TECHNOLOGY (CHONGQING) CO LTD

Self-adaptive time modeling driven motion scene human body posture estimation method and medium

The invention discloses a motion scene human body posture estimation method driven by adaptive time modeling and a medium, and the method comprises the steps: constructing an adaptive time modeling motion scene 3D human body posture estimation network model based on the motion speed degree; the robustness of the model to a motion scene is improved by using motion prior based on physical constraints and predicting the uncertainty of 3D human body articulation points; human body motion videos with different motion speeds are used as input, a corresponding 2D human body joint point sequence is obtained on the basis of an existing high-quality 2D human body posture estimation model, the 2D human body joint point sequence is input into the network model, and a trained model is obtained; according to the method, the receptive field can be adaptively adjusted according to the movement speed, the small receptive field is used for capturing transient changes for fast movement, the large receptive field is used for capturing long-term dependence for slow movement, the method adapts to movement at different speeds, and therefore the human body posture estimation accuracy in movement scenes at different speeds is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

An event camera and vision camera cooperative-based motion training evaluation system and method

The application discloses a motion training evaluation system and method based on cooperation of an event camera and a visual camera, and is applied to the field of computer vision, and aims at the problems that an existing training evaluation system based on an optical camera cannot capture fast motion trajectories and limb shapes, cannot effectively evaluate action quality, and cannot provide effective training guidance; the application provides continuous human motion events in the form of event streams, can achieve millisecond-level motion response, is not affected by motion blur effects of high-speed moving objects, and can provide a higher dynamic range, and can provide more effective motion training evaluation in scenes with strong light, backlight, and sharp changes in brightness. Through cooperation of the visual camera and the event camera, not only the accuracy of fast actions can be evaluated, but also standard action teaching and training in the form of visual scene interaction can be performed.
Owner:CHENGDU UNIV

XR Virtual-Real Interactive Control Device with Optical Motion Capture and Inertial Navigation Fusion

This invention belongs to the field of virtual reality technology, specifically relating to an XR virtual-real interaction control device that integrates optical motion capture and inertial navigation. It includes: a metasurface polarization-encoded marker array, distributed and installed at preset tracking point positions on the user's hand and head, used to generate reflected light carrying polarization identification codes when receiving illumination light; a miniature inertial measurement unit array, rigidly connected to each metasurface polarization-encoded marker in the metasurface polarization-encoded marker array, used to collect acceleration and angular velocity data at each tracking point; and a near-eye display device, including an illumination module and a polarization camera module. The illumination module projects linearly polarized illumination light onto the user's hand and head areas, and the polarization camera module receives the reflected light generated by the metasurface polarization-encoded marker array and outputs multi-channel polarization images. This invention achieves stable tracking in occlusion and rapid motion scenarios through tightly coupled visual-inertial fusion.
Owner:SICHUAN WUTONG TECH CO LTD

Whole body attitude estimation method and system based on downward fisheye and enhanced EgoPoseFormer model

The invention relates to a whole body attitude estimation method and system based on a downward fisheye and an enhanced EgoPoseFormer model, and the method enlarges the human body capture range through the layout of the bottom visual angle of a head-mounted display, compensates geometric nonlinearity in a feature extraction stage through distortion perception convolution, cooperates with an attention mechanism of an embedded pose compensation item, and achieves the estimation of the whole body attitude. And the feature space consistency is maintained when the camera dynamically shakes. Meanwhile, in combination with cross-frame time sequence gating and a multi-modal filtering algorithm, inertial data is utilized to perform kinematics correction on a visual predicted value, and smooth track output is realized in a shielding or rapid motion scene. According to the scheme, the computing power overhead is reduced through model quantification and operator fusion, and low-delay real-time whole body attitude tracking is realized in a mobile terminal chip environment.
Owner:HANGZHOU WUZHI MIXED REALITY TECHNOLOGY CO LTD

Embedded firmware camera data modification methods, devices, media and equipment

This invention discloses an embedded firmware camera data modification method, apparatus, medium, and device, relating to the field of computer vision technology. The method includes: real-time acquisition of image data and preliminary processing; using frame difference and background modeling algorithms to divide the pre-processed image data into regions and assign priorities, generating a region priority mapping table; extracting feature points of high-priority motion regions in the optimized image data using optical flow, and calculating motion vectors between consecutive frames; triggering an event-driven frame processing mechanism based on the calculated motion vectors between consecutive frames, automatically adjusting the frame rate and processing accuracy; and using the motion vectors extracted by optical flow to calculate the object's motion speed in real time, dynamically adjusting the frame rate and processing accuracy according to the motion speed, ensuring sufficient detail is captured during rapid motion while saving processing resources when stationary.
Owner:深圳市云希谷科技有限公司

Video frame optimization processing method based on computer vision

The invention discloses a video frame optimization processing method based on computer vision, and the method comprises the following steps: decoding an input video stream, obtaining a frame sequence, and carrying out the preprocessing of the frame sequence, and obtaining a basic correction frame; constructing an initial noise spore graph, and executing time sequence iteration according to a diffusion and attenuation equation to generate a noise spore evolution field; dividing the evolution field into three types of noise spores, determining a dominant noise spore type according to a competition rule, and forming a competed noise spore field; adjusting enhancement network parameters pixel by pixel to generate a noise self-adaptive enhancement frame; and performing time sequence consistency correction and quality evaluation on the enhanced frame, and outputting a final optimized video frame sequence. According to the method, the noise spore field is constructed, and pixel-by-pixel enhancement and time sequence correction are performed based on the noise type and density, so that the video can obtain a more stable, clearer and cross-frame consistent optimization result in a complex noise and rapid motion scene.
Owner:SHENZHEN JIPAI TECHNOLOGY CO LTD

Display method for prompting image switching according to image deflection angle

The invention discloses a display method for prompting image switching according to an image deflection angle, and the method comprises the steps: obtaining attitude sensor data, calculating attitude stability reflecting a jitter degree, and constructing an attitude stability state machine for distinguishing a motion state from a stable state according to the attitude stability; in a stable state, dynamically calculating a hysteresis adjustment coefficient by the system according to a current stability numerical value and historical mode preference of a user so as to stretch or translate a switching threshold; and in the motion state, forcibly locking the switching threshold value of the previous frame, and generating a continuous self-adaptive hysteresis threshold value sequence. And judging and switching the display mode in combination with the self-adaptive threshold value and the frame synchronization signal. According to the invention, through the coupling of the state machine gating and the adaptive hysteresis algorithm, the problem of wrong switching in the rapid movement process of the equipment is effectively solved, the adaptive matching of different hand shake degrees and user habits is realized, and the fluency and accuracy of display mode switching are ensured.
Owner:NANJING TUGE HEALTHCARE CO LTD

Multi-person posture fusion and interaction display system based on multi-path image data

The invention relates to a multi-user posture fusion and interaction display system based on multi-path image data, which belongs to the technical field of computer vision and man-machine interaction and comprises a multi-camera acquisition and analysis module, a data coupling module, a real-time transmission module and a three-dimensional reduction and self-adaptive correction module. Wherein the multi-camera acquisition and analysis module calls an attitude estimation engine to obtain three-dimensional coordinates of skeleton key points, and calls the segmentation image matting module to obtain a target mask; the data coupling module is used for associating a user identifier based on the multi-view geometry and time sequence consistency score, and generating a skeleton pose and jitter score; the real-time transmission module executes back pressure transmission according to the queue backlog depth; the three-dimensional reduction and self-adaptive correction module drives a virtual character and shielding rendering in a three-dimensional rendering engine; according to the method, the problem of identity loss under multi-person shielding and rapid movement is solved, and the recognition accuracy is remarkably improved.
Owner:HANGZHOU XINJUE TECH CO LTD

A method for training a system for automated detection of the stroboscopic effect

PCT designated stageWO2026015345A1Image enhancementImage analysisLight activationFrame sequence
A stroboscopic device to use a camera and a light source with software that generates synthetic and real video frames, compares the frames via a discriminator, repeat until a convergenc point is achieved where the discriminator can't reliably tell the real video frames from the synthetic video frames, analyzes a dense optical flow field between a first and second frame in a sequence of video frames to detect a motion, and controls a light activation to achieve a stroboscopic effect, wherein once achieved, a system identifies an object and automatically detects movement or vibration of the idenfied object using a trained model, enabling precise visualization and analysis of fast motions in video sequences.
Owner:IOT TECHNOLOGIES LLC

Fast relocalization method based on sparse semantic anchor points and inertial sensor tight coupling

This invention discloses a fast relocalization method based on tight coupling between sparse semantic anchors and inertial sensing, belonging to the fields of augmented reality and computer vision. It constructs a lightweight sparse semantic map by extracting sparse feature points with stable geometric and semantic attributes from the environment as sparse semantic anchors. When device tracking is unstable or lost, inertial navigation and visual matching threads based on sparse anchors are activated in parallel. The core PnP algorithm is used to quickly recover the visual pose, and the method is tightly coupled with inertial data for optimization, ultimately outputting a high-precision, smooth six-DOF pose. This solves the problem of excessively long relocalization time and experience interruption in traditional visual SLAM scenarios with fast motion and weak textures, achieving millisecond-level, user-unnoticed tracking recovery, significantly improving the robustness and user experience of augmented reality systems.
Owner:CHONGQING AEROSPACE POLYTECHNIC COLLEGE

Event scene text recognition method and system based on chain thinking reasoning

The invention discloses an event scene text recognition method and system based on chain thinking reasoning, and belongs to the technical field of event cameras, and the system comprises an event feature extraction module, a query alignment module, a text generation and thinking chain reasoning module, a joint training module and a chain thinking data generation module. According to the method, the problem that the recognition performance of traditional text recognition based on RGB images is reduced in low-illumination, fast-motion and high-dynamic scenes is solved, interpretability of event stream text recognition is achieved by introducing a chain reasoning mechanism, and the method simultaneously comprises a'text recognition result 'and a'thinking chain reasoning result' and is high in recognition efficiency. Therefore, the transparency and credibility of model decision making are improved. The context reasoning ability of the model is enhanced through the design of alignment of visual and language features, so that the model can still deduce correct text content under the condition that visual information is incomplete or noise exists, and the robustness under the scenes of low illumination, uneven exposure and motion blurring is remarkably improved.
Owner:ANHUI UNIV

Lightweight visual-inertial three-dimensional reconstruction method, system, medium and device for fast motion scene

This invention relates to the fields of computer vision and autonomous robot navigation, and discloses a lightweight vision-inertial 3D reconstruction method, system, medium, and device for fast-moving scenarios. The method includes: simultaneously acquiring continuous image sequences and high-frequency inertial data; obtaining the initial pose prediction value of the camera in the current frame based on the high-frequency inertial data; parameterizing the initial spatial point cloud using compact 3D Gaussian primitives to construct an initial lightweight 3D Gaussian map; extracting the opacity parameter, single scalar scale parameter, and historical observation information of the compact 3D Gaussian primitives within the current effective view frustum to generate a binarized reliability hard mask to constrain residual calculation and gradient backpropagation regions; constructing a multimodal joint tracking loss function to update the current frame camera pose; performing deduplication detection using spatial hashing, instantiating new compact 3D Gaussian primitives, removing redundant or low-contribution primitives, and reclaiming corresponding memory addresses to generate a lightweight 3D Gaussian map with controlled memory usage.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Video decoupling network and method for video target segmentation

The invention provides a video decoupling network and method for video target segmentation, and the network comprises a visual element decoupling module, a unified prior space-time decoupler, and a self-adaptive expert hybrid reconstruction module, and the method comprises the steps: decomposing a historical video frame into a motion representation, a scene representation, and an instance representation through the visual element decoupling module, and constructing a corresponding prior; then single element representation is constructed and analyzed through a unified prior space-time decoupler, finally, multiple visual clues are adaptively integrated through an adaptive expert hybrid reconstruction module, and multiple decoupled object clues are connected to generate global feature representation; multi-source information is fused to reconstruct a high quality object mask. According to the method, the modeling capability of fine-grained motion and target boundaries can be enhanced while the global consistency of the video is kept, so that the accuracy and robustness of video object segmentation are effectively improved, and complex scenes such as occlusion and rapid motion are effectively processed through combination of long-term and short-term reasoning and an attention mechanism; and meanwhile, the stability and reliability of a segmentation result are improved.
Owner:LANZHOU UNIV

A visual-inertial slam method based on deep learning

PendingCN122415739AAlgorithmVisual perception
This invention discloses a deep learning-based visual-inertial SLAM method. It acquires visual data from a monocular camera and inertial data from an IMU, inputting them into a pre-constructed visual-inertial SLAM model to obtain camera pose estimation results. The method inputs image sequences into a SuperPoint feature extraction network to extract SuperPoint feature points and descriptors for each frame. The feature points and descriptors of consecutive frames are then input into a LightGlue feature matching network to obtain initial matching point pairs. A RANSAC matching enhancement module based on descriptor consistency constraints is used for pre-screening to obtain optimized matching point pairs. These optimized pairs are then combined with IMU pre-integration results for preliminary pose estimation. Joint optimization is performed through sliding window nonlinear optimization, and loop closure constraints are introduced via a loop closure detection module to output the final camera pose estimation result. This invention solves the problems of unstable feature extraction, low matching accuracy, and easy trajectory drift or even tracking failure in traditional visual-inertial SLAM systems under complex environments such as low light, rapid motion, and texture degradation.
Owner:CHONGQING UNIV OF TECH

Object detection method and system based on adaptive time window and time threshold

The application discloses an object detection method and system based on adaptive time window and time threshold, and the method comprises the following steps: acquiring event camera data, acquiring an event set in an optimal time window through an adaptive time window method; eliminating a part of events generated by self-motion of the event camera from the event set in the optimal time window through an adaptive compensation algorithm, and acquiring a compensated event set; constructing an event frame image through an adaptive threshold method according to the acquired optimal time window and the compensated event set; and performing iterative image matrix fitting on the constructed event frame image to acquire a final detection result. The application provides an object detection method based on adaptive time window and time threshold, realizes detection of fast-moving objects through an adaptive time window, self-motion compensation of a camera, a threshold of a time image and an iterative image matrix fitting method, thereby realizing the requirement of obstacle avoidance of fast-moving obstacles in the environment and improving the obstacle avoidance performance of the unmanned aerial vehicle.
Owner:WUHAN UNIV

An unsupervised rockfall monitoring method based on memory-augmented network

The application discloses a kind of based on memory enhancement network's unsupervised rockfall monitoring method, it is related to rockfall monitoring technical field, mainly includes: the memory enhancement network and total loss function including encoder, memory module, motion weighting module and decoder are constructed, memory enhancement network is trained using total loss function and normal mountain monitoring video dataset to obtain rockfall monitoring model;The video frame sequence after pre-processing is predicted using the rockfall monitoring model to obtain target prediction frame, according to this, calculate abnormal score, and dynamically update memory bank using double threshold time series gating strategy, according to target mountain monitoring video, pixel-level reconstruction error and abnormal score generate multi-modal monitoring result.The implementation based on memory enhancement network's unsupervised rockfall monitoring method provided in the application can improve the monitoring sensitivity of small target fast motion and enhance the robustness of short-time anomaly under the condition of labeled rockfall data scarcity.
Owner:CHINA UNIV OF GEOSCIENCES (WUHAN)

Robust visual SLAM system based on fuzzy classification and differential deblurring

The invention discloses a robust visual SLAM system based on fuzzy classification and differential deblurring, and belongs to the technical field of image processing, multi-sensor fusion and visual SLAM. The system constructs a closed-loop processing framework of fuzzy discrimination-differential deblurring-feature matching enhancement-SLAM integration, and comprises an image and IMU data acquisition module, a fuzzy discrimination module, a differential deblurring module, an improved GMS feature matching module and a visual SLAM core module. The blurring discrimination module combines image gradient features and IMU motion information to realize blurring degree discrimination of the image, and further distinguishes global blurring and local blurring for repairable blurring; the differential deblurring module introduces IMU constraint deblurring or lightweight deblurring processing for different types of blurring so as to meet the SLAM feature extraction requirement; the improved GMS feature matching module improves the matching stability in a fuzzy scene through a multi-scale and adaptive neighborhood mechanism. And the visual SLAM core module fuses the clear image and the deblurred image to realize high-precision positioning and mapping. According to the method, the robustness and real-time performance of visual SLAM in dynamic and fast motion scenes are effectively improved, and the method is suitable for application scenes such as robot navigation, automatic driving, unmanned aerial vehicles and AR / VR.
Owner:KUNMING UNIV OF SCI & TECH

Cable diameter detection method and system based on machine vision

The invention relates to the technical field of machine vision, and discloses a cable diameter detection method and system based on machine vision. According to the method and system, spatial-temporal feature modeling is carried out by introducing the pulse neural network, cable images can be captured and analyzed in real time under the conditions of complex illumination and motion blur, it is ensured that the error of cable diameter measurement is controlled within the high-precision range smaller than or equal to 0.05 mm, and the measurement accuracy is improved. The design of a shielding compensation mechanism effectively solves the problem of image missing caused by rapid movement and a shielding object, so that the detection reliability is further improved, and the embedded design of the system not only reduces the power consumption by 40%, but also enables the system to adapt to a high-speed production environment, and meets the requirements of the modern industry for automation and intelligence.
Owner:NAT INSTR SMART EYES (CHONGQING) TECH CO LTD

Monocular video multi-person 3D human body motion reconstruction method based on multi-module fusion

The invention relates to a monocular video multi-person 3D human body motion reconstruction method based on multi-module fusion, and solves the defect that discontinuous or unstable reconstruction is generated when a target appears again due to the fact that tracking is lost when the target is partially shielded or moves rapidly compared with the prior art. The method comprises the following steps: acquiring a single-frame or multi-frame motion video sequence; constructing a motion perception semantic tracking module; generating a coherent human body grid sequence; predicting future attitude features in the occlusion scene; and outputting the 3D motion state. According to the invention, 3D motion reconstruction with robust shielding, stable identity and consistent time sequence is realized through collaborative design of motion perception tracking, time sequence enhancement reconstruction, motion prediction and adaptive fusion.
Owner:ANHUI UNIV

Method and system for scalable scene 3d reconstruction based on uav onboard localization

The application discloses a kind of based on unmanned aerial vehicle airborne positioning extensible scene three-dimensional reconstruction method and system, based on the color image and depth image obtained by depth camera carried on unmanned aerial vehicle are carried out three-dimensional reconstruction, including in the unit ball in 6D state space, particle is uniformly sampled, obtains particle group template;Surface measurement is carried out to each frame depth image input, the pixel-by-pixel projection of depth image is projected into three-dimensional space and the normal of each three-dimensional point is calculated, while calculating segmented normal map;According to the correlation between three-dimensional point normal, adaptively allocate reconstruction voxel memory in GPU;Only using depth image and the three-dimensional model that exists in current GPU active space has been constructed, the camera pose of fast motion is tracked;With the continuous transmission of depth image frame sequence, based on truncated signed distance field TSDF fusion measurement value with sensor noise, construct dense three-dimensional point cloud model, extract three-dimensional surface when needing visualization, generate three-dimensional mesh model.
Owner:WUHAN UNIV

Lossless compression method for eliminating inter-frame redundancy based on partitioning and probability matching

The invention relates to the technical field of lossless video coding, and provides a lossless compression method for eliminating inter-frame redundancy based on partitioning and probability matching, which comprises the following steps of: dividing a current frame into a plurality of sub-blocks, and classifying each sub-block by calculating the similarity between a current block and a matching block; the classified sub-blocks are subjected to matched coding operation through a differentiation processing strategy, and compressed data are generated; and carrying out entropy coding on the compressed data, and outputting a lossless compressed code stream. The video frame is divided into a plurality of sub-blocks through the technical means of combining block processing, probability matching, multi-level classification and reversible integer predictive coding, the sub-blocks are classified according to the inter-frame similarity, differentiated processing strategies are adopted for different types, and finally residual and motion information is compressed through entropy coding. The inter-frame redundancy is effectively eliminated, the compression efficiency is remarkably improved, and meanwhile, the robustness of fast motion and scene switching is kept.
Owner:CHANGZHOU WHISPER TECH CO LTD

Multi-view-angle three-dimensional human body posture estimation method based on dynamic time sequence fusion and view angle selection

The invention discloses a multi-view-angle three-dimensional human body posture estimation method based on dynamic time sequence fusion and view angle selection, belongs to the technical field of human body posture estimation, and aims to solve the problems of insufficient accuracy and poor robustness when a traditional method is used for processing complex scenes such as serious shielding, rapid movement and large view angle quality difference. The method comprises the following steps: firstly, constructing a space-time correlation model based on feature difference dynamic time sequence fusion, and realizing adaptive adjustment of time sequence fusion weight by introducing feature difference measurement and a motion state judgment mechanism between adjacent frames and combining a time attenuation factor and inter-frame motion consistency; secondly, designing a dynamic selection strategy based on view angle quality evaluation, and constructing a multi-view angle weight distribution model; and finally, realizing high-precision three-dimensional human body posture reconstruction through cross-view feature alignment and space-time consistency constraint.
Owner:JILIN UNIVERSITY

Video generation method and device, electronic equipment and computer readable storage medium

The invention provides a video generation method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of computer vision. According to the scheme, after a prompt text is obtained, static information is coded to obtain a first text feature, and motion information is coded to obtain a second text feature; the first text feature can be coded to generate a first potential representation and decoded to obtain an RGB image frame, and the second text feature can be coded to generate a second potential representation and decoded to obtain an optical flow field; and generating a target video by using the RGB image frame and the optical flow field. According to the scheme, the static information in the prompt text is used for executing RGB image reconstruction to provide rich visual appearance information, the motion information in the prompt text is used for executing optical flow field reconstruction to provide natural, smooth and real motion change information, mutual complementation can be achieved, the video generation requirements of large and fast motion change amplitude or complex actions are met, and the video generation efficiency is improved. The video generation quality is improved, the deviation is reduced, and the video content better conforms to the physical law.
Owner:JINGDONG CITY BEIJING DIGITS TECH CO LTD

A high-speed high-maneuver three-dimensional trajectory prediction method and system combining event imaging and motion decoupling

The present application belongs to the technical field of trajectory prediction, and discloses a high-speed high-maneuver three-dimensional trajectory prediction method and system combining event imaging and motion decoupling. The method comprises the following steps: S1, converting an event stream collected at any target time into an event stream at a reference time; S2, calculating the three-dimensional space coordinates corresponding to each target object in the event stream at the reference time; S3, repeating S2 to obtain the three-dimensional space coordinates of each target object in different time windows, i.e. the three-dimensional motion trajectory of each target object; S4, calculating the motion distance and motion direction of the prediction point by using the first plurality of trajectory points in the three-dimensional motion trajectory, thereby realizing trajectory prediction. Through the present application, the problems of difficult three-dimensional trajectory prediction and low prediction accuracy caused by the fast motion speed and random direction of high-speed high-maneuver targets are solved.
Owner:HUAZHONG UNIV OF SCI & TECH

An unmanned ship vision-inertial positioning method based on moving targets and near-shore features

The application particularly relates to a kind of unmanned ship vision inertial positioning methods based on mobile target and nearshore features, first using camera positioning overcomes the unstable problem of GPS signal in nearshore scene;Secondly, the water and land area division of camera image is carried out by introducing semantic segmentation, which avoids the influence of unstable feature points on water surface on visual positioning;Then, the IMU sensor is introduced to overcome the disadvantage of picture blur when camera is in fast motion;Finally, the Mask R-CNN algorithm is used, and the function of using the position information of mobile target is added, which improves the accuracy of the algorithm.The method has high-precision positioning function in static environment and dynamic environment with mobile target, and overcomes the defect that traditional vision inertial positioning algorithm does not fully utilize the position information of mobile target.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

A marker point detection method, device, equipment and medium

ActiveCN121366197BMarking outClosed loop
The application discloses a kind of mark point detection method, device, equipment and medium, it is related to equipment manufacturing technical field, the method comprises: according to the preset reference position of mark point on the object to be processed, the initial coordinates of listening point in machine coordinate system are determined, and then the shortest movement route is planned according to it, the inflection point of broken line in route is converted into arc, the new coordinates of listening point are obtained by optimizing trajectory and adjusting;Subsequently, send new coordinates to control component, send moving instruction to motion control motor, make crossbeam move according to optimized trajectory, camera is photographed by control component;The offset of mark point relative to listening point is obtained by picture recognition, to obtain the coordinates of mark point in machine coordinate system.This realizes smooth and fast movement, reduces sampling preparation time;Sampling is avoided in movement Trigger, and waiting for shutdown is improved sampling efficiency, and sampling and identification are asynchronous, and identification does not interrupt movement and subsequent sampling, form parallel closed loop, guarantee sampling continuous and efficient.
Owner:HANGZHOU IECHO SCI & TECH CO LTD