Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

62 results about "Fast motion" patented technology

The opposite of fast motion is slow motion. Cinematographers refer to fast motion as undercranking since it was originally achieved by cranking a handcranked camera slower than normal. Overcranking produces slow motion effects. How time-lapse works. Film is often projected at 24 frame/s, meaning 24 images appear on the screen every second.

Power cross-operation dynamic risk prediction and early warning method, system and device based on multi-modal data fusion and medium

The invention discloses an electric power cross-operation dynamic risk prediction and early warning method, system, equipment and medium based on multi-modal data fusion, and relates to the technical field of intelligent electric power systems, and the method comprises the steps: collecting multi-source data of a cross-operation site in real time, including personnel spatial position data, equipment state data, environmental parameters and operation flow information; performing multi-modal fusion processing on the multi-source data, and constructing space-time risk association features; constructing a dynamic risk factor matrix, performing fusion calculation on the static risk reference value and the dynamic correction value, and dynamically updating the weight of each factor along with time; predicting a risk change trend based on the combination of a time sequence prediction model and a graph structure modeling method; and generating a three-dimensional visual risk thermodynamic diagram based on a prediction result, and triggering a multi-level early warning mechanism when a risk value exceeds a threshold value. The image is reconstructed through multi-modal fusion processing, and key edge features can be recovered in complex environments such as low illumination, high dynamic range and fast motion.
Owner:YUNNAN POWER GRID CO LTD KUNMING POWER SUPPLY BUREAU

Video image segmentation method

The invention relates to the field of image processing, and discloses a video image segmentation method which is used for improving the precision, robustness and real-time performance of video image segmentation in a complex dynamic scene. The video image segmentation method comprises the following steps: generating an entropy generation rate map through weighted combination of light flow divergence and rotation, and quantifying motion irreversibility; a double-virtual-form prime concentration field is constructed, next-frame texture prediction is realized through iterative evolution, and the dynamic background adaptability is enhanced; fusing the projection entropy generation rate graph and the enhanced texture residual graph, dynamically distributing motion and appearance weights, and generating a high-precision boundary response graph; and extracting and persistent filtering are carried out, topological consistency maintenance of segmentation masks is realized, multi-scale collaborative segmentation and dynamic computing resource scheduling are supported, and efficiency and precision are balanced. The segmentation precision of the method is obviously superior to that of a traditional method in complex scenes such as illumination variation and rapid motion, and the method is suitable for the fields with high real-time requirements such as monitoring, medical treatment and automatic driving.
Owner:JIANGSU HUIHANG DIGITAL TECHNOLOGY CO LTD

Target identification tracking method based on self-supervision mechanism

The invention belongs to the technical field of computer vision, and discloses a target identification tracking method based on a self-supervision mechanism, and the method comprises the steps: enhancing a self-supervision pre-training module through causality, constructing a causal sample pair through unlabeled video data, learning universal features through combining with comparison loss, and achieving the high-precision tracking without large-scale manual labeling. A multi-modal feature fusion and dynamic calibration mechanism further reduces dependence on annotated data, is especially suitable for industrial inspection, field monitoring and other scenes where data acquisition is difficult, significantly reduces time and labor costs in a data preparation stage, and broadens the application range of the technology in resource limited scenes; a causal reasoning and physical constraint mechanism is introduced, a dynamic relation between targets is modeled through a space-time causal graph, unreasonable tracks are filtered in combination with a physical rule, and complex conditions such as shielding, rapid movement and extreme weather are effectively dealt with; the dynamic feature calibration module corrects feature drift in real time, and ensures stable model performance in long-term tracking.
Owner:ZHONGSHOU DIGITAL TECH CO LTD

Multi-target tracking method and device, electronic equipment and computer readable storage medium

The invention provides a multi-target tracking method and device, electronic equipment and a computer readable storage medium, and the method comprises the steps: carrying out the target detection of a video sequence, and obtaining a detection frame of a current frame; under the condition that the detection frame of the current frame is successfully matched with the prediction frame, calculating a track consistency score; for the target trajectory of which the trajectory consistency score exceeds a set score, extracting feature representation of a historical frame by using a time sequence attention mechanism; and calculating feature similarity according to the feature representation and a set model, and determining the identity of the target trajectory. According to the method, the video frame is detected through the target detection technology, the corresponding high-confidence detection box is screened out, the purposes of accurately positioning multiple targets and reducing false detection can be achieved, and reliable input is provided for follow-up track association. The target position is predicted through the trajectory prediction technology, matching of a detection frame and a prediction frame is achieved, the purpose of maintaining trajectory continuity in a shielding or rapid motion scene is achieved, and the target identifier switching rate is remarkably reduced.
Owner:SHANGHAI JIDOU TECH CO LTD

A UAV target tracking method

This invention discloses a target tracking method for unmanned aerial vehicles (UAVs), relating to the field of computer vision, aiming to solve the problems of target occlusion, rapid movement, and efficiency in dynamic scenes. The method includes: acquiring the target's bounding box and confidence score through a detector, and predicting the target's position and state using a noisy adaptive Kalman filter; proposing a confidence-based feature extraction and judgment strategy, performing IoU matching between the trajectory and the detection box; for targets with a matching degree below a threshold and a confidence score change rate above a threshold, extracting appearance features using a lightweight feature re-identification network (Rep-OSNet); otherwise, reusing features from the previous frame; designing a cost function that fuses motion direction and appearance similarity to perform cascaded matching between the target and the confirmed trajectory; performing secondary association matching on targets and trajectories that mismatch in the cascaded matching; updating the matched trajectory state using a noisy adaptive Kalman filter, deleting long-term lost mismatched trajectories, and outputting the target trajectory prediction box and ID.
Owner:WEIYUAN SHENGXIANG COMPOSITE MATERIAL CO LTD

Nanoscale high-precision positioning method based on multi-degree-of-freedom motion workbench

The invention discloses a nanoscale high-precision positioning method based on a multi-degree-of-freedom motion workbench, and relates to the field of large-stroke and high-precision positioning, and the method comprises the steps: S1, initializing the motion workbench, and guaranteeing that a system is in a zero position state; s2, sending a target position instruction by the upper computer, and driving the macro moving platform to quickly move to a position near a target; s3, acquiring the current position, calculating the deviation from the target position according to the target displacement and the current actual feedback displacement value, and generating micropositioner control parameters; s4, a control instruction is generated, and nanoscale displacement compensation is executed; and S5, iteratively correcting the displacement of the micropositioner until the deviation fed back by the laser interferometer meets the requirement. According to the invention, the positioning precision of large-stroke (hundred millimeters) movement is improved to a nanometer level, the device is suitable for the fields of photoetching machines, ultra-precision machining, biological micromanipulation and the like, and the equipment performance and the process level are improved. In addition, the system adopts a modular design, can adapt to different driving and sensing units according to requirements, and has strong expansibility.
Owner:TONGLING UNIV

Space-time fringe analysis method based on physical prior guidance

The invention discloses a space-time fringe analysis method based on physical prior guidance. The method comprises the following steps: firstly, acquiring a stripe sequence taking a target frame as a center, and separating background light through a background estimation network to obtain background light intensity; then, under the assumption of constant brightness and small displacement, an optical flow pixel alignment module is used for estimating inter-frame optical flow, and sub-pixel-level registration is carried out on the subsequences; and finally, inputting the alignment sequence into Attention U-Net with an attention mechanism to carry out space-time phase demodulation, and outputting a target frame phase. And then phase unwrapping is carried out, and three-dimensional reconstruction is completed in combination with system calibration parameters. In the training stage, a space-time consistency loss function and a frequency domain consistency loss function are introduced to improve the generalization ability. The method has high phase demodulation precision and time sequence stability in a low signal-to-noise ratio and fast motion scene, and is suitable for high-speed three-dimensional measurement engineering application.
Owner:NANJING UNIV OF SCI & TECH

Recursively-Cascading Diffusion Model for Image Interpolation

Despite recent progress, existing frame interpolation methods still struggle with extremely high resolution images and challenging cases such as repetitive textures, thin objects, and fast motion. To address these issues, provided is a cascaded diffusion frame interpolation approach that excels in these scenarios while achieving competitive performance on standard benchmarks.
Owner:GOOGLE LLC

RGB-T target tracking method and system based on space-time state evolution

The invention relates to an RGB-T target tracking method and system based on spatio-temporal state evolution, and the method comprises the steps: constructing a double-branch RGB-T target tracking model, enabling RGB and TIR branches to share the weight of a ViT encoder, employing an iterative processing frame, and transmitting spatio-temporal context information between frames through an updatable context memory Tokens. In each iteration, after RGB and TIR modal features are extracted respectively, cross-modal time sequence context modeling is carried out through a modal perception time sequence Mamba module, and the module realizes long-time target perception representation learning through a cross-modal coupling state transition mechanism and a prompt guide strategy; and carrying out intra-modal and inter-modal feature fusion in a spatial dimension through a cross-modal Mama aggregation module, and finally outputting a target position through a prediction head. According to the method, lasting cross-modal state evolution can be realized with linear complexity, the spatial-temporal characteristics of visible light and infrared light are effectively fused, and the tracking robustness and accuracy are improved in complex scenes such as rapid target movement, shielding or modal degradation.
Owner:XIAMEN UNIV OF TECH

A motion estimation based video recognition acceleration method

This invention provides a motion estimation-based method for accelerating video recognition, comprising: identifying keyframes and non-keyframes in a video sequence; extracting features from keyframes using a Bayer domain basic model to obtain perceptual features; calculating motion vectors between non-keyframes and a reference frame (the frame preceding the non-keyframe) using a fast motion estimation module, wherein the fast motion estimation module employs a pyramid block structure and performs multi-level matching search from coarse to fine under a GPU parallel architecture; deforming the features of the reference frame using the motion vectors to obtain propagation features; predicting the perceptual residual of the current frame using a perceptual residual correction network, numerically correcting the propagation features using the residual, and outputting the corrected propagation features, wherein the perceptual residual correction network is a lightweight network structure; and performing video recognition based on the corrected propagation features. This invention achieves faster and more efficient video recognition.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Handheld holder for long-focus shooting device

The application discloses a handheld gimbal capable of carrying a long-focus shooting device, which comprises a handheld part, a yaw shaft assembly and a pitch shaft assembly arranged in layers; the yaw shaft assembly comprises a yaw shaft shell, a yaw shaft driving device installed in the yaw shaft shell and a yaw shaft motor base connected with the yaw shaft driving device, and the top of the yaw shaft motor base is connected with the pitch shaft assembly; the pitch shaft assembly comprises a shaft arm fixedly connected with the yaw shaft motor base, a pitch shaft shell rotationally connected with the shaft arm, a pitch shaft driving device installed in the pitch shaft shell and fixedly connected with the shaft arm, a harmonic reducer connected with the output end of the pitch shaft driving device and fixedly connected with the pitch shaft shell; and the pitch shaft assembly further comprises a quick release part fixedly connected with the pitch shaft shell. The application has the beneficial effect that the "visual target detection / tracking" and the "high-precision execution structure of the two-axis gimbal" are integrated in a closed loop, so that the camera can keep stable and robust in tracking under the condition of rapid movement / occlusion.
Owner:DAWEI HONGYI ROBOT TECHNOLOGY (CHONGQING) CO LTD

Self-adaptive time modeling driven motion scene human body posture estimation method and medium

The invention discloses a motion scene human body posture estimation method driven by adaptive time modeling and a medium, and the method comprises the steps: constructing an adaptive time modeling motion scene 3D human body posture estimation network model based on the motion speed degree; the robustness of the model to a motion scene is improved by using motion prior based on physical constraints and predicting the uncertainty of 3D human body articulation points; human body motion videos with different motion speeds are used as input, a corresponding 2D human body joint point sequence is obtained on the basis of an existing high-quality 2D human body posture estimation model, the 2D human body joint point sequence is input into the network model, and a trained model is obtained; according to the method, the receptive field can be adaptively adjusted according to the movement speed, the small receptive field is used for capturing transient changes for fast movement, the large receptive field is used for capturing long-term dependence for slow movement, the method adapts to movement at different speeds, and therefore the human body posture estimation accuracy in movement scenes at different speeds is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

An event camera and vision camera cooperative-based motion training evaluation system and method

The application discloses a motion training evaluation system and method based on cooperation of an event camera and a visual camera, and is applied to the field of computer vision, and aims at the problems that an existing training evaluation system based on an optical camera cannot capture fast motion trajectories and limb shapes, cannot effectively evaluate action quality, and cannot provide effective training guidance; the application provides continuous human motion events in the form of event streams, can achieve millisecond-level motion response, is not affected by motion blur effects of high-speed moving objects, and can provide a higher dynamic range, and can provide more effective motion training evaluation in scenes with strong light, backlight, and sharp changes in brightness. Through cooperation of the visual camera and the event camera, not only the accuracy of fast actions can be evaluated, but also standard action teaching and training in the form of visual scene interaction can be performed.
Owner:CHENGDU UNIV

XR Virtual-Real Interactive Control Device with Optical Motion Capture and Inertial Navigation Fusion

This invention belongs to the field of virtual reality technology, specifically relating to an XR virtual-real interaction control device that integrates optical motion capture and inertial navigation. It includes: a metasurface polarization-encoded marker array, distributed and installed at preset tracking point positions on the user's hand and head, used to generate reflected light carrying polarization identification codes when receiving illumination light; a miniature inertial measurement unit array, rigidly connected to each metasurface polarization-encoded marker in the metasurface polarization-encoded marker array, used to collect acceleration and angular velocity data at each tracking point; and a near-eye display device, including an illumination module and a polarization camera module. The illumination module projects linearly polarized illumination light onto the user's hand and head areas, and the polarization camera module receives the reflected light generated by the metasurface polarization-encoded marker array and outputs multi-channel polarization images. This invention achieves stable tracking in occlusion and rapid motion scenarios through tightly coupled visual-inertial fusion.
Owner:SICHUAN WUTONG TECH CO LTD

Whole body attitude estimation method and system based on downward fisheye and enhanced EgoPoseFormer model

The invention relates to a whole body attitude estimation method and system based on a downward fisheye and an enhanced EgoPoseFormer model, and the method enlarges the human body capture range through the layout of the bottom visual angle of a head-mounted display, compensates geometric nonlinearity in a feature extraction stage through distortion perception convolution, cooperates with an attention mechanism of an embedded pose compensation item, and achieves the estimation of the whole body attitude. And the feature space consistency is maintained when the camera dynamically shakes. Meanwhile, in combination with cross-frame time sequence gating and a multi-modal filtering algorithm, inertial data is utilized to perform kinematics correction on a visual predicted value, and smooth track output is realized in a shielding or rapid motion scene. According to the scheme, the computing power overhead is reduced through model quantification and operator fusion, and low-delay real-time whole body attitude tracking is realized in a mobile terminal chip environment.
Owner:HANGZHOU WUZHI MIXED REALITY TECHNOLOGY CO LTD

Embedded firmware camera data modification methods, devices, media and equipment

This invention discloses an embedded firmware camera data modification method, apparatus, medium, and device, relating to the field of computer vision technology. The method includes: real-time acquisition of image data and preliminary processing; using frame difference and background modeling algorithms to divide the pre-processed image data into regions and assign priorities, generating a region priority mapping table; extracting feature points of high-priority motion regions in the optimized image data using optical flow, and calculating motion vectors between consecutive frames; triggering an event-driven frame processing mechanism based on the calculated motion vectors between consecutive frames, automatically adjusting the frame rate and processing accuracy; and using the motion vectors extracted by optical flow to calculate the object's motion speed in real time, dynamically adjusting the frame rate and processing accuracy according to the motion speed, ensuring sufficient detail is captured during rapid motion while saving processing resources when stationary.
Owner:深圳市云希谷科技有限公司

Video frame optimization processing method based on computer vision

The invention discloses a video frame optimization processing method based on computer vision, and the method comprises the following steps: decoding an input video stream, obtaining a frame sequence, and carrying out the preprocessing of the frame sequence, and obtaining a basic correction frame; constructing an initial noise spore graph, and executing time sequence iteration according to a diffusion and attenuation equation to generate a noise spore evolution field; dividing the evolution field into three types of noise spores, determining a dominant noise spore type according to a competition rule, and forming a competed noise spore field; adjusting enhancement network parameters pixel by pixel to generate a noise self-adaptive enhancement frame; and performing time sequence consistency correction and quality evaluation on the enhanced frame, and outputting a final optimized video frame sequence. According to the method, the noise spore field is constructed, and pixel-by-pixel enhancement and time sequence correction are performed based on the noise type and density, so that the video can obtain a more stable, clearer and cross-frame consistent optimization result in a complex noise and rapid motion scene.
Owner:SHENZHEN JIPAI TECHNOLOGY CO LTD

Event imaging and motion decoupling combined high-speed high-maneuvering three-dimensional trajectory prediction method and system

The invention belongs to the technical field related to trajectory prediction, and discloses a high-speed high-mobility three-dimensional trajectory prediction method and system combining event imaging and motion decoupling. The method comprises the following steps: S1, converting an event stream collected at any target moment into an event stream at a reference moment; s2, calculating a three-dimensional space coordinate corresponding to each target object in the event stream at the reference moment; s3, repeating S2 to obtain three-dimensional space coordinates of each target object in different time windows, namely obtaining a three-dimensional motion track of each target object; and S4, calculating a motion distance and a motion direction of a prediction point by using a plurality of first track points in the three-dimensional motion track, thereby realizing track prediction. According to the invention, the problems of difficult three-dimensional trajectory prediction and low prediction precision caused by high movement speed and random direction of the high-speed and high-maneuvering target are solved.
Owner:HUAZHONG UNIV OF SCI & TECH

Display method for prompting image switching according to image deflection angle

The invention discloses a display method for prompting image switching according to an image deflection angle, and the method comprises the steps: obtaining attitude sensor data, calculating attitude stability reflecting a jitter degree, and constructing an attitude stability state machine for distinguishing a motion state from a stable state according to the attitude stability; in a stable state, dynamically calculating a hysteresis adjustment coefficient by the system according to a current stability numerical value and historical mode preference of a user so as to stretch or translate a switching threshold; and in the motion state, forcibly locking the switching threshold value of the previous frame, and generating a continuous self-adaptive hysteresis threshold value sequence. And judging and switching the display mode in combination with the self-adaptive threshold value and the frame synchronization signal. According to the invention, through the coupling of the state machine gating and the adaptive hysteresis algorithm, the problem of wrong switching in the rapid movement process of the equipment is effectively solved, the adaptive matching of different hand shake degrees and user habits is realized, and the fluency and accuracy of display mode switching are ensured.
Owner:NANJING TUGE HEALTHCARE CO LTD

Virtual character driving method and system based on real-time facial feature point detection and Kalman filtering

The invention discloses a virtual character driving method based on real-time facial feature point detection and Kalman filtering, and the performance of virtual character driving is improved through the four aspects of multi-stage data optimization, high-robustness feature point selection, a lightweight communication protocol and time sequence association optimization. A two-stage filtering framework is combined with Kalman filtering and PD control, real-time denoising and smoothing are carried out on coordinates of feature points, jitter is effectively restrained, stability is improved, and therefore the problem of tracking instability of a traditional method in a fast movement or shielding scene is solved. Then, by selecting high-stability feature points and combining improved dynamic mixed function design, the adaptability of the system to complex scenes is enhanced, meanwhile, feature point detection errors are reduced, and cross-domain generalization performance is improved; then, a lightweight binary communication protocol and an efficient real-time driving architecture are designed, data transmission efficiency and model calculation performance are optimized, millisecond-level interaction requirements are met, the CPU occupancy rate is reduced, and high-frame-rate operation of a mobile terminal is supported.
Owner:SHENZHEN POLYTECHNIC

A fast-moving small target tracking method based on dual-modal fusion

The present invention discloses a dual-modal fusion fast-moving small target tracking method, which solves the problem that fast-moving small targets are easily lost and difficult to locate, and belongs to the field of computer vision. The method comprises: based on the first frame label of each video sequence, cutting out a template area and a search area of ​​the target; performing data enhancement, feature extraction and feature fusion on the visible light image sequence in the search area and the corresponding event voxel grid data set to obtain a fused feature map; outputting a robust feature representation of target positioning through an encoder and a decoder; calculating the target prediction bounding box area and the comprehensive confidence; taking the center point of the target prediction bounding box of the previous frame as a reference and generating a prediction search area of ​​the current frame with a preset multiple thereof; when the comprehensive confidence is higher than an update threshold and the predicted search area contains the target, updating the template area to obtain a constructed fast-moving small target tracking network. The present invention realizes continuous and stable tracking of fast-moving small targets.
Owner:PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV

Multi-person posture fusion and interaction display system based on multi-path image data

The invention relates to a multi-user posture fusion and interaction display system based on multi-path image data, which belongs to the technical field of computer vision and man-machine interaction and comprises a multi-camera acquisition and analysis module, a data coupling module, a real-time transmission module and a three-dimensional reduction and self-adaptive correction module. Wherein the multi-camera acquisition and analysis module calls an attitude estimation engine to obtain three-dimensional coordinates of skeleton key points, and calls the segmentation image matting module to obtain a target mask; the data coupling module is used for associating a user identifier based on the multi-view geometry and time sequence consistency score, and generating a skeleton pose and jitter score; the real-time transmission module executes back pressure transmission according to the queue backlog depth; the three-dimensional reduction and self-adaptive correction module drives a virtual character and shielding rendering in a three-dimensional rendering engine; according to the method, the problem of identity loss under multi-person shielding and rapid movement is solved, and the recognition accuracy is remarkably improved.
Owner:HANGZHOU XINJUE TECH CO LTD

A method, apparatus, device, and storage medium for determining intermediate frames in a video.

This invention discloses a method, apparatus, device, and storage medium for determining intermediate video frames. The method includes: acquiring adjacent video frames of the frame to be interpolated; performing branch feature extraction on the adjacent video frames to obtain video depth features, wherein the branch feature extraction includes at least coarse-grained feature extraction and fine-grained feature extraction; performing optical flow estimation on the video depth features to obtain optical flow motion information; and constructing a target intermediate frame based on the optical flow motion information and the adjacent video frames. This invention's technical solution can better capture motion trajectories in fast-moving small object scenes, ensuring the smoothness of the video intermediate frame interpolation result, solving the error problem when small objects are moving, and significantly improving the interpolation effect.
Owner:SUZHOU GAIDE PHOTOELECTRIC TECH CO LTD

A method for training a system for automated detection of the stroboscopic effect

PCT designated stageWO2026015345A1Image enhancementImage analysisLight activationFrame sequence
A stroboscopic device to use a camera and a light source with software that generates synthetic and real video frames, compares the frames via a discriminator, repeat until a convergenc point is achieved where the discriminator can't reliably tell the real video frames from the synthetic video frames, analyzes a dense optical flow field between a first and second frame in a sequence of video frames to detect a motion, and controls a light activation to achieve a stroboscopic effect, wherein once achieved, a system identifies an object and automatically detects movement or vibration of the idenfied object using a trained model, enabling precise visualization and analysis of fast motions in video sequences.
Owner:IOT TECHNOLOGIES LLC

Fast relocalization method based on sparse semantic anchor points and inertial sensor tight coupling

This invention discloses a fast relocalization method based on tight coupling between sparse semantic anchors and inertial sensing, belonging to the fields of augmented reality and computer vision. It constructs a lightweight sparse semantic map by extracting sparse feature points with stable geometric and semantic attributes from the environment as sparse semantic anchors. When device tracking is unstable or lost, inertial navigation and visual matching threads based on sparse anchors are activated in parallel. The core PnP algorithm is used to quickly recover the visual pose, and the method is tightly coupled with inertial data for optimization, ultimately outputting a high-precision, smooth six-DOF pose. This solves the problem of excessively long relocalization time and experience interruption in traditional visual SLAM scenarios with fast motion and weak textures, achieving millisecond-level, user-unnoticed tracking recovery, significantly improving the robustness and user experience of augmented reality systems.
Owner:CHONGQING AEROSPACE POLYTECHNIC COLLEGE

Event scene text recognition method and system based on chain thinking reasoning

The invention discloses an event scene text recognition method and system based on chain thinking reasoning, and belongs to the technical field of event cameras, and the system comprises an event feature extraction module, a query alignment module, a text generation and thinking chain reasoning module, a joint training module and a chain thinking data generation module. According to the method, the problem that the recognition performance of traditional text recognition based on RGB images is reduced in low-illumination, fast-motion and high-dynamic scenes is solved, interpretability of event stream text recognition is achieved by introducing a chain reasoning mechanism, and the method simultaneously comprises a'text recognition result 'and a'thinking chain reasoning result' and is high in recognition efficiency. Therefore, the transparency and credibility of model decision making are improved. The context reasoning ability of the model is enhanced through the design of alignment of visual and language features, so that the model can still deduce correct text content under the condition that visual information is incomplete or noise exists, and the robustness under the scenes of low illumination, uneven exposure and motion blurring is remarkably improved.
Owner:ANHUI UNIV

Foot contact detection method and system based on multi-scale space-time diagram convolutional network

The invention belongs to the technical field of three-dimensional reconstruction in computer vision, and discloses a foot contact detection method and system based on a multi-scale space-time diagram convolutional network, and the method comprises the steps: collecting video data, and converting each motion sequence into an image sequence; the image sequence is processed, and an input sequence with the target detection frame as the center is intercepted; generating a subsequence based on a time scale; processing the sub-sequence of each time scale, and extracting the features of the high-dimensional spatial-temporal feature map; performing weighted aggregation on the joint features through a spatial attention mechanism, and converting the spatial-temporal feature map into global representation; cross-time-scale information exchange and complementation are realized, and spatial-temporal features are fused; and outputting a four-dimensional contact probability vector, and generating a binary contact label after threshold judgment to obtain a foot contact state. According to the invention, accurate single-view foot contact detection can be realized in both fast movement and slow movement scenes, and the method is applied to various scenes such as movement capture optimization and gait analysis.
Owner:HANGZHOU YILAN TECH CO LTD

Miniature quadruped piezoelectric robot based on resonant and non-resonant fusion driving

The invention discloses a resonant and non-resonant driving fused miniature quadruped piezoelectric robot, relates to the field of multi-legged miniature piezoelectric robots, and solves the problems that the piezoelectric robot is compact in structure and high in speed and resolution. The resonant and non-resonant driving fused miniature quadruped piezoelectric robot comprises driving legs 1, a connecting part 2, a control system 3 and a camera shooting system 4, linear motion of an X axis and a Z axis and rotary motion around the Z axis can be achieved by applying different time sequence signal combinations to different driving legs, the robot has the high-speed motion characteristic during resonant motion, and the robot has the high-speed motion characteristic during resonant motion. The rapid movement around the X axis and the Z axis can be realized; during non-resonant motion, the robot has the high-resolution motion characteristic, nanoscale micro motion along the Z axis can be achieved, the robot is provided with a micro camera, obstacles can be effectively avoided, the surrounding environment can be finely observed, and efficient exploration of the narrow environment where human beings are difficult to enter is achieved.
Owner:NORTHEAST FORESTRY UNIV

Lightweight visual-inertial three-dimensional reconstruction method, system, medium and device for fast motion scene

This invention relates to the fields of computer vision and autonomous robot navigation, and discloses a lightweight vision-inertial 3D reconstruction method, system, medium, and device for fast-moving scenarios. The method includes: simultaneously acquiring continuous image sequences and high-frequency inertial data; obtaining the initial pose prediction value of the camera in the current frame based on the high-frequency inertial data; parameterizing the initial spatial point cloud using compact 3D Gaussian primitives to construct an initial lightweight 3D Gaussian map; extracting the opacity parameter, single scalar scale parameter, and historical observation information of the compact 3D Gaussian primitives within the current effective view frustum to generate a binarized reliability hard mask to constrain residual calculation and gradient backpropagation regions; constructing a multimodal joint tracking loss function to update the current frame camera pose; performing deduplication detection using spatial hashing, instantiating new compact 3D Gaussian primitives, removing redundant or low-contribution primitives, and reclaiming corresponding memory addresses to generate a lightweight 3D Gaussian map with controlled memory usage.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Video decoupling network and method for video target segmentation

The invention provides a video decoupling network and method for video target segmentation, and the network comprises a visual element decoupling module, a unified prior space-time decoupler, and a self-adaptive expert hybrid reconstruction module, and the method comprises the steps: decomposing a historical video frame into a motion representation, a scene representation, and an instance representation through the visual element decoupling module, and constructing a corresponding prior; then single element representation is constructed and analyzed through a unified prior space-time decoupler, finally, multiple visual clues are adaptively integrated through an adaptive expert hybrid reconstruction module, and multiple decoupled object clues are connected to generate global feature representation; multi-source information is fused to reconstruct a high quality object mask. According to the method, the modeling capability of fine-grained motion and target boundaries can be enhanced while the global consistency of the video is kept, so that the accuracy and robustness of video object segmentation are effectively improved, and complex scenes such as occlusion and rapid motion are effectively processed through combination of long-term and short-term reasoning and an attention mechanism; and meanwhile, the stability and reliability of a segmentation result are improved.
Owner:LANZHOU UNIV