Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

65 results about "Motion modeling" patented technology

Multi-target tracking method combining camera motion compensation and pseudo depth estimation

The invention discloses a multi-target tracking method combining camera motion compensation and pseudo depth estimation, belongs to the field of computer vision, and is suitable for a complex automatic driving road environment. The method comprises the following steps: constructing a training set and a test set; detecting the image by using a deep learning detector and extracting features; a Kalman filter is adopted to correct a motion modeling state vector, and the target position and size prediction precision is improved; solving a homography matrix through feature point matching, performing global camera motion compensation, and reducing camera jitter and displacement interference; target pseudo depth information is calculated, hierarchical cascade matching is carried out, and association performance in dense and shielding scenes is optimized; a three-level cascade strategy is adopted to complete high confidence degree, low confidence degree and residual target matching in sequence; and finally, outputting a tracking result with a detection frame and identity information to obtain a trained model. According to the invention, accurate detection and stable tracking of multi-category targets can be realized in a complex environment, and identity switching is effectively reduced.
Owner:CHANGCHUN UNIV OF SCI & TECH

Reference video object segmentation method and system based on motion modeling and multi-modal interaction

The invention discloses a reference video object segmentation method and system based on motion modeling and multi-modal interaction, and the method comprises the steps: taking a video sequence and natural language description as input, and generating a preliminary segmentation mask through a text coding and mask decoder; a Kalman filtering motion modeling module is introduced to predict the motion trail of the target object, and time sequence consistency optimization is carried out on the preliminary segmentation mask; fusing the historical track of the object and the action semantics in the semantic features on the basis of a key action semantic coding module to realize action semantic alignment and mask dynamic correction; the segmentation quality of the current frame is subjected to multi-dimensional scoring based on a representative frame screening mechanism, the representative frame is screened out to update a memory bank, and the long-term tracking stability is improved. According to the method, the problems of target drift, insufficient semantic alignment and memory pollution in a complex dynamic scene in the prior art are effectively solved, and the segmentation precision, robustness and semantic consistency are remarkably improved while the light weight of the model is kept.
Owner:ZHONGKE (SHENZHEN) WIRELESS SEMICON CO LTD

Scene three-dimensional reconstruction and vector information extraction method and device based on vehicle return data, equipment and storage medium

The invention discloses a scene three-dimensional reconstruction and vector information extraction method and device based on vehicle return data, equipment and a storage medium, and the method comprises the steps: carrying out the preprocessing of a time sequence image returned by a vehicle and corresponding pose data, and obtaining a semantic mask and a camera pose of each frame of image in the time sequence image; based on the semantic mask and the camera pose, performing three-dimensional geometric reconstruction on the static scene area in the time sequence image to obtain a static three-dimensional scene model; a dynamic target is separated from the time sequence image according to the semantic mask, three-dimensional motion modeling and independent three-dimensional reconstruction are carried out on the dynamic target, and a dynamic target three-dimensional model is generated; and generating a dynamic three-dimensional scene model and vector labeling information based on the static three-dimensional scene model and the dynamic target three-dimensional model. According to the method, the limitation of automatic driving shadow mode data is overcome, the three-dimensional reconstruction of the scene and the automatic generation of the true value information are realized, the data acquisition cost is further reduced, and the efficiency of automatic driving research and development is improved.
Owner:FOSS (HANGZHOU) INTELLIGENT TECH CO LTD

Multi-target tracking method combining time convolutional network and self-attention mechanism

The invention relates to the technical field of computer vision, in particular to a multi-target tracking method combining a time convolution network and a self-attention mechanism, which comprises a motion model based on a TCN, dynamic ReID feature updating based on track confidence, and a rematching strategy used for correcting wrong association caused by target interaction shielding. A TCN network is combined with an attention mechanism to be used for motion modeling, and nonlinear motion prediction of a track is achieved. In addition, a track confidence coefficient is provided to measure the robustness of the track, and the ReID features are dynamically updated based on the track confidence coefficient. Meanwhile, in order to make up for wrong association caused by target interaction shielding, a rematching module is designed, and the effectiveness of the rematching module is proved through experiments. In addition, each module of the system has a modular characteristic and a relatively strong generalization ability, so that the system is suitable for a wider real scene, and future research work is possibly stimulated.
Owner:NANJING UNIV OF POSTS & TELECOMM

Multi-sensor cooperative robot target positioning system

The invention discloses a multi-sensor cooperative robot target positioning system, and relates to the technical field of robot positioning, in particular to the multi-sensor cooperative robot target positioning system. The objective of the invention is to solve the problems of reduced positioning precision and insufficient robustness caused by sensor data failure, insufficient cooperation mechanism and motion interference in an extreme environment. The system comprises a multi-source sensing module, a collaborative fusion module, a motion modeling module and a closed-loop control module, a robot pose adjustment instruction is generated through multi-sensor data acquisition, space-time alignment and confidence weighted fusion, dynamic motion modeling and error compensation, and closed-loop collaboration of positioning and motion is achieved. According to the invention, the positioning precision, robustness and continuous operation capability of the system in an extreme environment are effectively improved.
Owner:RATE OF CHANGE CHANGSHA INFORMATION TECH CO LTD

Method and system for calculating influence of dynamic response of separation support device and buffer coupling device on ship pose in floating support installation process

The invention relates to the technical field of ocean engineering operation motion modeling and control, and particularly discloses a method and a system for calculating influence of dynamic response of a separation supporting device and a buffer coupling device on a ship pose in a floating support installation process. The method comprises the following steps: acquiring initial position information of a separation support device and a buffer coupling device at a previous moment; acquiring vertical position and attitude information of the dynamic positioning ship and the upper module; calculating the supporting force of the separation supporting device and the buffer coupling device at the current moment; calculating the vertical position and attitude information of the upper module and the dynamic positioning ship at the current moment based on the supporting force and the external ballast water load; outputting relevant parameters and judging whether the installation is completed or not. According to the method, a multi-degree-of-freedom coupling dynamic model is constructed, high-precision real-time simulation of device dynamic response and ship motion attitude in the continuous unsteady-state load transfer process is realized, key parameter monitoring and risk early warning are provided for floating mounting operation, and the operation safety and efficiency are improved.
Owner:HARBIN ENG UNIV

Three-dimensional target tracking method, system and device based on space-time enhancement and medium

The invention discloses a three-dimensional target tracking method, system and device based on space-time enhancement and a medium, and relates to the technical field of automatic driving. The method comprises the following steps: inputting three-dimensional point cloud sequence data into a trained SMTrack network for processing to obtain a final bounding box prediction result. Wherein the SMTrack network comprises a target specific encoder, an STFM module and an STT module which are connected in sequence; wherein the target specific encoder is used for extracting target specific features from a point cloud sequence; the STFM module is used for modeling appearance information and motion information in stages according to the specific features of the target, and generating a preliminary bounding box prediction result; the STT module is used for optimizing the preliminary bounding box prediction result; according to the method, a twin structure is adopted, features are extracted from a historical frame sequence and a current frame sequence, motion modeling of part sensing and a coarse-to-fine bounding box regression mechanism are introduced, and the target positioning precision can be greatly improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Intelligent ship digital twin motion modeling and parallel anti-interference path tracking controller

The invention provides an intelligent ship digital twinning motion modeling and parallel anti-interference path tracking controller. The intelligent ship digital twinning motion modeling and parallel anti-interference path tracking controller comprises a guidance signal module, a speed optimization module, a parallel path tracking controller module, a data-driven learning predictor module and an actual controller module. The data-driven learning predictor module receives position information, course angle information and speed information from the actual intelligent ship system module, and the data-driven learning predictor module receives parallel path tracking control force and torque from the parallel path tracking controller module; the data driving learning predictor module sends estimation information of the position and course angle of a virtual intelligent ship and estimation information of the speed of the virtual intelligent ship to the guidance signal module and the parallel path tracking controller module. According to the method, the self-adaptive capacity of the ship under the complex sea condition is enhanced, the collision risk is reduced fundamentally, and an effective solution is provided for the navigation scene with the high safety requirement.
Owner:DALIAN MARITIME UNIVERSITY

A reference video object segmentation method and system based on motion modeling and multi-modal interaction

The application discloses a reference video object segmentation method and system based on motion modeling and multi-modal interaction, which comprises the following steps: taking a video sequence and a natural language description as input, generating a preliminary segmentation mask through a text encoding and mask decoding module; introducing a Kalman filter motion modeling module to predict the motion trajectory of the target object and optimize the time sequence consistency of the preliminary segmentation mask; based on a key action semantic coding module, fusing the action semantics in the historical trajectory and semantic features of the object, realizing action semantic alignment and dynamic mask correction; based on a representative frame screening mechanism, multi-dimensional scoring is performed on the segmentation quality of the current frame, and representative frames are screened out to update the memory bank, thereby improving the long-term tracking stability. The application effectively solves the problems of target drift, insufficient semantic alignment and memory pollution in the prior art in a complex dynamic scene, while keeping the model lightweight, significantly improves the segmentation accuracy, robustness and semantic consistency.
Owner:ZHONGKE (SHENZHEN) WIRELESS SEMICON CO LTD

Multi-sensor fusion evaluation algorithm and device based on continuous time series

According to the multi-sensor fusion evaluation algorithm and device based on the continuous time sequence, modeling is carried out based on continuous time, physical reality is better met, motion modeling with better precision is achieved, intra-frame motion can be accurately described, smoother and more accurate track estimation can be provided, and the method and the device are suitable for being applied to multi-sensor fusion evaluation. Especially, the advantages are obvious in high-speed and high-frequency vibration scenes; meanwhile, according to the technical scheme disclosed by the invention, when LiDAR / Camera matching fails, the constraints of the IMU and the GNSS are seamlessly and continuously applied to the whole track through a continuous time model instead of only acting on a discrete frame, so that error accumulation is greatly inhibited, system collapse is avoided, the robustness of a degraded scene is enhanced, application in specific fields such as surveying and mapping is greatly facilitated, and the method is suitable for popularization and application. And the method has the capability of flexibly adding constraints.
Owner:BEIJING GREEN VALLEY TECH CO LTD +4

Human body surface dynamic reconstruction method based on monocular video

PendingCN121921446AAccurately restore garment wrinklesAccurate recovery of muscle movementsImage enhancementImage analysisHuman bodyMorphing
The invention discloses a human body surface dynamic reconstruction method based on a monocular video, and the method comprises the steps: extracting a key frame of human body motion from the monocular video, and obtaining an RGB image of the key frame, a mask, parameters of an SMPL human body template, and internal and external parameters of a camera; extracting a real normal vector diagram corresponding to the key frame, and constructing a basic data set; the basic data set is used for training a multi-stage progressive human body reconstruction framework, the framework comprises human body overall motion modeling serving as a first stage, local surface dynamic deformation modeling serving as a second stage and surface appearance and illumination modeling serving as a third stage, and an optimal reconstruction model is obtained through training; and extracting a new view angle frame of human body motion from the monocular video, obtaining an RGB image of the new view angle frame, parameters of the SMPL human body template and internal and external parameters of the camera, and inputting the RGB image, the parameters of the SMPL human body template and the internal and external parameters of the camera into the optimal reconstruction model to generate a human body dynamic reconstruction result under the conditions of a new view angle, a new posture and new illumination, thereby realizing high-quality rendering output with geometric details and appearance consistency.
Owner:SOUTH CHINA UNIV OF TECH

Unsupervised monocular image-based online 3D scene reconstruction method and device

The present disclosure provides a monocular image-based unsupervised three-dimensional scene online reconstruction method and device. The method of the present disclosure comprises: obtaining a static background mask and a mask of each dynamic object instance of a current frame monocular image through semantic segmentation, obtaining a surround view synthesis image through a static multi-view generator, determining the motion parameters of each dynamic object instance through motion modeling, obtaining a multi-view depth map based on the surround view synthesis image through a depth estimation network obtained through self-supervised training, obtaining a local depth map of each dynamic object instance through a local depth estimation network obtained through self-supervised training, and obtaining a complete 3D scene representation through the multi-view depth map, the local depth map of each dynamic object instance, and the motion parameters thereof. The present disclosure can avoid true value dependence, effectively reduce hardware cost, and at the same time improve the reliability and robustness of monocular image three-dimensional scene reconstruction.
Owner:BEIJING TRUNK TECHNOLOGY CO LTD

Coding method and device in AIGC generation

The application provides a coding method and device in AIGC generation. The decoding method in AIGC generation of the application comprises the following steps: obtaining first feature information and motion information of a current frame based on generated feature information output by an AIGC generation model, wherein the motion information is used to represent inter-frame correlation; performing motion modeling according to the first feature information and the motion information of the current frame to obtain second feature information of the current frame, wherein the motion modeling is used to compensate the first feature information of the current frame based on the motion information; and obtaining reconstructed content of the current frame according to the second feature information of the current frame. The application can improve the time domain compression efficiency and the quality of the reconstructed content in AIGC.
Owner:HUAWEI TECH CO LTD

Crowd motion modeling method and system based on large language model

PendingCN122290052ALinguistic modelData set
This invention relates to the field of traffic simulation and pedestrian flow modeling technology, and particularly to a method and system for crowd movement modeling based on a large language model. The method includes: constructing a multi-dimensional interactive feature system of pedestrians, small groups, and the environment; extracting micro-behavioral features of small pedestrian groups through micro-unit segmentation and dynamic group partitioning methods; proposing a pedestrian movement decision-making framework based on a large language model and thought chain; designing a hybrid update decision-making strategy to balance high-level semantic decision-making and local obstacle avoidance; and training and testing pedestrian evacuation at traffic hubs. Instance verification of pedestrian small group movement decision-making at traffic hubs is also conducted. The proposed model is applied to hub scenarios and its empirical dataset for training and testing. Results show that it achieves better motion realism, traffic efficiency, and group structure stability than baseline models.
Owner:BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Object-oriented periodic dynamic motion 4D Gaussian splash reconstruction method

The invention discloses an object-oriented periodic dynamic motion 4D Gaussian splash reconstruction method, and relates to the technical field of computer vision and graphics. The method comprises the following steps: identifying and tracking a dynamic object in a foreground, generating a 3D mask and extracting sparse point clouds, performing time sequence alignment to generate a 4D point cloud sequence, identifying periodic motion and extracting global trajectory codes; initializing a standard state 3D Gaussian, constructing a periodic deformation field for a periodic object, generating a dynamic object Gaussian in combination with a global trajectory, coding a non-periodic object only by using the global trajectory, and using the standard state Gaussian for a background; and a reconstruction image is obtained through rendering, reconstruction loss is calculated and back propagation optimization is carried out on normative state 3D Gaussian, a periodic deformation field and global trajectory coding, adaptive density control is executed at an object level in the process, and 4D Gaussian splash reconstruction is completed. According to the method, high-fidelity, editable and efficient-storage 4D dynamic scene reconstruction is realized, and the representation compactness, semantic controllability and motion modeling precision are improved.
Owner:SICHUAN VOCATIONAL & TECHN COLLEGE OF COMM

Method and device for modeling disturbed motion of underwater vehicle based on hydrodynamic similarity guidance

PendingCN122389695AFeature setMotion prediction
The application discloses a method and device for modeling disturbed motion of a submersible based on water dynamic similarity guidance, and belongs to the technical field of motion control of submersibles. The method comprises the following steps: determining a similarity parameter representing the degree of water dynamic essential similarity between a source submersible and a target submersible according to the distribution correlation of a source dimensionless feature set and a target dimensionless feature set in a dimensionless feature space; inputting the target dimensionless feature set into a teacher network with fixed parameters to obtain intermediate layer features and interference parameter representations output by the teacher network; inputting the intermediate layer features, the interference parameter representations and the similarity parameter into a similarity attention mechanism module to generate a knowledge distillation signal; and inputting the target dimensionless feature set into a student network comprising a lightweight adaptive layer to train the student network so as to deploy a disturbed motion prediction model of the target submersible. The method can improve the accuracy of motion prediction of the submersible in a disturbance environment such as a strong current and a turbulent flow.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

3D multi-target tracking method and system based on monocular vision

The invention provides a 3D multi-target tracking method and system based on monocular vision, and the method comprises the steps: obtaining a continuous video image sequence collected by a monocular camera, and carrying out the preprocessing of the continuous video image sequence, and obtaining an input image sequence; performing target detection and two-stage depth estimation and fusion on the input image sequence to obtain a target detection list with fusion depth information; constructing a multi-dimensional correlation cost matrix, and performing optimal correlation on the current detection and the existing tracking trajectory based on the target detection list and the multi-dimensional correlation cost matrix; and updating the tracking trajectory state according to the association result, and carrying out life cycle management on the tracking trajectory state. According to the monocular camera depth detection technology provided by the invention, low cost and easy integration are ensured, and the depth dimension information is injected for multi-target tracking, so that the defects of a traditional algorithm in shielding processing, similar target distinguishing and three-dimensional motion modeling can be effectively overcome, and the monocular camera depth detection technology is an optimal choice for balancing performance and cost.
Owner:CHINA NAT BUILDING MATERIALS TECH CO LTD +4

4D Gaussian structuring-based lunar surface / deep space scene target reconstruction and editing method

The invention relates to the technical field of three-dimensional graphic processing, and discloses a lunar surface / deep space scene target reconstruction and editing method based on 4D Gaussian structuring. The method is particularly suitable for dynamic target modeling and editing in deep space exploration tasks, including detectors, robots and related equipment on the lunar surface, the Mars or other celestial bodies. The method comprises the following steps: synthesizing a plurality of camera visual angles through a diffusion model based on a monocular video of an object, and reconstructing static three-dimensional Gaussian point representation in a standard space; extracting a skeleton structure of the object based on the surface grid, wherein the skeleton structure comprises a skeleton node position and a topological relation; on the basis of the skeleton structure, each Gaussian point is connected with a plurality of skeleton nodes through a linear hybrid skin mechanism, and skeleton-driven rigid deformation is achieved; non-rigid deformation is compensated through feature extraction and a regression network; and in combination with the rigid deformation and the non-rigid deformation, rendering to generate a dynamic Gaussian point model which can change along with time and can be edited. According to the method, motion modeling of an object is explicitly split into rigid motion driven by a framework and non-rigid correction, so that the definition and interpretability of motion representation are greatly improved, and particularly, higher stability and control precision are shown when challenges such as weak texture, violent illumination change and high noise are faced in a deep space environment. According to the method, object behavior modeling in a detection task is more flexible, and the method can be widely applied to scenes such as task planning, target recognition and dynamic editing.
Owner:UNIV OF SCI & TECH OF CHINA

Slow action generation method and device based on motion focus area analysis and medium

The invention provides a slow action generation method based on motion focus area analysis. The method comprises the following steps: determining a video of a slow action to be generated; detecting a motion focus area in the video, and grading the motion focus area; and for different grades of motion focus areas, different strategies are adopted to carry out motion modeling and differential interpolation processing, and a slow motion frame sequence is generated. According to the invention, through identification and grading processing of the motion focus area, high-precision calculation of the core motion object, simplified calculation of the secondary area and simplest processing of the static background, compared with a traditional scheme of uniform and dense calculation of the whole frame, the method has the advantages of high calculation complexity and lower calculation amount, and guarantees the visual effect of the core area.
Owner:CHENGDU SOBEY DIGITAL TECH CO LTD

Motion deblurring method and system based on taylor expansion and 3d gaussian sputtering

This invention proposes a motion deblurring method and system based on Taylor expansion and 3D Gaussian sputtering. The method includes: extracting the dynamic scene from a multi-view motion-blurred video sequence and parameterizing it into an optimizable time-varying 3D Gaussian sputtering representation; constructing a motion model based on Taylor expansion; obtaining an initial velocity field by differentiating and projecting the Taylor expansion motion model, and then correcting it to obtain a corrected 2D pixel-level velocity field; synthesizing a synthetic physically blurred image consistent with the input video frame by simulating the physical imaging process during camera exposure; finally, training and optimization are performed, and after optimization, a final 3D Gaussian scene representation is obtained. This final 3D Gaussian scene representation is then used to render a clear image, thus achieving motion deblurring. This invention improves the physical interpretability and accuracy of motion modeling by explicitly modeling continuous motion trajectories using Taylor expansion.
Owner:JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Inspection robot autonomous navigation system based on deep learning

The invention relates to the technical field of inspection robots, and particularly discloses an inspection robot autonomous navigation system based on deep learning, which comprises an environment sensing module, a dynamic map construction module, a path planning module, an obstacle avoidance control module and an execution feedback module. According to the scheme, a convolutional neural network and an attention mechanism are introduced, deep feature extraction and dynamic target processing are performed on multi-modal data of an inspection environment, and high-precision global dynamic map construction is realized; equipment importance and abnormal position detection are introduced to form a path priority, an important target is preferentially accessed, and the dynamic adaptability of an inspection path is improved according to a global dynamic map and a path risk value updated in real time, so that the inspection robot can efficiently complete a task in a complex dynamic environment; real-time obstacle avoidance and local path re-planning in a dynamic environment are realized, the future collision time is predicted by combining relative motion modeling with a kinematic model, the method has strong local path adaptability, and the stability of robot inspection control is improved.
Owner:BAIYIN YINZHU ELECTRIC POWER GRP CO LTD

A sequence image fusion enhancement method based on three-axis slide table displacement sensing

This invention discloses a sequential image fusion enhancement method based on three-axis slide table displacement sensing, belonging to the field of machine vision technology. The method includes the following steps: S1, constructing discretized state and output equations to complete the three-axis slide table motion modeling, and establishing the correlation between motion parameters and image blur kernel and pixel position offset; S2, using the inverse function of the camera response function combined with Taylor linearization, performing exposure normalization processing on multiple frames of observed images to obtain scene illumination estimates with a unified radiance scale; S3, obtaining multi-dimensional fusion weights for the image by setting a fusion weight calculation mechanism that correlates the quality features of the fused image with the slide table position; S4, completing the fusion of multiple frames of images through a weighted fusion model to obtain a high-quality single-frame fused image. Using the above method, the problems of image motion blur and uneven multi-frame imaging quality caused by three-axis slide table motion are effectively solved, achieving accurate fusion and deblurring enhancement of sequential images.
Owner:UNIV OF SCI & TECH BEIJING

Vegetation coverage area soil salt estimation method, electronic device, and program product

This application provides a method, electronic device, and program product for estimating soil salinity in vegetated areas. By introducing two synergistic constraint mechanisms, it effectively overcomes the problem of unstable unmixing in traditional NMF (Natural Motion Modeling) in high vegetation cover (FVC>53%) scenarios. **Unique Spectral Constraint on Vegetation Endmembers:** Based on a pre-constructed vegetation calibration spectral library, similarity constraints are applied only to the spectra of vegetation endmembers, ensuring stable extraction of vegetation spectral features while cleverly avoiding excessive restriction on the decomposition space of soil endmembers, thus fully preserving key spectral information of soil salinity. **Smoothing Spatial Constraint on Abundance Matrix:** Utilizing the spatial correlation of hyperspectral imagery, a smoothing constraint is applied to the abundance of adjacent pixels through a Laplacian matrix regularization term, significantly reducing error accumulation and enhancing the spatial consistency and overall stability of the unmixing results.
Owner:SHANDONG NORMAL UNIV

An interactive multi-model based composite seeker jamming track recognition method

The application discloses a composite seeker jamming track recognition method based on an interactive multi-model, which comprises the following steps: firstly, initializing system parameters, assuming the motion parameters of a target and jamming; secondly, determining a state model used for describing the motion form of the target, and then modeling a multi-model system; thirdly, processing the state model to obtain updated model probability and system state estimation; fourthly, calculating a jamming decision threshold according to the updated model probability, so as to determine the appearing jamming track; fifthly, further identifying the specific jamming; and finally, outputting a jamming track serial number. The application introduces the motion modeling of jamming into the composite seeker, effectively improves the jamming recognition probability of the track gradual change, does not need to introduce an extra high-dimensional feature calculation, removes the jamming track in the track level dimension, and has the advantages of small engineering implementation cost, easy realization and the like.
Owner:JIANGXI HONGDU AVIATION IND GRP

Multi-modal time sequence fusion 3D target detection method based on instantiation sparse representation

The invention provides a multi-modal time sequence fusion 3D target detection method based on instantiation sparse representation. The method comprises the following steps of: 1, extracting each modal feature through an independent backbone network and generating a complementary enhanced instance to initialize query; step 2, adopting semi-explicit-implicit mixed motion modeling to align historical instances; step 3, performing time sequence perception enhancement; and a fourth step of adaptively fusing cross-modal and cross-space-time instance information through a lightweight cross attention mechanism to generate a 3D detection result. The method is high in detection accuracy, low in false detection rate and short in single-frame processing time, and the robustness and efficiency of multi-modal time sequence fusion 3D detection are remarkably improved.
Owner:LIAONING TECHNICAL UNIVERSITY

Unmanned aerial vehicle multi-target tracking method and device based on trajectory occlusion synthesis and hierarchical multi-cue association

This invention discloses a method and apparatus for multi-target tracking of unmanned aerial vehicles (UAVs) based on trajectory occlusion synthesis and hierarchical multi-cue association. The method includes: based on a trajectory occlusion synthesis module (TOSM), actively mining spatiotemporally similar historical trajectory pairs and generating long-term synthetic occlusion samples, and combining a dual-branch training architecture to enhance the model's motion modeling ability for occlusion while ensuring the stability of the feature space; based on a hierarchical multi-cue association module (HMAM), employing a four-level cascaded matching strategy, adaptively fusing motion prediction, deep re-identification features, color histograms, and template matching features for detection results with different confidence levels. This invention effectively solves the problems of identity switching caused by long-term occlusion and low-confidence target tracking loss caused by severe motion blur in complex dynamic scenes of UAVs by actively introducing occlusion enhancement during the training phase to learn robust representations and designing a hierarchical mechanism to flexibly call multiple complementary cues during the inference phase.
Owner:SOUTH CHINA AGRICULTURAL UNIVERSITY

Real-time video anti-shake method based on adaptive motion modeling and artificial intelligence prediction

The present application relates to the field of image processing and computer vision, and particularly relates to a real-time video anti-shake method based on adaptive motion modeling and artificial intelligence prediction, comprising: establishing a grid coordinate system mapping relationship; extracting feature points from a video frame and calculating preliminary displacement of grid vertices; constructing a sliding time window, dynamically adjusting and optimizing parameters according to scene characteristics, establishing a time domain smooth optimization objective function, and introducing a frequency prediction constraint when periodic motion is detected; using a parallel iterative algorithm to solve the objective function to obtain steady-state displacement; and generating a stable frame based on the steady-state displacement. The present application realizes real-time processing of a video on a general CPU through a pre-computation caching strategy and vectorized parallel solving; and realizes adaptive prediction compensation of periodic motion and intelligent identification of a scene through FFT frequency detection and a lightweight machine learning model.
Owner:SICHUAN JIUZHOU SOFTWARE CO LTD +1

System and method for multi-frame video frame interpolation

Systems and methods for multi-frame video frame interpolation are provided. High order motion modeling, such as third order motion modeling, enables prediction of intermediate optical flow between multiple interpolated frames by relaxing constraints imposed by loss functions used in initial optical flow estimation. A temporal pyramid optical flow refinement module coarsely-to-finely refines optical flow maps used to generate intermediate frames, thereby concentrating proportionally more refinement attention on optical flow maps of high error intermediate frames. A temporal pyramid pixel refinement module coarsely-to-finely refines generated intermediate frames, thereby concentrating proportionally more refinement attention on the high error intermediate frames. A generative adversarial network (GAN) module computes loss functions used to train neural networks used in the optical flow estimation module, the temporal pyramid optical flow refinement module, and / or the temporal pyramid pixel refinement module.
Owner:HUAWEI TECH CO LTD

Music-driven dance generation method and device based on rhythm perception feature representation of gating enhancement

PendingCN121415809ASpeech analysisBiological modelsFeature codingRhythm perception
The invention discloses a music-driven dance generation method and equipment based on gating enhanced rhythm perception feature representation, and belongs to the technical field of music-driven dance generation. In order to solve the problem that dance generated by an existing generation method is poor in expressive force and rhythm coherence, the method comprises the following steps of: encoding motion of an upper body and motion of a lower body by using VQ-VAE to obtain an upper body code and a lower body code; performing channel dimension alignment on the music features, the upper body codes and the lower body codes by using independent linear layers to obtain respective corresponding embedded features, splicing the embedded features, obtaining rhythm enhancement features in combination with phase features representing music rhythms, and processing the rhythm enhancement features by using a time gating causal attention module; and generating respective probabilities of each action in an upper body code book and a lower body code book in a current prediction time frame by combining a parallel Mama motion modeling module, respectively selecting feature codes of the upper body action and the lower body action with the maximum probabilities in the code books, and decoding and outputting the actions by utilizing a VQ-VAE decoder.
Owner:HARBIN ENG UNIV

Intelligent body measurement method based on dynamic capture and motion modeling

The invention relates to the field of clothing body measurement, in particular to an intelligent body measurement method based on dynamic capture and motion modeling. The method comprises the following specific steps: step 1, capturing dynamic postures and extracting key points: shooting in-situ stepping actions of a user at different angles through at least two RGB-D cameras, and capturing data of the key points once every four frames from in-situ stepping video dynamic postures; step 2, three-dimensional coordinate conversion and size calculation: converting the key points of the 2D image into three-dimensional space coordinates based on the focal length, the principal point coordinates and the external parameters of the camera; 3, a space error compensation algorithm is carried out, wherein the three-dimensional space coordinates obtained in the step 2 are compensated through space error compensation, and errors generated by different distances between the person and the camera are eliminated; step 4, generating a curve: gathering all positions of part of key points in the three-dimensional space coordinates after the space error compensation algorithm to generate the curve; and 5, extracting the following data, and arranging the data into a table.
Owner:SHAOXING BOYA FASHION CO LTD