Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

46543 results about "Computer vision" patented technology

Computer vision is an interdisciplinary scientific field that deals with how computers can be made to gain high-level understanding from digital images or videos. From the perspective of engineering, it seeks to automate tasks that the human visual system can do.

Multi-modal medical image data intelligent processing system

The invention discloses a multi-modal medical image data intelligent processing system, relates to the field of medical image analysis, and is applied to multi-modal medical image whole-process analysis of CT, MRI, PET, ultrasound and the like. According to the system, different modal image features are extracted and fused through a cross-modal manifold fusion network; a semantic guidance dynamic registration engine optimizes registration parameters to ensure that the registration error is less than or equal to 1.5 mm; the multi-task collaborative diagnosis network realizes multiple tasks such as disease classification; the clinical knowledge embedding and interpretable module generates a structured report and is in butt joint with an HIS system. Meanwhile, the model is optimized through a federated learning architecture, the adaptability of newly added data is improved by more than or equal to 20%, and intelligent processing and analysis of multi-modal medical images are realized.
Owner:SHANDONG JUNKANGLIN MEDICAL TECHNOLOGY CO LTD

System and method for industrial risk assessment via computer vision

A device, system and method comprising computer vision techniques for fire prevention / detection and risk assessment, as well as for determining deviations from an ideal operational state. The present invention includes for example systems and methods which leverage data collected by camera systems composed of infrared and visible light sensors to detect and / or prevent a fire from starting, and additionally, use this data to determine a risk assessment for the building. The present invention also provides for example a system and method for monitoring and controlling safety risks in indoor industrial environments by determining deviations from an ideal operational state using computer vision techniques and game-theoretic competitive ranking frameworks.
Owner:INNOVIRE AG

Weak supervision target detection method guided by cross-modal pseudo tag

The invention relates to the technical field of computer vision and multi-modal learning, in particular to a weak supervision target detection method guided by cross-modal pseudo labels. According to the method, a labeled source domain data set is constructed to train an image classification teacher model, and a teacher-student network structure is constructed; clustering the regional features of the target domain image, allocating pseudo tags to each cluster by optimizing the allocation cost between the source domain category and the target domain cluster, and constructing a pseudo tag pool; and training a student model on the pseudo label pool for region feature detection of the target domain image. According to the method, a cross-modal attention mechanism is introduced, so that more accurate semantic alignment between a source category label and a target domain feature is realized; the stability of label distribution is improved by a structure keeping regular term; the generalization ability of the model is further enhanced by multiple rounds of pseudo-label confidence learning. The method can be widely applied to tasks such as target detection, cross-domain transfer learning and open world recognition, and efficient and accurate weak supervision target detection is realized.
Owner:DATA SPACE RES INST

Temporal bone disease classification method and system based on multi-modal medical image fusion technology

The invention relates to the field of image analysis, in particular to a temporal bone disease classification method and system based on a multi-modal medical image fusion technology. The method comprises the following steps: acquiring a multi-modal image of a patient, performing adaptive distortion correction, and generating a standardized image set; performing layer-by-layer anatomical structure semantic segmentation and multi-modal image fusion on the standardized image set to construct an image fusion framework; according to the image fusion framework, performing intelligent recognition on the fine structure of the temporal bone, and constructing a personalized temporal bone anatomical structure chart; performing tissue function state analysis and digital pathology dynamic simulation based on the personalized temporal bone anatomical structure chart, and constructing a digital pathology model; and performing intelligent pathological feature classification based on the digital pathological model to obtain an intelligent classification report. According to the method, rapid, efficient and accurate temporal bone disease classification is realized.
Owner:EYE & ENT HOSPITAL SHANGHAI MEDICAL SCHOOL FUDAN UNIV

Optical remote sensing image salient target detection method based on progressive attention enhancement

The invention discloses an optical remote sensing image salient target detection method based on progressive attention enhancement, and belongs to the technical field of computer vision. The method comprises the following steps: preprocessing an original data set; inputting the preprocessed image into a hierarchical progressive fusion encoder, capturing a global irregular topological structure and local fine-grained image details, and realizing cross-hierarchical feature fusion; inputting the output characteristics of the encoder into a global context enhancement module, and capturing multi-level context information by adopting a parallel multi-branch structure; and inputting the output features of the hierarchical progressive fusion encoder and the global context enhancement module into a multi-scale progressive attention enhancement decoder, carrying out hierarchical decoding on the input features by adopting a saliency-guided attention mechanism, and gradually aggregating deep semantic information and shallow detail features to realize coarse-to-fine progressive optimization, so as to improve the robustness of the multi-scale progressive attention enhancement decoder. And finally generating a saliency map. The method can effectively improve the processing performance of an irregular topological structure and a complex context relationship in the optical remote sensing image.
Owner:SHIJIAZHUANG TIEDAO UNIV

Unmanned aerial vehicle positioning system and method based on multi-source position signal fusion

The invention discloses an unmanned aerial vehicle positioning system and method based on multi-source position signal fusion, and relates to the technical field of unmanned aerial vehicle positioning, and the system comprises a plurality of positioning data collection units which are carried on an unmanned aerial vehicle, and each type is specially used for collecting positioning data of a single source; the positioning data processing unit is suitable for calculating a confidence coefficient value based on the attribute parameter of each positioning data, and determining a weight of each positioning data according to the confidence coefficient value; and the positioning data fusion unit is suitable for fusing all the positioning data by applying a fusion algorithm and combining the weights of the positioning data to generate fused positioning data. According to the system, the confidence of each source positioning data is evaluated in real time, and the fusion weight is dynamically optimized according to the confidence, so that the dynamic evaluation and optimal fusion of the multi-source positioning data are realized, and the positioning precision and reliability of the unmanned aerial vehicle in a complex environment are improved.
Owner:BEIJING ZHIWANG YILIAN TECH CO LTD

3D gausians splatting in scene description

Some embodiments of a method may include: obtaining information for a three-dimensional (3D) Gaussian model corresponding to a 3D scene, wherein the information comprises a set of attributes of the 3D Gaussian model; parsing the information for a first attribute of the set of attributes, wherein the first attribute corresponds to a position of the 3D Gaussian model; parsing the information for a second attribute of the set of attributes, wherein the second attribute corresponds to a covariance of the 3D Gaussian model; parsing the information for third, fourth, and fifth attributes, wherein the third, fourth, and fifth attributes correspond to first, second, and third sets of spherical harmonics coefficients associated with the 3D Gaussian model; parsing the information for a sixth attribute of the set of attributes, wherein the sixth attribute is an alpha coefficient for the 3D Gaussian model; and rendering the 3D scene using the parsed attributes.
Owner:INTERDIGITAL CE PATENT HOLDINGS SAS

AI-based animation sub-mirror script automatic generation and visual preview method and system

The invention discloses an AI-based animation split script automatic generation and visual preview method and system, and the method comprises the following steps: 1, receiving a natural language script text inputted by a user, the natural language script text comprising scene description, role action, dialogue and shot indication information; step 2, performing semantic analysis and structured analysis on the script text based on a natural language processing technology, and identifying and extracting key narrative elements; by introducing an artificial intelligence technology, end-to-end automatic generation and interactive optimization from a character script to a dynamic split rehearsal video are realized, the system can deeply understand scenes, actions, role emotions and shot languages in the script, corresponding visual elements are automatically matched and generated, and the dynamic split rehearsal effect is improved. And the timeline and the rhythm conforming to the film and television grammar are constructed, so that the efficiency and the consistency of the split creation are greatly improved, and the professional threshold and the manufacturing cost are reduced.
Owner:NEW AXIS ANIMATION TECHNOLOGY DEVELOPMENT (BEIJING) CO LTD

Multi-mode-based training method and system for cervical pathology image classification model

The invention relates to the technical field of image classification, in particular to a training method and system of a cervical pathological image classification model based on multiple modes. The method comprises the following steps: acquiring a cervical tissue image and carrying out tissue structure segmentation, forming a nucleus-interstitial-epithelium three-distribution framework, collecting development historical data, confirming a prediction trend of each layer, carrying out environment field simulation through the image, generating a simulated cervical environment field, and carrying out hierarchical evolution prediction on the framework. Evolution mapping images are generated according to the evolution data and classified, finally, a visual basic model is obtained through combined modeling training, image-text fusion is achieved, and a cross-center deployment model system is generated. According to the method, vision-language combined modeling is realized, and the stability and controllability of the whole model structure in image space deformation modeling, semantic cross-modal alignment construction and task-level response flow scheduling are improved.
Owner:GUANGZHOU JINRUI TECHNOLOGY CO LTD

Generative ai models for image rendering and inverse rendering

Embodiments of the present disclosure relate to rendering and inverse rendering using one or more generative models. “Rendering” refers to the process of generating a final visual image, video frame, or animation from a 2D or 3D model. “Inverse rendering” is a process that involves deducing or estimating the properties (e.g., material maps or other properties such as geometry, lighting, and textures) of a scene from observed images or visual data. Essentially, it aims to reverse the traditional rendering process. Various aspects of the present disclosure introduce editable light and material controls into generative models to allow for artistic creation. Various embodiments integrate generative models as a renderer for classic rendering pipelines to upcycle and enhance the style of rendered content.
Owner:NVIDIA CORP

System and method of three-dimensional object cleanup and text annotation

Some examples of the disclosure are directed to object manipulators and associated processes for manipulating an object representation in a three-dimensional environment. The object representation may correspond to a scan of a real-world object in a real-world environment. The object manipulators may include an object cleanup manipulator and a text annotation manipulator. The object cleanup manipulator may be selectable to display one or more control affordances providing functionality for selectively removing portions of the object representation in the three-dimensional environment and / or selectively adjusting one or more parameters of the object representation in the three-dimensional environment. The text annotation manipulator may be selectable to display one or more control affordances providing functionality for selectively generating one or more text labels in the three-dimensional environment. The one or more text labels may be associated with the object representation in the three-dimensional environment.
Owner:APPLE INC

Medical image segmentation method and system based on guiding information and multi-dimensional attention mechanism

The invention discloses a medical image segmentation method and system based on guidance information and a multi-dimensional attention mechanism. The method comprises the following steps: collecting an original dermatoscope image for preprocessing; constructing a segmentation model, wherein the segmentation model comprises a double-path image encoder, a guide information encoder and a mask decoder; the two-way image encoder is used for extracting local detail features and global context semantic information in the image; the guide information encoder is used for converting a coarse segmentation mask predicted by the last round of network into guide feature information; the mask decoder fuses the image feature information and the guide feature information, gradually restores and refines the coarse-grained feature map, and finally outputs an accurate lesion segmentation mask; constructing a loss function, and training the segmentation model by using the preprocessed data; and inputting a to-be-segmented original dermatoscope image into the trained segmentation model, and outputting a lesion region segmentation mask of the image. According to the method, the segmentation precision and the model generalization ability can be improved, and the multi-scale lesion processing ability is enhanced.
Owner:ZHEJIANG UNIV +1

Multi-modal large language model fine tuning method, system, equipment and medium

The invention relates to a multi-mode large language model fine tuning method, system and device and a medium, and belongs to the technical field of artificial intelligence and computer vision crossing. The fine tuning method comprises the steps that an original business scene image is acquired and preprocessed, and a preprocessed image is obtained; performing bounding box coordinate labeling and semantic label definition on the entity target in the preprocessed image through a labeling tool, and outputting a structured labeling file; based on the preprocessed image and the structured annotation file, constructing a training sample set comprising multiple rounds of image-text dialogues; loading the pre-trained multi-modal large language model, configuring low-rank matrix decomposition parameters, and generating a fine tuning instruction set; and inputting the training sample set into a pre-trained multi-modal large language model, carrying out joint training operation based on the fine tuning instruction set, and outputting the fine-tuned multi-modal large language model. According to the method, the identification accuracy, the interaction capability and the system availability of the visual question-answering system in an actual application scene are improved.
Owner:GOLDEN TIMES CULTURE COMM

Remote sensing image cultivated land segmentation method and system fusing context and boundary perception

The invention discloses a remote sensing image cultivated land segmentation method and system fusing context and boundary perception, and belongs to the technical field of remote sensing image processing and agricultural information. Constructing a cultivated land segmentation initial model composed of a backbone network, a feature enhancement module, a multi-scale feature fusion de-wharf module and a mask prediction module; training set data are input into the initial model, a composite loss function value is calculated, back propagation is executed, and a cultivated land segmentation model with boundary sensing ability is obtained through multi-round iterative optimization; and inputting the remote sensing image into the trained cultivated land segmentation model, and outputting a binary segmentation image representing the cultivated land position. Visual state space modeling and large receptive field convolution are combined, deep and shallow layer information is fused through feature injection, boundary perception supervision and composite loss are introduced, cultivated land boundary discrimination is improved, remote sensing image cultivated land high-precision extraction is achieved, and the method is suitable for agricultural interpretation and monitoring.
Owner:SANYA SCI & EDUCATION INNOVATION PARK WUHAN UNIV OF TECH

End-to-end automatic driving method based on dynamic multi-modal fusion in complex scene

The invention discloses an end-to-end automatic driving method based on dynamic multi-modal fusion in a complex scene, and belongs to the technical field of automatic driving. In order to solve the problems of sensor perception deficiency, cross-modal feature mismatching, unstable trajectory planning and the like easily occurring in night, low-illumination and complex dynamic environments in the existing end-to-end automatic driving method, texture details of a camera mode and geometric structure features of a laser radar mode are respectively enhanced through a double-flow feature refining mechanism; the characteristic difference between different modes is relieved; an information-driven dynamic fusion strategy is designed, the fusion weight is adaptively adjusted according to scene factors such as environment illumination and obstacle density, and the scene sensitivity and discrimination ability of the model are improved; asymmetric convolution and a low-rank-sparse decoupling technology are introduced, multi-order reconstruction of key channels is carried out on the multi-modal features, and the path change modeling capability is enhanced; and in combination with time sequence dependence of waypoints, outputting a future trajectory through an autoregression decoder to realize high-precision trajectory prediction and stable decision control.
Owner:ZHONGBEI UNIV

Geometry and topology collaborative guidance medical image segmentation method

The invention provides a medical image segmentation method based on geometry and topology cooperative guidance. The medical image segmentation method comprises the following steps of image preprocessing and data enhancement; a shared encoder; a dual-path cooperative decoder; carrying out multi-mode deformation iterative refining; and a multi-objective composite loss function and an optimization strategy. The method has the beneficial effects that the performance can be remarkably improved: through a unique geometry and topology collaborative refining mechanism, the segmentation precision and the boundary definition are far superior to those in the prior art, the topology correctness of an anatomical structure can be actively maintained and repaired, clinically unacceptable errors are remarkably reduced, and the reliability of a result is improved; in addition, operation can be simplified, stability and generalization are enhanced, and advanced application is promoted.
Owner:JIANGSU SHIYU INTELLIGENT MEDICAL TECH CO LTD +1

Crack segmentation method and system based on dynamic receptive field and multi-scale semantic aggregation

The invention discloses a crack segmentation method and system based on a dynamic receptive field and multi-scale semantic aggregation, and relates to the technical field of computer vision. The method comprises the following steps: inputting a crack image, and simultaneously capturing local details and global structural features of a crack through a dynamic snakelike Mama module: dynamically adjusting the shape of a convolution kernel to adapt to the geometric change of the crack, inputting the obtained features into a spatial pyramid pooling layer to extract multi-scale context information, and outputting feature representation fused with a long-range dependency relationship; based on the feature representation fused with the long-range dependency relationship, an interaction relationship between local details and global semantics is established through a multi-scale semantic aggregation module, background noise interference is suppressed through a parallel supervision attention mechanism, and a pixel-level crack segmentation result is generated through a lightweight segmentation head. According to the method, fine cracks can be segmented more accurately in a complex background environment, and meanwhile, the conditions of wrong segmentation and missing segmentation are effectively relieved.
Owner:SOUTHWEST JIAOTONG UNIV

Brain tumor multi-modal large model construction method and device, equipment and storage medium

The invention discloses a brain tumor multi-mode large model construction method, device and equipment and a storage medium, and is applied to the technical field of brain tumor imagines.The method comprises the steps that pixel-concept level alignment is conducted on a multi-mode MRI image and a pathological text; constructing a multi-modal feature fusion network for fusing image features and text features by adopting an attention mechanism of pathology perception and combining medical semantic information; training the multi-modal feature fusion network to generate an analysis report and a segmentation result; according to the technical scheme of multi-task cooperation, cross-modal pathological semantic accurate alignment, pathological knowledge graph injection and lightweight and continuous optimization parallelization, full-process coverage of brain tumor accurate segmentation, analysis report generation and prognosis prediction is achieved, the problems that a traditional model lacks pathological semantic support and is insufficient in clinical adaptability are solved, and the clinical adaptability of the traditional model is improved. And the deployment feasibility and the dynamic optimization capability are also considered.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Commodity display interaction visualization method and device

The invention relates to the field of commodity visualization, in particular to a commodity display interaction visualization method and device. The method comprises the following steps: collecting a multi-azimuth image of a commodity, carrying out three-dimensional texture modeling, and constructing a three-dimensional texture mapping model; performing material light rendering on the three-dimensional texture mapping model to generate a material rendering result; collecting an environment detection image of a commodity display environment, and performing environment illumination adaptation compensation on a material rendering result to obtain an illumination compensation rendering commodity; carrying out attribute information visual layout on the illumination compensation rendering commodity to obtain a commodity visual space; and carrying out interaction response animation analysis according to the commodity visualization space, carrying out multi-target parallel rendering, and executing commodity interaction visualization operation. The form and surface details of the commodity in the real world are accurately restored, the visual reality sense is improved, and the interactive experience feeling of browsing the commodity by a user is enhanced.
Owner:SHENZHEN XIAOYI SHUZHI TECH CO LTD

Neurological disease detection and analysis method and system

The invention discloses a nerve disease detection and analysis method and system, and the method comprises the steps: obtaining a bracelet collection signal, a sphygmomanometer collection signal, a movement behavior image and behavior test data, and extracting tremor intensity features, gait symmetry features and autonomic nerve rhythm features through multi-band decomposition of the bracelet collection signal; analyzing the motion behavior image and the standardized motion test to obtain a motion function score; carrying out heart rate variability analysis to identify a neural function abnormality mode; constructing a neural function state map and calculating a feature weight; predicting a disease progress trend in combination with historical monitoring data; and dynamically adjusting a prediction result through subsequent feedback correction information. Through a mode of combining short-time intensive monitoring and long-term intermittent acquisition, long-term trend prediction and dynamic correction based on initial data are realized, and the reliability and practicability of nerve disease risk assessment in a home scene are remarkably improved.
Owner:THE FIRST AFFILIATED HOSPITAL OF FUJIAN MEDICAL UNIV

Adaptive mask medical image segmentation method based on self-supervised mask and deep reinforcement learning

The invention discloses an adaptive mask medical image segmentation method based on a self-supervised mask and deep reinforcement learning, and the method comprises the steps: employing a classic encoder-decoder architecture for a self-supervised mask reconstruction network, fusing a Swin Transform encoder, and carrying out the feature fusion of local image blocks through a self-attention mechanism; according to the self-adaptive mask model, a PPO deep reinforcement learning algorithm is adopted, a strategy network and a value network are constructed, mask actions are dynamically regulated and controlled, reconstruction errors are gradually reduced, a mask strategy is continuously optimized in multiple times of strategy updating for self-adaptive optimization, and high-quality reconstruction of a medical image influenced by missing information is achieved; according to the method, high-quality feature representation can be obtained in an unlabeled data environment, and relatively high precision and accuracy are presented on a public data set.
Owner:YUNNAN UNIV

Style transfer using generative diffusion features

The present invention sets forth techniques for performing style transfer from multiple supplied style images to a supplied content image to generate novel images that include style elements from the multiple supplied style images and content elements from the supplied content image. The techniques include guiding one or more self-attention and cross-attention layers included in a machine learning model based on the multiple supplied style images, such that content elements and style elements included in the style images are not entangled when generating the novel images. The techniques also distill a small subset of representative attention map values from multiple style images, improving performance while reducing computational costs compared to processing all attention map values from the multiple style images.
Owner:DISNEY ENTERPRISES INC

Industrial image anomaly detection method based on deep learning

The invention discloses an industrial image anomaly detection method based on deep learning, and particularly relates to the technical field of industrial visual detection. The problems of high false alarm rate, fuzzy fine defect positioning, insufficient real-time response capability, difficulty in model increment updating and the like caused by data distribution drift in an industrial scene are solved. According to the method, robust features are extracted through a multi-scale feature fusion auto-encoder, and a dynamic memory bank is constructed to update a normal sample prototype online; a dual-path detection mechanism is adopted to cooperate with a pixel-level reconstruction error and attention weighted feature matching deviation; efficient edge reasoning is realized in combination with block parallel processing and model compiling optimization; and designing an elastic incremental learning framework to prevent disastrous forgetting. And finally, false alarms caused by environmental changes are reduced, accurate positioning of pixel-level defects is realized, millisecond-level detection requirements of high-resolution images are met, safe and efficient model online evolution is supported, and adaptability and reliability of an industrial quality inspection system are comprehensively improved.
Owner:SHANXI UNIV

Cover film defect intelligent detection method and system based on multi-feature fusion

The invention provides a multi-feature fusion-based cover film defect intelligent detection method and system, and the method comprises the steps: firstly obtaining a plurality of groups of image units of a to-be-detected cover film under different shooting parameters to form an image data set, carrying out the feature screening of the image data set, and obtaining a candidate feature set of a potential defect region; the method comprises the following steps: selecting a candidate feature set comprising regional gray features and morphological structure features, then performing association mapping on the candidate feature set, establishing an association relationship between the features to obtain an association feature spectrum, then calling a pre-constructed defect identification model to analyze the association feature spectrum, and generating an identification result of a defect prediction category identifier and a regional range parameter; and finally, generating a detection report containing defect position coordinates based on an identification result, and sending the detection report to a detection management system. Therefore, the accuracy and efficiency of cover film defect detection are improved.
Owner:SHENZHEN BANGZHENG PRECISION MACHINERY CO LTD

Water conservancy inspection robot inspection method and system based on multi-modal data fusion

The invention discloses a water conservancy inspection robot inspection method based on multi-modal data fusion, and relates to the field of intelligent inspection, and the method comprises the steps: obtaining the historical inspection information of a target dam; generating a water conservancy inspection strategy based on the historical inspection information and the dike data twinborn model; collecting multi-mode inspection information of the target dam; extracting multi-modal inspection features of the dam inspection area based on the multi-modal inspection information; and grade division of the multi-mode inspection features is completed, and dam abnormity grades are obtained. The accuracy of the inspection result can be effectively improved.
Owner:湖北亿立能科技股份有限公司

Short video intelligent editing method and system based on multi-modal analysis

The invention discloses a short video intelligent editing method and system based on multi-modal analysis, and relates to the technical field of video editing. The method is used for improving editing efficiency and visual experience and comprises the following steps: extracting lip motion features of a character, visual saliency features of a commodity and a voice emotion intensity value from a target short video stream to form multi-modal time sequence data; afterwards, the voice stream is recorded, a product keyword timestamp is extracted, the alignment degree is calculated through dynamic time warping in combination with a visual saliency peak value, and a preliminary editing point set is generated through weighted evaluation in combination with an emotional intensity value; constructing an editing decision optimization model based on deep reinforcement learning, taking the multi-modal features as state input, adjusting the retention probability of editing points through a joint reward function, and selecting an optimal transition mode; and the lip movement and voice synchronization error before and after the editing point and the emotional and visual continuity of the transition section are analyzed, the discontinuous region is smoothed, and the edited finished product is output, so that precise short video intelligent editing is realized.
Owner:ANHUI XINGBANG DIGITAL TECHNOLOGY GROUP CO LTD

Disease tracking management system and method based on lingual face diagnosis instrument

The invention discloses an illness state tracking management system and method based on a lingual face diagnosis instrument, and belongs to the technical field of traditional Chinese medicine tongue diagnosis and modern information technology fusion. The system comprises a multi-modal data acquisition module, a data processing and analysis module and an augmented reality visualization module; according to the method, illness state tracking is achieved through the steps of multi-modal data acquisition, data preprocessing and feature fusion, personalized digital twinborn model establishment, augmented reality visualization presentation and the like, multi-dimensional data such as tongue picture macroscopic features and tongue surface microorganism distribution can be integrated, dynamic association between the data is revealed, the health state and the intervention effect are visually displayed, and the method is suitable for being popularized and applied. The method is suitable for the fields of traditional Chinese medicine health management and chronic disease monitoring.
Owner:NANJING DAJING TCM INFORMATION TECH CO LTD

Three-dimensional attitude estimation method combining global modeling and local refinement

The invention discloses a three-dimensional attitude estimation method combining global modeling and local refinement, which comprises the following steps of: firstly, extracting a two-dimensional attitude sequence by using a human body video data set; secondly, inputting the two-dimensional attitude sequence into a structural modeling main branch, modeling a spatial topological relation and a time sequence dynamic state between joints, and outputting a global three-dimensional attitude sequence; and inputting the two-dimensional attitude sequence into a local refining branch, modeling dynamic change and detail information of a local area, and outputting a local three-dimensional attitude sequence. And finally, fusing the global three-dimensional attitude sequence and the local three-dimensional attitude sequence, generating a three-dimensional attitude sequence output, and completing three-dimensional attitude estimation. According to the method, the problem of insufficient cross-frame information transmission in a traditional method is relieved, and the accuracy and robustness of attitude estimation in a dynamic complex scene are remarkably improved.
Owner:HANGZHOU DIANZI UNIV

Automatic calibration method for 4D millimeter wave radar and visual camera in combination with space projection error analysis

The invention provides a space projection error analysis-combined 4D millimeter wave radar and visual camera automatic calibration method. The method comprises the steps of constructing a sensor internal reference model; designing a composite multifunctional calibration board and a three-dimensional adjustable holder, performing rigid geometrical relationship modeling, and automatically associating physics-features-coordinates; collecting multi-frame synchronization frame pair data, and performing time sequence fine grit alignment and interpolation compensation to obtain a space-time alignment frame pair sequence; extracting cross-modal spatial constraints of the features and automatically screening and matching point pairs to obtain a high-confidence point pair set; performing point-point constraint and point-point error, point-plane error and pixel reprojection error minimization to obtain a standardized error term; and constructing an overall objective function, and performing spatial constraint and time deviation estimation to obtain corrected external parameters and synchronization deviation. According to the invention, through a full-automatic standardized process, the probability of artificial participation and error occurrence is greatly reduced; and multi-modal and multi-scene high-precision calibration is supported, and the flexibility and robustness are greatly improved.
Owner:YANCHENG INST OF TECH

Medical image segmentation method based on adaptive anisotropic convolution

ActiveCN120726076AImage enhancementImage analysisData setRenal tumor
The invention provides a medical image segmentation method based on adaptive anisotropic convolution, and the method comprises the steps: obtaining a three-dimensional medical CT data set comprising images and labels of a plurality of abdominal organs and kidney tumors, and carrying out the preprocessing of the data set; dividing a data set into a training set and a test set for model training and evaluation; designing a three-dimensional medical image segmentation network model based on an adaptive anisotropic convolutional layer, and inputting the preprocessed training set into the three-dimensional medical image segmentation network model, the three-dimensional medical image segmentation network model is trained through parallel multi-modal convolution, adaptive attention weight generation, weighted feature dynamic fusion and multi-stage deep supervision, and model parameters are optimized; and applying the optimized three-dimensional medical image segmentation network model to a test set, generating a three-dimensional segmentation result with clear boundary and complete reserved details, and providing support for clinical diagnosis and treatment planning.
Owner:NANCHANG CAMPUS OF EAST CHINA UNIV OF TECH